A myocardial infarction localization method based on cross-modal knowledge fusion

By adopting a cross-modal knowledge fusion method in myocardial infarction localization, the spatial and timing features of the electrocardiogram vector graph are extracted using graph convolutional neural network and ResNet18 architecture, and feature alignment is performed, the shortcomings of single modal localization in the existing technology are solved, and the localization accuracy of myocardial infarction is significantly improved.

CN119477845BActive Publication Date: 2025-06-13GENERAL HOSPITAL OF SOUTHERN THEATRE COMMAND OF PLA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411558871.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2025-06-13
Estimated Expiration
2044-11-04

AI Technical Summary

Technical Problem

The prior art has bottlenecks in characterization learning from a single mode in the localization of myocardial infarction, and fails to best utilize the spatial and plane correlations of the electrocardiogram vector map, especially in the diagnosis of acute myocardial infarction in the posterior wall position, with insufficient diagnostic performance.

Method used

Using a myocardial infarction positioning method based on cross-modal knowledge fusion, the teacher model and student model are constructed, and the spatial features of 2D-VCG images and the timing features of 1D-VCG signals are extracted using graph convolutional neural network and ResNet18 architecture, and feature alignment and transformation are performed through cross-modal generators to achieve cross-modal knowledge fusion.

Benefits of technology

It significantly improves the localization accuracy of myocardial infarction, improves the semantic abundance and recognition accuracy of features, avoids potential negative migration caused by modal differences, and ensures the efficiency and accuracy of data integration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119477845B_ABST
    Figure CN119477845B_ABST
Patent Text Reader

Abstract

The present invention provides a myocardial infarction localization method based on cross-modal knowledge fusion, which includes collecting target 1D-VCG signals and constructing a 2D-VCG image dataset; extracting spatial features, temporal features, and pseudo-labels from 2D-VCG and 1D-VCG signals respectively through a teacher model and a student model; aligning and transforming the spatial features and temporal features through a cross-modal generator; and achieving cross-modal knowledge fusion and MI localization through the cross-modal generator. The present invention not only improves the accuracy of myocardial infarction localization through knowledge enhancement and student enhancement strategies, but also optimizes the overall framework of cross-modal learning through an adaptive weight method; the present invention is implemented in the form of knowledge distillation, successfully fuses the spatial features of 2D-VCG into the 1D-VCG temporal analysis, and significantly improves the localization accuracy of myocardial infarction by combining the collaborative work of the teacher model, the student model, and the cross-modal generator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of myocardial infarction localization, and in particular to a myocardial infarction localization method based on cross-modal knowledge fusion. Background Art

[0002] Myocardial infarction (MI) is a serious heart disease that can cause rapid and irreversible damage to the myocardium. Timely diagnosis and treatment are crucial for preventing the deterioration of the condition.

[0003] As non-invasive detection means, 12-lead electrocardiogram (ECG) and 3-lead vectorcardiogram (VCG) are often used clinically to localize MI. Given the high readability and wide application of the 12-lead electrocardiogram, some studies analyze the 12-lead electrocardiogram based on artificial intelligence methods to localize MI. Although positive results have been achieved using electrocardiograms and the possibility of computer intelligent-assisted localization of MI has been confirmed, its advantages in localizing certain types of MI still need to be improved. For example, in the diagnosis of acute myocardial infarction in the posterior wall position, better diagnostic performance can be shown using VCG. Part of the reason is that the electrocardiogram is a secondary projection of the VCG in the frontal plane and the horizontal plane and cannot fully capture the spatial information of cardiac electrical activity.

[0004] In response, the artificial intelligence algorithm analysis of the vectorcardiogram obtained by monitoring cardiac activity through the Frank orthogonal lead system has become a major research trend in the field of computer intelligent-assisted MI localization system development in recent years. Given that the vectorcardiogram is a three-dimensional electrical activity representation composed of the signals of three orthogonal leads, depicting the direction and intensity of cardiac electrical activity at each time point. It is a vector representation of cardiac electrical activity in space. In the field of artificial intelligence-based cardiac electrical activity analysis, the developed models are usually for one-dimensional time series or two-dimensional images. In response, existing research on VCG usually establishes a one-dimensional machine learning model to obtain time series features or spectral features of the three orthogonal leads, or establishes a two-dimensional deep learning architecture to obtain spatial features of the projections of the VCG on the three orthogonal planes of the frontal plane, the horizontal plane, and the sagittal plane, so as to localize myocardial infarction.

[0005] Although the above research has made remarkable progress in recent years, the paradigm of myocardial infarction localization and characterization learning from a single modality has gradually encountered bottlenecks. Among them, extracting non-linear features from one-dimensional vectorcardiograms may be complex and cumbersome, and the spatial and inter-plane correlations of vectorcardiograms are not optimally utilized. Secondly, for 2D plane projections, as the signal duration of the vectorcardiogram increases, its internal spatial structure becomes more complex, and the 2D projection may miss the subtle details inherent in the 3D vectorcardiogram, posing major challenges for myocardial infarction localization that prioritizes vectorcardiogram segments rather than obvious vectorcardiogram beats and the lightweight deployment of the model at the edge. Summary of the Invention

[0006] Aiming at the deficiencies of the prior art, the present invention provides a myocardial infarction localization method based on cross-modal knowledge fusion. The present invention makes full use of the richer spatial features contained in the projection images of VCG on three orthogonal planes.

[0007] The technical solution of the present invention is as follows: A myocardial infarction localization method based on cross-modal knowledge fusion, including the following steps:

[0008] S1), Collect the target 1D-VCG signal and construct a 2D-VCG image dataset, and then divide the 2D-VCG image dataset into a training set and a test set;

[0009] S2), Construct a teacher model and train it. Extract the spatial features and pseudo-label Logits of the 2D-VCG image through the teacher model;

[0010] S3), Construct a student model and train it. Obtain the temporal features and pseudo-label Logits from the 1D-VCG signal through the student model;

[0011] S4), Construct a cross-modal generator, and align and transform the spatial features extracted by the teacher model and the temporal features extracted by the student model through the cross-modal generator;

[0012] S5), Achieve cross-modal knowledge fusion and MI localization through the cross-modal generator.

[0013] Preferably, in step S2), the teacher model is iteratively optimized based on the graph convolutional neural network MVCNN. The teacher model obtains the features of each view from multiple parallel ResNet18 architectures; and intercepts and weighted aggregates the core spatial features of each view through a multi-view collaborative attention fusion module.

[0014] Preferably, in step S2), the multi-view collaborative attention fusion module includes an attention mechanism in two stages, and enhances the semantic information of the feature maps from different perspectives through the attention mechanism in two stages.

[0015] Preferably, in step S2), the relevant features from different perspectives are aggregated through an attention mechanism in one stage. Specifically:

[0016] S211), Compress the channel features of the feature tensors (f i 1 , f i 2 , f i 3 ) from three perspectives through a convolutional layer, and combine them into a tensor Fi Ang ; The calculation expression is:

[0017] f i Ang = ReLU(Conv(Cat(f i 1 , f i 2 , f i 3 )))

[0018] In the formula, (f i 1 , f i 2 , f i 3 ) are tensors from three different perspectives. Cat represents concatenating the tensor views (f i 1 , f i 2 , f i 3 ) along the channels; Conv represents the convolution operation; ReLU is the activation function; the subscript i represents the index of the currently processed sample, used to distinguish different samples;

[0019] S212), calculating the attention factor through the sigmoid function; that is:

[0020] u A = σ(f i Ang × 1 U A1 × 2 U A2 × 3 U A3 )

[0021] u B = σ(f i Ang × 1 U B1 × 2 U B2 × 3 U B3 )

[0022] u C = σ(f i Ang × 1 U C1 × 2 U C2 × 3 U C3 )

[0023] wherein, u A , u B , u C is the attention factor; representing the feature weights from different perspectives, calculated by using the sigmoid function, and its role is to adjust the importance of each feature from different perspectives; σ represents the sigmod function; {U Ai , U Bi , U Ci} represents the parameter matrix of the attention factor calculation module, indicating the weight allocation for different feature channels. Through these matrices, the model can capture the importance of the feature tensor on different channels; the parameter matrices are respectively used for three different perspective feature channels; × represents modulo multiplication; A, B, and C represent the labels of the three different perspective feature channels, and A, B, and C are used to distinguish features from different perspectives, enabling the attention factor to perform independent weighted sum and fusion for each perspective;

[0024] S213), by multiplying the attention factor with the corresponding feature tensor element by element, an enhanced feature tensor F i Ang is obtained; that is:

[0025] F i Ang = f i 1 ⊙ u A + f i 2 ⊙ u B + f i 3 ⊙ u c}

[0026] wherein, F i Ang represents the feature tensor of the enhanced first-stage attention; ⊙ represents element-by-element multiplication;

[0027] Then, the semantic information of the feature map is enhanced from different angles through two parallel internal attention feature fusion modules.

[0028] Preferably, in step S2), the attention mechanism in the second stage further interacts with the feature tensor from the first stage, and multiple multi-view feature tensors complement each other to form the final output.

[0029] Preferably, in step S2), the attention mechanism in the second stage processes the feature tensor in the first stage to obtain the final output, which specifically includes the following steps:

[0030] S221), through max-pooling and average-pooling operations, F i A1and F i A2 , that is:

[0031] F i A1 = Conv(ReLU(Conv(AvgPool)(F i Ang ))))

[0032] f i A2 = Conv(ReLU(Conv(MaxPool)(F i Ang ))))

[0033] In the formula, F i Ang represents the feature tensor of the attention mechanism in the first stage of enhancement, MaxPool and AvgPool respectively represent the max pooling and average pooling operations; Conv represents the convolution operation; ReLU is the activation function; F i A1 and F i A2 respectively represent the feature tensors after the max pooling MaxPool and average pooling AvgPool operations, aiming to further extract the saliency information of the features; MaxPool retains the main information of the feature map by taking the maximum value in the local window to highlight the significant features; AvgPool smooths the feature map by calculating the average value of the local window to reduce the detailed noise of the features;

[0034] S222), combined with the self-attention mechanism, and generates the final feature tensor F i ′Ang , that is:

[0035] F i Ang = F u Ang ⊙σ(F i A1 + F i A2 )

[0036] F i ′A1 = avg(F i Ang )

[0037] f i ′ A2 = max(f i Ang )

[0038] f i′Ang = F i Ang ⊙σ(Conv(Cat(F i ′A1 , F i ′A2 )))

[0039] In the formula, max represents taking the average maximum value of elements along the channel dimension, Cat represents concatenating two feature tensors along the channel dimension to integrate the feature information of different pooling operations; avg represents taking the average value of elements along the channel dimension; F i ′A1 , F i ′A2 respectively represent the results after performing average pooling and max pooling operations on the feature tensor F i Ang ; where F i A1 = avg(F i Ang ), representing the result of performing average pooling on the feature tensor F i Ang for smoothing features; F i A2 = avg(F i Ang ) represents the result of performing max pooling on the feature tensor F i Ang for extracting significant features; ⊙ represents element-wise multiplication; σ represents the sigmod function; Conv represents the convolution operation.

[0040] Preferably, in step S3), the student model uses the ResNet18 architecture as a feature extractor and a myocardial infarction localization classifier to extract the time-frequency features of the 1D-VCG time series signal. Through the max pooling layer and the fully connected layer of the student model, the time series features and the output pseudo-labels are obtained respectively for subsequent knowledge enhancement optimization and student enhancement optimization in cross-modal knowledge distillation.

[0041] Preferably, in step S3), the student model can effectively extract and learn the time features related to myocardial infarction by combining cross-modal distillation technology, and uses a deep convolutional network structure to capture the subtle time series changes of 1D-VCG data.

[0042] Preferably, in step S4), the cross-modal generator adopts a linear transformation strategy to ensure the precise mapping and alignment of features from different modalities in the Euclidean space.

[0043] Preferably, in step S4), the cross-modal generator converts the spatial feature f of 2D-VCG through a parameterized transformation matrix W 2D into a form aligned with the temporal feature f of 1D-VCG 1D :

[0044] f align = W·f 2D + b

[0045] wherein, W is the transformation matrix; b is the bias term; f align is the transformed feature.

[0046] Preferably, in step S4), after feature alignment, the cross-modal generator uses a variational autoencoder (VAE) to further optimize the features. The variational autoencoder (VAE) reconstructs a new feature representation consistent with the temporal feature space, while ensuring the information abundance of the features and the adaptability to unseen data. The following objective function is optimized by minimizing the reconstruction error and KL divergence:

[0047]

[0048] wherein, represents the knowledge enhancement loss of the variational autoencoder (VAE). This loss is used to ensure that the feature f align retains necessary information during the alignment process and adapts to unseen data; Encoder and Decoder respectively represent the variational autoencoder (VAE) and the decoder. The encoder maps the aligned feature f align to the latent space, and the decoder maps it back to the original feature space to achieve feature reconstruction; KL represents the Kullback-Leibler divergence, which is used to measure the difference between the learned distribution and the standard normal distribution. This term encourages the distribution in the latent space to be close to the normal distribution so as to achieve regularization; β is a hyperparameter for adjusting the importance of the KL divergence, which controls the balance between the reconstruction error and the distribution matching. A larger β value will pay more attention to the distribution matching, while a smaller value will pay more attention to the reconstruction accuracy; f 1D represents the temporal feature of 1D-VCG, which is used to calculate the reconstruction error; represents the standard normal distribution, which is used as the target of the latent space distribution to help optimize the latent representation of the VAE.

[0049] Preferably, in step S5), the cross-modal knowledge fusion includes knowledge enhancement optimization and student enhancement optimization.

[0050] Preferably, in step S5), the knowledge enhancement optimization refers to transforming the spatial features of the cross-modal images obtained in 2D-VCG into the Euclidean space, and through a point generator, independently transforming the features obtained by the teacher model and the student model into spatio-temporal domain features, and performing loss supervision through a variational autoencoder to obtain the knowledge enhancement loss

[0051] Preferably, in step S5), the student enhancement optimization includes the following steps:

[0052] S51), input the 1D-VCG into the encoder of the student model to generate a temporal feature representation f ts , and further optimize it through a feature mapping layer U:

[0053] f optimized = Uf ts

[0054] In the formula, f optimized is the feature optimized by the feature mapping layer; U is the feature mapping layer; f ts generates the temporal feature representation;

[0055] S52), use the optimized feature f optimized to train the classifier to generate a class prediction P pred ; and calculate the cross-entropy loss through the class prediction P pred ;

[0056]

[0057] In the formula, CrossEntropy represents the cross-entropy loss function, which is used to calculate the loss between the prediction result and the true label in the classification task, y true represents the true label, which is used to compare with the model prediction result; P pred represents the class prediction result of the model, which is used to calculate the loss with the true label;

[0058] S53), use the distillation loss to optimize:

[0059]

[0060] In the formula, P teacher is the output of the teacher model; λ is a trade-off factor, which is used to adjust the weight between the cross-entropy loss and the MSE loss; MSE represents the Mean Squared Error Loss, which is used to calculate the difference between the prediction result of the student model and the output of the teacher model to help the student model better imitate the teacher model;

[0061] S54), the final loss of the cross-modal generator is as follows:

[0062]

[0063] In the formula, α, β, and γ represent the weight coefficients of different loss terms, which are used to adjust the influence degree of each loss term in the final loss function; α corresponds to the weight of the knowledge enhancement loss ; β corresponds to the weight of the cross-entropy loss ; γ corresponds to the weight of the distillation loss ; is the knowledge enhancement loss; is the cross-entropy loss; is the distillation loss.

[0064] The beneficial effects of the present invention are as follows:

[0065] 1. The present invention enhances the orthogonal lead signal analysis of VCG by using the visual images of 2D-VCG, so as to effectively obtain auxiliary knowledge from the view images;

[0066] 2. The present invention develops a teacher-student framework and formulates cross-modal learning as a knowledge distillation problem; the distribution difference between different modalities is eliminated through a new classifier enhancement criterion, effectively avoiding potential negative transfer;

[0067] 3. Through the knowledge enhancement and student enhancement strategies, the present invention not only improves the accuracy of myocardial infarction localization, but also optimizes the overall framework of cross-modal learning by the adaptive weight method;

[0068] 4. The present invention is implemented in the form of knowledge distillation, successfully fuses the spatial features of 2D-VCG into the 1D-VCG time series analysis, combines the collaborative work of the teacher model, the student model and the cross-modal generator, and significantly improves the localization accuracy of myocardial infarction;

[0069] 5. The present invention adopts a multi-view convolutional neural network as the teacher model, and weighted merges the image features of different perspectives through a collaborative attention fusion module. This strategy enhances the model's ability to capture the spatial heterogeneity of cardiac activities, and improves the semantic abundance and recognition accuracy of the features;

[0070] 6. The cross-modal generator of the present invention realizes the precise alignment of different modality data through a simple and effective linear or convolutional layer configuration, combined with variational autoencoder supervision; not only optimizes the consistency at the feature level, but also avoids potential negative transfer caused by modality differences, ensuring the efficiency and accuracy of data integration. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 is a schematic framework diagram of the method of the present invention;

[0072] Figure 2 This is the framework diagram of the teacher model of the present invention;

[0073] Figure 3 This is the framework diagram of the student model of the present invention;

[0074] Figure 4 This is the framework diagram of the cross-modal generator of the present invention. Detailed implementation manners

[0075] The following further describes the detailed implementation manners of the present invention with reference to the accompanying drawings:

[0076] As Figure 1 shown, this embodiment provides a myocardial infarction localization method based on cross-modal knowledge fusion, including the following steps:

[0077] S1), construct a 2D-VCG image dataset, and divide the 2D-VCG image dataset into a training set and a test set;

[0078] In this embodiment, an electrocardiogram vectorgram instrument or a high-density body surface potential acquisition device is used to collect the target 1D-VCG signal, and 2D-VCG images are obtained through downsampling, filtering, and projection mapping.

[0079] In this embodiment, an electrocardiogram vectorgram dataset of myocardial infarction patients and normal people provided by Physikalisch-Technische Bundesanstalt (PTB) is used; the original sampling frequency of the data is 1000 Hz.

[0080] Among them, the downsampling is to downsample the sampling frequency to 512 Hz to reduce the computational complexity; and the filtering uses a second-order Butterworth high-pass filter with a cut-off frequency set to 0.5 Hz to eliminate baseline drift;

[0081] Timing signal acquisition: Use the Pan-Tompkins algorithm to detect the R wave, and take 0.25 seconds to the left and 0.75 seconds to the right centered on the R peak to obtain 0.75 seconds of non-overlapping 1D-VCG timing signals;

[0082] Image signal acquisition: 2D-VCG images are obtained by projecting the electrocardiogram vectorgram on three orthogonal planes, namely the frontal plane, the horizontal plane, and the sagittal plane. Each image is non-overlapping and consistent with the timing signal, and consists of a single projection map composed of 0.75 seconds.

[0083] S2), construct a teacher model and train it, and extract the spatial features of the 2D-VCG image through the teacher model;

[0084] In this embodiment, as Figure 2As shown, the teacher model uses the graph convolutional neural network MVCNN. The teacher model obtains the feature of each view through multiple parallel ResNet18 architectures. And the teacher model intercepts and aggregates the core spatial features of each view through the multi-view collaborative attention fusion module.

[0085] The multi-view collaborative attention fusion module includes an attention mechanism in two stages. Through the attention mechanism in two stages, the semantic information of the feature maps from different perspectives is enhanced for obtaining spatial features and pseudo-labels.

[0086] Among them, the relevant features from different perspectives are aggregated through the attention mechanism in the first stage. Specifically:

[0087] S211), compress the channel features of the feature tensors (f i 1 , f i 2 , f i 3 ) from three perspectives through a convolutional layer, and combine them into a tensor F i Ang ; The calculation expression is:

[0088] f i Ang = ReLU(Conv(cat(f i 1 , f i 2 , f i 3 )))

[0089] In the formula, (f i 1 , f i 2 , f i 3 ) are the feature tensors from three different perspectives. Cat means concatenating the feature tensors (f i 1 , f i 2 , f i 3 ) from three different perspectives along the channels; Conv represents the convolutional operation; ReLU is the activation function; the subscript i represents the index of the currently processed sample, used to distinguish different samples;

[0090] S212), calculate the attention factor through the sigmoid function; that is:

[0091] u A = σ(f i Ang ×1 U A1 × 2 U A2 × 3 U A3 )

[0092] u B =σ(f i Ang × 1 U B1 × 2 U B2 × 3 U B3 )

[0093] u C =σ(f i Ang × 1 U C1 × 2 U C2 × 3 U C3 )

[0094] Wherein, u A , u B , u C are attention factors, representing the feature weights from different perspectives, calculated by using the sigmoid function, and the role is to adjust the importance of each feature from different perspectives; σ represents the sigmod function; {U Ai , U Bi , U Ci} is the parameter matrix of the attention factor calculation module, representing the weight distribution for different feature channels. Through these matrices, the model can capture the importance of the feature tensor on different channels; the parameter matrices are respectively used for three different perspective feature channels; × represents modulo multiplication; A, B, and C respectively represent the markers of three different perspective feature channels, and A, B, and C are used to distinguish the features from different perspectives, so that the attention factor can perform independent weighted sum and fusion for each perspective;

[0095] S213), by multiplying the attention factor with the corresponding feature tensor element by element, the enhanced feature tensor F i Ang is obtained; that is:

[0096] F i Ang =f i 1 ⊙u A +f i 2 ⊙u B +f i 3 ⊙uc}

[0097] Wherein, F i Ang represents the feature tensor of the attention mechanism in the first stage of enhancement; ⊙ represents element-wise multiplication;

[0098] Then, the semantic information of the feature map is enhanced from different angles through two parallel internal attention feature fusion modules.

[0099] In this embodiment, the attention mechanism in the second stage forms the final output by further interacting with the feature tensors from the attention mechanism in the first stage, and multiple multi-view feature tensors complement each other, specifically including the following steps:

[0100] S221), obtain F i A1 and F i A2 , that is:

[0101] F i A1 = Conv(ReLU(Conv(AvgPool)(F i Ang ))))

[0102] F i a2 = Conv(ReLU(Conv(MaxPool)(F o Ang ))))

[0103] Wherein, F o Ang represents the feature tensor of the attention mechanism in the first stage of enhancement, MaxPool and AvgPool respectively represent the max pooling and average pooling operations; Conv represents the convolution operation; ReLU is the activation function; F o A1 and F o A2 respectively represent the feature tensors after the max pooling MaxPool and average pooling AvgPool operations, aiming to further extract the saliency information of the features; MaxPool retains the main information of the feature map by taking the maximum value in the local window to highlight the significant features; AvgPool smooths the feature map by calculating the average value of the local window to reduce the detailed noise of the features;

[0104] S222), combine the self-attention mechanism and generate the final feature tensor F i ′Ang , that is:

[0105] F i Ang = F i Ang ⊙σ(F i A1 + F i A2 )

[0106] F i ′A1 = avg(F i Ang )

[0107] f i ′A2 = max(f i Ang )

[0108] f i ′Ang = F i Ang ⊙σ(Conv(Cat(F i ′A1 , F i ′A2 )))

[0109] In the formula, max represents taking the average maximum value of elements along the channel dimension, Cat represents concatenating two feature tensors along the channel dimension to integrate the feature information of different pooling operations; avg represents taking the average value of elements along the channel dimension; F i ′A1 and F i ′A2 respectively represent the results after performing average pooling and max pooling operations on the feature tensor F i Ang ; where F i A1 = avg(F i Ang ) represents the result of performing average pooling on the feature tensor F i Ang for smoothing features; F i A2 = avg(F i Ang ) represents the result of performing max pooling on the feature tensor F i Ang for extracting significant features; ⊙ represents element-wise multiplication to combine the feature tensor with the attention factor to adjust the feature weights; σ represents the sigmod function for calculating the attention factor to help adjust the importance of different features; Conv represents the convolution operation for further extracting and enhancing feature information;

[0110] In addition, the dimension of the feature tensor of the teacher model in this embodiment is: 3 (number of channels) × 256 (height) × 256 (width); and the number of filters in the convolutional layer is 32, and the convolutional kernel size is set to 3×3 to ensure rich feature information extraction. The activation function uses ReLU, and the pooling size of both the max pooling and average pooling is 2×2 to effectively reduce the size of the feature map. The initial learning rate is set to 0.001, the number of training epochs is 50, and the loss function combines cross-entropy loss (L CE ) and contrastive loss (L CL ) to optimize the model performance.

[0111] S3), construct a student model to train it, and obtain temporal features from the 1D-VCG signal through the student model;

[0112] In this embodiment, as Figure 3 shown, the student model uses the ResNet18 architecture as a feature extractor and a myocardial infarction location classifier to extract the time-frequency features of the 1D-VCG time series signal. Through the max pooling layer and fully connected layer of the student model, temporal features and the output pseudo-labels are obtained respectively for knowledge enhancement and student enhancement in subsequent cross-modal knowledge distillation.

[0113] The student model can effectively extract and learn the time features related to myocardial infarction by combining cross-modal distillation technology. The student model uses a deep convolutional network structure to capture the subtle temporal changes in 1D-VCG data.

[0114] The student model includes multiple convolutional layers and fully connected layers. The convolutional layer is configured with 16 filters, with a size of 1×5 and a stride of 1; without padding to ensure precise processing of the time dimension; the activation function uses ReLU to provide non-linear processing capabilities. Batch normalization and max pooling are applied after each convolution, and the pooling size is 2×1 to enhance the generalization ability of the model.

[0115] And the student model uses a loss function specifically optimized for time series data, combined with the soft label output of the teacher model, to form a distillation loss

[0116]

[0117] where α and λ are hyperparameters, and their optimal values are determined by adaptive tuning through Bayesian optimization; CrossENtropy represents the cross-entropy loss function, which is used to calculate the loss between the prediction result and the true label in the classification task. y true represents the true label, which is used to compare with the model prediction result; P pred represents the class prediction result of the model, which is used to calculate the loss with the true label;

[0118] During the training process, the Adam optimizer is adopted, and the initial learning rate is set to 0.0001 to ensure a stable learning progress.

[0119] S4), Construct a cross-modal generator, and align and transform the spatial features extracted by the teacher model and the temporal features extracted by the student model through the cross-modal generator.

[0120] In this embodiment, as Figure 4 shown, the cross-modal generator converts the spatial feature f of 2D-VCG into a form aligned with the temporal feature f of 1D-VCG through a parameterized linear layer W 2D as follows: 1D f

[0121] f align = W·f 2D + b

[0122] where W is the transformation matrix; b is the bias term; f align is the transformed feature.

[0123] After the feature alignment, the cross-modal generator uses a variational autoencoder (VAE) to further optimize the features. The variational autoencoder (VAE) reconstructs a new feature representation consistent with the temporal feature space, while ensuring the information abundance of the features and the adaptability to unseen data. The knowledge enhancement loss is optimized by minimizing the reconstruction error and the KL divergence

[0124]

[0125] where Encoder and Decoder represent the variational autoencoder (VAE) and the decoder respectively, KL represents the Kullback-Leibler divergence, which is used to measure the difference between the learned distribution and the standard normal distribution, and β is a hyperparameter that adjusts the importance of the KL divergence; where represents the knowledge enhancement loss of the variational autoencoder (VAE). This loss is used to ensure that the feature f align retains the necessary information during the alignment process and adapts to unseen data; Encoder and Decoder represent the variational autoencoder (VAE) and the decoder respectively. The encoder maps the aligned feature f align to the latent space, and the decoder maps it back to the original feature space to achieve feature reconstruction; KL represents the Kullback-Leibler divergence, which is used to measure the difference between the learned distribution and the standard normal distribution. This term encourages the distribution in the latent space to be close to the normal distribution Regularization is thus achieved; β is a hyperparameter that adjusts the importance of the KL divergence, controlling the balance between the reconstruction error and the distribution matching. A larger β value places more emphasis on distribution matching, while a smaller value places more emphasis on reconstruction accuracy; f 1D represents the temporal features of 1D-VCG and is used to calculate the reconstruction error; represents the standard normal distribution, serving as the target for the latent space distribution to help optimize the latent representation of the VAE;

[0126] S5), Cross-modal knowledge fusion and MI localization;

[0127] The spatial features extracted by the teacher model from the 2D-VCG and the temporal features extracted by the student model from the 1D-VCG are fused through a cross-modal generator to achieve high-precision localization of myocardial infarction.

[0128] Among them, the cross-modal knowledge fusion includes knowledge enhancement optimization and student enhancement optimization. The knowledge enhancement optimization refers to transforming the spatial features of the cross-modal images obtained from the 2D-VCG into the Euclidean space, and through a point generator, independently transforming the features obtained by the teacher model and the student model into spatio-temporal domain features, and performing loss supervision through a variational autoencoder to obtain the knowledge enhancement loss Specifically:

[0129] The 2D-VCG image features extracted by the teacher model are first adjusted by a transformation matrix W to adapt to the temporal feature space of the 1D-VCG of the student model:

[0130] f align = W·f 2D + b

[0131] where W is the transformation matrix and b is the bias term. These parameters are obtained through training and learning, aiming to optimize the spatial alignment of the features;

[0132] The transformed feature f align is further optimized through a specifically designed variational autoencoder VAE; the loss function of the VAE includes the reconstruction loss and the KL divergence, and the formula is:

[0133]

[0134] The student enhancement optimization includes the following steps:

[0135] S51), Input the 1D-VCG into the encoder of the student model to generate the temporal feature representation f ts , and further optimize it through a feature mapping layer U:

[0136] f optimized = Uf ts

[0137] In the formula, f optimized is the optimized feature of the feature mapping layer; U is the feature mapping layer; f ts generates a temporal feature representation;

[0138] S52), using the optimized feature f optimized to train a classifier to generate a class prediction P pred ; and calculating the cross-entropy loss through the class prediction P pred ;

[0139]

[0140] In the formula, CrossEntropy represents the cross-entropy loss function, which is used to calculate the loss between the prediction result and the true label in the classification task. y true represents the true label and is used to compare with the model prediction result; P pred represents the class prediction result of the model and is used to calculate the loss with the true label;

[0141] S53), using the distillation loss for optimization:

[0142]

[0143] In the formula, P teacher is the output of the teacher model; λ is a trade-off factor used to adjust the weight between the cross-entropy loss and the MSE loss; MSE represents the Mean Squared Error Loss, which is used to calculate the difference between the prediction result of the student model and the output of the teacher model to help the student model better imitate the teacher model;

[0144] S54), the final loss of the cross-modal generator is:

[0145]

[0146] Among them, each weighting coefficient builds a probability model of the objective function through Bayesian optimization, uses the existing data points to optimize the weight allocation to adapt to different teacher-student architectures, and finally completes the localization of myocardial infarction through a classifier.

[0147] In addition, as shown in Table 1, this embodiment uses accuracy, specificity, sensitivity, and F1 score for performance evaluation. The method of this embodiment achieves an identification accuracy of 98.56% in cross-subject myocardial infarction recognition; the inter-patient test scheme for the 6-class MI localization task shows excellent performance, with an accuracy of 68.68%, significantly superior to the existing myocardial infarction localization technology.

[0148] Table 1 Comparison of the positioning results between the method of this embodiment and the existing methods

[0149]

[0150]

[0151] The above embodiments and descriptions in the specification only illustrate the principles and the best embodiments of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the claimed present invention.

Claims

1. A method for locating myocardial infarction based on cross-modal knowledge fusion, characterized in that: The steps include: S1), collecting target 1D-VCG signals and constructing a 2D-VCG image dataset, dividing the 2D-VCG image dataset into a training set and a test set; S2), constructing a teacher model and training it, extracting spatial features and pseudo-label Logits of 2D-VCG images through the teacher model; S3), constructing a student model to train it, and obtaining timing features and pseudo-label Logits from the 1D-VCG signal through the student model; S4), constructing a cross-modal generator, aligning and transforming the spatial features extracted by the teacher model and the temporal features extracted by the student model through the cross-modal generator; S5), realizing cross-modal knowledge fusion and MI positioning through a cross-modal generator; The teacher model is iteratively optimized based on the graph convolutional neural network MVCNN, and the teacher model obtains the features of each view by multiple parallel ResNet18 architectures; The core spatial features of each view are extracted and weighted aggregated through the multi-view collaborative attention fusion module; The multi-view collaborative attention fusion module includes a two-stage attention mechanism, which enhances the semantic information of feature maps from different perspectives through the two-stage attention mechanism; The attention mechanism in the first stage aggregates relevant features from different perspectives, specifically: S211), the feature tensors from three perspectives The channel features are compressed through the convolution layer and combined into a tensor The calculation expression is: In the formula, For three tensors of different perspectives, Cat represents the tensor views of three different perspectives Concatenate along the channel; Conv represents the convolution operation; ReLU is the activation function; the subscript i represents the index of the sample currently being processed, which is used to distinguish different samples; S212), calculate the attention factor through the sigmoid function; that is: In the formula, u A ,u B ,u C is the attention factor, which indicates the feature weights under different perspectives; σ represents the sigmoid function; {U Ai ,U Bi ,U Ci } is the parameter matrix of the attention factor calculation module, which indicates the weight allocation to different feature channels; × indicates modular multiplication; A, B, and C indicate the labels of three different view feature channels respectively; S213), by multiplying the attention factor by the corresponding feature tensor element by element, an enhanced feature tensor is obtained Right now: In the formula, Represents the feature tensor of the first stage of enhanced attention; ⊙ represents element-by-element multiplication; Then two parallel internal attention feature fusion modules are used to enhance the semantic information of the feature map from different perspectives.

2. The method for locating myocardial infarction based on cross-modal knowledge fusion according to claim 1, characterized in that: In step S2), the feature tensors from the first stage are further interacted through the attention mechanism of the second stage, and multiple multi-view feature tensors complement each other to form the final output; specifically, the following steps are included: S221), through the maximum pooling and average pooling operations, obtain and Right now: In the formula, Represents the feature tensor of the first stage of the enhanced attention mechanism, MaxPool and AvgPool represent the maximum pooling and average pooling operations respectively; Conv represents the convolution operation; ReLU is the activation function; and They represent the feature tensors after the maximum pooling MaxPool and average pooling AvgPool operations, respectively, to further extract the saliency information of the features; S222), combined with the self-attention mechanism, and the final feature tensor F is generated through convolution operation i ′Ang ,Right now: In the formula, max means taking the average maximum value of elements along the channel dimension; Cat means concatenating two feature tensors along the channel dimension to integrate the feature information of different pooling operations; avg means taking the average value of elements along the channel dimension; Respectively represent the feature tensors The result after average pooling and maximum pooling operations; ⊙ represents element-by-element multiplication; σ represents the sigmoid function; Conv represents the convolution operation.

3. The method for locating myocardial infarction based on cross-modal knowledge fusion according to claim 1, characterized in that: In step S3), the student model adopts the ResNet18 architecture as a feature extractor and myocardial infarction localization classifier, and extracts the time-frequency features of the 1D-VCG timing signal by combining the cross-modal distillation technology. The maximum pooling layer and the fully connected layer of the student model are used to obtain the timing features and the output pseudo-labels, respectively, for subsequent knowledge enhancement optimization and student enhancement optimization in cross-modal knowledge distillation.

4. The method for locating myocardial infarction based on cross-modal knowledge fusion according to claim 1, characterized in that: In step S5), the cross-modal knowledge fusion includes knowledge enhancement optimization and student enhancement optimization.

5. The method for locating myocardial infarction based on cross-modal knowledge fusion according to claim 4, characterized in that: In step S5), the knowledge enhancement optimization refers to converting the spatial features of the cross-modal image obtained in the 2D-VCG into the Euclidean space, and independently transforming the features obtained by the teacher model and the student model into spatiotemporal features through the point generator, and performing loss supervision through the variational autoencoder to obtain the knowledge enhancement loss 6. The method for locating myocardial infarction based on cross-modal knowledge fusion according to claim 5, characterized in that: First, the 2D-VCG image features extracted by the teacher model are adjusted to fit the 1D-VCG temporal feature space of the student model; The cross-modal generator adopts a linear transformation strategy to ensure that the features from different modalities are accurately mapped and aligned in the Euclidean space; the cross-modal generator transforms the spatial features f of the 2D-VCG into 2D Converted to the timing feature f of 1D-VCG 1D Alignment form: f align =W·f 2D +b Where W is the conversion matrix; b is the bias term; f align is the transformed feature; The cross-modal generator uses a variational autoencoder (VAE) to further optimize features after feature alignment. The variational autoencoder (VAE) reconstructs a new feature representation consistent with the temporal feature space while ensuring the information richness of the features and adaptability to unseen data. The following objective function is optimized by minimizing the reconstruction error and KL divergence: in, represents the knowledge enhancement loss of the variational autoencoder (VAE); Encoder and Decoder represent the variational autoencoder VAE and decoder respectively, KL represents the Kullback-Leibler divergence, and β is a hyperparameter for adjusting the importance of KL divergence; f 1D Represents the timing characteristics of 1D-VCG; represents the standard normal distribution as the target for the latent space distribution.

7. The method for locating myocardial infarction based on cross-modal knowledge fusion according to claim 6, characterized in that: In step S5), the student enhancement optimization comprises the following steps: S51), input 1D-VCG into the encoder of the student model to generate a temporal feature representation f ts , and further optimized through a feature mapping layer U: f optimized =Uf ts In the formula, f optimized is the optimized feature of the feature mapping layer; U is the feature mapping layer; f ts Generate temporal feature representation; S52), the optimized feature f optimized Used to train the classifier and generate category prediction P pred ; and predict P by category pred Calculating cross entropy loss In the formula, CrossEntropy represents the cross entropy loss function, which is used to calculate the loss between the predicted result and the true label in the classification task, and y true Represents the true label, which is used to compare with the model prediction results; P pred Represents the category prediction result of the model, which is used to calculate the loss with the true label; S53), use distillation loss To optimize: Where P teacher is the output of the teacher model; λ is a trade-off factor used to adjust the weight between the cross entropy loss and the MSE loss; MSE stands for mean squared error loss, which is used to calculate the difference between the student model prediction results and the teacher model output to help the student model better imitate the teacher model; S54), the final loss of the cross-modal generator is: In the formula, α, β, and γ represent the weight coefficients of different loss terms, which are used to adjust the influence of each loss term in the final loss function; α corresponds to the knowledge enhancement loss The weight of β corresponds to the cross entropy loss. The weight of γ corresponds to the distillation loss. The weight of Enhance loss for knowledge; is the cross entropy loss; is the distillation loss.