A method and related device for predicting asymptomatic myocardial infarction based on a migratable spatiotemporal mask transformer

By using a method based on transferable spatiotemporal mask Transformer, combined with self-supervised pre-training and transfer learning, the problems of weak feature extraction capability and coarse multimodal fusion in UMI screening are solved, and early accurate prediction and stable identification of asymptomatic myocardial infarction are achieved.

CN121617642BActive Publication Date: 2026-04-21AFFILIATED HOSPITAL OF GUANGDONG MEDICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
AFFILIATED HOSPITAL OF GUANGDONG MEDICAL UNIV
Filing Date
2026-02-03
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing screening methods for asymptomatic myocardial infarction (UMI) rely on examinations such as electrocardiograms, which have problems such as weak feature extraction ability, poor model transferability, coarse multimodal fusion, and low system integration, making it difficult to identify UMI in its early stages.

Method used

We employ a method based on transferable spatiotemporal masking Transformer, which combines self-supervised pre-training and transfer learning with generative masking mechanism and missing modality adaptive multimodal fusion to improve the electrocardiogram feature extraction capability and achieve early and accurate prediction of asymptomatic myocardial infarction.

Benefits of technology

It improves the accuracy and stability of identifying asymptomatic myocardial infarction, reduces reliance on expensive labeled data, enhances the adaptability and universality of the ECG feature extractor, and supports clinical screening and assisted diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121617642B_ABST
    Figure CN121617642B_ABST
Patent Text Reader

Abstract

This application discloses a method and related device for predicting asymptomatic myocardial infarction based on a transferable spatiotemporal mask Transformer. The method includes: performing spatiotemporal mask modeling and self-supervised pre-training on an electrocardiogram (ECG) feature sequence constructed from ECG data to obtain a pre-trained ECG feature extractor; fine-tuning the pre-trained ECG feature extractor using a small sample dataset to construct a target ECG feature extractor for asymptomatic myocardial infarction identification; extracting ECG feature representations of the target object using the target ECG feature extractor; extracting and modeling features from cardiac magnetic resonance imaging (MRI) and clinical structured data of the target object based on the feature extraction module; adaptively fusing multimodal features; and finally generating an asymptomatic myocardial infarction risk prediction result based on the fused features. This application can achieve early and accurate prediction of asymptomatic myocardial infarction and can be widely applied in the field of artificial intelligence technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method and related device for predicting asymptomatic myocardial infarction based on a transferable spatiotemporal mask Transformer. Background Technology

[0002] Currently, unrecognized myocardial infarction (UMI) is a type of ischemic necrosis of the myocardium that lacks typical chest pain, chest tightness, and other symptoms in the acute phase, making it difficult to identify in a timely manner.

[0003] In related technologies, traditional UMI screening methods mainly rely on examinations such as electrocardiogram (ECG), cardiac magnetic resonance imaging (CMR), or coronary angiography. However, ECG recognition methods based on manual interpretation or traditional rules often fail when faced with weak and atypical waveforms in UMI. In recent years, automatic ECG recognition models based on convolutional neural networks (CNN) or recurrent neural networks (RNN) have been proposed, significantly improving the accuracy of recognizing typical diseases such as arrhythmias and acute myocardial infarction. However, for diseases like UMI, which have insidious onset, weak features, and high heterogeneity, current models still suffer from problems such as weak feature extraction capabilities, poor model transferability, coarse multimodal fusion, and low system integration.

[0004] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention

[0005] The embodiments of this application aim to at least partially address one of the technical problems in the related art. Therefore, the main objective of the embodiments of this application is to propose a method and related device for predicting asymptomatic myocardial infarction based on a transferable spatiotemporal mask Transformer, which can achieve early and accurate prediction of asymptomatic myocardial infarction, assisting relevant medical personnel in preliminary screening and diagnostic decision-making for asymptomatic myocardial infarction.

[0006] To achieve the above objectives, one aspect of this application proposes a method for predicting asymptomatic myocardial infarction based on a transferable spatiotemporal mask Transformer, the method comprising the following steps:

[0007] Obtain raw multimodal training data; wherein, the raw multimodal training data includes raw electrocardiogram data, raw cardiac magnetic resonance imaging data, and raw clinical structured data;

[0008] The original multimodal training data is preprocessed to obtain target multimodal training data; wherein, the target multimodal training data includes electrocardiogram feature sequences and small sample datasets;

[0009] The ECG feature sequence is input into a transferable spatiotemporal mask Transformer, and spatiotemporal mask modeling is performed on the ECG feature sequence using a generative masking mechanism to generate spatiotemporal mask features.

[0010] The transferable spatiotemporal mask Transformer is self-supervised pre-trained based on the spatiotemporal mask features, and then fine-tuned based on the small sample dataset using transfer learning to obtain an electrocardiogram feature extractor.

[0011] The current electrocardiogram (ECG) data of the object to be predicted is obtained, and features are extracted from the current ECG data using the ECG feature extractor to generate a current ECG feature representation.

[0012] A missing modality adaptive multimodal fusion mechanism is adopted to fuse the current electrocardiogram feature representation, the current cardiac magnetic resonance imaging feature representation corresponding to the object to be predicted, and the current clinical structured feature representation to generate a multimodal fused feature representation;

[0013] The multimodal fusion feature representation is input into the asymptomatic myocardial infarction risk classifier to generate asymptomatic myocardial infarction risk prediction results.

[0014] To achieve the above objectives, another aspect of this application proposes an asymptomatic myocardial infarction prediction device based on a transferable spatiotemporal mask Transformer, the device comprising the following modules:

[0015] A multimodal data acquisition module is used to acquire raw multimodal training data; wherein, the raw multimodal training data includes raw electrocardiogram data, raw cardiac magnetic resonance imaging data, and raw clinical structured data;

[0016] A multimodal data preprocessing module is used to preprocess the original multimodal training data to obtain target multimodal training data; wherein, the target multimodal training data includes electrocardiogram feature sequences and small sample datasets;

[0017] The spatiotemporal mask modeling module is used to input the electrocardiogram feature sequence into a transferable spatiotemporal mask Transformer, and combine a generative masking mechanism to perform spatiotemporal mask modeling on the electrocardiogram feature sequence to generate spatiotemporal mask features.

[0018] The self-supervised pre-training and transfer fine-tuning module is used to perform self-supervised pre-training of the transferable spatiotemporal mask Transformer based on the spatiotemporal mask features, and to combine transfer learning to perform transfer fine-tuning of the self-supervised pre-trained transferable spatiotemporal mask Transformer based on the small sample dataset to obtain an electrocardiogram feature extractor.

[0019] The current electrocardiogram feature extraction module is used to obtain the current electrocardiogram data of the object to be predicted, and to extract features from the current electrocardiogram data through the electrocardiogram feature extractor to generate a current electrocardiogram feature representation;

[0020] An adaptive multimodal fusion module is used to fuse the current electrocardiogram feature representation, the current cardiac magnetic resonance imaging feature representation corresponding to the object to be predicted, and the current clinical structured feature representation using a missing modality adaptive multimodal fusion mechanism to generate a multimodal fused feature representation;

[0021] The asymptomatic myocardial infarction risk prediction module is used to input the multimodal fusion feature representation into the asymptomatic myocardial infarction risk classifier to generate asymptomatic myocardial infarction risk prediction results.

[0022] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method.

[0023] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0024] To achieve the above objectives, another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0025] The embodiments of this application include at least the following beneficial effects: This application provides a method and related device for predicting asymptomatic myocardial infarction based on a transferable spatiotemporal mask Transformer. This scheme inputs the electrocardiogram feature sequence into a transferable spatiotemporal mask Transformer and combines a generative masking mechanism to perform spatiotemporal masking modeling on the electrocardiogram feature sequence. This can effectively improve the Transformer model's ability to identify features of non-obvious asymptomatic myocardial infarction, thereby improving the accuracy of model prediction. Furthermore, this method combines the long-distance dependency modeling advantages of the Transformer structure, overcoming the performance degradation limitations of traditional convolutional neural networks or recurrent neural networks when processing long sequences; by using spatiotemporal masking features to perform spatiotemporal masking on the transferable spatiotemporal mask Transformer... The algorithm performs self-supervised pre-training on the ECG feature extractor and then uses a small sample dataset to perform transfer learning on the transferable spatiotemporal mask Transformer to obtain an ECG feature extractor with good generalization ability. This allows the ECG feature extractor to more sensitively characterize the fine-grained changes in ECG associated with asymptomatic myocardial infarction, thereby enhancing its ability to represent occult myocardial ischemia and subclinical pathological changes, reducing dependence on expensive labeled data, and enhancing the adaptability and universality of the ECG feature extractor. In the actual prediction process, by introducing a missing modality adaptive multimodal fusion mechanism to achieve semantic alignment and dynamic weighted fusion of different modalities, the stability and accuracy of early identification of asymptomatic myocardial infarction can be improved, providing effective support for clinical screening and auxiliary diagnosis. Attached Figure Description

[0026] Figure 1 This is a flowchart illustrating the steps of a method for predicting asymptomatic myocardial infarction based on a transferable spatiotemporal mask Transformer, as provided in an embodiment of this application.

[0027] Figure 2 This is a flowchart illustrating a method for predicting asymptomatic myocardial infarction based on a transferable spatiotemporal mask Transformer, as provided in an embodiment of this application.

[0028] Figure 3 This is a schematic diagram of the structure of an asymptomatic myocardial infarction prediction device based on a transferable spatiotemporal mask Transformer provided in an embodiment of this application;

[0029] Figure 4 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0031] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”

[0032] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.

[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0034] Currently, asymptomatic myocardial infarction (UMI) is a type of ischemic necrosis of the myocardium that lacks typical chest pain and tightness during the acute phase, making it difficult to identify in a timely manner. It is commonly seen in high-risk cardiovascular populations such as the elderly, those with hypertension, and those with diabetes. Due to its "silent" onset, the rate of missed diagnosis is high, and intervention is delayed. It is often discovered retrospectively during follow-up examinations or when combined with other major events, seriously affecting the long-term prognosis of patients. This poses a significant challenge to accurate early warning and proactive intervention in clinical cardiovascular medicine.

[0035] Traditional UMI screening methods mainly rely on examinations such as electrocardiogram (ECG), cardiac magnetic resonance imaging (CMR), or coronary angiography. Among these, ECG, as a non-invasive and low-cost tool, is often used for preliminary risk assessment. However, ECG identification methods based on manual interpretation or traditional rules often fail when faced with weak or atypical UMI waveforms, especially slight ST segment depression (the isoelectric line between the end of the QRS complex and the beginning of the T wave in the ECG waveform) or T wave inversion (the rounded, asymmetrical waveform that follows the ST segment), which are scattered and inconsistent in distribution across multiple leads and are difficult to accurately identify using traditional algorithms or human experience.

[0036] In recent years, automatic ECG recognition models based on convolutional neural networks (CNNs) or recurrent neural networks (RNNs) have been proposed, significantly improving the accuracy of recognizing typical diseases such as arrhythmias and acute myocardial infarction. However, for diseases like UMI, which have insidious onset, weak features, and high heterogeneity, current algorithms still suffer from problems such as "weak feature extraction ability, poor model transferability, coarse multimodal fusion, and low system integration." Specific problems are analyzed below:

[0037] (1) Relying solely on a single ECG modality limits the information available: The mechanisms of UMI are complex and often involve the interaction of various imaging, laboratory tests, and past medical history. Modeling with a single ECG signal cannot fully characterize the risk of UMI.

[0038] (2) Ignoring spatial coupling and temporal evolution between leads: Traditional ECG modeling mostly processes one-dimensional time series, which fails to effectively encode the spatial distribution and potential synergistic relationship between leads, and does not introduce a spatiotemporal joint modeling mechanism.

[0039] (3) Poor model transferability and reliance on high-quality labels: Due to the low incidence of UMI and sparse labeled samples, the current model often suffers a sharp drop in performance in actual clinical deployment due to data distribution drift, and lacks self-supervised or pre-trained transfer ability.

[0040] (4) The multimodal fusion method is crude and it is difficult to achieve collaborative decision-making: Some studies have attempted to fuse ECG and clinical data, but most of them use simple splicing, averaging and other strategies to achieve information fusion. They lack a fusion mechanism based on semantic alignment and dynamic weights, resulting in weak model expression ability and difficulty in capturing deep connections between heterogeneous information.

[0041] (5) Low system integration, unable to support clinical intelligent closed loop: At present, most models remain at the static offline prediction level, lacking the integrated implementation of data collection, modeling, risk assessment, interpretability analysis and clinical advice output, which cannot meet the needs of simple operation, intelligent feedback and clinical support in actual medical scenarios.

[0042] In view of this, this application provides a method and related device for predicting asymptomatic myocardial infarction based on a transferable spatiotemporal mask Transformer. This method inputs electrocardiogram (ECG) feature sequences into a transferable spatiotemporal mask Transformer and combines a generative masking mechanism to perform spatiotemporal masking modeling of the ECG feature sequences. This effectively improves the Transformer model's ability to identify features of inconspicuous asymptomatic myocardial infarction, thereby improving the accuracy of model prediction. Furthermore, this method leverages the long-distance dependency modeling advantages of the Transformer structure, overcoming the performance degradation limitations of traditional convolutional neural networks or recurrent neural networks when processing long sequences. The method uses spatiotemporal masking features to perform spatiotemporal masking on the transferable spatiotemporal mask Transformer... Self-supervised pre-training was performed, and based on the pre-training, a small sample dataset was used to perform transfer learning on the transferable spatiotemporal mask Transformer to obtain an ECG feature extractor with good generalization ability. This enabled the ECG feature extractor to more sensitively characterize the fine-grained ECG changes associated with asymptomatic myocardial infarction, thereby enhancing its ability to represent occult myocardial ischemia and subclinical pathological changes, reducing dependence on expensive labeled data, and enhancing the adaptability and universality of the ECG feature extractor. In the actual prediction process, by introducing a missing modality adaptive multimodal fusion mechanism to achieve semantic alignment and dynamic weighted fusion of different modalities, the stability and accuracy of early identification of asymptomatic myocardial infarction can be improved, which can provide effective support for clinical screening and auxiliary diagnosis.

[0043] This application provides a method for predicting asymptomatic myocardial infarction based on a transferable spatiotemporal mask Transformer, relating to the field of artificial intelligence technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or in-vehicle terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the asymptomatic myocardial infarction prediction method based on a transferable spatiotemporal mask Transformer, but is not limited to the above forms.

[0044] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0045] Please see Figure 1 , Figure 1 This is an optional flowchart of a method for predicting asymptomatic myocardial infarction based on a transferable spatiotemporal mask Transformer, provided in an embodiment of this application. Figure 1 The method may include, but is not limited to, steps S101 to S107.

[0046] Step S101: Obtain raw multimodal training data; wherein, the raw multimodal training data includes raw electrocardiogram data, raw cardiac magnetic resonance imaging data, and raw clinical structured data;

[0047] In the specific implementation, large-scale multi-lead electrocardiograms from public datasets (such as PTB-XL, CPSC2018, and PhysioNet2017) and pre-collected real clinical data from hospitals (real clinical data from hospitals can be referred to as in-hospital data) are combined to construct a multi-source, multi-modal dataset covering different pathological states, which is also the original multi-modal training data. Among them, the public electrocardiogram data is mainly used for general representation learning in the self-supervised pre-training stage, and the real clinical data from hospitals is mainly used to construct a small sample dataset of UMI (asymptomatic myocardial infarction) and to adapt the task in the transfer fine-tuning stage.

[0048] The raw multimodal training data can be categorized into raw electrocardiogram data, raw cardiac magnetic resonance imaging data, and raw clinical structured data. In some optional embodiments, the raw multimodal training data can be imported via a terminal interface or API, and can undergo format conversion and storage management for subsequent training and inference. This application does not limit this aspect.

[0049] It should be noted that in each specific implementation of this application, when it is necessary to collect sensitive medical data from hospitals or other institutions or to process sensitive data related to the target of prediction, permission or consent from the user or relevant institution will be obtained first. Moreover, the collection, use and processing of this data will comply with relevant laws, regulations and standards.

[0050] Step S102: Preprocess the original multimodal training data to obtain target multimodal training data; wherein, the target multimodal training data includes electrocardiogram feature sequences and small sample datasets;

[0051] In some embodiments, step S102 may include: standardizing the first original electrocardiogram (ECG) data, the original cardiac magnetic resonance imaging (MRI) data, and the original clinical structured data to obtain first standardized data; wherein the first standardized data includes first standardized ECG data, standardized cardiac MRI data, and standardized clinical structured data; labeling the first standardized data to obtain an asymptomatic myocardial infarction label; wherein the labeling reference information used for labeling includes at least one of the following: clinical diagnostic records, cardiac MRI results, and myocardial injury-related examination indicators; constructing a small sample dataset based on the first standardized data and the asymptomatic myocardial infarction label; standardizing the second original ECG data to obtain second standardized ECG data; performing data segmentation on the second standardized ECG data to obtain several spatiotemporal patches; performing linear projection processing on each spatiotemporal patch to obtain an embedding vector corresponding to each spatiotemporal patch; and superimposing each embedding vector with position encoding and lead encoding respectively to obtain an ECG feature sequence.

[0052] Optionally, the raw electrocardiogram (ECG) data includes a first raw ECG data set and a second raw ECG data set. These two sets have different sources and purposes, and are used in different model training phases. The first raw ECG data set refers to ECG data extracted solely from pre-collected real clinical data within the hospital. Furthermore, the first raw ECG data set, the raw cardiac magnetic resonance imaging (MRI) data, and the raw clinical structured data all originate from the same subject and are used to construct a small-sample dataset containing asymptomatic myocardial infarction labels. The second raw ECG data set refers to ECG data obtained solely from a publicly available ECG dataset (i.e., the aforementioned target ECG dataset). This publicly available ECG dataset does not contain pre-collected real clinical data from the hospital, and the second raw ECG data set is not involved in the construction of the small-sample dataset; instead, it is used for self-supervised pre-training or pre-training modeling of the ECG feature extractor.

[0053] The standardization process can include, but is not limited to, resampling, noise filtering, and normalization. In this embodiment, resampling is mainly applied to ECG data (using a uniform sampling frequency) and some CMR data (e.g., pixel spacing normalization); noise filtering is mainly used for ECG data (e.g., baseline drift, power line interference), but CMR data undergoes slight filtering; normalization is required for all three types of data: ECG data is normalized using Z-score or min-max, CMR data is normalized using pixel intensity, and clinical data is normalized using mean / variance. It should be noted that the standardization process for the first and second original ECG data is the same.

[0054] The first standardized data includes first standardized electrocardiogram data, standardized cardiac magnetic resonance imaging data, and standardized clinical structured data.

[0055] In this embodiment, the pre-collected real clinical data from within the hospital undergoes a rigorous quality control process to remove samples with excessive noise or missing leads. Combined with annotations from clinicians, the data is categorized into two groups: a "normal / healthy group" and an "abnormal / disease group" for asymptomatic myocardial infarction. The normal / healthy group is the non-UMI group, referring to subjects without evidence of myocardial infarction and excluded from UMI through clinical evaluation. The abnormal / disease group is the UMI group, referring to subjects diagnosed with asymptomatic myocardial infarction based on a combination of clinical examination and imaging / laboratory evidence.

[0056] The target multimodal training data includes electrocardiogram feature sequences and small sample datasets.

[0057] For small sample datasets, which refer to UMI case samples actually collected and labeled by the hospital, these include: preprocessed 12-lead ECG data, i.e., first standardized ECG data; preprocessed clinical structured data (age, gender, blood pressure, blood lipids, history of diabetes, smoking history, etc.), i.e. standardized clinical structured data; preprocessed cardiac magnetic resonance imaging (CMR) images (DICOM (Digital Imaging and Communications in Medicine) format), i.e. standardized cardiac magnetic resonance imaging data; and UMI labels confirmed by doctors based on examinations, hospitalization records, MACE (Major Adverse Cardiovascular Events), or medical records, i.e., asymptomatic myocardial infarction labels.

[0058] For electrocardiogram feature sequences, it refers to the process of standardizing the second original electrocardiogram data to obtain the second standardized electrocardiogram data, then performing data segmentation on the second standardized electrocardiogram data to obtain several spatiotemporal patches, then performing linear projection on each spatiotemporal patch to obtain the embedding vector corresponding to each spatiotemporal patch, and finally superimposing each embedding vector with the position code and lead code to obtain the feature sequence.

[0059] Specifically, the implementation process of data sharding and embedding is as follows:

[0060] First, the ECG signal after data preprocessing Each lead signal Divided into non-overlapping spatiotemporal patches, among which, ECG signals Indicates containing The time series of leads, with a length of ; This represents the batch size, i.e., the number of ECG samples included in a single model input. Specifically, the calculation expression for data partitioning is as follows:

[0061] ;

[0062] in, This represents a spatiotemporal patch obtained by dividing electrocardiogram signals in chronological order. Indicates the number of spacetime patches. This represents the number of time steps corresponding to each spacetime patch; ,in, This indicates the total number of time steps of the electrocardiogram signal in a single lead. This represents the number of time steps corresponding to each spacetime patch.

[0063] Next, each spatiotemporal patch is transformed into an embedding vector through linear projection, and position embedding and lead embedding are added to each embedding vector to finally obtain the electrocardiogram feature sequence:

[0064] ;

[0065] ;

[0066] in, The linear projection embedding vector representing the spatiotemporal patch. and These represent location embedding and lead embedding, respectively. This represents the dimension index in the embedded vector. This indicates the position index of the spatiotemporal patch in the sequence. This represents the total dimension of the embedding vector.

[0067] By segmenting the electrocardiogram (ECG) signal into spatiotemporal patches and combining location embedding and lead embedding to obtain the ECG feature sequence, the model can efficiently capture the spatiotemporal relationship of the ECG signal based on the ECG feature sequence, providing a general representation for subsequent classification tasks.

[0068] In step S102, by performing corresponding preprocessing procedures (including resampling, noise filtering, normalization, and time slicing) on ​​the original multimodal training data, the consistency and comparability of multi-source data can be ensured.

[0069] Step S103: Input the ECG feature sequence into the transferable spatiotemporal mask Transformer, and perform spatiotemporal mask modeling on the ECG feature sequence in combination with the generative masking mechanism to generate spatiotemporal mask features.

[0070] In some embodiments, step S103 may include: inputting the electrocardiogram feature sequence into a transferable spatiotemporal masking Transformer; performing a random masking operation on the electrocardiogram feature sequence according to a preset ratio through the input layer in the transferable spatiotemporal masking Transformer to obtain a masked patch feature sequence and an unmasked patch feature sequence; inputting the unmasked patch feature sequence into the encoder in the transferable spatiotemporal masking Transformer to perform global feature modeling and generate spatiotemporal mask features.

[0071] In the model building phase, this embodiment employs a self-supervised learning framework, fully utilizing unlabeled data to develop a spatiotemporal mask generative learning model for spatiotemporal feature extraction from electrocardiograms (ECGs) and generating spatiotemporal mask features. The unlabeled data refers to ECG signal data from publicly available ECG datasets. This data does not contain asymptomatic myocardial infarction labels and is only used for spatiotemporal feature learning in the self-supervised pre-training phase, without participating in the construction of small sample datasets. In the input phase, the ECG signal, after time-slicing, is mapped into a feature sequence through segment embedding and spatiotemporal position encoding. This feature sequence serves as the input to the encoder in the self-supervised learning process for learning spatiotemporal context features. Specifically, in the input stage, the multi-lead electrocardiogram (ECG) signal is divided into several spatiotemporal patches according to a preset time window, and an ECG feature sequence is generated through patch embedding and spatiotemporal positional encoding. Subsequently, a time mask (T-Mask) is performed on the ECG feature sequence according to a preset ratio to obtain the masked spatiotemporal patch sequence and the unmasked spatiotemporal patch sequence. The unmasked spatiotemporal patch sequence is then input into a temporal encoder for global temporal dependency modeling, and a temporal representation is output. The temporal encoder is used to perform global temporal dependency modeling on the unmasked spatiotemporal patch sequence and generate a temporal representation.

[0072] In the self-supervised pre-training stage, to construct a generative mask reconstruction learning task, the temporal feature representation output by the temporal encoder and the mask position information are input into the temporal decoder. The temporal decoder uses temporal padding and introduces mask tokens as placeholders for the masked spatiotemporal patch sequence. Combined with the temporal position encoding and decoding side Transformer layers, it reconstructs the target features corresponding to the masked spatiotemporal patch sequence to form a self-supervised pre-training objective and optimize the model parameters.

[0073] It should be noted that the encoder-decoder structure is only used for feature learning in the self-supervised pre-training stage; in the subsequent transfer fine-tuning and asymptomatic myocardial infarction risk prediction stages, only the time encoder is retained as an electrocardiogram feature extractor.

[0074] In some optional embodiments, the transferable spatiotemporal masking Transformer can be implemented using a multi-layer Transformer coding structure. Position encoding, lead encoding, and mask marking can be implemented in a fixed or learnable manner, but are not limited to this. Optionally, the temporal position embedding can be implemented using a fixed sine wave position embedding method, while other special embeddings (such as mask embedding, lead embedding, and SEP (Separator) embedding) can be learned during training. Regarding pre-training, embodiments of this application mask and reconstruct 75% of randomly selected slices. Generative pre-training uses a mean squared error loss function in some optional embodiments, and in other optional embodiments, it can be further combined with a contrastive learning loss function for joint optimization.

[0075] During the self-supervised pre-training phase, the length is... Patch sequence by mask ratio Perform random occlusion. Let the set of indices of the occluded patches be... The set of unmasked patch indices is Then there is , ,and , The encoder input is an unmasked patch sequence. .

[0076] The formula for calculating the hidden feature H generated by the encoder is as follows:

[0077] ;

[0078] This calculation formula defines the feature encoding process, where This represents the global spatiotemporal features output by the encoder. This represents a temporal encoder, which consists of a multi-layered Transformer structure. This is the set of embedded features from the unmasked input segments. The encoder obtains a high-dimensional spatiotemporal feature representation for subsequent tasks by performing contextual modeling on the unmasked segments. This refers to the hidden features. Next, the decoder uses the hidden features output by the time encoder. The masked patches are then reconstructed using the mask location information. The calculation formula for the reconstruction process is as follows:

[0079] ;

[0080] in, This represents the reconstruction of the feature vector. This represents the temporal decoder; the decoder uses the hidden features output by the encoder. Reconstructing the embedded representation of the occluded patch ,in The reconstruction loss is defined using the mean squared error:

[0081] ;

[0082] in, Indicates the mask ratio, The total number of patches obtained by partitioning the input sequence. For the set of indices of the masked patch, satisfying ; This indicates that the reconstruction error corresponding to the masked patch is summed and normalized. Indicates the reconstruction loss. Indicates the first output of the decoder The reconstructed feature vector of a masked patch, This represents the corresponding real patch embedding feature vector.

[0083] In step S103, the electrocardiogram feature sequence corresponding to the 12-lead ECG signal is modeled as a spatiotemporal image. This breaks through the limitations of traditional one-dimensional CNN or RNN that ignore the spatial structure and temporal dependence of leads. By learning global context features through a masking mechanism, the ability to capture weak features of asymptomatic myocardial infarction (UMI) can be enhanced, and finally, spatiotemporal mask features are generated.

[0084] Step S104: Perform self-supervised pre-training on the transferable spatiotemporal mask Transformer based on the spatiotemporal mask features, and combine transfer learning to fine-tune the self-supervised pre-trained transferable spatiotemporal mask Transformer based on the small sample dataset to obtain the electrocardiogram feature extractor.

[0085] In some embodiments, step S104 may include: inputting spatiotemporal mask features and occlusion patch feature sequences into a transferable spatiotemporal mask Transformer; reconstructing the occlusion patch feature sequences based on the spatiotemporal mask features using a generative mask reconstruction mechanism; and jointly optimizing and training the transferable spatiotemporal mask Transformer using a contrastive learning strategy; and fine-tuning the jointly optimized and trained transferable spatiotemporal mask Transformer using a small sample dataset using transfer learning to obtain an electrocardiogram feature extractor.

[0086] In practical applications, insufficient sample size is a common challenge in UMI research. This application's embodiments alleviate this problem by mining latent patterns in the data through self-supervised learning. During the self-supervised pre-training stage, no manual labels are relied upon; effective feature representations are learned through a pre-defined self-supervised pre-training task. In the downstream task, the feature representations extracted by the electrocardiogram feature extractor are input into a classifier to predict the risk of asymptomatic myocardial infarction.

[0087] In the specific implementation, the Masked Patch Prediction (MPP) task is introduced during the self-supervised pre-training stage. By randomly masking parts of the ECG spatiotemporal segments, the model is required to recover the embedding representation of the masked segments, thus prompting the model to simultaneously model local morphological features and global temporal dependencies. The corresponding mask patch prediction loss function is defined as:

[0088] ;

[0089] in, This represents the mask patch prediction loss function, used to measure the model's reconstruction error of the masked ECG spatiotemporal patch during the self-supervised pre-training phase; The index representing the spatiotemporal patch of the obscured electrocardiogram; A set of indices representing the masked patch; Indicates the number of patches that are hidden; Indicates the first The true embedded feature vector corresponding to each masked patch This represents the model's predicted embedding feature vector for the patch. In this way, the model can learn a discriminative spatiotemporal representation of the electrocardiogram under unsupervised conditions.

[0090] Furthermore, contrastive learning is introduced as an auxiliary optimization objective to enhance the model's ability to model the discriminative relationships between different samples. The contrastive learning loss function is defined as follows:

[0091] ;

[0092] in, This represents the contrastive learning loss function; This indicates the number of samples in a batch, i.e., the number of samples participating in the contrastive learning. Indicates the first Feature representation of each electrocardiogram sample; This represents the feature representation that forms a positive sample pair with it; This represents the characteristics of other samples in the same batch and is used as a negative sample. The similarity measure between positive and negative samples is usually calculated using cosine similarity. The calculation formula is as follows:

[0093] ;

[0094] In the transfer learning phase, the model weights pre-trained on a large-scale dataset are first used as initialization parameters, which are defined as follows:

[0095] ;

[0096] in, These represent the initialization parameters for the migration fine-tuning phase. This refers to the model weight parameters obtained through self-supervised pre-training on publicly available large-scale electrocardiogram (ECG) datasets, such as parameters obtained by pre-training on ECG datasets like PhysioNet, PTB-XL, or CPSC.

[0097] Building upon this, step S104 further includes a transfer fine-tuning process based on electrocardiogram (ECG) data extracted from pre-collected real clinical data within the hospital. Specifically, firstly, the feature extraction module in the transferable spatiotemporal mask Transformer is pre-trained under self-supervised supervision using publicly available ECG datasets such as PTB-XL, CPSC2018, and PhysioNet, enabling the model to acquire general spatiotemporal modeling capabilities across devices and populations. Subsequently, the feature extraction module in the transferable spatiotemporal mask Transformer is fine-tuned using ECG data extracted from pre-collected real clinical data within the hospital to adapt to different equipment types, acquisition processes, and population distribution characteristics within the hospital, ultimately resulting in an ECG feature extractor.

[0098] The electrocardiogram (ECG) data extracted from pre-collected real clinical data within hospitals can be labeled or unlabeled. During fine-tuning, the embedding layer and the first few Transformer coding layers can be frozen, and only the parameters of the last few layers are updated to reduce the risk of overfitting. Through this method, the ECG feature extractor, fine-tuned from the ECG data extracted from real clinical data within hospitals, can more sensitively characterize the fine-grained ECG changes associated with asymptomatic myocardial infarction (UMI), thereby enhancing its ability to represent occult myocardial ischemia and subclinical pathological changes.

[0099] In this embodiment, considering the scarcity of UMI data, self-supervised learning (SSL) and transfer learning techniques are employed for model pre-training and fine-tuning: after pre-training on public or large-scale ECG datasets, the model is transferred to a small-sample UMI dataset to build an adaptation mechanism in low-resource environments, improving the model's generalization ability and robustness. Ultimately, through efficient representation learning and inference in small-sample scenarios, the accurate operation of the early UMI prediction model is achieved, providing feasible technical support for clinical screening and intervention. It should be noted that the training objective of step S104 is to obtain an ECG feature extractor for the UMI prediction task.

[0100] In this embodiment, separate Transformer feature extraction modules are also constructed for feature extraction and modeling of cardiac magnetic resonance imaging (MRI) and clinical structured data. Specifically, cardiac MRI data is feature extracted using a cardiac MRI feature extraction module, which uses a Transformer coding structure to model image sequence or slice features to obtain cardiac MRI feature representations. Clinical structured data is feature extracted using a clinical structured feature extraction module, which uses a Transformer coding structure to model structured variable sequences to obtain clinical structured feature representations.

[0101] Step S105: Obtain the current electrocardiogram (ECG) data of the object to be predicted, and extract features from the current ECG data using the ECG feature extractor to generate a current ECG feature representation;

[0102] In the specific implementation process, the current electrocardiogram (ECG) data of the object to be predicted is input into the ECG feature extractor obtained in step S104. Through the encoding structure of the transferable spatiotemporal mask Transformer, the corresponding ECG feature representation (ECG spatiotemporal feature sequence or low-dimensional embedding representation) is extracted. The ECG feature representation is used to characterize the ECG activity pattern and potential pathological risk characteristics of the object to be predicted at the current detection time.

[0103] In some embodiments, the feature representation output by the ECG feature extractor may be a fixed-dimensional vector embedding; in other embodiments, the feature representation output by the ECG feature extractor may also be a feature sequence that preserves the temporal structure, so that subsequent multimodal attention mechanisms can perform temporal alignment and fusion.

[0104] It should be noted that step S105 only extracts and represents the electrocardiogram (ECG) features and does not involve the final risk classification decision. The ECG feature representation obtained in step S105 serves as the input to the subsequent multimodal fusion module, and is dynamically fused with the clinical structured feature representation and / or cardiac magnetic resonance imaging feature representation of the same subject to be predicted. In subsequent steps, it completes the prediction and hierarchical output of asymptomatic myocardial infarction risk.

[0105] Step S106: Using a missing modality adaptive multimodal fusion mechanism, the current electrocardiogram feature representation, the current cardiac magnetic resonance imaging feature representation corresponding to the object to be predicted, and the current clinical structured feature representation are fused to generate a multimodal fused feature representation;

[0106] In some embodiments, step S106 may include: constructing a first modality availability indicator vector corresponding to the current electrocardiogram feature representation; constructing a second modality availability indicator vector corresponding to the current cardiac magnetic resonance imaging (MRI) feature representation; constructing a third modality availability indicator vector corresponding to the current clinical structured feature representation; semantically aligning the current electrocardiogram feature representation, the current MRI feature representation, and the current clinical structured feature representation using a contrastive learning strategy to obtain current electrocardiogram alignment features, current MRI alignment features, and current clinical structured feature representation, respectively; and aligning the current electrocardiogram alignment features, the current MRI alignment features, and the current clinical structured feature representation. The homogeneous feature input is used to an adaptive fusion network, and combined with the first modality availability indicator vector, the second modality availability indicator vector, and the third modality availability indicator vector, the current ECG alignment feature, the current cardiac magnetic resonance imaging alignment feature, and the current clinical structured alignment feature are dynamically weighted to generate a multimodal fusion feature representation. Among them, when there is unavailable target modality data, the target modality feature corresponding to the target modality data is set as a preset mask vector, and during the dynamic weighting of the current ECG alignment feature, the current cardiac magnetic resonance imaging alignment feature, and the current clinical structured alignment feature, the weight of the target modality data is constrained to zero or approximately zero based on the preset mask vector.

[0107] In the process of multimodal information fusion, the missing modality adaptive multimodal fusion process based on modality availability perception and adaptive attention mechanism mainly includes steps such as single modality feature encoding, modality availability modeling, representation-level dynamic fusion, and decision-level adaptive fusion.

[0108] (1) Single-modal feature encoding: First, feature encoding is performed on different modalities of data. Specifically, electrocardiogram (ECG) data is extracted using the ECG feature extractor obtained in step S104 to obtain ECG feature representations (completed in step S105); cardiac magnetic resonance imaging (MRI) data is feature-extracted using the cardiac MRI feature extraction module, which uses a Transformer encoding structure to model image sequence or slice features to obtain cardiac MRI feature representations; clinical structured data is feature-extracted using the clinical structured feature extraction module, which uses a Transformer encoding structure to model structured variable sequences to obtain clinical structured feature representations. Each modal feature extraction module outputs a fixed-dimensional modal feature representation, which is used as input to the subsequent missing modality adaptive multimodal fusion module.

[0109] (2) Modal availability modeling: Given that different modal data may be missing, unavailable, or inconsistent in quality in real clinical environments, a corresponding modal availability indicator vector is constructed for each modality to characterize whether the modality is available in the current sample and its information completeness level. Among them, the modal availability indicator vector is used as prior information to participate in the subsequent fusion process, enabling the model to perceive the availability status of different modalities.

[0110] (3) Representation-level missing modality adaptive fusion: In the representation-level fusion stage, a contrastive learning mechanism is introduced to semantically align the feature representations of different modalities (ECG feature representation, cardiac magnetic resonance imaging feature representation, and clinical structured feature representation), so that image features and clinical features have consistent representations in the shared feature space. The cross-modal contrastive learning loss function is defined as:

[0111] ;

[0112] in, This represents the cross-modal contrastive learning loss function. Indicates the first The modal characteristics of electrocardiograms or cardiac imaging corresponding to each sample; This indicates the corresponding clinical features; This represents the number of samples in a batch, i.e., the number of samples participating in the comparative learning. Similarity function. Using the definition of cosine similarity:

[0113] ;

[0114] After completing cross-modal semantic alignment, the features of each modality are input into an adaptive fusion network, and modality availability information is combined to dynamically weight the features of different modalities, resulting in a multimodal fusion feature representation. The calculation process is as follows:

[0115] ;

[0116] in, This represents a fully connected mapping layer or its equivalent nonlinear transformation structure used for multimodal feature fusion.

[0117] In some optional embodiments, when unavailable target modal data exists, the target modal features corresponding to the target modal data are set as a preset mask vector. During the dynamic weighting of the current ECG alignment features, current cardiac MRI alignment features, and current clinical structured alignment features, the weight of the target modal data is constrained to zero or approximately zero based on the preset mask vector. This approach ensures the continuity and stability of the multimodal data fusion process when some modal data is missing, avoids interference from invalid data, and improves the accuracy of the fusion results and the clinical applicability of the solution.

[0118] Step S107: Input the multimodal fusion feature representation into the asymptomatic myocardial infarction risk classifier to generate asymptomatic myocardial infarction risk prediction results.

[0119] Among them, the asymptomatic myocardial infarction risk classifier may include, but is not limited to, classification models based on multilayer perceptron or Softmax structures, for generating asymptomatic myocardial infarction risk prediction results.

[0120] In the specific implementation, the multimodal fusion feature output in step S106 is represented as... Input an in-modal prediction branch (such as a Softmax classification predictor), and output the prediction result of asymptomatic myocardial infarction risk. Its output format is:

[0121] ;

[0122] in, The predicted probability vector represents the risk of target asymptomatic myocardial infarction and is used to represent the probability distribution of a sample belonging to different risk categories. This represents the multimodal fusion feature representation vector. and These represent the weight matrix and bias vector of the classification layer, respectively.

[0123] In some embodiments, after step S107, the method may further include: determining the asymptomatic myocardial infarction risk level based on the asymptomatic myocardial infarction risk prediction results; performing post-processing analysis on the multimodal fusion feature representation based on the attention weights of the asymptomatic myocardial infarction risk classifier and the SHAP (SHapley Additive exPlanations) method to generate visual interpretation results of key leads and key features; generating personalized intervention recommendations based on the asymptomatic myocardial infarction risk prediction results, the asymptomatic myocardial infarction risk level, and the visual interpretation results; generating an asymptomatic myocardial infarction risk prediction report based on the asymptomatic myocardial infarction risk prediction results, the asymptomatic myocardial infarction risk level, the visual interpretation results, and the personalized intervention recommendations, and displaying the asymptomatic myocardial infarction risk prediction report on the visualization interface.

[0124] In its implementation, the Attention map (attention weight map) of the target asymptomatic myocardial infarction risk classifier and SHAP are used to perform interpretability analysis on the multimodal fusion feature representation, generating visualization results of key leads and key features to enhance medical personnel's trust in the model results. Based on this, preliminary intervention suggestions or personalized intervention suggestions are generated according to the risk prediction results, interpretability analysis conclusions, and risk levels, achieving a complete closed loop from risk prediction to clinical decision support. Finally, an asymptomatic myocardial infarction risk prediction report is generated by integrating all output results. This report may include, but is not limited to, prediction time, risk prediction results, and interpretable analysis results such as Attention heatmaps and SHAP feature contribution maps.

[0125] To enable practical application of the model, embodiments of this application can embed the model into an intelligent diagnostic system, developing a convenient and efficient intelligent analysis platform for automatic electrocardiogram analysis and multi-source multimodal data fusion. The system adopts a modular architecture design, including modules for data import, real-time analysis, and diagnostic output. The system supports real-time uploading and processing of in-hospital data, displaying diagnostic results and explanatory information through an intuitive user interface to provide decision support for clinicians. During the clinical testing phase, the intelligent diagnostic system can be deployed in multiple in-hospital scenarios (such as outpatient clinics, inpatient wards, and telemedicine centers) for pilot applications, collecting feedback from doctors and patients, and iteratively optimizing the model and system interface based on actual needs.

[0126] It should be noted that the asymptomatic myocardial infarction risk prediction results generated by the asymptomatic myocardial infarction risk classifier, as well as the subsequent asymptomatic myocardial infarction risk prediction reports containing interpretable analysis results such as intervention recommendations, are mainly used as an auxiliary reference for medical practitioners to conduct clinical assessments and do not constitute a final diagnosis of asymptomatic myocardial infarction. The diagnosis, condition assessment, and treatment plan formulation are all determined by qualified doctors or medical personnel based on a comprehensive assessment of clinical symptoms, signs, other examination and test results, and clinical experience, according to the actual application scenario. The prediction results and related reports in this application are used to provide technical support for clinical decision-making.

[0127] Steps S101 to S107 of this embodiment, by inputting the ECG feature sequence into a transferable spatiotemporal mask Transformer and combining it with a generative masking mechanism to perform spatiotemporal masking modeling of the ECG feature sequence, can effectively improve the Transformer model's ability to identify features of inconspicuous asymptomatic myocardial infarction, thereby improving the accuracy of model prediction. Furthermore, this method, combined with the long-distance dependency modeling advantages of the Transformer structure, overcomes the performance degradation limitations of traditional convolutional neural networks or recurrent neural networks when processing long sequences. Self-supervised pre-training of the transferable spatiotemporal mask Transformer is performed using spatiotemporal masking features, and based on the pre-training... Furthermore, a small sample dataset is used to perform transfer learning on the transferable spatiotemporal mask Transformer to obtain an ECG feature extractor with good generalization ability. This enables the ECG feature extractor to more sensitively characterize the fine-grained changes in ECG associated with asymptomatic myocardial infarction, thereby enhancing its ability to represent occult myocardial ischemia and subclinical pathological changes, reducing dependence on expensive labeled data, and enhancing the adaptability and universality of the ECG feature extractor. In the actual prediction process, by introducing a missing modality adaptive multimodal fusion mechanism to achieve semantic alignment and dynamic weighted fusion of different modalities, the stability and accuracy of early identification of asymptomatic myocardial infarction can be improved, which can provide effective support for clinical screening and auxiliary diagnosis.

[0128] To explain in detail the principles of the technical solution of this application, the overall process of this application will be described below with reference to some specific embodiments. It is easy to understand that the following is an explanation of the technical principles of this application and should not be regarded as a limitation of this application.

[0129] This application addresses the problems of weak feature extraction ability, poor model transferability, coarse multimodal fusion, and low system integration in current methods for identifying asymptomatic myocardial infarction (UMI). It proposes an asymptomatic myocardial infarction prediction method based on a transferable spatiotemporal mask Transformer, which enables deep collaborative modeling of ECG, imaging, and clinical data, effectively improving the accuracy of early risk identification and model generalization ability, and constructing a complete integrated clinical support system encompassing "model building-prediction-interpretation-recommendation". The core of this application lies in three major structural modules: a transferable spatiotemporal mask Transformer model, a multimodal alignment and fusion mechanism, and an integrated intelligent prediction system. Specifically, the overall architecture of the asymptomatic myocardial infarction prediction system corresponding to the transferable spatiotemporal mask Transformer includes a data access module, an ECG preprocessing and spatiotemporal patching module, a spatiotemporal mask Transformer modeling module, a self-supervised pre-training and transfer learning module, a multimodal feature fusion module, a UMI risk prediction module, an interpretive analysis and recommendation generation module, and a front-end display and data management module. This application provides a method for predicting asymptomatic myocardial infarction based on a transferable spatiotemporal mask Transformer. The method mainly covers data acquisition and access, spatiotemporal mask modeling, model pre-training and transfer learning, multimodal fusion, risk prediction, interpretation and analysis, and front-end display. The specific details are as follows:

[0130] Firstly, regarding data acquisition and access, the system's data access module supports multimodal input. Multimodal data can include, but is not limited to, raw 12-lead ECG data (.mat or .txt format), clinical structured data (such as gender, age, blood pressure, history of diabetes, blood lipids, smoking history, etc.), and DICOM images from cardiac magnetic resonance imaging (CMR). Multimodal data can be acquired through interface integration or manual input. Interfaces can include, but are not limited to, hospital information systems or image archiving systems.

[0131] Subsequently, the system first preprocesses the ECG signal, including resampling, noise filtering, and normalization. After preprocessing, the ECG signal of each lead is divided into fixed-length spatiotemporal patches, and the spatiotemporal patches are mapped into feature sequences through lead embedding and time position embedding. Then, the feature sequences are input into the spatiotemporal mask Transformer modeling module, and combined with the self-supervised learning mechanism of random masking segments, the global spatiotemporal features of the ECG signal are modeled, and finally a 128-dimensional global ECG representation vector is generated.

[0132] During the model training phase, the system uses large-scale public datasets such as PTB-XL and CPSC2018 for self-supervised pre-training, and adopts mask reconstruction loss and contrastive learning for joint optimization. Subsequently, by freezing low-level embeddings and fine-tuning high-level parameters, the model is transferred to the UMI small sample dataset to complete the adaptation for specific tasks, and finally the trained risk classifier is obtained. In the multimodal fusion module, cardiac magnetic resonance imaging (MRI) data undergoes feature extraction through the MRI feature extraction module. This module uses a Transformer coding structure to model the image sequence or slice features, resulting in an MRI feature representation. Clinical structured data undergoes feature extraction through the clinical structured feature extraction module. This module uses a Transformer coding structure to model the structured variable sequence, resulting in a clinical structured feature representation. The MRI feature representation, the clinical structured feature representation, and the ECG feature representation output by the ECG feature extractor are all input into a cross-modal attention mechanism / adaptive fusion network for alignment and fusion. Multimodal attention pooling is then used to obtain the fused comprehensive feature representation (i.e., the multimodal fusion feature representation). This comprehensive feature representation serves as the input to a risk classifier, which outputs the asymptomatic myocardial infarction risk prediction result.

[0133] In the risk prediction and stratification phase, the system uses a risk prediction network to output the probability of asymptomatic myocardial infarction and makes risk judgments based on preset thresholds; these thresholds are configurable and can be dynamically adjusted according to clinical needs. For interpretable analysis, the system combines Transformer attention icons to annotate key leads and time windows, uses the SHAP algorithm to perform interpretability analysis on the fused features, and generates personalized health recommendations based on the predicted risk level and clinical characteristics using a large language model, including recommendations for follow-up examinations, lifestyle interventions, and referrals. The front-end display and doctor workstation modules are developed based on Vue or React, intuitively presenting raw signals, risk scores, interpretation results, and recommendation text. They support PDF report export, case archiving, historical tracking, and batch assessment of multiple patients, facilitating clinical application and research expansion for doctors.

[0134] It is understood that the asymptomatic myocardial infarction (UMI) prediction method based on transferable spatiotemporal mask Transformer provided in this application follows the overall process of "ECG feature extractor construction → in-hospital fine-tuning and adaptation → ECG feature representation extraction → multimodal fusion → risk prediction output", specifically including:

[0135] (1) Public data pre-training stage: First, based on public electrocardiogram datasets such as PTB-XL and CPSC2018, the electrocardiogram feature extractor is pre-trained in a self-supervised manner to learn a spatiotemporal representation with good generalization ability.

[0136] (2) In-hospital data fine-tuning stage: Subsequently, the electrocardiogram data extracted from the pre-collected real clinical data in the hospital is used to perform migration fine-tuning of the electrocardiogram feature extractor, so as to adapt it to the distribution of in-hospital data and form an electrocardiogram feature representation for the asymptomatic myocardial infarction prediction task;

[0137] (3) ECG feature representation extraction: Based on the transfer-adjusted ECG feature extractor, the ECG feature representation of the object to be predicted is extracted;

[0138] (4) Multimodal fusion and prediction stage: After obtaining the finely tuned electrocardiogram feature representation, it is dynamically fused with the clinical structured features and / or cardiac magnetic resonance imaging (CMR) features of the same subject, and finally outputs the prediction results and risk stratification of asymptomatic myocardial infarction.

[0139] Please see Figure 2 , Figure 2 This is a flowchart illustrating a method for predicting asymptomatic myocardial infarction based on a transferable spatiotemporal mask Transformer, as provided in an embodiment of this application. Figure 2 As shown, based on the basic architecture of the asymptomatic myocardial infarction prediction system, the overall implementation process of the asymptomatic myocardial infarction prediction method is as follows:

[0140] Step 1: Data acquisition and input in the data access module: This system supports the acquisition and input of multimodal medical data, including 12-lead standard electrocardiogram (ECG, e.g., 500Hz sampling rate, 10-second recording time), cardiac magnetic resonance imaging (CMR, DICOM format), and clinical structured information (e.g., gender, age, history of diabetes, blood lipids, blood pressure, smoking status, etc.). All data, after being uploaded to the system front-end, undergoes standardized format conversion and is stored in the system database to ensure consistency and usability for subsequent analysis.

[0141] Step 2: ECG preprocessing and patch construction are performed in the ECG preprocessing and spatiotemporal patching module: the raw ECG signal is processed into a three-dimensional tensor. ,in Indicates the number of samples in the batch. Indicates the number of leads in the electrocardiogram (ECG) signal. This indicates the number of sampling points (500Hz × 10 seconds). During signal preprocessing, the system divides the timing signal of each lead into a fixed window size. Step length The time sequence is divided into 50 non-overlapping temporal patches. Each temporal patch is then mapped to a corresponding embedding vector and fused with positional encoding and lead encoding in the embedding space to form a feature sequence representation that can be input into the model.

[0142] Step 3: Spatiotemporal masking Transformer modeling is performed in the spatiotemporal masking Transformer modeling module. This stage borrows the structure of the Visual Transformer (ViT). First, a random masking operation is performed on the patch feature sequence obtained in Step 2 at the input layer, with a masking ratio of 75%. The remaining unmasked patches are used as the input to the encoder. The encoder part of the Visual Transformer model consists of 12 stacked standard Transformer Blocks, each layer containing a multi-head attention mechanism and a feedforward network for global feature modeling of the unmasked patches. The output layer uses a temporal decoder structure to reconstruct the masked patches, with the training objective being to minimize the mean squared error of the reconstruction. Furthermore, the Visual Transformer model introduces a contrastive learning branch, further optimizing the discriminative ability of the feature representation by defining positive and negative sample pairs and using NT-Xent loss.

[0143] Step 4: Pre-training and fine-tuning of the ECG feature extractor with in-hospital data in the self-supervised pre-training and transfer learning module: During model training, firstly, a transferable spatiotemporal mask Transformer is self-supervised pre-trained based on large-scale public ECG datasets such as PTB-XL and CPSC2018, using mask reconstruction error and contrastive learning loss as training objectives to learn ECG spatiotemporal feature representations with good generalization ability. Subsequently, the ECG feature extractor constructed by the self-supervised pre-training is fine-tuned using ECG data extracted from pre-collected real clinical data in the hospital to adapt to differences in in-hospital equipment, collection procedures, and population distribution. During the fine-tuning process, the embedding layer and the first few Transformer layers can be frozen, and only the parameters of the last few layers are updated, thereby obtaining ECG feature extraction capabilities for the asymptomatic myocardial infarction prediction task under limited sample conditions.

[0144] Step 5: Multimodal data alignment and fusion are performed in the multimodal feature fusion module. During multimodal information fusion, ECG features are output as a 128-dimensional embedded representation by the ECG feature extractor, which was pre-trained on the publicly available data in Step 4 and fine-tuned through in-hospital data migration. Cardiac magnetic resonance imaging (MRI) data undergoes feature extraction through the cardiac MRI feature extraction module, which outputs a cardiac MRI feature representation using a Transformer encoding structure. Clinical structured information undergoes feature extraction through the clinical structured feature extraction module, which outputs a clinical structured feature representation using a Transformer encoding structure. These three types of feature representations achieve semantic alignment and dynamic fusion under the influence of cross-modal attention mechanisms and modality availability modeling, generating a multimodal fused feature representation for subsequent prediction.

[0145] Step 6: Perform risk prediction and model interpretation in the UMI risk prediction and interpretive analysis module: Input the multimodal fusion feature representation into the risk prediction network (such as the Softmax classifier), output the predicted probability of asymptomatic myocardial infarction risk, and perform interpretive analysis based on the intermediate features of the model and attention weight information to identify ECG leads and time series regions that contribute significantly to the prediction results, providing interpretable evidence for clinical decision-making.

[0146] Step 7: Based on the completion of asymptomatic myocardial infarction risk prediction, in some optional embodiments, interpretive analysis and auxiliary intervention suggestions can be generated from the prediction results. Specifically, the system can combine the patient's structured feature information, risk level, and model output results to generate personalized management suggestions or follow-up prompts; personalized management suggestions can be output in natural language text form, and their generation method can be combined with a rule engine or language generation model, but is not limited to this. In some optional embodiments, the method proposed in this application can be integrated into a clinical decision support system, displaying risk prediction results and interpretive information through a graphical interface, and supporting the storage, export, or report generation of results; the system architecture, front-end display method, report format, and deployment method can be implemented according to specific application scenarios, and this application does not limit them.

[0147] It should be noted that this embodiment is only a brief illustrative description of the overall process of a method for predicting asymptomatic myocardial infarction based on a transferable spatiotemporal mask Transformer. Detailed descriptions of each step can be found in the relevant content of the foregoing embodiments, and will not be repeated here. It is understood that this application does not impose any limitations on this.

[0148] In this embodiment, the system deployment method can include local deployment and cloud deployment. Local deployment supports deployment on a hospital intranet GPU (Graphics Processing Unit) server (with configurable GPU computing power to meet training and inference needs), or running a lightweight model on an edge server (Jetson, MiniPC). Cloud deployment supports containerizing the model into a Docker image and deploying it on a cloud platform such as a hospital private cloud, and connecting to the hospital's HIS (Hospital Information System) system through an HTTPS (Hypertext Transfer Protocol Secure) interface.

[0149] For example, key parameters and configuration examples are shown in Table 1:

[0150] Table 1: Key Parameters and Configuration Examples

[0151]

[0152] It should be noted that the model structure and parameter configuration listed in Table 1 are merely exemplary implementations used to illustrate the feasible implementation of the method of this application, and do not constitute a limitation on the scope of protection of this application. Those skilled in the art can make equivalent substitutions or adjustments to the model structure and parameters without departing from the technical concept of this application.

[0153] Furthermore, during the experimental validation process, the model will undergo performance testing on both publicly available datasets (such as PTB-XL, CPSC2018, and PhysioNet2017) and in-hospital data to comprehensively evaluate its performance in the UMI classification task. The experiments employ multidimensional metrics (such as accuracy, F1 score, and AUROC (Area Under the Receiver Operating Characteristic Curve)) to evaluate model performance and compare it with traditional methods (such as convolutional neural networks and long short-term memory networks) and other self-supervised learning models (such as contrastive learning and masked autoencoders) to verify the superiority of the proposed method. Table 2 shows a comparison of the experimental results of the proposed method on publicly available datasets and in-hospital data with some related methods, illustrating the effectiveness and feasibility of the proposed method in the asymptomatic myocardial infarction prediction task.

[0154] Table 2: Comparison of Data Between This Application and OpenECG

[0155]

[0156] As shown in Table 2, the experimental results demonstrate that the proposed method outperforms the baseline method in predicting multiple publicly available electrocardiogram datasets and in-hospital data, thus validating the effectiveness and feasibility of the proposed method in predicting asymptomatic myocardial infarction. Further experiments show that the proposed method maintains good predictive stability and generalization ability across different datasets, providing a more reliable technical means for the early identification of asymptomatic myocardial infarction.

[0157] In summary, the key points of the asymptomatic myocardial infarction prediction method based on transferable spatiotemporal mask Transformer provided in this application are:

[0158] (1) Spatiotemporal masking Transformer ECG signal modeling method: The multi-lead ECG signal is transformed into a "spatiotemporal patch matrix", realizing the joint modeling of spatial coupling and temporal dynamics; at the same time, a generative masking mechanism is introduced, which uses the unmasked part to predict the masked region, effectively improving the model's ability to identify non-obvious UMI features such as weak ST-T segment anomalies. This method combines the long-distance dependency modeling advantages of the Transformer structure and overcomes the performance degradation of traditional RNN and CNN in long sequence processing.

[0159] (2) Self-supervised pre-training and transfer learning strategy: To address the challenge of small sample sizes in UMI (Understanding Measuring Injuries), this application designs a self-supervised pre-training strategy based on large-scale ECG datasets (such as PTB-XL and PhysioNet), employing mask reconstruction and cross-patient contrastive learning methods to pre-train the feature extraction network. During the fine-tuning stage, only a very small number of UMI labeled samples are needed to achieve near-fully supervised performance, and it possesses good transfer capabilities across hospitals and populations. This strategy significantly reduces the model's dependence on expensive labeled data and enhances the model's adaptability and universality.

[0160] (3) Multimodal fusion mechanism: This application realizes multimodal fusion of ECG time-series signals, cardiac magnetic resonance imaging (CMR), and clinical structured data. Through a cross-modal attention mechanism, semantic alignment of different modalities is achieved, and the contribution weight of each modality to the task can be dynamically learned, thereby achieving precise weighted fusion. In clinical application scenarios, this method helps to improve the stability of risk discrimination for borderline cases through multimodal collaborative modeling, demonstrating the technical advantages of multimodal fusion analysis.

[0161] (4) Optional explanation and suggestion output: Visual explanation results can be generated based on methods such as attention weight and SHAP, and further intervention suggestions or report outputs can be generated; the explanation method and suggestion generation method can be implemented according to actual needs, and this application does not limit them.

[0162] (5) Lightweight deployment and rapid rollout capability: To meet the needs of different application scenarios, this application designs a model compatible with the lightweight Transformer architecture, which can run on GPU servers or be deployed on local edge devices. The system interface is standardized, enabling rapid connection to existing hospital PACS and HIS systems to achieve low-latency prediction and large-scale batch analysis. This feature ensures that the system has good clinical scalability and industrial application prospects.

[0163] The multimodal asymptomatic myocardial infarction (UMI) prediction system provided in this application, which integrates electrocardiogram (ECG), cardiac magnetic resonance imaging (CMR), and structured clinical indicators, possesses spatiotemporal modeling capabilities and cross-domain transferability. By employing a transferable spatiotemporal mask Transformer model as the core modeling engine, and combining it with a multimodal fusion mechanism, a self-supervised pre-training strategy, and a clinical suggestion generation module, a practical and intelligent UMI early identification system is constructed. This system can achieve accurate early prediction of UMI and possesses interpretability, deployability, and clinical applicability, providing proactive identification and intelligent intervention support for asymptomatic high-risk individuals. Compared with related technologies, this application has the following beneficial effects:

[0164] (1) A Transformer ECG modeling method based on spatiotemporal masking mechanism is proposed to comprehensively improve the ability to identify weak signals. The Transformer ECG modeling method adopts the "time slice + lead spatial embedding" mechanism to map multi-lead ECG signals into a spatiotemporal patch matrix, thereby realizing the joint modeling of spatial coupling and temporal dynamics of electrical signals. Secondly, a generative masking learning mechanism is introduced to enhance the model's sensitivity to slight abnormal features in the ST-T segment by predicting the unmasked part. Finally, by taking advantage of the Transformer architecture's ability to model long-range transconductance connections globally, the limitations of traditional RNN and CNN in performance degradation under long sequences are overcome.

[0165] (2) Introducing self-supervised pre-training and transfer learning mechanisms to solve the problem of difficult UMI small sample modeling: During the model training process, self-supervised mask reconstruction and cross-patient contrastive learning strategies are first used on large-scale ECG datasets such as PTB-XL and PhysioNet to pre-train the feature extraction network, thereby effectively capturing the global and individual differences in ECG signals. On this basis, only a very small number of UMI labeled samples are needed in the fine-tuning stage to obtain good model performance, reducing the excessive reliance on high-cost manual annotation, and enabling transferability across hospitals and populations.

[0166] (3) Constructing a multimodal alignment and fusion mechanism to achieve deep collaborative analysis of structured data, images, and ECG: This application further introduces a cross-modal attention mechanism to semantically align CMR images, ECG time-series signals, and structured clinical indicators, and can dynamically learn the contribution of each modality in the final task, thereby achieving precise weighted fusion. Under this mechanism, the model can not only integrate the complementary advantages of different modal information, but also adaptively adjust the modal weights for specific cases.

[0167] (4) Constructing a complete closed-loop integrated system of "model building-prediction-intervention" with deployability and clinical applicability: This application constructs a practically implementable intelligent clinical auxiliary decision-making system. The system front-end can automatically receive ECG and related examination data from the HIS system and output patient risk scores in the prediction module. At the same time, the interpretation module generates visual results based on attention weight maps and SHAP values, providing doctors with transparent and interpretable model evidence. On this basis, the intervention suggestion generation module combines structured information to call large models (such as DeepSeek and ChatGLM) to generate individualized management and intervention suggestions, further improving the operability of clinical applications. The system also supports automatic report export (such as PDF format), data archiving, and hierarchical user permission management to ensure security, compliance, and ease of use. This system can help improve doctors' work efficiency and has the potential for promotion and application in various scenarios such as communities and outpatient clinics.

[0168] (5) Supports lightweight deployment and large-scale promotion, adapting to various practical application scenarios: The model structure designed in this application is compatible with the lightweight Transformer architecture, which can be flexibly deployed on GPU servers and local edge devices, ensuring adaptability and scalability in multiple scenarios. At the same time, the system interface adopts a standardized design, which can be easily connected to the hospital's existing information systems (such as PACS (Picture Archiving and Communication Systems) and HIS), seamlessly integrating into the clinical workflow. In actual deployment, the system has a short average deployment time and low single-case prediction latency, which can meet the needs of large-scale batch evaluation and support real-time clinical analysis applications, fully demonstrating its engineering implementation value and clinical usability.

[0169] In summary, this application represents a significant technological advancement in model structure, training strategy, fusion method, and system integration. It addresses a series of key issues, including difficulties in early UMI identification, challenges in small-sample modeling, weak fusion analysis, and low clinical applicability, demonstrating promising potential for clinical translation.

[0170] Please see Figure 3 This application also provides an asymptomatic myocardial infarction prediction device 300 based on a transferable spatiotemporal mask Transformer, which can implement the above-mentioned method. The device includes the following modules:

[0171] The multimodal data acquisition module 301 is used to acquire raw multimodal training data; wherein, the raw multimodal training data includes raw electrocardiogram data, raw cardiac magnetic resonance imaging data, and raw clinical structured data;

[0172] The multimodal data preprocessing module 302 is used to preprocess the original multimodal training data to obtain target multimodal training data; wherein, the target multimodal training data includes electrocardiogram feature sequences and small sample datasets;

[0173] The spatiotemporal mask modeling module 303 is used to input the electrocardiogram feature sequence into a transferable spatiotemporal mask Transformer, and perform spatiotemporal mask modeling on the electrocardiogram feature sequence in combination with a generative masking mechanism to generate spatiotemporal mask features.

[0174] The self-supervised pre-training and transfer fine-tuning module 304 is used to perform self-supervised pre-training on the transferable spatiotemporal mask Transformer based on the spatiotemporal mask features, and combine transfer learning to perform transfer fine-tuning on the self-supervised pre-trained transferable spatiotemporal mask Transformer based on the small sample dataset to obtain an electrocardiogram feature extractor.

[0175] The current electrocardiogram feature extraction module 305 is used to acquire the current electrocardiogram data of the object to be predicted, and to extract features from the current electrocardiogram data through the electrocardiogram feature extractor to generate a current electrocardiogram feature representation.

[0176] The adaptive multimodal fusion module 306 is used to fuse the current electrocardiogram feature representation, the current cardiac magnetic resonance imaging feature representation corresponding to the object to be predicted, and the current clinical structured feature representation using a missing modality adaptive multimodal fusion mechanism to generate a multimodal fused feature representation;

[0177] The asymptomatic myocardial infarction risk prediction module 307 is used to input the multimodal fusion feature representation into the asymptomatic myocardial infarction risk classifier to generate asymptomatic myocardial infarction risk prediction results.

[0178] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0179] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0180] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0181] Please see Figure 4 , Figure 4 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0182] The processor 401 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0183] The memory 402 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 402 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 402 and is called and executed by the processor 401 using the methods described in the embodiments of this application.

[0184] Input / output interface 403 is used to implement information input and output;

[0185] The communication interface 404 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0186] Bus 405 transmits information between various components of the device (e.g., processor 401, memory 402, input / output interface 403, and communication interface 404);

[0187] The processor 401, memory 402, input / output interface 403 and communication interface 404 are connected to each other within the device via bus 405.

[0188] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0189] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0190] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0191] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0192] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0193] This application provides a method and related device for predicting asymptomatic myocardial infarction based on a transferable spatiotemporal mask Transformer. By inputting electrocardiogram (ECG) feature sequences into a transferable spatiotemporal mask Transformer and combining a generative masking mechanism to perform spatiotemporal mask modeling on the ECG feature sequences, this method effectively improves the Transformer model's ability to identify features of inconspicuous asymptomatic myocardial infarction, thereby enhancing the accuracy of model prediction. Furthermore, this method leverages the long-distance dependency modeling advantages of the Transformer structure, overcoming the performance degradation limitations of traditional convolutional neural networks or recurrent neural networks when processing long sequences. The transferable spatiotemporal mask Transformer is self-monitored using spatiotemporal mask features. Supervised pre-training was performed, and based on the pre-trained data, a transferable spatiotemporal mask Transformer was used for transfer learning to obtain an ECG feature extractor with good generalization ability. This enabled the ECG feature extractor to more sensitively characterize the fine-grained ECG changes associated with asymptomatic myocardial infarction, thereby enhancing its ability to represent occult myocardial ischemia and subclinical pathological changes, reducing dependence on expensive labeled data, and enhancing the adaptability and universality of the ECG feature extractor. In the actual prediction process, the introduction of a missing modality adaptive multimodal fusion mechanism to achieve semantic alignment and dynamic weighted fusion of different modalities can improve the stability and accuracy of early identification of asymptomatic myocardial infarction, and can provide effective support for clinical screening and auxiliary diagnosis.

[0194] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for predicting asymptomatic myocardial infarction based on a transferable spatiotemporal mask Transformer, characterized in that, The method includes the following steps: Obtain raw multimodal training data; wherein, the raw multimodal training data includes raw electrocardiogram data, raw cardiac magnetic resonance imaging data, and raw clinical structured data; The original multimodal training data is preprocessed to obtain target multimodal training data; wherein, the target multimodal training data includes electrocardiogram feature sequences and small sample datasets; The ECG feature sequence is input into a transferable spatiotemporal mask Transformer, and spatiotemporal mask modeling is performed on the ECG feature sequence using a generative masking mechanism to generate spatiotemporal mask features. The transferable spatiotemporal mask Transformer is self-supervised pre-trained based on the spatiotemporal mask features, and then fine-tuned based on the small sample dataset using transfer learning to obtain an electrocardiogram feature extractor. The current electrocardiogram (ECG) data of the object to be predicted is obtained, and features are extracted from the current ECG data using the ECG feature extractor to generate a current ECG feature representation. A missing modality adaptive multimodal fusion mechanism is adopted to fuse the current electrocardiogram feature representation, the current cardiac magnetic resonance imaging feature representation corresponding to the object to be predicted, and the current clinical structured feature representation to generate a multimodal fused feature representation; The multimodal fusion feature representation is input into the asymptomatic myocardial infarction risk classifier to generate asymptomatic myocardial infarction risk prediction results; The step of inputting the electrocardiogram (ECG) feature sequence into a transferable spatiotemporal mask Transformer, and performing spatiotemporal mask modeling on the ECG feature sequence using a generative masking mechanism to generate spatiotemporal mask features includes: The electrocardiogram feature sequence is input into the transferable spatiotemporal mask Transformer; Through the input layer of the transferable spatiotemporal masking Transformer, a random masking operation is performed on the electrocardiogram feature sequence according to a preset ratio to obtain the masked patch feature sequence and the unmasked patch feature sequence. The unmasked patch feature sequence is input into the encoder in the transferable spatiotemporal mask Transformer to perform global feature modeling and generate the spatiotemporal mask features. The process of performing self-supervised pre-training on the transferable spatiotemporal mask Transformer based on the spatiotemporal mask features, and then fine-tuning the self-supervised pre-trained transferable spatiotemporal mask Transformer using the small sample dataset in conjunction with transfer learning to obtain an electrocardiogram feature extractor, includes: The spatiotemporal mask features and the masking patch feature sequence are input into the transferable spatiotemporal mask Transformer. The masking patch feature sequence is reconstructed based on the spatiotemporal mask features through a generative mask reconstruction mechanism. The transferable spatiotemporal mask Transformer is then jointly optimized and trained using a contrastive learning strategy. By combining transfer learning with the small sample dataset, the transferable spatiotemporal mask Transformer, which has been jointly optimized and trained, is fine-tuned to obtain the electrocardiogram feature extractor.

2. The method according to claim 1, characterized in that, The original electrocardiogram (ECG) data includes first original ECG data and second original ECG data. The first original ECG data, along with the original cardiac magnetic resonance imaging (MRI) data and the original clinical structured data, originates from the same subject and is used to construct the small sample dataset containing asymptomatic myocardial infarction labels. The second original ECG data originates from the target ECG dataset and is used for self-supervised pre-training of the transferable spatiotemporal mask Transformer. The preprocessing of the original multimodal training data to obtain the target multimodal training data includes: The first original electrocardiogram data, the original cardiac magnetic resonance imaging data, and the original clinical structured data are standardized to obtain the first standardized data; wherein, the first standardized data includes the first standardized electrocardiogram data, the standardized cardiac magnetic resonance imaging data, and the standardized clinical structured data. The first standardized data is labeled to obtain the asymptomatic myocardial infarction label; wherein the labeling reference information used for labeling includes at least one of the following: clinical diagnostic records, cardiac magnetic resonance imaging results, and myocardial injury-related examination indicators; Based on the first standardized data and the asymptomatic myocardial infarction label, the small sample dataset is constructed; The second raw electrocardiogram data is standardized to obtain the second standardized electrocardiogram data. The second standardized electrocardiogram data was processed by data fragmentation to obtain several spatiotemporal patches; Perform linear projection processing on each of the spatiotemporal patches to obtain the embedding vector corresponding to each spatiotemporal patch; Each of the embedding vectors is superimposed with the position code and lead code respectively to obtain the electrocardiogram feature sequence.

3. The method according to claim 1, characterized in that, The missing modality adaptive multimodal fusion mechanism fuses the current electrocardiogram feature representation, the current cardiac magnetic resonance imaging feature representation corresponding to the object to be predicted, and the current clinical structured feature representation to generate a multimodal fused feature representation, including: Construct the first modality availability indicator vector corresponding to the current electrocardiogram feature representation; Construct the second modality availability indicator vector corresponding to the current cardiac magnetic resonance imaging feature representation; Construct the third modality availability indicator vector corresponding to the current clinical structured feature representation; A contrastive learning strategy is used to semantically align the current electrocardiogram feature representation, the current cardiac magnetic resonance imaging feature representation, and the current clinical structured feature representation, resulting in current electrocardiogram aligned features, current cardiac magnetic resonance imaging aligned features, and current clinical structured feature alignment features, respectively. The current ECG alignment features, the current cardiac MRI alignment features, and the current clinical structured alignment features are input into an adaptive fusion network. Combined with the first modality availability indicator vector, the second modality availability indicator vector, and the third modality availability indicator vector, the current ECG alignment features, the current cardiac MRI alignment features, and the current clinical structured alignment features are dynamically weighted to generate the multimodal fusion feature representation. When unavailable target modality data exists, the target modality feature corresponding to the target modality data is set as a preset mask vector. During the dynamic weighting of the current ECG alignment features, the current cardiac MRI alignment features, and the current clinical structured alignment features, the weight of the target modality data is constrained to zero or approximately zero based on the preset mask vector.

4. The method according to claim 1, characterized in that, After inputting the multimodal fusion feature representation into the asymptomatic myocardial infarction risk classifier to generate asymptomatic myocardial infarction risk prediction results, the method further includes: Based on the asymptomatic myocardial infarction risk prediction results, the asymptomatic myocardial infarction risk level is determined; Based on the attention weights and SHAP method of the asymptomatic myocardial infarction risk classifier, the multimodal fusion feature representation is post-processed and analyzed to generate a visual interpretation of key leads and key features. Based on the asymptomatic myocardial infarction risk prediction results, the asymptomatic myocardial infarction risk level, and the visualization interpretation results, personalized intervention recommendations are generated. Based on the asymptomatic myocardial infarction risk prediction results, the asymptomatic myocardial infarction risk level, the visualization interpretation results, and the personalized intervention recommendations, an asymptomatic myocardial infarction risk prediction report is generated and displayed on the visualization interface.

5. A device for predicting asymptomatic myocardial infarction based on a transferable spatiotemporal mask Transformer, characterized in that, The device includes the following modules: A multimodal data acquisition module is used to acquire raw multimodal training data; wherein, the raw multimodal training data includes raw electrocardiogram data, raw cardiac magnetic resonance imaging data, and raw clinical structured data; A multimodal data preprocessing module is used to preprocess the original multimodal training data to obtain target multimodal training data; wherein, the target multimodal training data includes electrocardiogram feature sequences and small sample datasets; The spatiotemporal mask modeling module is used to input the electrocardiogram feature sequence into a transferable spatiotemporal mask Transformer, and combine a generative masking mechanism to perform spatiotemporal mask modeling on the electrocardiogram feature sequence to generate spatiotemporal mask features. The self-supervised pre-training and transfer fine-tuning module is used to perform self-supervised pre-training of the transferable spatiotemporal mask Transformer based on the spatiotemporal mask features, and to combine transfer learning to perform transfer fine-tuning of the self-supervised pre-trained transferable spatiotemporal mask Transformer based on the small sample dataset to obtain an electrocardiogram feature extractor. The current electrocardiogram feature extraction module is used to obtain the current electrocardiogram data of the object to be predicted, and to extract features from the current electrocardiogram data through the electrocardiogram feature extractor to generate a current electrocardiogram feature representation; An adaptive multimodal fusion module is used to fuse the current electrocardiogram feature representation, the current cardiac magnetic resonance imaging feature representation corresponding to the object to be predicted, and the current clinical structured feature representation using a missing modality adaptive multimodal fusion mechanism to generate a multimodal fused feature representation. The asymptomatic myocardial infarction risk prediction module is used to input the multimodal fusion feature representation into the asymptomatic myocardial infarction risk classifier to generate asymptomatic myocardial infarction risk prediction results. Specifically, the spatiotemporal mask modeling module is used for: The electrocardiogram feature sequence is input into the transferable spatiotemporal mask Transformer; Through the input layer of the transferable spatiotemporal masking Transformer, a random masking operation is performed on the electrocardiogram feature sequence according to a preset ratio to obtain the masked patch feature sequence and the unmasked patch feature sequence. The unmasked patch feature sequence is input into the encoder in the transferable spatiotemporal mask Transformer to perform global feature modeling and generate the spatiotemporal mask features. The self-supervised pre-training and transfer learning fine-tuning module is specifically used for: The spatiotemporal mask features and the masking patch feature sequence are input into the transferable spatiotemporal mask Transformer. The masking patch feature sequence is reconstructed based on the spatiotemporal mask features through a generative mask reconstruction mechanism. The transferable spatiotemporal mask Transformer is then jointly optimized and trained using a contrastive learning strategy. By combining transfer learning with the small sample dataset, the transferable spatiotemporal mask Transformer, which has been jointly optimized and trained, is fine-tuned to obtain the electrocardiogram feature extractor.

6. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 4.

8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 4.

Citation Information

Patent Citations

  • Aspect level sentiment analysis method and device based on transfer learning

    CN114912423A

  • Cross-domain time sequence completion method integrating pre-training language model

    CN119884600A