This application relates to a supervised pre-training-based multimodal electrocardiogram (ECG)
signal representation learning method. The method includes: entity extraction and standardized mapping of clinical text reports to obtain structured diagnostic labels; extraction of ECG features from raw ECG data through cross-channel
slicing and routing aggregation using a multi-
granularity ECG
encoder; inputting the structured diagnostic labels and ECG features into a multimodal fusion network to complete text semantic extraction and cross-
modal interaction, obtaining fused features; constructing modality consistency loss and classification loss to jointly optimize
model parameters; and inputting the ECG data to be tested into the optimized model to complete diagnostic prediction. This method can fully
exploit the value of clinical text, achieve fine-grained alignment of cross-
modal features, reduce the computational complexity of long sequences, and take into account the multi-scale features of ECG signals, effectively improving the accuracy, robustness, and generalization ability of
ECG analysis models.