The invention discloses a multi-
label electrocardiogram classification method based on self-supervised pre-training and multi-
modal semantic alignment, which belongs to the technical field of
artificial intelligence, and comprises the following steps: realizing self-supervised pre-training of unlabeled data through a single-
modal contrast enhancement network, generating global and local contrast views by adopting a multi-scale random
cutting strategy, and classifying the global and local contrast views in a multi-scale random
cutting mode; in combination with a teacher-student
network architecture, the potential invariance features of the ECG signals are learned while
negative sample dependence is avoided, the problem of
annotation data scarcity is effectively relieved, and the feature robustness is improved. A multi-
modal fusion mechanism based on
label semantic guidance is provided, a
time domain signal and a
frequency domain time-frequency graph are mapped to a unified
semantic space through fine-grained
semantic alignment, local feature enhancement and cross-modal complementary
information fusion are realized by using a cross attention mechanism, and the problem of
semantic difference caused by modal heterogeneity in a traditional method is overcome. A multi-
label comparison
loss function based on a
disease co-occurrence relation is proposed, a category discrimination boundary is dynamically optimized by modeling a label co-
occurrence probability, the feature separability of a
tail category is improved while the head category discrimination ability is enhanced, and the problem of sample category imbalance in a multi-label scene is remarkably relieved.