Multi-label electrocardiogram classification method based on self-supervised pre-training and multi-modal semantic alignment

By employing self-supervised pre-training and multimodal semantic alignment, the limitations of labeled data size and semantic differences between multimodal data in ECG classification were addressed, resulting in better feature extraction and multi-label classification performance and improved ECG diagnostic efficacy.

CN120850033APending Publication Date: 2025-10-28YANSHAN UNIV
View PDF 0 Cites 11 Cited by

Patent Information

Application Number
CN202510949703.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing ECG classification methods face challenges such as limited labeled data size, semantic differences between multimodal data, and imbalanced sample classes, resulting in insufficient model generalization ability and poor feature fusion performance.

Method used

We employ a self-supervised pre-training and multimodal semantic alignment approach, generating a contrastive view through contrastive learning and data augmentation. Combining a teacher-student network architecture, we utilize multi-scale convolutional attention and channel attention networks for feature extraction, and fuse multimodal features through label semantic guidance and cross-attention mechanisms. We also design a multi-label contrastive loss function to optimize category discrimination.

Benefits of technology

It improves the modeling capability of ECG signals, enhances the diagnostic effect of multi-lead electrocardiograms, alleviates the problem of scarce labeled data, strengthens feature robustness, and optimizes multi-label classification performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120850033A_ABST
    Figure CN120850033A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-label electrocardiogram classification method based on self-supervised pre-training and multi-modal semantic alignment, which belongs to the technical field of artificial intelligence, and comprises the following steps: realizing self-supervised pre-training of unlabeled data through a single-modal contrast enhancement network, generating global and local contrast views by adopting a multi-scale random cutting strategy, and classifying the global and local contrast views in a multi-scale random cutting mode; in combination with a teacher-student network architecture, the potential invariance features of the ECG signals are learned while negative sample dependence is avoided, the problem of annotation data scarcity is effectively relieved, and the feature robustness is improved. A multi-modal fusion mechanism based on label semantic guidance is provided, a time domain signal and a frequency domain time-frequency graph are mapped to a unified semantic space through fine-grained semantic alignment, local feature enhancement and cross-modal complementary information fusion are realized by using a cross attention mechanism, and the problem of semantic difference caused by modal heterogeneity in a traditional method is overcome. A multi-label comparison loss function based on a disease co-occurrence relation is proposed, a category discrimination boundary is dynamically optimized by modeling a label co-occurrence probability, the feature separability of a tail category is improved while the head category discrimination ability is enhanced, and the problem of sample category imbalance in a multi-label scene is remarkably relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence technology, specifically relating to a multi-label electrocardiogram classification method based on self-supervised pre-training and multimodal semantic alignment. Background Art

[0002] In recent years, deep learning-based ECG classification methods have made significant progress. Convolutional Neural Networks (CNNs), due to their powerful local feature extraction capabilities, have become the mainstream method for processing ECG signals, effectively capturing spatial information within the signal. For example, ECG classification networks based on deep CNNs have achieved good results on both noisy and noiseless single-lead datasets. Furthermore, combining CNNs with Recurrent Neural Networks (RNNs) for spatiotemporal modeling further improves model performance. All these methods rely on high-quality labeled datasets; for supervised learning, the quality of the dataset often directly impacts the model's performance.

[0003] In the fields of computer vision and natural language processing, self-supervised learning (SSL) methods based on unlabeled datasets have attracted widespread attention and demonstrated great potential. Unlike traditional supervised learning, SSL does not rely on large-scale labeled datasets. Instead, it generates pseudo-labels through pre-designed tasks for training, effectively mitigating the dependence on data labeling, especially in domains where labeled data is scarce or expensive to obtain, such as medical imaging and industrial inspection. Contrastive learning (CL) is the most commonly used method in SSL. Its basic idea is to generate multiple views of the same data sample through data augmentation techniques, and then use a contrastive loss function to optimize the distance between different data sample views, thereby obtaining data invariant features. This process allows the model to be effectively pre-trained on unlabeled data and transfer the learned prior knowledge to downstream tasks, thus improving the model's performance when labeled data is scarce.

[0004] In ECG classification tasks, SSL has begun to demonstrate its application value. For example, the paper "Le D, Truong S, Brijesh P, et al. scl-st: Supervised contrastive learning with semantic transformations for multiple lead ecg arrhythmia classification[J]. IEEE journal of biomedical and health informatics, 2023, 27(6): 2818-2828" discloses ECG classification using the SimCLR framework, proving that the CL method can effectively improve the performance of ECG classification tasks. Another example is the paper "Liu W, Pan S, Li Z, et al. Bootstrap each lead's latent: A novel method for self-supervised learning of multilead electrocardiograms[J]. Computer Methods and Programs in Biomedicine, 2024, 257: 108452", which proposes a multilead ECG classification network based on BYOL, avoiding the use of negative samples and improving the computational efficiency of the model. Despite the promising potential of SSL in the ECG field, current methods still face some challenges, such as how to effectively utilize the latent features in unlabeled data, especially in complex multimodal environments.

[0005] With the continuous development of computer vision technology, image processing algorithms have begun to be applied to the analysis of ECG signals. Traditional ECG signals are usually presented in the form of one-dimensional time-series data, which makes it difficult to extract frequency domain features and capture global and local features in ECG signals. To solve this problem, recent studies have begun to attempt to convert ECG signals into two-dimensional images, thereby utilizing image processing techniques for feature extraction, achieving significant results. For example, the Short Time Fourier Transform (STFT) is used to convert ECG signals into spectrograms, and 2D-CNN is used for feature extraction, achieving good recognition results. Another example is the use of Gram angle field (GAF) and Markov transform field (MTF) image encoding methods for signal conversion, achieving 97% and 85% respectively on dual-lead ECG data. These methods, by converting ECG signals into two-dimensional images, not only better capture frequency domain features but also utilize classic image processing network models, such as VGG and ResNet, for more efficient feature extraction.

[0006] Despite the achievements of image processing-based methods, they still face several challenges. First, the complementarity of time and frequency domain information may be weakened after ECG signals are converted into images, resulting in insufficient capture of temporal features. Second, existing image processing methods often rely on standard deep learning frameworks, lacking semantic alignment of ECG signals across different modalities, leading to poor feature fusion between different modalities. Furthermore, traditional image processing methods are highly dependent on the quality of image conversion; the choice of different conversion methods can significantly impact model performance.

[0007] In summary, although self-supervised learning and image processing methods have shown great potential in ECG classification tasks, they still face challenges such as data dependence, insufficient feature extraction, and poor modality fusion. Therefore, there is a need for a multi-label ECG classification method based on self-supervised pre-training and multimodal semantic alignment. Summary of the Invention

[0008] The purpose of this invention is to provide a multi-label electrocardiogram (ECG) classification method based on self-supervised pre-training and multimodal semantic alignment. This method aims to address the problems of insufficient model generalization ability due to limited labeled dataset size; the impact of semantic differences between multimodal data on feature fusion performance; and the decline in tail category recognition accuracy due to severe imbalance of sample categories. This invention improves the ability to model ECG signals, thereby enhancing the diagnostic effect of multi-lead ECGs.

[0009] To achieve the above objectives, the technical solution adopted by this invention is: a multi-label electrocardiogram classification method based on self-supervised pre-training and multimodal semantic alignment, comprising the following steps:

[0010] Self-supervised pre-training phase based on contrastive learning:

[0011] Step 1: Perform signal processing on the original ECG signal to obtain denoised one-dimensional sequence data, and generate a two-dimensional time-frequency graph using CWT;

[0012] Step 2: Apply different data augmentation strategies to the ECG data of the two modalities to enhance the key features of each modality;

[0013] Step 3: Use two different cropping methods to generate a comparison view on the enhanced modal data to complete the view construction;

[0014] Step 4: Input the constructed view into a single-modal contrast enhancement network to capture modality-invariant features through contrastive learning;

[0015] Semantic-guided multimodal fusion stage:

[0016] Step 5: Perform semantic alignment between modalities. Use a fine-grained label alignment method to reduce the differences in semantic representation between different modalities and improve the collaborative representation between modalities by aligning the feature representations of specific categories.

[0017] Step 6: Perform intramodal semantic enhancement. Through semantic enhancement attention, the label-specific tag-level representation is fused with the global representation of the entire object. The two complement each other to obtain the enhanced label-level representation.

[0018] Step 7: Perform multimodal semantic fusion. Deep fusion of features from different modalities is achieved through a cross-attention mechanism. This process can dynamically generate attention weights based on the features of each modality to capture complementary information between modalities. At the same time, information from different modalities can interact effectively in a unified semantic space to generate more robust multimodal feature representations.

[0019] Multi-label classification stage based on contrastive learning:

[0020] Step 8: The enhanced feature information is fed into a multi-label classifier to obtain multi-label prediction results, and then the network is optimized using the binary cross-entropy loss (BCE) function and the proposed multi-label contrastive loss function.

[0021] A further improvement to the technical solution of the present invention is as follows: Step 1 is specifically performed as follows:

[0022] Step 1.1: Given a raw 12-lead ECG signal Power frequency filters are used to remove power supply noise, followed by the use of high-order Butterworth bandpass filters to remove electromyographic noise and baseline drift. Finally, the denoised signal representation is obtained.

[0023] Step 1.2: Select a mother wavelet function, generate multiple sub-wavelets through scaling and translation operations, and convolve these sub-wavelets with the signal to obtain the time-frequency plot of the signal; for the input signal The mathematical definition of CWT is:

[0024]

[0025] Where ψ(t) is the mother wavelet function, a is the scaling parameter, b is the translation parameter, and ψ * (t) is the complex conjugate of the wavelet; since the Mexican Hat waveform is closer to the ECG signal, it is chosen as the mother wavelet function; the time-frequency plot obtained through CWT is X. img , where X img ∈R 12×h×w , where h is the frequency range and w is the number of sampling points.

[0026] A further improvement to the technical solution of this invention lies in the following: In step 2, for the original one-dimensional ECG sequence data, two data augmentation methods are employed: random pruning and noise addition. The specific steps are as follows:

[0027] Step 2.1: The augmented ECG one-dimensional sequence data is as follows:

[0028]

[0029] Where F crop For random cropping, F gaussian To add Gaussian noise;

[0030] Step 2.2: Frequency domain masking and time domain masking are used to randomly mask the frequency or time domain of the time-frequency plot, so as to enable the model to learn the local features of the signal in different frequency bands or time periods. The time-frequency plot after data augmentation is as follows:

[0031]

[0032] Where F freq For frequency domain masking, F time Obscured by time.

[0033] A further improvement of the technical solution of this invention lies in the fact that the view construction process in step 3 does not rely on negative samples. The specific steps are as follows: by randomly cropping the input image at different scales, a global view and a local view for comparison are generated. The global view refers to an image area greater than 50% of the original image after random cropping, while the local view refers to an image area less than 50% after random cropping. The global view contains the overall features of the data, which helps to learn global context information, while the local view focuses on the local areas of the image to help the model capture detailed features. By comparing the global view and the local view, the model can avoid relying solely on simple feature matching to obtain consistency, thereby promoting the learning of better feature representations.

[0034] A further improvement of the technical solution of the present invention is as follows: In step 4, in the single-modal contrast enhancement network, the single-modal contrast view obtained after data augmentation and view construction will be input into the teacher and student networks constructed using the same network structure for feature extraction. Then, the outputs of the two networks are projected onto the same feature space for comparative learning. The network parameters for feature extraction are optimized by minimizing the difference between the two. Finally, the trained student network parameters are saved for subsequent supervised contrast learning.

[0035] A further improvement to the technical solution of this invention lies in the following: the specific composition and setup process of each network used in step 4 are as follows:

[0036] Feature extraction network for one-dimensional time series modalities: For one-dimensional time series modalities, ResNet1d50 is selected as the backbone network. A multi-scale convolutional attention mechanism is introduced into the backbone network to perform multi-scale analysis of the signal. For multi-lead ECG signals, each lead represents the electrical activity of different parts of the heart. Multi-lead analysis can integrate information from different perspectives, thereby revealing the complex characteristics of heart disease. In the multi-scale convolutional attention network, the feature map is uniformly divided into four feature map subsets after 1×1 convolution. These subsets have the same size as the original feature map, but the number of channels changes. Subsequently, multi-branch hierarchical convolution with a residual structure is used to obtain feature map combinations with different numbers and field sizes, thereby extracting multi-scale features of the one-dimensional signal. In the channel attention network, multiple leads of the ECG signal are regarded as channels, and the relationship between multi-lead signals is weighted by the SE network.

[0037] Feature extraction network for 2D time-frequency map modality: For 2D time-frequency map modality, ResNet50 is selected as the backbone network. Combining the characteristics of image data, multi-scale convolutional attention and channel-spatial attention are modified to capture the time-frequency features of ECG signals in a single lead and the correlation information between multiple leads. In channel-spatial attention, firstly, adaptive max pooling and adaptive average pooling are used to extract global and local features of the image along the spatial dimension of the time-frequency map. Then, these two parts are added by channel squeezing excitation operation. The feature map containing key lead information is then spliced ​​after max pooling and average pooling along the channel dimension. Then, the focal information in the image space is obtained by convolutional attention. In multi-scale convolutional attention, firstly, multiple convolutional kernels of 1×1, 3×3 and 5×5 are used to extract receptive field information of different sizes. Then, the features of different receptive fields are cross-mixed by channel shuffling to break the fixed branch structure and enhance the interaction between multi-scale features.

[0038] The entire single-modal contrast enhancement network: For one-dimensional sequence data, its contrast view can be represented as:

[0039]

[0040] in and These are collections of global views and local views, respectively. It contains two global views. Contains 8 partial views, F G_crop and F L_crop These are random cropping operations for the global view and the local view, respectively.

[0041] Both student and teacher networks employ one-dimensional time-series modality feature extraction networks. First, the student network parameters are initialized, and then copied to the teacher network to complete its initialization. The teacher network then extracts features from each global view, while the student network extracts features from all views, including both global and local views. The view representations learned by the teacher and student networks are mapped to a high-dimensional feature space via a projection network for comparative learning. The projected features undergo L2 normalization to ensure that comparisons between different features are unaffected by scale differences, while also maintaining the stability of the student and teacher network outputs during training.

[0042]

[0043] Among them Teacher 1D () represents the teacher network, Student 1D () represents the student network, and Project() represents the projection network; finally, comparative learning is performed using the projection features of the teacher and student networks, this process iterates through each teacher output. and each student's output Calculate the cross-entropy loss between them;

[0044] During training, the parameters of the student network are updated using standard backpropagation and gradient descent, while the parameters of the teacher network are updated using exponential moving average (EMA). The EMA method ensures smooth updates for the teacher network, avoiding the instability of the learning objectives caused by drastic changes in the student network and effectively preventing local overfitting during the student network's learning process. The parameter update formula for the teacher network is as follows:

[0045]

[0046] Where k represents the current training step number, λ k Represents the parameter smoothing coefficient. This represents the current teacher network parameters. This indicates the current student network parameters. This indicates the updated teacher network parameters.

[0047] A further improvement to the technical solution of this invention lies in the following steps in step 5: First, the feature embedding of the corresponding modality is obtained through a pre-trained single-modality coding network:

[0048] F sig =Encoder 1D (X sig ), F img=Encoder 2D (X img )

[0049] Encoder 1D () and Encoder 2D () represent pre-trained single-modal coding networks, F sig and F img These are feature maps for different modalities.

[0050] For tag semantics, given a tag set L = (l1, l2, ..., l m ), where m is the number of labels, and the semantic embeddings of the label set are obtained using the pre-trained natural language model BERT:

[0051] Y = (y1, y2, ..., y m =BERT(L)

[0052] Where y j This represents the semantic embedding of a single tag.

[0053] F sig and F img The feature maps of the two modalities are respectively obtained through label-specific attention under the guidance of label semantics. First, the correlation matrix between each ECG feature and the semantic embedding of a single label in the corresponding modality is calculated:

[0054]

[0055] Where P, U, V, and b are all learnable parameters, and ⊙ represents element-wise multiplication. It is a learnable feedforward neural network that uses a low-rank bilinear computation model to reduce computational complexity.

[0056] Then, the label-level representation of the corresponding modality is obtained using the correlation matrix:

[0057]

[0058] This method achieves implicit semantic alignment of the two modalities at the label fine-grained level through label-specific attention.

[0059] A further improvement to the technical solution of the present invention is as follows: Step 6 is specifically as follows: by using semantically enhanced attention, the local information in the label-level representation and the global representation of the entire object are fused together, which not only preserves the fine information of the local features, but also introduces the contextual relationship of the global information;

[0060] First, let's define the tag-level representation as h. sig,j and global representation f sig,i ∈Fsig Obtain the query Q, key K, and value V through linear transformation:

[0061] [Q,K,V]=[ω q h sig,j ,ω k f sig,i ,ω v f sig,i ]

[0062] Then, the enhanced representation under a single label is obtained through attention calculation:

[0063]

[0064] Where u represents the learnable attentional bias.

[0065] A further improvement to the technical solution of this invention lies in: Step 7 specifically involves the following steps: representing the enhanced tag level... Compared with the original feature map f sig,i By splicing, we get c sig,i ;

[0066] Cross-attention mechanisms can achieve semantic association and feature fusion between different modalities. For the feature representation C of two modalities of ECG signals... sig =(c sig,1 ,c sig,2 ,...,c sig,n ) and C img =(c img,1 ,c img,2 ,...,c img,n First, the one-dimensional time series mode C sig As a query Q t Two-dimensional image modality C img As key K v Sum V v :

[0067] [Q t ,K v V v ]=[ω q h sig,j ,ω k f sig,i ,ω v f sig,i ]

[0068] Then, the fused two-dimensional image modality is obtained through cross-attention:

[0069]

[0070] Similarly, the fused one-dimensional time series mode representation is as follows:

[0071] Finally, the results obtained through cross-attention and The fusion vectors of the two modalities are concatenated to obtain the multimodal joint representation C.

[0072] The further improvement of the technical solution of the present invention is as follows: Step 8 is as follows: Combining the label co-occurrence relationship and supervised contrastive learning in the multi-label problem, a multi-label contrastive loss function is proposed, which aims to bring the distance between the representation vectors of categories with high label co-occurrence probability closer and expand the distance between categories with low label co-occurrence probability or no co-occurrence relationship, so that the model can better utilize the label relationship to optimize the distance between categories in the discrimination process and alleviate the problem of sample class imbalance.

[0073] The feature vector C obtained by multimodal fusion is projected onto the label category through a linear classifier, and then the final predicted label is obtained through the sigmoid activation function; for each sample Its label set is L = (l1, l2, ..., l m If BCE is calculated as follows:

[0074]

[0075] Where y ik For the sample The true value of the k-th label, The corresponding predicted value;

[0076] First, a tag co-occurrence matrix A is constructed using the tag co-occurrence frequency, where elements A... ij This represents the co-occurrence frequency of label i and label j in the training set; then, based on A, the label relationship weight matrix W is obtained:

[0077]

[0078] Where N is the total number of samples in the current mini-batch, g(i) is all samples in the current mini-batch except i, τ is the temperature coefficient, and W ij Let W be the weight of the label relationship between samples i and j. For a sample pair (i,j), the greater the label relevance, the higher the weight W. ij The larger the value, the larger the corresponding loss term, thus allowing the distance between them to be optimized closer.

[0079] The overall training loss is a weighted sum of the BCE and the multi-label contrastive loss:

[0080] L = L BCE +λL con

[0081] Where λ is a weighting parameter that controls the proportions of the two losses.

[0082] The technological advancements achieved by this invention, due to the adoption of the aforementioned technical solutions, are as follows: It proposes applying self-supervised pre-training on unlabeled datasets to enhance single-modal feature extraction capabilities. Through the generation of contrastive views from unlabeled data and a single-modal contrast enhancement network, it utilizes data augmentation strategies and a teacher-student architecture to capture the modality invariance features of ECG signals. This method not only alleviates the bottleneck of limited labeled data size but also enhances robustness against noise interference by fusing inter-lead correlations through multi-scale convolutional attention and channel attention networks. Compared to existing methods that rely on supervised learning for feature extraction, this invention achieves superior feature learning while reducing label dependency.

[0083] This paper proposes a label-guided multimodal semantic alignment and fusion method. It obtains label semantic embeddings through a pre-trained text model, utilizes low-rank bilinear computation to achieve fine-grained cross-modal semantic alignment, and dynamically fuses multimodal features from time-series signals and time-frequency maps using a cross-attention mechanism. Compared to existing multimodal methods, this invention addresses the semantic discrepancies caused by modal heterogeneity in traditional multimodal methods by complementing global context and local features through label semantic constraints. This method improves the classification performance of diseases such as AF and IAVB, which depend on frequency variations, and verifies the effectiveness of intermodal semantic alignment in capturing key pathological features.

[0084] We design a multi-label contrastive loss function based on label co-occurrence relationships. By constructing a label co-occurrence matrix, we dynamically adjust the inter-class distance weights and integrate prior knowledge of label associations in supervised contrastive learning. Compared to the traditional binary cross-entropy loss, this loss function clusters high-probability co-occurrence class representations in the embedding space while optimizing the discrimination boundary of tail classes. It solves the head-class bias problem caused by imbalanced samples in existing methods, achieving more balanced multi-label classification performance. Attached Figure Description

[0085] Figure 1 This is a schematic diagram of the overall classification method;

[0086] Figure 2 This is a flowchart of the self-supervised pre-training phase;

[0087] Figure 3 This is a schematic diagram of a feature extraction network structure for a one-dimensional time series mode.

[0088] Figure 4 This is a schematic diagram of the feature extraction network structure for two-dimensional time-frequency graph modes;

[0089] Figure 5 This is a flowchart of the supervised learning phase;

[0090] Figure 6This is a schematic diagram of the confusion matrix for eight types of cardiac arrhythmia on the CPSC2018 dataset;

[0091] Figure 7 This is a schematic diagram of the overall confusion matrix on the CPSC2018 dataset;

[0092] Figure 8 This is a loss graph for the two modalities before and after SSL pre-training. Detailed Implementation

[0093] The present invention will be further described in detail below with reference to embodiments:

[0094] like Figure 1 As shown, this invention addresses the problems faced by existing deep learning-based electrocardiogram (ECG) classification methods in practical applications, such as limited labeled data scale, semantic differences between multimodal data, and sample class imbalance. It proposes a multi-label ECG classification algorithm based on self-supervised pre-training and multimodal semantic alignment.

[0095] First, this invention achieves self-supervised pre-training of unlabeled data through a single-modal contrast enhancement network. A multi-scale pruning strategy is employed to generate global and local contrast views. Combined with a teacher-student network architecture, it learns the latent invariant features of ECG signals while avoiding negative sample dependence, effectively alleviating the problem of scarce labeled data and improving feature robustness. Then, a label semantic-guided multimodal fusion mechanism is designed. Fine-grained semantic alignment maps time-domain signals and frequency-domain time-frequency maps to a unified semantic space. A cross-attention module is introduced to achieve local feature enhancement and cross-modal complementary information fusion, overcoming the semantic differences caused by modal heterogeneity in traditional methods. Finally, a multi-label contrast loss function based on disease co-occurrence relationships is proposed. By modeling label co-occurrence probabilities, the class discrimination boundary is dynamically optimized, strengthening the head class discrimination capability while improving the feature separability of the tail class, significantly alleviating the sample class imbalance problem in multi-label scenarios.

[0096] like Figure 2 The diagram shows the self-supervised pre-training phase based on contrastive learning, including the following steps:

[0097] Step 1: Perform signal processing on the original ECG signal to obtain denoised one-dimensional sequence data, and generate a two-dimensional time-frequency graph using CWT;

[0098] Given a raw 12-lead ECG signal For single-lead signals Assuming a sampling rate of 100Hz and a duration of 10s, then Raw ECG signals are typically affected by various noises and interferences, including power supply noise, electromyography (EMG) noise, and baseline drift. These factors can degrade signal quality, thus requiring signal denoising. First, a power frequency filter is used to remove power supply noise. Then, a high-order Butterworth bandpass filter is used to remove EMG noise and baseline drift. Finally, the denoised signal representation is obtained.

[0099] Convolutional Wavelet Map (CWT) is a classic time-frequency analysis method. By convolving a signal with a wavelet function, it decomposes the signal into a series of local frequency components, thereby revealing the local features of the signal at different time scales or frequency ranges. In CWT image conversion, a suitable mother wavelet function is first selected, and multiple sub-wavelets are generated through scaling and translation operations. These sub-wavelets are then convolved with the signal to obtain the final time-frequency map of the signal. For the input signal... The mathematical definition of CWT is:

[0100]

[0101] Where ψ(t) is the mother wavelet function, a is the scaling parameter, b is the translation parameter, and ψ * (t) is the complex conjugate of the wavelet. Common mother wavelet functions for CWT include Morlet, Mexican Hat, Haar, and Gaus. Mexican Hat is chosen as the mother wavelet function because its waveform more closely resembles the ECG signal. The time-frequency plot obtained through CWT is X. img , where X img ∈R 12 ×h×w , where h is the frequency range and w is the number of sampling points.

[0102] Step 2: Apply different data augmentation strategies to the ECG data of the two modalities to enhance the key features of each modality;

[0103] For the two-dimensional time-frequency plot, considering the polarity variations of the ECG signal across different leads, instead of the common rotation or flipping operations, frequency domain masking and time domain masking were employed. Random masking of the frequency or time domain of the time-frequency plot was used to encourage the model to learn the local features of the signal in different frequency bands or time periods. The data-augmented time-frequency plot is shown below:

[0104]

[0105] Where F freq For frequency domain masking, F time Obscured by time.

[0106] Step 3: Use two different cropping methods to generate a comparison view on the enhanced modal data to complete the view construction;

[0107] In constructing the comparison views, this invention does not rely on negative samples. Instead, it generates global and local views for comparison by randomly cropping the input image at different scales. The global view refers to an image whose area after random cropping is greater than 50% of the original image, while the local view refers to an image whose cropping ratio is less than 50%. The global view contains the overall features of the data, which helps in learning global contextual information, while the local view focuses on local regions of the image to help the model capture detailed features. By comparing the global and local views, the model avoids relying solely on simple feature matching to achieve consistency, thereby promoting the learning of better feature representations.

[0108] Step 4: Input the constructed view into a single-modal contrast enhancement network to capture modality-invariant features through contrastive learning;

[0109] During the pre-training phase, given the differences in data structure and feature space between the two modalities of ECG signals, the two modalities were trained separately. The advantage of this strategy is that it allows each modal model to focus on learning its own modal features, avoiding the complexity of cross-modal fusion interfering with single-modal learning performance. Furthermore, it effectively reduces overfitting issues caused by excessive reliance on a single modality during multimodal joint training.

[0110] In the unimodal contrast enhancement network, the unimodal contrast view obtained after data augmentation and view construction is input into teacher and student networks built with the same network structure for feature extraction. Then, the outputs of the two networks are projected onto the same feature space for contrastive learning. The network parameters for feature extraction are optimized by minimizing the difference between the two. Finally, the trained student network parameters are saved for subsequent supervised contrastive learning.

[0111] For one-dimensional time series modalities, ResNet1d50 was selected as the backbone network. Considering the differences in heart rate among individuals and the different focuses on long-term trend changes and short-term fluctuations (such as instantaneous changes in QRS complexes) for different disease types, a multi-scale convolutional attention mechanism was introduced into the backbone network to perform multi-scale signal analysis.

[0112] For multi-lead ECG signals, each lead represents the electrical activity of different parts of the heart. Multi-lead analysis can synthesize information from different perspectives, thereby revealing the complex characteristics of heart disease. Furthermore, previous studies have shown that the interrelationships between leads are crucial in ECG signal analysis. Therefore, we added a channel attention mechanism to the backbone network to effectively model the correlations between multi-leads.

[0113] like Figure 3 As shown, this network adds multi-scale convolutional attention and channel attention to the backbone network. In the multi-scale convolutional attention network, the feature map is uniformly divided into four feature map subsets after a 1×1 convolution. These subsets have the same size as the original feature map, but the number of channels changes. Subsequently, multi-branch hierarchical convolution with a residual structure is used to obtain feature map combinations with different numbers and field sizes, thereby extracting multi-scale features of the one-dimensional signal. In the channel attention network, multiple leads of the ECG signal are treated as channels, and the relationships between the signals of multiple leads are weighted by the SE network.

[0114] For the two-dimensional time-frequency map modality, ResNet50 was selected as the backbone network. Based on the characteristics of the image data, multi-scale convolutional attention and channel-spatial attention were modified to capture the time-frequency features of ECG signals within a single lead and the correlation information between multiple leads. In channel-spatial attention, adaptive max pooling and adaptive average pooling are first used along the spatial dimension of the time-frequency map to extract global features (e.g., overall waveform disorder across multiple leads) and local features (e.g., QRS complex anomalies in certain leads). These two parts are then added together through channel compression excitation. Next, feature maps containing key lead information are concatenated after max pooling and average pooling along the channel dimension, and then focal information in the image space is obtained through convolutional attention. In multi-scale convolutional attention, multiple convolutional kernels of 1×1, 3×3, and 5×5 are first used to extract receptive field information of different sizes. Then, channel shuffling is used to cross-mix features from different receptive fields, breaking the fixed branch structure and enhancing the interaction between multi-scale features. The specific network structure is as follows: Figure 4 As shown.

[0115] Here, taking a one-dimensional time-series mode of an ECG signal as an example, the entire single-mode contrast enhancement network is described in detail. For one-dimensional sequence data, its contrast view can be represented as:

[0116]

[0117] in and These are collections of global views and local views, respectively. It contains two global views. Contains 8 partial views, F G_crop and F L_crop These are random cropping operations for the global view and the local view, respectively.

[0118] Both student and teacher networks use the following methods: Figure 3The diagram illustrates a feature extraction network for a one-dimensional time-series modality. First, the student network's parameters are initialized, and then copied to the teacher network to initialize it. Next, the teacher network extracts features from each global view, while the student network extracts features from all views, including both global and local views. The view representations learned by the teacher and student networks are mapped to a high-dimensional feature space via a projection network for comparative learning. The projected features undergo L2 normalization to ensure that comparisons between different features are unaffected by scale differences, while also maintaining the stability of the student and teacher network outputs during training.

[0119]

[0120] Among them Teacher 1D () represents the teacher network, Student 1D () represents the student network, and Project() represents the projection network. Finally, comparative learning is performed using the projection features of the teacher and student networks. This process iterates through each teacher output. and each student's output Calculate the cross-entropy loss between them.

[0121] During training, the parameters of the student network are updated using standard backpropagation and gradient descent, while the parameters of the teacher network are updated using Exponential Moving Average (EMA). The teacher network's EMA update ensures smooth changes, avoiding the unstable learning objectives caused by drastic changes in the student network and effectively preventing local overfitting during the student network's learning process. The parameter update formula for the teacher network is as follows:

[0122]

[0123] Where k represents the current training step number, λ k Represents the parameter smoothing coefficient. This represents the current teacher network parameters. This indicates the current student network parameters. This indicates the updated teacher network parameters.

[0124] like Figure 5 As shown, the semantically guided multimodal fusion stage:

[0125] Step 5: Perform semantic alignment between modalities. Use a fine-grained label alignment method to reduce the differences in semantic representation between different modalities and improve the collaborative representation between modalities by aligning the feature representations of specific categories.

[0126] First, feature embeddings for the corresponding modality are obtained through a pre-trained single-modality coding network:

[0127] F sig =Encoder 1D (X sig ), F img =Encoder 2D (X img )

[0128] Encoder 1D () and Encoder 2D () represent pre-trained single-modal coding networks, F sig and F img These are feature maps for different modalities.

[0129] For tag semantics, given a tag set L = (l1, l2, ..., l m ), where m is the number of labels, and the semantic embeddings of the label set are obtained using the pre-trained natural language model BERT:

[0130] Y = (y1, y2, ..., y m =BERT(L)

[0131] Where y j This represents the semantic embedding of a single tag.

[0132] F sig and F img The feature maps of the two modalities, guided by label semantics, obtain their respective label-level representations through label-specific attention. First, the correlation matrix between each ECG feature and the semantic embedding of a single label in the corresponding modality is calculated:

[0133]

[0134] Where P, U, V, and b are all learnable parameters, and ⊙ represents element-wise multiplication. It is a learnable feedforward neural network that uses a low-rank bilinear computation model to reduce computational complexity.

[0135] Then, the label-level representation of the corresponding modality is obtained using the correlation matrix:

[0136]

[0137] This method achieves implicit semantic alignment of the two modalities at the label fine-grained level through label-specific attention.

[0138] Step 6: Perform intramodal semantic enhancement. Through semantic enhancement attention, the label-specific label-level representation is fused with the global representation of the entire object. The two complement each other to obtain the enhanced label-level representation. By using semantic enhancement attention to fuse the local information in the label-level representation with the global representation of the entire object, the fine information of the local features is preserved, while the contextual relationship of the global information is introduced.

[0139] Taking a one-dimensional time series mode as an example, the label-level representation h is first... sig,j and global representation f sig,i ∈F sig Obtain the query Q, key K, and value V through linear transformation:

[0140] [Q,K,V]=[ω q h sig,j ,ω k f sig,i ,ω v f sig,i ]

[0141] Then, the enhanced representation under a single label is obtained through attention calculation:

[0142]

[0143] Where u represents the learnable attentional bias.

[0144] Step 7: Perform multimodal semantic fusion. Deep fusion of features from different modalities is achieved through a cross-attention mechanism. This process can dynamically generate attention weights based on the features of each modality to capture complementary information between modalities. At the same time, information from different modalities can interact effectively in a unified semantic space to generate more robust multimodal feature representations.

[0145] Enhanced label-level representation Compared with the original feature map f sig,i By splicing, we get c sig,i .

[0146] Cross-attention mechanisms can achieve semantic association and feature fusion between different modalities. For the feature representation C of two modalities of ECG signals... sig =(c sig,1 ,c sig,2 ,...,c sig,n ) and C img =(c img,1, c img,2 ,...,c img,n First, the one-dimensional time series mode C sig As a query Q t Two-dimensional image modality C img As key Kv Sum V v :

[0147] [Q t ,K v V v ]=[ω q h sig,j ,ω k f sig,i ,ω v f sig,i ]

[0148] Then, the fused two-dimensional image modality is obtained through cross-attention:

[0149]

[0150] Similarly, the fused one-dimensional time series mode representation is as follows:

[0151] Finally, the results obtained through cross-attention and The fusion vectors of the two modalities are concatenated to obtain the multimodal joint representation C.

[0152] Multi-label classification stage based on contrastive learning:

[0153] Step 8: The enhanced multimodal features are fed into a multi-label classifier to obtain multi-label prediction results. Then, the network is optimized using the binary cross-entropy loss (BCE) function and the proposed multi-label contrastive loss function. While BCE is widely used in multi-label classification tasks, it treats each label as an independent binary classification task, thus failing to capture the correlation between labels. Therefore, when facing imbalanced samples, it often leads to biased learning of certain head labels, affecting the discrimination of the minority class.

[0154] To address this, this invention combines label co-occurrence relationships and supervised contrastive learning in multi-label problems to propose a multi-label contrastive loss function. This function aims to shorten the distance between representation vectors of categories with high label co-occurrence probabilities and widen the distance between categories with low label co-occurrence probabilities or no co-occurrence relationship. This allows the model to better utilize label relationships to optimize the distance between categories during the discrimination process, thereby alleviating the problem of class imbalance.

[0155] The feature vector C obtained from multimodal fusion is projected onto the label category through a linear classifier, and then the final predicted label is obtained through a sigmoid activation function. For each sample... Its label set is L = (l1, l2, ..., l m If BCE is calculated as follows:

[0156]

[0157] Where y ik For the sample The true value of the k-th label, This is the corresponding predicted value.

[0158] First, a tag co-occurrence matrix A is constructed using the tag co-occurrence frequency, where elements A... ij This represents the co-occurrence frequency of labels i and j in the training set. Then, based on A, the label relationship weight matrix W is obtained:

[0159]

[0160] Where N is the total number of samples in the current mini-batch, g(i) is all samples in the current mini-batch except i, τ is the temperature coefficient, and W ij Let W be the weight of the label relationship between samples i and j. For a sample pair (i,j), the greater the label relevance, the higher the weight W. ij The larger the value, the larger the corresponding loss term, thus allowing the distance between them to be optimized closer.

[0161] The overall training loss is a weighted sum of the BCE and the multi-label contrastive loss:

[0162] L = L BCE +λL con

[0163] Where λ is a weighting parameter that controls the proportions of the two losses.

[0164] The above method was verified using experiments.

[0165] The experimental environment for this invention was a server running the Linux distribution Ubuntu 23 operating system, with an Intel Xeon Gold 6133 processor, an NVIDIA RTX 4090 24G graphics card, and 128GB of RAM. The experimental framework used the PyTorch machine learning library.

[0166] To evaluate the effectiveness of the proposed algorithm, this invention selected the publicly available CPSC2018 dataset for experiments. This dataset, derived from the 2018 China Physiological Signal Challenge, was collected from 11 hospitals and contains 6877 12-lead electrocardiogram (ECG) records, including 3699 male samples and 3178 female samples. The sampling rate was 500Hz, and the recording duration ranged from 6 to 60 seconds. The class statistics of the dataset are shown in Table 1. From the class distribution, the dataset exhibits class imbalance, and co-occurrence relationships exist between labels. During the experiment, the 6877 samples in the CPSC2018 dataset were divided into 10 equal parts, and 10-fold cross-validation was used to evaluate the algorithm's performance.

[0167] For multi-label classification tasks, we typically focus on three metrics: precision, recall, and F1 score. These metrics are defined as follows:

[0168]

[0169] TP, FP, and FN represent the number of true positives, false positives, and false negatives, respectively. F1 is the harmonic mean of precision and recall. A higher F1 indicates better generalization ability of the algorithm.

[0170] Table 1. Category statistics of the experimental dataset

[0171]

[0172]

[0173] The present invention has been experimentally tested and the classification effect on the publicly available 12-lead electrocardiogram dataset CPSC2018 is ideal, consistent with the design expectations.

[0174] For the CPSC2018 dataset, this invention selected seven existing algorithms as benchmarks and compared them with the proposed algorithm to verify its effectiveness. The compared algorithms cover the latest research results in different directions, including multi-scale feature extraction, spatiotemporal information fusion, and multimodal fusion, specifically as follows: Multi-scale CNN uses a multi-scale fusion CNN for feature extraction and designs a multi-scale loss joint optimization function to gradually achieve complementary learning of multi-scale features; ATI-CNN has the ability to classify ECG signals of different lengths, and combines deep CNN, LSTM, and attention modules to achieve spatiotemporal information fusion; MLC-CNN utilizes the correlation of ECG abnormality labels, and further enhances the feature representation effect by hierarchically fusing features and combining a multi-scale receptive field module; 1D-CNN proposes an ECG classification algorithm based on 1D-CNN to verify the model's classification effect in each single lead and multi-lead, and uses the SHAP method for interpretability analysis; MMnet supplements the original one-dimensional ECG by introducing a time-frequency plot, which only contains time-domain information, and combines a multi-scale feature extraction network to achieve accurate representation of the dynamic changes of ECG signals. By comparing and analyzing with these advanced methods, we can more comprehensively evaluate the performance of the proposed algorithm on the ECG signal classification task and further verify its advantages in feature extraction, information fusion and other aspects.

[0175] Table 2 shows the overall performance comparison of the SS-MSA algorithm and the comparison algorithms on the CPSC2018 dataset. The experimental results show that the SS-MSA algorithm achieved the best performance with 85.1% F1, 88.6% Precision and 82.9% Recall, which is better than other comparison algorithms. From the comparative analysis, different methods show their own characteristics in terms of performance: (1) Traditional CNN methods (1D-CNN and ATI-CNN) mainly rely on CNN for feature extraction, but their F1 scores are 81.3% and 81.2% respectively, indicating that their feature extraction ability in complex ECG signal classification tasks is limited; (2) Among the methods that integrate multi-scale information, MLC-CNN combines the correlation between multi-scale receptive fields and anomaly labels, and its F1 score reaches 82.7%, but its Recall (81.6%) is slightly inferior to other methods. MMnet supplements the time domain information by using time-frequency plots, which improves the model's ability to characterize the dynamic changes of ECG signals. Its Precision reaches 84.9%, which is better than Multi-scale CNN.

[0176] In comparison, SS-MSA outperformed all benchmark models in F1, Precision, and Recall, especially in Precision, where it improved by 3.7% compared to MMnet (88.6% vs. 84.9%), significantly reducing the false positive rate. This improved Precision makes it more valuable in false-positive sensitive applications such as outpatient screening. High Precision means fewer misdiagnoses, reducing the risk of patients undergoing unnecessary examinations or treatments, which meets clinical needs in medical diagnosis.

[0177] Table 2 Overall performance comparison on the CPSC2018 dataset

[0178]

[0179] The F1 score is crucial in medical diagnosis and multi-label tasks because it integrates precision and recall, ensuring robust detection capabilities in positive categories (especially rare disease categories). Therefore, to further evaluate the performance of the proposed algorithm, it is necessary to analyze the performance of SS-MSA on the F1 score and compare it with comparative algorithms.

[0180] Table 3 shows the F1 score comparison results for each category in the CPSC2018 dataset. Experimental results show that, compared with all comparison methods, the proposed algorithm achieves the best performance in the AF, IAVB, LBBB, STD, and STE categories, improving the F1 score by 0.7%, 0.5%, 3.4%, 2.1%, and 3.4% respectively compared to the best result among the comparison algorithms. For categories with a large number of samples, such as RBBB, AF, and STD, the SS-MSA algorithm performs well in classification. Furthermore, the SS-MSA algorithm also achieves the best F1 score in the tail categories of LBBB and STE, which have the fewest samples, demonstrating the effectiveness of the multimodal fusion module utilizing label semantics and multi-label contrastive learning based on label co-occurrence relationships. By leveraging label semantics and co-occurrence relationships, the SS-MSA algorithm enhances the feature representation ability for a minority of categories while also optimizing the discriminant distance between related categories. Compared with the MMnet algorithm, which also uses multimodal information, the SS-MSA algorithm outperforms the MMnet algorithm in all categories except PAC and PVC. Specifically, it improves performance by 4.0% and 0.5% in AF and IAVB, two disease categories that focus more on frequency changes, respectively, which to some extent demonstrates the effectiveness of the proposed label-based multimodal fusion method.

[0181] Table 3. Comparison of multi-class F1 scores on the CPSC2018 dataset.

[0182]

[0183]

[0184] Figure 6 The proposed SS-MSA algorithm is demonstrated to classify eight types of cardiac arrhythmias. It can be seen that the algorithm achieves the highest true negative rate in the AF, LBBB, and RBBB categories, while its true negative rate is the lowest in the STE category. Combined with... Figure 7 The overall confusion matrix shown reveals that the STE category is easily misclassified as the Normal category, primarily due to the smaller sample size for the STE category, leading to insufficient training of the algorithm on this category. Combined with... Figure 6 and Figure 7 Analysis of the confusion matrix results shows that the proposed algorithm achieved good results in the arrhythmia classification task, further verifying the effectiveness of the proposed method.

[0185] Figure 8 The loss curves for the two modalities before and after SSL pre-training are shown on the CPSC2018 dataset. Figure 8 Subfigures (a) and (b) show the changes in training and validation set losses as the training process progresses for the one-dimensional time series mode and the two-dimensional time-frequency plot mode, respectively, without SSL pre-training. It can be observed that the one-dimensional time series mode stabilizes earlier than the two-dimensional time-frequency plot mode. This may lead to the algorithm being easily influenced by a single mode during mode fusion, failing to fully exploit the complementary features between modes. Figure 8 Subgraphs (c) and (d) illustrate the loss changes of the two modalities after SSL pre-training. The results show that after pre-training, the number of training epochs at which the two modalities stabilize is relatively close, ensuring that the model can fully integrate multimodal features during training, thus avoiding insufficient feature learning caused by premature convergence of a single modality. This phenomenon indicates that SSL pre-training can effectively perform preliminary feature extraction for both modalities, alleviate the imbalance problem between modalities, and enable the model to learn complementary information more efficiently during the multimodal semantic fusion stage.

[0186] To verify the impact of the proposed components on the performance of the SS-MSA algorithm, ablation experiments were conducted on the CPSC2018 dataset under the same experimental environment. Specifically, the effects of self-supervised pre-training (SSL), inter-modal semantic alignment (IMSA), intra-modal semantic enhancement (IMSE), and multi-label contrastive loss (MCL) on the overall model performance were evaluated. The effectiveness of each module was verified by removing different modules and observing the performance changes.

[0187] The ablation experiments of the proposed algorithm on the CPSC2018 dataset are shown in Table 4, where "T" represents using only the one-dimensional time series mode, "G" represents using only the two-dimensional time-frequency plot mode, and "T+G" represents multi-modal fusion. Experimental results show that, regardless of whether SSL is applied, the one-dimensional time series mode (T) consistently outperforms the two-dimensional time-frequency plot mode (G). However, multi-modal fusion (T+G) consistently outperforms the single modality, indicating that although the original time series mode already possesses good classification capabilities, the time-frequency plot mode can still provide effective supplementation, thereby enhancing the model's feature extraction capabilities.

[0188] Table 4 Ablation experiments on the CPSC2018 dataset

[0189]

[0190]

[0191] Further analysis of the impact of SSL on different modalities is presented. In the single-modal scenario, SSL improves the F1 score of the T modality by 0.3% and the G modality by 2.7%, indicating that SSL can effectively mine latent features in unlabeled data and improve the classification performance of a single modality. In the multimodal fusion scenario, SSL improves the F1 score by 0.4%, demonstrating that SSL not only enhances the complementarity and synergy of cross-modal features but also further improves the robustness of the algorithm. Overall, the experimental results verify the effectiveness of multimodal fusion and the advantages of SSL in both single-modal and multimodal tasks, providing stronger feature learning capabilities for ECG signal analysis.

[0192] Table 4 shows the impact of IMSA, which brought a 0.8% improvement in F1 score, validating the effectiveness of the label-based semantic modality alignment method in enhancing cross-modal feature fusion. However, a 0.9% decrease in precision was observed, while recall improved by 2.0%. This phenomenon stems from IMSA's primary focus on category-related local information, with weaker learning of global features, leading to a decrease in model recognition accuracy. However, the supplementary local information across different modalities brought about by multimodal alignment improves the model's recall and reduces the false negative rate.

[0193] To address the shortcomings of IMSA in modeling global information, this invention proposes IMSE. IMSE enhances the perception of the overall feature distribution by fusing global features with local semantics. Table 4 shows that after introducing IMSE, Precision and F1 scores improved by 3.0% and 0.5%, respectively, further demonstrating the effectiveness of IMSE in multimodal feature fusion.

[0194] In summary, IMSA and IMSE complement each other, jointly constructing a multimodal fusion method that takes into account both local and global features, thereby improving the robustness and classification performance of the model.

[0195] The introduction of MLC during the classification phase significantly improved the Precision and F1 scores, demonstrating that MLC effectively alleviates the sample imbalance problem in multi-label classification by incorporating prior knowledge of label distribution. MLC combines label co-occurrence relationships with supervised contrastive learning, making classes with high co-occurrence probabilities closer together, while increasing the distance between classes with low co-occurrence probabilities or no co-occurrence relationships. This mechanism not only optimizes the representation distribution between classes but also enhances the model's discriminative ability on minority classes, thereby improving overall classification performance.

Claims

1. A multi-label electrocardiogram classification method based on self-supervised pre-training and multimodal semantic alignment, characterized in that: The steps include the following: Self-supervised pre-training phase based on contrastive learning: Step 1: Perform signal processing on the original ECG signal to obtain denoised one-dimensional sequence data, and generate a two-dimensional time-frequency graph using CWT; Step 2: Apply different data augmentation strategies to the ECG data of the two modalities to enhance the key features of each modality; Step 3: Use two different cropping methods to generate a comparison view on the enhanced modal data to complete the view construction; Step 4: Input the constructed view into a single-modal contrast enhancement network to capture modality-invariant features through contrastive learning; Semantic-guided multimodal fusion stage: Step 5: Perform semantic alignment between modalities. Use a fine-grained label alignment method to reduce the differences in semantic representation between different modalities and improve the collaborative representation between modalities by aligning the feature representations of specific categories. Step 6: Perform intramodal semantic enhancement. Through semantic enhancement attention, the label-specific tag-level representation is fused with the global representation of the entire object. The two complement each other to obtain the enhanced label-level representation. Step 7: Perform multimodal semantic fusion. Deep fusion of features from different modalities is achieved through a cross-attention mechanism. This process can dynamically generate attention weights based on the features of each modality to capture complementary information between modalities. At the same time, information from different modalities can interact effectively in a unified semantic space to generate more robust multimodal feature representations. Multi-label classification stage based on contrastive learning: Step 8: The enhanced feature information is fed into a multi-label classifier to obtain multi-label prediction results, and then the network is optimized using the binary cross-entropy loss (BCE) function and the proposed multi-label contrastive loss function.

2. The multi-label electrocardiogram classification method based on self-supervised pre-training and multimodal semantic alignment according to claim 1, characterized in that: Step 1: The specific steps are as follows: Step 1.1: Given a raw 12-lead ECG signal Power frequency filters are used to remove power supply noise, followed by the use of high-order Butterworth bandpass filters to remove electromyographic noise and baseline drift. Finally, the denoised signal representation is obtained. Step 1.2: Select a mother wavelet function, generate multiple sub-wavelets through scaling and translation operations, and convolve these sub-wavelets with the signal to obtain the time-frequency plot of the signal; for the input signal The mathematical definition of CWT is: Where ψ(t) is the mother wavelet function, a is the scaling parameter, b is the translation parameter, and ψ * (t) is the complex conjugate of the wavelet; since the Mexican Hat waveform is closer to the ECG signal, it is chosen as the mother wavelet function; the time-frequency plot obtained through CWT is X. img , where X img ∈R 12×h×w , where h is the frequency range and w is the number of sampling points.

3. The multi-label electrocardiogram classification method based on self-supervised pre-training and multimodal semantic alignment according to claim 1, characterized in that: In step 2, two data augmentation methods are used for the original one-dimensional ECG sequence data: random pruning and noise addition. The specific steps are as follows: Step 2.1: The augmented ECG one-dimensional sequence data is as follows: Where F crop For random cropping, F gaussian To add Gaussian noise; Step 2.2: Frequency domain masking and time domain masking are used to randomly mask the frequency or time domain of the time-frequency plot, so as to enable the model to learn the local features of the signal in different frequency bands or time periods. The time-frequency plot after data augmentation is as follows: Where F freq For frequency domain masking, F time Obscured by time.

4. The multi-label electrocardiogram classification method based on self-supervised pre-training and multimodal semantic alignment according to claim 1, characterized in that: Step 3, the view construction process, does not rely on negative samples. The specific steps are as follows: By randomly cropping the input image at different scales, global and local views are generated for comparison. The global view refers to an image area larger than 50% of the original image after random cropping, while the local view refers to an image area smaller than 50% after random cropping. The global view contains the overall features of the data, which helps to learn global contextual information, while the local view focuses on local regions of the image to help the model capture detailed features. By comparing the global and local views, the model can avoid relying solely on simple feature matching to achieve consistency, thereby promoting the learning of better feature representations.

5. The multi-label electrocardiogram classification method based on self-supervised pre-training and multimodal semantic alignment according to claim 1, characterized in that: In step 4, in the unimodal contrast enhancement network, the unimodal contrast view obtained after data augmentation and view construction is input into the teacher and student networks constructed using the same network structure for feature extraction. Then, the outputs of the two networks are projected onto the same feature space for contrast learning. The network parameters for feature extraction are optimized by minimizing the difference between the two. Finally, the trained student network parameters are saved for subsequent supervised contrast learning.

6. The multi-label electrocardiogram classification method based on self-supervised pre-training and multimodal semantic alignment according to claim 5, characterized in that: The specific components and setup procedures of each network used in step 4 are as follows: Feature extraction network for one-dimensional time series modalities: For one-dimensional time series modalities, ResNet1d50 is selected as the backbone network. A multi-scale convolutional attention mechanism is introduced into the backbone network to perform multi-scale analysis of the signal. For multi-lead ECG signals, each lead represents the electrical activity of different parts of the heart. Multi-lead analysis can integrate information from different perspectives, thereby revealing the complex characteristics of heart disease. In the multi-scale convolutional attention network, the feature map is uniformly divided into four feature map subsets after 1×1 convolution. These subsets have the same size as the original feature map, but the number of channels changes. Subsequently, multi-branch hierarchical convolution with a residual structure is used to obtain feature map combinations with different numbers and field sizes, thereby extracting multi-scale features of the one-dimensional signal. In the channel attention network, multiple leads of the ECG signal are regarded as channels, and the relationship between multi-lead signals is weighted by the SE network. Feature extraction network for 2D time-frequency map modality: For 2D time-frequency map modality, ResNet50 is selected as the backbone network. Combining the characteristics of image data, multi-scale convolutional attention and channel-spatial attention are modified to capture the time-frequency features of ECG signals in a single lead and the correlation information between multiple leads. In channel-spatial attention, firstly, adaptive max pooling and adaptive average pooling are used to extract global and local features of the image along the spatial dimension of the time-frequency map. Then, these two parts are added by channel squeezing excitation operation. The feature map containing key lead information is then spliced ​​after max pooling and average pooling along the channel dimension. Then, the focal information in the image space is obtained by convolutional attention. In multi-scale convolutional attention, firstly, multiple convolutional kernels of 1×1, 3×3 and 5×5 are used to extract receptive field information of different sizes. Then, the features of different receptive fields are cross-mixed by channel shuffling to break the fixed branch structure and enhance the interaction between multi-scale features. The entire single-modal contrast enhancement network: For one-dimensional sequence data, its contrast view can be represented as: in and These are collections of global views and local views, respectively. It contains two global views. Contains 8 partial views, F G_crop and F L_crop These are random cropping operations for the global view and the local view, respectively. Student and teacher networks: Both employ one-dimensional time-series modality feature extraction networks. First, the student network parameters are initialized, and then the student network parameters are copied to the teacher network to complete the teacher network initialization. Then, the teacher network extracts features from each global view, while the student network extracts features from all views, including both global and local views. The view representations learned by the teacher and student networks need to be mapped to a high-dimensional feature space through a projection network to facilitate comparative learning. The projected features need to be L2 normalized to ensure that the comparison between different features is not affected by scale differences, while ensuring the stability of the outputs of the student and teacher networks during the training process. Among them Teacher 1D () represents the teacher network, Student 1D () represents the student network, and Project() represents the projection network; finally, comparative learning is performed using the projection features of the teacher and student networks, this process iterates through each teacher output. and each student's output Calculate the cross-entropy loss between them; During training, the parameters of the student network are updated using standard backpropagation and gradient descent, while the parameters of the teacher network are updated using exponential moving average (EMA). The EMA method ensures smooth updates for the teacher network, avoiding the instability of the learning objectives caused by drastic changes in the student network and effectively preventing local overfitting during the student network's learning process. The parameter update formula for the teacher network is as follows: Where k represents the current training step number, λ k Represents the parameter smoothing coefficient. This represents the current teacher network parameters. This indicates the current student network parameters. This indicates the updated teacher network parameters.

7. The multi-label electrocardiogram classification method based on self-supervised pre-training and multimodal semantic alignment according to claim 1, characterized in that: Step 5 involves the following steps: First, obtain the feature embeddings of the corresponding modality through a pre-trained single-modality coding network: F sig =Encoder 1D (X sig ),F img =Encoder 2D (X img ) Encoder 1D () and Encoder 2D () represent pre-trained single-modal coding networks, F sig and F img These are feature maps for different modalities. For tag semantics, given a tag set L = (l1, l2, ..., l m ), where m is the number of labels, and the semantic embeddings of the label set are obtained using the pre-trained natural language model BERT: Y=(y1,y2,...,y m )=BERT(L) Where y j This represents the semantic embedding of a single tag. F sig and F img The feature maps of the two modalities are respectively obtained through label-specific attention under the guidance of label semantics. First, the correlation matrix between each ECG feature and the semantic embedding of a single label in the corresponding modality is calculated: Where P, U, V, and b are all learnable parameters, and ⊙ represents element-wise multiplication. It is a learnable feedforward neural network that uses a low-rank bilinear computation model to reduce computational complexity. Then, the label-level representation of the corresponding modality is obtained using the correlation matrix: This method achieves implicit semantic alignment of the two modalities at the label fine-grained level through label-specific attention.

8. The multi-label electrocardiogram classification method based on self-supervised pre-training and multimodal semantic alignment according to claim 1, characterized in that: Step 6 is as follows: By using semantically enhanced attention, the local information in the label-level representation and the global representation of the entire object are fused together, which not only preserves the fine information of the local features, but also introduces the contextual relationship of the global information. First, let's define the tag-level representation as h. sig,j and global representation f sig,i ∈F sig Obtain the query Q, key K, and value V through linear transformation: [Q,K,V]=[ω q h sig,j ,oh k f sig,i ,oh v f sig,i ] Then, the enhanced representation under a single label is obtained through attention calculation: Where u represents the learnable attentional bias.

9. The multi-label electrocardiogram classification method based on self-supervised pre-training and multimodal semantic alignment according to claim 1, characterized in that: Step 7 involves the following steps: Represent the enhanced tag level. Compared with the original feature map f sig,i By splicing, we get c sig,i ; Cross-attention mechanisms can achieve semantic association and feature fusion between different modalities. For the feature representation C of two modalities of ECG signals... sig =(c sig,1 ,c sig,2 ,...,c sig,n ) and C img =(c img,1 ,c img,2 ,...,c img,n First, the one-dimensional time series mode C sig As a query Q t Two-dimensional image modality C img As key K v Sum V v : [Q t ,K v ,V v ]=[ω q h sig,j ,oh k f sig,i ,oh v f sig,i ] Then, the fused two-dimensional image modality is obtained through cross-attention: Similarly, the fused one-dimensional time series mode representation is as follows: Finally, the results obtained through cross-attention and The fusion vectors of the two modalities are concatenated to obtain the multimodal joint representation C.

10. The multi-label electrocardiogram classification method based on self-supervised pre-training and multimodal semantic alignment according to claim 1, characterized in that: Step 8 is as follows: Combining the label co-occurrence relationship and supervised contrastive learning in the multi-label problem, a multi-label contrastive loss function is proposed. It aims to shorten the distance between the representation vectors of categories with high label co-occurrence probability and widen the distance between categories with low label co-occurrence probability or no co-occurrence relationship. This allows the model to better utilize label relationships to optimize the distance between categories during the discrimination process and alleviate the problem of sample class imbalance. The feature vector C obtained by multimodal fusion is projected onto the label category through a linear classifier, and then the final predicted label is obtained through the sigmoid activation function; for each sample X rawi Its label set is L = (l1, l2, ..., l m If BCE is calculated as follows: Where y ik For the sample The true value of the k-th label, The corresponding predicted value; First, a tag co-occurrence matrix A is constructed using the tag co-occurrence frequency, where elements A... ij This represents the co-occurrence frequency of label i and label j in the training set; then, based on A, the label relationship weight matrix W is obtained: Where N is the total number of samples in the current mini-batch, g(i) is all samples in the current mini-batch except i, τ is the temperature coefficient, and W ij Let W be the weight of the label relationship between samples i and j. For a sample pair (i,j), the greater the label relevance, the higher the weight W. ij The larger the value, the larger the corresponding loss term, thus allowing the distance between them to be optimized closer. The overall training loss is a weighted sum of the BCE and the multi-label contrastive loss: L=L BCE +λL con Where λ is a weighting parameter that controls the proportions of the two losses.

Citation Information

Cited By

  • Multi-source data classification method and device based on multi-channel comparison, equipment and medium

    CN121093098A

  • Multi-source data classification method and device based on multi-channel contrast, equipment and medium

    CN121093098B

  • Low-temperature boat hanging frame fault diagnosis method and system based on multi-mode dynamic fusion

    CN121117981A

  • Pseudo-CT cross-modal conversion method and system of multi-modal feature coupling diffusion model

    CN121120367A

  • High-frequency electrocardiogram coronary artery criminal blood vessel positioning method and system

    CN121120622A