Electrocardiogram classification method based on multi-modal feature fusion and fuzzy integral decision

Through the ECG classification method of multimodal feature fusion and fuzzy integral decision-making, the problems of insufficient multimodal fusion and insufficient decision-making reliability are solved, and the accuracy and interpretability of ECG diagnosis are improved.

CN120336930APending Publication Date: 2025-07-18NANTONG UNIV

Patent Information

Application Number
CN202510500075.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing electrocardiogram analysis methods have problems such as insufficient multimodal fusion and insufficient decision-making reliability, resulting in low diagnostic efficiency and poor reproducibility of results.

Method used

The electrocardiogram classification method based on multimodal feature fusion and fuzzy integral decision is adopted to extract clinically interpretable features through discrete wavelet transformation, transconductor spatiotemporal features are extracted using Transformer depth encoder, and feature fusion is achieved through a two-way cross attention mechanism, combining fuzzy integral decisions to improve diagnostic accuracy.

Benefits of technology

It significantly improves the accuracy and interpretability of electrocardiogram diagnosis, enhances the deep interaction ability of multimodal feature fusion and the robustness of diagnostic decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336930A_ABST
    Figure CN120336930A_ABST
Patent Text Reader

Abstract

The invention provides an electrocardiogram classification method based on multi-modal feature fusion and fuzzy integral decision, belongs to the technical field of artificial intelligence and medical treatment, and solves the technical problems of insufficient multi-modal fusion and low decision reliability in the prior art. According to the technical scheme, the method comprises the steps that time domain and frequency domain statistical features are extracted through discrete wavelet transform, and trans-lead spatial-temporal features are extracted through a Transform-based depth encoder; a bidirectional cross attention mechanism is adopted, and complementarity between modes is captured; then three classifiers are constructed, fuzzy integration is introduced for decision fusion, the dependency relationship between the classifiers is modeled through fuzzy measurement, and weights are dynamically distributed; the method has the beneficial effects that the clinical interpretable features are combined with the data driving model, the multi-modal feature fusion depth is improved through a bidirectional attention mechanism, the decision robustness is enhanced by using the fuzzy integral, and the accuracy and interpretability of automatic diagnosis of cardiovascular diseases are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and medical health technologies, and particularly to an electrocardiogram classification method based on multi-modal feature fusion and fuzzy integral decision-making. Background Art

[0002] As one of the diseases with the highest lethality rate globally, the early and accurate diagnosis of cardiovascular diseases is of great significance for reducing the mortality rate of patients. Electrocardiogram (ECG), as a non-invasive, economical, and widely used detection method, plays an irreplaceable role in the clinical diagnosis of cardiovascular diseases. However, traditional ECG analysis mainly relies on the subjective experience of doctors, not only with low diagnostic efficiency, but also having problems such as inconsistent diagnostic criteria and poor result repeatability. With the development of artificial intelligence technology, the automatic diagnosis methods of ECG based on machine learning and deep learning have gradually become a research hotspot, but there are still several key problems to be solved in the existing technologies.

[0003] In terms of feature extraction, the existing technologies are mainly divided into two categories: one is the method based on traditional machine learning, such as support vector machine (SVM) and principal component analysis (PCA). This kind of method requires manual design of features, not only relying on the domain knowledge of experts, but also having poor robustness of the designed features when facing noise interference or signal distortion; the other is the method based on deep learning, such as convolutional neural network (CNN) and Transformer model. Although they can automatically extract features, the extracted features lack clinical interpretability and it is difficult to establish a corresponding relationship with the diagnostic logic of doctors. In terms of multi-modal feature fusion, the existing technologies generally adopt simple feature splicing or weighted average strategies, and fail to fully consider the non-linear interaction relationship between different feature modalities. In addition, the existing methods often only focus on the fusion of high-level abstract features, while ignoring the important diagnostic information contained in low-level features, resulting in information loss. In terms of decision fusion, most existing ECG classification systems adopt independent classifiers or majority voting strategies, and fail to fully consider the dependence relationship between classifiers. The traditional weighted average method, due to its inability to model the cooperation or competition relationship between classifiers, will show a significant decline in its classification performance when facing data imbalance or noise interference. Summary of the Invention

[0004] The present invention provides an electrocardiogram classification method based on multi-modal feature fusion and fuzzy integral decision-making, which aims to solve the technical problems such as insufficient multi-modal fusion and insufficient decision reliability existing in the existing ECG automatic classification technology.

[0005] The main inventive concept of the present invention is as follows: The core of the present invention lies in constructing an electrocardiogram classification method that integrates clinical knowledge guidance, multi-modal feature dynamic interaction, and fuzzy integral decision-making. This method first obtains statistical features with clear clinical significance from ECG signals through a knowledge-guided feature extraction module, and extracts deep features representing signal patterns and structures through a neural network; then uses a bidirectional cross-attention mechanism to achieve deep interaction and fusion of different modal features; finally, adopts a decision-making method based on fuzzy integral, fully considering the dependence relationship between classifiers, and significantly improving the accuracy and reliability of the diagnostic results.

[0006] To achieve the above-mentioned inventive purpose, the present invention adopts the following technical solutions: An electrocardiogram classification method based on multi-modal feature fusion and fuzzy integral decision-making, comprising the following steps:

[0007] S1. The system preprocesses and analyzes the original multi-lead electrocardiogram signal, and extracts a set of statistical features with interpretability and discriminability based on clinical medical knowledge. This module uses discrete wavelet decomposition and reconstruction to comprehensively reflect the dynamic change characteristics of cardiac electrical activity, providing interpretable input information for the subsequent diagnostic model;

[0008] S2. The system uses a neural network to perform end-to-end modeling on the original electrocardiogram signal. By constructing a time-series segmentation structure with leads as units and introducing positional encoding and lead embedding, the signal is transformed into a unified multi-dimensional representation. Subsequently, this input is fed into a deep encoder based on Transformer to extract spatio-temporal features across leads and time slices, and output a deep feature vector representing signal patterns and structures. This module can automatically learn the pathological information hidden in time-series data and improve the recognition ability of the model;

[0009] S3. The system constructs a bidirectional cross-attention mechanism to map the statistical features and deep encoded features to the same feature space, perform bidirectional attention interaction, and capture the correlation and complementarity between modalities. The fusion module outputs a joint feature representation, providing a more comprehensive diagnostic basis for the subsequent decision-making layer;

[0010] S4. The system constructs classifiers for statistical features, deep features, and fusion features respectively, and outputs multiple prediction results. By introducing Choquet fuzzy integral, the outputs of multiple classifiers are weighted and fused using fuzzy measure, thereby enhancing the stability and decision-making reliability of the system;

[0011] S5. The system integrates the above modules into an end-to-end diagnostic process, trains and validates through a standard electrocardiogram dataset, and uses a joint optimization strategy to synchronously learn the model parameters and fuzzy measure, and finally outputs the target disease category and performs performance evaluation.

[0012] Further, S1 specifically includes the following steps:

[0013] S11. Collect the standard 12-lead electrocardiogram (ECG) signals. The length of each lead signal is T, and there are a total of L = 12 leads, forming the original input matrix where N is the number of samples;

[0014] S12. Perform unified preprocessing on all lead signals. Use a 4th-order Butterworth band-pass filter with a frequency band of 0.67 - 48 Hz to remove high-frequency noise, electromyogram interference, and low-frequency baseline drift, and perform normalization processing to unify the signal amplitude to the range of [-1, 1], obtaining the signal input matrix

[0015] S13. Use discrete wavelet transform to perform 7-layer wavelet decomposition on the second lead signal. Use the Daubechies wavelet as the mother wavelet function to extract 7 groups of detail coefficients (D1 - D7) and 1 group of approximation coefficients (A7) to obtain the multi-resolution information of the signal;

[0016] S14. In the original signal and the reconstructed signal, calculate the following statistical features respectively:

[0017] Time-domain features: RR interval, PR interval, QRS duration, QT interval, heart rate (HR), heart rate variability (SDNN, RMSSD), etc.;

[0018] Frequency-domain features: power spectral entropy, Shannon entropy, sample entropy, signal complexity, kurtosis, and skewness, etc.;

[0019] S15. Stack the D1-dimensional statistical features extracted from the second lead in order to form a statistical feature matrix And use the Z-score normalization method to unify its dimension, so that each dimension of the feature has a distribution with a mean of 0 and a standard deviation of 1;

[0020] S16. Take F stat as the output of the statistical feature extraction module for subsequent use by the fusion module and the classifier, and retain the lead dimension information for subsequent feature space matching.

[0021] Further, S2 specifically includes the following steps:

[0022] S21. For the signals preprocessed in step S1 For the signals of each lead Divide them into N patches respectively, and the length of each patch is p;

[0023] S22. Project each patch to a fixed dimension d using a fully connected layer patch , forming an embedded feature representation;

[0024] S23. Design two embedding mechanisms to enhance the input structure information:

[0025] Position embedding Used to distinguish the temporal position of patches in the lead signals, and adopt learnable vectors. Lead embedding Assign a unique learnable vector to each lead to represent its spatial position;

[0026] S24. Concatenate the patches of all leads into an input tensor Input it into the Transformer encoder module, and construct multiple layers of encoders to model the temporal dependencies within patches and the cross-lead spatial correlations between patches, where d emb Is the dimension of the concatenated embedding;

[0027] S25. After encoding, aggregate the output vectors of each lead by average pooling to form the representation vector of each sample As the output of the deep modeling module, where D2 is the dimension of the representation vector;

[0028] Furthermore, the specific steps of S3 are as follows:

[0029] S31. Map the statistical feature F stat And the deep feature F ecg To a unified dimension d fusion Respectively through linear transformation to obtain the corresponding query, key, and value:

[0030] Q1, K1, V1 = F stat W1

[0031] Q2, K2, V2 = F ecg W2

[0032] Where Q1, K1, V1 are the query, key, and value obtained by the linear mapping of the statistical feature respectively, and W1 is the learnable matrix of the linear mapping of the statistical feature; Q2, K2, V2 are the query, key, and value obtained by the linear mapping of the deep feature respectively, and W2 is the learnable matrix of the linear mapping of the deep feature.

[0033] S32. Calculate the attention of the deep feature F ecg Of the ECG signal relative to the statistical feature F stat Calculate the attention by using the query Q1 from the statistical feature and the keys K2 and values V2 from the deep feature to achieve the attention alignment of the deep feature guiding the statistical feature, and obtain the intermediate representation F 12 :

[0034]

[0035] S33. Inverse calculate the statistical features F of the ECG signal stat Attention relative to the depth feature F ecg Calculate the attention by using the query Q2 from the depth feature and the keys K1 and values V1 from the statistical features, realizing the attention alignment of the statistical features guiding the depth features, and obtaining the intermediate representation F 21 :

[0036]

[0037] S34. Concatenate F 12 and F 21 in the feature dimension, and integrate the concatenated features into a unified fused feature representation F fusion :

[0038] F fusion = Concat(F 12 , F 21 )W3

[0039] where Concat(.) represents the concatenation operation of matrices.

[0040] S35. Output F fusion for the subsequent sub-model classification and fuzzy integration process.

[0041] Furthermore, the S4 specifically includes the following steps:

[0042] S41. Input F stat , F ecg and F fusion into the classifier respectively, and output representing the prediction probabilities for C disease categories;

[0043] S42. Concatenate the vectors to form the classifier output matrix

[0044]

[0045] The matrix element y ij ∈Y, the row index i ∈ {1, 2, 3} respectively corresponds to the prediction results of the statistical features, the depth features, and the fused features, and the column index j ∈ {1, 2,..., C} represents the disease category number.

[0046] S43. Construct the fuzzy measure. First, define the differentiable fuzzy measure function g θ : 2 {1,2,3} → [0, 1], where the indices 1, 2, 3 respectively correspond to the prediction results of the statistical features the prediction results of the depth features and the prediction results of the fused features The fuzzy measure function is dynamically modeled by the parameter vector θ = [θ1, θ2, …, θ7], with a parameter dimension of 7, covering the measure values of all non-empty subsets. To satisfy the mathematical constraints of the fuzzy measure, g θ (.) represents the mapping from the set to the parameter θ i and the following constraints can be adopted:

[0047] 1. Boundary condition constraint: By setting and g θ ({1, 2, 3}) = 1, the completeness of the measure space is ensured;

[0048] 2. Monotonicity constraint: A regularization term is introduced into the loss function to penalize the parameter update direction that violates monotonicity, where λ is a hyperparameter.

[0049] S45. Fuzzy measure parameterization. First, define the learnable parameters α1, α2, and α3, which are mapped to independent weights through the sigmoid function to quantify the independent decision contribution degrees of each model. Define the interaction parameter β 12 , β 21 and β 23 , and use the non-linear function φ(x) = ln(1 + e x ) to calculate the joint weight increment. The joint weight of model i and model j, i < j, is defined as:

[0050] g({i, j}) = min(θ i + θ j + φ(β ij ), 1)

[0051] Then, normalize the weights to satisfy the constraint that the measure of the entire fuzzy measure set is 1.

[0052] S46. To implement gradient descent optimization, design a differentiable fuzzy integral calculation process:

[0053] First, extract the classifier output matrix Y and generate a matrix by sorting the column elements in descending order Then, based on the matrix S, calculate the Choquet integral weights:

[0054] w1 = g θ ({σ1})

[0055] w2 = g θ ({σ1, σ2}) - g θ ({σ1})

[0056] w3 = 1 - g θ ({σ1, σ2})

[0057] where σ i represents the model index corresponding to the i-th sorting position determined by the matrix S.

[0058] Finally, calculate the fused prediction value for class j

[0059]

[0060] S46. Calculate the prediction probability Calculate the cross-entropy loss function between the prediction probability and the true label, calculate the gradient, and update the parameter θ through backpropagation.

[0061] Furthermore, the S5 specifically includes the following steps:

[0062] S51. Combine the S1 to S4 modules into an end-to-end trainable network. The input is the 12-lead ECG, and the output is the classification result of the target cardiovascular disease.

[0063] S52. Use multiple public datasets (such as PTB-XL, Chapman, CPSC2018) for training and evaluation. Standardize the samples and divide them into training set, validation set, and test set.

[0064] S53. Use the weighted cross-entropy as the loss function, the AdamW optimizer, set the initial learning rate to 1e-4, the batch size to 64, train for 100 epochs, and use early stopping to prevent overfitting.

[0065] S54. Perform independent inference on each input sample in the test phase, output the final classification result, and evaluate it in five dimensions: accuracy, sensitivity, specificity, F1 score, and AUROC value.

[0066] Compared with the prior art, the beneficial effects of the present invention are:

[0067] 1. The depth interaction ability of multi-modal feature fusion is improved. Through the bidirectional cross-attention mechanism, the system can achieve dynamic interaction between statistical features (time-domain indicators, signal entropy, etc.) and deep spatio-temporal features (inter-lead dependence relationships), solving the problem of insufficient utilization of inter-modal information complementarity caused by feature concatenation or weighted averaging in traditional methods, and enhancing the representation ability of the fused features.

[0068] 2. Optimization of interpretability and robustness of diagnostic decisions. By combining knowledge-guided feature extraction (wavelet decomposition for signal reconstruction, heart rate variability metrics) with a fuzzy integral dynamic fusion strategy, the system integrates clinical diagnostic logic while retaining the advantages of data-driven approaches. The statistical feature matrix directly corresponds to commonly used diagnostic metrics by doctors (such as QT interval, heart rate variability), supporting the interpretability analysis of the decision-making process; the fuzzy integral suppresses noise interference through dynamic weight allocation, enhancing robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 It is a block diagram of an electrocardiogram classification system for multi-modal feature fusion and fuzzy integral decision-making in the present invention;

[0070] Figure 2 It is a block diagram of signal decomposition and reconstruction based on discrete wavelet transform in the present invention;

[0071] Figure 3 It is a block diagram of feature fusion based on a bidirectional cross-attention mechanism in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0072] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Of course, the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0073] Embodiment 1

[0074] This embodiment provides an electrocardiogram classification method based on multi-modal feature fusion and fuzzy integral decision-making, including the following steps:

[0075] S1. First, perform the operation of collecting and preprocessing the original multi-lead electrocardiogram signal, which provides standardized and clinically significant input data for subsequent feature extraction and diagnostic models. The specific steps are as follows:

[0076] First, collect the standard 12-lead electrocardiogram (ECG) signals, resample the signals to a frequency of 250 Hz, and intercept 10 seconds of continuous ECG activity data for each sample. To improve the signal quality, the system performs unified band-pass filtering and notch processing on the original ECG signals. Among them, the band-pass filter uses a fourth-order Butterworth filter, and the set passband range is from 0.67 Hz to 48 Hz, effectively filtering out electromyogram interference and low-frequency baseline drift. To eliminate the mains interference, the system further performs 50 Hz notch processing on each lead signal. Subsequently, the system performs artifact removal and normalization operations. For possible sudden spikes and baseline fluctuation artifacts, a method combining wavelet packet threshold denoising and moving average strategy is adopted to smooth the waveform profile and improve the signal stability of the ECG. To achieve a consistent input amplitude across samples, the system normalizes the maximum and minimum values of each lead signal, scaling all signals uniformly to the [-1, 1] interval. Finally, a signal input matrix is obtained. where L represents the number of leads, N represents the total number of samples, and T represents the number of sampling points for each lead.

[0077] After the above preprocessing operations are completed, the system selects the second lead as the standard lead for statistical feature extraction. Based on the commonly used clinical method, the signal of this lead is decomposed by seven-layer Daubechies wavelet transform to obtain seven groups of detail coefficients D1, D2, …, D7 and a group of approximation coefficients A7. After wavelet transform, the system performs inverse transform reconstruction on the signals of each scale respectively, so as to separate the signal components of different frequency bands, laying a foundation for subsequent frequency-domain and time-domain feature analysis.

[0078] Next, the system extracts statistical features with clear clinical interpretation significance from the original signals and wavelet reconstructed signals, specifically including but not limited to: RR interval, PR interval, QRS width, QT interval, average heart rate (HR), heart rate variability (SDNN, RMSSD), power spectrum entropy, Shannon entropy, sample entropy, signal complexity, kurtosis, and skewness. The extraction of all features is based on the international ECG diagnostic standards and key indicators defined by clinicians, ensuring their interpretability and medical consistency.

[0079] After the statistical feature extraction is completed, the system stacks the above features in a predetermined order to form a feature matrix. The system calculates the mean and standard deviation for each dimension feature of the feature matrix column by column and performs Z-score normalization processing. The conversion formula is:

[0080]

[0081] where μ j , σ j are the mean and standard deviation of the j-th column feature in the training samples respectively, and x ijThey are the elements in the transformed and original feature matrices respectively. The standardization process aims to eliminate the influence of different dimensions on model training and improve the model convergence speed and performance stability.

[0082] Through the above steps, the system has completed the normalization preprocessing of the original ECG signal and the extraction operation of statistical features with medical guiding significance, and outputs a standardized statistical feature matrix and the lead signal matrix after preprocessing for use by the subsequent deep feature extraction module and feature fusion module.

[0083] S2. Learning and modeling operations of deep features. This step aims to fully mine the non-linear structural information, local waveform variability and cross-lead temporal patterns existing in the ECG signal, and its operation process is as follows:

[0084] First, the system divides each preprocessed lead signal into several equally long segments (patches) in chronological order, and the length of each segment is set to 250 sampling points. For a 10-second, 2500-point input signal, this process divides each lead into 20 non-overlapping patches. This segmentation strategy not only retains the local dynamic features of the signal but also provides a modeling basis for the subsequent model in the time window.

[0085] For each divided patch, the system first uses a one-dimensional convolution operation to extract its local context features. The convolution kernel length is 3, the stride is 1, and the number of output channels is 64. Then it is sent to a fully connected layer to be further mapped to a 64-dimensional embedding space.

[0086] To enhance the temporal order information and lead discrimination information, the system adds two learnable embedding vectors to each patch respectively: Position Embedding and Lead Embedding. The position embedding represents the temporal position of the patch in the signal sequence, and the lead embedding identifies the ECG lead number it belongs to. After adding the above three, the final embedded representation tensor is obtained.

[0087] Subsequently, the system inputs this multi-dimensional representation into a sequence modeling module composed of multiple layers of Transformer encoders. The Transformer module contains 6 stacked self-attention encoders, each layer contains 8 attention heads, the hidden layer dimension is 256, and a combination of residual connection and feed-forward network structure is used to realize high-order non-linear interaction modeling across patches and leads.

[0088] After the system obtains the encoder output, it performs an average pooling operation on the output results of each lead, thereby generating a set of deep temporal feature representations for each sample. This feature matrix is denoted as

[0089] Through the above operations, the system not only retains the information of local ECG signal segments but also models the context dependencies between different leads, providing important support for constructing a unified representation in the subsequent multi-modal fusion stage.

[0090] S3. After completing the extraction of statistical features and deep features, the system enters the multi-modal feature fusion stage. The core objective of this stage is to unify the expression and complementary enhancement of feature vectors from two different modalities, thereby improving the discriminative ability of downstream classifiers. The fusion process is implemented using a bidirectional cross-attention mechanism and includes the following steps:

[0091] First, the system projects the statistical feature matrix F stat and the deep feature matrix F ecg into a unified dimensional space through two linear mapping modules respectively to obtain intermediate feature representations. The dimensions of the linear mapping matrices W1 and W2 are 50×128 and 256×128 respectively.

[0092] Subsequently, the system constructs a bidirectional cross-attention structure: on the one hand, using statistical features as query terms, and deep features as key and value terms; on the other hand, using deep features as query terms, and statistical features as key and value terms. Its expression is as follows:

[0093] Q1, K1, V1 = F stat W1

[0094] Q2, K2, V2 = F ecg W2

[0095] where Q1, K1, V1 are the query, key, and value obtained by linearly mapping the statistical features respectively, and W1 is the learnable matrix for linearly mapping the statistical features; Q2, K2, V2 are the query, key, and value obtained by linearly mapping the deep features respectively, and W2 is the learnable matrix for linearly mapping the deep features.

[0096] Through the above calculations, the system obtains output vectors in two fusion directions:

[0097]

[0098] These two vectors respectively represent the response features of statistical information to the deep representation and the response features of the deep representation to the statistical features.

[0099] Next, the system concatenates F 12 and F 21 in the feature dimension to form a fusion matrix, and projects it to the final fusion dimension through a third linear mapping to obtain a fusion feature matrix.

[0100] Through the above-mentioned bidirectional attention mechanism and projection transformation, the system has completed the unified modeling of multi-source modal information and output a feature matrix. This will be used as the input for the subsequent classifier.

[0101] S4. In the fourth step of the present invention, the system uses three types of feature expressions obtained in the statistical feature extraction, deep feature encoding, and multi-modal fusion stages to construct three independent classification models respectively, and introduces a differentiable Choquet fuzzy integral mechanism to dynamically weight and integrate their prediction outputs to enhance the discriminative power and robustness of the model.

[0102] First, the system establishes a two-layer feedforward neural network structure as a classifier for each type of feature, specifically as follows: for statistical features (dimension 50), a network structure of 128→64→5 is constructed, and the ReLU activation function is used between the fully connected layers; for deep features (dimension 256), a network structure of 256→128→5 is constructed; for fusion features (dimension 256), a network structure of 256→128→5 is constructed; where the five-dimensional vector output by the last layer represents the prediction probabilities for the five types of cardiovascular diseases, satisfying the normalization condition.

[0103] The system respectively denotes the prediction vectors output by each classifier as To achieve the final decision output, the system introduces a decision function based on the Choquet fuzzy integral.

[0104] To further illustrate the specific operation method of this fuzzy integral fusion strategy, the following gives a numerical calculation example of a simplified three-classification task. Suppose the system performs a three-classification task for a certain electrocardiogram sample, and the three classes are: Class 1 is Normal, Class 2 is Atrial Fibrillation, and Class 3 is AVBlock.

[0105] The prediction results of the system's three classifiers for this sample are as follows in the table:

[0106] Table 1 Prediction results of different classifiers for the sample

[0107] Classifier Category 1 Category 2 Category 3 Statistical feature classifier 0.20 0.60 0.30 Deep feature classifier 0.10 0.80 0.50 Fusion feature classifier 0.15 0.75 0.60

[0108] Now, taking Class 2 as an example, the fuzzy integral calculation is as follows:

[0109] First, sort the prediction values from largest to smallest to obtain: S 12 = 0.80, S 22 = 0.75, S 32 = 0.60.

[0110] The fuzzy measure parameters obtained through systematic learning are: g({2}) = 0.40, g({2,3}) = 0.75, g({1,2,3}) = 1.

[0111] Substitute into the Choquet integral formula:

[0112]

[0113] Therefore, after the system fuses the three models, the predicted probability that the sample belongs to class 2 is 0.7325. Similarly, the fusion calculation can be performed on classes 1 and 3 according to this rule, and finally the fusion output vector is obtained: The system will use to make a final classification judgment on this sample.

[0114] Through the above classifier construction and fuzzy fusion mechanism, while maintaining a high accuracy rate, this system significantly improves the recognition ability for complex signal patterns, especially showing stronger stability and interpretability when the sample class distribution is uneven or the model prediction divergence is large.

[0115] S5. After completing the design of the fusion mechanism, the system enters the joint training and optimization stage of the model. In this step, the system constructs all the aforementioned sub-modules, including statistical feature extraction, deep encoding, cross-attention fusion, three types of classifiers, and fuzzy measure parameters, into an end-to-end differentiable network structure to achieve joint optimization.

[0116] To ensure the stable training and convergence of the model, the system uses the cross-entropy loss function as the main optimization objective. In terms of the optimization method, the system adopts the AdamW optimizer, sets the initial learning rate to 1×10 -4 , and the weight decay coefficient to 1×10 -5 to control overfitting. The model updates the parameters in a mini-batch manner with a batch size of 64 during training, and the maximum number of training epochs is 100. If the performance on the validation set does not improve for 10 consecutive epochs, the training is stopped early (early stopping strategy).

[0117] After the training is completed, the system conducts performance evaluation on the reserved test set, and the metrics include: classification accuracy (Accuracy), (Precision), recall rate (Recall), F1 score, and AUROC (area under the curve) for each class. Table 1 shows the electrocardiogram classification results on the PTBXL dataset. In terms of all performance metrics, the proposed model outperforms the comparison methods.

[0118] Table 2 Electrocardiogram classification results on the PTBXL dataset

[0119]

[0120] The above results show that in the multi-lead electrocardiogram signal classification task, the present invention can achieve stable and efficient end-to-end learning and inference performance, and has significant engineering practical value.

[0121] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An electrocardiogram classification method based on multi-modal feature fusion and fuzzy integral decision-making, characterized in that Including the following steps: S1. Collect and preprocess the original multi-lead electrocardiogram (ECG) signals, and extract time-domain and frequency-domain statistical features through discrete wavelet transform; S2. Use a neural network to perform end-to-end modeling on the ECG signals, and extract deep features across leads and across time; S3. Fuse the statistical features and deep features through a bidirectional cross-attention mechanism to obtain a joint feature representation; S4. Construct multiple classifiers based on the statistical features, deep features, and fused features respectively, and use Choquet fuzzy integral to fuse multiple classification results; S5. Based on the end-to-end model optimized by training, automatically classify and diagnose cardiovascular diseases.

2. The electrocardiogram classification method based on multi-modal feature fusion and fuzzy integral decision according to claim 1, wherein, In the step S1, the statistical feature extraction includes the following steps: S11. Perform multi-layer discrete wavelet decomposition on the ECG signals of the second lead to obtain several groups of detail coefficients and approximation coefficients; S12. Extract a set of clinically interpretable features from the original signal and the wavelet reconstructed signal, including RR interval, PR interval, QRS duration, QT interval, heart rate, heart rate variability, power spectral entropy, Shannon entropy, sample entropy, complexity, kurtosis, and skewness; S13. Stack the features of the above step S12 in a preset order to form a feature matrix and perform Z-score normalization processing, where D1 is the dimension of the statistical feature.

3. The electrocardiogram classification method based on multi-modal feature fusion and fuzzy integral decision according to claim 1, wherein In the step S2, the extraction of the depth features includes: dividing each lead signal into multiple segments of length p, and mapping each segment to a representation space of a fixed dimension d through a fully connected network patch ; the system adds learnable position embedding vectors and lead embedding vectors to each segment respectively, combines them into an input tensor, and inputs it into a model built by a multi-layer Transformer encoder; the depth feature representation obtained after the encoder output is averaged and pooled is used as the output of the depth modeling module, where D2 is the dimension of the characterization vector.

4. The electrocardiogram classification method based on multi-modal feature fusion and fuzzy integral decision according to claim 1, wherein, The step S3 realizes multi-modal fusion through a bidirectional cross-attention mechanism, including the following steps: First, the statistical feature F stat and the depth feature F ecg are respectively mapped to queries, keys, and values through linear transformations, specifically as follows: Q1, K1, V1 = F stat W1 Q2, K2, V2 = F ecg W2 Where Q1, K1, and V1 are the query, key, and value obtained by linearly mapping the statistical features respectively, and W1 is the learnable matrix of the linear mapping of the statistical features; Q2, K2, and V2 are the query, key, and value obtained by linearly mapping the deep features respectively, and W2 is the learnable matrix of the linear mapping of the deep features; Then, calculate the attention representation F of the statistical features to the depth features 12 : And the attention representation F of the depth features to the statistical features 21 : where d fusion is the dimensionality of the feature representation obtained by linearly transforming the statistical features and the depth features, and softmax is the normalized exponential function; Finally, F 12 and F 21 are concatenated by the function Concat(.) and projected into the fused feature F by the learnable matrix W3 fusion , that is: F fusion = Concat(F 12 , F 21 )W3。 5. The electrocardiogram classification method based on multi-modal feature fusion and fuzzy integral decision according to claim 1, wherein, In the step S4, the fuzzy integral fusion adopts the Choquet integral to integrate the outputs of multiple classifiers, including: respectively denoting the three prediction vectors obtained based on statistical features, depth features, and fusion features as concatenating them by rows into a matrix C represents the number of disease categories; for each column vector y of the matrix Y j generating a matrix by arranging them in descending order of element size and constructing a fuzzy measure function g that satisfies boundary and monotonicity θ :2 {1,2,3} →[0,1], and finally, the fusion output for class j is: where w1 = g θ ({σ1}), w2 = g θ ({σ1,σ2}) - g θ ({σ1}) and w3 = 1 - g θ ({σ1,σ2}) represents the weight matrix, σ i represents the model index corresponding to the i-th sorting position determined by the matrix S, S ij and y ij represent the elements in the matrices S and Y respectively.

6. A system for implementing the method according to any one of claims 1 to 5, characterized in that, Including: Clinical knowledge-guided statistical feature extraction module; Deep feature extraction module of Transformer encoder; multi-modal fusion module of bidirectional cross-attention; Classifier and fuzzy integral decision-making module; The clinical knowledge-guided statistical feature extraction module obtains time-domain and frequency-domain statistical features through discrete wavelet transform; The deep feature extraction module of Transformer encoder obtains deep features representing signal patterns and structures; The multi-modal fusion module of bidirectional cross-attention fuses statistical features and deep features to obtain a joint feature representation; The classifier and fuzzy integral decision-making module, the classifier obtains the prediction results of the statistical features, deep features, and joint features, and the fuzzy integral decision-making module fuses multiple classification results to obtain the final prediction result.

Citation Information

Patent Citations

  • Feature selection method based on FSA-Choquet fuzzy integration

    CN111709440A

  • Multi-lead multi-scale electrocardiogram detection method and system based on deep learning

    CN114587376A

  • Multi-lead pulse signal intelligent identification method and system based on deep learning

    CN117281528A

  • Detection Of Disease Conditions And Comorbidities

    US20170251985A1

Cited By

  • Long-sequence electrocardiosignal disease recognition system based on Transform architecture

    CN120983046A