Heart sound and electrocardio combined coronary heart disease detection method based on time-frequency fusion
Through the joint detection method of heart-sound electrocardiogram based on time-frequency fusion, combined with multi-scale convolution and attention mechanism, the signal weight is dynamically adjusted, which solves the problems of unreasonable signal weight allocation and insufficient frequency domain information in the prior art, and improves the accuracy and reliability of coronary heart disease classification.
Patent Information
- Application Number
- CN202510209569.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-25
AI Technical Summary
When the existing heart-sound electrocardiogram signal fusion method processes signals with poor quality, the weight allocation is unreasonable, resulting in poor signal fusion effect and affecting the accuracy of coronary heart disease classification. At the same time, existing deep learning technologies have failed to fully explore the value of frequency domain information in the classification of coronary heart disease.
The joint detection method of heart-sound electrocardiogram based on time-frequency fusion is adopted, and time-frequency features are extracted through multi-scale convolution and gated cycle units, combining channel and spatial attention mechanisms, self-attention and cross-attention mechanisms, and time-domain and frequency-domain information are fused. Dynamically adjust the weights of PCG and ECG signals and fuse them according to signal quality and confidence.
The accuracy of coronary heart disease classification and the accuracy and reliability of diagnostic results are improved. Especially in scenarios with poor single-modal signal quality, the dynamic fusion module can adaptively optimize the feature fusion process to reduce the impact of noise.
Smart Images

Figure CN120036791A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical signal processing, and specifically provides a method for joint detection of heart sound and electrocardiogram for coronary heart disease based on time-frequency fusion. Background Technique
[0002] Coronary heart disease is known as the "number one killer in the world" and is a major global public health challenge, with characteristics of concealment, progressive stability, and fatality. The latest data from the Chinese Center for Disease Control and Prevention shows that there are approximately 11.39 million coronary heart disease patients in China. It is a highly prevalent and fatal disease that endangers the health of middle-aged and elderly people. 28% of urban and rural residents' deaths are caused by coronary heart disease, imposing a heavy social and economic burden. As the gold standard for coronary heart disease detection, coronary angiography cannot be used as a routine detection method due to its invasiveness and high cost. Among them, phonocardiogram (PCG) and electrocardiogram (ECG) signals have been widely studied by domestic and foreign scholars because they reflect the electrical activity and mechanical state of the heart, and it has been found that the two can effectively reveal the physiological state and abnormalities of the heart. The combination of PCG and ECG provides more comprehensive information about the heart, and the two belong to different types of signals and can complement each other. The information of PCG and ECG can improve the diagnostic effect of coronary heart disease and has significant advantages in the diagnosis of coronary heart disease.
[0003] At present, the fusion of PCG and ECG signals has been widely concerned in the classification of coronary heart disease. The methods for diagnosing coronary heart disease using PCG and ECG signals adopt traditional manually designed features and use deep learning for classification, but they are time-consuming and laborious and do not capture features comprehensively enough. Extracting features of PCG and ECG signals using deep learning algorithms and using machine learning algorithms for classification has become the mainstream research direction. However, the existing methods still have many problems: First, when the quality of a single signal is poor, if the weights of the signals cannot be dynamically adjusted, inaccurate features may be extracted. Specifically, a signal with poor quality may be assigned too large a weight, while a signal with good quality may be assigned too small a weight. This unreasonable weight allocation will lead to poor signal fusion effect, thus affecting the accuracy of the final classification result; Second, the PCG and ECG signals of coronary heart disease patients are often accompanied by an increase in low-frequency components and a decrease in high-frequency components in the frequency domain, and the existing deep learning technologies have not fully explored the value of frequency domain information for coronary heart disease classification. Therefore, if the above problems can be effectively solved, the accuracy of coronary heart disease classification will be further improved, and the accuracy and reliability of the diagnostic results will be enhanced. Summary of the Invention
[0004] (1) Technical Problems to be Solved
[0005] Aiming at the deficiencies of the existing technology, the present invention provides a method for joint detection of heart sound and electrocardiogram for coronary heart disease based on time-frequency fusion, which solves the problems raised in the above background technique.
[0006] (2) Technical Solution
[0007] The present invention specifically adopts the following technical solutions to achieve the above object:
[0008] A method for joint detection of coronary heart disease by combining phonocardiogram (PCG) and electrocardiogram (ECG) based on time-frequency fusion, comprising the following steps:
[0009] S1: Perform data preprocessing on the original PCG and ECG signals, including noise reduction, baseline drift elimination, power line interference elimination, etc., and normalize the PCG and ECG signals. Use the Welch method to generate the power spectrum signals of the heart sound and electrocardiogram, and divide the time-domain signals and power spectrum signals of the PCG and ECG into equal-length segments;
[0010] S2: Construct PCG and ECG time-frequency feature extraction modules with the same structure, introduce multi-scale convolutional layers to capture shallow and deep features in the time-frequency information, thereby generating features in the spatial dimension. Use gated recurrent units to capture long-term dependencies in the time-frequency information and generate features in the time dimension. Introduce global pooling layers and convolutional layers to further capture global features. Introduce a channel attention module to allocate weights to each channel, and output comprehensive spatio-temporal feature results through convolution;
[0011] S3: Construct PCG and ECG time-frequency fusion modules with the same structure for joint processing and information supplementation. Use position encoding to embed position information into the time-domain and frequency-domain features. Respectively adopt self-attention mechanisms for the time-domain signals and power spectrum signals to extract and capture the global dependencies between features, and adopt cross-attention mechanisms to integrate the frequency-domain information as supplementary features into the time-domain features to generate time-frequency fusion features. The fused features are further non-linearly transformed through a feed-forward neural network, and at the same time, residual connections and layer normalization are used to accelerate training and stabilize the optimization process of the network;
[0012] S4: Dynamically fuse PCG and ECG features according to signal quality, construct PCG and ECG dynamic fusion modules with the same structure, and perform weighted fusion on the time-frequency fused features according to their respective contributions. This module introduces modal confidence, dynamically adjusts the weights of the PCG and ECG, reflects their respective influences on the final classification result, and this module can adaptively optimize the feature fusion process and generate accurate fused features to be input into the classifier for classifying coronary heart disease.
[0013] Further, the steps of generating the power spectrum by the Welch method in S1 are as follows: Divide the heart sound and electrocardiogram signals into N sub-signals of length M according to the data length, add a window function to each sub-signal to reduce the signal truncation effect, and perform discrete Fourier transform on each windowed signal to obtain the spectrum of each segment:
[0014] X m F(f) = F{x m (t)}
[0015] where X m F(f) is the frequency-domain representation of the m-th segment;
[0016] The power spectral density of each frequency point is obtained by averaging the squared magnitude of the Fourier transform result:
[0017]
[0018] where P(f) is the power spectrum estimate of frequency f, and w(t) is the window function.
[0019] Furthermore, in S2, the generated time-domain signal and power spectrum signal are respectively input into the time-frequency feature extraction module, which contains five branches to extract spatio-temporal features. Among them, three branches process the input signal through convolution kernels of different sizes to extract multi-scale features; one branch uses global pooling operation to capture global feature information; one branch captures the long-term dependencies in the signal sequence through a gated recurrent unit; the features of multiple channels are combined to achieve the extraction of multi-scale spatio-temporal features, capturing patterns of different frequencies or durations. After each convolutional layer, a ReLU layer and a BN layer are added for non-linear transformation and accelerating training. The channel attention module adaptively adjusts the weights of the channels, highlighting key features and suppressing irrelevant features.
[0020] Furthermore, the formula of the channel attention module is as follows:
[0021] α c = σ(W c × f c (X))
[0022] where α c is the weight of the C-th channel, W c is the feature extraction matrix of the channel weights, f c (X) is the channel feature extraction function, and σ is the activation function;
[0023] After being calculated by the channel attention module, the corresponding weights of each channel are obtained, and the formula is as follows:
[0024]
[0025] where is the C-th channel feature after being processed by the channel attention mechanism, and X c is the C-th channel feature before being processed by the channel attention mechanism.
[0026] Furthermore, in the time-frequency fusion module constructed in S3, position encoding is introduced, and the encoded features and the extracted features are concatenated and input into the self-attention module; the position encoding uses sine and cosine encoding to embed the sequence of temporal information into the features, enabling the model to utilize the sequential information of the sequence. Through the self-attention module, each part of the features can pay attention to each other, enabling the model to automatically focus on important features;
[0027] Among them, the position encoding formulas for sine and cosine are as follows:
[0028]
[0029] Among them, pos represents the position in the sequence, i represents the index of the feature dimension, d model is the total feature dimension of the model, is a scaling factor that controls the sine and cosine waves of different frequencies to ensure the uniqueness of the encodings at different positions;
[0030] The position-encoded features and the features of the feature extraction module are superimposed to output new features, expressed as:
[0031]
[0032] Among them, is the output result of the time-frequency feature extraction module;
[0033] In the time-frequency fusion module, the time-domain features and frequency-domain features with position encoding are input into the self-attention module, and the following formula is used for calculation:
[0034]
[0035] Among them, Q, K, and V respectively represent the query, key, and value of the features; the softmax function is used for normalization to ensure that the sum of the output attentions is 1;
[0036] The results obtained from the time-domain and frequency-domain through the self-attention mechanism are input into the cross-attention mechanism. Taking the time-domain features as the main features and integrating the frequency-domain information as supplementary features into the time-domain features, the calculation formula is as follows:
[0037]
[0038] Among them, K time represents the key of the time-domain features; Q frequency , V frequency respectively represent the query and value of the time-domain features.
[0039] Furthermore, in the time-frequency fusion module in S3, the feedforward neural network consists of two fully connected layers and one ReLU layer, mapping the features extracted by the attention mechanism to a higher-dimensional feature space and using the ReLU layer to alleviate the problem of gradient disappearance; skip connections are added to the residual connection to directly add the input to the output, enabling the model to retain the original features while learning new features and preventing the loss of feature information; layer normalization adjusts the mean and variance of the features of each layer to keep the output in a stable distribution.
[0040] Furthermore, in the PCG and ECG dynamic fusion module in S4, based on the optimization principle of the error upper bound, the relationship between the confidence and loss of the PCG signal and the ECG signal is used to dynamically adjust the weights. The fusion weights of different modalities should be negatively correlated with the losses of the corresponding modalities, that is:
[0041] Cov(ω m ,l m )<0
[0042] where ω m represents the fusion weight of modality m, and l m represents the loss of the corresponding modality;
[0043] The probability of predicting the coronary heart disease category is used as the weight of a certain signal after fusion:
[0044]
[0045] where P true is the probability of predicting the coronary heart disease category, and W m represents the weight of modality m;
[0046] In the PCG and ECG dynamic fusion module, a relative calibration model is constructed to predict the fusion weights. Binary classification of coronary heart disease is selected, and the distribution uniformity of binary classification is defined:
[0047]
[0048] where Softmax(f m (x m ))) i represents the prediction probability of modality m for class i.
[0049] Furthermore, in the PCG and ECG dynamic fusion module, considering the environmental changes, the uncertainties of the ECG and PCG should be relative. One modality should dynamically perceive the changes of the other modality and modify its relative contribution to the system. Relative calibration is introduced to calibrate the relative uncertainties of each modality:
[0050]
[0051] Among them, m and n represent ECG and PCG. If m is the ECG signal, then n is the PCG signal;
[0052] Dynamically select to reduce or maintain the modal weight according to the relative uncertainty of the modality. In RC m < 1, this modality has greater uncertainty, multiply the predicted accuracy rate and RC m to reduce the contribution of this modality. In RC m > 1, this modality has less uncertainty and an accurate prediction rate, maintain the contribution of this modality. Therefore, the following asymmetric calibration term is defined:
[0053]
[0054] Obtain the calibrated modal weight according to the asymmetric calibration term and the probability of the modality predicting the coronary heart disease category. The formula is as follows:
[0055]
[0056] Furthermore, the PCG and ECG dynamic fusion module uses the relative calibration strategy to dynamically adjust and allocate the weight W m , and then fuse the results of the time-frequency characteristics of PCG and ECG. The dynamic fusion result is:
[0057] F = W ECG × f ECG + W PCG × f PCG
[0058] where W ECG and W PCG are the fusion weights of the ECG signal and the PCG signal respectively, and f ECG and f PCG are the characteristics after the time-frequency fusion of the ECG signal and the PCG signal respectively.
[0059] (III) Beneficial effects
[0060] Compared with the prior art, the present invention provides a method for jointly detecting coronary heart disease by heart sound and electrocardiogram based on time-frequency fusion, which has the following beneficial effects:
[0061] 1. The present invention adopts a time-frequency convolution method, which cascades multi-scale convolution and gated recurrent units. First, convolution layers with different kernel sizes, max pooling layers, and gated recurrent units operate in parallel to extract spatio-temporal features and global features at different time scales; the channel and spatial attention modules adaptively adjust the weights of different channels and spatial positions by combining channel attention and spatial attention; and then the multi-scale spatio-temporal features are further synthesized through convolution. By adopting multi-scale convolution and gated recurrent units, the model enhances the multi-scale feature capture ability, is suitable for extracting features from a long time range, and is especially suitable for processing long sequence signals such as PCG and ECG.
[0062] 2. The cross-attention mechanism proposed by the present invention incorporates frequency-domain features as supplementary information into time-domain features and introduces positional encoding, so that the fused time-frequency features contain positional information. This can not only perceive the absolute position in the sequence, but also capture the relative relationship between positions, and is especially more likely to capture long-distance dependencies when processing long sequence signals such as PCG and ECG. In addition, the self-attention operation is performed on the features containing positional information and time-frequency information by using the attention mechanism to further explore the internal correlation of the features.
[0063] 3. The present invention dynamically determines the weights of PCG and ECG signals through the single-modal prediction rate and the full-modal prediction rate, and adopts a relative calibration strategy, so that the weight of each modality can be dynamically adjusted according to the quality of other modalities. The dynamic fusion module can adjust the modality weights according to the real-time quality of the data, thereby improving the reliability and stability of the fusion result under noise conditions and reducing the misjudgment risk caused by single-modal dependence. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 is a schematic diagram of the overall network framework of an embodiment of the present invention;
[0065] Figure 2 is a schematic diagram of the time-frequency feature extraction module of an embodiment of the present invention;
[0066] Figure 3 is a schematic diagram of the time-frequency fusion module of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0067] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0068] Embodiment
[0069] As Figures 1-3As shown in the figure, a method for jointly detecting coronary heart disease based on time-frequency fusion of heart sound and electrocardiogram proposed by an embodiment of the present invention includes the following steps:
[0070] S1: Preprocess the original PCG and ECG signals, including noise reduction, elimination of baseline drift, power line interference, etc. Normalize the PCG and ECG signals, generate the power spectrum signals of heart sound and electrocardiogram using the Welch method, and divide the time domain signals and power spectrum signals of PCG and ECG into equal-length segments.
[0071] Among them, for the preprocessing part, a Butterworth band-pass filter with a frequency range of 0.5 - 60 Hz is used to reduce noise and eliminate baseline drift of the ECG signal; a 25 Hz high-pass Butterworth filter is used to eliminate low-frequency noise and baseline drift of the PCG signal; a 50 Hz IIR notch filter is used to eliminate power line interference in the ECG and PCG signals.
[0072] The steps of generating the power spectrum by the Welch method are as follows: Divide the heart sound and electrocardiogram signals into N sub-signals of length M according to the data length, add a window function to each sub-signal to reduce the truncation effect of the signal, and perform a discrete Fourier transform on each windowed signal to obtain the spectrum of each segment:
[0073] X m (f) = F{x m (t)}
[0074] Among them, X m (f) is the frequency domain representation of the m-th segment.
[0075] Average the squared magnitude of the Fourier transform result to obtain the power spectral density at each frequency point:
[0076]
[0077] Among them, P(f) is the power spectrum estimate at frequency f, and w(t) is the window function.
[0078] S2: Construct PCG and ECG time-frequency feature extraction modules with the same structure, introduce multi-scale convolutional layers to capture shallow and deep features in time-frequency information, thereby generating features in the spatial dimension. Use gated recurrent units to capture long-term dependencies in time-frequency information and generate features in the time dimension. Introduce global pooling layers and convolutional layers to further capture global features. Introduce a channel attention module to allocate weights to each channel and output comprehensive spatio-temporal feature results through convolution.
[0079] Such as Figure 2As shown, the generated time-domain signal and power spectrum signal are respectively input into the time-frequency feature extraction module. This module contains five branches to extract spatio-temporal features. Among them, three branches process the input signal through convolutional kernels of different sizes to extract multi-scale features; one branch uses global pooling operation to capture global feature information; one branch captures the long-term dependencies in the signal sequence through a gated recurrent unit. The features of multiple channels are combined to achieve the extraction of multi-scale spatio-temporal features and capture patterns of different frequencies or durations.
[0080] Specifically, the convolutional kernel sizes of the three one-dimensional convolutions are 1, 3, and 5 respectively, and the stride is 1; the convolutional kernel size of the pooling layer is 3, the stride is 1, the convolutional kernel size of the convolutional layer is 1, and the stride is 1; the number of hidden units of the gated recurrent unit is 32. And after each convolutional layer, a ReLU layer and a BN layer are added for non-linear transformation and accelerating training. The channel attention module adaptively adjusts the weights of the channels, highlighting key features and suppressing irrelevant features. Finally, a convolutional layer with a convolutional kernel size of 7 and a stride of 2 is introduced for feature dimensionality reduction to improve computational efficiency.
[0081] The formula of the channel attention module is as follows:
[0082] α c =σ(W c ×f c (X))
[0083] where, α c is the weight of the C-th channel, W c is the feature extraction matrix of the channel weight, f c (X) is the channel feature extraction function, and σ is the activation function.
[0084] After being calculated by the channel attention module, the corresponding weights of each channel are obtained, and the formula is as follows:
[0085]
[0086] where, is the C-th channel feature after being processed by the channel attention mechanism, and X c is the C-th channel feature before being processed by the channel attention mechanism.
[0087] S3: Construct PCG and ECG time-frequency fusion modules with the same structure for joint processing and information supplementation. Position encoding is used to embed position information into time-domain and frequency-domain features. The self-attention mechanism is respectively adopted for the time-domain signal and the power spectrum signal to capture the global dependencies among features, and the frequency-domain information is integrated into the time-domain features as supplementary features to generate time-frequency fusion features. The fused features are further non-linearly transformed through a feed-forward neural network, and at the same time, residual connections and layer normalization are used to accelerate training and stabilize the optimization process of the network.
[0088] As Figure 3 shown, the obtained features are input into the constructed time-frequency fusion module. Position encoding is introduced, and the encoded features and the extracted features are concatenated and input into the self-attention module. Specifically, the position encoding uses sine and cosine encoding to embed the sequence of timing information into the features, enabling the model to utilize the sequential information of the sequence; through the self-attention module, each part of the feature itself pays attention to each other, enabling the model to automatically focus on important features.
[0089] Among them, the position encoding formulas of sine and cosine are as follows:
[0090]
[0091] Among them, pos represents the position in the sequence, i represents the index of the feature dimension, d model is the total feature dimension of the model, is a scaling factor that controls the sine and cosine waves of different frequencies to ensure the uniqueness of the encodings at different positions.
[0092] The position-encoded features and the features of the feature extraction module are superimposed to output new features, expressed as:
[0093]
[0094] Among them, is the output result of the time-frequency feature extraction module.
[0095] In the time-frequency fusion module, the time-domain features and frequency-domain features with position encoding are input into the self-attention module, and are calculated using the following formula:
[0096]
[0097] Among them, Q, K, and V respectively represent the query, key, and value of the feature; the softmax function is used for normalization to ensure that the sum of the output attentions is 1.
[0098] The results obtained from the time domain and frequency domain through the self-attention mechanism are input into the cross-attention mechanism. The time domain features are used as the main features, and the frequency domain information is integrated into the time domain features as supplementary features. The calculation formula is as follows:
[0099]
[0100] Among them, K time represents the key of the time domain features; Q frequency , V frequency represent the query and value of the time domain features respectively.
[0101] In the PCG and ECG time-frequency fusion module, the feed-forward neural network consists of two fully connected layers and a ReLU layer, which maps the features extracted by the attention mechanism to a higher-dimensional feature space and uses the ReLU layer to alleviate the problem of gradient disappearance; skip connections are added to the residual connection to directly add the input to the output, enabling the model to retain the original features while learning new features and preventing the loss of feature information; layer normalization adjusts the mean and variance of each layer of features to keep the output in a stable distribution.
[0102] S4: Dynamically fuse PCG and ECG features according to the signal quality, construct PCG and ECG dynamic fusion modules with the same structure, and perform weighted fusion on the time-frequency fused features according to their respective contributions. This module introduces modal confidence and reflects the influence of each on the final classification result by dynamically adjusting the weights of PCG and ECG. This module can adaptively optimize the feature fusion process, generate accurate fusion features and use an SVM classifier to classify coronary heart disease.
[0103] Based on the optimization principle of the error upper bound in the PCG and ECG dynamic fusion module, the relationship between the confidence and loss of the PCG signal and the ECG signal is used to dynamically adjust the weights. The fusion weights of different modalities should be negatively correlated with the losses of the corresponding modalities, that is:
[0104] Cov(ω m ,l m )<0
[0105] Among them, ω m represents the fusion weight of modality m, and l m represents the loss of the corresponding modality.
[0106] The probability of predicting the coronary heart disease category is used as the weight of a certain signal after fusion:
[0107]
[0108] Among them, P true is the probability of predicting the coronary heart disease category, and W m represents the weight of modality m.
[0109] In the PCG and ECG dynamic fusion module, a relative calibration model is constructed to predict the fusion weight. The present invention selects to perform binary classification on coronary heart disease, so the distribution uniformity of binary classification is defined as:
[0110]
[0111] Among them, Softmax(f m (x m ))) i represents the predicted probability of modality m for class i.
[0112] Considering the changes in the environment, in the PCG and ECG dynamic fusion module, the uncertainties of ECG and PCG should be relative. One modality should dynamically sense the changes of the other modality and modify its relative contribution to the system, and relative calibration is introduced to calibrate the relative uncertainty of each modality:
[0113]
[0114] Among them, m and n represent ECG and PCG. If m is the ECG signal, then n is the PCG signal.
[0115] Dynamically select to reduce or maintain the weight of this modality according to the relative uncertainty of the modality. When RC m < 1, this modality has greater uncertainty, and multiply the prediction accuracy rate and RC m to reduce the contribution of this modality. When RC m > 1, this modality has smaller uncertainty and accurate prediction rate, and maintain the contribution of this modality. Therefore, the following asymmetric calibration term is defined:
[0116]
[0117] According to the asymmetric calibration term and the probability of the modality predicting the coronary heart disease category, the calibrated modality weight is obtained, and the formula is as follows:
[0118]
[0119] The PCG and ECG dynamic fusion module dynamically adjusts and allocates the weight W m using the relative calibration strategy, and then fuses the results of the time-frequency characteristics of PCG and ECG. The dynamic fusion result is:
[0120] F = W ECG × f ECG + W PCG × f PCG
[0121] Among them, W ECG and WPCG The fusion weights f of the ECG signal and the PCG signal respectively ECG and f PCG are the features after time-frequency fusion of the ECG signal and the PCG signal respectively.
[0122] Finally, the fused result is input into a classifier for coronary heart disease classification.
[0123] Aiming at the problem that the current method for detecting coronary heart disease using heart sound and electrocardiogram does not consider the weight distribution of PCG and ECG, and the existing deep learning methods do not fully utilize the frequency domain information, the present invention provides a method for jointly detecting coronary heart disease based on time-frequency fusion of heart sound and electrocardiogram. The proposed method can dynamically adjust the weights of PCG and ECG signals, especially suitable for scenarios where the quality of single-modal signals is poor. Moreover, the method combines positional encoding and self-attention mechanism to introduce positional information, focuses on internal correlations, and fully utilizes the frequency domain information of the signals. Through the cross-attention mechanism, the frequency domain features are incorporated into the time domain features as supplementary information to improve the accuracy of subsequent classification. Therefore, the method proposed in this application can be used as an effective and feasible method for classifying coronary heart disease.
[0124] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for detecting coronary heart disease based on time-frequency fusion of heart sound and electrocardiogram, characterized by: The steps include: S1: Preprocess the PCG and ECG raw signals, including noise reduction, elimination of baseline drift and power line interference, and normalize the PCG and ECG signals. Use the Welch method to generate the power spectrum signals of heart sounds and electrocardiograms, and divide the time domain signals and power spectrum signals of PCG and ECG into segments of equal length. S2: Construct a PCG and ECG time-frequency feature extraction module with the same structure. A multi-scale convolution layer is introduced to capture the shallow and deep features in the time-frequency information, thereby generating features in the spatial dimension. A gated recurrent unit is used to capture the long-term dependencies in the time-frequency information and generate features in the temporal dimension. A global pooling layer and a convolution layer are introduced to further capture global features. A channel attention module is introduced to assign weights to each channel, and the comprehensive spatiotemporal feature results are output through convolution. S3: Construct a PCG and ECG time-frequency fusion module with the same structure for joint processing and information supplementation. Use position encoding to embed position information into time domain and frequency domain features. Use self-attention mechanism to extract and capture the global dependency between features for time domain signals and power spectrum signals respectively. Use cross-attention mechanism to integrate frequency domain information into time domain features as supplementary features to generate time-frequency fusion features. The fused features are further transformed nonlinearly through feedforward neural network. Use residual connection and layer normalization to accelerate training and stabilize the optimization process of the network. S4: Dynamically fuse PCG and ECG features according to signal quality, build a PCG and ECG dynamic fusion module with the same structure, and perform weighted fusion on the time-frequency fused features according to their respective contributions. This module introduces modal confidence and dynamically adjusts the weights of PCG and ECG to reflect their respective influence on the final classification results. This module can adaptively optimize the feature fusion process and generate accurate fusion features to input into the classifier for classifying coronary heart disease.
2. The method for detecting coronary heart disease based on time-frequency fusion of heart sound and electrocardiogram according to claim 1, characterized in that: The steps of generating the power spectrum by the Welch method in S1 are as follows: divide the heart sound and electrocardiogram signals into N sub-signals of length M according to the data length, add a window function to each sub-signal to reduce the truncation effect of the signal, and perform a discrete Fourier transform on each windowed signal to obtain the spectrum of each segment: X m (f)=F{x m (t)} Among them, X m (f) is the frequency domain representation of the mth segment; The power spectrum density at each frequency point is obtained by averaging the squared amplitude of the Fourier transform result: Where P(f) is the power spectrum estimate of frequency f and w(t) is the window function.
3. The method for detecting coronary heart disease based on time-frequency fusion of heart sound and electrocardiogram according to claim 1, characterized in that: In S2, the generated time domain signal and power spectrum signal are respectively input into the time-frequency feature extraction module, which includes five branches to extract spatiotemporal features, three of which process the input signal through convolution kernels of different sizes to extract multi-scale features; and one branch uses a global pooling operation to capture global feature information; One branch captures the long-term dependencies in the signal sequence through a gated recurrent unit; the features of multiple channels are combined to extract multi-scale spatiotemporal features and capture patterns of different frequencies or durations. ReLU layers and BN layers are added after each convolution layer for nonlinear transformation and accelerated training. The channel attention module adjusts the channel weights through adaptive adjustments to highlight key features and suppress irrelevant features.
4. The method for detecting coronary heart disease based on time-frequency fusion of heart sound and electrocardiogram according to claim 3, characterized in that: The formula of the channel attention module is as follows: a c =σ(W c ×f c (X)) Among them, α c is the weight of the Cth channel, W c is the feature extraction matrix of channel weights, f c (X) is the channel feature extraction function, σ is the activation function; The corresponding weight of each channel is calculated by the channel attention module. The formula is as follows: in, is the Cth channel feature after the channel attention mechanism, X c It is the Cth channel feature before being processed by the channel attention mechanism.
5. The method for detecting coronary heart disease based on time-frequency fusion of heart sound and electrocardiogram according to claim 4, characterized in that: The time-frequency fusion module constructed in S3 introduces position coding, and inputs the encoded features and the extracted features in series into the self-attention module; the position coding uses sine and cosine coding to embed the sequence of time series information into the features, so that the model can use the order information of the sequence, and through the self-attention module, let the various parts of the features pay attention to each other, so that the model can automatically focus on important features; Among them, the position encoding formulas of sine and cosine are as follows: Among them, pos represents the position in the sequence, i represents the index of the feature dimension, and d model is the total feature dimension of the model, is a scaling factor that controls the sine and cosine waves of different frequencies to ensure that the codes at different locations are unique; The position encoding features and the features of the feature extraction module are superimposed to output new features, which are expressed as: in, is the output result of the time-frequency feature extraction module; In the time-frequency fusion module, the time domain features and frequency domain features with position encoding are calculated through the self-attention module using the following formula: Among them, Q, K, and V represent the query, key, and value of the feature respectively; the softmax function is used for normalization to ensure that the sum of the output attention is 1; The results of the time domain and frequency domain obtained through the self-attention mechanism are input into the cross-attention mechanism, with the time domain feature as the main feature and the frequency domain information as a supplementary feature integrated into the time domain feature. The calculation formula is as follows: Among them, K time The key representing the time domain characteristics; Q frequency 、V frequency They represent the query and value of time domain features respectively.
6. The method for detecting coronary heart disease based on time-frequency fusion of heart sound and electrocardiogram according to claim 1, characterized in that: In the time-frequency fusion module of S3, the feedforward neural network consists of two fully connected layers and one ReLU layer, which maps the features extracted by the attention mechanism to a higher-dimensional feature space, and uses the ReLU layer to alleviate the problem of gradient vanishing. Skip connections are added to the residual connections so that the input is directly added to the output, so that the model can retain the original features when learning new features and prevent the loss of feature information. Layer normalization adjusts the mean and variance of the features of each layer to maintain a stable distributed output.
7. The method for detecting coronary heart disease based on time-frequency fusion of heart sound and electrocardiogram according to claim 1, characterized in that: The PCG and ECG dynamic fusion module in S4 uses the optimization principle based on the upper limit of the error to dynamically adjust the weights by using the relationship between the confidence and loss of the PCG and ECG signals. The fusion weights of different modalities should be negatively correlated with the loss of the corresponding modality, that is: Among them, ω m represents the fusion weight of mode m, Represents the loss of the corresponding mode; The probability of predicting the coronary heart disease category is used as the weight of a certain signal after fusion: Among them, P true To predict the probability of coronary heart disease category, W m represents the weight of mode m; A relative calibration model is built in the PCG and ECG dynamic fusion module to predict the fusion weight, and coronary heart disease is selected for binary classification, and the distribution uniformity of the binary classification is defined: Among them, Softmax(f m (x m ))) i Represents the predicted probability of modality m for category i.
8. The method for detecting coronary heart disease based on time-frequency fusion of heart sound and electrocardiogram according to claim 7, characterized in that: In the PCG and ECG dynamic fusion module, considering the changes in the environment, the uncertainties of ECG and PCG should be relative. One modality should dynamically perceive the changes of the other modality and modify its relative contribution to the system. Relative calibration is introduced to calibrate the relative uncertainty of each modality: Wherein, m and n represent ECG and PCG, if m is an ECG signal, then n is a PCG signal; According to the relative uncertainty of the mode, the modal weight is dynamically selected to be reduced or maintained. m <1, the mode has greater uncertainty, which will affect the prediction accuracy and RC m Multiply, reduce the contribution of this mode, in RC m When >1, the mode has smaller uncertainty and accurate prediction rate, and the contribution of this mode is maintained. Therefore, the following asymmetric calibration term is defined: The calibrated modal weight is obtained based on the asymmetric calibration term and the probability of modal prediction of the coronary heart disease category. The formula is as follows:
9. The method for detecting coronary heart disease based on time-frequency fusion of heart sound and electrocardiogram according to claim 8, characterized in that: The PCG and ECG dynamic fusion module uses a relative calibration strategy to dynamically adjust the distribution weight W m , and then fuse the results of PCG and ECG time-frequency features, the dynamic fusion result is: F=W ECG ×f ECG +W PCG ×f PCG Among them, W ECG and W PCG are the fusion weights f of ECG signal and PCG signal respectively ECG and f PCG They are the features of ECG signal and PCG signal after time-frequency fusion.
Citation Information
Patent Citations
Heart sound signal classification method and terminal equipment
CN108742697A
Mental stress analysis method based on video non-contact measurement
CN111714144A
Method and device for monitoring arrhythmia event
CN113164072A
Heart sound and electrocardio combined diagnosis device and system based on deep learning
CN115177262A
Heart sound classification method based on deep residual neural network
CN116030829A
Cited By
Electrical impedance cancer detection method based on frequency domain information enhancement
CN120114034A
A method for electrical impedance cancer detection based on frequency domain information enhancement
CN120114034B
Non-inductive sleep ambulatory blood pressure rhythm analysis method and system, equipment and medium
CN121080939A
A method and system for analyzing the dynamic blood pressure rhythm of sleep without induction, equipment and medium
CN121080939B
Heart sound classification method and device and computer readable storage medium
CN121583290A