Intelligent heart sound recognition method, system and equipment and medium
Through the PCA-Transformer model and independent component analysis method combined with deep learning model, the problem of relying on artificial characteristics and limited generalization capabilities of existing heart sound signal analysis methods is solved, and efficient, accurate identification and robustness of heart sound signals are achieved.
Patent Information
- Application Number
- CN202510420081.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-25
AI Technical Summary
The existing heart sound signal analysis methods rely on expert manual design characteristics and are difficult to capture complex signals. In addition, deep learning models require a large amount of labeled data and professional knowledge, and have limited generalization capabilities, making it difficult to adapt to different acquisition equipment and environments.
The PCA-Transformer model is used to filter out noise, combine independent component analysis to extract local and time-frequency features, and the fusion recognition results of the first neural network, the second neural network and the Transformer model are achieved to achieve efficient identification of heart sound signals.
It improves the accuracy and robustness of heart sound signal recognition, can adapt to complex environments and non-Gaussian noise, and has wide application prospects.
Smart Images

Figure CN120375867A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, system, device and medium for intelligent heart sound recognition. Background Art
[0002] Heart diseases are one of the main causes of death globally. Early detection and timely treatment are crucial for improving the survival rate and quality of life of patients. As a vibration signal generated by the mechanical activity of the heart, heart sound signals contain rich physiological and pathological information and are an important basis for evaluating heart function.
[0003] In recent years, with the development of computer technology and signal processing technology, automatic analysis methods based on heart sound signals have gradually become a research hotspot. Existing methods mainly include feature extraction methods based on signal processing, but this method usually requires relying on experts to manually design features and is difficult to capture relatively complex heart sound signals.
[0004] In addition, existing methods include methods for analyzing heart sound signals based on deep learning. However, deep learning models usually require a large amount of labeled data for training, and high-quality data labeling requires professional medical knowledge and experience, resulting in high labeling costs. At the same time, the generalization ability of existing deep learning models on different datasets is limited and it is difficult to adapt to heart sound signals under different acquisition devices, different populations and different environments. Summary of the Invention
[0005] Based on the above deficiencies of the prior art, the present invention provides a method, system, device and medium for intelligent heart sound recognition, which can efficiently and accurately recognize heart sound signals.
[0006] To solve the above technical problems, the first aspect of the present invention discloses a method for intelligent heart sound recognition, including:
[0007] Obtain a heart sound signal and preprocess the heart sound signal;
[0008] Filter the noise of the heart sound signal through a PCA-Transformer model;
[0009] Extract the local features of the heart sound signal through independent component analysis, and combine short-time Fourier transform to extract the time-frequency features of the heart sound signal;
[0010] Input the heart sound signal features into a heart sound recognition model, which includes a first neural network, a second neural network, and a Transformer model; the first neural network is used to recognize the time-frequency features to obtain a first recognition result; the second neural network is used to process the local features to obtain a second recognition result; the Transformer model is used to process the heart sound signal to obtain a third recognition result; fuse the first recognition result, the second recognition result, and the third recognition result to obtain a heart sound recognition result.
[0011] In some embodiments, the PCA-Transformer model includes a Transformer encoder, a learnable PCA projection layer, and a reconstruction decoder;
[0012] Filter the noise of the heart sound signal through the PCA-Transformer model, including:
[0013] Convert the heart sound signal into a time-frequency spectrum through STFT by the Transformer encoder, extract the mel-sscale energy features as positional encoding, and encode through the encoder to obtain an encoded signal;
[0014] Input the encoded signal into the PCA projection layer to dynamically generate principal component basis vectors; introduce gated weighted projection to dynamically generate retention thresholds, and adaptively select the number of principal components according to the heart sound characteristics;
[0015] Extract the spectrum of the heart sound signal based on the reconstruction decoder, perform phase correction through a neural network, and perform inverse STFT on the corrected phase and spectrum to obtain a heart sound signal with noise filtered out.
[0016] In some embodiments, the PCA-Transformer model uses heart sound signals without noise and heart sound signals with added noise as the training set, and is trained with the goal of minimizing the loss function, which includes a reconstruction loss function and an orthogonal regularization loss function;
[0017] The reconstruction loss function is:
[0018]
[0019] where X is the original heart sound signal, is the reconstructed heart sound signal; STFT(X) is the frequency domain representation after performing short-time Fourier transform on the original heart sound signal, is the frequency domain representation after performing short-time Fourier transform on the reconstructed heart sound signal; α, β are weight coefficients used to balance the contributions of time domain and frequency domain losses; ||·||2 is the L2 norm used to calculate the mean square error of the time domain signal, ||·|| Fis the Frobenius norm, which is used to calculate the difference of matrices;
[0020] The orthogonal regularization loss function is:
[0021]
[0022] where W is the matrix of principal component basis vectors in the learnable PCA projection layer, with dimensions d×k, d being the original feature dimension and k being the dimension after dimensionality reduction; w T is the transpose matrix of W, I is the identity matrix, and ||·|| F is the Frobenius norm, which is used to calculate the difference of matrices.
[0023] In some embodiments, local features of the heart sound signal are extracted by independent component analysis, including:
[0024] Centering and whitening the heart sound signal, and retaining key frequency bands according to the energy distribution of the heart sound frequency segment;
[0025] Adopting the FastICA algorithm of maximizing negative entropy, adaptively selecting non-linear functions and matching the non-Gaussianity of heart sounds; decomposing the components of different independent sources in the heart sound signal by iteratively optimizing the demixing matrix;
[0026] Selecting discriminative independent components from the components and extracting local features in the independent components;
[0027] Combining short-time Fourier transform and mel-frequency cepstrum to extract the time-frequency features of the heart sound signal, including:
[0028] Pre-emphasizing, framing, and windowing the heart sound signal;
[0029] Converting each frame of the heart sound signal into the frequency domain through short-time Fourier transform to obtain a time-frequency spectrum matrix;
[0030] Filtering the time-frequency spectrum matrix through a mel filter bank to obtain a mel spectrum;
[0031] Taking the logarithm and discrete cosine transform of the mel spectrum to obtain the mel-frequency cepstrum, and using the mel-frequency cepstral coefficients and their first-order and second-order differences as the time-frequency features of the heart sound signal.
[0032] In some embodiments, the first neural network takes the time-frequency features as input, extracts features of different modalities through convolutional layers, and fuses the time-frequency features based on a cross-modal attention mechanism;
[0033] A backbone network is constructed by using multi-layer heart sound adaptive convolution kernels. The backbone network generates dynamic kernels based on MCFF features, performs feature fusion through a hierarchical feature pyramid, and obtains feature vectors.
[0034] The feature vectors are mapped through a fully connected layer to obtain a first recognition result.
[0035] In some embodiments, the second neural network takes each of the independent components as input, and through a convolutional layer and a GRU encoder, obtains independent coding results for each of the independent components.
[0036] The independent coding results are stacked into a three-dimensional tensor, and the inter-component correlation is calculated through multi-head attention to obtain the time-series features after component fusion.
[0037] The time-series features are input into a time-series attention layer to locate the significant regions of pathological features and output the focused global feature vectors.
[0038] The global feature vectors are mapped through a fully connected layer to obtain a second recognition result.
[0039] In some embodiments, the Transformer model includes a multi-scale feature extraction layer, a period perception encoding layer, a period perception encoding layer adaptive feature decoupling layer, and a multi-task output layer. The multi-scale feature extraction layer performs short-term, medium-term, and long-term analysis and feature fusion on the heart sound features. The period perception encoding layer predicts the probability of each time point belonging to the key stage of the heartbeat and dynamically adjusts the position encoding. The adaptive decoupling layer estimates the noise interference intensity at each time point, generates a noise confidence map, and performs feature purification and pathological feature enhancement. The multi-task output layer is used to output a third recognition result, and the third recognition result includes the probability of the pathological type of the heart sound signal and the auxiliary beat evaluation result.
[0040] In a second aspect, a heart sound intelligent recognition system is disclosed, including:
[0041] A heart sound acquisition module that acquires a heart sound signal and preprocesses the heart sound signal.
[0042] A noise filtering module that filters the noise of the heart sound signal through a PCA-Transformer model.
[0043] A feature extraction module that extracts local features of the heart sound signal through independent component analysis and combines short-time Fourier transform to extract time-frequency features of the heart sound signal.
[0044] The heart sound recognition module inputs the heart sound signal features into a heart sound recognition model, which includes a first neural network, a second neural network, and a Transformer model. The first neural network is used to recognize the time-frequency features and obtain a first recognition result. The second neural network is used to process the local features and obtain a second recognition result. The Transformer model is used to process the heart sound signal and obtain a third recognition result. The first recognition result, the second recognition result, and the third recognition result are fused to obtain a heart sound recognition result.
[0045] In a third aspect, a computer device is disclosed, including: a processor and a memory. Among them, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of a heart sound intelligent recognition method as described in any one of the above.
[0046] In a fourth aspect, a computer storage medium is disclosed, on which a computer program is stored. When the computer program is executed by a processor, it implements a heart sound intelligent recognition method as described in any one of the above.
[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0048] By combining the PCA-Transformer model, independent component analysis (ICA), and short-time Fourier transform (STFT), the present invention realizes efficient noise filtering and feature extraction of heart sound signals. By inputting the extracted features into a heart sound recognition model including a first neural network, a second neural network, and a Transformer model, the time-frequency features, local features, and original heart sound signals are processed respectively, and finally multiple recognition results are fused, significantly improving the accuracy and robustness of heart sound recognition. The present invention can not only adapt to complex heart sound signal environments, but also effectively cope with non-Gaussian noise and sudden interferences, and has broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 It is a schematic flow chart of a heart sound intelligent recognition method provided by the present invention;
[0050] Figure 2 It is a schematic flow chart of step S2 of a heart sound intelligent recognition method provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] For better understanding and implementation, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0052] The terms "including" and "having" and any variations thereof in the embodiments of the present invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules does not necessarily have to be limited to those clearly listed steps or modules, but may include other steps or modules not clearly listed or inherent to these processes, methods, products, or devices.
[0053] An embodiment of the present invention discloses a method for intelligent heart sound recognition, which can efficiently and accurately recognize heart sound signals.
[0054] As Figure 1 shown, this method includes:
[0055] Step S1, obtain a heart sound signal and preprocess the heart sound signal;
[0056] The heart sound signal can be collected by a collection device, and the collection device includes an electronic stethoscope and other devices that can collect heart sound signals. Different collection devices may use different sampling rates, such as 4000Hz, 8000Hz, etc. In order to unify the data format, all data needs to be converted to the same sampling rate. For example, the librosa.resample function can be used to resample the audio data to the target sampling rate. In order to eliminate the differences brought by different collection devices and environments, the data is normalized. For example, each data point can be divided by the amplitude of the maximum absolute value in the audio data to scale the data range to between [-1, 1].
[0057] The original heart sound signal is normalized and stored in the data storage module. The original heart sound signal is generally stored in the form of an audio file, including formats such as WAV and MP3. In order to ensure the quality and compatibility of the data, the lossless compressed WAV format can be used.
[0058] When identification is required, preprocessing is performed on the original heart sound signal. The preprocessing includes operations such as signal framing, windowing, DC component removal, and pre-emphasis, which improve the quality of the heart sound signal and prepare for subsequent feature extraction and model training. The heart sound signal is a non-stationary signal, and its statistical characteristics change over time. To analyze the time-varying characteristics of the heart sound signal, the signal is divided into multiple short-time frames. Specifically, signal framing means dividing the continuous heart sound signal into multiple short-time frames for subsequent processing. Usually, the length of each frame is 20 - 40 milliseconds, and the frame shift is less than the frame length. Framing is performed through the following formula:
[0059] x_i(n) = x(n + i*N_s), 0 ≤ n < N, 0 ≤ i < L
[0060] where x_i(n) represents the heart sound signal of the i-th frame, x(n) is the original heart sound signal, N is the frame length, N_s is the frame shift, and L is the total number of frames.
[0061] Windowing is to apply a window function, such as a Hamming window, a Hanning window, etc., to each frame of the signal to reduce spectral leakage. The formula for the Hamming window is:
[0062] w(n) = 0.54 - 0.46 * cos(2πn / (N - 1)), 0 ≤ n ≤ N - 1
[0063] The windowed signal can be expressed as: x'_i(n) = x_i(n) * w(n)
[0064] DC component removal is to remove the DC component in the heart sound signal, which can eliminate baseline drift. Through the following formula:
[0065] x”(n) = x(n) - mean(x)
[0066] where x”(n) is the heart sound signal after DC component removal, and mean(x) is the mean value of the heart sound signal.
[0067] Pre-emphasis is to enhance the high-frequency part of the heart sound signal through a first-order high-pass filter to balance the spectrum. It is achieved through the following formula:
[0068] y(n) = x(n) - α * x(n - 1)
[0069] where y(n) is the signal after pre-emphasis, and α is the pre-emphasis coefficient, usually taking a value of 0.97.
[0070] Step S2, filter the noise of the heart sound signal through the PCA-Transformer model.
[0071] The PCA-Transformer model (PCA, principal components analysis; Transformer, self-attention mechanism model) includes a Transformer encoder, a learnable PCA projection layer, and a reconstruction decoder. Among them, the Transformer encoder is used for the long-term dependence and temporal patterns of heart sound signals, the learnable PCA projection layer is used to dynamically generate principal component basis vectors to achieve non-linear dimensionality reduction and noise suppression, and the reconstruction decoder is used to map the low-dimensional representation back to the original space and optimize the output heart sound signal through a complex spectrum correction module.
[0072] Specifically, as Figure 2 shown, filtering the noise of the heart sound signal by the PCA-Transformer model includes:
[0073] Step S21: Convert the heart sound signal into a time-frequency spectrum through STFT by the Transformer encoder, extract the mel-scale energy feature as the position encoding, and encode it through the encoder to obtain an encoded signal.
[0074] Frame the heart sound signal, with the length of each frame being a fixed value, which can be 256 sampling points or 512 sampling points, etc. Perform short-time Fourier transform (STFT) on each frame of the heart sound signal to extract the time-frequency spectrum, and extract the mel-scale energy feature from the time-frequency spectrum as the position encoding. Insert a multi-scale dilated convolution module in front of the self-attention layer of the Transformer encoder to capture time-domain features of different granularities. The dilation rate of the multi-scale dilated convolution module is set to [1, 2, 4], corresponding to different time scales respectively. Concatenate the output features of the multi-scale dilated convolution module with the original heart sound signal as the input of the Transformer encoder.
[0075] Use multiple layers of Transformer encoders, each layer containing a multi-head self-attention mechanism and a feed-forward neural network. In the self-attention mechanism, use the Mel-scale energy feature as the position encoding to enhance the perception ability of the heart sound harmonic structure. The Transformer encoder outputs high-dimensional temporal features, that is, the encoded signal, with dimensions including [batch_size, sequence_length, feature_dim] (batch size, sequence length, feature dimension).
[0076] Step S22: Input the encoded signal into the PCA projection layer to dynamically generate principal component basis vectors; introduce gated weighted projection to dynamically generate a retention threshold, and adaptively select the number of principal components according to the heart sound characteristics.
[0077] The learnable PCA projection layer can dynamically generate the principal component basis vectors to achieve non-linear dimensionality reduction and noise suppression. A fully connected neural network (FCN) is used to generate the principal component basis vectors V, replacing the traditional fixed PCA basis. The principal component vector basis is a set of orthogonal vectors extracted from high-dimensional data by PCA, and these vectors are sorted according to the magnitude of the data variance. Generally, it includes the first principal component, the second principal component, and the k-th principal component. The first principal component is the direction with the largest data variance, the second principal component is the direction orthogonal to the first principal component and with the second largest variance, and the k-th principal component is the direction orthogonal to the previous k-1 principal components and with the k-th largest variance. The principal component basis vectors form a new coordinate system, and projecting into this coordinate system can achieve dimensionality reduction or feature extraction. The principal component basis vectors are used to extract the main features of the data, remove redundant information, and help separate the heart sound signal and noise.
[0078] The Transformer encoder outputs high-dimensional time-series features, and heart sound features for generating the retention threshold are extracted from the high-dimensional time-series features. The high-dimensional time-series features are input into a fully connected neural network (FCN) or a lightweight convolutional network, and the output is the retention probability p i ∈[0,1]. The specific formula is:
[0079] p i =σ(W g *z i +b g )
[0080] where z i is the i-th principal component of the input, W g and b g are the weights and biases of the gating network, and σ is the activation function that maps the output to the range of [0,1]. The retention threshold is determined by the product of the principal component variance and the maximum variance. Each principal component basis vector is weighted according to the retention probability to retain the important components of the principal component basis vector and suppress the noise components. The input high-dimensional time-series features are projected into the weighted principal component vector basis to obtain a low-dimensional representation. According to the retention probability and the retention threshold, the number of principal components is dynamically selected. When the retention probability is greater than the retention threshold, this principal component is retained.
[0081] Step S23: Based on the reconstruction decoder, extract the spectrum of the heart sound signal, perform phase correction through a neural network, and perform inverse STFT on the corrected phase and the spectrum to obtain the heart sound signal with noise filtered out.
[0082] The reconstruction decoder maps the low-dimensional representation back to the original space and optimizes the output heart sound signal through the complex spectrum correction module. A fully connected neural network (FCN) is used to map the low-dimensional representation back to the original space. The noisy signal is subjected to STFT to extract the spectrum, which includes amplitude and phase. Phase correction is performed through a convolutional neural network, and the corrected phase is combined with the amplitude of the target signal for inverse STFT to restore the time-domain signal, that is, the heart sound signal with noise filtered out.
[0083] During the training process of the PCA-Transformer model, noise-free heart sounds and heart sounds with added noise are used as data pairs for training. The noise includes in-vivo noise such as intestinal peristalsis, vascular murmurs, and external limb noise such as stethoscope friction and breath sounds. The learning rate is dynamically adjusted using CosineAnnealingLR. After multiple rounds of training, the model with the minimum loss is found. During the training process, the loss function includes a reconstruction loss function and an orthogonal regularization loss function;
[0084] The reconstruction loss function L recon is:
[0085]
[0086] where X is the original heart sound signal, is the reconstructed heart sound signal; STFT(X) is the frequency-domain representation after performing short-time Fourier transform on the original heart sound signal, is the frequency-domain representation after performing short-time Fourier transform on the reconstructed heart sound signal; α and β are weight coefficients used to balance the contributions of time-domain and frequency-domain losses; ||·||2 is the L2 norm used to calculate the mean square error of the time-domain signal, and ||·|| F is the Frobenius norm used to calculate the difference of matrices.
[0087] The orthogonal regularization loss function constrains the orthogonality of the learnable PCA basis vectors W to ensure that they satisfy the orthogonal condition, specifically:
[0088]
[0089] where W is the matrix of principal component basis vectors in the learnable PCA projection layer, with dimensions d×k, d is the original feature dimension, and k is the dimension after dimensionality reduction; W T is the transpose matrix of W, I is the identity matrix, and ||·|| F is the Frobenius norm used to calculate the difference of matrices.
[0090] Step S3: Extract the local features of the heart sound signal through independent component analysis, and combine short-time Fourier transform to extract the time-frequency features of the heart sound signal.
[0091] Independent Component Analysis (ICA) is a signal separation technique that separates independent source signals from mixed signals. Specifically, local features of the heart sound signal are extracted through Independent Component Analysis, including:
[0092] Center and whiten the heart sound signal, and retain the key frequency bands according to the energy distribution of the heart sound frequency band;
[0093] Adopt the FastICA algorithm that maximizes negentropy and adaptively selects a non-linear function.
[0094] Decompose the components of different independent sources in the heart sound signal by iteratively optimizing the demixing matrix;
[0095] Screen out the discriminative independent components from the components and extract the local features in the independent components.
[0096] Center and whiten the heart sound signal with noise removed to eliminate the correlation between channels. According to the frequency band energy distribution of the heart sound signal, that is, the principal components obtained by the principal component analysis method, retain the key frequency bands to suppress environmental noise. Adopt the FastICA algorithm that maximizes negentropy and adaptively selects a non-linear function to obtain the components representing different independent sources in the heart sound signal. Each component may correspond to different cardiac activities or other physiological processes. Select the independent components that are sensitive to the heart health status and have discriminability according to the waveform and frequency characteristics of the components, and separate the independent components such as the heart sound cycles S1, S2, and noise. According to the time-domain periodicity (such as the S1-S2 interval) and frequency-domain energy concentration (30 - 150 Hz) of the heart sound events, screen out the physiologically relevant components and extract the statistical features such as mean, variance, kurtosis, skewness, etc. In this application, the FastICA algorithm is applied to each denoised heart sound signal to extract 4 independent components. Calculate the mean, variance, kurtosis, and skewness for each independent component to obtain 16 local features.
[0097] Mel Frequency Cepstral Coefficient (MFCC) is an audio feature that combines the auditory characteristics of the human ear. Combine the short-time Fourier transform and the Mel frequency cepstrum series to extract the time-frequency features of the heart sound signal, including:
[0098] Pre-emphasize, frame, and window the heart sound signal;
[0099] Convert each frame of the heart sound signal into the frequency domain through the short-time Fourier transform to obtain the time-frequency spectrum matrix;
[0100] Filter the time-frequency spectrum matrix through the Mel filter bank to obtain the Mel spectrum;
[0101] Take the logarithm and discrete cosine transform of the Mel spectrum to obtain the Mel-frequency cepstrum. Use the Mel-frequency cepstral coefficients, their first-order differences, and second-order differences as the time-frequency features of the heart sound signal.
[0102] Pre-emphasize the heart sound signal to enhance the high-frequency part of the signal. Perform short-time Fourier transform on each independent component using a Hanning window with a window length of 25 milliseconds and a frame shift of 10 milliseconds. Calculate the spectral energy for each time frame to obtain a time-frequency spectrum matrix, i.e., the STFT feature. Apply a Mel filter bank to the time-frequency spectrum matrix obtained by STFT to obtain the Mel spectrum. Take the logarithm of the Mel spectrum and perform discrete cosine transform to obtain 13 Mel-frequency cepstral coefficients (MFCCs). Calculate the 13 Mel-frequency cepstral coefficients (MFCCs), their first-order differences, and second-order differences for each time frame to obtain 39 MFCC features.
[0103] Step S4: Input the heart sound signal features into the heart sound recognition model. The heart sound recognition model includes a first neural network, a second neural network, and a Transformer model. The first neural network is used to recognize the time-frequency features to obtain a first recognition result. The second neural network is used to process the local features to obtain a second recognition result. The Transformer model is used to process the heart sound signal to obtain a third recognition result. Fuse the first recognition result, the second recognition result, and the third recognition result to obtain the heart sound recognition result.
[0104] Train the heart sound recognition model using a training dataset. The training dataset includes feature vectors and labels of multiple heart sound samples. Specifically, the feature vectors include local features and time-frequency features. The time-frequency features include STFT features and MFCC features. The labels include normal or abnormal. Divide the training dataset into 5 folds for 5-fold cross-validation. The training dataset is obtained by collecting heart sound signals using an electronic stethoscope with a sampling rate of 4000 Hz and a duration of 5 seconds for each sample. 1000 samples are collected, including 500 normal samples and 500 abnormal samples, such as samples of patients with heart valve diseases.
[0105] Train a primary learner using a 4-fold training dataset. The primary learner includes a first neural network, a second neural network, and a Transformer model. The first neural network uses a classification network of multi-modal CNN. The first neural network includes dual-branch feature input mixing, multi-scale feature alignment, multi-layer heart sound adaptive convolution, and a fully connected layer. Using the time-frequency features as input, the input feature of the first input branch is 13-dimensional MFCC features with a feature dimension of (Batch, 1, 13), which are expanded to a high-dimensional representation through 1D convolution. The second input branch is the STFT feature with a feature dimension of (Batch, 1, 64, 128).
[0106] Extract the features of each modality through the convolutional layer, and fuse the time-frequency features based on the cross-modal attention mechanism; use a multi-layer heart sound adaptive convolutional kernel to form the backbone network. The backbone network generates a dynamic kernel according to the MCFF features, and performs feature fusion through the hierarchical feature pyramid to obtain a feature vector; map the feature vector through the fully connected layer to obtain the first recognition result, and use Dropout(0.5) to reduce overfitting and enhance the generalization ability of the model.
[0107] The second neural network is an RNN network, which uses a multi-channel gated recurrent unit (MC-GRU) combined with a cross-component attention mechanism to classify heart sound pathology by parallel processing of multiple independent components decomposed by ICA and dynamically fusing local features. The second neural network takes each of the independent components as input, and through the convolutional layer and the GRU encoder, obtains the independent coding result of each independent component; stacks the independent coding results into a three-dimensional tensor, calculates the correlation degree between components through multi-head attention, and obtains the time series features after component fusion; inputs the time series features into the time series attention layer to locate the significant region of pathological features and output the focused global feature vector; maps the global feature vector through the fully connected layer to obtain the second recognition result.
[0108] The Transformer model uses a Transformer architecture enhanced with physiological features to process heart sound signals with noise. It includes a multi-scale feature extraction layer, a period perception encoding layer, an adaptive feature decoupling layer, and a multi-task output layer. The multi-scale feature extraction layer performs short-term, medium-term, and long-term analysis and feature fusion. Short-term analysis uses a fine-grained convolutional kernel (15 milliseconds level) to capture the transient features of heart sound segments, such as the short sound of heart valve closure; medium-term analysis refers to using medium-scale convolution (about 0.5 second window) to identify the complete cycle pattern of a single heartbeat; long-term analysis is to use large-scale convolution (window above 2 seconds) to perceive the overall regularity of heart rhythm. Concatenate the features with three different time resolutions to form a comprehensive feature containing local details and global rhythm.
[0109] The period perception encoding layer predicts the probability that each time point belongs to the key stage of the heartbeat (such as S1 / S2) through a lightweight network to generate a heat map; dynamically adjusts the position encoding according to the beat probability, so that the model adopts a differential feature processing strategy at different heartbeat phases. Allows high-probability regions to pay attention to each other, which not only reduces the computational amount but also strengthens the key signals.
[0110] The adaptive decoupling layer estimates the noise interference intensity at each time point through a branch network to generate a noise confidence map of 0-1. Use the noise mask to reverse-weight the original features to suppress the signal influence in high-noise regions. Perform secondary convolution on the purified features to amplify the specific patterns related to diseases.
[0111] The multi-task output layer is used to output a third recognition result, which includes the probability of the pathological type of the heart sound signal and the auxiliary beat evaluation result. The main task is a classification task, which performs weighted averaging on the global features and outputs the probability of the pathological type. The auxiliary task is beat evaluation, which additionally predicts the stability of heart sound beats to assist in judging abnormal conditions such as arrhythmia.
[0112] The secondary learner then uses a logistic regression model. During the training process, the trained primary learner is used to identify the 1-fold data to obtain the recognition result. All the prediction results are concatenated as the input of the secondary learner for training. The recognition results of the above first neural network, second neural network, and Transformer model are used as inputs for weighted fusion to obtain the final heart sound recognition result. The calculation formula of the logistic regression model is
[0113] P(y=1|x)=σ(w^T*x+b)
[0114] where x is the prediction result vector of the primary learner, w is the weight vector, b is the bias term, and σ is the Sigmoid function. The weight vector w is used to weight the prediction results of the primary learner. Each weight vector w represents the influence degree of the prediction result of the primary learner on the final output. The weight vector is determined by training to minimize the loss function. The bias term b adjusts the position of the decision boundary to improve the flexibility of the model. It is independent of the prediction results of each primary learner, ensuring that the model can produce reasonable outputs even if all inputs are 0. The bias term b is adjusted in a similar way to the optimization process of the weight vector, also by minimizing the loss function. The importance of the prediction results of different primary learners in the logistic regression is reflected by the weight vector w. If the prediction results of some primary learners are more reliable, their corresponding weight values will be larger.
[0115] The loss function uses cross-entropy loss, and the calculation formula is:
[0116]
[0117] y k is the one-hot encoded vector of the true label. For the sample x, if its true label is the category k, then y k =1, otherwise y k =0.
[0118] Evaluation metrics for the heart sound recognition model are accuracy, precision, recall, F1-score, ROC curve and AUC value, PR curve and AP value. Accuracy is the ratio of the number of samples predicted correctly by the model to the total number of samples. Precision refers to the proportion of samples that are truly positive among the samples predicted as positive by the model. Recall is the proportion of samples predicted as positive by the model among all samples that are truly positive; F1-score is the harmonic mean of precision and recall. The ROC curve is plotted with the false positive rate (FPR) on the x-axis and the true positive rate (TPR) on the y-axis. The AUC value is the area under the ROC curve, and its value range is [0,1]. The larger the AUC value, the better the performance of the model. The PR curve is plotted with recall on the x-axis and precision on the y-axis. The AP value is the area under the PR curve.
[0119] To avoid missed diagnoses, recall is a key metric. A high recall rate indicates that the model can identify more abnormal heart sounds and reduce missed diagnoses. At the same time, the F1-score is the harmonic mean of precision and recall, comprehensively considering the accuracy and coverage of positive class predictions. Therefore, when evaluating the model, mainly consider models with high F1-score and recall rate, and high AUC value on the graph.
[0120] The heart sound recognition model after training is used to identify heart sound signals, obtain heart sound recognition results, realize intelligent classification of heart signals, and improve the accuracy of heart sound classification. This application can be applied to institutions such as hospitals, physical examination centers, and community health service centers to assist doctors in early screening and diagnosis of heart diseases, improve the detection rate and treatment rate of heart diseases, reduce medical costs, and has important social significance and economic value.
[0121] Based on the same inventive concept, this application also provides a heart sound intelligent recognition system, including:
[0122] A heart sound acquisition module that acquires heart sound signals and preprocesses the heart sound signals;
[0123] A noise filtering module that filters the noise of the heart sound signals through a PCA-Transformer model;
[0124] A feature extraction module that extracts local features of the heart sound signals through independent component analysis and combines short-time Fourier transform to extract time-frequency features of the heart sound signals;
[0125] The heart sound recognition module inputs the heart sound signal features into a heart sound recognition model, which includes a first neural network, a second neural network, and a Transformer model. The first neural network is used to recognize the time-frequency features and obtain a first recognition result. The second neural network is used to process the local features and obtain a second recognition result. The Transformer model is used to process the heart sound signal and obtain a third recognition result. The first recognition result, the second recognition result, and the third recognition result are fused to obtain a heart sound recognition result.
[0126] Based on the same inventive concept, the present invention also provides a computer device, including: a processor and a memory. Wherein, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of the above heart sound intelligent recognition method.
[0127] The processing method of the computer device can refer to the description of the above method and will not be elaborated here.
[0128] The embodiment of the present application also provides a non-transitory machine-readable storage medium, on which an executable program is stored. When the executable program runs on a microprocessor, the processor is enabled to execute a heart sound intelligent recognition method provided in the above embodiment.
[0129] The embodiment of the present invention discloses a computer-readable storage medium, which stores a computer program for electronic data exchange. Wherein, the computer program enables a computer to execute the described heart sound intelligent recognition method.
[0130] The embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute the described heart sound intelligent recognition method.
[0131] The above-described embodiments are merely illustrative. The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0132] Through the above specific description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, which includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disc memories, tape memories, or any other computer-readable medium that can be used to carry or store data.
[0133] Finally, it should be noted that: The embodiments disclosed in the present invention are only the preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, rather than to limit them; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: They can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An intelligent heart sound recognition method, characterized in that, Including: Obtain a heart sound signal and preprocess the heart sound signal; Filter the noise of the heart sound signal through a PCA-Transformer model; Extract the local features of the heart sound signal through independent component analysis, and extract the time-frequency features of the heart sound signal in combination with short-time Fourier transform; Input the heart sound signal features into a heart sound recognition model, where the heart sound recognition model includes a first neural network, a second neural network, and a Transformer model; the first neural network is used to recognize the time-frequency features to obtain a first recognition result; the second neural network is used to process the local features to obtain a second recognition result; the Transformer model is used to process the heart sound signal to obtain a third recognition result; fuse the first recognition result, the second recognition result, and the third recognition result to obtain a heart sound recognition result.
2. The intelligent heart sound recognition method according to claim 1, characterized in that The PCA-Transformer model includes a Transformer encoder, a learnable PCA projection layer, and a reconstruction decoder; Filtering the noise of the heart sound signal through a PCA-Transformer model includes: Convert the heart sound signal into a time-frequency spectrum through STFT by the Transformer encoder, extract the mel-sscale energy feature as the position encoding, and encode through the encoder to obtain an encoded signal; Input the encoded signal into the PCA projection layer to dynamically generate principal component basis vectors; Introduce gated weighted projection, dynamically generate a retention threshold, and adaptively select the number of principal components according to the heart sound characteristics; Extract the spectrum of the heart sound signal based on the reconstruction decoder, perform phase correction through a neural network, and perform inverse STFT on the corrected phase and the spectrum to obtain a heart sound signal with noise filtered out.
3. The intelligent heart sound recognition method according to claim 2, characterized in that The PCA-Transformer model uses a heart sound signal without noise and a heart sound signal with added noise as the training set, and is trained with the goal of minimizing the loss function, where the loss function includes a reconstruction loss function and an orthogonal regularization loss function; The reconstruction loss function is: where X is the original heart sound signal, is the reconstructed heart sound signal; STFT(X) is the frequency domain representation after performing short-time Fourier transform on the original heart sound signal, is the frequency domain representation after performing short-time Fourier transform on the reconstructed heart sound signal; α and β are weight coefficients used to balance the contributions of time domain and frequency domain losses; ||·||2 is the L2 norm used to calculate the mean square error of the time domain signal, and ||·|| F is the Frobenius norm used to calculate the difference of matrices; The orthogonal regularization loss function is: Among them, W is the matrix of principal component basis vectors in the learnable PCA projection layer, with dimensions d×k, where d is the original feature dimension and k is the dimension after dimensionality reduction; W T is the transpose matrix of W, I is the identity matrix, and ||·|| F is the Frobenius norm, which is used to calculate the difference between matrices.
4. The method for intelligent heart sound recognition according to claim 3, characterized in that Extracting the local features of the heart sound signal through independent component analysis includes: Center and whiten the heart sound signal, and retain the key frequency bands according to the energy distribution of the heart sound segment; Adopt the FastICA algorithm of maximizing negative entropy, adaptively select the nonlinear function and match the non-Gaussianity of the heart sound; decompose the components of different independent sources in the heart sound signal by iteratively optimizing the demixing matrix; Screen out the discriminative independent components from the components and extract the local features in the independent components; Extracting the time-frequency features of the heart sound signal in combination with short-time Fourier transform and mel-frequency cepstrum includes: Perform pre-emphasis, framing, and windowing on the heart sound signal; Convert each frame of the heart sound signal into the frequency domain through short-time Fourier transform to obtain a time-frequency spectrum matrix; Filter the time-frequency spectrum matrix through a mel filter bank to obtain a mel spectrum; Take the logarithm and discrete cosine transform of the Mel spectrum to obtain the Mel frequency cepstrum. Use the Mel frequency cepstral coefficients, their first-order differences, and second-order differences as the time-frequency features of the heart sound signal.
5. The intelligent heart sound recognition method according to claim 4, wherein The first neural network takes the time-frequency features as input, extracts features of each modality through convolutional layers, and fuses the time-frequency features based on the cross-modal attention mechanism. Adopt a multi-layer heart sound adaptive convolutional kernel to form the backbone network. The backbone network generates dynamic kernels according to the MCFF features, and performs feature fusion through a hierarchical feature pyramid to obtain feature vectors. Map the feature vectors through a fully connected layer to obtain the first recognition result.
6. The intelligent heart sound recognition method according to claim 4, wherein The second neural network takes each independent component as input, and through convolutional layers and a GRU encoder, obtains the independent coding results of each independent component. Stack the independent coding results into a three-dimensional tensor, calculate the correlation degree between components through multi-head attention, and obtain the time-series features after component fusion. Input the time-series features into the time-series attention layer to locate the significant regions of pathological features and output the focused global feature vectors. Map the global feature vectors through a fully connected layer to obtain the second recognition result.
7. A heart sound intelligent recognition method according to claim 4, characterized in that, The Transformer model includes a multi-scale feature extraction layer, a period perception coding layer, an adaptive feature decoupling layer for the period perception coding layer, and a multi-task output layer. The multi-scale feature extraction layer performs short-term, medium-term, and long-term analysis and feature fusion on the heart sound features. The period perception coding layer predicts the probability of each time point belonging to the key stage of the heartbeat and dynamically adjusts the position coding. The adaptive decoupling layer estimates the noise interference intensity at each time point, generates a noise confidence map, and performs feature purification and pathological feature enhancement. The multi-task output layer is used to output the third recognition result, which includes the probability of the pathological type of the heart sound signal and the auxiliary beat evaluation result.
8. An intelligent heart sound recognition system, characterized in that, Comprises: A heart sound acquisition module that acquires a heart sound signal and preprocesses the heart sound signal. A noise filtering module that filters the noise of the heart sound signal through a PCA-Transformer model. A feature extraction module that extracts the local features of the heart sound signal through independent component analysis and combines the short-time Fourier transform to extract the time-frequency features of the heart sound signal. A heart sound recognition module that inputs the heart sound signal features into a heart sound recognition model. The heart sound recognition model includes a first neural network, a second neural network, and a Transformer model. The first neural network is used to recognize the time-frequency features and obtain the first recognition result. The second neural network is used to process the local features and obtain the second recognition result. The Transformer model is used to process the heart sound signal and obtain the third recognition result. The first recognition result, the second recognition result, and the third recognition result are fused to obtain the heart sound recognition result.
9. A computer device, characterized in that, Comprises: A processor and a memory. Among them, the memory stores a computer program, and the computer program is adapted to be loaded and executed by the processor to perform the steps of a heart sound intelligent recognition method according to any one of claims 1-7.
10. A computer storage medium, characterized in that, A computer program is stored thereon. When the computer program is executed by a processor, the steps of a heart sound intelligent recognition method according to any one of claims 1-7 are implemented.
Citation Information
Cited By
Risk prediction method and device based on heart sound and electrocardio, equipment and medium
CN122208159A