Explanatable deep learning method for non-invasive detection of pulmonary arterial hypertension from heart sound
By analyzing heart sound recording using a parameterized deep neural network, separating S2 sounds into A2 and P2 components, and combining post-interpretation methods, the problems of high cost, strong invasiveness and insufficient accuracy of detecting pulmonary hypertension in the prior art are solved, and high accuracy, low cost and non-invasive detection effects are achieved.
Patent Information
- Application Number
- CN202380063247.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-02
- Filing Date
- 2023-09-01
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art has problems such as high cost, strong invasiveness, insufficient accuracy and lack of explanatory detection of pulmonary hypertension (PH), especially in low- and middle-income areas.
A parameterized deep neural network was used to analyze digital heart sound recordings, and the S2 sound was separated into the aortic (A2) and pulmonary (P2) components through pre-processing steps, and combined with the post-hoc interpretation method, a prediction model with high accuracy and low resource demand was provided.
High-accuracy pulmonary hypertension detection is achieved, and the area under the ROC curve reaches 0.95, which is better than the Gaussian hybrid model of the prior art, and is characterized by low cost and non-invasiveness.
Smart Images

Figure CN119947653A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an explainable deep learning method for non-invasive detection of pulmonary hypertension from heart sounds. Background Art
[0002] Pulmonary hypertension (PH) is an underrecognized disease with an unmet need for diagnostic and treatment recommendations in low- and middle-income settings [1]. PH is associated with high mortality and early detection in screening programs improves outcomes. Existing PH detection tools are not adequately tailored to the needs of low- and middle-income settings.
[0003] Right heart catheterization is the gold standard for PH detection, but it is costly, invasive, not suitable for screening programs, and requires a specialized team of trained clinicians.
[0004] Implantable pressure sensors are placed in the pulmonary arteries, which are extremely expensive and highly invasive. They provide the best blood pressure readings, but are only available to a small number of subjects in wealthy countries.
[0005] Doppler echocardiography is widely used in clinical screening for PH, but the noisy nature of its measurements requires additional methods to improve reliability [2, 3]. It only provides predictions of pulmonary artery pressure in some subjects. Ultrasound techniques also require highly trained technicians and expensive equipment [4].
[0006] Other tests that can aid in PH detection include blood gas analysis and cardiac magnetic resonance imaging, chest X-ray, and pulmonary angiography [5]. Several constraints limit the applicability of these tests, including low predictive performance of the tests, lack of interpretability, higher cost, or more invasive nature. Furthermore, these tests can support PH detection but are not definitive tests for PH detection. Recently, automated PH detection using cardiac auscultation data has emerged as a non-invasive and low-cost alternative that outperforms physicians [6], however these tests are also limited by low predictive performance, lack of interpretability, and the requirement to use simultaneous electrocardiogram (EKG) and heart sound signals.
[0007] Detection of PH based on heart sounds focuses on analyzing the second heart sound, S2, which itself consists of two mixed heart sound signals: the louder aortic valve closure (A2) and the quieter pulmonary valve closure (P2) [7]. Peak-to-peak analysis in the time domain shows that subjects with PH have a larger distance and larger amplitude difference between the A2 and P2 peaks [6].
[0008] Automatic diagnosis of PH based on heart sounds includes manual analysis [8] and traditional machine learning [6, 9]. In related fields, the application of deep convolutional neural networks (CNNs) is useful in pediatric heart murmur detection [4] and heart sound segmentation
[10] .
[0009] Existing automatic methods for assessing PH or PAP based on heart sounds do not provide sufficiently accurate results.
[0010] Known deep learning methods do not explain which part of the heart sounds is responsible for the output.
[0011] Late diagnosis of pulmonary hypertension (PH) results in adverse patient outcomes, highlighting the need for early, non-invasive PH detection. Cardiac auscultation offers a non-invasive and cost-effective alternative to right heart catheterization, CardioMEMS, and Doppler analysis in PH analysis, however it represents an indirect measurement of pulmonary artery pressure and therefore requires appropriate validation in different settings.
[0012] These facts are disclosed to illustrate the technical problems solved by the present disclosure. Summary of the invention
[0013] This document proposes the detection of pulmonary hypertension (PH) by analyzing digital heart sound recordings using an over-parameterized deep neural network. A preprocessing step aimed at separating the S2 sound into aortic (A2) and pulmonary artery (P2) components, as well as an explanation of the predictions are also disclosed. A deep neural network architecture, an optional compression step, and an optional alternative training method to obtain a prediction model with high accuracy and low resource requirements are also disclosed.
[0014] Its area under the ROC curve is 0.95, which is an improvement of 0.17 over the state-of-the-art Gaussian mixture model PH detector. Post hoc interpretation and analysis indicate that the availability of separated A2 and P2 components contributes significantly to the prediction.
[0015] Analyzing stethoscope heart sound recordings using deep networks is an effective, low-cost, and non-invasive solution for detecting pulmonary hypertension.
[0016] The method uses a deep convolutional neural network (CNN) [11, 12] commonly used for image analysis to analyze the audio data. The method also uses post hoc attribution methods commonly applied to CNN outputs, such as integrated gradients
[13] , to develop an explanation for the PH detection of the subjects' overall audio signal data and individual heartbeats.
[0017] It is known that deep networks usually require large training datasets, and previous state-of-the-art methods Gaussian mixture models (GMM) and support vector machines (SVM) cannot be scaled up to large datasets. Therefore, one of the novel aspects of the present disclosure regarding PH detection is to propose a deep network that uses datasets of any size, with specific optimizations for using deep networks on small data.
[0018] These optimizations include:
[0019] — alternative training mechanisms based on fixed-weight neural networks and / or non-iterative (i.e. one training step) optimization;
[0020] — Alternatively, batch gradient descent instead of mini-batch gradient descent;
[0021] Optionally, zero-pad the heartbeats to provide the same number of heartbeats for all subjects, i.e. pass images of the same size to the CNN;
[0022] —Optionally, channel normalization to stabilize gradient backpropagation via the equation:
[0023] (x) / channel std (x)
[0024] Where x is the "color channel" of the 3-channel input passed to the CNN.
[0025] Disclosed are: applying deep networks to heart sound recording analysis to provide strong predictive performance; and post hoc interpretation validating the roles of the proposed A2 and P2 components in the second heart sound.
[0026] In one embodiment, the method uses a physiologically relevant feature corresponding to at least one domain knowledge, preferably, the physiologically relevant feature is a characteristic of the P2 component and its relationship with respect to the A2 component.
[0027] Advantages of the disclosed interpretation method include:
[0028] — Enhance the credibility of a model for a given topic by verifying that (a) the model behaves according to domain knowledge and (b) its predictions are not different from other predictions made by the model. Explanations enhance credibility and are essential for decision making.
[0029] - Per-beat interpretation of regions of interest across time and across channels (i.e. proposed A2, proposed P2, S2);
[0030] - Aggregated per-subject interpretations of regions of interest across time and channels. Thus, per-beat interpretations can be aggregated to provide an overall impression of the prediction in the context of the subject.
[0031] This document discloses a computer-implemented method for non-invasively assessing pulmonary hypertension (PH) based on heart sound signals, the method comprising the following steps:
[0032] receiving a sound signal collected from a beating heart of a subject within a predetermined period of time (S2);
[0033] generating one or more 2D feature maps, the feature maps comprising a 2D feature map having the received sound signal (S2), wherein a first axis of the map is arranged over time and a second axis of the map is arranged over individual heartbeats;
[0034] A pre-trained neural network is applied to associate the generated one or more 2D feature maps with a previously acquired training data set and generate training 2D feature maps for a group of PH subjects and a group of non-PH subjects to obtain an indication of the presence of pulmonary hypertension.
[0035] This approach with a single input channel (ie not separated) is at least as accurate, significantly faster and has lower resource usage than the following approach with the additional step of separating the sound signal (S2) into the proposed A2 and P2 signals.
[0036] Although deep learning methods are almost always updated via backpropagation, in one embodiment, the method does not have backpropagation because the network is "wide" rather than "deep".
[0037] Optionally, the method introduces dimensionality reduction (e.g., principal component analysis (PCA)) into the deep network architecture, and introduces serial processing of parallel convolutional blocks that makes RAM usage scalable for a given device. Therefore, it can be evaluated and trained on mobile devices, such as laptops.
[0038] In an alternative embodiment, a computer-implemented method for non-invasively assessing pulmonary hypertension (PH) based on heart sound signals is also disclosed, the method comprising the following steps:
[0039] receiving a sound signal collected from a beating heart of a subject within a predetermined period of time (S2);
[0040] Separating the sound signal (S2) into an aorta sound signal (A2) and a pulmonary artery sound signal (P2);
[0041] generating one or more 2D feature maps, the feature maps comprising a 2D pulmonary artery feature map having a pulmonary artery sound signal (P2), wherein a first axis of the map is arranged in time and a second axis of the map is arranged in individual heartbeats;
[0042] A pre-trained neural network is applied to associate the generated one or more 2D feature maps with a previously acquired training data set, separating and generating training 2D feature maps for a group of PH subjects and a group of non-PH subjects, thereby obtaining an indication of the presence of pulmonary hypertension.
[0043] In one embodiment, the one or more 2D feature maps include a 2D aortic feature map having an aortic acoustic signal (A2), wherein a first axis of the map is arranged over time and a second axis of the map is arranged over individual heartbeats;
[0044] In one embodiment, the one or more 2D feature maps include a 2D full signal feature map having the received sound signal (S2), wherein a first axis of the map is arranged over time and a second axis of the map is arranged over individual heartbeats;
[0045] In one embodiment, the method further comprises combining the one or more 2D feature maps into a multi-channel input to the neural network, wherein each of the one or more 2D feature maps is combined into a channel of the multi-channel input.
[0046] In one embodiment, the method further comprises a multi-channel input to the neural network, wherein for the case of three feature maps, each channel is considered to be a color channel of the generated image, and wherein for the case of one feature map, each channel is considered to be a grayscale image.
[0047] In one embodiment, the method further includes providing the image to the neural network, wherein the one or more feature maps are combined into color channels of the generated image.
[0048] In one embodiment, the method includes dividing the collected sound signal into a plurality of time windows with a predetermined duration, each time window including a heartbeat sound signal peak, preferably, the predetermined duration is 200 milliseconds.
[0049] In one embodiment, the method includes aligning the segmented sound signal time windows by aligning heartbeat sound signal peaks of the segmented sound signal time windows.
[0050] In one embodiment, the method further comprises calculating a saliency attribution, preferably via integrated gradients or gradient times corresponding to the generated one or more 2D feature maps.
[0051] In one embodiment, the method comprises pre-processing the acquired sound signal by filtering, spike removal, normalization, alignment or segmentation or a combination thereof.
[0052] In one embodiment, the neural network is a convolutional neural network (CNN).
[0053] In one embodiment, the neural network is an over-parameterized deep neural network.
[0054] In one embodiment, the neural network is an extreme learning machine.
[0055] In one embodiment, the neural network is a deep or wide neural network with fixed weights. Fixed weights mean that the weights are not modified by optimization (not modified by training).
[0056] In one embodiment, the separation is performed using an alternating optimization of a least squares problem.
[0057] In one embodiment, the method further comprises, after separating the heart sound signals, filtering with a second order Butterworth filter, in particular filtering with a Butterworth filter having cutoff frequencies of 25 Hz and 400 Hz, resampling to 1 kHz, and cleaning by removing spikes.
[0058] In one embodiment, the method includes acquiring a heart sound signal at a point in the subject's lung, preferably at the second left intercostal space.
[0059] Also disclosed is a computer-implemented method for training a neural network to perform non-invasive assessment of pulmonary hypertension (PH) based on heart sound signals, the method comprising the following steps for both a PH subject group and a non-PH subject group:
[0060] receiving a sound signal collected from a beating heart of a subject within a predetermined period of time (S2);
[0061] Separate the heart sound signal (S2) into the aortic sound signal (A2) and the pulmonary artery sound signal (P2);
[0062] generating one or more 2D feature maps, the feature maps comprising a 2D pulmonary artery feature map having a pulmonary artery sound signal (P2), wherein a first axis of the map is arranged in time and a second axis of the map is arranged in individual heartbeats;
[0063] A pre-trained neural network is applied to associate the generated one or more 2D feature maps with a previously acquired training data set, separate and generate training 2D feature maps for the PH subject group and the non-PH subject group, and obtain an indication of the presence of pulmonary hypertension.
[0064] Also disclosed is a computer-implemented system for non-invasively assessing pulmonary hypertension (PH) based on heart sound signals, the system comprising an electronic data processing machine arranged to implement the following steps:
[0065] receiving a sound signal collected from a beating heart of a subject within a predetermined period of time (S2);
[0066] Separating the sound signal (S2) into an aorta sound signal (A2) and a pulmonary artery sound signal (P2);
[0067] generating one or more 2D feature maps, the feature maps comprising a 2D pulmonary artery feature map having a pulmonary artery sound signal (P2), wherein a first axis of the map is arranged in time and a second axis of the map is arranged in individual heartbeats;
[0068] A pre-trained neural network is applied to associate the generated one or more 2D feature maps with a previously acquired training data set, separate and generate training 2D feature maps for the PH subject group and the non-PH subject group, and obtain an indication of the presence of pulmonary hypertension.
[0069] In one embodiment, the system comprises a digital stethoscope for collecting sound signals of a beating heart, wherein the digital stethoscope is connected to an electronic data processor for transmitting the collected sound signals of a beating heart.
[0070] In one embodiment, the electronic data processor is further arranged to segment the collected sound signal into a plurality of time windows with a predetermined duration, each time window including a heartbeat sound signal peak, preferably, the predetermined duration is 200 milliseconds.
[0071] In one embodiment, the electronic data processing machine is further arranged to align the segmented sound signal time windows by aligning the heartbeat sound signal peaks of the segmented sound signal time windows.
[0072] In one embodiment, the electronic data processor is contained on a mobile device. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] The following drawings are provided to illustrate preferred embodiments of the present disclosure and should not be construed as limiting the scope of the present invention.
[0074] Figure 1 : Illustration of the S2 audio signal, the proposed A2 audio signal, the proposed P2 audio signal and the average feature attributes over time.
[0075] Figure 2A , Figure 2B , Figure 2C : Illustration of an implementation of a subject's audio data displayed as a 3-channel image and corresponding interpretation.
[0076] Figure 3: Includes illustrations of ROC curves for the disclosed method, referred to herein as DeepPHDet, GMM, and SVM models, which demonstrate the superior predictive performance of DeepPHDet over the prior art baselines.
[0077] Figure 4 : Flow chart of an embodiment of a method for non-invasively assessing pulmonary hypertension (PH) based on heart sound signals. DETAILED DESCRIPTION
[0078] An algorithm for non-invasive detection of pulmonary hypertension (PH) is disclosed. Heart sounds are collected from a subject with a digital stethoscope and subsequently passed as audio input to the algorithm. The output is an estimate of the PH value and an interpretation of the region of interest for the subject's heart sounds. The heart sound audio recording is collected from a point in the subject's lungs, for example, on the left hand side of the sternum, in the second intercostal space. The algorithm has a training phase, in which the algorithm acquires domain knowledge from multiple subject recordings, and an evaluation phase, in which the algorithm makes predictions for individual subjects.
[0079] In one embodiment, predictions are made for subjects that were not considered during the training phase.
[0080] In one embodiment, the algorithm begins with a set of preprocessing steps that include filtering, spike removal, and segmentation of S2 sounds from heart sound recordings. A source separation algorithm is then applied to the recorded S2 sounds to separate each S2 sound into its aortic (A2) and pulmonary artery (P2) components. The obtained S2 sounds are organized and normalized together with the corresponding A2 and P2 components into a feature map of the same shape as the 3-channel image provided as input to the deep neural network. During training, several steps were taken to ensure that the algorithm is applicable to small data, including the use of batch gradient descent, channel normalization to unit variance, and zero padding of inputs to ensure that all inputs have the same number of heartbeats. When evaluated, the algorithm predicts whether the subject has PH. In addition, the components of heart sounds, moments, and heartbeats that are most informative for prediction are highlighted by post hoc explanation attribution and aggregation methods.
[0081] Alternatively, either of the two optimization methods used to train the network is applied to ensure that the algorithm works well with small data. The first optimization method utilizes the following steps: using batch gradient descent instead of mini-batch gradient descent or stochastic gradient descent (the parameter update step occurs once per iteration over the entire dataset, but mini-batches can be used to accumulate gradients), channel normalization to unit variance, and in some cases, zero padding or cropping of the inputs to ensure that all inputs have the same number of cut heartbeats. The second method also uses the following steps: the neural network is assigned fixed parameter weights that are not updated due to training, and the optimization method uses the analytically derived solution to train the final classifier layer or model; sparsity regularization can be used; compression can be used.
[0082] Architectural choices were also made to ensure that the model works effectively with the dataset, including: choosing a convolutional layer architecture to compute parallel convolutions serially to reduce RAM usage, choosing a kernel width large enough to match the sampling rate of the sound signal, and careful use of pooling to implement the mathematical transformation of the input signal into a compressible latent space, as well as compression methods to reduce the size of the embedding space.
[0083] The predictions of the model depend critically on the proposed P2 heart sound, especially in the region of maximum variation between 30 ms and 60 ms.
[0084] Table 1: Dataset summary.
[0085]
[0086] A private dataset of 42 subjects was collected at Centro Hospitalar Universitário do Porto, Portugal. Summary statistics in Table 1 show 29 subjects with PH and 13 without PH. Among the affected subjects, the majority (20 of 29) were female. Age and heart rate were similar between the positive and affected populations. Inclusion and exclusion criteria were unknown. PH was defined as positive when the subject had a mean pulmonary artery pressure (MPAP) above 25 mm Hg or a pulmonary artery systolic pressure (PASP) above 30 mmHg. For each subject, true pulmonary artery pressure was obtained from right heart catheterization accompanied by a 5-min PCG heart sound recording. The recordings were obtained in a relatively quiet clinical environment with the subjects in the supine and resting position. The heart rate was measured using a 3D-scan 3D-scan radiograph connected to a Rugloop Waves 3D scanner. ® The system's custom-made cable stethoscope was used to auscultate the left second intercostal space. Heart sounds were recorded at a sampling rate of 8 kHz and their amplitudes were quantified at 16-bit resolution. The dataset was not made public to protect privacy.
[0087] In each five-minute audio signal, for each heartbeat, the S2 sound was segmented and extracted into 200ms windows, where the start time of the window was chosen so that the peaks of all S2 sounds for the subject were aligned in time. The S2 signal was filtered using a second-order Butterworth filter with cutoff frequencies of 25Hz and 400Hz, resampled to 1kHz, cleaned by removing spikes using the method in
[14] , and separated into the proposed A2 and P2 components according to
[15] . The source separation assumes that the aortic and pulmonary components maintain approximately the same waveform between heartbeats, and assumes that the delay between the components within a heartbeat varies due to changes in chest pressure during different respiratory phases. The two components are retrieved by alternating optimization of a least squares problem.
[0088] In one embodiment, the duration of each window of the audio signal is a predetermined parameter set by the user, for example, 200 ms for a sampling rate of 1 kHz.
[0089] In one example, the alignment and segmentation results in a multi-channel 2-D representation of audio data containing S2, the proposed A2, and the proposed P2 components. Each 2-D channel has 200 columns representing 200ms windows and as many rows as heartbeats. Channels are then prepared for all subjects with the same shape by zero padding the 454 rows, and each of the 3 channels for each subject is independently normalized to unit variance. Normalization to unit variance helps stabilize gradient backpropagation by reducing the risk of gradient vanishing or gradient exploding.
[0090] In one embodiment, it is considered to be DenseNet121, ResNet18, and EfficientNet-b0 architectures.
[0091] In one embodiment, pre-trained deep network initialization improves performance, especially for small datasets.
[0092] In one embodiment, random and ImageNet initializations are considered.
[0093] DenseNet trained from random initialization, ResNet18 trained from ImageNet initialization, and EfficientNet-b0 trained from standard adversarial ImageNet initialization are disclosed. These models are trained with batch gradient descent with a learning rate of 0.0001 and momentum of 0.5 for 150 epochs. Deep networks are often trained on large datasets using mini-batch gradient descent. To stabilize the gradient updates, batch gradient descent is implemented, which means that a gradient update is applied per iteration on the dataset. To be applicable to datasets of arbitrary size, we compute a gradient update for each sample and keep a sum or running average until a gradient update occurs. The loss is a weighted loss with positive class balance. The weighted binary cross entropy of .
[0094] To evaluate the predictive performance of deep networks relative to classical methods, Gaussian mixture models (GMM) and support vector machines (SVM) were implemented.
[0095] The GMM implementation of the present invention adopts the prior art work in [6], where one GMM is trained for the positive class and the other for the negative class. The test sample class is the GMM model with the higher posterior negative log-likelihood. In order to obtain the best performance with this baseline, different preprocessing pipelines were developed and the GMM model was optimized accordingly to have two components and spherical covariance. The SVM used an RBF kernel and relaxation parameter C=1. For preprocessing, only the S2 channel was used. The proposed addition of A2 and P2 channels had a negative impact on the performance due to overfitting. Each heartbeat (each row of the S2 channel) was transformed by a 1-d short-time Fourier transform using an FFT window of 64 samples and a jump length of two samples, and the energy spectrum was calculated by absolute value. The subject data, a tensor of shape (H,33,101), was simplified to (33,101) by calculating the 98% quantile of H heartbeats. The channels are padded with zeros to 454 rows, normalized to unit variance, and flattened into vectors before being passed to the SVM and GMM models.
[0096] All models were evaluated using 10-fold stratified cross validation. To report performance, the predicted probabilities from the validation set for each fold were stored. There was one predicted probability for each subject. The area under the ROC curve (ROC AUC) and standard classification metrics are reported. Classification metrics require the selection of a threshold to convert probabilities into categories. A threshold Tk was selected for each K-th fold that maximized the difference between the true positive rate minus the false positive rate on the K-th fold training set ROC curve. This threshold optimized the balanced accuracy score for the training set. Validation performance was calculated within each fold, and the metrics were then aggregated by averaging across folds and epochs 100 to 150.
[0097] To better understand which parts of the proposed A2 and proposed P2 channels contribute to PH detection, the integrated gradient attribution method
[13] was applied.
[0098] In one embodiment, after training the DenseNet121 model on ten folds, ten independently trained models are obtained. Therefore, ten attributions are calculated for each heartbeat in the dataset and then averaged to obtain one attribution per channel or summed to obtain an importance score for each heartbeat. For better visualization, the attributions are converted to magnitudes by absolute values and then clipped to 1% to 99% of their values. Clipping helps with visualization because gradient-based attribution methods generate some outliers.
[0099] Table 2: DeepPHDet* presents the state-of-the-art results
[0100]
[0101] The results in Table 2 show that for the considered PH dataset, the DenseNet121 and EfficientNet-b0 deep networks outperform the state-of-the-art machine learning models by a large margin. The DenseNet121 model has the highest performance of 0.95 ROC AUC, the highest balanced accuracy (BAcc), and the highest Matthew correlation coefficient (MCC). The two best performing models are DenseNet121 and EfficientNet-b0.
[0102] The bottom row of Table 2 shows that the availability of S2, A2, and P2 channels improves the performance compared to using only S2. The motivation for deep learning is to overcome the need for preprocessing through data-driven feature generation and larger datasets. In the small data regime, as in this paper, preprocessing is observed to improve the performance. Moreover, the over-parameterized nature of deep networks requires a rethinking of the existing technical explanations of underfitting and overfitting. Classical methods, such as SVM and GMM, overfit using the extra parameters from A2 and P2 channels, while deep networks improve this situation.
[0103] Figure 1 Shown are graphs of the S2 audio signal, the proposed A2 audio signal, the proposed P2 audio signal, and average feature attributes over time.
[0104] Figure 1 The top 3 rows of show the heart sound data of one subject. Each line represents one heartbeat. The top row shows the S2 signal. The second and third rows show the proposed source separation signals A2 and P2. The signals shown are normalized to unit variance to represent the input passed to the prediction model.
[0105] Normalization has been empirically found to improve performance; normalization makes the quieter P2 have a similar amplitude to the louder A2. Since the heartbeats are aligned based on their peaks, the A2 signal is very clearly defined. The distance between the A2 and P2 components varies depending on factors such as whether the subject is inhaling or exhaling and the presence of PH. Therefore, current domain knowledge is consistent with visual observation that the average P2 signal should not be located very well in time. In this example, it was observed that P2 has the most variable behavior between 30ms and 60ms. Current domain knowledge expects PH to be related to changes in the timing and amplitude of P2.
[0106] Figure 1 The bottom plot in shows the average attribution and 99.9% confidence interval for all heartbeats. For this example, the attribution to P2 dominates and also coincides with the time period between 30ms and 60ms where the P2 behavior is most variable. Both observations indicate that the deep network is consistent with the domain knowledge. The attribution to A2 is strongest at the peak, just before 25ms. The attribution suggests that the availability of separated components helps in prediction.
[0107] Figure 2A , Figure 2B , Figure 2C A diagram showing an embodiment of a subject's audio data displayed as a 3-channel image and corresponding interpretation is shown.
[0108] The first row is the input of the CNN, and the second and third rows are the outputs of the attribution method. The first three columns are the S2, the proposed A2, and the proposed P2 components. The fourth column represents the aggregated view of all inter-beat waveforms. In the figure, all beats are zero-padded, and the second and third rows use inputs normalized to unit variance before calculating the attribution.
[0109] Figure 3 A graphical representation including ROC curves of the disclosed method, referred to herein as DeepPHDet, GMM and SVM models is shown, demonstrating that DeepPHDet outperforms the prior art baseline in predictive performance.
[0110] It is found that deep networks improve detection performance; separating S2 into A2 and P2 improves performance and improves the interpretability of the model and the analyzed S2 signal; the proposed A2 and P2 are consistent with domain knowledge; and post hoc interpretation verifies the domain knowledge and utility of the A2 and P2 segmentation.
[0111] Then, the present disclosure contributes to the advancement of the prior art of automatic detection of pulmonary hypertension (i.e., pulmonary arterial hypertension) from heart sounds. It includes several advantages, such as: high prediction performance; suitability for training using both small and large datasets; interpretation of the predicted A2 and S2 components; and only requires heart sound data for inference.
[0112] Deep networks trained on a private dataset of preprocessed digital stethoscope recordings have been shown to achieve ROC AUC scores of 0.95 and 0.93, providing improvements of +0.17 and +0.15 over the above prior art adaptations based on Gaussian mixture models, and +0.07 and +0.05 over prior art machine learning implementations.
[0113] Post hoc interpretation and improved performance suggest that separation of S2 sounds into the proposed A2 and P2 components aids detection.
[0114] Figure 4 A flow chart showing an embodiment of a method for non-invasively assessing pulmonary hypertension (PH) based on heart sound signals.
[0115] In one embodiment, the entire model is trained via back-propagation.
[0116] In another embodiment, during training, the convolutional network weights are initialized and fixed (never modified), and the remaining steps are obtained through techniques from extreme learning machines or regression models.
[0117] The test was conducted on 3 datasets obtained by stethoscope (PCG) and seismocardiogram (SCG) devices, including human and pig recordings. That is, human dataset, PCG:
[0118] 42 human subjects who underwent right heart catheterization;
[0119] recording of heart sounds using a digital stethoscope;
[0120] 13 had no PH, and 29 had PH;
[0121] Pig dataset, PCG+SCG:
[0122] Ten pigs, each undergoing right heart catheterization;
[0123] The dataset size is 125 “pig patients” (according to the sampling period from 10 pigs);
[0124] Each pig underwent multiple bouts of chemically induced hypertension. Heart sounds were recorded at selected intervals;
[0125] Recording devices: phonocardiography (PCG) and shockcardiography (SCG);
[0126] Human dataset, SCG:
[0127] 73 human subjects underwent right heart catheterization and episiocardiography.
[0128] Each dataset is evaluated separately (via cross-validation) and "cross-domain generalization" is evaluated (training on one dataset and evaluating on the other). Methods are evaluated using records of different record lengths.
[0129] Table 3: Different record lengths for person (PCG) data
[0130]
[0131] Each number describes the performance of 12 independently trained models, each with 10-fold cross validation. Macro averages are taken for each fold. Micro describes the performance for each sample. auROC is the area under the ROC curve. AP is the average precision score (area under the PR curve).
[0132] Table 3: Test results of the method without separation step.
[0133]
[0134]
[0135] Each number describes the performance of 12 independently trained models, each with 10-fold cross validation. Macro averages are taken for each fold. Micro describes the performance on each sample. auROC is the area under the ROC curve. AP is the average precision score (area under the PR curve). Model hyperparameters were tuned to the training portion of the baseline human (PCG) and pig (PCG+SCG) datasets. Percentiles compare cross-domain performance to baseline performance.
[0136] Results for methods having the additional steps of separating the sound signal (S2) into an aortic sound signal (A2) and a pulmonary artery sound signal (P2) and applying a pre-trained neural network to correlate the generated one or more 2D feature maps with a training dataset of previously acquired, componentized and generated training 2D images were similar, and in some cases slightly lower.
[0137] AP is the average precision (area under the precision-recall curve), and auROC is the area under the ROC curve. These two metrics together give a good judgement of the algorithm's performance, including in the presence of class imbalance. The model hyperparameters were the same for all experiments.
[0138] Analyzing stethoscope heart sound data using deep networks is an effective, low-cost, and non-invasive solution for detecting pulmonary hypertension.
[0139] The present disclosure uses a deep network to analyze heart sounds, has low resource cost and is suitable for early screening.
[0140] The present disclosure includes advantages such as interpretation of regions of interest in an individual heartbeat and all heartbeats as a whole, and enhancing the medical credibility of models for subject-specific predictions.
[0141] Whenever used in this document, the term “comprising” is intended to indicate the presence of mentioned features, integers, steps, components, but does not exclude the presence or addition of one or more other features, integers, steps, components or groups thereof.
[0142] The present disclosure should not be considered in any way as being limited to the described embodiments, and a person skilled in the art will foresee numerous possibilities for modifications thereof.The above-described embodiments are combinable.
[0143] The following claims further describe specific embodiments of the present disclosure.
[0144] References
[0145] [1] Hasan B, Hansmann G, Budts W, Heath A, Hoodbhoy, Jing ZC,Koestenberger M, Meinel K, Mocumbi AO, Radchenko GD, et al. Challenges and special aspects of pulmonary hypertension in middle-to low-income regions: Jacc state-of-the-art review. Journal of the American College of Cardiology2020;75(19):2463–2477.
[0146] [2] Lau EM, Humbert M, Celermajer DS. Early detection of pulmonary arterial hypertension. Nature Reviews Cardiology 2015;12(3):143–155.
[0147] [3] Taleb M, Khuder S, Tinkel J, Khouri SJ. The diagnostic accuracyof d oppler echocardiography in assessment of pulmonary artery systolicpressure: A meta-analysis. Echocardiography 2013;30(3):258–265.
[0148] [4] Oliveira JH, Renna F, Costa P, Nogueira D, Oliveira C, Fer258reira C, Jorge A, Mattos S, Hatem T, Tavares T, Elola A, Rad A, Sameni R,Clifford GD, Coimbra MT. The circordigiscope dataset: From murmur detectionto murmur classification. IEEE Journal of Biomedical and Health Informatics2021;1–1.
[0149] [5] Lang IM, Plank C, Sadushi-Kolici R, Jakowitsch J, Klepetko W,Maurer G. Imaging in pulmonary hyper tension. JACC Cardiovascular Imaging2010;3(12):1287–1295.
[0150] [6] Kaddoura T, Vadlamudi K, Kumar S, Bobhate P, Guo L, Jain S,Elgendi M, Coe JY, Kim D, Taylor D, et al. Acoustic diagnosis of pulmonaryhypertension: automated speech recognition-inspired classification algorithmoutperforms physicians. scientific reports 2016;6(1):1–11.
[0151] [7] Xu J, Durand L, Pibarot P. Nonlinear transient chirp signalmodeling of the aortic and pulmonary components of the second heart sound.IEEE Transactions on Biomedical Engineering 2000;47(10):1328–1335.
[0152] [8] Andreev V, Gramovich V, Krasikova M, Korolkov A, Vyborov O,Danilov N, Martynyuk T, Rodnenkov O, Rudenko O. Time–frequency analysis ofthe second heart sound to assess pulmonary artery pressure. AcousticalPhysics 2020; 66(5):542–547.
[0153] [9] Dennis A, Michaels AD, Arand P, Ventura D. Noninvasive diagnosisof pulmonary hypertension using heart sound analysis. Computers in Biologyand Medicine 2010; 40(9):758–764.
[0154]
[10] Renna F, Oliveira J, Coimbra MT. Deep convolutional neuralnetworks for heart sound segmentation. IEEE journal of biomedical and healthinformatics 2019;23(6):2435–2445.
[0155]
[11] Huang G, Liu Z, Van Der Maaten L, Weinberger KQ. Denselyconnected convolutional networks. In Proceedings of the IEEE conference oncomputer vision and pattern recognition. 2017; 4700–4708.
[0156]
[12] Tan M, Le Q. Efficientnet: Rethinking model scaling forconvolutional neural networks. In International conference on machinelearning. PMLR, 2019; 6105–6114.
[0157]
[13] Sundararajan M, Taly A, Yan Q. Axiomatic attribution for deepnetworks. In International conference on machine learning. PMLR, 2017; 3319–3328.
[0158]
[14] Schmidt SE, Holst-Hansen C, Graff C, Toft E, Struijk JJ.Segmentation of heart sound recordings by a duration dependent hidden markovmodel. Physiological measurement 2010;31(4):513.
[0159]
[15] Renna F, Plumbley MD, Coimbra M. Source separation of the secondheart sound via alternating optimization. In 2021 Computing in Cardiology(CinC), volume 48. IEEE, 2021; 1–4。
Claims
1. A computer-implemented method for non-invasively assessing pulmonary hypertension (PH) based on heart sound signals, the method comprising the following steps: receiving a sound signal collected from a beating heart of a subject within a predetermined period of time (S2); generating one or more 2D feature maps, the feature maps comprising a 2D feature map having the received sound signal (S2), wherein a first axis of the map is arranged over time and a second axis of the map is arranged over individual heartbeats; A pre-trained neural network is applied to associate the generated one or more 2D feature maps with a previously acquired training data set and generate training 2D feature maps for the PH subject group and the non-PH subject group, and obtain an indication of the presence of pulmonary hypertension.
2. The method for non-invasively assessing pulmonary hypertension according to the preceding claims, further comprising the following steps: The sound signal (S2) is separated into an aortic sound signal (A2) and a pulmonary artery sound signal (P2); one or more 2D feature maps are generated, wherein the feature maps include a 2D pulmonary artery feature map having a pulmonary artery sound signal (P2), wherein a first axis of the map is arranged in time and a second axis of the map is arranged in each heartbeat; a pre-trained neural network is applied to associate the generated one or more 2D feature maps with a previously acquired training data set, separate and generate training 2D feature maps for a PH subject group and a non-PH subject group, and obtain an indication of the presence of pulmonary hypertension.
3. A method for non-invasively assessing pulmonary hypertension according to the preceding claim, wherein the one or more 2D feature maps include a 2D aortic feature map having an aortic sound signal (A2), wherein a first axis of the map is arranged over time and a second axis of the map is arranged over individual heartbeats.
4. A method for non-invasively assessing pulmonary hypertension according to claim 2 or 3, wherein the one or more 2D feature maps include a 2D complete signal feature map with the received sound signal (S2), wherein a first axis of the map is arranged in time and a second axis of the map is arranged in individual heartbeats.
5. The method for non-invasively assessing pulmonary hypertension according to claim 3 or 4 further comprises combining one or more 2D feature maps into a multi-channel input to a neural network, wherein each of the one or more 2D feature maps is combined into a channel of the multi-channel input.
6. A method for non-invasively assessing pulmonary hypertension according to any one of the preceding claims, comprising: The collected sound signal is divided into a plurality of time windows with a predetermined duration, each time window including a heartbeat sound signal peak.
7. A method for non-invasively assessing pulmonary hypertension according to the preceding claims, comprising: The segmented sound signal time windows are aligned by aligning the heartbeat sound signal peaks of the segmented sound signal time windows.
8. The method for non-invasive assessment of pulmonary hypertension with interpretability according to any one of the preceding claims, further comprising: Saliency attributions are calculated, preferably via integrated gradients or gradient times corresponding to the generated one or more 2D feature maps.
9. A method for non-invasively assessing pulmonary hypertension according to any one of the preceding claims, comprising: The acquired sound signal is pre-processed by filtering, spike removal, normalization, alignment or segmentation or a combination thereof.
10. The method for non-invasively assessing pulmonary hypertension according to any one of the preceding claims, wherein the neural network is a convolutional neural network (CNN).
11. A method for non-invasively assessing pulmonary hypertension according to any one of the preceding claims, wherein the neural network is an over-parameterized deep neural network.
12. The method for non-invasively assessing pulmonary hypertension according to any one of claims 1 to 9, wherein the neural network is an extreme learning machine.
13. The method for non-invasive assessment of pulmonary hypertension according to any one of claims 2 to 12, wherein the separation is performed using alternating optimization of a least squares problem.
14. The method for non-invasively assessing pulmonary hypertension according to any one of claims 2-13 further comprises, after separating the heart sound signal, filtering with a second-order Butterworth filter, in particular filtering with a Butterworth filter with cutoff frequencies of 25 Hz and 400 Hz, resampling to 1 kHz, and purifying by removing spikes.
15. A method for non-invasively assessing pulmonary hypertension according to any one of the preceding claims, comprising: The heart sound signal is acquired at a lung site of the subject, preferably at the second left intercostal space.
16. A computer-implemented method for training a neural network to perform non-invasive assessment of pulmonary hypertension (PH) based on heart sound signals, for both a PH subject group and a non-PH subject group, the method comprising the steps of: receiving a sound signal collected from a beating heart of a subject within a predetermined period of time (S2); generating one or more 2D feature maps, the feature maps comprising a 2D feature map having the received sound signal (S2), wherein a first axis of the map is arranged over time and a second axis of the map is arranged over individual heartbeats; A pre-trained neural network is applied to associate the generated one or more 2D feature maps with a previously acquired training data set and generate training 2D feature maps for the PH subject group and the non-PH subject group, and obtain an indication of the presence of pulmonary hypertension.
17. A method for training a neural network according to the preceding claims, said method further comprising the steps of: Separate the heart sound signal (S2) into the aortic sound signal (A2) and the pulmonary artery sound signal (P2); generating one or more 2D feature maps, the feature maps comprising a 2D pulmonary artery feature map having a pulmonary artery sound signal (P2), wherein a first axis of the map is arranged in time and a second axis of the map is arranged in individual heartbeats; A pre-trained neural network is applied to associate the generated one or more 2D feature maps with a previously acquired training data set, separate and generate training 2D feature maps for the PH subject group and the non-PH subject group, and obtain an indication of the presence of pulmonary hypertension.
18. A computer-implemented system for non-invasively assessing pulmonary hypertension (PH) based on heart sound signals, the system comprising an electronic data processor arranged to implement the following steps: receiving a sound signal collected from a beating heart of a subject within a predetermined period of time (S2); generating one or more 2D feature maps, the feature maps comprising a 2D feature map having the received sound signal (S2), wherein a first axis of the map is arranged over time and a second axis of the map is arranged over individual heartbeats; A pre-trained neural network is applied to associate the generated one or more 2D feature maps with a previously acquired training data set and generate training 2D feature maps for the PH subject group and the non-PH subject group, and obtain an indication of the presence of pulmonary hypertension.
19. The computer-implemented system according to the preceding claim, wherein the electronic data processing machine is further arranged to implement the following steps: Separating the sound signal (S2) into an aorta sound signal (A2) and a pulmonary artery sound signal (P2); generating one or more 2D feature maps, the feature maps comprising a 2D pulmonary artery feature map having a pulmonary artery sound signal (P2), wherein a first axis of the map is arranged in time and a second axis of the map is arranged in individual heartbeats; A pre-trained neural network is applied to associate the generated one or more 2D feature maps with a previously acquired training data set, separate and generate training 2D feature maps for the PH subject group and the non-PH subject group, and obtain an indication of the presence of pulmonary hypertension.
20. A computer-implemented system according to any one of claims 18-19, the system comprising a digital stethoscope for collecting sound signals of a beating heart, wherein the digital stethoscope is connected to the electronic data processor for transmitting the collected sound signals of a beating heart.
21. A computer-implemented system according to any one of claims 18-20, wherein the electronic data processor is further arranged to segment the collected sound signal into a plurality of time windows with predetermined durations, each time window including a heartbeat sound signal peak, preferably, the predetermined duration is 200 milliseconds.
22. The computer-implemented system of any one of claims 18-21, wherein the electronic data processing machine is further arranged to align the segmented sound signal time windows by aligning heartbeat sound signal peaks of the segmented sound signal time windows.
Citation Information
Cited By
Pulmonary arterial hypertension recognition device and method based on multiple paths of electrocardio and heart sound signals
CN121714242A