An explainable deep learning method for noninvasive detection of pulmonary hypertension from heart sounds
A deep learning method using CNNs to analyze heart sound components A2 and P2 provides accurate and interpretable PH detection, addressing the limitations of existing technologies with improved performance and resource efficiency.
Patent Information
- Application Number
- JP2025513071
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-09-02
- Filing Date
- 2023-09-01
- Publication Date
- 2025-08-28
AI Technical Summary
Current methods for detecting pulmonary hypertension (PH) are invasive, costly, or lack accuracy and interpretability, particularly in low-resource settings, and existing automated methods do not effectively utilize heart sound data for reliable non-invasive detection.
A deep learning method using convolutional neural networks (CNNs) to analyze heart sound recordings, specifically separating the aortic (A2) and pulmonary (P2) components, with post-hoc attribution techniques to provide explanations for PH detection, optimized for small datasets and resource-efficient implementation.
Achieves high predictive performance with improved accuracy and interpretability, enabling non-invasive and cost-effective detection of PH, suitable for low-resource settings, with ROC AUC scores of 0.95 and 0.93, surpassing previous state-of-the-art methods.
Smart Images

Figure 2025528500000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an explainable deep learning method for non-invasive detection of pulmonary hypertension from heart sounds. [Background technology]
[0002] Pulmonary hypertension (PH) is an underrecognized disease with unmet need for diagnosis and treatment recommendations in low- and middle-income communities (Non-Patent Document 1). PH disease has a high mortality rate, and early detection in screening programs can improve outcomes. Existing tools for PH detection are not fully optimized for the needs of low- and middle-income communities.
[0003] Right heart catheterization is the gold standard for detecting PH, but it is very expensive, highly invasive, not suitable for screening programs, and requires a dedicated team of well-trained clinicians.
[0004] Implantable pressure sensors are placed in the pulmonary artery and are extremely expensive and highly invasive. Although implantable pressure sensors provide the best blood pressure readings, they are only used in a very small percentage of subjects in wealthy countries.
[0005] Doppler echocardiography is widely used for clinical screening of PH, but the noisy nature of its measurements requires additional modalities (medical devices, pharmaceuticals) to improve reliability (Non-Patent Documents 2, 3). Doppler echocardiography only provides a predictive value for pulmonary artery pressure in a subset of subjects. Ultrasound techniques also require trained technicians and expensive equipment (Non-Patent Document 4).
[0006] Other tests useful for detecting PH include blood gas analysis and imaging with cardiac magnetic resonance, chest X-ray, and pulmonary angiography (Non-Patent Document 5). Several limitations limit the applicability of these tests, including their low predictive performance, lack of interpretability, higher cost, or more invasive nature. Furthermore, although these tests can assist in PH detection, they are not definitive tests for PH detection. Automated PH detection using auscultatory cardiac data has recently emerged as a non-invasive and low-cost alternative that can surpass physicians (Non-Patent Document 6). However, these tests are also limited by their low predictive performance, lack of interpretability, and the need to use a synchronized electrocardiogram (EKG) simultaneously with the heart sound signal.
[0007] Detection of PH from heart sounds focuses on the analysis of the second heart sound (S2), which itself consists of a mixture of two acoustic signals: the louder aortic valve closure (A2) and the quieter pulmonary valve closure (P2) (Non-Patent Document 7). Peak-to-peak analysis in the time domain has shown that subjects with PH exhibit a longer (temporal) distance and a larger amplitude difference between the A2 and P2 peaks (Non-Patent Document 6).
[0008] Automated diagnosis of PH from heart sounds involves manual analysis (Non-Patent Document 8) and traditional machine learning (Non-Patent Documents 6, 9). In a related area, the application of deep convolutional neural networks (CNNs) has been useful in heart murmur detection in children (Non-Patent Document 4) and heart sound segmentation (Non-Patent Document 10).
[0009] Existing automated methods for estimating PH or PAP (Pulmonary Alveolar Proteinosis) from heart sounds do not provide sufficiently accurate results.
[0010] Known deep learning methods do not explain which part of the heart sound the output is attributed to.
[0011] Poor patient outcomes as a result of delayed diagnosis of pulmonary hypertension highlight the need for earlier, noninvasive detection of PH. Cardiac auscultation offers a noninvasive and cost-effective alternative to right heart catheterization. However, auscultation represents an indirect measurement of pulmonary artery pressure and therefore needs to be properly validated in different scenarios.
[0012] These facts are disclosed to explain the technical problem that the present invention addresses. [Prior art documents] [Non-patent literature]
[0013] [Non-Patent Document 1] Hasan B, Hansmann G, Budts W, Heath A, Hoodbhoy, Jing ZC, Koestenberger M, Meinel K, Mocumbi AO, Radchenko GD, et al, “Challenges and special aspects of pulmonary hypertension in middle-to-low-income regions: Jacc state-of-the-art review”, Journal of the American College of Cardiology 2020;75(19):2463-2477 [Non-patent document 2] Lau EM, Humbert M, Celermajer DS, “Early detection of pulmonary artificial hypertension”, Nature Reviews Cardiology 2015;12(3):143-155 [Non-patent document 3] Taleb M, Khuder S, Tinkel J, Khouri SJ, “The diagnostic accuracy of doppler echocardiography in the assessment of pulmonary artery systolic pressure: A meta-analysis”, Echocardiography 2013;30(3):258-265
Outdoor Tools 4
Direct Environment 5
Outdoor Configuration6
Direct Environment 7
Outdoor Track 8
Outdoor Tools9
[0014] Overview This document proposes the detection of pulmonary hypertension by analyzing digital heart sound recordings with a deep neural network with excessive parameters. A preprocessing step aimed at separating the S2 sound into aortic (A2) and pulmonary (P2) components, and a description of the prediction are further disclosed. The deep neural network architecture, an optional comparison step, and optional alternative training methods that yield highly accurate and less resource-demanding prediction models are also disclosed.
[0015] An area under the curve with a ROC (Radius of Curvature) of 0.95 was obtained, an improvement of 0.17 over current state-of-the-art Gaussian Mixture Model (Gaussian Mixture Model, Normal Mixture Model) PH detectors. Post-hoc description and analysis show that the availability of separated A2 and P2 contributes significantly to the prediction.
[0016] Deep network analysis of stethoscope heart sound recordings is a low-cost, non-invasive solution for the detection of pulmonary hypertension.
[0017] This method employs deep convolutional neural networks (CNNs) (11, 12), which are typically used for image analysis and for analyzing audio data. This method also employs a posteriori attribution techniques, such as integral gradients (13), typically applied to CNN outputs, to develop explanations for PH detection and individual heart rate on the subject's overall acoustic signal data.
[0018] It is known that deep networks generally require large training databases, and previous state-of-the-art methods, such as Gaussian Mixture Models (GMMs) and Support Vector Machines (SVMs), do not scale to large databases. Therefore, one of the novel aspects of the present invention for PH detection is to propose a deep network that can be used with databases of any size, with specific optimizations for using the deep network on small data.
[0019] These optimizations include: Alternative learning mechanisms based on neural networks with fixed weights and / or non-iterative (i.e., in a single learning step) optimization; -Batch gradient descent instead of mini-batch gradient descent; - Optionally, zero-padding heartbeats to give all subjects the same heartbeat, i.e., to feed the CNN images of the same size; -Optionally, channel normalization to stabilize gradient backpropagation via the following equation: (x) / Channel std (x) where x is the "color channel" of the 3-channel input passed to the CNN, and Channel std () denotes channel normalization within ().
[0020] Application of deep networks to heart sound recordings gives strong predictive performance; post-hoc explanations are shown to substantiate the proposed role of the A2 and P2 components in the second heart sound.
[0021] In one preferred embodiment, the method uses physiologically relevant features that represent at least one domain knowledge, preferably the physiologically relevant features being properties of the P2 component and the relationship of the P2 component to the A2 component.
[0022] The disclosed illustrative methods include: - Strengthening the credibility of a model for a given subject by verifying that (a) the model behaves according to domain knowledge and (b) the predictions are not different from other predictions made by this model. Explanations strengthen credibility and are essential for decision making. - Beat-by-beat description of the region of interest over time and across channels, i.e., proposed A2, proposed P2, S2; - Aggregation of subject-specific descriptions of areas of interest across time and channels. Thus, beat-by-beat descriptions can be aggregated to give an overall impression of the prognosis associated with the subject.
[0023] This document discloses a computer-implemented method for non-invasively estimating pulmonary hypertension, PH, from heart sound signals, the method comprising: receiving an acoustic signal (S2) acquired from the subject's beating heart over a predetermined period of time; generating one or more two-dimensional (2D) feature maps comprising a 2D feature map having the received acoustic signal (S2), wherein a first axis of the map is a time axis and a second axis of the map is an axis spanning a plurality of individual heartbeats; and applying the pre-trained neural network to correlate the generated one or more 2D maps with a training dataset of previously acquired and generated training 2D feature maps of a group of PH and non-PH subjects to obtain an indication of the presence of pulmonary hypertension.
[0024] This method is at least as accurate as, significantly faster than, and has lower resource usage than the following method, which has an additional step of splitting the acoustic signal (S2) into the proposed A2 and P2 signals with a single input channel (i.e., no splitting).
[0025] Deep learning methods are almost always updated by backpropagation, but in one preferred embodiment, the method does not have backpropagation and the network is "wide" rather than "deep."
[0026] Optionally, the method includes dimensionality reduction, e.g., Principal Component Analysis (PCA), in the deep network architecture, and sequential processing of parallel (parallel-averaged) convolution blocks, allowing RAM usage to be scalable for a given device. Thus, the method can be evaluated and trained on mobile devices, e.g., laptops.
[0027] In an alternative preferred embodiment, a computer-implemented method for non-invasively estimating pulmonary hypertension, PH, from heart sound signals is disclosed, the method comprising: receiving an acoustic signal (S2) acquired from the subject's beating heart over a predetermined period of time; Splitting the acoustic signal (S2) into an aortic acoustic signal (A2) and a pulmonary acoustic signal (P2); generating one or more 2D feature maps including a 2D lung feature map having a lung acoustic signal (P2), wherein a first axis of the map is a time axis and a second axis of the map is an axis spanning a plurality of individual heartbeats; and applying the pre-trained neural network to correlate the generated one or more 2D feature maps with a training dataset of previously acquired and segmented training 2D feature maps for a group of PH and non-PH subjects, thereby obtaining an indication of the presence of pulmonary hypertension.
[0028] In one embodiment, the one or more 2D feature maps include a 2D aortic feature map having an aortic acoustic signal (A2), a first axis of the map being a time axis and a second axis of the map being an axis spanning multiple individual heartbeats.
[0029] In one embodiment, the one or more 2D feature maps include a 2D full signal feature map having a received acoustic signal (S2), a first axis of the map being a time axis, and a second axis of the map being an axis spanning multiple individual heartbeats.
[0030] In one embodiment, the method further comprises combining the one or more 2D feature maps as a multi-channel input to a neural network, each of the one or more 2D feature maps being combined as one channel of the multi-channel input.
[0031] In one embodiment, the method further includes a multi-channel input for the neural network, where in the case of three feature maps, each channel is considered a color channel of the generated image, and in the case of one feature map, each channel is considered as a grayscale image.
[0032] In one embodiment, the method further includes generating an image for the neural network, combining one or more feature maps as color channels of the generated image.
[0033] In one preferred embodiment, the method includes dividing the acquired acoustic signal into a plurality of time windows of a predetermined duration, each time window including a peak of the heartbeat sound signal, the predetermined duration being preferably 200 milliseconds.
[0034] In one embodiment, the method includes aligning the segmented acoustic signal time windows by aligning peaks of the heartbeat sound signals in the segmented acoustic signal time windows.
[0035] In one embodiment, the method further comprises the step of calculating a saliency attribution, preferably by integrated gradients or gradient times corresponding to the generated one or more 2D feature maps.
[0036] In one embodiment, the method includes pre-processing the acquired acoustic signals by filtering, spike removal, normalization, registration and / or segmentation.
[0037] In one embodiment, the neural network is a convolutional neural network.
[0038] In one preferred embodiment, the neural network is a deep neural network with excess parameters.
[0039] In one preferred example, the neural network is an extreme learning machine.
[0040] In one preferred embodiment, the neural network is a deep or wide area neural network with fixed weights, meaning that the weights are not modified by optimization (not modified by learning).
[0041] In one preferred embodiment, the partitioning is performed using alternating optimization of a least squares problem.
[0042] In one embodiment, the method includes the steps of splitting the heart sound signal, then filtering it with a second-order Butterworth filter, in particular with a Butterworth filter having cutoff frequencies of 25 Hz and 400 Hz, resampling at 1 kHz, and cleaning it by removing spikes.
[0043] In one embodiment, the method includes acquiring a heart sound signal at a lung spot of the subject, preferably over the second intercostal space on the left side.
[0044] Also disclosed is a computer-implemented method for training a neural network for non-invasive estimation of pulmonary hypertension, PH, from heart sound signals, the method comprising the steps of: receiving an acoustic signal (S2) acquired from the subject's beating heart over a predetermined period of time; Splitting the acoustic signal (S2) into an aortic acoustic signal (A2) and a pulmonary acoustic signal (P2); generating one or more 2D feature maps including a 2D lung feature map having a lung acoustic signal (P2), wherein a first axis of the map is a time axis and a second axis of the map is an axis spanning a plurality of individual heartbeats; and applying the pre-trained neural network to correlate the generated one or more 2D feature maps with a training dataset of previously acquired and segmented training 2D feature maps for a group of PH and non-PH subjects to obtain an indication of the presence of pulmonary hypertension.
[0045] Further disclosed is a computer-implemented system for non-invasively estimating pulmonary hypertension, PH, from a heart sound signal, the system comprising an electronic data processor, the electronic data processor comprising: receiving an acoustic signal (S2) acquired from the subject's beating heart over a predetermined period of time; Splitting the acoustic signal (S2) into an aortic acoustic signal (A2) and a pulmonary acoustic signal (P2); generating one or more 2D feature maps including a 2D lung feature map having a lung acoustic signal (P2), wherein a first axis of the map is a time axis and a second axis of the map is an axis spanning a plurality of individual heartbeats; applying the pre-trained neural network to correlate the generated one or more 2D feature maps with a training dataset of previously acquired and segmented training 2D feature maps for a group of PH and non-PH subjects to obtain an indication of the presence of pulmonary hypertension; is configured to run
[0046] In one embodiment, the system includes a digital stethoscope for acquiring pulsatile heart sound signals, the digital stethoscope being connected to an electronic data processor for transmitting the acquired pulsatile heart sound signals.
[0047] In one preferred example, the electronic data processor is further configured to divide the acquired acoustic signal into a plurality of time windows of predetermined duration, each time window including a peak of the heartbeat sound signal, the predetermined duration being preferably 200 milliseconds.
[0048] In one embodiment, the electronic data processor is further configured to align the segmented acoustic signal time windows by aligning peaks of the heartbeat sound signals in the segmented acoustic signal time windows.
[0049] In one embodiment, the electronic data processor is embodied on a mobile device.
[0050] The following drawings provide preferred embodiments for illustrating the present disclosure and should not be viewed as limiting the scope of the present invention. [Brief explanation of the drawings]
[0051] [Figure 1]1 is a graphical representation of the S2 speech signal, the proposed A2 speech signal, the proposed P2 speech signal, and the average feature attributions over time. [Figure 2] 2A, 2B, and 2C are graphical representations of an embodiment in which the subject's speech data and their respective descriptions are visualized as a three-channel image. [Figure 3] 1 is a graphical representation including ROC curves for the disclosed method (herein named DeepPHDet), GMM, and SVM models, demonstrating the superior predictive performance of DeepPHDet over existing state-of-the-art baselines. [Figure 4] 1 is a flowchart representation of an embodiment of a method for non-invasively estimating pulmonary hypertension, PH, from heart sound signals. DETAILED DESCRIPTION OF THE INVENTION
[0052] Detailed Description An algorithm for noninvasively detecting pulmonary hypertension (PH) is disclosed. Heart sounds are collected from a subject with a digital stethoscope and then passed to the algorithm as audio input. The output is an estimate of PH and a description of the relevant region of the subject's heart sounds. Audio recordings of heart sounds are collected from a spot on the subject's lung, for example, within the second intercostal space to the left of the sternum. The algorithm has a training phase and an evaluation phase, where the training phase acquires domain knowledge from multiple subject recordings, and the evaluation phase makes predictions for individual subjects.
[0053] In one embodiment, predictions are made for subjects not considered during the learning phase.
[0054] In one embodiment, the algorithm begins with a set of preprocessing steps, including filtering, spike removal, and segmentation of S2 sounds from the heart sound recording. A source separation algorithm is then applied to the S2 sounds in the recording to separate each S2 sound into its aortic (A2) and pulmonary (P2) components. The acquired S2 sounds, along with their corresponding A2 and P2 components, are organized and normalized into a feature map of the same shape as the three-channel image provided as input to the deep neural network. During training, several steps are performed to ensure the algorithm works with small data sets, including the use of batch gradient descent, channel normalization for unit variance, and zero-padding to ensure all inputs have the same heart rate. During evaluation, the algorithm predicts whether a subject has PH. Additionally, the heart sound components, moments, and heart rates that are most informative for prediction are emphasized using post-hoc explanatory attribution and ensemble techniques.
[0055] Instead, one of two optimization methods is applied to train the network to ensure that the algorithm works on small data sets. The first optimization method uses the following steps: using batch gradient descent rather than mini-batch gradient descent or stochastic gradient descent (the parameter update step occurs once per iteration over the entire dataset, but gradients can be accumulated using mini-batches); channel normalization to unit variance; and in some cases, zero-padding or cropping of inputs to ensure all inputs have the same number of segmented heartbeats. The second optimization method also uses the following steps: assigning fixed parameter weights to the neural network that are not updated by training, and this optimization method uses an analytically derived solution to train the final classification layer or model; sparse regularization can be used; and compression can be used.
[0056] To ensure that the model works efficiently on the above dataset, we also made architectural choices, including: choosing a convolutional layer architecture to compute parallel convolutions sequentially to reduce RAM usage; choosing a kernel width large enough to match the sampling rate of the acoustic signal; and judicious use of pooling to perform a mathematical transformation of the input signal into a compressible latent space, and a compression method to reduce the size of the embedding space.
[0057] The model's predictions depend primarily on the proposed P2 heart sound, especially in the region of most variability between 30 ms and 60 ms.
[0058] [Table 1]
[0059] A personal dataset of 42 subjects was acquired at the University Hospital Center of Porto, Portugal. Summary statistics in Table 1 show 29 subjects with PH and 13 subjects without PH. Of the affected subjects, the majority (20 of 29) were women. The age and heart rate of both the positive and affected populations were similar. Inclusion and exclusion criteria were unknown. PH was defined as positive if the subject had a mean pulmonary arterial pressure (MPAP) greater than 25 mmHg or a pulmonary arterial systolic pressure (PASP) greater than 30 mmHg. For each subject, ground truth pulmonary arterial pressure was obtained by right heart catheterization, accompanied by a 5-minute PhotoCardioGram (PCG) recording. This recording was obtained with the subject resting in a supine position in a relatively quiet clinical environment. Auscultation was performed over the left second intercostal space using a custom-made cabled stethoscope connected to a Rugloop Waves® system. Heart sounds were recorded at a sample rate of 8 kHz, and heart sound amplitude was quantified with 16-bit resolution. This dataset is not publicly available to protect privacy.
[0060] For each 5-minute speech signal, heartbeats were segmented and extracted within a 200-ms window for each S2 sound of the heartbeat. The start time of the window was selected so that the peak positions of all S2 sounds for that subject were aligned in time (i.e., the peak times were aligned). The S2 signal was filtered with a second-order Butterworth filter at cutoff frequencies of 25 Hz and 400 Hz, resampled at 1 kHz, and cleaned by removing spikes using the method in non-patent literature 14. Then, the A2 and P2 components were separated as proposed by non-patent literature 15. Source separation assumes that the aortic and pulmonary components maintain the same waveform across the entire heartbeat, and that the delay between components within a single heartbeat varies due to changes in intrathoracic pressure in different respiratory layers. These two components are recovered by alternating optimization of a least-squares problem.
[0061] In one embodiment, the duration of each window of the audio signal is a user-predefined set of parameters, for example 200 ms for a 1 kHz sample rate.
[0062] In one example, alignment and segmentation produces a multi-channel 2D representation of the speech data, including S2, proposed A2, and proposed P2. Each 2D channel has 200 columns representing a 200 ms window, and as many rows as there are heartbeats. The channels are then zero-padded to 454 rows for all subjects with the same shape. Normalization to unit variance helps stabilize gradient backpropagation by reducing the risk of vanishing or exponential gradients.
[0063] In one embodiment, we considered the DenseNet121, ResNet18, and EfficientNet-B0 architectures.
[0064] In one embodiment, pre-trained deep network initialization improves performance, especially for small datasets.
[0065] In one embodiment, random and ImageNet initializations were considered.
[0066] We present DenseNet trained from random initialization, ResNet18 from ImageNet initialization, and EfficientNet-B0 from standard adversarial ImageNet initialization. These models were all trained with batch gradient descent at a learning rate of 0.0001 and momentum of 0.5 for 150 epochs. Deep networks are typically trained with mini-batch gradient descent on large databases. To stabilize gradient updates, we perform batch gradient descent, which means performing a gradient update once per iteration across the entire dataset. To work with datasets of any size, we compute a gradient update once per sample and maintain a sum or running average until a gradient update occurs. The loss is a balancing weight for the positive class. (outside 1) Weighted binary cross entropy using JPEG2025528500000003.jpg1718.
[0067] To evaluate the predictive performance of deep networks against classical methods, Gaussian mixture models (GMM) and support vector machines (SVM) were implemented.
[0068] Our GMM implementation adapts the current state-of-the-art work in [6], where one GMM is trained for the positive class and the other for the negative class. The test sample class is the GMM model with a higher posterior negative log-likelihood. To achieve the best performance in this baseline, we developed a different preprocessing pipeline and correspondingly optimized the GMM model to have two components and spherical covariance. The SVM uses a radial basis function (RBF) kernel and a slack parameter C = 1. Only the S2 channel was used for preprocessing. The addition of the proposed A2 and P2 channels adversely affects performance due to overfitting. Each row of each heartbeat, i.e., the S2 channel, was transformed with a 1D short-time Fourier transform (SFT) using a 64-sample FFT and a hop length of 2 samples, and the energy spectrum was calculated using the absolute value. The subject data, a tensor of shape (H,33,101), was simplified to (33,101) by computing the 98% quantile over H heartbeats. The channels were zero-padded to 454 rows, normalized to unit variance, and then flattened as vectors before being passed to the SVM and GMM models.
[0069] All models were evaluated using stratified 10-fold cross-validation. To report performance, predicted probabilities for the validation set were stored for each fold. There was one predicted probability per subject. The area under the receiver operating characteristic (ROC) curve (ROC AUC) and standard classification metrics were reported. Classification metrics require a threshold to convert probabilities to classes. A threshold Tk was chosen for each kth fold to maximize the difference between the true positive rate and the false positive rate on the training set ROC curve for the kth fold. This threshold optimizes the balanced accuracy score for the training set. Validation performance was calculated within each fold, and then metrics were aggregated by average across folds and epochs 100-150.
[0070] To better understand which portions of the proposed A2 and P2 channels contribute to PH detection, we applied the integrated gradient attribution method (Non-Patent Document 13).
[0071] In one embodiment, a DenseNet121 model is trained on 10 folds, resulting in 10 independently trained models. Therefore, 10 attributions are calculated for each heartbeat in the dataset, and then averaged to obtain one attribution per channel, or summed to obtain one importance score per heartbeat. For better visualization, the attributions are converted to magnitudes using absolute values, and then clipped to 1% and 99% of their values (with 1% and 99% as the lower and upper bounds). Clipping aids visualization because gradient-based attribution methods generate some outliers.
[0072] [Table 2]
[0073] The results in Table 2 show that the DenseNet121 and EfficientNet-B0 deep networks outperform state-of-the-art machine learning models by a large margin on the considered PH dataset. The DenseNet121 model has the best performance with a ROC AUC of 0.95, the highest Balanced Accuracy (BAcc), and the highest Matthews Correlation Coefficient (MCC). The two best-performing models are DenseNet121 and EfficientNet-B0.
[0074] The bottom row of Table 2 shows that the availability of S2, A2, and P2 channels improves performance relative to the use of S2 alone. The motivation for deep learning is to overcome the need for data-driven feature generation and preprocessing with larger datasets. In small-data regimes, such as our case, preprocessing has been observed to improve performance. Furthermore, the over-parameterized nature of deep networks has necessitated a rethinking of the current state-of-the-art interpretation of under- and overfitting. While classical models such as SVM and GMM overfit with additional parameters from the A2 and P2 channels, deep networks improve.
[0075] Figure 1 shows a graphical representation of the S2 speech signal, the proposed A2 speech signal, the proposed P2 speech signal, and the average feature attributions over time.
[0076] The top three rows of Figure 1 visualize the heart sound data of one subject. Each line represents a single heartbeat. The top row shows the S2 signal. The second and third rows show the proposed source-separated signals A2 and P2. The signals shown are normalized to unit variance and represent the inputs passed to the predictive model.
[0077] Empirically, normalization has been found to improve performance; normalization ensures that the quieter P2 has a similar amplitude to the louder A2. The A2 signal is very well defined due to the alignment of the heartbeat based on its peak. The distance between the A2 and P2 components varies depending on factors such as whether the subject is inhaling or exhaling, as well as the presence of PH. Therefore, current domain knowledge is consistent with the visual perception that the mean value of the P2 signal should exhibit less variability over time. In this example, P2 is observed to have the most variable behavior between 30 ms and 60 ms. Current domain knowledge expects PH to be related to changes in the timing and amplitude of P2.
[0078] The bottom plot in Figure 1 shows the average attribution across all heartbeats, along with the 99.9% confidence interval. In this example, the attribution to P2 dominates, coinciding with the most variability in P2 behavior between 30 ms and 60 ms. Both observations suggest that the deep network is consistent with domain knowledge. The attribution to A2 is strongest at the peak just before 25 ms. The attributions demonstrate that the availability of separate components facilitates prediction.
[0079] 2A, 2B, and 2C show graphical representations of an embodiment in which the subject's speech data and their respective descriptions are visualized as a three-channel image.
[0080] The first row is the input to the CNN, and the second and third rows are the output of the attribution method. The first three columns are the S2 component, proposed A2 component, and proposed P2 component. The fourth column represents a summary plot of the waveform across all beats. In these plots, all beats are zero-padded, and the second and third rows use inputs normalized to unit variance before calculating attributions.
[0081] Figure 3 shows graphical representations including ROC curves for the disclosed method (herein named DeepPHDet), GMM, and SVM models, demonstrating the superior predictive performance of DeepPHDet over existing state-of-the-art baselines.
[0082] It is found that deep networks improve detection performance; separating S2 into A2 and P2 can improve performance and improve the explainability of the model and analyzed S2 signals; the proposed A2 and P2 are consistent with domain knowledge; and post-hoc explanations validate the usefulness of domain knowledge and the segmentation of A2 and P2.
[0083] Therefore, the present invention contributes to the advancement of the current state of the art in the automated detection of pulmonary hypertension, i.e., pulmonary arterial hypertension, from heart sounds. The present invention has several advantages, including: high predictive performance; amenability to training with small and large datasets; description of A2 and P2 components to explain prediction; and requiring only heart sound data for estimation.
[0084] Deep networks trained on a private dataset of preprocessed digital stethoscope recordings achieved ROC AUC scores of 0.95 and 0.93, providing improvements of +0.17 and +0.15 over previous state-of-the-art adaptations based on Gaussian mixture models, and +0.07 and +0.05 over current machine learning implementations.
[0085] Post-hoc explanations and improved performance indicate that separating the S2 signal into A2 and P2 components aids detection.
[0086] FIG. 4 shows a flow chart representation of an embodiment of a method for non-invasively estimating pulmonary hypertension, PH, from heart sound signals.
[0087] In one embodiment, the entire model is trained by backpropagation.
[0088] In other embodiments, during training, the weights of the convolutional network are initialized and fixed (never modified), and the remaining steps are obtained by extreme learning machine or regression model techniques.
[0089] Tests were performed on three datasets acquired by stethoscope (PCG) and seismocardiogram (PCG) devices, including human and pig recordings. The human dataset, PCG, is: 42 human subjects underwent right heart catheterization; Heart sounds were recorded with a digital stethoscope; 13 patients without PH and 29 patients with PH; The pig dataset, PCG+SCG: 10 pigs, each undergoing right heart catheterization; The dataset size is 125 "pig patients" (with sampling sessions from 10 pigs); Each pig undergoes multiple episodes of chemically induced hypertension. Heart sounds are recorded at selected intervals; Recording devices: phonocardiography (PCG) and psychocardiography (SCG); Seventy-three human insureds underwent right heart catheterization and psychocardiogram testing.
[0090] Each dataset was evaluated (by cross-validation) and also for "cross-domain generalization" (training on one dataset and evaluating on other datasets). The method was evaluated on records of varying record length.
[0091] [Table 3]
[0092] Each value describes the performance of 12 independently trained models, each with 10-fold cross-validation. Macro is the average across all folds. Micro describes the performance for each sample. auROC (area under Radius of Curvature) is the area under the ROC curve. AP (Average Precision) is the average precision score (area under the PR (Precision-Recall) curve).
[0093] [Table 4]
[0094] Each number describes the performance of 12 independently trained models, each with 10-fold cross-validation. Macro is the average across all folds. Micro describes the performance for each sample. auROC is the area under the ROC curve. AP is the average precision score (area under the PR curve). Model hyperparameters were tuned for the training partitions of the baseline human (PCG) and pig (PCG+SCG) datasets. Percentages compare cross-domain performance to baseline performance.
[0095] Results are similar, and in some cases slightly lower, for methods with the additional step of splitting the acoustic signal (S2) into an aortic acoustic signal (A2) and a pulmonary acoustic signal (P2) and applying a pre-trained neural network to correlate one or more generated 2D maps with a training dataset of previously acquired and split training 2D feature maps.
[0096] AP is the average precision (area under the precision-recall curve), and auROC is the area under the ROC curve. These two metrics, combined, give a good sense of the algorithm's performance, including the presence of class imbalance. The model hyperparameters are the same for all experiments.
[0097] Analysis of stethoscope heart sound data with deep networks is an effective, low-cost, and non-invasive solution for the detection of pulmonary hypertension.
[0098] The present invention uses deep networks to analyze heart sounds, has low resource costs, and is suitable for early screening.
[0099] The present invention has several advantages, such as description of regions of interest at individual heartbeats and across all heartbeats, and medical reliability of the model for subject-specific predictions.
[0100] Whenever used in this document, the word "comprising" is intended to indicate the presence of stated features, values, steps or components, but does not exclude the presence or addition of one or more other features, values, steps, components or groups thereof.
[0101] The present invention should in no way be seen as limited to the described embodiments, and those of ordinary skill in the art will envision many possibilities for modification thereof. The above-described embodiments can be combined.
[0102] The following claims further present particular embodiments of the present invention.
[0103] References [1] Hasan B, Hansmann G, Budts W, Heath A, Hoodbhoy, Jing ZC, Koestenberger M, Meinel K, Mocumbi AO, Radchenko GD, et al, “Challenges and special aspects of pulmonary hypertension in middle-to-low-income regions: Jacc state-of-the-art review”, Journal of the American College of Cardiology 2020;75(19):2463-2477 (Non-patent document 1) [2] Lau EM, Humbert M, Celermajer DS, “Early detection of pulmonary artificial hypertension”, Nature Reviews Cardiology 2015;12(3):143-155(Non-patent Document 2) [3] Taleb M, Khuder S, Tinkel J, Khouri SJ, “The diagnostic accuracy of doppler echocardiography in assessment of pulmonary artery systolic pressure: A meta-analysis”, Echocardiography 2013;30(3):258-265(Non-patent Document 3) [4] Oliveira JH, Renna F, Costa P, Nogueire D, Oliveira C, Fer258 reira C, Jorge A, Mattos S, Hatem T, Tavares T, Elola A, Rad A, Sameni R, Clifford GD, Coimbra MT, “The circordigiscope dataset: From number detection to number classification”, IEEE Journal of Biomedical and Health Informatics 2021;1-1(Non-patent Document 4) [5] Lang IM, Plank C, Sadushi-Kolici R, Jakowitsch J, Klepetko W, Maurer G, “Imaging in pulmonary hypertension”, JACC Cardiovascular Imaging 2010;3(12):1287-1295(Non-patent Document 5) [6] Kaddoura T, Vadlamudi K, Kumar S, Bobhate P, Guo l, Jain S, Elgendi M, Coe JY, Kim D, Taylor D, et al 2016;6(1):1-11(Vol.6) [7] Xu J, Durand L, Pibarot P, “Nonlinear transient chirp signal modeling of the aortic and pulmonary components of the second heart sound“, IEEE Transactions on Biomedical Engineering 2000;47(10):1328-1335(Chapter 7); [8] Andreev V, Gramovich V, Krasikova M, Korolkov A, Vyborov O, Danilov N, Martynyuk T, Rodnenkov O, Rudenko O, “Time-frequency analysis of the second heart sound to assess pulmonary artery pressure”, Acoustical Physics 2020;66(5):542-547(Chapter 8) [9] Dennis A, Michaels AD, Arand P, Ventura D, “Noninvasive diagnosis of pulmonary hypertension using heart sound analysis“.
[10] Renna F, Oliveira J, Coimbra MT, “Deep convolutional neural networks for heart sound segmentation”, IEEE journal of biomedical and health informatics 2019;23(6)2435-2445 (Non-Patent Document 10)
[11] Haug G, Liu Z, Van Der Maaten L, Weinberger KQ, “Densely connected convolutional networks”, In Proceedings of the IEEE conference on computer vision and pattern recognition, 2017;4700-4708 (Non-patent document 11)
[12] Tan M, Le Q, “Efficientnet: Rethinking model scaling for conventional neural networks”, In International conference on machine learning, PMLR, 2019;6105-6114 (Non-patent document 12)
[13] Sundararajan M, Taly A, Yan Q, “Acoustic attribution for deep networks”, In International conference on machine learning, PMLR, 2017;3319-3328 (Non-Patent Document 13)
[14] Schmidt SE, Holst-Hansen C, Graff C, Toft E, Struijk JJ, “Segmentation of heart sound recordings by a duration-dependent hidden Markov model”, Physiological measurement 2013;31(4):513 (Non-patent document 14)
[15] Renna F, Plumbley MD, Coimbra M, “Source separation of the second heart sound via alternating optimization”, In 2021 Computing in Cardiology (CinC), volume 48, IEEE, 2021; 1-4 (Non-Patent Document 15)
Claims
1. 1. A method for non-invasively estimating pulmonary hypertension (PH) from a heart sound signal, the method comprising: receiving an acoustic signal (S2) acquired from the beating heart of the subject over a predetermined period of time; generating one or more 2D feature maps including a 2D feature map having the received acoustic signal (S2), wherein a first axis of the 2D feature map is a time axis and a second axis of the 2D map is an axis spanning a plurality of individual heartbeats; applying a pre-trained neural network to correlate the generated one or more 2D maps with a training dataset of previously acquired and generated training 2D feature maps for a group of PH and non-PH subjects to obtain an indication of the presence of pulmonary hypertension; A non-invasive method for estimating pulmonary hypertension, including:
2. dividing the acoustic signal (S2) into an aortic acoustic signal (A2) and a pulmonary acoustic signal (P2); generating one or more 2D feature maps including a 2D lung feature map having the lung acoustic signal (P2), wherein a first axis of the 2D lung feature map is a time axis and a second axis of the 2D lung feature map is an axis spanning a plurality of individual heartbeats; applying a pre-trained neural network to correlate the generated one or more 2D feature maps with a training dataset of previously acquired and segmented training 2D feature maps for a group of PH and non-PH subjects to obtain an indication of the presence of pulmonary hypertension; The method for non-invasively estimating pulmonary hypertension according to claim 1 , further comprising:
3. 3. The method for non-invasively estimating pulmonary hypertension according to claim 2, wherein the one or more 2D feature maps include a 2D aortic feature map having the aortic acoustic signal (A2), a first axis of the 2D aortic feature map being a time axis, and a second axis of the 2D aortic feature map being an axis spanning multiple individual heartbeats.
4. 4. The method for non-invasively estimating pulmonary hypertension according to claim 2 or 3, wherein the one or more 2D feature maps comprise a 2D full-signal feature map having the received acoustic signal (S2), a first axis of the 2D full-signal feature map being a time axis, and a second axis of the 2D full-signal feature map being an axis spanning multiple individual heartbeats.
5. 5. The method for non-invasively estimating pulmonary hypertension according to claim 3 or 4, or claims 3 and 4, further comprising the step of combining the one or more 2D feature maps as a multi-channel input to the neural network, each of the one or more 2D feature maps being combined as one channel of the multi-channel input.
6. 6. The method for non-invasively estimating pulmonary hypertension according to claim 1, further comprising the step of dividing the acquired acoustic signal into a plurality of time windows of predetermined duration, each of the time windows including a peak of the heartbeat sound signal.
7. 7. The method of claim 6, further comprising aligning the time windows of the segmented acoustic signals by aligning peaks of the heartbeat sound signal in the time windows of the segmented acoustic signals.
8. Preferably, the method for non-invasively estimating pulmonary hypertension according to any one of claims 1 to 7 further comprises the step of calculating saliency attribution by integrated gradients or gradient times corresponding to the generated one or more 2D feature maps.
9. 9. The method for non-invasively estimating pulmonary hypertension according to any one of claims 1 to 8, comprising pre-processing the acquired acoustic signals by filtering, spike removal, normalization, registration and / or segmentation.
10. The method for non-invasively estimating pulmonary hypertension according to any one of claims 1 to 9, wherein the neural network is a convolutional neural network.
11. The method for non-invasively estimating pulmonary hypertension according to any one of claims 1 to 10, wherein the neural network is a deep neural network with excess parameters.
12. The method for non-invasively predicting pulmonary hypertension according to any one of claims 1 to 9, wherein the neural network is an extreme learning machine.
13. The method for non-invasively estimating pulmonary hypertension according to any one of claims 2 to 12, wherein the division is performed using alternating optimization of a least squares problem.
14. 14. The method for non-invasive estimation of pulmonary hypertension according to any one of claims 2 to 13, further comprising the steps of filtering the heart sound signal after splitting it with a second order Butterworth filter, in particular with a Butterworth filter with cut-off frequencies of 25 Hz and 400 Hz, resampling at 1 kHz and cleaning it by removing spikes.
15. 15. The method for non-invasively estimating pulmonary hypertension according to any one of claims 1 to 14, comprising the step of acquiring the heart sound signal at a lung spot of the subject, preferably over the second left intercostal space.
16. 1. A computer-implemented method for training a neural network for non-invasively estimating pulmonary hypertension (PH) from heart sound signals, comprising: For both the PH and non-PH subject groups: receiving an acoustic signal (S2) acquired from the beating heart of the subject over a predetermined period of time; generating one or more 2D feature maps including a 2D feature map having the received acoustic signal (S2), wherein a first axis of the 2D feature map is a time axis and a second axis of the 2D feature map is an axis spanning a plurality of individual heartbeats; applying a pre-trained neural network to correlate the generated one or more 2D feature maps with a training dataset of previously acquired and generated training 2D feature maps for a group of PH and non-PH subjects to obtain an indication of the presence of pulmonary hypertension; How to train neural networks, including:
17. dividing the acoustic signal (S2) into an aortic acoustic signal (A2) and a pulmonary acoustic signal (P2); generating one or more 2D feature maps including a 2D lung feature map having the lung acoustic signal (P2), wherein a first axis of the 2D lung feature map is a time axis and a second axis of the 2D lung feature map is an axis spanning a plurality of individual heartbeats; applying a pre-trained neural network to correlate the generated one or more 2D feature maps with a training dataset of previously acquired and segmented training 2D feature maps for a group of PH and non-PH subjects to obtain an indication of the presence of pulmonary hypertension; 17. The method of training a neural network of claim 16, further comprising:
18. 1. A computer-implemented system for non-invasively estimating pulmonary hypertension (PH) from heart sound signals, comprising: The system comprises an electronic data processor, the electronic data processor comprising: receiving an acoustic signal (S2) acquired from the beating heart of the subject over a predetermined period of time; generating one or more 2D feature maps including a 2D feature map having the received acoustic signal (S2), wherein a first axis of the 2D feature map is a time axis and a second axis of the 2D feature map is an axis spanning a plurality of individual heartbeats; applying a pre-trained neural network to correlate the generated one or more 2D feature maps with a training dataset of previously acquired and generated training 2D feature maps for a group of PH and non-PH subjects to obtain an indication of the presence of pulmonary hypertension; 2. A computer-implemented system configured to:
19. said electronic data processor dividing the acoustic signal (S2) into an aortic acoustic signal (A2) and a pulmonary acoustic signal (P2); generating one or more 2D feature maps including a 2D lung feature map having the lung acoustic signal (P2), wherein a first axis of the 2D lung feature map is a time axis and a second axis of the 2D lung feature map is an axis spanning a plurality of individual heartbeats; applying a pre-trained neural network to correlate the generated one or more 2D feature maps with a training dataset of previously acquired and segmented training 2D feature maps for a group of PH and non-PH subjects to obtain an indication of the presence of pulmonary hypertension; 20. The computer-implemented system of claim 18, further configured to:
20. 20. The computer-implemented system of claim 18 or 19, comprising a digital stethoscope for acquiring heart sound signals of the beating heart, the digital stethoscope connected to the electronic data processor for transmitting the acquired heart sound signals of the beating heart.
21. 21. A computer-implemented system according to any one of claims 18 to 20, wherein the electronic data processor is further configured to divide the acquired acoustic signal into a plurality of time windows of predetermined duration, each of the time windows including a peak of the heartbeat sound signal, and wherein the predetermined duration is preferably 200 milliseconds.
22. 22. The computer-implemented system of claim 18, wherein the electronic data processor is further configured to align the time windows of the segmented acoustic signals by aligning peaks of the heartbeat sound signals in the time windows of the segmented acoustic signals.