Systems, devices, and methods for computer-aided auscultation

WO2026174404A1PCT designated stage Publication Date: 2026-08-27ESI IMAGING INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CA2026/050281
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-24
Filing Date
2026-02-24
Publication Date
2026-08-27

Smart Images

  • Figure CA2026050281_27082026_PF_FP_ABST
    Figure CA2026050281_27082026_PF_FP_ABST
Patent Text Reader

Abstract

A system for diagnosing a health condition of a patient, a diagnostic device for diagnosing a health condition of a patient, and a method of diagnosing a health condition of a patient, are disclosed. A computer-implemented method of diagnosing a health condition of a patient comprises: receiving an acoustic signal obtained from a diagnostic device measuring an internal parameter of the patient; inputting a frequency domain representation of the acoustic signal into a machine learning model trained to determine a presence of the health condition in the patient based on the acoustic signal; and generating an output for display on a user device based on a result from the machine learning model. A diagnostic device comprises a signal acquisition module configured to acquire an acoustic signal based on structure-borne vibration from a patient, and a controller coupled to the signal acquisition module.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS, DEVICES, AND METHODS FOR COMPUTER-AIDED AUSCULTATIONCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and the benefit of US Provisional Application No. 63 / 762,152, filed on February 24, 2025, the entire contents of which is incorporated herein by reference for all purposes.TECHNICAL FIELD

[0002] The present disclosure relates to analyzing different sounds from inside the human body to assist with diagnosing various health conditions, in particular for diagnosing heart and / or respiratory conditions.BACKGROUND

[0003] Analysis of different sounds from inside the human body can be used to diagnose various health conditions. Different devices are used for this purpose and one of the most commonly used diagnostic devices is a stethoscope, which is used by doctors as part of physical examinations of patients. A stethoscope is a medical device used by doctors and other healthcare professionals to listen to sounds produced by the human body. When assessing a patient’s heart, doctor’s examine different areas of the chest where murmurs and clicks are likely to radiate from. Among other information, doctors must recognize the duration, shape and intensity of the murmur or click in order to effectively give a diagnosis. Doctors also use the stethoscope to detect lung conditions by examining the chest area for wheezing and crackles. In order to identify these conditions, doctors must have excellent diagnostic listening skills, known as auscultation, through extensive clinical experience.

[0004] Despite its importance, studies show that doctors and healthcare professionals generally struggle with identifying common heart and lung ailments using stethoscopes. This may be attributed to the doctor’s age and experience as well as the performance and fitting of the stethoscope. Current digital stethoscopes are expensive and typically just present acoustic data for doctors to interpret. The use of more advanced medical procedures such as electrocardiograms and magnetic resonance imaging is also more invasive and expensive.

[0005] Accordingly, systems, devices, and methods that assist with auscultation remain highly desirable.SUMMARY

[0006] In accordance with one aspect of the present disclosure, a computer-implemented method of diagnosing a health condition of a patient is disclosed, comprising: receiving an acoustic signal obtained from a diagnostic device measuring an internal parameter of the patient; inputting a frequency domain representation of the acoustic signal into a machine learning model trained to determine a presence of the health condition in the patient based on the acoustic signal; and generating an output for display on a user device based on a result from the machine learning model.

[0007] In some aspects, the acoustic signal is received in a time domain, and the method further comprises generating the frequency domain representation of the acoustic signal.

[0008] In some aspects, the acoustic signal is received in the frequency domain.

[0009] In some aspects, the method further comprises pre-processing the acoustic signal to extract one or more features from the acoustic signal, and the one or more features are input into the machine learning model.

[0010] In some aspects, the one or more features comprise a peak between 0-150 Hz for heart sound, a sharp decrease after the peak for heart sound, a local maxima around 200 Hz for heart murmur, or combinations thereof.

[0011] In some aspects, the machine learning model comprises one of naive Bayes, decision tree, or logistic regression.

[0012] In some aspects, the acoustic signal obtained from the diagnostic device is from a chest area of the patient.

[0013] In some aspects, the health condition is associated with a heart or lung of the patient.

[0014] In accordance with another aspect of the present disclosure, a diagnostic device is disclosed, comprising: a signal acquisition module configured to acquire an acoustic signal based on structure-borne vibration from a patient; and a controller coupled to the signal acquisition module, comprising: an input / output interface configured to receive the acoustic signal; and a communication module configured to communicate the acoustic signal to an external device.

[0015] In some aspects, the signal acquisition module comprises a stethoscope diaphragm and a microphone.

[0016] In some aspects, the signal acquisition module comprises a piezoelectric microphone.

[0017] In some aspects, the system further comprises a signal pre-processing module configured to pre-process the acoustic signal prior to sending to the controller.

[0018] In some aspects, the signal pre-processing module comprises an amplifier module to amplify the acoustic signal and a filter module to filter the acoustic signal.

[0019] In some aspects, the filter module comprises a low-pass filter with a cutoff frequency of 1000 Hz.

[0020] In some aspects, the communication module enables Bluetooth communication.

[0021] In some aspects, the system further comprises a portable battery to power the device.

[0022] In some aspects, the controller further comprises an analog-to-digital converter to convert the acoustic signal from an analog signal to a digital signal.

[0023] In some aspects, the diagnostic device is configured as a digital stethoscope.

[0024] In some aspects, the diagnostic device is wearable by a user.

[0025] In accordance with another aspect of the present disclosure, a system for diagnosing a health condition of a patient is disclosed, comprising: the diagnostic device of any one of the above aspects; and a computer communicatively coupled to the diagnostic device and configured to implement the computer-implemented method of any one of the above aspects.BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Further features and advantages of the present disclosure will become apparent from the following detailed description, taken in combination with the appended drawings, in which:

[0027] FIG. 1A shows a block diagram of a system for computer-aided auscultation for diagnosing a health condition of a patient;

[0028] FIG. 1 B is a block diagram of a computer device comprising part of the system depicted in FIG. 1A;

[0029] FIG. 2A shows an amplifier and filter schematic;

[0030] FIG. 2B shows a plot of the signal pre-processing schematic frequency response;

[0031] FIG. 3A shows an audio signal of a heartbeat;

[0032] FIG. 3B shows audio signal activity shapes for common murmurs;

[0033] FIG. 3C depicts a slope detection method that may be used for feature extraction when assessing a heart condition;

[0034] FIGs. 4A-C show frequency domain plots of different audio signals;

[0035] FIG. 4D shows plots from a 5-second pulmonary recording;

[0036] FIG. 5 shows a computer-implemented method for diagnosing a health condition of a patient based on acoustic data;

[0037] FIG. 6 shows a hierarchal navigation flow that may be used for an application provided by the system; and

[0038] FIG. 7 shows an example of a user interface showing a diagnostic recording page of the application.

[0039] It will be noted that throughout the appended drawings, like features are identified by like reference numerals.DETAILED DESCRIPTION

[0040] The present disclosure provides a diagnostic device, in particular a bioacoustic auscultation device, and related systems and computer-implemented methods for diagnosing a health condition of a patient, which can be used to help healthcare professionals detect heart murmurs and clicks as well as wheezing and crackles in the lungs. The devices, systems, and methods disclosed herein provide an inexpensive, comprehensive, and fast solution that may be used for detecting various heart conditions, such as precursors to heart attacks, artery blockage, heart failure monitoring, hemodynamic instability detection (blood flow instability), etc., and / or respiratory system conditions such as pneumonia, asthma, pulmonary edema detection, Chronic Obstructive Pulmonary Disease (COPD) exacerbation, etc.

[0041] In some embodiments, the system may be implemented as a machine learning-enabled auscultation system that operates across multiple timeframes. At short timescales, milliseconds to seconds, high-resolution acoustic signals (optionally along with data from other sensors for multi-sensor fusion) are processed using frame-level feature extraction and convolutional models to detect subtle events such as murmurs, crackles, wheezes, and abnormal heart sound timing. At intermediate timescales, seconds to minutes, sequence models analyze beat-to-beat and breath-to-breath variability to identify persistent patterns, irregular rhythms, and evolving acoustic features that may not be evident within a single cycle. For long-term monitoring over weeks or months, attention-based models such as transformers analyze session-level embeddings derived from each recording. These models learn which prior timepoints are most relevant, enabling detection of gradual disease progression, rare but clinically significant events, and treatment response trends. By modeling non-local temporal relationships without relying on fixed memory decay, attention mechanisms could support robust longitudinal risk stratification and early identification of cardiopulmonary deterioration.

[0042] The diagnostic device and associated systems and methods can thus reduce human interpretation by processing acoustic signals from a body of a patient (e.g. from the chest area) and recognizing abnormal patterns in the acoustic signal. The results can be displayed on a user device and an application interface may be used to present the output and provide various other functionality such as playback and / or visualization of the audio recordings for further analysis.

[0043] Accordingly, the systems, devices, and methods in accordance with the present disclosure can help healthcare professionals with the detection of common heart and lung ailments by using a diagnostic device with sensors around the chest area to capture and amplify sounds, convert them to digital signals, and analyze the signals to detect possible conditions. The systems, devices, and methods in accordance with the present disclosure can reduce or even possibly eliminate human interpretation, and assists in detection of various health conditions with more accuracy and objectivity. An advantage of the systems, devices, and methods disclosed herein over current alternatives is that it is less invasive and expensive compared to advanced medical equipment and may even be configured as a wearable device. Due to the manner in which data is analyzed and output the system can also be used in the training of young medical professionals.

[0044] Embodiments are described below, by way of example only, with reference to Figures 1A-7.

[0045] FIG. 1A shows a block diagram of a system 100 for computer-aided auscultation for diagnosing a health condition of a patient. The system 100 comprises a bio-acoustic diagnostic device 101 used for measuring a parameter of a patient for assisting with diagnosis, and a user device 102 used for analyzing data received from the diagnostic device 101, which are in communication with each other over a communications network. The system 100 can be used for detecting various heart and lung conditions such as murmurs and wheezing, as described in more detail below.

[0046] Signal Acquisition module 110 is used to measure an internal parameter of a patient. The signal acquisition module 110 may comprise a stethoscopediaphragm 112 and a microphone 114. For example, the stethoscope diaphragm 112 may pick up cardiac sounds of the patient to audible levels, which is then fed to a microphone 114. In some embodiments, the microphone 114 may be an air-coupled microphone, where air pressure of the sound waves change the microphone’s capacitance via its impedance, which produces a voltage swing proportional to input sound waves. In other embodiments, the signal acquisition module 110 may comprise a piezoelectric contact microphone (i.e. a piezoelectric disc coupled to a diaphragm) such as a CM-01B piezoelectric contact microphone, which transduces strain (i.e. structure-borne vibration) into voltage. In still other embodiments, the signal acquisition module 110 may comprise accelerometer I inertial sensors to provide accelerometer-based chest-wall sensing that captures mechanical vibrations directly.

[0047] The signal acquisition system is important because the real-time signal analysis depends on the quality of the signal. For signal acquisition module 110 comprising a diaphragm and air-coupled microphone, the microphone may be an electret microphone that is attached to the chest piece for capturing the vibrations that the diaphragm produces. The chest piece comprises a diaphragm that is placed against patients for detecting sounds. Sounds from the human body vibrate the diaphragm creating acoustic pressure waves, which travel up the tubing to the microphone. Several criteria may be taken into consideration for the selection of the stethoscope, with the main criteria of concern being the performance and durability of the chest piece, as the diaphragm of the stethoscope should be able to detect the heart sounds with high accuracy. Electret microphones are commonly used in designing digital stethoscope due to their small size, excellent frequency response, and reasonable cost. An “electret” is a thin, Teflon™-like material with a fixed charge bonded to its surface. The electret is contained between two electrodes, and the structure forms a capacitor which contains a fixed charge. The charge on the microphone is fixed, and varying the capacitance causes the voltage on the capacitor to change. Electret condenser microphones have an internal JFET which buffers the microphone capacitor. The voltage signal produced by sound modulates the gate voltage of the JFET causing a change in the current flowing between the drain and source of the JFET.

[0048] For a signal acquisition module 110 comprising a piezoelectric contact microphone, a central advantage for a digital stethoscope use-case is that the sensor naturally emphasizes vibrations transmitted through tissue and the chest wall rather than airborne sound. A piezoelectric contact microphone with an integrated buffer amplifier may provide for body-sound pickup with a low-frequency cutoff on the order of 8 Hz and useful bandwidth extending into the kHz range, which overlaps key content for heart sounds (dominant energy often below 150 Hz for S1 / S2, with murmurs extending higher) and lung sounds (commonly 100 Hz to 1000 Hz with higher-frequency components for certain adventitious sounds). Use of a piezoelectric contact microphone may be particularly compelling for a wearable and / or portable digital stethoscope because it targets common failure modes of air-coupled microphones, such as:Ambient-noise rejection (physics-level): structure-borne pickup suppresses airborne speech / room noise relative to chest vibrations, reducing SNR dependence on a perfect acoustic seal.Compact mechanics: a small contact puck plus adhesive / strap mount can be simpler than a traditional bell / diaphragm cavity while still coupling strongly to the body.Wearable feasibility: low power draw at the sensor front-end (buffered piezo film) and direct vibration pickup suit long-term monitoring designs.

[0049] One drawback of using a piezoelectric contact microphone is motion artifact when the sensor or cable moves relative to the skin, causing low-frequency and impulsive artifacts to dominate. Accordingly, mounting, strain relief, and artifact removal are important considerations to ensure system performance.

[0050] Accelerometer-based chest-wall sensing captures mechanical vibrations directly and can be relatively immune to airborne noise. Recent wearable mechano-acoustic systems highlight that skin-coupled vibration sensing can enable robust physiological monitoring (including cardiopulmonary signals) in practical settings. However, limitations include motion-induced artifacts during gross movement and the need for careful mechanical coupling and mounting compliance to preserve diagnostic bands.

[0051] Signal Pre-Processing module 120 comprises an amplifier module and a filter module. An operational amplifier is a DC-coupled high gain electronic voltage amplifierwith a differential input and usually a single-ended output. Since the acoustic signal coming out of the microphone is in the order of millivolts, it is not significant enough to be converted to digital signal for further signal processing. Thus, an operational amplifier is used to enlarge the gain of the signal. A low pass filter passes signals with a frequency lower than a certain cutoff frequency and attenuates signals with frequencies higher than the cutoff frequency. Statistical analysis (discussed below) proved that the frequency for abnormal heart and lung sounds is typically under 1000 Hz. Accordingly, the results conclude that the frequencies above 1000 Hz are not significant in determining whether patients suffer from heart or lung conditions. Thus, a low pass filter having cut-off frequency of 1000 Hz may be used.

[0052] FIG. 2A shows an amplifier and filter schematic 200. FIG. 2B shows a plot 210 of the signal pre-processing schematic frequency response.

[0053] An electret condenser microphone, such as that used as an air-coupled microphone, has an internal JFET, which is biased by two 1kQ resistors (i.e. Ri and R2) by connecting them in parallel to the microphone output 201. The resistors value is calculated using the supply voltage (VDD), the microphone operating voltage (VMIC) and current consumption (Is) using the following equation.R = (VDD ~ VMICy / Is

[0054] The DC offset of the microphone output is eliminated by passing the signal through the 10pF capacitor Ci . Since the input range of the analog-to-digital converter of the microcontroller is 0 to 5V, to avoid saturation the circuit is biased at 2.5V. The resistors (R3 and R4 ) are connected in parallel to bias the circuit at 2.5 V. To rise the amplitude of the signal, a non-inverting TL081 operational amplifier may be used.

[0055] The circuit gain for frequencies in the range of 20-1000 Hz is determined by the following equation.

[0056] A capacitor C3 and resistor Re are connected in series in the feedback loop of the op-amp to restrict the amplification of the 2.5 V bias at the positive node of the op-amp by creating a unity gain system. The capacitors C4 and Cs are added in parallel to VDD and ground of the TL081 op-amp for the stability of the circuit. The 10 kQ resistor is added at the output of the circuit to protect the analog input pin of the microcontroller.

[0057] The resistor Rs and capacitor C2 form a low-pass filter as shown. The corner frequency of this filter is 1000 Hz. To design this filter, the following equation is used to calculate the capacitor value when Rs is 1 MQ and fcutoff is 1000 Hz.159.15 pF

[0058] A value of 160 pF is selected for designing the low pass filter. The resulting frequency response of the circuit shown in FIG. 2A is displayed in the plot 210 of FIG. 2B. Forgetting the frequency response, the microphone was replaced by a function generator in Nl Multisim.

[0059] Referring back to FIG. 1A, Power System module 130 comprises a portable battery 132, which can be used to provide power to the bio-acoustic diagnostic device 101, including to microcontroller 140.

[0060] The microcontroller 140 comprises input / output (I / O) interface / ports 141 for receiving inputs from the signal pre-processing module 140 and power from the power system 130. The microcontroller 140 also comprises a voltage / power regulator 142, an analog-to-digital converter (ADC) 144, a data I / O software module 146, and a Bluetooth™ communications module 148. The power regulator 142 of the microcontroller 140 is used to regulate and provide power to the signal acquisition module 110 (e.g. the electret microphone condenser) and the signal pre-processing module 120. An analog input pin of the microcontroller takes the filtered and amplified signal as input and converts it into digital form using the analog-to-digital converter (ADC) 144. The software module 146 of the microcontroller is responsible forcoordinating the read from input ports, quantizing the signal, and writing to output ports. The software module 146 also interfaces with the Bluetooth communications module 148 in order to transmit data to and receive commands from a user device 102, which may be running a corresponding application for receiving data from the device 101, processing / analyzing the data, and outputting results. For example, the Bluetooth module 148 is used to transmit data wirelessly from the microcontroller to the external application, and this module may receive commands from the external application to signal recording of audio.

[0061] Bluetooth Low Energy (BLE) may be a suitable communications means compared to other alternatives because of its low power consumption, cost, and sufficient range, however it will be appreciated that alternative types of communication networks could be used with a corresponding type of communication module, and accordingly the use of a Bluetooth communication module is not limiting. The microcontroller 140 is used to process information from the signal pre-processing module 120 before sending to the user device 102 for analysis / diagnostics. The information received from the pre-processing module 120 is in analog form and needs to be converted into digital form for further signal analysis. This task is accomplished by the ADC converter 144, which is built-in in most microcontrollers. If BLE is selected as the wireless technology, the microcontroller needs to have an inbuilt BLE module. An Arduino Primo may for example be used as the microcontroller 140.

[0062] Additionally or alternatively, in some embodiments the microcontroller 140 may comprise another processing module with one or more processors that analyze the data on the microcontroller 140 (i.e. some or all of the data analysis is performed at the diagnostic device 101 instead of at the external device 102). However, to support this processing may require a larger microcontroller and ma thus increase the overall size and / or cost of the device 101.

[0063] A diagnostic device 101 in accordance with the present disclosure may thus comprise the signal acquisition module, signal pre-processing module, power system, and microcontroller of the system 100 shown in FIG. 1. The signal acquisition module is used to convert the heart and lungs sounds of the patient to analog signal and the signal pre-processing module is used to amplify the signal by a gain, e.g. of150 dB. The signal pre-processing module also filters the signal at a cutoff frequency (e.g. 1000 Hz) to remove the noise and unwanted frequencies before sending the analog signal to the microcontroller. The signal conversion system utilizes built in analog-to-digital converter on the microcontroller.

[0064] In manufacturing, since the device should be portable (and in some embodiments, wearable) and meet appropriate size specifications the components may be stacked on top of each other to form the device. The device should preferably be designed to be inexpensive to manufacture, and allows for heart murmur diagnosis accuracy because it returns a digital signal that is encoded using a moderately high bit resolution (e.g. 14 bits) and a relatively low percentage error relative to the Force Sensitive Resistor (FSR) of the input. The signal transmission subsystem can make use of an Arduino microcontroller, nRF52832 wireless technology, and a simple application add-on to successfully transfer required acoustic signal data from the microcontroller to the user device 102. The device can be tailored to desired functional requirements, such as device range (e.g. minimum 3 metres) and latency (e.g. maximum 10 seconds). In addition, the device can be optimized according to other requirements, such as cost, size, weight and battery life.

[0065] External Application module 150 represents an application running on an external user device 102 (e.g. a computer, a tablet, a smartphone, etc.) that receives and analyzes the acoustic signal data. The application provides a user interface for a user device 102 to analyze data received from the bio-acoustic device 101 and to allow a user to interact with the diagnostic device 101. The application is responsible for scanning, connecting and receiving data from the diagnostic device 101. A networking interface component 152 of the external application 150 is responsible for communicating with the Bluetooth module 148 of the diagnostic device 101 and receiving acoustic signal data. The data is then passed to the signal processing module comprising one or both of a time-domain feature extraction module 154 and a frequency-domain feature extraction module 156, which extracts features from the acoustic signal data and passes these to a classification module 158 for analysis. The classification module 158 analyzes and recognizes abnormal patterns and sounds of certain heart and lung conditions based on the extracted features.Once analysis is complete, the application presents diagnosis data and optionally the audio recording in a user-friendly graphical user interface (GUI) 160 for presentation to a user (e.g. a healthcare provider).

[0066] FIG. 1 B is a block diagram of a computer device that may be the user device 102 running the external application 150. The computer device comprises a processor 162 that controls the device’s overall operation. The processor 162 is communicatively coupled to and controls several subsystems. These subsystems comprise user input devices 164, which may comprise, for example, any one or more of a keyboard, mouse, touch screen, voice control, etc.; random access memory (“RAM”) 166, which stores computer program code for execution at runtime by the processor 162; non-volatile storage 168, which stores the computer program code executed by the RAM 166 at runtime; a display controller 170, which is communicatively coupled to and controls a display 172; and a communications / network interface 174, which facilitates network communications over the communications network with the diagnostic device 101. The non-volatile storage 168 has stored on it computer program code that is loaded into the RAM 166 at runtime and that is executable by the processor 162. When the computer program code is executed by the processor 162, the processor 162 causes the user device 102 to implement a method for diagnosing a health condition of a patient based on acoustic data, such as is described in more detail herein below.

[0067] As further described below, the memory or storage can store algorithms, such as machine learning models. The server 108 may also comprise graphical processing unit(s) (GPU) to control a display and to run the machine learning models, which may comprise multiple GPUs that are run in parallel to reduce processing time. The processors (e.g. a central processing unit (CPU) and the GPU(s)) may be one or more processors or microprocessors, which are examples of suitable processing units, and may additionally or alternatively include an artificial intelligence (Al) accelerator, programmable logic controller, a microcontroller (which includes both a processing unit and a non-transitory computer readable medium), neural processing unit (NPU), orsystem-on-a-chip (SoC).

[0068] The diagnostic device 101 may use the Generic Attribute Profile (GATT) to define how the data is organized and exchanged between applications. The data in GATT is organized in services. Each GATT service is organized in GATT profiles, where each profile contains multiple services. A 16-bit UUID is used to distinguish the services. It provides a standardized structure for applications of a common type to exchange information. The diagnostic device 101 may act as a GATT server and the application on user device 102 acts as a client. The GATT server broadcasts its services, which are filtered by UUID to prevent other BLE device from showing up in the scan results. The characteristics of the device are used for receiving the data of the device. The diagnostic device 101 may use Generic Access Profile (GAP) to prevent connection from other devices. The Generic Access Profile (GAP) provides the framework for devices to broadcast data, discover other devices and create secure connections.

[0069] The external application 150 may start with checking if the Bluetooth and location services are enabled for the diagnostic device 101. If these services are disabled then it prompts the user to turn them on. Once the services are enabled, the scanning interface displays the diagnostic device 101 , while other BLE devices in the vicinity are filtered out. The user could touch the listed device to establish a connection. Once the connection is established, analysis is performed on the data transmitted by the diagnostic device 101, as described further herein below. The diagnostic device 101 may remain connected to the user device 102 even if the application is running in the background but the data transmission may be stopped to save power consumption of the device. This prevents the overhead of reconnecting each time when the user switches to a different application for a few minutes. The connection may be terminated between the diagnostic device 101 and a user device 102 if the application is closed or running in the background for more than a threshold period of time.

[0070] In order to diagnose health conditions such as heart murmurs, supervised learning may be used to train a machine learning model forming part of the classification module 158. The training dataset may include audio files for normal and extra heart sounds, murmurs, as well as miscellaneous sounds, however it willbe appreciated that different training datasets can be used for training the model to detect different types of health conditions, and thus the example types of health conditions disclosed herein is not limiting. Since the classification of the dataset is known, supervised learning may be used to obtain a function mapping a set of inputs to a known output. In general, 70% of the dataset may be used to train the model and the remaining 30% may be used for validation. Another important aspect of machine learning is determining the set of features to input into the model. The following paragraphs analyze features in both the time and frequency domain of signal data.

[0071] The heart comprises several valves that open and close to allow blood to flow through the body. During this process, blood is collected from veins and is dumped into the Right Atrium. Blood then moves to the Right Ventricle chamber through the Tricuspid valve where it awaits oxygenation. Next, it passes through the Pulmonic valve and into both lungs where it picks up oxygen and releases carbon dioxide. Once the blood is refreshed, it enters the heart again through an artery and is dumped into the Left Atrium. To end the process, blood flows to the Left Ventricle chamber through the Mitral Valve and finally to the rest of the body through the Aortic Valve.

[0072] FIG. 3A shows an audio signal 300 of a heartbeat. During a normal heartbeat, two distinctive sounds can be heard. The first heart sound or S1 is produced by the shutting of the Tricuspid and Mitral valves. The Pulmonic and Aortic valves open shortly after this happens. The second heart sound or S2 is produced when the Pulmonic and Aortic valves close and similarly, the Tricuspid and Mitral valves open right after. These pairs of valves work simultaneously to collect used blood from veins and pump out fresh blood into the human body in a periodic manner. The region between the first heart sound and the second heart sound is called Systole and the region between the second heart sound and the next first heart sound is called Diastole. As shown in FIG. 3A, the duration of the Systole is shorter than that of the Diastole.

[0073] There are several different types of murmurs and they all give an indication of irregular blood flow through the heart. They can be identified by the presence of whooshing, clicks, or extra sounds outside of the first and second heartsounds. The cause of murmurs can usually be attributed to stenosis or regurgitation in the valves of the heart. Stenosis occurs when there’s abnormal blood flow through valves that are open in the heart and regurgitation occurs when there’s any blood flow through valves that are closed. A summary of some common murmurs and their audio signal activity shapes are shown in FIG. 3B, which depicts a normal heart signal 310, an aortic stenosis signal 320, a mitral stenosis signal 330, a mitral regurgitation signal 340, and an aortic regurgitation signal 350.

[0074] In order to detect cardiac conditions, it is desirable to detect and segment the cardiac cycle (S1, systole, S2, diastole) to enable feature extraction for murmurs, split sounds, and extra sounds (S3 / S4). For pulmonary conditions, detection / classification of wheezes, crackles, and other adventitious sounds is typically framed as time-frequency event detection and cycle-aware classification.

[0075] From the activity of a heart signal, it can be seen that well defined peaks occur in relatively periodic intervals. It is unlikely that miscellaneous sounds such as a door closing or people talking would have such a characteristic. As a result, one time-domain based feature that can distinguish heart signals from miscellaneous signals is the presence of peaks in amplitude that occur periodically. These peaks represent the S1 and S2 sounds in a heart signal. A peak detection algorithm may be applied on all samples of the audio data. Audio data in time-domain is represented by an array of values where the index represents time and data at a particular index represents the amplitude of the signal.

[0076] Given an array A with n samples of data, A[i] is a peak if A[i-1] < A[i] > A[i+1] where A[0] and A[n] being the first and last data samples are ignored. In other words, a data sample is defined to be a peak if it is not smaller than its neighbours. One problem that arises from this algorithm is that it may find peaks that are not the first or second heart sounds, which define the bounds of the Systole and Diastole regions. To solve this problem, a threshold based on the top 5% of peaks may be introduced to filter out peaks that do not meet the minimum amplitude requirement to be considered a first or second heart sound. The absolute value of the signal is also taken since there are both positive and negative amplitude values.

[0077] Under the assumption that Systole is shorter than Diastole, a Systole is always followed by a Diastole and the durations of Systoles and Diastoles remain fairly constant over a recording, the two regions can be separated for feature extraction. The first differences of the peak indices give the length of the Systolic and Diastolic regions at either even or odd indices depending on when the recording was started. In a typical recording of a heart signal, these regions occur multiple times and their detection is crucial in distinguishing heart sounds from miscellaneous ones.

[0078] Once this feature is extracted from the dataset, its validity is verified using a Naive Bayes classifier. The Naive Bayes classifier is chosen due to its simplicity and ability to perform well given a small dataset. The accuracy results in using this feature for a Naive Bayes classifier is given in Table 1 below.

[0079] Table i.

[0080] Accordingly, peak detection of the first and second heart sounds does a decent job in spotting heart signals however it fails to identify murmurs effectively. Another classification algorithm that could be used in addition to or as an alternative to peak detection is slope detection. In order to detect slopes in the data and classify features as described above, a slope detection algorithm may be implemented. FIG.3C depicts a slope detection method 360 that may be used for feature extraction when assessing a heart condition.

[0081] Given an array of audio samples, the slope detection algorithm may be required to recognize crescendos, decrescendos and whooshing sounds with constant intensity. These features can be extracted by calculating the slope over audio samples. In order for an increase or decrease in loudness to qualify as a feature, it needs to meet some requirements. An increase or decrease in sound intensity must last for a minimum length of time and their slopes must exceed a minimum value.Whooshing sounds with constant intensity also need to be differentiated from ambient noise during recordings.

[0082] When determining slope, the absolute value of the audio samples is taken and a windowing technique may be applied. The length of the window is the minimum length of time that an increase or decrease of sound intensity must persist. In other words, data samples within the window must either be increasing or decreasing. If this behaviour is observed, the window can be increased to determine the duration of the crescendo or decrescendo. Otherwise, the window is moved to the next set of data samples.

[0083] With reference to FIG. 3C, the slope detection method 360 comprises receive an array of audio data (362) and segmenting the audio data into Systoles and Diastoles (362), which may be performed using the the peak detection algorithm described above.

[0084] The peak detection algorithm is then run on individual segments to detect a slope (366), i.e. to determine if crescendos, decrescendos or continuous sound with constant slope exist. If definitive slopes are not found (No at 366), then the data may be classified as a normal and healthy heartbeat (368). If extra sounds are detected in the Systole (Yes, in Systole, at 366), Aortic Stenosis or Mitral Regurgitation murmur may be present (370). On the other hand, if extra sounds are detected in the Diastole (Yes, in Diastole, at 366), Mitral Stenosis or Aortic Regurgitation may be present (372). These classifications may be based on the signal activity of common murmurs as described above. Murmurs can be further characterized based on the location of the chest where the signal is obtained. For example, if extra sounds are detected in Systole near the Aortic area of the chest, then the murmur may be identified as Aortic Stenosis.

[0085] However, in addition or as an alternative to the foregoing signal analysis techniques, it is worth noting that cardiopulmonary signals are non-stationary, so timefrequency representations can be useful for visualization and feature extraction. As discussed below, features in the frequency-domain are found to be more powerful for detecting and classifying health conditions based on acoustic signals.

[0086] Frequency domain characteristics analysis of acoustic signals is used to observe how the signal’s energy is distributed over a range of frequencies and find signal characteristics that can be used to differentiate heart murmur audio from normal heart audio and any other noise. A Fourier transform can be used to convert the signal from time to frequency domain. Fourier transform decomposes a function into sum up to an infinite number of sine wave frequency components and the spectrum of frequency components is the frequency domain representation of the signal. Other types of tools can also be used for frequency-domain analysis, such as Short-Time Fourier Transform (STST), wavelet transforms, and S-transforms.

[0087] The design process of extracting useful features from the frequency domain of audio signals involved the following steps:- Identifying relevant characteristics and patterns that has the potential to be different for noise, normal heart audio, and heart murmur audio; and- Performing tests on sample audio files to confirm how effective each characteristic or pattern is in identifying sounds from: normal heart, heart murmurs, and other noises.

[0088] Two design stages were followed during the process. Firstly, features in the frequency domain were identified by observing plots of samples audio files and tested individually using a Bayesian machine learning model. The next stage involved finding a statistical parameter that can help increase detection accuracy by making use of magnitude values in the frequency spectrum.

[0089] The first part involved finding features in the frequency domain for the different audio signals. Four main features were observed mainly through observation of frequency domain plots of several sample audio signals containing random noises, normal heart audio, and heart murmur audio. In addition, the knowledge of the fact that frequency of heart sounds is low in range between 20 and 150 Hz helped focus on important areas of the plots. The features were the following:1. For audio signal of heart sounds, the peak magnitude occurs between 0 to 150 Hz. For other noise audio signals, the peak magnitude occurs after 150 Hz.2. For heart murmur audio signal, there is a significant local maximum around the 200 Hz frequency area. Scanning this area for increase in magnitude values can help identify heart murmur signals.3. Heart audio signals have a sharp decrease compared to the gradual decrease of noise audio signals after the peak.4. After the peak, the heart signals decreased continuously with very few local maxima. However, the noise audio signals are unpredictable with many local maxima. To quantify the pattern, the frequency plot can be sampled every 30 Hz after the peak to ensure that the magnitude is decreasing in every step with the increase in frequency. The 30 Hz sampling gets rid of sampling through small local maxima. If there is not a constant decrease, the signal can be classified as noise.

[0090] Frequency domain plots of one audio sample from each class, i.e. a frequency domain plot 410 of a normal heart audio signal, a frequency domain plot 420 of a heart murmur audio signal, and a frequency domain plot 430 of an audio signal with a sample of noises, are shown in FIGs. 4A-C, respectively.

[0091] It can be seen that for both the normal heart audio signal shown in FIG.4A and the heart murmur audio signal shown in FIG. 4B, the peak magnitude is at a frequency between 0 to 150 Hz. Moreover, the magnitude sharply decreases after the peak and there are few local maxima after the peak. For the murmur signal in FIG.4B, a significant local maxima is observed around the 200 Hz frequency area, which is consistent with the second feature identified. FIG. 4C contains a noise audio signal and it contrasts the other signals discussed. The peak magnitude is after the 150 Hz point, the decrease after the peak is gradual, and there are many local maxima as the signal magnitude decreases with increasing frequency.

[0092] After identifying the three features in the frequency domain, the patterns were tested to find out if they are successful in differentiating between heart sound and noise. The testing was performed by training a Bayesian machine learning model with the features identified. There were several test cases followed: testing of each feature by itself to identify the success rate of each pattern and testing of acombination of the features to identify the degree to which each feature compliments others. Each test was performed on sample of 50 normal heart, 35 heart murmur and 50 noise audio signals. Each test was performed 10 times with sample files rearranged each time and the average accuracy of the 10 tests was used to ensure quality of test results. The results for each of the test cases performed are shown in Table 2.

[0093] Table 2. The accuracy in Table 2 is calculated using the following: Accuracy (%)=((Number of Correct Identification-Number of False Positives-Number of False Negatives)-(Number of Samples))*100%.

[0094] From Table 2, it can be seen that two of the features identified ((i) Peak between 0-150 Hz for heart sound + Local Maxima around 200 Hz for murmur, and (ii) Sharp decrease for heart sound + Local Maxima around 200 Hz for murmur) had a relatively good success rate at differentiating heart audio from other noises. Both ofthe features combined produced even better results. The feature of heart sounds having continuous decrease with few local maxima proved to provide poor results and was dropped as a feature to use in the final prototype. The feature of heart murmur audio signals having a local maxima around 200 Hz was the only feature used to differentiate normal heart sounds from murmur heart sounds. It was used in combination with features that performed well in differentiating from heart audio and noise, and provided around 65% accuracy in detection of heart murmurs.

[0095] Statistical parameters in combination with the features identified above may be used to analyze magnitude data of the frequency domain for a range of frequency and improve the accuracy in detecting heart sounds and heart murmurs. Five statistical parameters were considered: mean, median, standard deviation, skew (measure of symmetry of the distribution of magnitude values), and Kurtosis (measure of whether data are heavy-tailed or light-tailed relative to a normal distribution. After the selection of statistical parameters, testing was performed by training a Bayesian machine learning model with each statistical parameter and different combinations of the statistical parameters. Each test was performed on sample of 50 normal heart, 35 heart murmur and 50 noise audio signals. Each test was performed 10 times with sample files rearranged each time and the average accuracy of the 10 tests was used to ensure quality of test results. Table 3 shows the results for each of the test cases performed.

[0096] Table 3.

[0097] From Table 3, it can be seen that using a combination of mean, skew and kurtosis of magnitude values over the entire frequency spectrum produced the most accurate results. Further, it was found that combing mean, skew and kurtosis of magnitude values with the features of peak between 0-150 Hz for heart sound, sharp decrease for heart sound, and local maxima around 200 Hz for murmur resulted in an accuracy of 81 % for differentiating heart audio from noise and an accuracy of 75% for detection of heart murmur audio.

[0098] It will be appreciated that while the foregoing specifically looks at features for heart sounds, similar analysis may be performed to extract relevant features for lung sounds, which are typically 100 Hz to 1000 Hz, with higher-frequency sounds typically being used for detecting wheezes, crackles, and other adventitious components. FIG. 4D shows plots from a 5-second pulmonary recording using a piezoelectric contact microphone coupled to a microcontroller. Plot 442 shows a raw time-domain signal of voltage vs. time. Plot 444 shows a FFT representation of the signal. Plot 446 shows a STFT representation of the signal. Plot 448 shows a S-transform representation. The frequency domain representations highlight the cyclic time-frequency structure and transient components of the signal.

[0099] Accordingly, it will be appreciated that features can be extracted from the frequency domain representation of an acoustic signal and used for detecting health conditions in a patient. As such, a machine learning model may be trained to determine whether a converted analog signal from the stethoscope corresponds to either a patient with a normal heartbeat, or a patient with a murmur. The input to the machine learning model may be a frequency domain representation of the converted analog acoustic signal, optionally along with features extracted using one or more deterministic algorithms as part of a pre-processing of the acoustic signal such asthose described above, as well as any other features that may be useful to extract for classification, such as band energies, wavelet coefficients, Mel-frequency cepstral coefficients (MFCCs), etc. The machine learning algorithm may use a supervised learning approach, in which the correct results being predicted were already known prior to the experiments, and serve as a training set to train the model to predict other instances where results are not known.

[0100] To improve the machine learning model classification, artifact control, removal, and / or robustness should be considered in the input data. Artifact control and removal may be particularly important for wearable auscultation, especially for contact microphones.

[0101] For example, mechanical I hardware-level artifact control may be achieved using several techniques. Mounting and compliance control uses consistent pressure and a compliant interface to stabilize coupling; overly rigid mounting can amplify motion transients, while overly soft mounting can attenuate higher frequencies. As many modules include a cable, cable strain relief control uses strain relief and anchoring to reduce microphonics and rubbing artifacts (often the dominant “false” low-frequency energy). Front-end headroom and anti-aliasing ensures the analog chain does not saturate under motion transients; soft-clipping and clamp networks can prevent ADC rail hits but must be designed to avoid distorting the diagnostic band.

[0102] Signal processing may be performed for motion and friction artifacts. Processing strategies of an acoustic signal prior to input to the machine learning model may include one or more of:Band-limiting: Apply a high-pass (to suppress very-low-frequency drift) and a physiologically motivated band-pass for the target task (cardiac-focused vs. lung-focused). This is usually the first line of defense.Notch filtering: Remove power-line contamination (50 / 60 Hz) and harmonics if present; whether this is necessary depends on grounding and the analog chain.Time-frequency masking I gating: Use spectrogram-domain suppression (e.g., attenuate broadband impulsive bursts) when artifacts have distinct structure in the TF plane. This is commonly used in audio denoising and can be adapted to auscultation.Cycle-aware rejection: For heart sounds, segmentation algorithms (e.g., LR-HSMM, Logistic Regression Hidden. Semi-Markov Model) provide phase boundaries; segments with implausible morphology or extreme energy can be rejected / flagged for reacquisition

[0103] Robustness to artifacts can also be achieved with the machine learning models. When building ML models, robustness may be improved via: (i) data augmentation: additive noise, simulated friction impulses, gain variation, and time warping to match wearable variability; and / or (ii) multi-task or hierarchical labels: separate “signal quality” from pathology labels (“noisy / usable” vs. anomaly classes), mirroring how PhysioNet-style challenges treat uncertainty and noise.

[0104] Three choices for machine learning algorithms were evaluated for their ability and performance to detect various conditions in response to an input signal: Decision Trees, Naive Bayes Classification, and Logistic Regression.

[0105] The Decision T ree algorithm was tested to determine whether or not this algorithm would successfully classify whether a wave file belonged to a normal patient or a patient with a murmur. One of the main advantages of using the decision tree model is the ease in which the model can be visualized. Another advantage is that the algorithm has a relatively quick processing time when it comes to determining the correct classification for a wave file, as the tree traversal time is relatively fast, regardless of the scale of the dataset. However the disadvantages to using this model is that the initial nodes carry more weight in deciding the final classification of a particular item, as a result if the data is changed slightly, this may result in different branches being traversed as the data will be interpreted differently with regards to the top nodes, which will result in a different classification for items in the dataset. Furthermore based on the different criteria for leaves, decision trees may be too small or too large to make an accurate classification.

[0106] The advantages of using Bayes Theorem is that it is very easy to implement, as it is a simple model that relies on simple counts to calculate the membership probabilities for an item. Furthermore, the classifier does not require large amounts of data, as it will converge very quickly given that a conditional independence assumption (where it assumed that xi...Xd is conditionally independent given y) holds. However, the disadvantage is that the algorithm can’t determine answers to problems in which features are related to each other, as that would violate the conditional independence requirement that optimizes the algorithm’s performance.

[0107] A Logistic Regression Model is a model that relies on analyzing a dataset on the results of one or more independent predictor variables, these variables can output binary values. Thus, the purpose of logistic regression is to find the relationship between the outputs of the dependent variables with the set of independent predictor variables. The logistic regression algorithm will return a set of coefficients that define a logistic function that maps the predicted result to either 0 or 1. An advantage to using the Logistic Regression Model is that a user can easily tune the function to avoid overfitting by modifying the returned coefficients. Additionally, the model can update itself easily to handle any new incoming data through the usage of stochastic gradient descent, which is an iterative method of finding the minima of the function representing the difference between the new logistic function being fitted and the old one.

[0108] The criteria for assessing a machine learning algorithm were simplicity, robustness, speed and accuracy. Accuracy may be the most important criteria, because predicting the correct result as to whether or not a waveform corresponds to a murmur or a normal heartbeat is the ultimate objective, and the achievable accuracy must be maximized in order to justify the validity of the systems and methods disclosed herein. Next, Runtime is an important factor, because with constant additions of new data to the model, it is important that the model can train itself quickly with the new dataset in a short amount of time. Robustness describes how well the model adapts to the addition of new data. This criteria is important, because if the process of adding new data completely changes previous predictions then the modelis inconsistent and has no validity. Finally, simplicity was considered as a criteria because a simple model is both easy to implement and comprehend.

[0109] Simulations were performed simulations using a dataset of 124 unique wav files, and the accuracy is based off of over 100 random splits of files into training or test sets. The benchmark tests can be seen below in Table 4.

[0110] Table 4.

[0111] All the models are all relatively simple to understand and easy to implement, however another key factor that differentiates them is how robust each of the solutions are. The decision tree is lacking in this category, because any slight change in data could cause the final result to change dramatically, as the final result depends heavily on the results of the earlier stages of nodes. However with Naive Bayes this is not a problem because the algorithm simply depends on the features used to analyze a dataset being completely independent from each other, however it’s pitfall is that if the tester is wrong and the features are not independent of each other, then the predictions begin to lose value. Next, logistic regression is robust and handles change relatively well, because it implements a stochastic gradient descent in which it attempts to fit a new projected model by minimizing the squared difference between the new logistic function and the old one.

[0112] Accordingly, the Naive Bayes Classifier method is considered the most optimal choice of the models tested, as it offers the best accuracy in generating the correct classification. However, it exchanges speed for this increase in accuracy, though the model makes up for this flaw with it’s simplicity and robustness. It will be appreciated that other machine learning models could also be used, including deep learning networks such as convolutional neural networks (CNNs), recurrent neuralnetworks (RNNs), gated recurrent units (GRUs), hybrid CNN-LSTM (Long Short-Term Memory), audio transformer or attention-based models, self-supervised pretraining, foundation audio models, hybrid time-frequency attention, confidence scoring based on signal quality, transformer encoders, self-attention mechanisms, pretrained audio embeddings, etc. Multimodal fusion models may also be used to combine other sensor data with the acoustic signal based on multi-modal physiological sensing, e.g. acoustic data plus motion data (e.g. from an accelerometer) plus ECG data, using appropriate data fusion algorithms and / or fusion-based classification. In some embodiments secondary sensors may be used to enable artifact detection.

[0113] Consistent with the foregoing, FIG. 5 shows a computer-implemented method 500 for diagnosing a health condition of a patient based on acoustic data. The method may be performed at the user device 102 shown in FIG. 1, based on acoustic data received from the bio-acoustic diagnostic device 101. However, it will also be appreciated that some processing may instead be performed at the diagnostic device 101 with an appropriate processing unit, and the results could be output to the user device 102, though this would increase the size and cost of the diagnostic bio-acoustic device 101. Instructions for executing the method may be stored on a non-transitory computer-readable medium which, when executed by a processing unit, configure the processing unit to perform the method 500.

[0114] The method 500 comprises receiving an acoustic signal obtained from a diagnostic device measuring an internal parameter of the patient (502). In some aspects, the acoustic signal is received in a time domain, and the method further comprises generating the frequency domain representation of the acoustic signal (504). In other aspects, the acoustic signal may be received in the frequency domain, for example if the transformation to the frequency domain is performed at the diagnostic device.

[0115] The frequency domain representation of the acoustic signal is input into a machine learning model trained to determine a presence of the health condition in the patient based on the acoustic signal (506). For example, the machine learning model may be one of naive Bayes, decision tree, or logistic regression classifiers. The acoustic signal obtained from the diagnostic device may be from a chest area of thepatient. Accordingly, the health condition that the machine learning model is trained to identify may be associated with a heart or lung of the patient.

[0116] In some aspects, the acoustic signal may be pre-processed to extract one or more features from the acoustic signal, and the one or more features are input into the machine learning model. For example, the one or more features may comprise a peak between 0-150 Hz for heart sound, a sharp decrease after the peak for heart sound, a local maxima around 200 Hz for heart murmur, or combinations thereof.

[0117] An output is generated for display on a user device based on a result from the machine learning model (508).

[0118] While the above method 500 may particularly be useful for single session monitoring and diagnosis of a patient, it will also be appreciated that the method can be expanded upon for trend-based diagnosis by storing data generated by the method over time and performing further analysis of the stored data. Such further analysis could comprise: longitudinal trend modeling; change detection over time; patient-specific baselines; and / or risk trajectory prediction, and in particular to comparing current acoustic signatures to historical patient data, determining personalized adaptive thresholds, early detection via delta modeling, etc.

[0119] Various information may be displayed in the user interface, including the output from the machine learning model. An application presenting the user interface should allow users (i.e. medical professionals) to access data for all their patients and run the diagnostic tool for each patient. Various example types of data that may be displayed to a user and a corresponding description of the data that is displayed is listed in Table 5.

[0120] Table 5.

[0121] Different personas may be developed that affect how data is displayed in the application. Personas are fictional characters with different needs, experiences, behaviours and goals that might use a service or product. Personas help identify and empathize with the users of the system. For example, personas can be tailored to how doctors, nurses, and medical students may use the application in their daily routines.

[0122] Examples of personas are provided below.Andy is a fifty-five year-old family physician with more than thirty years of experience in the medical field. Prior to being a physician, Andy was a medical officer for the Canadian Armed Forces. While on duty, Andy suffered hearing damage and can no longer use his stethoscope effectively. He now helps patients out of a quiet office environment. Andy understands technology and enjoys using it in his personal and professional time.Beth is a thirty year-old nurse with the city of Toronto and has about five years of work experience in the profession. Beth is constantly on the move for home visits using her car. As a result, she prefers carrying as little weight as possible. Beth uses herstethoscope everyday but often in noisy environments. Beth is not technologically savvy and is limited to using a flip phone to contact patients.Cory is a twenty-five year-old medical student at the University of Calgary. Cory is training to be a doctor but his auscultation skills are inadequate. He uses his stethoscope infrequently to practice diagnosing heart conditions. Cory has access to patients but does not always have the supervision of a trained doctor. Cory uses his computer and tablet every day for school.

[0123] The application may be structured using a card sorting technique. Information can be sorted categorically, hierarchically, chronologically, lexicographically or based on location. FIG. 6 depicts a hierarchal navigation flow that may be used by an application provided by the system, however it will be appreciated that various different application designs may be supported and that additional and / or different features may be implemented. After the initial splash screen 602 indicating the app is launched, a home screen 604 is presented to the user. The home screen may include call-to-actions to view the list of patients in the system 610, run a diagnostic recording 620, and view the application settings 630. From the list of patients, users can view a specific patient 612 and pull up his or her agreement form 614. Users can also perform a diagnostic recording 620 fora specific patient from the patient info page. The recording page also gives access to playback and visualization of the audio data and presents the diagnostic results. Furthermore, the settings page 630 allows users to navigate to a device pairing page 632 to pair their tablet / device with the signal acquisition device, and also view software information 634. An example of a user interface 700 showing a diagnostic recording page of the application is shown in FIG. 7. In the user interface 700 is displayed a patient’s name, audio recording data, audio visualization, and diagnostic results.

[0124] It would be appreciated by one of ordinary skill in the art that the system and components shown in the figures may include components not shown in the drawings. For simplicity and clarity of the illustration, elements in the figures are not necessarily to scale and are only schematic. It will be apparent to persons skilled in the art that a number of variations and modifications can be made without departing from the scope of the invention as described herein.

[0125] It is contemplated that any part of any aspect or embodiment discussed in this specification can be implemented or combined with any part of any other aspect or embodiment discussed in this specification.

[0126] It should be recognized that features and aspects of the various examples provided above can be combined into further examples that also fall within the scope of the present disclosure.

[0127] When used in this specification and claims, the terms "comprises" and "comprising" and variations thereof mean that the specified features, steps, or components are included. The terms are not to be interpreted to exclude the presence of other features, steps, or components.

[0128] The invention may also broadly consist in the parts, elements, steps, examples and / or features referred to or indicated in the specification individually or collectively in any and all combinations of two or more said parts, elements, steps, examples, and / or features. In particular, one or more features in any of the embodiments described herein may be combined with one or more features from any other embodiment(s) described herein.

Claims

CLAIMS:

1. A computer-implemented method of diagnosing a health condition of a patient, comprising:receiving an acoustic signal obtained from a diagnostic device measuring an internal parameter of the patient;inputting a frequency domain representation of the acoustic signal into a machine learning model trained to determine a presence of the health condition in the patient based on the acoustic signal; and generating an output for display on a user device based on a result from the machine learning model.

2. The computer-implemented method of claim 1 , wherein the acoustic signal is received in a time domain, and the method further comprises generating the frequency domain representation of the acoustic signal.

3. The computer-implemented method of claim 1 , wherein the acoustic signal is received in the frequency domain.

4. The computer-implemented method of any one of claims 1 to 3, further comprising pre-processing the acoustic signal to extract one or more features from the acoustic signal, and wherein the one or more features are input into the machine learning model.

5. The computer-implemented method of claim 4, wherein the one or more features comprise a peak between 0-150 Hz for heart sound, a sharp decrease after the peak for heart sound, a local maxima around 200 Hz for heart murmur, or combinations thereof.

6. The computer-implemented method of any one of claims 1 to 5, wherein the machine learning model comprises one of naive Bayes, decision tree, or logistic regression.

7. The computer-implement method of any one of claims 1 to 6, wherein the acoustic signal obtained from the diagnostic device is from a chest area of the patient.

8. The computer-implemented method of any one of claims 1 to 7, wherein the health condition is associated with a heart or lung of the patient.

9. A diagnostic device, comprising:a signal acquisition module configured to acquire an acoustic signal based on structure-borne vibration from a patient; anda controller coupled to the signal acquisition module, comprising:an input / output interface configured to receive the acoustic signal; and a communication module configured to communicate data based on the acoustic signal to an external device.

10. The diagnostic device of claim 9, wherein the signal acquisition module comprises a stethoscope diaphragm and a microphone.

11. The diagnostic device of claim 9, wherein the signal acquisition module comprises a piezoelectric microphone.

12. The diagnostic device of any one of claims 9 to 11 , further comprising a signal pre-processing module configured to pre-process the acoustic signal prior to sending to the controller.

13. The diagnostic device of claim 12, wherein the signal pre-processing module comprises an amplifier module to amplify the acoustic signal and a filter module to filter the acoustic signal.

14. The diagnostic device of claim 13, wherein the filter module comprises a low- pass filter with a cut-off frequency of 1000 Hz.

15. The diagnostic device of any one of claims 9 to 14, wherein the communication module enables Bluetooth communication.

16. The diagnostic device of any one of claims 9 to 15, further comprising a portable battery to power the device.

17. The diagnostic device of any one of claims 9 to 16, wherein the controller further comprises an analog-to-digital converter to convert the acoustic signal from an analog signal to a digital signal.

18. The diagnostic device of any one of claims 9 to 17, wherein the diagnostic device is configured as a digital stethoscope.

19. The diagnostic device of any one of claims 9 to 17, wherein the diagnostic device is wearable by a user.

20. A system for diagnosing a health condition of a patient, comprising:the diagnostic device of any one of claims 9 to 19; anda computer communicatively coupled to the diagnostic device and configured to implement the computer-implemented method of any one of claims 1 to 8.