Diagnosing medical conditions using voice recordings and internal listening
By recording and analyzing audio and acoustic signals to calculate transfer functions, the method effectively detects pulmonary edema and other conditions, facilitating early intervention and reducing hospitalization risks.
Patent Information
- Application Number
- JP2022548568
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-03
- Filing Date
- 2021-02-21
- Publication Date
- 2025-09-18
- Estimated Expiration
- 2041-02-21
AI Technical Summary
Existing methods for detecting pulmonary edema, particularly in the early stages of decompensation in heart failure patients, are inadequate for frequent and convenient monitoring, often leading to delayed detection and severe fluid accumulation in the lungs.
A method involving the recording of audio and acoustic signals from a patient's speech and thoracic sounds, calculating a transfer function between these signals, and evaluating deviations from baseline functions to detect fluid accumulation and other conditions like interstitial lung disease.
Provides sensitive and frequent monitoring of fluid levels, enabling early detection and treatment of pulmonary edema, reducing the need for hospitalization by allowing patients or caregivers to administer timely interventions.
Smart Images

Figure 0007741558000026 
Figure 0007741558000027 
Figure 0007741558000028
Abstract
Description
[Technical Field]
[0001] The present invention relates generally to systems and methods for medical diagnosis, and more particularly to the detection and assessment of pulmonary edema. [Background technology]
[0002] Pulmonary edema is a common consequence of heart failure, resulting in fluid accumulation in the lung parenchyma and airspaces, which can lead to impaired gas exchange and cause respiratory failure.
[0003] Patients with heart failure can remain stable ("compensated") for long periods of time by taking appropriate medications. However, various unexpected changes can cause the patient's condition to become unstable and lead to "decompensation." At the beginning of the decompensation process, fluid leaks from the pulmonary capillaries into the interstitial space around the alveoli. When fluid pressure in the interstitial space increases, fluid leaks from the interstitial space into the alveoli, making breathing difficult. It is important to detect and treat decompensation early, before respiratory distress begins.
[0004] Various methods for detecting fluid accumulation in the lungs are known in the art. For example, PCT International Application Publication No. WO2017 / 060828 (Patent Document 1), the disclosure of which is incorporated herein by reference, describes an apparatus in which a processor receives the voice of a subject suffering from a pulmonary condition associated with excessive fluid accumulation. The processor analyzes the voice to identify one or more voice-related parameters, assesses the status of the pulmonary condition in response to the voice-related parameters, and generates an output indicative of the status of the pulmonary condition.
[0005] As another example, Mulligan et al. described the use of acoustic responses to detect lung fluid in their article titled "Detecting Local Lung Characteristics Using Audio Transfer Functions of the Respiratory System" (Non-Patent Document 1), presented at the 2009 Annual International Conference of the IEEE Engineering in Medicine and Biology Society (IEEE, 2009). The authors developed an instrument to measure changes in the distribution of lung fluid in the respiratory system. The instrument consists of a speaker that inputs a 0-4 kHz white Gaussian noise (WGN) signal into the patient's mouth and an array of four electronic stethoscopes linked via a fully adjustable harness, which are used to recover signals from the thoracic surface. The software system for processing the data utilizes adaptive filtering principles to obtain transfer functions that describe the input-output relationship of the signal as the amount of fluid in the lungs changes. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] WO2017 / 060828 [Patent Document 2] U.S. Patent Application Serial No. 16 / 299,178 [Non-patent literature]
[0007] [Non-Patent Document 1] Mulligan et al., "Detection of Local Lung Characteristics Using Respiratory System Audio Transmission Features," presented at the 2009 Annual International Conference of the IEEE Engineering in Medicine and Biology Society (IEEE, 2009). Summary of the Invention
[0008] Embodiments of the present invention described herein below provide improved methods and devices for detecting pulmonary conditions.
[0009] According to one embodiment of the present invention, there is provided a method for medical diagnosis comprising the steps of: recording an audio signal resulting from sounds spoken by a patient; and recording an acoustic signal output simultaneously with the audio signal by an acoustic transducer in contact with the patient's thorax. A transfer function is calculated between the recorded audio signal and the recorded acoustic signal, or between the recorded acoustic signal and the recorded audio signal. The calculated transfer function is evaluated to assess the patient's medical condition.
[0010] In some embodiments, evaluating the calculated transfer function comprises evaluating a deviation between the calculated transfer function and a baseline transfer function; and detecting a change in the patient's medical condition in response to the evaluated deviation. In one embodiment, detecting the change comprises detecting fluid accumulation in the patient's thorax. The method includes administering treatment to the patient in response to detecting the change to reduce the amount of fluid accumulated in the thorax.
[0011] Alternatively or additionally, evaluating the calculated transfer function comprises evaluating the patient for interstitial lung disease.
[0012] In one disclosed embodiment, the method includes administering a treatment to the patient to treat the assessed condition.
[0013] In some embodiments, recording the acoustic signal comprises removing heart sounds from the acoustic signal output by the acoustic transducer before calculating the transfer function, hi one embodiment, removing heart sounds comprises detecting intervals of occurrence of extraneous sounds, including heart sounds, in the acoustic signal and removing from the acoustic signal the intervals used to calculate the transfer function.
[0014] Alternatively or additionally, removing heart sounds comprises filtering out heart sounds from the recorded acoustic signals before calculating the transfer function. In one disclosed embodiment, recording acoustic signals comprises receiving at least first and second acoustic signals from at least first and second acoustic transducers, respectively, in contact with the rib cage, and filtering out heart sounds comprises applying a delay to the arrival of heart sounds in the second acoustic signal relative to the first acoustic signal when combining the first and second acoustic signals while filtering out the heart sounds.
[0015] Further alternatively or additionally, calculating the transfer function comprises calculating spectral components of each of the recorded speech signal and the recorded acoustic signal at a set of frequencies and calculating a set of coefficients representing a relationship between the respective spectral components, which in one embodiment are cepstral representations.
[0016] In some embodiments, calculating the transfer function comprises calculating a set of coefficients representing the relationship between the recorded speech signal and the recorded acoustic signal for an infinite impulse response filter.
[0017] Alternatively or additionally, calculating the transfer function comprises calculating a set of coefficients representing a relationship between the recorded speech signal and the recorded acoustic signal with respect to a predictor in the time domain, hi one embodiment, calculating the set of coefficients comprises applying a prediction error of the relationship when calculating adaptive filter coefficients associated with the recorded speech signal and the recorded acoustic signal.
[0018] In one disclosed embodiment, calculating the transfer functions comprises dividing the spoken sound into a plurality of different types of speech units and calculating separate respective transfer functions for the different types of speech units.
[0019] In some embodiments, calculating the transfer function comprises calculating a set of time-varying coefficients representing a temporal relationship between the recorded speech signal and the recorded acoustic signal. In one disclosed embodiment, calculating the set of time-varying coefficients comprises identifying the pitch of the spoken speech signal and constraining the time-varying coefficients to be periodic with the same period as the identified pitch.
[0020] Alternatively or additionally, calculating the transfer function comprises calculating a set of coefficients representing a relationship between the recorded speech signal and the recorded acoustic signal, and evaluating the deviation comprises calculating a distance function between the coefficients of the calculated transfer function and a baseline transfer function. In one embodiment, calculating the distance function comprises calculating respective differences between pairs of coefficients, each pair having a first coefficient in the calculated transfer function and a corresponding second coefficient in the baseline transfer function, and calculating a norm of all the respective differences. Further alternatively or additionally, calculating the distance function comprises observing differences between the calculated transfer functions in different health states and selecting a distance function in response to the observed differences.
[0021] According to yet another embodiment of the present invention there is provided an apparatus for medical diagnosis, comprising: An apparatus is provided having a memory configured to store a recorded audio signal from sounds spoken by a patient and a recorded acoustic signal output simultaneously with the audio signal by an acoustic transducer in contact with the patient's thorax. The processor is configured to calculate a transfer function between the recorded voice signal and the recorded acoustic signal, or between the recorded acoustic signal and the recorded voice signal, and evaluate the calculated transfer function to assess the patient's condition.
[0022] Additionally, according to one embodiment of the present invention, there is provided a computer software product having a non-transitory computer-readable medium having stored thereon program instructions that, when read by a computer, cause the computer to: receive an audio signal representing a sound spoken by a patient and an acoustic signal output by an acoustic transducer in contact with the patient's thorax simultaneously with the audio signal; calculate a transfer function between the recorded audio signal and the recorded acoustic signal, or between the recorded acoustic signal and the recorded audio signal; and evaluate the calculated transfer function to assess the patient's condition. [Brief explanation of the drawings]
[0023] The present invention will be more fully understood from the following detailed description of the embodiments thereof, taken in conjunction with the following drawings: [Figure 1] 1 is a schematic diagram of a system for detecting a pulmonary condition, according to one embodiment of the present invention. [Figure 2] 2 is a block diagram that schematically illustrates details of elements of the system of FIG. 1, in accordance with one embodiment of the present invention. [Figure 3] 1 is a flow chart that schematically illustrates a method for detecting a pulmonary condition, in accordance with an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0024] (overview) The early stages of decompensation in patients with heart failure may be asymptomatic. By the time symptoms appear and the patient experiences signs of distress, the patient's condition may progress rapidly. Often, by the time the patient seeks medical attention, is seen, and begins treatment, fluid accumulation in the lungs has become severe enough to require hospitalization and prolonged medical intervention. Therefore, frequent (even daily) monitoring of patients is desirable to detect early signs of fluid accumulation in the thorax. Monitoring techniques must be simple enough to be administered by the patient or their family, yet sensitive enough to detect small, subtle changes in fluid levels.
[0025] The embodiments of the invention described herein address the need for frequent and convenient monitoring by recording sounds spoken by a patient and comparing them with sounds transmitted through the patient's rib cage to an acoustic transducer in contact with the thoracic body surface. (Such transducers are used in electronic stethoscopes known in the art, and the process of listening to and recording sounds at the body surface is called auscultation.) Fluid accumulation is known to affect both speech and chest sounds. Techniques using each of these types of sounds alone have been developed to detect pulmonary edema. However, in this embodiment, the relationship between these two types of sounds in a given patient is monitored to provide a much more sensitive indicator of changes in fluid levels.
[0026] Specifically, in disclosed embodiments, a patient or caregiver attaches one or more acoustic transducers to the patient's thorax at one or more predetermined locations. The patient then speaks into a microphone. A recording device, such as a cell phone running an appropriate application, records the audio signal from the microphone (in the form of a digitized electrical signal) and simultaneously records the digitized acoustic signal output by the acoustic transducer. A processor (either within the recording device or a remote computer) calculates a profile of correspondence between the audio and acoustic signals in the form of a transfer function between the recorded audio signal and the recorded acoustic signal, or between the recorded acoustic and audio signal.
[0027] The term "transfer function" is used herein and in the claims in the same sense as in the field of communications to refer to the functional relationship between two time-varying signals. As shown in the embodiments described below, transfer functions can be linear or nonlinear. To calculate a transfer function, one of the signals (a recorded speech signal or a recorded acoustic signal) is treated as the input signal, and the other is treated as the output signal. (In contrast to actual communications signals, the choice of input and output signals is arbitrary in this case.) Transfer functions are typically represented by a set of coefficients that can be calculated based on the "input" and "output" in either the time domain or the frequency domain. Various types of transfer functions that can be used for this purpose, including both time-invariant and time-varying transfer functions, are described below, along with methods for calculating them.
[0028] The processor examines the transfer function to detect changes in the patient's condition, particularly fluid accumulation in the thorax, which may prompt medical personnel to administer treatment to the patient, such as initiating or increasing the dosage of appropriate medications, such as diuretics or beta-blockers.
[0029] Transfer function tests can be patient-independent or patient-specific. Patient-independent tests use knowledge gathered by testing the transfer functions of large numbers of people with various health conditions to determine characteristics that distinguish the transfer functions of healthy people from those of people with a particular medical condition. For example, if the transfer function is represented in the frequency domain, a discriminatory characteristic might include the ratio between the average power of the transfer function in two different frequency bands.
[0030] For patient-specific tests, the processor evaluates deviations between the calculated transfer function and a baseline transfer function. This baseline can include or be derived from one or more transfer functions calculated for this same patient during a healthy period. Additionally or alternatively, the baseline can be based on samples collected across a larger patient population. Significant deviations may indicate a change in the patient's condition, particularly fluid accumulation within the thorax.
[0031] In some embodiments, a second, "pulmonary edema" baseline transfer function can be compared to the calculated baseline function. This second baseline transfer function may be equal to or derived from a transfer function calculated for this same patient during a period of pulmonary edema. Additionally or alternatively, the second baseline transfer function can be based on samples collected across a larger patient population when those patients experienced pulmonary edema. Small deviations from the "pulmonary edema" baseline may indicate a change in the patient's condition, particularly fluid accumulation in the thorax. In some cases, the only available baseline may be the "pulmonary edema" baseline transfer function. For example, if patient monitoring begins when the patient is hospitalized for acute pulmonary edema. In this case, an alert is generated if the pulmonary edema deviation from the baseline becomes too small. In other cases, both the "stable" baseline transfer function and the "pulmonary edema" baseline transfer function are available, and an alert is generated if the deviation from the pulmonary edema baseline becomes too small and the deviation from the "stable" baseline becomes too large.
[0032] As described above, embodiments of the present invention are particularly useful for detecting and treating changes in body fluid levels due to heart failure. Additionally or alternatively, these techniques can be applied to the diagnosis and treatment of other conditions that can cause pulmonary edema, such as high altitude, adverse drug reactions, etc. For example, if a patient is about to travel to a high altitude or be treated with a medication that carries a potential risk of pulmonary edema, a baseline can be obtained before entering the at-risk state (i.e., while still at low altitude or before taking the medication). The patient can then be monitored using the methods described above with a check frequency appropriate for their condition.
[0033] In addition to pulmonary edema, there are other conditions that can alter the acoustic conductance properties of the lung, such as interstitial lung disease, which causes the alveolar walls to thicken and stiffen, and these conditions affect the transfer function and can therefore be detected using the methods of the present invention.
[0034] (System Description) Reference is now made to Figures 1 and 2, which show a schematic illustration of a system 20 for detecting a pulmonary condition, in accordance with an embodiment of the present invention. Figure 1 is a diagram, and Figure 2 is a block diagram showing details of elements of the system.
[0035] In the illustrated embodiment, the patient 22 emits sounds into an audio microphone 24, such as a microphone that is part of a headset 26 connected to a user device 30, such as a smartphone, tablet, or personal computer. The patient can be prompted to emit a particular sound, for example, via an earpiece of the headset 26 or the screen of the user device 30, or he can speak freely. Alternatively, the microphone 24 can be built into the user device 30, or it can be a freestanding unit connected to the user device 30 by a wired or wireless connection.
[0036] The acoustic transducer 28 is placed in contact with the patient's thorax before the patient begins speaking. The acoustic transducer 28 may be included in an electronic stethoscope, such as the Littmann® Electronic Stethoscope manufactured by 3M (Maplewood, Minnesota), which is held in place by the patient or caregiver. Alternatively, the acoustic transducer 28 may be a special-purpose device that can be attached to the thorax using adhesive, a suction cup, or a suitable belt or harness. While this type of acoustic transducer is shown in the figures with only a single acoustic transducer positioned on the subject's thorax, in alternative embodiments, one or more acoustic transducers may be placed in different locations around the thorax, such as on the subject's back. Additionally or alternatively, the acoustic transducer 28 may be permanently fixed to the patient's body, for example, as part of a subcutaneous control unit for a pacemaker or intracardiac defibrillator.
[0037] As shown in FIG. 2 , the acoustic transducer 28 includes a microphone 36, such as a piezoelectric microphone, which contacts the skin of the chest directly or through a suitable interface. A front-end circuit 38 amplifies, filters, and digitizes the acoustic signal output by the microphone 36. In an alternative embodiment (not shown), the same front-end circuit 38 also receives and digitizes the audio signal from the audio microphone 24. A communications interface 40, such as a Bluetooth® wireless interface, transmits the resulting stream of digital samples to the user device 30. Alternatively, the front-end circuit 38 can communicate the acoustic signal in analog form to the user device 30 via a wired interface.
[0038] The user device 30 includes a communications interface 42 that receives the audio signals output by the microphone 24 and the acoustic signals output by the acoustic transducer 28 via a wired or wireless link. A processor 44 of the user device 30 records the signals as data in a memory 46, such as random access memory (RAM). Typically, recordings of the signals from the microphone 24 and the acoustic transducer 28 are synchronized with one another. This synchronization can be achieved by synchronizing the sampling circuitry used to acquire and digitize the signals, or perhaps by using the same sampling circuitry for both microphones 24 and 36, as described above. Alternatively, the processor 44 can synchronize the recordings based on acoustic events occurring in both the audio and acoustic signals, such as part of the patient's speech or artificially added sounds, such as clicks generated at regular intervals by the audio speaker of the user device 30. A user interface 48 of the user device 30 outputs instructions to the patient or caregiver, for example, via the headset 26 or on a display screen.
[0039] In this embodiment, the processor 44 transmits the recorded signals as data over a network 34, such as the Internet, to a server 32 for further analysis. Alternatively, or additionally, the processor 44 can perform at least a portion of the analysis locally within the user device 30. The server 32 includes a network interface controller (NIC) 50 that receives and passes the data to the processor 52 and communicates the data to the server's memory 54 for storage and subsequent analysis. While FIG. 1 shows only a single patient 22 and user device 30, in practice the server 32 will typically communicate with multiple user devices and provide services to multiple patients.
[0040] As described in more detail below, processor 52 calculates a transfer function between the recorded voice signal and the recorded acoustic signal, or between the recorded acoustic signal and the recorded voice signal. Processor 52 evaluates the deviation between the calculated transfer function and a baseline transfer function and reports the results to patient 22 and / or a caregiver. Based on this deviation, processor 52 can detect a change in the patient's condition, such as an increase in fluid accumulation in the patient's thorax. In this case, server 32 typically alerts a medical professional, such as the patient's physician, who can then prescribe treatment to reduce the fluid accumulation.
[0041] Processor 44 and processor 52 typically comprise general-purpose computer processors that perform the functions described herein under the control of appropriate software. This software may be downloaded to the processors in electronic form, for example, via network 34. Additionally or alternatively, the software may be stored on a tangible, non-transitory computer-readable medium, such as an optical, magnetic, or electronic memory medium. Still additionally or alternatively, at least some of the functions of processors 44 and 52 may be performed by dedicated digital signal processors or hardware logic circuitry.
[0042] (Methods of signal analysis and evaluation) 3 is a flow chart that schematically illustrates a method for detecting a pulmonary condition, in accordance with an embodiment of the present invention. For clarity and convenience, the method is described with reference to elements of system 20 as shown in FIGS. 1-2 and described above. Alternatively, the principles of the method may be implemented in virtually any system capable of simultaneously recording and analyzing speech and chest sounds for both the detection of pulmonary edema and other medical conditions. All such alternative implementations are considered within the scope of the present invention.
[0043] The method begins with the acquisition of an input signal. Microphone 24 acquires sounds spoken by patient 22 and outputs an audio signal, in an audio acquisition step 60. Simultaneously, acoustic transducer 28 is held in contact with the patient's thorax, acquires chest sounds, and outputs a corresponding audio signal, in a simultaneous internal sound listening step 62. Processor 44 records the signals in digital form in memory 46. As previously described, the audio and audio signals are synchronized by processor 44 by synchronous sampling during acquisition or thereafter, for example, by aligning the acoustic features of the recorded signals.
[0044] In this embodiment, processor 44 transmits the digitized raw signal to server 32 for further processing. Accordingly, the steps following Figure 4 are described below with reference to elements of server 32. Alternatively, some or all of these processing steps may be performed locally by processor 44.
[0045] The processor 52 stores the data received from the user device 30 in memory 54 and filters the data to remove background sounds and other noises. The processor 52 filters the audio signal in an audio filtering step 64 using audio processing methods known in the art to remove interference from background noise. The processor 52 filters the audio signal from the audio transducer 28 in an acoustic filtering step 66 to remove heart sounds, digestive peristalsis, and chest sounds not directly related to the patient's speech, such as wheezing. For example, in steps 64 and 66, the processor 52 may detect abnormal sounds in the audio and / or audio signals and simply ignore the time intervals in which the abnormal sounds occurred. Alternatively or additionally, the processor 52 may actively suppress background sounds and noise.
[0046] Detection of abnormal sounds can be done in several ways. In some cases, the inherent acoustic characteristics of the abnormal sound may be used. For example, in the case of a heartbeat, its normal periodicity can be used. The period and acoustic characteristics of the heartbeat can be detected during periods of silence when the patient is not speaking and then used to detect the heartbeat during speech.
[0047] As explained below, the transfer function can be expressed as a predictor of the chest signal using the microphone signal. The prediction error is the difference between the actual chest signal and the predicted value. In some embodiments, the prediction error is calculated, and a significant increase in its power, or its power in a particular frequency band, indicates the presence of an extraneous signal.
[0048] When multiple acoustic transducers are used, sound waves emitted from an internal source arrive at each transducer with slightly different delays and attenuations (which may vary across different frequency bands). These differences in delay and attenuation vary depending on the location of the source. Therefore, extraneous sounds arriving from sources such as the heart or digestive system can be detected because their relative delays differ from the relative delay of sound. Based on this, in some embodiments, processor 52 receives signals from multiple acoustic transducers attached to the patient's body and uses the relative delays to combine the signals while filtering out irrelevant sounds. In some embodiments with multiple acoustic transducers, beamforming techniques known in the field of microphone arrays can be used to suppress the gain of extraneous sounds arriving from directions different from the sound.
[0049] In one embodiment, for example, processor 52 detects heart sounds in the acoustic signal output by acoustic transducer 28, thus determining the heart rate. Based on this, processor 52 calculates a matched filter in step 66 that is matched to the spectrum of the heart sounds in the spectral or time domain, and uses the matched filter to suppress the contribution of the heart sounds to the acoustic signal.
[0050] In another embodiment, for example, processor 52 uses an adaptive filter to predict the acoustic signal caused by the heartbeat in the acoustic signal of the previous heartbeat and subtracts the predicted heartbeat from the recorded signal, thereby substantially canceling the effect of the heartbeat.
[0051] (Transfer function estimation) After filtering the signals, processor 52 calculates a transfer function between the recorded speech signal and the recorded acoustic signal in a correspondence calculation step 68. As explained above, the transfer function is conveniently represented as a transfer function h(t), which predicts one of the two signals as a function of the other. In the following discussion, the speech signal x output by microphone 24 is M (t) is the relation x S = h * x M The acoustic signal x output by the acoustic transducer 28 is S For computational purposes, the acoustic signal can be arbitrarily delayed by a short period, say a few milliseconds, if desired. Alternatively, the procedure described below can be used, mutatis mutandis, to predict x S x as a function of M This can be applied to calculate a transfer function that predicts
[0052] In some embodiments, the processor 52 calculates the transfer function H(ω) in the spectral domain, where the transfer function is S The spectral components at a set of frequencies {ω} of (ω) are expressed as the audio signal X M In terms of frequency, the signal is sampled at a specific sampling frequency, so the frequency components of the signal and the transfer function can be calculated as a set of coefficients representing the spectral components of the signal H(e iω ),X M (e iω ),X S (e iω), where |ω| ≦ π, where ω is the normalized frequency (equal to 2π times the actual frequency divided by the sampling frequency). The transfer function coefficients for each frequency component ω are given by:
number
[0053] Usually, X S and X M The frequency components of are calculated at N discrete frequencies using an appropriate transform function such as the Discrete Fourier Transform (DFT). The quotient in equation (1) is e 2πin / N , n = 0,…,N-1 denotes the coefficients of H at N equally spaced points on the unit circle, defined as
[0054] Alternatively, H can be expressed more compactly in terms of a cepstrum, e.g., in the form of cepstral coefficients. k , -∞ < k < ∞ is log((H(e iω ) are the Fourier coefficients of the signal x M and x S Since is real-valued, the sequence of cepstral coefficients is conjugate symmetric, i.e.,
number
number
number
[0055] In an alternative embodiment, the processor 52 calculates a transfer function in terms of a set of coefficients that represents the relationship between the recorded speech signal and the recorded acoustic signal as an infinite impulse response filter:
number
number
number
number
[0056] The above equations implicitly assume that a single time-invariant transfer function is calculated between the audio signal recorded by microphone 24 and the acoustic signal from acoustic transducer 28. However, some embodiments of the present invention do not rely on this assumption.
[0057] From a physical perspective, the process of speech production consists of three main stages: excitation, modulation, and propagation. Excitation occurs when airflow from the lungs is restricted or intermittently blocked, generating an excitation signal. Excitation can be caused by the vocal cords intermittently blocking the airflow, or by higher articulatory organs such as the tongue or lips blocking or constricting the airflow at various points in the vocal tract. The excitation signal is modulated by reverberation within the vocal tract and possibly also in the tracheobronchial space. Finally, the modulated signal propagates through both the nose and mouth, where it is received by microphone 24, and through the lungs and chest wall, where it is received by acoustic transducer 28. The transfer function between the microphone and acoustic transducer varies depending on the location of the excitation; therefore, the transfer function may be different for different phonemes.
[0058] The term "phoneme" generally refers to distinct phonetic elements of speech. To clarify terminology, "phonetic" refers to sounds produced by the subject's respiratory system and can be captured with a microphone placed in front of the subject. "Speech" is sound representing a specific syllable, word, or sentence. Our paradigm is based on having subjects speak, i.e., produce prescribed text or sounds freely selected by the subject. However, in addition to speech, recorded speech may include a variety of additional, often involuntary, non-speech sounds, such as wheezing, coughing, yawning, interjections ("um," "hmm"), and sighs. Such sounds are generally captured by the acoustic transducer 28 and result in characteristic transfer functions depending on the location of the excitation that produces them. In embodiments of the present invention, these non-speech sounds, to the extent they occur, can be treated as additional speech units with their characteristic transfer functions.
[0059] Thus, in one embodiment, the processor 52 divides the spoken sound into multiple different types of speech units and calculates separate transfer functions for each of the different types of speech units. For example, the processor 52 may calculate phoneme-specific transfer functions. To this end, the processor 52 may identify phoneme boundaries by using a reference speech signal of the same linguistic content for which the phoneme boundaries are known. Such a reference speech signal may be based on speech previously recorded from the patient 22, or on speech by another person or synthesized speech. The signals from the microphone 24 and acoustic transducer 28 are nonlinearly aligned to the reference signal (e.g., using dynamic time warping), and then the phoneme boundaries are mapped back from the reference signal to the current signal. Methods for identifying and aligning phoneme boundaries are further described in U.S. Patent Application No. 16 / 299,178, filed March 12, 2019, the disclosure of which is incorporated herein by reference.
[0060] After separating the input signal into phonemes, processor 52 calculates transfer functions for each phoneme individually or for groups of similar types of phonemes. For example, processor 52 can group together phonemes produced by excitation at the same location in the vocal tract. Such grouping allows processor 52 to reliably estimate transfer functions over relatively short recording times. Processor 52 can then calculate one transfer function for all phonemes in the same group, such as all glottal consonants or all alveolar consonants. In either case, the correspondence between the signals from microphone 24 and acoustic transducer 28 is defined by multiple phoneme- or phoneme-type-specific transfer functions. Alternatively, processor 52 can calculate transfer functions for other types of speech units, such as diphones or triphones.
[0061] In the above embodiment, the processor 52 calculates the transfer function between the signal from the microphone 24 and the signal from the acoustic transducer 28 in terms of a set of linear, time-invariant coefficients (in either the time or frequency domain). This type of calculation can be performed efficiently and results in a compact numerical representation of the transfer function.
[0062] However, in an alternative embodiment, at least some of the coefficients of the transfer function calculated by processor 52 are time-varying, representing the temporal relationship between the speech signal recorded by microphone 24 and the acoustic signal recorded by acoustic transducer 28. This type of time-varying representation is useful for analyzing voiced sounds, particularly vowels. In these sounds, the vocal cords are active, periodically cycling between opening and closing at rates of over 100 times per second. When the vocal cords are open, the tracheobronchial tree and the vocal tract become one continuous space, allowing sound to resonate between them. On the other hand, when the vocal cords are closed, the subglottic space (the tracheobronchial tree) and the supraglottic space (the vocal tract above the vocal cords) are separated, preventing sound from reverberating. Therefore, for voiced sounds, the transfer function is not time-invariant.
[0063] In voiced sounds, excitation of the supraglottal space is periodic, with a period corresponding to one cycle of the vocal folds opening and closing (corresponding to the "fundamental frequency" of the sound). Therefore, the excitation can be modeled as a train of uniform pulses, with spacing equal to the period of vocal fold vibration between successive pulses. (The spectral shaping caused by the vocal folds is effectively concentrated in the modulation of the vocal tract.) Excitation of the subglottal space is also caused by the vocal tract, and so can be modeled by the same train of uniform pulses. In the frequency domain, speech and acoustic signals are the product of the excitation signal by the supraglottal and subglottal transfer functions, respectively, so their spectra also consist of pulses of the same frequency as the excitation pulses, with amplitudes proportional to the respective transfer functions.
[0064] Thus, in one embodiment, processor 52 applies this model in estimating the spectral envelope of the speech and acoustic signals, and thus the transfer function of the vocal tract, H VT (e iω ), and the transfer function H of the tracheobronchial tree (including the lung wall)TB (e iω ) is estimated. The transfer function of the entire system is given by:
number
[0065] The processor 52 uses methods from the field of intra-speech recognition to generate the spectral envelope H VT (e iω ) and H TB (e iω ) can be derived. For example, for each signal X M (e iω ), X S (e iω ) are computed by linear predictive coding (LPC), respectively, and the spectral envelope is derived using equation (3) above. In effect, by considering only the spectral envelope, processor 52 obtains a time-invariant approximation.
[0066] The temporal variations of voiced sounds occur at a frequency that is a function of pitch, i.e., the frequency of vibration of the vocal cords. Thus, in some embodiments, processor 52 identifies the pitch of the spoken sound and calculates the time variation coefficient of the transfer function between the signal from microphone 24 and acoustic transducer 28, constraining the time variation to be periodic with a period corresponding to the pitch. To this end, equation (5) can be rewritten as:
number
number
[0067] The method described above requires estimating a relatively large number of coefficients, especially for low-pitched male voices. Reliably determining many of these coefficients requires multiple repetitions of a particular voiced phoneme, which can be difficult to obtain in routine medical monitoring. To alleviate this problem, the coefficients can be expressed as parametric functions that describe their time-varying behavior during the vocal cycle:
number
number
number
[0068] Alternatively, processor 52 can use more sophisticated forms of these parametric functions, which can more accurately represent the transfer function at the transition between vocal fold open and closed states. For example, B l (v), 0 ≦ l ≦ q and A k (v), 0 ≦ k ≦ p, may be a polynomial of fixed degree or a rational function (ratio of polynomials).
[0069] In another embodiment, the processor 52 applies an adaptive filtering approach in deriving the transfer function. M [n] is fed into a time-varying filter, which converts the acoustically transformed signal x S Predictors for [n]
number
number
number
[0070] Using this adaptive filtering approach, processor 52, for each sample of the patient's speech, derives a set of adaptive filter coefficients for that sample. Processor 52 can use this set of filter coefficients itself to characterize the transfer function. Alternatively, it may be desirable to reduce the amount of data that needs to be stored. For example, processor 52 can retain only every Tth set of filter coefficients, where T is a predetermined number (e.g., T = 100). As another alternative, processor 52 can retain a specific number of sets of filter coefficients per phoneme, for example, three: one at the beginning, one in the middle, and one at the end of the phoneme.
[0071] (Distance calculation) Returning now to FIG. 3, after calculating the transfer function between the signals from microphone 24 and acoustic transducer 28 (using any of the techniques described above or other techniques known in the art), processor 52 evaluates the deviation between the calculated transfer functions. "Distance" in this context is a numerical value calculated for the coefficients of the current transfer function and the baseline transfer function to quantify the difference between them. Any suitable type of distance measure can be used in step 70, and the distance need not be Euclidean, or even symmetric under the inversion of that argument. Processor 52 compares the distance to a predefined threshold in distance evaluation step 72.
[0072] As mentioned above, the baseline transfer function used as a reference in step 70 can be derived from previous measurements made on patient 22 or from measurements taken from a larger population. In some embodiments, processor 52 calculates a distance from the baseline that includes two or more reference functions. For example, processor 52 can calculate a vector of distances from a set of reference transfer functions and then select the minimum or average of the distances to evaluate in step 72. Alternatively, processor 52 can combine the reference transfer functions, for example, by averaging the coefficients and then calculating the distance from the current transfer function to the average function.
[0073] In one embodiment, these two approaches are combined: for example, using k-means clustering, the reference transfer functions are clustered based on similarity (meaning the distance between transfer functions in the same cluster is small). Processor 52 then synthesizes representative transfer functions for each cluster. Processor 52 calculates the distance between the current transfer function and representative transfer functions of different clusters, and then calculates the final distance based on these cluster distances.
[0074] The definition of the distance between the tested transfer function and the reference transfer function depends on the form of the transfer function. For example, f T = H T (e iω ) and f R = H R (e iω ) are the current and reference transfer functions, respectively, so that |ω|≦π as defined in equation (1) above. The distance d between these transfer functions can be written as:
number
[0075] In some embodiments, processor 52 uses a frequency domain transfer function H T (e iω ) and H R (e iω ) need not be explicitly calculated. Rather, as explained above, these functions can be described in terms of time-domain impulse responses or cepstral coefficients, and equation (12) can be expressed and evaluated, either exactly or approximately, in terms of operations on sequences of values, such as autocorrelations, cepstral coefficients, or impulse responses, that correspond to the transfer functions.
[0076] In some embodiments, processor 52 evaluates the distance by calculating each difference between pairs of coefficients of the current transfer function and the baseline transfer function, and then calculates the norm of all each difference. For example, in one embodiment, the distance is expressed as:
number
number
[0077] In the limit, p → ∞, so equation (12) becomes ∞ becomes the norm, which is simply an upper bound on the difference:
number
[0078] As another example, W(e iω ) = 1 and p = 2, the distance is reduced to the root mean square (RMS) of the difference between the current log spectrum and the baseline log spectrum.
[0079] Alternatively, the logarithm in equation (13) can be replaced by other monotonically non-decreasing functions, and p and W(e iω ) can also be used.
[0080] In other embodiments, a statistical maximum likelihood approach is used, such as the Itakura-Saito distortion, which is obtained by setting:
number
[0081] Alternatively or additionally, the distance function G(t,r,ω) can be selected based on empirical data, based on observing the actual transfer functions of a particular patient or many patients with different health conditions. For example, if the deterioration of health associated with a particular disease is log|H for ω in a particular frequency range Ω, T (e iω )||, and if the baseline transfer function corresponds to a patient's healthy and stable state, then the distance can be defined accordingly as:
number
[0082] As another example, if time-varying transfer function coefficients are used as in equations (7) and (8), where v = n / T, 0 ≦ n < T, then for each value of 0 ≦ v < 1, equation (7) defines a time-varying transfer function:
number
number
[0083] Finally, in embodiments where each transfer function includes multiple phoneme-specific transfer functions, processor 52 separately calculates the distance between each pair of corresponding phoneme-specific components of the current and baseline transfer functions using one of the techniques described above. The result is a set of phoneme-specific distances. Processor 52 applies a scoring procedure to these phoneme-specific distances to find a final distance value. For example, the scoring procedure could calculate a weighted average of the phoneme-specific distances, where phonemes that are more sensitive to changes in health (based on empirical data) are weighted higher.
[0084] In another embodiment, the scoring procedure uses rank statistics instead of averaging. The phoneme-specific distances are weighted according to their sensitivity to changes in health status and sorted into a sequence in ascending order. Processor 52 selects the value (e.g., the median) that appears at a particular location in this sequence as the distance value.
[0085] Regardless of which of the above distance measures is used, if processor 52 finds in step 72 that the distance between the current transfer function and the baseline transfer function is less than the maximum expected deviation, processor 52 records the measurement but typically does not initiate any further action. (Server 32 may notify the patient or caregiver that there is no change in the patient's condition, or even that the patient's condition has improved.) However, if the distance exceeds the maximum expected deviation, server 32 initiates action in action initiation step 76. The action may include, for example, issuing an alert in the form of a message to the patient's caregiver, such as the patient's physician. The alert typically indicates increased fluid accumulation in the patient's thorax and prompts the caregiver to take treatment, such as administering or modifying medication, to reduce fluid accumulation.
[0086] Alternatively, the server 32 may not proactively push alerts but may simply present (e.g., on a display or in response to a query) an indicator of the subject's condition, such as the level of pulmonary edema. The indicator may include, for example, a number based on the distance between the transfer functions representing an estimated level of pulmonary edema, assuming that a correlation between the distance between the transfer functions and pulmonary edema has been learned from previous observations of this or other subjects. Physicians may refer to this indicator along with other medical information in making diagnostic and treatment decisions.
[0087] In some embodiments, drug administration and dosage changes are performed automatically by controlling the drug delivery device without the need for a human caregiver in the loop. In such cases, step 76 may include changing the dosage level, with or without issuing an alert (or the alert may indicate that the dosage level has been changed).
[0088] In some cases, for example in a hospital or other clinic setting, the distance assessment in step 72 may indicate an improvement in the subject's condition rather than a deterioration. In this case, the action initiated in step 76 may indicate that the subject be moved out of the intensive care unit or released from the hospital.
[0089] The above-described embodiments are cited by way of example, and it will be understood that the present invention is not limited to what has been particularly shown and described above. Rather, the scope of the present invention includes both combinations and subcombinations of the various features described above, as well as variations and modifications thereof not disclosed in the prior art that will occur to those skilled in the art upon reading the foregoing description.
Claims
1. 1. A method for medical diagnosis executed by a processor in a computer, the processor comprising: recording an audio signal resulting from sounds spoken by the patient; recording an acoustic signal emitted by an acoustic transducer in contact with the patient's thorax simultaneously with the audio signal; calculating a transfer function between the recorded speech signal and the recorded acoustic signal, or between the recorded acoustic signal and the recorded speech signal; and evaluating the calculated transfer function to assess a medical condition of the patient; configured to perform A method characterized by:
2. The step of evaluating the calculated transfer function comprises: assessing the deviation between the calculated transfer function and a baseline transfer function; and detecting a change in the patient's condition in response to the assessed deviation; 2. The method of claim 1, comprising:
3. 3. The method of claim 2, wherein detecting a change in the patient's medical condition comprises detecting an accumulation of fluid in the patient's thorax.
4. 4. The method of claim 3, further comprising administering treatment to the patient under control of the processor in response to detecting a change in the patient's medical condition to reduce the amount of fluid accumulated in the thorax.
5. 10. The method of claim 1, wherein evaluating the calculated transfer function comprises evaluating the patient for interstitial lung disease.
6. 10. The method of claim 1, further comprising administering a treatment to the patient to treat the assessed medical condition.
7. 2. The method of claim 1, wherein recording the acoustic signal comprises removing heart sounds from the acoustic signal output by the acoustic transducer before calculating the transfer function.
8. 8. The method of claim 7, wherein removing the heart sounds comprises detecting intervals in the acoustic signal that include extraneous sounds, including the heart sounds, and removing the intervals from the acoustic signal used to calculate the transfer function.
9. 8. The method of claim 7, wherein removing the heart sounds comprises filtering the heart sounds from the recorded acoustic signal before calculating the transfer function.
10. 10. The method of claim 9, wherein recording the acoustic signals comprises receiving at least first and second acoustic signals from at least first and second acoustic transducers, respectively, in contact with the rib cage, and filtering out the heart sounds comprises applying a delay to the arrival of the heart sounds in the second acoustic signal compared to the first acoustic signal in a combination of the first and second acoustic signals while filtering out the heart sounds.
11. 11. The method of claim 1, wherein the step of calculating the transfer function comprises calculating the spectral components of the recorded speech signal and the recorded acoustic signal at a set of frequencies, and calculating a set of coefficients representing a relationship between the respective spectral components.
12. 12. The method of claim 11, wherein the coefficients are in a cepstral representation.
13. 11. The method according to claim 1, wherein the step of calculating the transfer function comprises calculating a set of coefficients representing the relationship between the recorded speech signal and the recorded acoustic signal in an infinite impulse response filter.
14. 11. The method according to claim 1, wherein the step of calculating the transfer function comprises the step of calculating a set of coefficients representing a relationship between the recorded speech signal and the recorded acoustic signal in a predictor in the time domain.
15. 15. The method of claim 14, wherein calculating the set of coefficients comprises applying a prediction error of the relationship in calculating adaptive filter coefficients associated with the recorded speech signal and the recorded acoustic signal.
16. 11. A method according to any preceding claim, wherein the step of calculating the transfer functions comprises dividing the spoken sound into a plurality of different types of speech units and calculating separate respective transfer functions for the different types of speech units.
17. 11. The method of claim 1, wherein the step of calculating the transfer function comprises calculating a set of time-varying coefficients representing the temporal relationship between the recorded speech signal and the recorded acoustic signal.
18. 18. The method of claim 17, wherein calculating the set of time-varying coefficients comprises identifying a pitch of the spoken speech signal; and constraining the time-varying coefficients to be periodic with the same period as the identified pitch.
19. 5. The method of claim 2, wherein the step of calculating the transfer function comprises calculating a set of coefficients representative of a relationship between the recorded speech signal and the recorded acoustic signal, and wherein the step of evaluating the deviation comprises calculating a distance function between the coefficients of the calculated transfer function and the baseline transfer function.
20. 20. The method of claim 19, wherein calculating the distance function comprises: calculating each difference between pairs of coefficients, each pair consisting of a first coefficient in the calculated transfer function and a second corresponding coefficient in the baseline transfer function; and calculating a norm over all of the respective differences.
21. 20. The method of claim 19, wherein calculating the distance function comprises observing differences between the transfer functions calculated for different health states, and selecting the distance function in response to the observed differences.
22. 1. An apparatus for medical diagnosis comprising: a memory configured to store a recorded audio signal from sounds spoken by the patient and a recorded acoustic signal output simultaneously with the audio signal by an acoustic transducer in contact with the patient's thorax; a processor configured to calculate a transfer function between the recorded voice signal and the recorded acoustic signal, or between the recorded acoustic signal and the recorded voice signal, and to evaluate the calculated transfer function to assess a medical condition of the patient; An apparatus comprising:
23. 23. The apparatus of claim 22, wherein the processor is configured to assess a deviation between the calculated transfer function and a baseline transfer function and to detect a change in a medical condition of the patient in response to the assessed deviation.
24. 24. The apparatus of claim 23, wherein the change detected by the processor comprises an accumulation of fluid in the patient's thorax.
25. 25. The apparatus of claim 24, wherein in response to detecting the change, the patient is treated to reduce the amount of fluid accumulated in the thorax.
26. 23. The apparatus of claim 22, wherein the processor is configured to assess interstitial lung disease in the patient in response to the calculated transfer function.
27. 23. The apparatus of claim 22, wherein the patient is treated to treat the assessed condition.
28. 23. The apparatus of claim 22, wherein the processor is configured to remove heart sounds from the acoustic signal output by the acoustic transducer before calculating the transfer function.
29. 29. The apparatus of claim 28, wherein the processor is configured to detect intervals of occurrence of extraneous sounds, including heart sounds, in the acoustic signal output by the acoustic transducer and to remove from the acoustic signal the intervals used in calculating the transfer function.
30. 30. The apparatus of claim 28, wherein the processor is configured to filter out heart sounds from the recorded acoustic signal before calculating the transfer function.
31. 31. The apparatus of claim 30, wherein the memory is configured to receive and store at least first and second acoustic signals from at least first and second acoustic transducers, respectively, in contact with the rib cage, and the processor is configured to apply a delay to the arrival of the heart sound in the second acoustic signal relative to the first acoustic signal in a combination of the first and second acoustic signals while filtering out the heart sound.
32. 32. The apparatus of claim 22, wherein the processor is configured to calculate the spectral components of each of the recorded speech signal and the recorded acoustic signal at a set of frequencies, and to calculate a set of coefficients representing a relationship between the respective spectral components.
33. 33. The apparatus of claim 32, wherein the coefficients are in a cepstral representation.
34. 32. The apparatus of claim 22, wherein the processor is configured to calculate a set of transfer function coefficients representing the relationship between the recorded speech signal and the recorded acoustic signal in an infinite impulse response filter.
35. 32. The apparatus according to any one of claims 22 to 31, characterized in that the processor is configured to calculate coefficients of a set of transfer functions representing a relationship between the recorded speech signal and the recorded acoustic signal with a predictor in the time domain.
36. 36. The apparatus of claim 35, wherein the processor is configured to apply a prediction error of the relationship when calculating adaptive filter coefficients associated with the recorded speech signal and the recorded acoustic signal.
37. 32. The apparatus of claim 22, wherein the processor is configured to divide the spoken sound into a plurality of different types of speech units and to calculate separate respective transfer functions for the different types of speech units.
38. 32. Apparatus according to any one of claims 22 to 31, characterized in that the processor is configured to calculate a set of time-varying coefficients of a transfer function representing the temporal relationship between the recorded speech signal and the recorded acoustic signal.
39. 39. The apparatus of claim 38, wherein the processor is configured to identify a pitch of the spoken audio signal and to constrain the time-varying coefficients to be periodic with the same period as the identified pitch.
40. 26. The apparatus of claim 23, wherein the processor is configured to calculate coefficients of a set of transfer functions representing a relationship between the recorded speech signal and the recorded acoustic signal, and to assess the deviation by calculating a distance function between the coefficients of the calculated transfer functions and the baseline transfer function.
41. 41. The apparatus of claim 40, wherein the processor is configured to calculate respective differences between pairs of coefficients, each pair consisting of a first coefficient of the calculated transfer function and a second corresponding coefficient of the baseline transfer function, and to calculate a distance function by calculating a norm over all of the respective differences.
42. 41. The apparatus of claim 40, wherein the processor is configured to calculate the distance function in response to observed differences between the transfer functions calculated for different health states.
43. 1. A non-transitory computer-readable medium having stored thereon program instructions that, when read by a computer, cause the computer to: receive an audio signal representing a sound spoken by a patient and an acoustic signal output by an acoustic transducer in contact with the patient's thorax simultaneously with the audio signal; calculate a transfer function between the recorded audio signal and the recorded acoustic signal, or between the recorded acoustic signal and the recorded audio signal; and evaluate the calculated transfer function to assess a medical condition of the patient.
Citation Information
Patent Citations
Diagnostic system and portable telephone device
CN1520271A
Phonopneumograph system
JP2001505085A
An application server for reducing ambient noise in auscultation signals while listening to a patient with an electronic stethoscope, and for recording comments.
JP2011527211A
Diagnostic techniques based on speech-sample alignment
US11011188B2
Telemedicine system for IMD patients using audio / video data
US20130218582A1