Virtual spirometry and other pulmonary function tests (PTFS) from speech
A two-stage neural network system effectively estimates lung volume and capacity from speech by isolating relevant respiratory signals, enhancing the accuracy of spirometry parameter estimation.
Patent Information
- Application Number
- PCT/EP2025/064428
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-30
- Filing Date
- 2025-05-26
- Publication Date
- 2025-12-04
AI Technical Summary
Existing methods struggle to accurately estimate lung volume and capacity from speech due to weak correlations between speech characteristics and lung functionality, with many factors unrelated to lung volume and capacity affecting speech characteristics.
A two-stage neural network approach is employed, where the first stage processes speech to produce a respiratory signal largely unaffected by uncorrelated audio features, and the second stage estimates spirometry parameters from this respiratory signal, using deep learning models like Large Language Models (LLMs) and deep recurrent neural networks.
This method provides more accurate estimation of spirometry parameters by minimizing the influence of speech characteristics unrelated to lung volume and capacity, resulting in reliable lung function assessments.
Smart Images

Figure 00000020_0000 
Figure 00000020_0001 
Figure 00000020_0002
Abstract
Description
VIRTUAL SPIROMETRY AND OTHER PULMONARY FUNCTION TESTS (PTFs)FROM SPEECHFIELD OF THE INVENTIONThis invention relates to the field of medical instruments, and in particular to a method and system for estimating Spirometry parameters based on a subject's speech.BACKGROUND OF THE INVENTIONSpirometry is the "gold standard" for assessing lung functionality. A variety of parameters related to lung functionality may be determined, including Forced Vital Capacity (FVC), the vital capacity (volume) from a maximally forced expiratory effort; Forced Expiratory Volume at 1 second (FEV1), the volume that has been exhaled at the end of the first second of forced expiratory effort; Peak Expiratory Flow (PEF), the highest expirator forced flow rate; and others. Using these parameters, a Flow- Volume loop may be determined.As indicated by the terms "forced expiratory effort" and "maximal inhalation effort", a spirometry test involves a strict protocol, typically under the supervision of a medical professional, wherein the subject is instructed to inhale as completely as possible, then exhale maximally into a device that measures the volume of the exhalation over time, then inhale deeply through the device to measure the volume of the inhalation over time. Alternative sequences are permitted, provided that the sequence provides a measure of the volume of air exhaled over time from the point of maximal inhalation, and a measure of the value of air inhaled over time from the point of maximal exhalation.FIG. 1 A illustrates an example spirogram and illustrates the volume of air in the lungs through an example maximum inhalation-exhalation breathing cycle.FIGs. 1 B-1 E illustrate a set of example Flow-Volume (FV) loops for subjects with different lung conditions. An FV loop is a plot of flow of air into and out of the lungs (Inspiratory - Expiratory Flow) v. the volume of air exhaled over a forced inhalation - exhalation breathing cycle (Inspiratory - Expiratory Volume), wherein apositive flow indicates air flow out of the lungs, and a negative flow indicates air flow into the lungs.Each FV loop provides spirometry parameters, including Peak Expiratory Flow (PEF) and Forced Vital Capacity (FVC). Between these points, the Forced Expiratory Volume (FEV) is plotted. An example FEV at 1 second (FEV1) is indicated along the FEV plot. Other spirometry parameters may be illustrated in the FV loops.FIG. 1 B illustrates a typical Flow-Volume loop of a subject exhibiting normal lung function. As illustrated, the flow decreases at a relatively constant rate from the peak flow PEF until the end of the forced expiratory effort, at which the vital capacity of air FVC has been exhaled from the lungs. The inhalation portion of the Flow- Volume loop exhibits a semi-circular flow as the subject performs a maximal inhalation effort.FIG. 1 C illustrates a typical Flow-Volume loop of a subject with asthma. As can be seen, the PEF is substantially lower, and the flow sharply decreases from the start of the forced expiratory effort.FIG. 1 D illustrates a typical Flow-Volume loop of a subject with upper airway obstruction. The PEF is also substantially lower, and in this case, the flow is relatively constant for the first part of the forced expiratory effort.FIG. 1 E illustrates a typical Flow-Volume loop of a subject with Chronic Obstructive Pulmonary Disease (COPD).In recent years, techniques for assessing a subject's respiratory condition, or other physiological parameters, based on an analysis of the subject's speech ("speech-to-respiration") have advanced significantly. Such techniques have become of increasing importance in the field of telehealth, particularly in view of Covid-19 and other contagious illnesses, wherein the subject partakes in a "remote visit" with a medical practitioner. While speaking with the subject, the practitioner may enable a speech-to-respiration application and receive a real-time assessment of the subject's respiration process.U.S. Patent Application Publication 2022 / 0257175, "SPEECH-BASED PULMONARY ASSESSMENT", serial number 17 / 592,777, filed 4 February 2022 for Vatanparvar et al. (hereinafter the '777 application), discloses "
[0019] ... performing an audio analysis of the individual's speech to identify and extract audio featuresfrom the individual's speech. Based on machine-determined correlations between the audio features and lung function parameters, the individual's pulmonary condition is determined. The lung function parameters can include FEV1 , FVC, and FEV1 / FVC, from which pulmonary conditions corresponding to airway obstructions and airway restrictions can be determined."The 777 application processes the input speech to determine temporal changes in time and frequency and uses a combination of a Convolutional Neural Network (CNN) model and a Long Short Term Memory (LSTM) model to estimate the lung function parameters, including FEV1 , FVC, and FEV1 / FVC. Mel-frequency cepstral coefficients (MFCCs), which represent the signal power in each band of the frequency domain of the audio signal, provide the feature vectors that are input to the CNN-LSTM model.However, although a correlation between speech and respiration may be strong, the correlation between speech and particular parameters related to lung volume and capacity is relatively weak. That is, persons with similar MFCCs may or may not have similar lung volume and capacity, and vice versa. Many factors unrelated to lung volume and capacity, such as the structure of the person's vocal chords, nasal cavity, etc. affect the characteristics (e.g. MFCCs) of a person's speech. Accordingly, it may be difficult to effectively train a neural network to distinguish between the features that are correlated or uncorrelated with lung volume and capacity.SUMMARY OF THE INVENTIONIt would be advantageous to provide a speech-to-spirometry system that maximizes the effect of speech characteristics that are likely to be related to lung volume and capacity, and minimizes the effect of speech characteristics that are relatively unrelated to lung volume and capacity.Embodiments of this invention comprise two neural network stages. A first neural network stage is configured to receive the speech signal, and to produce therefrom a respiratory signal corresponding to the speech signal. A second neural network stage is configured to receive the respiratory signal, and to produce therefrom values of a plurality of spirometer parameters (including PEF, FVC, FEV1 , FEV1 / FVC, etc.) based on the respiratory signal.By processing the speech input in the first neural network stage to produce respiratory parameters, auditory features that are not strongly correlated to respiration have minimal effect on the determined respiratory signal. Accordingly, the input to the second neural network stage is relatively free of these uncorrelated audio features.By processing the speech input in the first neural network stage to produce respiratory parameters, a variety of features beyond temporal changes in time and frequency (MFCCs) may be used to train the first neural network stage to optimize the performance of the network for determining respiration parameters. In embodiments of this invention, for example, the first neural network stage is based on Large Language Models(LLMs) that possess a deep contextual understanding of language and speech physiology apart from the spectral features of speech (MFCCs, log-Mel, etc).In embodiments of this invention, the speech-to-spirometry system comprises: a receiver circuit; a processor circuit; a first neural network circuit; a second neural network circuit; and an output circuit; wherein the receiver circuit is configured to receive a speech signal; wherein the processor is configured to process the speech signal to provide speech parameters to the first neural network circuit; wherein the first neural network circuit is configured to receive the speech parameters and to produce therefrom values of a respiratory signal corresponding to the speech signal; wherein the processor is configured to process the respiratory signal to provide respiratory parameters to the second neural network; wherein the second neural network circuit is configured to receive the respiratory parameters and to produce therefrom values of a plurality of spirometer parameters; wherein the spirometer parameters comprise Forced Vital Capacity (FVC); Forced Expiratory Volume (FEV); and Peak Expiratory Flow (PEF); and wherein the output circuit is configured to provide the values of the plurality of spirometer parameters to a user.In some embodiments, the output circuit is configured to provide, based at least in part on the values of the spirometer parameters, at least one of a Volume- Time loop and a Flow-Volume loop.In some embodiments, the first neural network circuit is a deep-learning neural network that is trained using speech inputs and respiration device outputs.In some embodiments, the second neural network circuit is a deep-learning neural network that is trained using respiration inputs and spirometer outputs.In some embodiments, the speech-to-respiration neural network is trained using speech inputs and respiratory device outputs; and the respiration-to- spirometer neural network is trained using outputs of the speech-to-respiration neural network and spirometer device outputs.In some embodiments, at least one of the first or second neural network circuit comprises a Bayesian neural network.In some embodiments the values of the FEV parameter comprise values of Forced Expiration Volume at 1 second (FEV1); Forced Expiration Volume at 25% (FEV25); and Forced Expiration Volume at 75% (FEV75).In some embodiments the respiratory parameters comprise respiration rate, inhalation duration, and exhalation duration.In some embodiments, the input to the respiration-to-spirometry stage comprises physical features of the subject (age, height, weight, gender, etc.).BRIEF DESCRIPTION OF THE DRAWINGSThe invention is explained in further detail, and by way of example, in embodiments that include a speech-to-respiration system with reference to the accompanying drawings wherein:FIGs. 1A-1 E illustrate an example spirogram and example Flow-Volume curves.FIG. 2 illustrates an example embodiment of a Speech-to-Spirometry system in accordance with principles of this invention.FIGs. 3A-3C illustrate an example embodiment of a speech-to-respiration neural network stage.FIGs. 4A-4B illustrate an example embodiment of a respiration-to-spirometry neural network stage.Throughout the drawings, the same reference numerals indicate similar or corresponding features or functions. The drawings are included for illustrative purposes and are not intended to limit the scope of the invention.DETAILED DESCRIPTIONIn the following description, for purposes of explanation rather than limitation, specific details are set forth such as the particular architecture, interfaces, techniques, etc., in order to provide a thorough understanding of the concepts of the invention. However, it will be apparent to those skilled in the art that the present invention may be practiced in other embodiments, which depart from these specific details. In like manner, the text of this description is directed to the example embodiments as illustrated in the Figures, and is not intended to limit the claimed invention beyond the limits expressly included in the claims. For purposes of simplicity and clarity, detailed descriptions of well-known devices, circuits, and methods are omitted so as not to obscure the description of the present invention with unnecessary detail.As used herein, the term "circuit" refers to a configuration of electrical elements that perform a given function or set of functions. The circuit may include discrete elements or integrated elements, and may include, for example, a processor configured to execute a program to perform at least a part of the function(s).Speech production involves a systematic outflow of air during exhalation characterized by linguistic content and prosodic factors of the utterance. Accordingly, there is a strong correlation between speech and respiration, and techniques have been developed to model this relationship and sense respiratory dynamics directly from the speech. These techniques generally apply "deep learning" to sense a breathing signal and breathing parameters from the speech. Estimating the breathing pattern from the speech provides information about the subject's respiratory parameters, thus enabling an understanding of the subject's respiratory health based on the subject's speech.However, as noted above, estimating the lung volume and capacity of a subject based directly upon the subject's speech is difficult because the direct correlation between speech and lung volume and capacity is relatively weak, primarily because a variety of speech characteristics are substantially unrelated to the subject's lung volume and capacity (spirometry parameters).Embodiments of this invention include a two-stage approach to estimate spirometry parameters from a subject's speech. In the first stage, the speech is analyzed to estimate the subject's repiratory activity. In the second stage, the subject's respiratory activity is analyzed to estimate the subject's lung volume and capacity. The subject's lung volume and capacity are presented to the medical practitioner (and / or the subject) using conventional spirometry characteristics, including Flow-Volume (FV) loops and other graphic illustrations.As illustrated in FIG. 2, in operation after training, an audio receiver 220 receives audio in the form of speech from any of a variety of sources 210, such as via a telephone 210i over a telephony network, a local microphone 2102, a smartphone 2103 via an internet connection, etc. The receiver 220 may include circuitry that converts radio-frequency signals into audio-frequency signals. As used herein, the terms audio signal and speech signal are used synonymously, and comprise an electronic signal that corresponds to the input speech.The speech signal 225 from the receiver 220 is provided to the speech-to- respiration stage 230. The speech-to-respiration stage 230 processes the speech signal 225 to create a speech-input feature set that is provided to a deep learning neural network that estimates the subject's respiration activity, typically in the form of an estimated respiration signal and associated parameters (breath rate, inhalation duration, exhalation duration, etc.) 235, as detailed with regard to FIG. 3.As noted above, during the training stage of the speech-to-respiration neural network, the neural network 'learns' which characteristics of the input speech affect the estimated respiration signal, and accordingly, which characteristics of the input speech have a minimal effect on the estimated respiration signal. In this manner, the estimated respiration signal is relative free of speech characteristics that are unrelated to respiration activity.This estimated respiration signal and parameters 235 are provided to the respiration-to-spirometry stage 240. The respiration-to-spirometry stage 240 processes the respiration signal and parameters 235 to create a respiration-input feature set to a deep learning neural network that estimates the subject's spirometry (lung volume and capacity) parameters 245, as detailed with regard to FIG. 4.Because the estimated respiration signal and parameters 235 are relatively free of speech characteristics that are unrelated to respiration activity, the corresponding respiration-input feature set is also relatively free of thesecharacteristics. Accordingly, the training of the respiration-to-spirometry neural network will be more efficient and will likely result in a more accurate estimate of the spirometry parameters than a direct speech-to-spirometry neural network.The estimated spirometry parameters 245 are analyzed 250 to provide an output 265 that is similar to the ouput provided by a conventional spirometry test, for presentation on an output device 260.FIGs. 3A-3C illustrate example configurations of a speech-to-respiration circuit as may be embodied as the speech-to-respiration stage 230 of FIG. 2, and an example comparison of a predicted respiration signal and the actual respiration signal.FIG. 3A illustrates an example configuration for training a deep neural network 370 that may be used to model the correlation between input speech 325 and a respiration signal 335 while the speech is being input. Typically, dozens of subjects, or more, are used to train the neural network 370.Respiration measurement belts 315 are strapped to a subject 310, and a transducer 330 converts the expansions and contractions into a respiration signal 335. The transducer 330 may also perform signal processing to provide a suitable respiration signal 335. In some embodiments, the transducer 330 may include the belts 315 that provide an electrical signal directly in response to being stretched.One of skill in the art will recognize that any of a variety of devices may be used to monitor the subject's respiration 335 while the subject is instructed to speak into a microphone 320 to provide a speech signal 325. To emulate breathing associated with spirometry tests, the subject may be instructed to inhale deeply than speak continuously for as long as possible before taking the next breath, then repeat a few times. For example, the subject may be instructed to count aloud as long as possible. A variety of training cycles may be conducted, wherein for example, the subject is instructed to repeat the counting aloud at different speeds, and at different loudness levels. The variety may include speaking other words, such as reciting the alphabet as long as possible before the next breath. Other variations may be performed that are not, per se, emulations of spirometry tests. For example, the subject may be instructed to recite a fixed text, such as "The Rainbow Passage" from the book "Voice and Articulation" by G. Fairbanks, at "normal" cadence andloudness. In like manner, the training may include training using the subject's speech during normal conversation with the person conducting the training.A trainer circuit 340 receives the speech signal 325 and corresponding respiration signal 335. A speech processor 350 pre-processes the speech signal 325 to provide a set of inputs ("features") 342 to the network 370. Such pre-processing may include spectral analysis, Mel-frequency scaling, normalization, etc., such as disclosed in "Deep learning architectures for estimating breathing signal and respiratory parameters from speech recordings", by Venkata Srikanth Nallanthighal, Zohreh Mostaani, Aki Harma, Helmer Strik, and Mathew Magimai-Doss; Neural Networks 141 (2021) pgs. 311-324, which is incorporated by reference herein. In some embodiments both the raw speech signal 325 and the pre-processed signals are input to the network 370. As used further herein, the term "speech signal" refers to raw speech signal, the pre-processed speech signal, or a combination of both, depending upon the particular embodiment of the input processor 350 and the deep neural network 370.If the training sessions included different variations of speech (spirometry emulation, text recitation, conversation, etc.), as described above, an identification of the particular variation of input speech may also be included in the feature set 342 that is provided to the network 370.The deep neural network 370 is designed to provide an estimate of a respiration signal corresponding to the input speech 325. The network 370 may comprise a deep recurrent neural network or other sequential regression algorithms, such as disclosed in U.S. patent application 17 / 071 ,312 (USPA 2021 / 0146082) by Aki Sakari Harma, Francesco Vicario, and Venkata Srikanth Nallanthighal, and incorporated by reference herein. Although illustrated as a single network 370, the network 370 may comprise a plurality of neural networks and an overall architecture that "fuses" the outputs of these neural networks to provide a potentially more accurate respiration signal.In embodiments, the input 342 to the network 370 may be discrete segments of the input speech 325, such as segments of a few seconds each. In some embodiments, to minimize transition anomalies, the segments may overlap, such that the beginning of a segment overlaps with the ending of a prior segment, and the ending of the segment overlaps with the beginning of a next segment, as describedin EP23155738, "Speech Processing of Audio Signal" by Aki Harma, filed 9 February 2023, which is incorporated by reference herein.Each of these segments may be pre-processed as detailed above. In a training mode, the network 370 receives the speech input 342 and provides an estimate of the respiration signal 372 to the trainer 340. The trainer 340 includes a feedback circuit 360 that compares the predicted respiration signal 372 with the known respiration signal 335, and provides feedback 344 to the network 370 based on the correspondence, or lack of correspondence, between the predicted 372 and known 335 respiration signals.Typically, the training of the network 370 continues for each subject until the feedback circuit 360 indicates that a sufficient level of correspondence between the predicted 372 and known 335 respiration signals has been achieved.Subsequent tests may be performed to confirm that the training has been successful. Variations of speech signals 325 are provided in a "test mode", wherein the feedback circuit 360 provides a measure of the correspondence between the predicted 372 and known 335 respiration signals, but the trainer circuit 340 does not provide the feedback signal 344 to the network 370. That is, the network 370 remains unchanged throughout the test period. If the testing is successful, the trained network is subsequently used for performing speech-to-respiration in an operational environment, wherein the subject does not wear the respiration belts 315.FIG. 3B illustrates an example configuration of the speech-to-respiration system in normal operation, after the training and testing of the network 370. In this configuration, the network 370 remains unchanged from its trained and tested state.The subject 310 speaks and produces a (new) speech signal 325', preferably following the same instructions that were given to subjects during the training period. This signal 325' is processed via the input processor 350, which performs the same processing as performed during the training process, including, if used during training, an identification of the particular variation of speech (spirometry emulation, text recitation, conversation, etc.) that the subject was instructed to use to produce the speech signal 325'. Similarly, the same segmentation of the speech signal 325 applied during training is applied to the current speech signal 325'. The resultant feature set 352 is provided to the network 370, which applies the input speech 352 to the trained model of the network 370. In response to the input speech, the network370 produces an estimated respiration signal 335' corresponding to the input speech signal 325'. This estimated respiration signal 335' is herein termed a "Virtual Respiration Belt" (VRB) signal, in that it is assumed to be the respiration belt signal 325 of FIG. 3A that would have been produced if the subject had been wearing the respiration belts 315.FIG. 3C illustrates an example comparison of a VRB signal 335' and an actual ("ground truth") respiration signal 335 that was measured while the speech signal 325' was input to the speech-to-respiration system. As can be seen, the correlation between the estimated respiration and the actual respiration is quite significant.It is significant to note that in FIG. 3A, the respiration signal 335 from the respiration belts 315 corresponds directly to the subject's breathing, and is unaffected by characteristics of the speech input 325 that do not affect the subject's respiration. As noted above, the VRB signal 335' corresponds to the respiration signal 335 that would have been produced had the subject been wearing the respiration belts 315. Accordingly the VRB signal 335' is substantially unaffected by characteristics of the speech input 325' that do not affect the subject's respiration.The VRB signal 335' may be subsequently processed to provide respiration related parameters, such as respiration rate, tidal volume, inhalation and exhalation moments and durations, etc. The VRB signal 335' and the respiration related parameters form the Respiration Signal and Parameters 235 of FIG. 2 that are provided to the Respiration to Spirometry stage 240.FIGs. 4A-4B illustrate an example configuration of a respiration-to-spirometry circuit as may be embodied as the speech-to-respiration stage 240 of FIG. 2.FIG. 4A illustrates the stage 240 of FIG. 2 during a training period. As contrast to the speech-to-respiration stage 230 of FIG. 2, the capture of the 'ground truth' spirogram 415 does not occur simultaneously with the captured speech, because the protocol for capturing a standard spirogram 415 does not permit speaking during the inhalation-exhalation cycles. This non-coincident capture is not significant, because a person's spirogram and corresponding spirometry parameters remain fairly constant over time.Although FIG. 4A illustrates a spirogram 415 as the input to the trainer, one of skill in the art will recognize that the aforementioned spirometry parameters may provide this input to replace or supplement the spirogram 415. The choice of theground truth input to the trainer 440 is dependent upon the architecture of the neural network 470 and the feedback circuit 460. In a straightforward embodiment, the form of the ground truth input corresponds to the form of the neural network output 472. Alternatively, the feedback circuit 460 may be configured to pre-process the ground truth input 415 to produce a data set that corresponds to the form of the neural network output 472, so as to enable the feedback of the differences between this data set and the output of the neural network. For ease of reference, the term "spirometry input" is defined as the ground truth input, regardless of the form of this input.The subject that provides the spirometry input may provide a plurality of respiration signals 235 and each signal is input to the trainer 440 while the spirometry input remains constant. Each signal 235 is pre-processed by a respiration processor 450 to create a feature set 442 that is input to the respiration-to- spirometry deep learning neural network 470. As noted above, the signal 235 may comprise the VRB signal 335' or derived respiration parameters from this signal 335', or, preferably a combination of both.If the signal 235 corresponds to a respiration signal of speech input that emulates a spirometry test input, as detailed above with respect to FIG. 3, the speech processor 450 may 'synchronize' the signal 235 to correspond to the spirometry input 415 as it prepares the feature set 442 that is provided to the neural network 470. Preferably, a neural network architecture suitable for sequence-to- sequence mapping is used, and may comprise, for example, a Recurrent Neural Network (RNN), a Convolutional Neural Network (CNN), or a combination of both.The processing of each signal 235 constitutes a training epoch and typically comprises a 20 second time window, including about 3-5 respiratory waves, and the sampling rate is nominally about 50Hz. As each signal 235 is processed, a predicted spirometry output 472 is producted and compared to the ground truth spirometry input 415. The difference 444 between the predicted spirometry output 472 and the ground truth spirometry input 415 is provided to the neural network, which uses this difference 444 to improve its performance, typically using back-propagation.The training continues by collecting and processing speech from multiple subjects 410. The speech of each subject is processed by the speech-to-respiration stage 230 of FIG. 2 to produce the respiration signals 235. Each subject 410 also undergoes a spirometry test to produce the ground truth spirometry input 415.Because a person's lung volume and capacity is typically a function of physical characteristics of the person, such as age, gender, height, weight, etc., some or all of these parameters may also be provided as part of the feature set 442 that is provided to the neural network 470 for each subject 410.As with the speech-to-respiration stage 230, the respiration-to-spirometry stage 240 is trained and tested until an acceptable predicted spirometry output is achieved.FIG. 4B illustrates an example respiration-to-spirometry stage 240 of FIG. 2 as used in an operational (non-training) mode. In this mode, a subject 310 of FIG. 3B provides one or more speech input samples 325, and the speech-to-respiration stage 230 provides estimated / predicted respiration signals and / or parameters 235 for each speech input sample, as described above with regard to FIG. 3B.The estimated respiration signal and / or parameters 235 are pre-processed by the respiration processor 450, detailed above, to produce a feature set 452 for the now-trained respiration-to-spirometry neural network 470. The network 470 processes this feature set 452 and produces an estimated spirometry output 415'.The estimated spirometry output 415' may comprise a spirogram or spirometry parameters, or a combination of both. The spirometry parameters may include Forced Expiratory Volume at 1 second (FEV1 ), Forced Vital Capacity (FVC), Peak Expiratory Flow (PEF), Vital Capacity (VC), Residual Lung Volume (RV), Maximum voluntary Minute Ventilation (MMV), Total Lung Capacity (TLC), and / or others. Some or all of these parameters may be determined by providing the spirometry output 415' to the spirometry analyzer 250 of FIG. 2.With reference to FIG. 2, if multiple speech samples 225 are captured for a subject, each corresponding estimated respiration signal 235 is provided to the respiration-to-spirometry stage 240 and a spirometry output 245 is produced. The respiration-to-spirometry stage 240 or the spirometry analyzer 250 may provide a consolidated set of respiration parameters based on, for example, a weighted or unweighted average.As noted above, the spirometry analyzer 250 may also provide graphic representations of lung volume and capacity, including a Flow-Volume (FV) loop 265 and others.While the invention has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive; the invention is not limited to the disclosed embodiments.For example, although the above descriptions address the training and use of a single deep-learning neural network model in each of the speech-to-respiration stage and the respiration-to-spirometry stage, one of skill in the art will recognize that each stage may include a plurality of models, each model having been trained using a different class of speakers, such as models based on gender, age, or other parameters found to influence the speech-to-respiration process and / or the respiration-to-spirometry process. These parameters would be input to the corresponding neural network stage at the start of the each process to enable the process to select and apply the appropriate model.Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. Also in the claims, the expression "at least one of A, B, and C" means "at least one of A, B, and / or C". A single processor or other unit may fulfill the functions of several items recited in the claims; in like manner, multiple processors may be used in lieu of a single processor. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage. A computer program may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems. Any reference signs in the claims should not be construed as limiting the scope.
Claims
We claim:
1. A speech-processing system comprising: a receiver circuit; a processor circuit; a first neural network circuit; a second neural network circuit; and an output circuit; wherein the receiver circuit is configured to receive a speech signal of a subject; wherein the processor is configured to process the speech signal to provide speech parameters to the first neural network circuit; wherein the first neural network circuit is configured to receive the speech parameters and to produce therefrom values of a respiratory signal corresponding to the speech signal; wherein the processor is configured to process the respiratory signal to provide respiratory parameters to the second neural network; wherein the second neural network circuit is configured to receive the respiratory parameters and to produce therefrom values of a plurality of spirometry parameters; wherein the spirometry parameters comprise at least:Forced Vital Capacity (FVC);Forced Expiratory Volume (FEV); and Peak Expiratory Flow (PEF); wherein the output circuit is configured to output the values of the plurality of spirometry parameters.
2. The system of claim 1 , wherein the output circuit is configured to provide, based at least in part on the values of the spirometry parameters, at least one of: a Volume-Time loop; and a Flow-Volume loop.
3. The system of claim 1 , wherein the first neural network circuit is a deep-learning neural network that is trained using speech inputs and respiration device outputs.
4. The system of claim 1 , wherein the second neural network circuit is a deeplearning neural network that is trained using respiration inputs and spirometry outputs.
5. The system of claim 1 , wherein: the first neural network is trained using speech inputs and respiratory device outputs; and the second neural network is trained using outputs of the first neural network and spirometry device outputs.
6. The system of claim 1 , wherein at least one of the first or second neural network circuit comprises a Bayesian neural network.
7. The system of claim 1 , wherein the values of the FEV parameter comprise values of:Forced Expiration Volume at 1 second (FEV1);Forced Expiration Volume at 25% (FEV25); andForced Expiration Volume at 75% (FEV75).
8. The system of claim 1 , wherein the respiratory parameters comprise respiration rate, inhalation duration, and exhalation duration.
9. The system of claim 1 , wherein at least one of the first and second neural networks receive one or more physical characteristics related to the subject as part of an input feature set associated with the at least one of the first and second neural networks.
10. A method of estimating a plurality of spirometry parameters based on a speech signal of a subject comprising: receiving the speech signal; processing the speech signal to provide a speech feature set to a first neural network circuit; wherein the first neural network circuit provides values of a respiratory signal corresponding to the speech signal based on the speech feature set; processing the respiratory signal to provide a respiratory feature set to a second neural network; wherein the second neural network circuit provides values of a spirometry output based on the respiratory feature set; processing the spirometry output to produce the plurality of spirometry parameters; wherein the plurality of spirometry parameters comprise at least:Forced Vital Capacity (FVC);Forced Expiratory Volume (FEV); and Peak Expiratory Flow (PEF); rendering, on an output device, the values of the plurality of spirometry parameters.11 . The method of claim 10, wherein the values of the spirometry parameters are rendered via at least one of: a Volume-Time loop; and a Flow-Volume loop.
12. The method of claim 10, wherein the first neural network circuit is a deeplearning neural network that is trained using speech inputs and respiration device outputs.
13. The method of claim 10, wherein the second neural network circuit is a deeplearning neural network that is trained using respiration inputs and spirometry outputs.
14. The method of claim 10, wherein:the first neural network is trained using speech inputs and respiratory device outputs; and the second neural network is trained using outputs of the first neural network and spirometry device outputs.
15. The method of claim 10, wherein at least one of the first or second neural network circuit comprises a Bayesian neural network.
16. The method of claim 10, wherein the values of the FEV parameter comprise values of:Forced Expiration Volume at 1 second (FEV1);Forced Expiration Volume at 25% (FEV25); andForced Expiration Volume at 75% (FEV75).
17. The method of claim 10, wherein the respiratory parameters comprise respiration rate, inhalation duration, and exhalation duration.
18. The method of claim 10, wherein at least one of the first and second neural networks receive one or more physical characteristics related to the subject as part of an input feature set associated with the at least one of the first and second neural networks.
19. A non-transitory computer-readable medium comprising instructions that, when executed by a processing system, performs the method of claim 10.
20. A Respiration-to-Spirometry circuit comprising: a respiration processor that is configured to receive at least one of a respiration signal and a plurality of respiration parameters corresponding to a speech signal of a subject, and to provide therefrom an input feature set; and a respiration-to-spirometry neural network circuit that is configured to receive the input feature set and to produce therefrom an estimated spirometry output; wherein the estimated spirometry output enables a determination of a plurality of spirometry parameters; wherein the plurality of spirometry parameters comprises at least:Forced Vital Capacity (FVC);Forced Expiratory Volume (FEV); and Peak Expiratory Flow (PEF).21 . The Respiration-to-Spirometry circuit of claim 20, wherein the spirometry parameters enable a generation of at least one of: a Volume-Time loop; and a Flow-Volume loop.
22. The Respiration-to-Spirometry circuit of claim 20, wherein the respiration-to- spirometry neural network circuit is a deep-learning neural network that is trained using respiration inputs and spirometry outputs.
23. The Respiration-to-Spirometry circuit of claim 20, wherein the values of the FEV parameter comprise values of:Forced Expiration Volume at 1 second (FEV1);Forced Expiration Volume at 25% (FEV25); and Forced Expiration Volume at 75% (FEV75).
24. The Respiration-to-Spirometry circuit of claim 20, wherein the plurality of respiratory parameters comprises respiration rate, inhalation duration, and exhalation duration.
25. The Respiration-to-Spirometry circuit of claim 20, wherein the respiration-to- spirometry neural network receives one or more physical characteristics related to the subject as part of the input feature set.
Citation Information
Patent Citations
Breathing signal-dependent speech processing of an audio signal
EP4414984A1
Speech-based breathing prediction
US20210146082A1
Speech-based pulmonary assessment
US20220257175A1
Systems and methods for estimation of forced vital capacity using speech acoustics
US20240049981A1