Speech-based breathing prediction

A machine learning model predicts inspiration times from speech patterns to enhance ventilation system efficiency and accuracy, addressing inefficiencies in gas delivery for respiratory support.

JP7775189B2Active Publication Date: 2025-11-25KONINKLIJKE PHILIPS NV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022525026
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-11-18
Filing Date
2020-11-17
Publication Date
2025-11-25
Estimated Expiration
2040-11-17

AI Technical Summary

Technical Problem

Ventilation systems for respiratory support may be inefficient due to slow respiratory rates and difficulties in accurately monitoring breathing patterns, leading to inadequate gas delivery and potential health issues for patients.

Method used

A method using a machine learning model to predict inspiration times based on speech patterns, allowing for precise control of gas delivery to align with the subject's breathing cycle, reducing the need for additional monitoring equipment.

Benefits of technology

Improves the efficiency and accuracy of gas delivery to patients, ensuring timely support during inhalation and reducing the burden of additional equipment setup.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007775189000001
    Figure 0007775189000001
  • Figure 0007775189000002
    Figure 0007775189000002
  • Figure 0007775189000003
    Figure 0007775189000003
Patent Text Reader

Abstract

In one embodiment, a method is described. The method includes obtaining an indication of a subject's voice pattern and using the indication to determine a time of expected inspiration by the subject. A machine learning model is used to predict a relationship between the subject's voice pattern and breathing pattern. The machine learning model is then used to determine a time of expected inspiration by the subject. The method further includes controlling delivery of gas to the subject based on the time of expected inspiration by the subject.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method, apparatus and tangible machine-readable medium for controlling the delivery of gas to a subject, such as a patient. [Background technology]

[0002] Subjects, such as ventilator-supported patients, are in need of supportive therapy for various respiratory disorders, such as chronic obstructive pulmonary disease (COPD). Respiratory support for such disorders is provided by a ventilation system for delivering a gas, such as therapeutic air, to the subject. The ventilation system can deliver the therapeutic air to the subject at a particular oxygen level and / or pressure selected for the subject's individual therapeutic requirements. The therapeutic air can be administered using an interface, such as a nasal cannula or an oral and / or nasal mask. In some cases, delivery of the therapeutic air is triggered by detecting an attempt by the subject to breathe spontaneously.

[0003] The lungs are Speech and breathing, which in the case of a ventilated patient Speech disability-related difficulties with social interaction, and / or Speech This can lead to health problems due to impaired gas exchange during the procedure.

[0004] The subject's breathing rate is usually faster when the subject is not talking than when the subject is not talking. Speech For example, in healthy subjects, the respiration rate is Speech During exhalation, the air is expelled by 50%. Speech Most of the intake is done during Speech Because this is done with a short pause in the inhalation, healthy people Speech The breathing pattern is asymmetrical due to the relatively long exhalation during the breath. SpeechThe corresponding temporary increase in carbon dioxide and decrease in oxygen levels in the lungs is generally not a problem for healthy subjects, but can cause discomfort in patients on some ventilators. Speech may require additional support (e.g., more oxygen) during Summary of the Invention [Problem to be solved by the invention]

[0005] However, ventilation systems for providing respiratory support to patients may suffer from, for example, slow respiratory rates and / or Speech Due to the relatively long exhalation period associated with Speech In some cases, such assistance may not be very effective or efficient. Furthermore, attempting to directly monitor the subject's breathing patterns involves the use of additional equipment, which places an additional burden on the subject in terms of setting up and using the additional equipment.

[0006] Therefore, the object of the present invention is to Speech Another object of the present invention is to improve the support provided to subjects receiving gas during Speech The goal is to improve the performance of gas delivery to subjects during the procedure. [Means for solving the problem]

[0007] Aspects or embodiments described herein may include: Speech improving the support provided to subjects receiving gas during Speech Aspects or embodiments described herein relate to improving gas delivery to a subject during an inhalation. Speech and / or assisting the subject during Speech This can eliminate one or more of the problems associated with delivering gas to a subject during inhalation.

[0008] In a first aspect, a method is described. The method comprises: SpeechThe method includes obtaining an indication of a speech pattern. The method further includes using the indication to determine a time of expected inspiration by the subject. The determination is performed by a processing circuit. The determination is based on the subject's Speech The method is based on a machine learning model for predicting a relationship between a breathing pattern and a breathing pattern. The method further comprises controlling delivery of gas to the subject based on the predicted duration of inspiration by the subject.

[0009] In some embodiments, the method includes deriving a respiratory signal from the display and using the respiratory signal as input to a machine learning model to predict (e.g., using processing circuitry) a time of inspiration by the subject.

[0010] In some embodiments, the machine learning model is obtained from multiple trainers. Speech The signal is constructed using a neural network configured to identify any correlation between the signal and the corresponding respiratory signal.

[0011] In some embodiments, the neural network is obtained from a trainer. Speech The method is configured to identify at least one of linguistic content and prosodic features of the signal to facilitate identifying said correlation.

[0012] In some embodiments, the method includes providing a ventilation system with: Delivering gas to the subject for a specific time period during the expected inspiration time The step of causing the specific The time period is pre-determined The time period is one that is adapted to the individual needs of the subject.

[0013] In some embodiments, the individual needs of the subject are determined by the subject's Speech The duration of the previous inspiration by the subject and the medical need of the subject are determined based on at least one of the linguistic content of the

[0014] In some embodiments, the method comprises: Speech Using change-point detection to predict a time of inspiration for the subject based on the subject's respiratory signal as predicted by the machine learning model based on the pattern.

[0015] In a second aspect, an apparatus is described, the apparatus comprising a processing circuit, the processing circuit comprising a prediction module, the prediction module detecting a time of predicted inspiration by the monitored subject, the time of the predicted inspiration by the monitored subject, and a time of predicted inspiration by the monitored subject. Speech The determination is configured to use an indication of the pattern of the subject. Speech The processing circuitry further includes a control module configured to control delivery of gas to the subject based on the predicted duration of inspiration by the subject.

[0016] In some embodiments, the device Speech Corresponding to patterns Speech It has an acoustic transducer configured to acquire a signal.

[0017] In a third aspect, a tangible machine-readable medium is described, the tangible machine-readable medium, when executed by at least one processor, causing the at least one processor to Speech The method stores instructions for determining the expected duration of inspiration by the subject from the pattern representation. Speech The instructions further cause the at least one processor to control delivery of gas to the subject based on the predicted duration of inspiration by the subject, the prediction being based on a machine learning model for predicting a relationship between a duration of inspiration and a breathing pattern.

[0018] In some embodiments, the machine learning model comprises a plurality of trainees. Speech The signal is trained with the corresponding respiratory signal.

[0019] In some embodiments, the input to the machine learning model is: specific Multiple time intervals Speech The input comprises a spectral representation of the signal and a representation of the corresponding respiratory signal. The input is sent to a neural network having multiple memory layers such that when the neural network is optimized to update the network weights based on the input, the machine learning model is updated accordingly.

[0020] In some embodiments, a plurality of Speech A spectral representation of each of the signals is obtained. In one embodiment, the spectral representations are Speech The signal is spectrally flattened, Speech Each Speech filtering the signal; Speech A Mel spectrogram is obtained by applying a Fourier transform to obtain a power spectrum corresponding to the signal, applying Mel frequency scaling to the power spectrum to obtain a Mel spectrogram, and selecting multiple time windows from the Mel spectrogram, each time window separated by a specified stride interval. A representation of the corresponding respiratory signal at the specified time intervals is obtained by acquiring a respiratory inductive plethysmography (RIP) signal from the training subject and determining the RIP signal value at the end of each time window that falls within the specified stride interval.

[0021] In some embodiments, the neural network comprises at least one of a recurrent neural network (RNN), an RNN-long short-term memory (RNN-LSTM) network, and a convolutional neural network (CNN).

[0022] In some embodiments, an attention mechanism using respiration rate as an auxiliary training parameter is used to optimize the neural network.

[0023] Aspects or embodiments described herein may include: Speech provide improved assistance to subjects when receiving gas during Speech For example, aspects or embodiments described herein may improve the performance of gas delivery to a subject during Speech and / or providing improved gas delivery to the subject to assist in Speech The present invention can provide improved delivery of gas to the subject to assist breathing during breathing.

[0024] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. [Brief explanation of the drawings]

[0025] Exemplary embodiments of the present invention will now be described, by way of example only, with reference to the following drawings, in which: [Figure 1] FIG. 1 illustrates a method for controlling gas delivery according to one embodiment. [Figure 2a] FIG. 2a is a schematic diagram of a ventilation system according to one embodiment. [Figure 2b] FIG. 2b is a schematic diagram of a ventilation system according to one embodiment. [Figure 3] FIG. 3 is a schematic diagram of a system for training and testing machine learning models according to one embodiment. [Figure 4] FIG. 4 is a graph of experimental results from testing the machine learning model referenced in FIG. [Figure 5] FIG. 5 illustrates a method for controlling gas delivery according to one embodiment. [Figure 6] FIG. 6 is a schematic diagram of an apparatus for controlling gas delivery according to one embodiment. [Figure 7] FIG. 7 is a schematic diagram of an apparatus for controlling gas delivery according to one embodiment. [Figure 8] FIG. 8 is a schematic diagram of a machine-readable medium for controlling delivery of gas according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0026] 1 illustrates a method 100 (e.g., a computer-implemented method) for controlling the delivery of a gas, e.g., therapeutic air, to a subject, e.g., a patient wearing a ventilator. Method 100 can be used to control the supply of gas provided by a ventilation system (an example of which is described in more detail below in connection with FIGS. 2a-2b). For example, method 100 provides instructions or other indications to a ventilation system to control how the ventilation system delivers gas. For example, the timing and / or duration of gas delivery can be controlled based on the instructions or other indications provided by method 100.

[0027] The method 100 begins at block 102 by measuring the subject's Speech obtaining a representation of the pattern of the subject. Speech The pattern can be obtained from an acoustic transducer, such as a microphone, for detecting sound and generating a signal representative of the detected sound. Speech A pattern comprises characteristic features, such as prosodic features and / or linguistic content, present in the signal produced by the acoustic transducer.

[0028] The method 100 includes, at block 104, Speech and using the indication to determine (e.g., using processing circuitry) a predicted time of inspiration by the subject based on a machine learning model for predicting a relationship between the pattern and the subject's breathing pattern.

[0029] The breathing pattern has two phases: inspiration (i.e., breathing in) and expiration (i.e., breathing out). Speech The pattern can be adapted (either voluntarily or unconsciously) by the subject. Speech The pattern can have characteristic features (e.g., prosodic features and / or linguistic content) that can indicate whether the subject is inhaling or exhaling. For example, Speech A pause in the middle may indicate that the subject is emotional or about to be emotional. Speech A change in pitch or rate of speech may indicate that the subject has just been moved or is about to be moved. Speech consists of sentences, during which the subject exhales, and in between, the subject inhales. Speech These are just a few examples of how some characteristic features of the pattern relate to the subject's breathing pattern.

[0030] in fact, Speech The patterns are complex and subjective Speech As it is difficult to design a reliable model to predict the relationship between the pattern and the subject's breathing pattern (e.g., Speech How the subject's breathing pattern changes (during the same time period or between different subjects) Speech The above example of pattern dependence simply refers to how the subject's breathing pattern affects the subject's Speech The illustrative hypotheses are related to the patterns, Speech The patterns should not be considered limiting due to the complexity and / or variability of breathing patterns.

[0031] For example, by monitoring signals generated by an airflow sensor, an air pressure sensor, and / or a microphone, Speech It is possible to detect pauses and attempts at inspiration during breathing. SpeechThe duration of inspiration in a ventilator typically lasts for several hundred milliseconds, which is too fast for certain ventilators (e.g., mechanical ventilators) to react within a short enough time frame to deliver gas once an attempted inspiration is detected. For example, a ventilator that delivers gas through an interface such as a nasal cannula connected to the ventilator via a hose will, upon receiving an indication that the ventilator is to deliver gas, require a certain amount of time to deliver gas that depends on the length of the hose (and the reaction speed of the ventilator). specific For example, for the moment of inspiration, which lasts for several hundred milliseconds, the subject receives the gas too late. Speech are not adequately supported by a ventilation system during Speech This can create artifacts in the air pressure and flow signals that can make it difficult to detect actual breathing. Furthermore, the moment of inspiration can be easily misread and misinterpreted. studies content and / or Speech Thus, for example, based on sensor data provided by an airflow sensor, an air pressure sensor and / or a microphone, Speech Attempting to detect pauses and attempts to inhale during Speech This does not necessarily allow for adequate support to be given to the subject during the procedure.

[0032] The machine learning model referenced in method 100 may be adapted to predict the breathing patterns of a subject with acceptable reliability. Speech This machine learning model is used to interpret the subject's breathing patterns to provide a prediction of the subject's breathing patterns. Speech As described in more detail herein, machine learning models may be used to interpret complex and / or variable patterns in the patterns. SpeechThe machine learning approach can be trained using information from a training dataset derived from the muscular patterns and breathing patterns, without building a model that relies on specific assumptions (e.g., the exemplary assumptions described above) that would otherwise result in erroneous predictions due to possible biases and / or errors in the assumptions. Speech It is possible to provide a simplified method for modeling breathing patterns and because machine learning models can avoid making or rely less on certain assumptions, predictions based on machine learning models are more reliable than models that rely on assumptions that would otherwise be subject to bias and / or error.

[0033] Method 100 further includes controlling delivery of gas to the subject based on a predicted time of inspiration by the subject at block 106. For example, method 100 generates an indication (e.g., an inspiration signal) that is received by a ventilator of a ventilation system to cause the ventilator to deliver gas to the subject based on the predicted time of inspiration.

[0034] The machine learning model is Speech The method 100 may be used to provide a prediction of the time of inspiration (e.g., the starting point and / or duration of an inspiration attempt) by a subject during a breathing session, so that delivery of gas by a ventilation system can be activated at the predicted time of inspiration. For example, if the ventilation system has a particular reaction time (e.g., due to the reaction speed of a ventilator and / or the length of a hose connecting the ventilator to an interface), the prediction may be Speech In other words, the machine learning model allows method 100 to determine the subject's ability to receive gas during a period of time to activate the delivery of gas in time to provide sufficient support to the subject. Speech This allows the ventilation system to proactively predict the time of inspiration based on the pattern. specificThis may provide sufficient time to respond to deliver gas within a time frame and / or may allow the ventilation system to deliver gas for a duration corresponding to the duration of inspiration by the subject. Additionally, method 100 may reduce or avoid the need for additional equipment, such as a body sensor for directly monitoring breathing, allowing an end user of the ventilation system (e.g., the subject themselves) to relatively easily configure the ventilation system (e.g., using a microphone or other acoustic detector). Speech Configuring a ventilation system to include monitoring data is believed to be relatively easy for end users to set up themselves.

[0035] Figure 2a schematically illustrates a ventilation system 200 according to an embodiment for at least partially implementing certain methods described herein, such as method 100 of Figure 1. In Figure 2a, a subject 202 is, in this embodiment, interfaced with nasal cannulae 204 for delivering gas to the subject 202 via hoses 206 connected to a ventilation device 208. The ventilation device 208 may be controlled (e.g., according to block 106 of method 100) such that certain gas parameters (e.g., gas flow rate, pressure, oxygen level, timing, and / or any other parameters related to the delivery of gas) are appropriate for the needs of the subject at a particular moment.

[0036] In this regard, the ventilation system 200 further includes a prediction module 210 for at least partially implementing certain methods described herein. For example, the prediction module 210 may implement at least one of blocks 102, 104, and 106 of the method 100. An input 212 to the prediction module 210 is the subject's 202 Speech The pattern can be provided to the prediction module 210. Speech Upon receiving the pattern, the prediction module 210 predicts the time of inspiration by the subject 202. This predicted time of inspiration is used to control the delivery of gas to the subject 202.

[0037] 2b shows a schematic representation of some of the modules of the prediction module 210. In this embodiment, and as will be explained in more detail below, the prediction module 210 performs a prediction (i.e., a calculation of the monitored breathing pattern) of the subject. Speech and converting the monitored data into a format suitable for input to a machine learning module 216 that outputs a predicted breathing pattern for the subject based on the monitored data. Speech , and a preprocessing module 214 for converting the predicted breathing pattern (e.g., a "predicted breathing wave") generated by the machine learning module 216. Based on the predicted breathing pattern (e.g., a "predicted breathing wave") generated by the machine learning module 216, the inspiration prediction module 218 can predict the time of inspiration by the subject 202 and generate a ventilator control signal 220 to cause the ventilator 208 to deliver gas to the subject at the time of the predicted inspiration by the subject 202.

[0038] Ventilator control signal 220 activates ventilator 208 at the start of the predicted inspiration to allow gas to flow to the subject's 202 lungs (e.g., by subject 202 inhaling gas or by gas being pushed by a gas pump). The amount (e.g., concentration or rate) of oxygen and / or pressure may also be adjusted according to the detected and / or predicted breathing rate to ensure a reduced minute ventilation and / or to prevent shortness of breath, hypoxemia, and / or hypercapnia. At the end of the predicted inspiration, ventilator control signal 220 deactivates ventilator 208, stopping the flow of gas and thus allowing subject 202 to exhale.

[0039] The output of the machine learning model (i.e., using the machine learning module 210) is a representation of the estimated or predicted respiratory signal. There may beIn one embodiment, a change-point detection algorithm (e.g., implemented by the inspiration prediction module 218) uses the estimated respiratory signal to predict the moment of inspiration of the subject 202. In one embodiment, the pump of the ventilator 208 is turned on a short time T (e.g., T=300 ms) before the expected (i.e., predicted) onset of inspiration so that gas is delivered to the subject 202 in time for inspiration. In one embodiment, the value of T may be individually optimized for each subject 202 (e.g., depending on the capacity of the ventilator 208, the preferred operating mode of the ventilator 208, and / or the individual requirements of the subject 202). The value of T is Speech Language studies In one embodiment, the duration of ventilation may depend on the subject's 202 context and / or may be based on previously observed inspiration pause durations. Speech Thus, in some embodiments, the value of T may be determined in advance based on the subject's Speech and / or selected based on a previous prediction of the subject's respiratory signal. value At least one of the following is true.

[0040] Figure 3 shows the results of the subjects Speech 2 b , a system 300 is shown in schematic form for training (and subsequently testing) a machine learning model 302 for predicting a subject's breathing pattern based on the trained respiratory rate (RRP) of the subject. The machine learning model 302 may be implemented by a machine learning module 216 as described in connection with FIG. 2 b. As described in more detail below, in some embodiments, the machine learning model 302 is based on a deep recurrent neural network or other sequentially recurrent algorithm. The machine learning model 302 may be implemented using a large amount of data collected from a trainer, for example, using airflow measurement sensors and / or body sensors. Speechand respiratory data. In the embodiment of FIG. 3, training respiratory signal 304 (i.e., "measured breathing pattern") is collected by a body sensor, which in this embodiment includes two respiratory elastic band sensors 306 arranged to monitor chest and / or abdominal movement of trainer 308 during breathing. Training respiratory signal 304 represents a respiratory inductive plethysmography (RIP) signal. In this embodiment, one of sensors 306 is placed around trainer 308's rib cage while the other of sensors 306 is placed around trainer 308's abdomen to detect chest and / or abdominal movement corresponding to trainer 308's breathing, although a different number of sensors (e.g., one or more) can be used and arranged as needed. As the trainer 308 breathes, movement of the trainer's chest and / or abdomen causes at least one of the respiratory elastic band sensors 306 to expand and / or contract, generating body movement signals 310 (e.g., thoracic signal 310a and abdominal signal 310b), which collectively represent the training breathing signal 304 (e.g., by combining the body movement signals 310).

[0041] In this embodiment, instead of or in addition to the microphone 312, Speech Although any other device for detecting Speech is detected using a microphone 312. The microphone 312 detects the trainer's 308 Speech Based on Speech Generate the data and Speech The data is processed by a training module (corresponding to the "preprocessing module 214" described in connection with FIG. 2). Speech The processing module 314 trains the data for input to the machine learning model 302. Speech The signal data 316 is processed. Speech The processing module 314 performs audio spectrum analysis (monitored by the microphone 312) Speech The data is converted into a format suitable for input to the machine learning model 302.

[0042] In this embodiment, Speech The data is processed as follows, using the values ​​shown: Speech The signal data 316 is divided into fixed time window lengths of 4 seconds with a stride of 10 milliseconds between adjacent windows (in FIG. 3, these windows are denoted by window lengths “ <ts>", strides are exaggerated in length for ease of understanding). Speech These windows of signal data 316 are: Speech The signal is processed with a filter (e.g., a pre-emphasis filter) to spectrally flatten the signal and increase the higher frequencies. A short-time Fourier transform (STFT) is computed using a short frame size of 25 ms, a stride of 10 ms, and a Hamming window to obtain a power spectrum. A Mel filter bank (in this embodiment, a Mel filter bank with n=40) is applied to the power spectrum to obtain a Mel spectrum. The Mel filter bank applies Mel frequency scaling, a perceptual scale that helps simulate the way the human ear and brain work to interpret sound. The Mel spectrum provides better resolution at low frequencies and lower resolution at relatively high frequencies. The training data is then used as input to the machine learning model 302. Speech A Log Mel spectrogram is generated to represent the spectral characteristics of the signal data 316. In another embodiment, the Log Mel spectrogram is generated by Speech When processing the data, different values ​​(eg, different window lengths, strides, and frame lengths) can be used.

[0043] To determine a training respiratory signal 304 to be used as another input to the machine learning model 302, the log-mel spectrogram is mapped with the training respiratory signal 304 at the end of a time window to train the machine learning model 302 with a stride of 10 milliseconds between the time windows. As shown in FIG. 3, each time window of the log-mel spectrogram is provided to the machine learning model 302, and the corresponding respiratory signal 304 at the end of these time windows is also provided to the machine learning model 302.

[0044] Thus, the input training data for training the machine learning model 302 is the Speech and training respiratory signal samples. Each trainer 308 is healthy (i.e., not suffering from a respiratory disease), and multiple trainers 308 are used to train the machine learning model 302. In an exemplary training session, 40 trainers 308 are instructed to read a phonetically balanced paragraph. In this example, the phonetically balanced paragraph read by the trainers 308 is known as "Rainbow Passage" (from Fairbanks, G. (1960). Voice and articulation drillbook, 2nd edn. New York: Harper & Row. Pp124-139), Speech This is a paragraph commonly used for training purposes.

[0045] In this embodiment, the machine learning model 302 is based on a recurrent neural network-long short-term memory (RNN-LSTM) network model. In this RNN-LSTM network model, the input training data is fed to a network with 128 hidden units and two long short-term memory layers with a learning rate of 0.001. The Adam optimizer is used as an optimization algorithm to iteratively update the network weights based on the input training data. The mean squared error is used as the regression loss function. The hyperparameters selected for the network are estimated through repeated experiments, although they could alternatively be selected randomly.

[0046] FIG. 4 is a graph showing experimental results of tests performed using the trained model 302 to estimate test subject respiratory signals 318 (i.e., "estimated breathing patterns" or "estimated respiratory signals") to cross-validate data from multiple trainers 308 (e.g., using "leave one subject" cross-validation). Thus, for each test subject, Speech The data can be, for example, training Speech A test module that can provide the same functionality as the processing module 314 Speech The processing module 320 is used to process the training from the remaining trainers 308. Speech 4 shows an exemplary comparison between the measured (or "actual") respiratory signal (e.g., "RIP signal") in the upper graph and the estimated respiratory signal (i.e., estimated or predicted respiratory signal 318) in the lower graph as a function of time (seconds) for a test subject.

[0047] Using the RNN-LSTM network model Speech Because estimating breathing patterns from data is a regression problem, two metrics are used to evaluate and compare the measured and estimated breathing signals. These metrics are the correlation and mean square error (MSE) of the estimated and measured breathing signals. Therefore, experimental results produced by a model providing a high correlation value and / or a low MSE can indicate that the trained model provides an acceptable or reliable estimation of the breathing signal. For example, experimental results from the test subject shown in FIG. 4 were found to estimate the test subject's breathing pattern with a correlation of 0.42 and an MSE of 0.0016 for the test subject's measured breathing signal. For example, experimental results from another test subject were found to estimate the test subject's breathing pattern with a correlation of 0.47 and an MSE of 0.0017.

[0048] Based on Model 302 training and testing, conversation Utterance of The trainer's breathing rate during speech was observed to be nearly half of the trainer's normal breathing rate (i.e., compared to the trainer's breathing rate when not speaking). Certain respiratory parameters, such as breathing rate and tidal volume, were determined for multiple trainers 308 based on their experimental results. As such, an average estimated breathing rate of 7.9 breaths per minute with an error of 5.6% was observed for multiple trainers 308. Furthermore, tidal volume was estimated with an error of 2.4%. As can be seen from FIG. 4 , specific respiratory events (e.g., inspiration and expiration points and their lengths) are evident from the estimated and measured respiratory signal. To determine specific respiratory events (e.g., inspiration points), an algorithm (e.g., a "change-point detection algorithm") may be implemented to identify peaks and / or troughs in the estimated respiratory signal. Thus, if the algorithm detects a change that appears to correspond to a respiratory event, this is compared to the measured respiratory signal to determine whether the detected change actually corresponds to that respiratory event. Based on experimental results from multiple trainers 308, inhalation events were identified with a sensitivity of 0.88, a precision of 0.82, and an F1 score of 0.8534.

[0049] The experimental results based on the above experiments show that the RNN-LSTM network model Speech Language studies We demonstrate that it is possible to learn and understand breathing dynamics based on semantic content and / or prosodic features. The model trained is: Speech The results presented above may be used to estimate the respiratory sensor value of the signal in real time (thus providing sufficient time for the ventilator to react to deliver gas at the time of inspiration). The results presented above may be used to ensure that the ventilator is adequately meeting the respiratory needs of the subject while the subject is speaking and / or Speech It has been demonstrated that the model 302 can be trained to provide sufficient sensitivity and / or accuracy to enable assisting the subject during the

[0050] In another embodiment, the recurrent neural network (RNN) is replaced with a convolutional neural network (CNN). Based on the same training and test data as described in the previous embodiment, the CNN was found to predict the actual respiratory signal with a correlation of 0.41 and a mean squared error of 0.00229. In another embodiment, for example, a memory network such as that described above can employ an attention mechanism, a multi-task learning-based approach, using respiratory rate (e.g., inhalation and exhalation rates) as an auxiliary training parameter to improve estimation of the predicted respiratory signal.

[0051] FIG. 5 is a flowchart of a method 500 for predicting a breathing pattern of a subject, according to one embodiment, for example, to enable control of a ventilation system (e.g., as described above in connection with FIG. 2) for delivering gas to the subject. If desired, certain blocks described in connection with method 500 may be omitted, and / or the arrangement / order of these blocks may be at least partially modified relative to the arrangement / order illustrated by FIG. 5. Method 500 may include at least one block corresponding to method 100 of FIG. 1. For example, block 502 may correspond to block 102 of FIG. 1, block 504 may correspond to block 104 of FIG. 1, and / or block 506 may correspond to block 106 of FIG. 1. Thus, method 500 may be implemented in conjunction with and / or include method 100 of FIG. 1. Furthermore, method 500 may be implemented by or in conjunction with certain modules or blocks as described herein and in connection with certain devices and systems (e.g., as shown in FIGS. 2 and / or 3). Thus, certain blocks described below may refer to certain features of other figures described herein.

[0052] The method 500 includes deriving a respiratory signal from the display at block 508. The respiratory signal is used as an input to a machine learning model that can be used (e.g., using processing circuitry) to predict the duration of inspiration by the subject.

[0053] As discussed above, multiple trainers can be used to train a machine learning model (e.g., machine learning model 302 of FIG. 3). In this regard, a machine learning model is obtained from multiple trainers at block 510. Speech The neural network is constructed using a neural network configured to identify any correlations between the respiratory signal and the corresponding respiratory signal. If any correlations are identified, the correlations are used to update the network weights of the neural network in order to improve predictions made based on the neural network. By using a neural network, potentially large amounts of training data used as input to the neural network are analyzed to improve predictions of the respiratory signal without using predetermined models that may be subject to bias (i.e., based on assumptions of human analysts) and / or without making erroneous assumptions about specific correlations. Speech It can identify hard-to-find patterns in data.

[0054] A neural network is obtained from the trainer in block 512 to facilitate the identification of the correlations. Speech Signal Language studies The method is configured to identify at least one of prosodic content and / or prosodic features. Speech Signal Language studies The prosodic content and / or prosodic features are Speech This can provide context, which is useful for determining specific correlations that may not be easy to identify without machine learning techniques. For example, Speech Signal Language studies The target content is potentially complex and variable, making it difficult for a human analyst to identify a model that makes sufficiently reliable predictions. Speech It has a significant amount of information with patterns.

[0055] The method 500, at block 514, includes providing a ventilation system (e.g., as described above) with: Over a specific time period during the expected inspiration time delivering a gas to the subject. specific The time period is pre-determined The time period is one that is adapted according to the individual needs of the subject. specific The start of the time period may begin when the ventilator is turned on. specific The time period may indicate the duration for which gas is delivered to the subject. specific The time period may or may not correspond to the expected duration of inspiration. For example, if there is a delay due to the reaction time of the ventilator, specific The time period may be longer than inspiration to account for that delay.

[0056] The aforementioned specific If the time period is adapted according to the individual needs of the subject, these needs are adjusted in block 516 to the subject's Speech Language studies The subject's current state is determined based on at least one of the subject's current state, the subject's previous inspiration duration, and the subject's medical need. For example, the subject is speaking a sentence, Speech Language studies Decisions can be made based on the context and / or duration of previous inhalations, and if the subject is predicted to inhale between sentences (or at any other point) for a certain duration, specific The time period may be adjusted accordingly. Furthermore, if the subject has a specific medical need (e.g., a target oxygen level in the subject's lungs or any other medical need), specific time period can be adapted accordingly to provide sufficient gas (eg, to reach the target oxygen level).

[0057] Although not shown in FIG. 5, the method 500 may be used to Speech The method may include using change-point detection (as described above) to predict the subject's inspiration time based on the subject's respiratory signal as predicted by a machine learning model based on patterns.

[0058] Figure 6 is a schematic diagram of an apparatus 600 according to an embodiment for performing certain methods described herein. Where necessary, the apparatus 600 will be described with reference to certain components of Figure 2 for ease of reference. The apparatus 600 includes processing circuitry 602 that may be implemented, for example, in at least one of the prediction module 210 and the ventilator 208 of Figure 2.

[0059] The processing circuit 602 includes a prediction module 604 that can at least partially implement certain methods described herein, for example, as described in connection with Figures 1 and / or 5, and / or provide at least partially the functionality described in connection with the systems of Figures 2 and / or 3. In this embodiment, the prediction module 604 predicts the subject's Speech of the monitored subject to determine an expected time of inspiration by said subject based on a machine learning model for predicting a relationship between the pattern and the breathing pattern. Speech It is configured to use a representation of the pattern.

[0060] Processing circuit 602 further includes a control module 606 that can implement certain methods described herein, for example, those described in connection with FIGS. 1 and / or 5, and / or provide functionality described in connection with the device or system of FIGS. 2 and / or 3. In this embodiment, control module 606 is configured to control delivery of gas to the subject based on the predicted time of inspiration by the subject. For example, control module 606 generates a ventilator control signal (as described in connection with FIG. 2) to cause a ventilator to deliver gas to the subject at the time of inspiration by the subject. In some embodiments, device 600 can form part of a ventilator such as described above in connection with FIG. 2. In some embodiments, device 600 can be a separate entity (e.g., a separate computer or server, etc.) communicatively coupled to a ventilator and configured to provide instructions or other indications to the ventilator so that the ventilator delivers gas at the time determined by device 600.

[0061] 7 is a schematic diagram of an apparatus 700 according to an embodiment for implementing certain methods described herein. In this embodiment, the apparatus 700 includes the processing circuit 602 of FIG. 6 and a subject's Speech Corresponding to patterns Speech The device 600 or 700 further includes a processing circuit 702 having an acoustic transducer 704 (e.g., a microphone) configured to acquire a signal. In some embodiments, the device 600 or 700 can further include a ventilation device as described in connection with FIG.

[0062] 8 schematically illustrates a machine-readable medium 800 (e.g., a tangible machine-readable medium) according to an embodiment that stores instructions 802 that, when executed by at least one processor 804, cause the at least one processor 804 to perform a particular method described herein (e.g., method 100 of FIG. 1 or method 500 of FIG. 5). The machine-readable medium 800 may be implemented in a computing system, e.g., a computer or server, for controlling a ventilator and / or may be implemented by the ventilator itself.

[0063] Command 802 is the subject's Speech and at least one processor 804 based on a machine learning model for predicting a relationship between the subject's breathing pattern and the subject's respiratory pattern. Speech From the pattern representation, instructions 806 are included to determine the expected time of inspiration by the subject.

[0064] The instructions 802 further include instructions 808 for causing the at least one processor 804 to control delivery of gas to the subject based on an expected duration of inspiration by the subject.

[0065] The training of the machine learning model used to predict inspiration time in accordance with instruction 802 above is described in more detail with reference to Figure 3 and the associated description. As noted above, the machine learning model is trained using multiple training data obtained from multiple trainers. Speech The signal and the corresponding respiratory signal can be used for training.

[0066] The input to the machine learning model is a set of multiple inputs (e.g., from each trainer). Speech The input may comprise a spectral representation of the signal, which may comprise a log-mel spectrogram as described above. specific It may further comprise a display of the corresponding respiratory signal at the time interval. Speech The input may comprise or represent a respiratory signal obtained at the end of each time window selected from the signal data. The input may be provided to a neural network (e.g., any of the neural networks described above) having multiple memory layers such that when the neural network is optimized to update the network weights based on the input, the machine learning model is updated accordingly.

[0067] The plurality of Speech The spectral representation of each signal is Speech By filtering the signal, Speech The signal is spectrally flattened to Speech This can be achieved by boosting the higher frequencies of the signal relative to the lower frequencies. A Fourier transform (e.g., STFT) can be applied to the spectral representation to obtain the Speech A power spectrum corresponding to the signal can be obtained. Mel frequency scaling can be applied to the power spectrum to obtain a mel spectrogram (which in some embodiments may be a log-mel spectrogram). Multiple time windows can be selected from the mel spectrogram, where each time window is separated by a specified stride interval. In the embodiment of FIG. 3, each time window has a duration of 4 seconds and is separated from subsequent time windows by a stride of 10 milliseconds. It is within this stride interval that a representation of the corresponding respiratory signal is obtained. In other embodiments, the length of the time windows and / or strides may differ from those shown in the above embodiment.

[0068] In this embodiment, specific The corresponding respiratory signal representations at time intervals are obtained by obtaining respiratory inductive plethysmography (RIP) signals from the trained subject, with RIP signal values ​​determined at the end of each time window (i.e., within a specified stride interval).

[0069] In one embodiment, the neural network may include at least one of a recurrent neural network (RNN), an RNN long-short-term memory (RNN-LSTM) network, and a convolutional neural network (CNN). Although described above, other neural networks may also be used.

[0070] In one embodiment, the neural network can be optimized using an attention mechanism with respiration rate as an auxiliary training parameter.

[0071] While the invention has been illustrated and described in detail in the drawings and foregoing description, such illustration and description is to be considered illustrative or exemplary and not restrictive; that is, the invention is not limited to the disclosed embodiments.

[0072] One or more features described in one embodiment may be combined with or substituted for features described in another embodiment. For example, the methods 100, 500 of Figures 1 and / or 5 may be modified based on features described in connection with the systems of Figures 2 and / or 3, and vice versa.

[0073] Embodiments of the present disclosure may be provided as a method, a system, or a combination of machine-readable instructions and processing circuitry. Such machine-readable instructions may be contained on a non-transitory machine (e.g., computer) readable storage medium (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) having computer-readable program code therein or thereon.

[0074] The present disclosure will be described with reference to flowcharts and block diagrams of methods, apparatuses, and systems according to embodiments of the present disclosure. Although the flowcharts described above show a specific order of execution, the order of execution may differ from that shown. Blocks described in connection with one flowchart may be combined with blocks in another flowchart. It should be understood that each block in the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, may be implemented by machine-readable instructions.

[0075] The machine-readable instructions may be executed, for example, by a general-purpose computer, a special-purpose computer, or an embedded processor of another programmable data processing device to implement the functions described above and in the figures. In particular, a processor or processing circuit, or modules thereof, may execute the machine-readable instructions. Thus, functional modules of ventilation system 200 (e.g., prediction module 210, pre-processing module 214, machine learning module 216, and / or inhalation prediction module 218) and / or functional modules of system 300 (e.g., training module 216) may be programmed to execute the machine-readable instructions. Speech Processing Module 314 and / or Test Speech The processing module 320) and apparatus are implemented by a processor that executes machine-readable instructions stored in a memory or operates according to instructions embedded in logic circuitry. The term "processor" should be interpreted broadly to include a CPU, processing unit, ASIC, logic unit, or programmable gate array, etc. All of the methods and functional modules may be performed by a single processor or may be divided among multiple processors.

[0076] Such machine-readable instructions may also be stored in a computer-readable storage device capable of directing a computer or other programmable data processing apparatus to operate in a particular mode.

[0077] Such machine-readable instructions may be loaded into a computer or other programmable data processing apparatus such that the computer or other programmable data processing apparatus performs a sequence of operations to generate computer-implemented processes, such that the instructions executed on the computer or other programmable apparatus implement the functions specified by the blocks in the flowcharts and / or block diagrams.

[0078] Furthermore, the teachings herein may be embodied in the form of a computer program product, the computer program product being stored on a storage medium and having a plurality of instructions for causing a computer device to perform the methods set forth in the embodiments of the present disclosure.

[0079] Elements or steps described in connection with one embodiment may be combined with or substituted for elements or steps described in connection with another embodiment. Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the invention, from studying the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other elements or steps, nor does it exclude a plurality of them, even if a plurality is mentioned. A single processor or other unit may fulfill the functions of several items recited in the claims. The mere fact that certain means are recited in mutually different dependent claims does not indicate that a combination of these means cannot be used to advantage. A computer program may be stored or distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, or in other forms, such as via the Internet or other wired or wireless telecommunication systems. Any reference signs in the claims should not be construed as limiting the scope.< / ts>

Claims

1. In an apparatus having a processing circuit, the processing circuit a prediction module configured to use a representation of the subject's speech pattern to determine a predicted time of inspiration by the subject based on a machine learning model trained to predict a relationship between the subject's speech pattern and breathing pattern; and a control module configured to control delivery of gas to the subject based on the predicted duration of inspiration by the subject. An apparatus having:

2. 10. The apparatus of claim 1, further comprising an acoustic transducer configured to acquire speech signals corresponding to the subject's speech patterns.

3. 3. The apparatus of claim 1 or 2, wherein the prediction module is configured to derive a respiratory signal from the indication and predict a time of inspiration by the subject.

4. 4. The apparatus of claim 1, wherein the machine learning model comprises a neural network configured to identify any correlations between speech signals and corresponding breathing signals obtained from a plurality of trainers.

5. The apparatus of claim 4 , wherein the neural network is configured to identify at least one of linguistic content and prosodic features of a speech signal obtained from the trainer to facilitate identifying the correlation.

6. 6. The apparatus of claim 1, wherein the control module is configured to cause the ventilation system to deliver gas to the subject for a specific period of time during the predicted time of inspiration, the specific period of time being one of a predetermined period of time or a period of time adapted according to the individual needs of the subject.

7. 7. The apparatus of claim 6, wherein the individual needs of the subject are determined based on at least one of the linguistic content of the subject's speech, the duration of a previous inspiration by the subject, and the medical needs of the subject.

8. 8. The apparatus of claim 1, wherein the prediction module is configured to predict a time of inspiration of the subject using a change-point detection algorithm, the change-point detection algorithm using a respiratory signal predicted by the machine learning model based on the speech pattern of the subject.

9. When executed by at least one processor, the method causes the at least one processor to: determining, from a representation of a speech pattern of a subject, a predicted time of inspiration by the subject based on a machine learning model trained to predict a relationship between the speech pattern and breathing pattern of the subject; and controlling delivery of gas to the subject based on the predicted duration of inspiration by the subject. A tangible, machine-readable medium that stores instructions.

10. 10. The tangible, machine-readable medium of claim 9, wherein the machine learning model is trained using input training data obtained from a plurality of trainers, the input training data comprising speech signals and corresponding breathing signals obtained from each trainer.

11. The input training data for the machine learning model is: a spectral representation of the speech signal obtained from each trainer; and a display of the corresponding respiratory signals obtained from each trainer; 11. The tangible, machine-readable medium of claim 10, wherein the input training data is provided to a neural network having multiple memory layers such that when the neural network is optimized to update network weights based on the input training data, the machine learning model is updated accordingly.

12. The spectral representation of each of the plurality of speech signals comprises: filtering each speech signal to spectrally flatten the speech signal and boost higher frequencies relative to lower frequencies of the speech signal; applying a Fourier transform to obtain a power spectrum corresponding to the speech signal; applying mel-frequency scaling to the power spectrum to obtain a mel spectrogram; and Selecting multiple time windows from the mel spectrogram each time window is separated by a specified stride interval, and the corresponding respiratory signal representation is obtained by: acquiring respiratory inductive plethysmography (RIP) signals from the subject to be trained; and Determine the RIP signal value at the end of each time window within the specified stride interval.

12. The tangible, machine-readable medium of claim 11 obtained by:

13. 13. The tangible, machine-readable medium of claim 11 or 12, wherein the neural network comprises at least one of a recurrent neural network (RNN), an RNN long short-term memory (RNN-LSTM) network, and a convolutional neural network (CNN).

14. 14. The tangible, machine-readable medium of claim 11, 12 or 13, wherein an attention mechanism using respiration rate as an auxiliary training parameter is used to optimize the neural network.

Citation Information

Patent Citations

  • Respiration detection type chemical substance presenting device and respiration detector

    JP2008086741A

  • Karaoke apparatus for supercharging oxygen from microphone to singer

    JP2010145952A

  • Patients need a system and method for ventilation proportional to their effort.

    JP2011522621A

  • Device, method and program for detecting ingressive in voice

    JP2012032557A

  • System and method for determining a person's breathing

    WO2014045257A1