Calculation of respiratory rate using touchless monitoring and ai
A depth-sensing vision system processes breathing signals with machine learning models to accurately predict RR, addressing inefficiencies in touchless monitoring and enabling early detection of health complications.
Patent Information
- Application Number
- PCT/IB2025/051909
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-29
- Filing Date
- 2025-02-21
- Publication Date
- 2025-09-04
AI Technical Summary
Existing methods for measuring respiratory rate (RR) in clinical settings are inefficient and lack accuracy, particularly in touchless monitoring systems, which are crucial for early detection of health complications such as respiratory tract infections and respiratory depression.
A system using a depth-sensing vision system processes breathing signals from a camera to generate an input feature matrix, which is trained with a machine learning model to predict RR, employing techniques like FFT, STFT, and wavelet transforms to filter and transform the breathing signal, and uses convolutional neural networks (CNN) for accurate prediction.
The system provides highly accurate RR predictions with a root mean square deviation of approximately three breaths per minute, enhancing early detection of health complications through touchless monitoring.
Smart Images

Figure IB2025051909_04092025_PF_FP_ABST
Abstract
Description
CALCULATION OF RESPIRATORY RATE USING TOUCHLESSMONITORING AND AlCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 559,409, filed February 29, 2024, the entire content of which is incorporated herein by reference.Background
[0002] Respiratory rate (RR) is one of the common vital signs that is measured in clinical settings. The levels of RR provide information about the health of patients. Similarly, any significant change in the levels of RR are often early indicators of major health complications such as respiratory tract infections, respiratory depression associated with opioid consumption, anesthesia and / or sedation, as well as respiratory failure. Patient’s RR can be measured in a number of different ways, including using a spirometer, pulse oximeters, etc. Alternative technology may use touchless monitoring of patients to generate depth information using cameras and extracting the RR from the depth information.Summary
[0003] Implementations described herein disclose a method including determining, based on an image signal received from a camera focused on at least a portion of a patient, a breathing signal, selecting a training segment of the breathing signal fortraining a machine learning (ML) model, processing the training segment of the breathing signal to generate an input feature matrix, inputting the input feature matrix to the machine learning (ML) model, and training the ML model, using the input feature matrix and an observed respiratory rate (RR) of the patient.
[0004] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0005] Other implementations are also described and recited herein.Brief Descriptions of the Drawings
[0006] A further understanding of the nature and advantages of the present technology may be realized by reference to the figures, which are described in the remaining portion of the specification.
[0007] FIG. 1 illustrates an example implementation of a system for calculating respiratory rate using artificial intelligence (Al) as disclosed herein.
[0008] FIG. 2 indicates example operations of the system for calculating respiratory rate using Al as disclosed herein.
[0009] FIG. 3 illustrates example operations for processing depth image stream for a neural network according to the system for calculating respiratory rate using Al as disclosed herein.
[0010] FIG. 4 illustrates example operations for using an ML model to predict RR based on real-time depth image stream.
[0011] FIG. 5 illustrates alternative example operations of a neural network for predicting respiratory rate according to the system for calculating respiratory rate using Al as disclosed herein.
[0012] FIG. 6 illustrates an example graph indicating predicted respiratory rate from a trained convolutional neural network (CNN) based model.
[0013] FIG. 7 shows a portable non-contact subject monitoring system that includes a noncontact detector and a computing device.
[0014] FIG. 8 shows a semi-portable non-contact subject monitoring system that includes a non-contact detector and a computing device.
[0015] FIG. 9 shows a non-portable non-contact subject monitoring system that includes a non-contact detector and a computing device.
[0016] FIG. 10 is a block diagram illustrating a system including a computing device, a server, and an image capture device.Detailed Descriptions
[0017] Respiratory rate (RR) is one of the common vital signs that is measured in clinical settings. RR may be indicated as the number of breaths by a patient over a time period, such as number of breaths per minute. The levels of RR provide information about the health of patients. Similarly, any significant change in the levels of RR are often early indicators of major health complications such as respiratory tract infections, respiratory depression associated with opioid consumption, anesthesia and / or sedation, as well as respiratory failure. A patient’s RR can be measured in a number of different ways, including using a spirometer, capnograph,pulse oximeter, microphone, etc. Alternative technology may use touchless monitoring of patients to generate depth information using cameras and extracting the RR from the depth information.
[0018] The technology disclosed herein uses output from a depth-sensing vision system to calculate RR in a touchless manner. Specifically, a training segment of the breathing signal generated by the depth-sensing vision system is processed to generate an input feature matrix. Subsequently, the input feature matrix is used to train a machine learning (ML) model together with observed RR of the patient. Once the ML model is trained, a real-time segment of the breathing signal may be input into the trained ML model to generate predicted RR. Specifically, segments of the real-time breathing signals may be selected and used to generate predicted RR.
[0019] In one implementation, the breathing signal generated by the depth-sensing vision system may be a breathing flow signal determined from the image signal received from a camera. For example, depth of frames of signals from the camera system 114 are analyzed and the changes in the depth of such frames together with the width of a portion of the subject 102 experiencing such changes in the depth are used to calculate the flow of air breathed by the subject.
[0020] Segments of the breathing volume signal, referred to as training segments, are processed for generating the input feature matrix for the ML model. For example, the breathing flow signal may be in terms of amount of air breathed by the patient per a time period, such a ml / sec. Alternatively, the breathing signal generated by the depth-sensing vision system may be a volume signal, wherein the volume signal may be determined as an integral of the flow signal determined from the image signal received from a camera and segments of the volume signal are processed for generating the input feature matrix for the NN.
[0021] The processing of the breathing signal may include, for example, filtering the breathing signal to remove high or low frequency noise outside of the expected respiratory frequency range, removing a mean of the breathing signal to center the breathing signal around zero, scaling the breathing signal so that the maximum and minimum excursions are set to limits such as +1 and -1 respectively, transforming the breathing signal using Fast Fourier Transform (FFT), applying a time-frequency transform such as an STFT or wavelet transform to the breathing signal, down sampling the breathing signal, etc.
[0022] The ML model may be implemented in a number of different manners such as by using a convolutional neural network (CNN), a feed forward neural network (FFNN), a long short-term memory (LSTM) neural network, a modular neural network, a recurrent neural network (RNN), a residual network (ResNet) based model, etc.
[0023] FIG. 1 illustrates a non-contact subject monitoring system 100 for a subject 102, in this particular example an infant in a crib. It is noted that the systems and methods described herein are not limited to a crib, but may be used with a bed, a bassinette, an incubator, an isolette, or any other place where the subject 102 is. Similarly, the subject 102 is illustrated to be an infant, in alternative implementations, the subject 102 may be an adult.
[0024] The system 100 includes a non-contact detector system 110 placed remote from the subject 102. In this embodiment, the detector system 110 includes a camera system 114, particularly, a camera that may include an infrared (IR) detection feature. The camera system 114 may be a depth sensing camera system, such as a Kinect camera from Microsoft Corp. (Redmond, Washington) or a RealSenseTM D415, D435 or D455 camera from Intel Corp. (Santa Clara, California).
[0025] The camera system 114 is remote from the subject 102, in that it is spaced apart from and does not physically contact the subject 102. The camera system 114 may be positioned in close proximity to or on the crib. The camera system 114 has a field of view F that encompasses at least a portion of the subject 102. The field of view F may be selected to be at least the upper torso of the subject 102. However, as it is common for young children and infants to move within the confines of their crib, bed or other sleeping area, the entire area potentially occupied by the subject 102 (e.g., the crib) may be the field of view F.
[0026] The camera system 114 includes a depth sensing camera that can detect a distance between the camera system 114 and objects in its field of view F. Such information can be used to determine that the subject 102 is within the field of view of the camera system 114 and determine a region of interest (ROI) to monitor on the subject. The ROI may be the entire field of view F or may be less than the entire field of view F. Once an ROI is identified, the distance to the desired feature is determined and the desired measurement(s) can be made.
[0027] The measurements (e.g., one or more of depth signal, RGB reflection, light intensity) are sent to a computing device 120 through a wired or wireless connection 121. The computing device 120 includes a display 122, a processor 124, and memory 126 for storing data, software, computer instructions, etc. Sequential image frames of the subject 102 are recorded by the video camera system 114 and sent to the computing device 120 for analysis by the processor 124. The display 122 may be remote from the computing device 120, such as a video screen positioned separately from the processor and memory. Other embodiments of the computing device 120 may have different, fewer, or additional components than shown in FIG. 1. In some embodiments, the computing device 120 may be a server. In other embodiments, the computing device 120 of FIG. 1 may be connected to a server. The captured images (e.g.,still images or video) can be processed or analyzed at the computing device 120 and / or at the server to create a topographical map or image to identify the subject 102 and any other objects within the ROI.
[0028] The signals collected from the camera system 114 may be stored in the memory 126. For example, the signals from the camera system 114 may be stored as a depth image data stream 142. For example, the depth image data stream 142 may include depth of frames as captured by the camera system. Alternatively, the depth image data stream 142 may include depth of frames in the form or RGB signals, infra-red (IR) signals, etc.
[0029] Furthermore, the memory 126 may store various computer programs, software, instructions, etc., to process the data including the depth image data stream 142. In one implementation, the depth image data stream 142 may be processed to generate a breathing signal 144. The breathing signal 144 may be in the form of a flow signal representing the volume of breathing by the patient 102 measured in terms of ml of air breathed over a period such as ml / sec. Alternatively, the breathing signal 144 may be in the form of a volume signal generated as an integral of the flow signal over a segment of time.
[0030] In the illustrated implementation, the memory 126 may store one or more computer instructions or programs to implement a preprocessor 146 that is configured to process training segments of the depth image data stream 142 and / or the breathing signal 144 to generate an input feature matrix 148. Specifically, the preprocessor 146 may be provided atraining segment of the breathing signal with the length of the training segment being 20s, 30s, 40s, etc. The length of the training segment may be limited by the RR that the ML model 150 is trying to detect. Specifically, the lowest RR that needs to be detected may be used to determine the length of the breathing signal segment.
[0031] For example, a segment of the breathing signal is selected such that it contains at least one or more breaths of the patient. This allows the ML model 150 to determine the RR. Thus, for example, if the lowest RR that needs to be detected is three (3) breaths per minute giving each breath a length of 20 seconds, the length of the breathing signal segment that is processed to be input into the ML Model 150 may be three breath periods, which is equal to 3 * 20 = 60 seconds. Alternatively, the length of the breathing signal segment that is processed to be input into the ML Model 150 may be two breath periods, which is equal to 2 * 20 = 40 seconds, or one and a half breath periods, which is equal to 1.5 * 20 = 30 seconds, or a single breath period, which is equal to 1 * 20 = 20 seconds, etc.
[0032] On the other hand, if the lowest RR that needs to be detected is five (5) breaths per minute, giving each breath a length of 12 seconds, the length of the breathing signal segmentthat is processed to be input into the ML Model 150 may be three breath periods, which is equal to 3 * 12 = 36 seconds. Alternatively, the length of the breathing signal segment that is processed to be input into the ML Model 150 may be two breath periods, which is equal to 2 * 12 = 24 seconds, or one and a half breath periods, which is equal to 1.5 * 12 = 18 seconds, or a single breath period, which is equal to 1 * 12 = 12 seconds, etc.
[0033] The preprocessor 146 is configured to process the selected segment of the breathing signal 144 for input to the ML model 150. For example, the preprocessor 146 may filter the breathing signal 144 to remove high or low frequency noise outside of the expected respiratory frequency range. In one implementation, the expected frequency range may be determined based on other data 160 about the patient 102. For example, the other data 160 may include the age, weight, sex, etc., of the patient 102 and the expected respiratory range of patient 102 may be selected based on the age of the patient 102. In this case, a bandpass filter may be used to remove the high or low frequency noise outside of the expected respiratory frequency range.
[0034] In another implementation, the preprocessor 146 may remove the mean from the breathing signal 144 to center the breathing signal 144 around zero. Centering the breathing signal around zero allows setting thresholds, removing bias from the breathing signal 144. Alternatively, the preprocessor 146 may scale the breathing signal 144 so that the maximum and minimum excursions of the breathing signal 144 are set to limits such as +1 and -1 respectively.
[0035] Furthermore, the breathing signal 144 may be transformed by applying a timefrequency transform such as Fast Fourier Transform (FFT) to generate frequency-based information as the input feature matrix 148. Such frequency-based input feature matrix 148 makes the ML model 150 more efficient. Specifically, people breathe more regularly when they breathe at higher rate compared to when they breathe at a lower rate and using the FFT transform of the breathing signal 144 uses this observed characteristic of breathing signal 144. Alternatively, the preprocessor 146 may also apply a time-frequency transform such as short- time Fourier transform (STFT) to the breathing signal 144 to generate time-frequency-based information as the input feature matrix 148. Subsequently, the resulting two-dimensional or time-frequency matrix of values in the input feature matrix 148 may be input into the ML model 150. Similarly, the preprocessor 146 may also apply a time-frequency transform such wavelet transform to the breathing signal 144 to generate time-frequency-based information as the input feature matrix 148. Furthermore, such time-frequency-based information may be used in addition to frequency-based information and the time-based waveforms discussed above.
[0036] In other implementations, the breathing signal 144 may be transformed by down sampling the breathing signal 144 to generate the input feature matrix 148. For example, the breathing signal 144 may be collected at 30Hz, but it may be down sampled to 10Hz or 5Hz to generate the input feature matrix 148. Doing such down sampling requires fewer parameters in the network and makes training of the ML model 150 and generating inference more efficient without any loss in performance.
[0037] During a training phase of the ML model 150, the input feature matrix 148 is input to the ML model together with an observed respiratory rate (RR) of the patient 102. The ML model may iteratively adjust the weights of various kernels in one or more of its convolution blocks (further described below in FIG. 5), to train itself such that the RR predicted based on the input feature matrix 148 is substantially similar to the observed respiratory rate (RR) of the patient 102.
[0038] While the above implementations disclose processing of the breathing signal 144 to generate the input feature matrix 148 that is input the ML model 150, in an alternative implementation, the depth image data stream 142 may be input directly into the ML model 150 to train the ML model 150. In such implementation, the input to the ML model 150 may be the depth image data stream 142 in the form of two-dimensional depth matrix provided for a number of consecutive frames over time. For example, the depth image data stream 142 for each computation of RR may be provided within a 3 -dimensional spatiotemporal volume segment.
[0039] The input feature matrix 148 may be used to train the machine learning (ML) model 150. The ML model 150 may be a NN-based model and may be implemented in a number of different manners such as by using a convolutional neural network (CNN), a feed forward neural network (FFNN), a long short-term memory (LSTM) neural network, a modular neural network, a recurrent neural network (RNN), a residual network (ResNet) based model, etc.
[0040] In one implementation, the memory 126 may also store motion data 162 regarding the motion by the patient 102. Such motion data 162 may be generated also by the detector system 110. Subsequently, the ML model 150 may be prompted to not produce an output RR when a motion of the patient 102 is detected. Alternatively, the ML model 150 may be trained to use the motion data 162 to determine when not to post the RR from the ML model 150. In another implementation, the ML model 150 may be trained on the raw depth image data stream 142 and motion data 162 to predict both the RR and gross motion of the patient 102.
[0041] In an alternative implementation, the recent outputs from the ML model 150 are stored in an array and fed back into the ML model 150 together with the breathing signal 144or the input feature matrix 148. This allows the ML model 150 to become more stable and less likely to fluctuate due to noise in the breathing signal 144.
[0042] In one implementation disclosed herein, the depth image data stream 142 stores breathing signals over time for a number of patients. In such an implementation, the ML model 150 may be trained using breathing signals collected over a time period of, say a week, a month, etc. Furthermore, the memory 126 may also store the actual respiratory rate (RR) of the patients that are related to the breathing signals. For example, breathing signals of patient X collected over a period of time may be associated with actual RR of patient X so that such breathing signals and the actual RR of the signal can be used to train the ML model 150. Furthermore, such training of the ML model 150 may be done with breathing signals and RR of a number of different patients situated in a number of different clinical settings, such as in intensive care units (ICUs), neonatal ICUs (NICUs), general care floors (GCFs), etc. Such training of the ML model over a number of different patients and different clinical settings makes the trained ML model more robust.
[0043] However, in alternative implementations, separate ML models may be trained for each clinical setting. For example, given the nature of the breathing signals for neonatal patients in NICUs may be quite different from other clinical settings and therefore, an ML model may be trained for neonatal patients in NICUs. On the other hand, another model may be trained for stable patients in GCFs.
[0044] Once trained, the trained ML model 150 may be deployed to predict RR by inputting processed real-time breathing signals in field. As additional data including breathing signals, predicted RR, and actual RR are collected over time when the ML model 150 is in use, such additional data may be used to fine-tune the ML model 150 over time.
[0045] FIG. 2 illustrates operations 200 of the system for calculating respiratory rate using Al as disclosed herein. One or more of the operations 200 may be stored as computer instructions in memory of a computing device, such as the computing device 120 illustrated in FIG. 1. An operation 202 captures a depth image data stream, such as the depth image data stream 142, from an imaging device. Such depth image data stream may be in the form of depth of frames 224 of a patient 222, as captured by a camera system. Alternatively, the depth image data stream may include depth of frames in the form or RGB signals, infra-red (IR) signals, etc.
[0046] Subsequently, an operation 204 processes the depth image data stream to generate a breathing signal waveform such as the breathing signal waveform 232. The breathing signal waveform 232 may be in the form of a flow signal representing the flow of breathing by thepatient 102 measured in terms of volume of air breathed over a period and measured in units such as ml / sec. Alternatively, the breathing signal waveform 232 may be in the form of a volume signal generated as an integral of the flow signal over a segment of time.
[0047] An operation 206 selects a segment of the breathing signal waveform 232. Specifically, the length of the selected segment may be limited by the RR that the ML model is trying to detect. For example, the lowest RR that needs to be detected may be used to determine the length of the breathing signal segment.
[0048] An operation 208 processes the breathing signal waveform 232 to generate an input feature matrix that may be input into an ML model to train the ML model. For example, such processing may include filtering the breathing signal to remove high or low frequency noise outside of the expected respiratory frequency range, removing a mean of the breathing signal to center the breathing signal around zero, scaling the breathing signal so that the maximum and minimum excursions are set to limits such as +1 and -1 respectively, transforming the breathing signal using Fast Fourier Transform (FFT), applying a time-frequency transform such as STFT to the breathing signal, down sampling the breathing signal, etc.
[0049] Subsequently, the input feature matrix generated by processing the depth image data stream may be used to train an ML model at an operation 210. For example, training the ML model may include iteratively determining values of internal weights and biases within the ML model. Example operations for training an ML model are illustrated in further detail below in FIG. 5. In one implementation, different ML models may be trained for different patients and clinical settings. For example, depth image streams and real-time RRs specific to patients in ICUs may be used to training an ICU specific ML model. Similarly, depth image streams and real-time RRs specific to patients in GCFs may be used to training a GCF specific ML model.
[0050] An operation 212 selects an ML model that may be used to generate RR for a patient. For example, if the patient is in the ICU, an ML model that is trained for ICU patients is selected to generate predicted RR. Subsequently, at an operation 214 the current breathing signal of a patient may be input into the trained ML model to output the inferred RR for the patient. For example, such an RR may be 20 breaths per minute (BPM) as illustrated by reference numeral 228. The predicted RR generated by the ML model may be illustrated by a RR waveform 226. An operation 216 may display the RR waveform 226 on a screen for a user, such as a healthcare professional.
[0051] In one implementation, optionally, the ML model may be updated on a periodic basis based on additionally collected breathing signal data and actual RR data observed over time . For example, at operation 218, the system may update the ML model, that was previouslytrained at operation 210, using such additional breathing signal data and actual RR data. Providing such periodic updates improves the performance of the ML model overtime.
[0052] FIG. 3 illustrates operations 300 for processing depth image stream for a neural network according to the system for calculating respiratory rate using Al as disclosed herein. Specifically, various of the operations 300 are used to generate an input feature matrix that can be input into an ML model to train the ML model. An operation 302 receives a depth image stream from a camera system. For example, such depth image data stream may include depth of frames as captured by the camera system. An operation 304 generates a breathing signal from the depth image stream received from the camera system.
[0053] One or more of the operations 306 - 318 processes the breathing signal to generate an input feature matrix that may be used to train an ML model. For example, an operation 306 may remove high or low frequency noise outside of the expected respiratory frequency range from the breathing signal, wherein the expected frequency range may be determined based on other data about the patient, such as the age, weight, sex, etc., of the patient. An operation 308 may center the breathing signal by subtracting the mean of the breathing signal from it so that the processed breathing signal is centered around zero. The centering of the breathing signal around zero removes biases from the breathing signal.
[0054] An operation 310 may rescale the breathing signal so that the maximum and minimum excursions ofthe breathing signal 144 are set to limits such as +xand -x, respectively. Another operation 312 may transform the breathing signal by performing an FFT transform on it. An operation 314 may, on the other hand, apply a time-frequency transform such as short- time Fourier transform (STFT) or wavelet transform to the breathing signal. In another implementation, an operation 316 may down-sample the breathing signal. Doing such down sampling requires fewer parameters in the network and makes training of the ML model and generating inference more efficient without any loss in performance.
[0055] The outputs from one or more of the processing operations 306-318 is input into an ML model to train the ML model as well as to generate inferences once the ML model is trained. In alternative implementations, two or more of the operations 306-318 may be performed on the breathing signal to generate the input feature matrix. Thus, for example, the breathing signal maybe processed to remove noise and per operation 306 and the resulting breathing signal maybe rescaled as per operation 310 and thus the breathing signal input into the ML model at operation 330 is both noise free and rescaled.
[0056] FIG. 4 illustrates operations 400 for using an ML model to predict RR based on a real-time depth image stream. Specifically, the operations 400 may be used to predict an RRfor a patient using an ML model that has been trained using breathing signals generated from depth images captured by a camera system. An operation 402 captures real-time depth image stream, wherein the stream may include a series of depth images of a patient captured by a depth sensing camera. The image stream is processed at operation 404 to generate breathing signal. The breathing signal generated at 404 may be in the form of a flow signal representing the flow of breathing by the patient measured in terms of volume of air breathed over a period such as ml / sec. Alternatively, the breathing signal may be in the form of a volume signal generated as an integral of the flow signal over a segment of time.
[0057] An operation 406 selects a segment of the breathing signal. Specifically, the length of the selected segment may be limited by the RR that the ML model is trying to detect. For example, the lowest RR that needs to be detected may be used to determine the length of the breathing signal segment. Subsequently, an operation 408 applies one or more procedures to the selected segment of the breathing signal to generate input feature matrix that may be input into an ML model. Examples of various procedures are described in further detail above in FIG. 3.
[0058] An operation 410 selects an ML model to be used to predict the RR. For example, the ML model may be selected based on the patient’s clinical settings. At operation 412 the input feature matrix generated at operation 408 is input into the ML model selected at operation 410. Subsequently, the ML model generates a predicted RR that may be displayed on a screen at operation 414.
[0059] FIG. 5 illustrates alternative operations 500 of a neural network for predicting respiratory rate according to the system for calculating respiratory rate using Al as disclosed herein. Specifically, the NN shown in FIG. 5 is a convolutional neural network (CNN). The breathing signal 502 input into the ML model may be a processed breathing signal. The ML model may include convolution blocks 504-512. Each of the convolution blocks 504-512 may include one or more convolutional layers that are configured to extract features from the breathing signal input feature matrix. Each of the convolution blocks may have a different kernel size (k) and number of output channels. Specifically, a kernel of the convolution blocks provides a matrix of weights that slides over the signal input into the ML model and performs element-wise multiplication with the part of the input it is currently on, and then sums up all the results into a single output value, thus generating a convolution.
[0060] For example, the convolution block 504 has three kernels and 8 output channels, the convolution block 506 has 25 kernels and 16 output channels, convolution block 508 has 25 kernels and 32 output channels, convolution block 510 has three kernels and 32 outputchannels, and convolution block 512 has three kernels and 64 output channels. The output from the convolution blocks 504-412 is input into a fully connected (FC) layer 514, which may generate the output respiratory rate (RR) at 516.
[0061] FIG. 6 illustrates a graph 600 indicating predicted respiratory rate from a trained CNN based model. Specifically, the graph 600 plots results of training a CNN based model to predict the respiratory rate from a real-world data set. The training data 602, validation data 604, and test results data 606 are all shown on the graph 600 by dots having different shades. Using the method described herein results in a root mean square deviation (RMSD) of the predicted RR compared to the reference RR of approximately three (3) breaths per minute for both the validation data set RR 606 and the test data set RR 604. Thus, the system disclosed herein to predict RR using an ML model generates highly accurate RR once it is trained using the method disclosed herein.
[0062] FIG. 7 shows a portable non-contact subject monitoring system 700 that includes a non-contact detector 710 and a computing device 720. In this embodiment, the non-contact detector 710 and the computing device 720 are generally fixed in relation to each other and the system 700 is readily moveable in relation to the subject to be monitored. The detector 710 and the computing device 720 are supported on a trolley or stand 702, with the detector 710 on an arm 704 that is pivotable in relation to the stand 702 as well as adjustable in height. The system 700 can be readily moved and positioned where desired.
[0063] The detector 710 includes a first camera 714 and a second camera 715, at least one of which includes an infrared (IR) camera feature. The detector 710 also includes an IR projector 716, which projects individual features (e.g., dots, crosses or Xs, lines, or a featureless pattern, or a combination thereof etc.).
[0064] The detector 710 may be wired or wireless connected to the computing device 720. The computing device 720 includes a housing 721 with a touch screen display 722, a processor (not seen), and hardware memory (not seen) for storing software and computer instructions.
[0065] FIG. 8 shows a semi-portable non-contact subject monitoring system 800 that includes a non-contact detector 810 and a computing device 820. In this embodiment, the noncontact detector 810 is in a fixed relation to the subject to be monitored and the computing device 820 is readily moveable in relation to a subject lying on a bed 830. Specifically, the bed 830 may have a headboard 832, a side rail 834, and a mattress 836.
[0066] The detector 810 is supported on an arm 801 that is attached to a bed, in this embodiment, a hospital bed, although the detector 810 and the arm 801 can be attached to a crib, a bassinette, an incubator, an isolette, or other bed-type structure. In some embodiments,the arm 801 is pivotable in relation to the bed as well as adjustable in height to provide for proper positioning of the detector 810 in relation to the subject.
[0067] The detector 810 may be wired or wireless connected to the computing device 820, which is supported on a moveable trolley or stand 802. The computing device 820 includes a housing 821 with a touch screen display 822, a processor (not seen), and hardware memory (not seen) for storing software and computer instructions.
[0068] FIG. 9 shows a non-portable non-contact subject monitoring system 900 that includes a non-contact detector 910 and a computing device (not seen in FIG. 9). In this embodiment, at least the non-contact detector 910 is generally fixed in a location, configured to have the subject to be monitored moved into the appropriate position to be monitored.
[0069] The detector 910 is supported on a stand 901 that is free standing, the stand having a base 903, a frame 905, and a gantry 907. The gantry 907 may have an adjustable height, e.g., movable vertically along the frame 905, and may be pivotable, extendible and / or retractable in relation to the frame 905. The stand 901 is shaped and sized to allow a bed or bed-type structure to be moved (e.g., rolled) under the detector 910.
[0070] FIG. 10 is a block diagram illustrating a system including a computing device 1000, a server 1025, and an image capture device 1085 (e.g., a camera, e.g., the camera system 114). In various embodiments, fewer, additional and / or different components may be used in the system.
[0071] The computing device 1000 includes a processor 1015 that is coupled to a memory 1005. The processor 1015 can store and recall data and applications in the memory 1005, including applications that process information and send commands / signals according to any of the methods disclosed herein. The processor 1015 may also display objects, applications, data, etc. on an interface / display 1010 and / or provide an audible alert via a speaker 1012. The processor 1015 may also or alternately receive inputs through the interface / display 1010. The processor 1015 is also coupled to a transceiver 1020. With this configuration, the processor 1015, and subsequently the computing device 1000, can communicate with other devices, such as the server 1025 through a connection 1070 and the image capture device 1085 through a connection 1080. For example, the computing device 1000 may send to the server 1025 information determined about a subject from images captured by the image capture device 1085, such as depth information of a subject or object in an image.
[0072] The server 1025 also includes a processor 1035 that is coupled to a memory 1030 and to a transceiver 1040. The processor 1035 can store and recall data and applications in the memory 1030. With this configuration, the processor 1035, and subsequently the server 1025,can communicate with other devices, such as the computing device 1000 through the connection 1070.
[0073] The computing device 1000 may be, e.g., the computing device 120 of FIG. 1. Accordingly, the computing device 1000 may be located remotely from the image capture device 1085, or it may be local and close to the image capture device 1085 (e.g., in the same room). The processor 1015 of the computing device 1000 may perform any or all of the various steps disclosed herein. In other embodiments, the steps may be performed on a processor 1035 of the server 1025. In some embodiments, the various steps and methods disclosed herein may be performed by both of the processors 1015 and 1035. In some embodiments, certain steps may be performed by the processor 1015 while others are performed by the processor 1035. In some embodiments, information determined by the processor 1015 may be sent to the server 1025 for storage and / or further processing.
[0074] The devices shown in the illustrative embodiment may be utilized in various ways. For example, either or both of the connections 1070, 1080 may be varied. For example, either or both the connections 1070, 1080 may be a hard-wired connection. A hard-wired connection may involve connecting the devices through a USB (universal serial bus) port, serial port, parallel port, or other type of wired connection to facilitate the transfer of data and information between a processor of a device and a second processor of a second device. In another example, one or both of the connections 1070, 1080 may be a dock where one device may plug into another device. As another example, one or both of the connections 1070, 1080 may be a wireless connection. These connections may be any sort of wireless connection, including, but not limited to, Bluetooth connectivity, Wi-Fi connectivity, infrared, visible light, radio frequency (RF) signals, or other wireless protocols / methods. For example, other possible modes of wireless communication may include near-field communications, such as passive radio-frequency identification (RFID) and active RFID technologies. RFID and similar near- field communications may allow the various devices to communicate in short range when they are placed proximate to one another. In yet another example, the various devices may connect through an internet (or other network) connection. That is, one or both of the connections 1070, 1080 may represent several different computing devices and network components that allow the various devices to communicate through the internet, either through a hard-wired or wireless connection. One or both of the connections 1070, 1080 may also be a combination of several modes of connection.
[0075] The configuration of the devices in FIG. 10 is merely one physical system on which the disclosed embodiments may be executed. Other configurations of the devices shown mayexist to practice the disclosed embodiments. Further, configurations of additional or fewer devices than the ones shown in FIG. 10 may exist to practice the disclosed embodiments. Additionally, the devices shown in FIG. 10 may be combined to allow for fewer devices than shown or separated such that more than the three devices exist in a system. It will be appreciated that many various combinations of computing devices may execute the methods and systems disclosed herein. Examples of such computing devices may include other types of infrared cameras / detectors, night vision cameras / detectors, other types of cameras, radio frequency transmitters / receivers, smart phones, personal computers, servers, laptop computers, tablets, RFID enabled devices, or any combinations of such devices.
[0076] In contrast to tangible computer-readable storage media, intangible computer- readable communication signals may embody computer readable instructions, data structures, program modules or other data resident in a modulated data signal, such as a carrier wave or other signal transport mechanism. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, intangible communication signals include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared and other wireless media.
[0077] The implementations described herein are implemented as logical steps in one or more computer systems. The logical operations may be implemented (1) as a sequence of processor-implemented steps executing in one or more computer systems and (2) as interconnected machine or circuit modules within one or more computer systems. The implementation is a matter of choice, dependent on the performance requirements of the computer system being utilized. Accordingly, the logical operations making up the implementations described herein are referred to variously as operations, steps, objects, or modules. Furthermore, it should be understood that logical operations may be performed in any order, unless explicitly claimed otherwise or a specific order is inherently necessitated by the claim language.
[0078] The above specification, examples, and data provide a complete description of the structure and use of exemplary embodiments of the invention. Since many implementations of the invention can be made without departing from the spirit and scope of the invention, the invention resides in the claims hereinafter appended. Furthermore, structural features of the different embodiments may be combined in yet another implementation without departing from the recited claims.
[0079] The following examples are illustrative of the techniques described herein.
[0080] Example 1. A method, comprising: determining, based on an image signal received from a camera focused on at least a portion of a patient, a breathing signal; selecting a segment of the breathing signal for predicting a respiratory rate (RR) for the patient; processing the selected segment of the breathing signal to generate an input feature matrix; inputting the input feature matrix to the machine learning (ML) model; and generating a predicted RR for the patient.
[0081] Example 2. The method of Example 1, wherein selecting the segment of the breathing signal further comprising selecting length of the segment based on the respiratory rate that needs to be detected.
[0082] Example 3. The method of Example 1, wherein the breathing signal is a flow signal determined from the image signal received from a camera.
[0083] Example 4. The method of Example 3, wherein the breathing signal is a volume signal determined from the flow signal.
[0084] Example 5. The method of Example 4, wherein selecting a segment of the breathing signal further comprising: combining the segment of the breathing signal where the breathing signal is a volume signal with the segment of the breathing signal where the breathing signal is a flow signal.
[0085] Example 6. The method of Example 1, wherein processing the selected segment of the breathing signal further comprising filtering the selected segment of the breathing signal to remove frequency noise outside of an expected respiratory rate frequency range.
[0086] Example 7. The method of Example 1, wherein processing the selected segment of the breathing signal further comprising subtracting a mean value of the breathing signal from the selected segment of the breathing signal.
[0087] Example 8. The method of Example 1, wherein processing the selected segment of the breathing signal further comprising performing a Fast Fourier Transform (FFT) on the selected segment of the breathing signal.
[0088] Example 9. The method of Example 1, wherein processing the selected segment of the breathing signal further comprising scaling the selected segment of the breathing signal between limits of -X to +X, X being a real number.
[0089] Example 10 A system comprising: memory; one or more processor units; a respiratory rate determining system stored in the memory and executable by the one or more processor units, the respiratory rate determining system encoding computer-executable instructions on the memory for executing on the one or more processor units a computer process, the computer process comprising: determining, based on an image signal receivedfrom a camera focused on at least a portion of a patient, a breathing signal; selecting a training segment of the breathing signal for training a machine learning (ML) model; processing the selected segment of the breathing signal to generate an input feature matrix; inputting the input feature matrix to the machine learning (ML) model; inputting observed respiratory rate (RR) of the patient in the ML model; training the ML model, using the input feature matrix and the observed respiratory rate (RR) of the patient; and inputting a real-time breathing signal into the trained ML model to generate an inferred RR of the patient.
[0090] Example 11. The system of Example 10, wherein selecting the training segment of the breathing signal further comprising selecting length of the segment based on a the respiratory rate that needs to be detected.
[0091] Example 12. The system of Example 10, wherein the breathing signal is at least one of a flow signal determined from the image signal received from a camera and a volume signal determined from the flow signal.
[0092] Example 13 The system of Example 12, wherein selecting the training segment of the breathing signal further comprising: combining the selected segment of the breathing signal where the breathing signal is a volume signal with the selected segment of the breathing signal where the breathing signal is a flow signal.
[0093] Example 14 The system of Example 10, wherein processing the selected segment of the breathing signal further comprising filtering the selected segment of the breathing signal to remove frequency noise outside of an expected respiratory rate frequency range.
[0094] Example 15 The system of Example 10, wherein processing the selected segment of the breathing signal further comprising subtracting a mean value of the breathing signal from the selected segment of the breathing signal.
[0095] Example 16. The system of Example 10, wherein processing the selected segment of the breathing signal further comprising performing a Fast Fourier Transform (FFT) on the selected segment of the breathing signal.
[0096] Example 17. A physical article of manufacture including one or more tangible computer-readable storage media encoding computer-executable instructions for executing on a computer system a computer process to determine respiratory rate of a patient, the computer process comprising: determining, based on an image signal received from a camera focused on at least a portion of a patient, a breathing signal; selecting a training segment of the breathing signal for training a machine learning (ML) model; processing the training segment of the breathing signal to generate an input feature matrix; inputting the input feature matrix to themachine learning (ML) model; and training the ML model, using the input feature matrix and an observed respiratory rate (RR) of the patient.
[0097] Example 18. The physical article of manufacture of Example 17, wherein the computer process comprising inputting a real-time breathing signal into the trained ML model to generate an inferred RR of the patient.
[0098] Example 19. The physical article of manufacture of Example 17, wherein selecting the segment of the breathing signal further comprising selecting length of the segment based on the respiratory rate that needs to be detected.
[0099] Example 20. The physical article of manufacture of Example 16, wherein the breathing signal is at least one of a flow signal determined from the image signal received from a camera and a volume signal determined from the flow signal.
Claims
ClaimsWHAT IS CLAIMED IS:
1. A method, comprising:Determining (204), based on an image signal received from a camera focused on at least a portion of a patient, a breathing signal; selecting (206) a segment of the breathing signal for predicting a respiratory rate (RR) for the patient; processing (208) the selected segment of the breathing signal to generate an input feature matrix; inputting (412) the input feature matrix to the machine learning (ML) model; and generating (214) a predicted RR for the patient.
2. The method of claim 1, wherein selecting the segment of the breathing signal further comprising selecting length of the segment based on the respiratory rate that needs to be detected.
3. The method of claims 1 and 2, wherein the breathing signal is a flow signal determined from the image signal received from a camera.
4. The method of claim 3, wherein the breathing signal is a volume signal determined from the flow signal.
5. The method of claim 4, wherein selecting a segment of the breathing signal further comprising: combining the segment of the breathing signal where the breathing signal is a volume signal with the segment of the breathing signal where the breathing signal is a flow signal.
6. The method of claim 1, wherein processing the selected segment of the breathing signal further comprising filtering the selected segment of the breathing signal to remove frequency noise outside of an expected respiratory rate frequency range.
7. The method of claim 1, wherein processing the selected segment of the breathing signal further comprising subtracting a mean value of the breathing signal from the selected segment of the breathing signal.
8. The method of claim 1, wherein processing the selected segment of the breathing signal further comprising performing a Fast Fourier Transform (FFT) on the selected segment of the breathing signal.
9. The method of claim 1, wherein processing the selected segment of the breathing signal further comprising scaling the selected segment of the breathing signal between limits of -X to +X, X being a real number.
10. A system comprising: memory; one or more processor units; a respiratory rate determining system (126) stored in the memory and executable by the one or more processor units, the respiratory rate determining system encoding computerexecutable instructions on the memory for executing on the one or more processor units a computer process, the computer process comprising:Determining (204), based on an image signal received from a camera focused on at least a portion of a patient, a breathing signal; selecting (206) a training segment of the breathing signal for training a machine learning (ML) model; processing (208) the selected segment of the breathing signal to generate an input feature matrix; inputting (412) the input feature matrix to the machine learning (ML) model; inputting observed respiratory rate (RR) of the patient in the ML model; training (210) the ML model, using the input feature matrix and the observed respiratory rate (RR) of the patient; and inputting (214) a real-time breathing signal into the trained ML model to generate an inferred RR of the patient.
11. The system of claim 10, wherein selecting the training segment of the breathing signal further comprising selecting length of the segment based on a the respiratory rate that needs to be detected.
12. The system of claim 10, wherein the breathing signal is at least one of a flow signal determined from the image signal received from a camera and a volume signal determined from the flow signal.
13. The system of claim 12, wherein selecting the training segment of the breathing signal further comprising: combining the selected segment of the breathing signal where the breathing signal is a volume signal with the selected segment of the breathing signal where the breathing signal is a flow signal.
14. A physical article of manufacture including one or more tangible computer- readable storage media encoding computer-executable instructions for executing on a computer system a computer process to determine respiratory rate of a patient, the computer process comprising:Determining (204), based on an image signal received from a camera focused on at least a portion of a patient, a breathing signal; selecting (206) a training segment of the breathing signal for training a machine learning (ML) model;Processing (208) the training segment of the breathing signal to generate an input feature matrix; inputting (210) the input feature matrix to the machine learning (ML) model; andTraining (210) the ML model, using the input feature matrix and an observed respiratory rate (RR) of the patient.
15. The physical article of manufacture of claim 14, wherein the computer process comprising inputting a real-time breathing signal into the trained ML model to generate an inferred RR of the patient.
Citation Information
Patent Citations
Non-Contact Breathing Activity Monitoring And Analyzing System Through Thermal On Projection Medium Imaging
US20200138337A1
Cardiopulmonary health monitoring using thermal camera and audio sensor
US20220133156A1
Non-contact monitoring for night tremors or other medical conditions
US20230310865A1