Artificial intelligence assisted monitoring of patients
The AI-assisted monitoring system addresses the challenge of tracheostomy tube complications by using a YOLO-NAS model to detect tube dislodgement, ensuring continuous, remote supervision and reducing hospitalization risks.
Patent Information
- Application Number
- PCT/US2025/024550
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-04-14
- Filing Date
- 2025-04-14
- Publication Date
- 2025-10-16
AI Technical Summary
There is a critical need for home-based monitoring systems for children with tracheostomies due to the high risk of complications from accidental dislodgement or obstruction of the tracheostomy tube, exacerbated by the shortage of home care nurses, leading to prolonged hospitalization or placement in long-term care facilities.
A method using artificial intelligence, specifically a YOLO-NAS model trained on video data to detect tracheostomy tube positions, providing real-time alerts for dislodgement or obstruction, integrated with a remote monitoring system.
Enables continuous, remote monitoring of tracheostomy tube status, reducing complications and potentially avoiding hospitalization by providing timely alerts and reducing the need for constant supervision.
Smart Images

Figure US2025024550_16102025_PF_FP_ABST
Abstract
Description
[0001] UNITED STATES PATENT APPLICATION FOR:
[0002] ARTIFICIAL INTELLIGENCE ASSISTED MONITORING OF PATIENTS
[0003] Inventors: Michael Dunham
[0004] RELATED APPLICATIONS
[0005] This application claims priority to the United States Provisional Application No. 63 / 633,635, filed April 12, 2024, and to the United States Provisional Application No. 63 / 788,264, filed April 14, 2025, the entire contents of each of which are incorporated herein by reference.
[0006] TECHNICAL FIELD
[0007] The disclosure herein involves the use of artificial intelligence to monitor use of medical device technologies.
[0008] BACKGROUND
[0009] Tracheostomy is a surgical procedure that involves creating an opening in the front of the neck into the trachea (windpipe) to provide an alternative airway passage. It is frequently performed in children to bypass an obstructed airway due to blockage in the upper throat or for chronic lung conditions related to prematurity where the child requires a ventilator (breathing machine) to breathe. Children undergoing tracheostomy require ongoing care and support to manage their airway and maintain optimal respiratory function. Thanks to advances in critical care, many of these patients now survive hospitalization and can continue living at home with their tracheostomy.
[0010] During the past decade, there has been a 300% increase in the number of children with a tracheostomy who require home health care, including supervised monitoring 24 hours a day, 365 days per year. Over 25% of children discharged home after a tracheostomy experience serious complications, including death or severe anoxic brain injury due to airway obstruction. Many of these complications are linked to accidental dislodgement or obstruction of the tracheostomy tube, mainly when the child is unattended. The nationwide shortage of home care nurses makes continuous observation of these children impractical. The lack of care and observation is especially problematic in disadvantaged and rural homes. As a result, current care strategies often require prolonged hospitalization or placement in long-term care facilities. There is a critical need for home-based monitoring systems for children living with a tracheostomy.
[0011] INCORPORATION BY REFERENCE
[0012] Each patent, patent application, and / or publication mentioned in this specification is herein incorporated by reference in its entirety to the same extent as if each individual patent, patent application, and / or publication was specifically and individually indicated to be incorporated by reference.
[0013] SUMMARY OF THE INVENTION
[0014] A method is described herein comprising receiving first video data of a first plurality of subjects, wherein the first video data tracks a tracheostomy tube position in the first plurality of subjects, wherein the tracheostomy tube position comprises either a first tracheostomy tube position or a second tracheostomy tube position, training a predictive model using information of the first video data, wherein the information of the first video data comprises annotated frames, wherein the annotated frames indicate a first or second tracheostomy tube position, wherein the trained predictive model distinguishes between the tracheostomy tube position states, and applying the trained predictive model to second video data of a subject, wherein the second video data tracks a tracheostomy tube position, wherein the trained predictive model detects a transition from the first tracheostomy tube position to the second tracheostomy tube position.
[0015] In embodiments, the training comprises constructing bounding boxes around tracheostomy tube positions in frames of the first video data.
[0016] In embodiments, the annotated frames comprise annotated bounding boxes.
[0017] In embodiments, the constructing includes implementing a software tool to automatically identify the bounding boxes.
[0018] In embodiments, the software tool includes at least one of Roboflow. Label Studio, VGG Image Annotator, and COCO Annotator.
[0019] In embodiments, the first video data is derived from recordings taken during tracheostomy tube changes. In embodiments, the plurality of subjects comprises a mannequin with an attached tracheostomy tube.
[0020] In embodiments, the plurality of subjects comprises a human subject using a tracheostomy tube.
[0021] In embodiments, the recordings include additional computer simulation to deidentify the at least one subject.
[0022] In embodiments, the recordings comprise computer generated imagery
[0023] In embodiments, the recordings comprise video captured under a plurality of recording conditions.
[0024] In embodiments, the plurality of conditions comprises variable lighting.
[0025] In embodiments, the plurality of conditions comprises variable camera angles.
[0026] In embodiments, the trained predictive model comprises a You Only Look Once Neural Architecture Search.
[0027] In embodiments, the method comprises providing an alert to a remote monitoring system upon detecting a transition from the first tracheostomy tube position to the second tracheostomy tube position.
[0028] In embodiments, the first tracheostomy tube position comprises an in-place condition.
[0029] In embodiments, the second tracheostomy tube position comprises a dislodged condition.
[0030] BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 shows a pediatric tracheostomy mannequin, under an embodiment.
[0032] Figure 2 shows a pediatric tracheostomy mannequin, under an embodiment.
[0033] Figure 3 shows training parameters for a YOLO-NAS model, under an embodiment.
[0034] Figure 4 describes a machine learning method for monitoring use of a tracheostomy tube in real time, under an embodiment.
[0035] Figure 5 shows a wearable sensing systems for infants, under an embodiment.
[0036] Figures 6A-6F shows code information that demonstrates a model for processing wearable sensor suit data, under an embodiment.
[0037] DETAILED DESCRIPTION We are addressing tracheostomy-related complications with "Eyes-On," a technology designed to enhance remote visual monitoring and telemedicine through artificial intelligence. This study evaluates the potential of home critical care monitoring and observation using machine learning, an approach not previously tested. Eyes-On employs deep learning algorithms tailored explicitly for pediatric patients with tracheostomies receiving home care.
[0038] The Eyes-On system works under diverse lighting conditions and backgrounds, from different angles, for a variety of positions and camera angles. The system tracks the tracheostomy tube position as well as any movements of the patient in real-time. It is trained on video recordings of children with a tracheostomy before, undergoing, and after tracheostomy tube changes. Under one embodiment, these videos may be video recordings of mannequins. Under another embodiment, videos may be of human subjects using computer software to deidentify any attributes of the subjects prior to use of any images in model training. Under yet another embodiment, videos used in training the model may be entirely computer simulated. Note that embodiments are described below with respect to children. However, the same framework may apply to subjects of any age.
[0039] The framework will provide real-time alerts to abnormal tube positions. An example of Eyes-On monitoring simulated with a pediatric tracheostomy mannequin is shown in Figures 1 and 2.
[0040] Systems and methods described herein include designing an Al-enabled computer vision system to monitor infants with a tracheostomy for dislodgment or obstruction of the tracheostomy tube. The project requires an adequate, labeled dataset of images and a suitable machine-learning model capable of classifying live video feeds of patients in real time.
[0041] The first step involves gathering a dataset of images, including video frames, that show infants with tracheostomy tubes in a variety of positions, lighting conditions, and camera angles. We will require images showing a tracheostomy tube that is in place and functioning normally, as well as cases of dislodgment or obstruction. As part of an IRB-approved protocol, LSU faculty in the Department of Otolaryngology acquires training data for the model using video recordings of children undergoing tracheostomy tube changes, a routine procedure required for tracheostomy care. Every sample frame taken from the video dataset is labeled with the appropriate classification (e.g., normal, dislodged, obstructed). A team of trained medical professionals familiar with the issues associated with tracheostomy label the images. For the video stream annotation model, deep learning model architectures adept at localizing, segmenting, and classifying video streams are used to detect tracheostomy tube dislodgement in real-time. Under an embodiment, a trained neural network monitors a live video stream and detects tracheostomy tube dislodgement.
[0042] The development and deployment of Eyes-On involves a data preparation (i.e., preparation of a data training set), model fine-tuning, evaluation, deployment, and ongoing monitoring. For "Eyes-On," the training dataset is derived from video frames capturing patients with tracheostomies. Training frames are manually or automatically sourced from footage of scheduled tracheostomy tube changes. Frames are selected to ensure a range of different tube types, patient demographics, lighting conditions, and camera angles.
[0043] Images for the training dataset depict scenarios where a tracheostomy tube is in its usual position and situations where it has become dislodged. This includes images of mannequins equipped with tracheostomy tubes and images of children undergoing tracheostomy tube changes, which are a standard part of their care. During a tracheostomy tube change, the caretaker removes and replaces the existing tube with a new one. Under an embodiment, images and / or videos of tracheostomy tube changes may be virtually created.
[0044] Under an embodiment, each frame in the dataset is labeled by human experts, with bounding boxes drawn around the tracheostomy tubes. These annotations define what the model needs to identify and classify. The labels locate the tracheostomy tube in the images and categorize them as either 'in-place' or 'dislodged,' providing the model with the context needed to understand the computational task.
[0045] Bounding boxes are created on images using annotation software. During the annotation process with bounding boxes, software captures the coordinates of the bounding box corners relative to the image's pixel locations. There are several image annotation tools available. We used one called Roboflow. Other tools include Label Studio, VGG Image Annotator, and COCO Annotator. Some programs offer semi-automated labeling features. For instance, in videos, adjacent frames are likely to have very similar bounding boxes for any identified areas of interest, and the software may utilize this similarity for tracking. Under an embodiment, the image annotation tools approximate bounding boxes which must then be adjusted by a user.
[0046] Several pre-trained models may be used to train Eyes-On, allowing it to learn the specific features of in-place and dislodged tracheostomy tubes. The current approach leverages the YOLO-NAS (You Only Look Once - Neural Architecture Search) model fine-tuned on the specialized tracheostomy tube object detection and classification dataset described above.
[0047] The architecture of the neural network, determined by the number and types of layers and the number of neurons in each layer, directly influences the model's ability to learn complex patterns. Further, the choice of the optimizer (such as SGD or Adam), loss function, and the method used for weight initialization are additional hyperparameters that significantly impact the training process and its outcomes. Under an embodiment, the weights are initialized with random values.
[0048] A loss function and an optimizer are two essential components that help to improve the performance of a model. A loss function measures the difference between the predicted output of a model and the actual output, while an optimizer adjusts the model's parameters to minimize the loss function.
[0049] A loss function, also known as a cost function, is used to measure the accuracy of a model’s predictions. It calculates the difference between the predicted output and the actual output for each training sample.
[0050] Once the loss function is defined, an optimizer is used to adjust the model’s parameters to minimize the loss function. It's also worth mentioning that these optimizers can be fine-tuned with different settings or Hyperparameter such as learning rate, momentum, decay rate etc.
[0051] Gradient descent is one of the most widely used optimizers. It adjusts the model’s parameters by taking the derivative of the loss function with respect to the parameters and updating the parameters in the direction of the negative gradient. Gradient descent is simple to implement, but it can be slow to converge when the loss function has many local minima.
[0052] SGD is an extension of gradient descent. It updates the model’s parameters after each training sample, rather than after each epoch. This makes it faster to converge, but it can also make the optimization process more unstable. Stochastic gradient descent is often used for problems with a large amount of data.
[0053] The goal of the model is to minimize the loss function. By minimizing the loss function, we are effectively trying to find the best set of parameters that will produce the most accurate predictions.
[0054] Optimization algorithms are the bread and butter of deep learning and are used for training any neural network and modifying the model parameters to minimize loss. But one algorithm stands out for training neural networks — Adam optimization. Adam, which stands for Adaptive Moment Estimation, is particularly well-suited for training deep neural networks because it computes individual adaptive learning rates for different parameters.
[0055] Adam is an adaptive learning rate algorithm designed to improve training speeds in deep neural networks and reach convergence quickly. Standard gradient descent lays the foundation for Adam, which is essentially an adaptive extension of the same algorithm. Standard gradient descent is represented by the following equation:
[0056] 0 = 0 -y* gt
[0057] Here, 0 = Model parameters, a = Learning rate, and gt= Gradient of the cost function with respect to the parameters.
[0058] This update changes the parameters 0 in the negative direction of the gradient to minimize the cost function. The learning rate a determines the size of the step.
[0059] In the standard gradient descent algorithm, the learning rate a is fixed, meaning we need to start at a high learning rate and manually change the alpha by steps or by some learning schedule. A lower learning rate at the onset would lead to very slow convergence, while a very high rate at the start might miss the minima.
[0060] Adam solves this problem by adapting the learning rate a for each parameter 0, enabling faster convergence compared to standard gradient descent with a constant global learning rate. It customizes each parameter’s learning rate based on its gradient history, and this adjustment helps the neural network learn efficiently as a whole.
[0061] Adam leverages the concepts of Momentum and Root Mean Square Propagation. Momentum speeds up training by accelerating gradients in the right directions by adding a fraction of the previous gradient to the current one. For example, let’s say a gradient has been consistently pointing in the same direction. The momentum term proportional to the previous gradients will accumulate and accelerate the optimization in that direction.
[0062] Model Selection and Fine-Tuning
[0063] As indicated above, several pre-trained models have been used to train Eyes-On, allowing it to learn the specific features of in-place and dislodged tracheostomy tubes. We currently leverage the YOLO-NAS (You Only Look Once - Neural Architecture Search) model fine-tuned on the specialized tracheostomy tube object detection and classification dataset described above. YOLO-NAS is known for its efficiency and accuracy in object detection tasks. Tn preliminary studies of manikins equipped with tracheostomy tubes, we achieved greater than 90% object recognition with frame inference latencies of less than 100 milliseconds. The latency must be minimal for medical monitoring and similar critical applications to ensure timely intervention and decision-making. For deployment in a remote monitoring system, latency needs are evaluated in the context of video capture delay, data transmission time, and other processing times.
[0064] During training, the model iteratively processes the dataset, learning to recognize the tracheostomy tube's appearances and positions. It adjusts its parameters to minimize errors, guided by a loss function quantifying the difference between its predictions and the actual labels. The following training parameters are optimized for accuracy and efficiency.
[0065] Input Resolution
[0066] Higher resolutions can improve detection accuracy for small objects (like tracheostomy tubes) but at the cost of computational resources and slower processing times. The current model is optimized for inference at 1080p (1920 x 1080) image resolution.
[0067] The input and output image resolutions are defined in the code. However, the model typically down-samples the images to lower resolutions during training, usually maintaining a 1 : 1 aspect ratio. For this model, we used 256x256 resolution.
[0068] Model Architecture Details
[0069] By adjusting the depth (number of layers) and width (number of units in each neural network layer), we have balanced the inference latency and detection accuracy to maximize model effectiveness. An advantage of using the YOLO-NAS base model is its neural network architecture search algorithm to find an optimal architecture and define the search space (e.g., types of layers, connections) and search strategy (e.g., reinforcement learning, evolutionary algorithms).
[0070] Initially set by the developer of the original model, the network architecture parameters can be modified during fine-tuning. YOLO-NAS employs a proprietary algorithm for neural network architecture search (NAS) to determine the layers' width and depth, connection patterns, and types of layers. Additionally, there are open-source NAS implementations available that we might consider as we further refine the model. The method described herein uses Yolo-NAS large. However, embodiments may use Yolo-NAS small or medium. Detection Bounding Boxes
[0071] The size and aspect ratios of bounding boxes, which the model uses as references for predicting object locations, must be carefully outlined on the training images. Carefully crafting these better to match tracheostomy tubes' typical size and shape improves detection accuracy.
[0072] Loss Function
[0073] The loss function includes terms for bounding box overlap (intersection-over-union), and classification accuracy. Tuning the hyperparameters related to these components balances the trade-off between detecting tubes accurately and classifying their status correctly.
[0074] The loss function, defined within the training hyperparameters, utilizes PPYoloELoss — a loss function determined by the base model's developer. This loss function relies on the intersection-over-union metric for bounding box annotations, which measures the overlap between the predicted bounding box and the actual label. The model also implements the Adam optimizer.
[0075] Learning Rate and Schedule
[0076] The initial learning rate and adjustment schedule are determined experimentally to avoid overshooting minima and optimize convergence during training.
[0077] The initial learning rate and momentum are detailed in the training hyperparameters. In this instance, we employ cosine decay, incorporating a cosine term into its formula. The initial learning rate is le-5.
[0078] Regularization Techniques
[0079] Dropout, L2 regularization, and batch normalization are experimentally optimized to avoid overfitting and improve generalization.
[0080] Dropout, regularization, and batch normalization are integrated into the base model's design through coding. Regularization, dropout, and batch normalization are techniques employed to enhance neural network performance and mitigate overfitting. Regularization involves penalizing outlier weights by reducing their magnitude, thereby encouraging a simpler model that generalizes better to unseen data. Dropout works by randomly removing some neurons from the network during training, reducing the dependency of neurons on the presence of others and promoting a more robust network that can generalize better. Batch normalization, on the other hand, involves scaling the inputs at each neural network layer to have a standard distribution, typically with a mean of 0 and a variance of 1 . This standardization helps in stabilizing the learning process and accelerates the convergence of the training phase.
[0081] Evaluation and Iteration
[0082] Post-training, the model's performance is evaluated using a separate set of labeled frames not seen during training. This evaluation assesses the model's accuracy in detecting and classifying the tracheostomy tubes. Metrics such as precision, recall, and the Fl score provide insights into its effectiveness.
[0083] For a computer vision model that performs localization of an object in an image, the following metrics for model performance are relevant:
[0084] Intersection over Union (loU): loU measures the overlap between the predicted bounding box and the ground truth box (human labeled). It is calculated as the area of overlap between the two boxes divided by the area of their union. A higher loU indicates a better match between the predicted and actual locations of objects. loU is often used as a threshold to decide whether a prediction is considered a true positive (TP), false positive (FP), or false negative (FN). Typically, an loU of 0.5 or better is considered an accurate prediction.
[0085] Precision: Precision is the ratio of correctly predicted positive observations (true positives) to the total predicted positives (true positives plus false positives). It reflects the model's accuracy in predicting positive (object-present) instances.
[0086] Recall (Sensitivity): Recall is the ratio of correctly predicted positive observations to all observations in actual class (true positives divided by the sum of true positives and false negatives). It measures the model's ability to detect all relevant instances.
[0087] Fl Score: The Fl score is the harmonic mean of precision and recall, providing a single metric to assess the balance between precision and recall. An Fl score reaches its best value at 1 (perfect precision and recall) and worst at 0.
[0088] Average Precision (AP): For each class, AP is calculated by plotting precision vs. recall at different loU thresholds and computing the area under the curve. This metric accounts for the model's precision and recall, providing a comprehensive view of its performance across different levels of confidence.
[0089] Mean Average Precision (mAP): mAP is the mean of the AP scores across all classes. In contexts where multiple object classes are detected, mAP provides a single performance metric that encapsulates the model's effectiveness across all classes. Depending on model performance, adjustments are made, including dataset expansion, model architecture alterations, and training parameters. This iterative refinement continues until the model meets the desired performance benchmarks.
[0090] Deployment
[0091] Once Eyes-On achieves satisfactory accuracy, it is deployed in a clinical setting. Deployment requires integrating the model with a telemonitoring system, allowing it to analyze live streams of patients with tracheostomies. The model continuously scans the video feed, applying its learned capabilities to detect and alert medical personnel of dislodgements.
[0092] Under an embodiment, the monitoring system is integrated into a telemonitoring system, analyzing live video streams from a camera attached to the infant's crib in real time. Should a tracheostomy tube become dislodged, it triggers an alarm locally or remotely. This approach is like that of many baby monitoring systems that are equipped with video cameras. High- resolution video (1080p or higher) will be utilized for the tracheostomy tracking system.
[0093] The deployment phase implements an infrastructure for real-time analysis and communication, ensuring the model's predictions are promptly communicated to the care team. Additionally, monitoring the model's performance is essential to catch any drift in its effectiveness due to changes in patient demographics, tube designs, or other variables.
[0094] Ongoing Improvement
[0095] Post-deployment, Eyes-On is subject to continuous improvement, with user feedback allowing for periodic retraining to incorporate new data, adapt to changing conditions, and incorporate technological advancements. This ensures that the model remains accurate and effective over time.
[0096] Other Applications
[0097] Eyes-On's core technology, which focuses on real-time object detection and classification within video streams, has additional applications in healthcare telemonitoring. In-home critical care settings where patients rely on mechanical ventilation, Eyes-On can be specifically trained to identify ventilator disconnections and malfunctions. It enables monitoring of patients for early signs of distress — such as changes in posture, excessive movements, and other indicators — without the need for constant physical supervision. Moreover, it supports monitoring seizure activity in susceptible individuals and detects falls among elderly or disabled patients. This technology follows a versatile development cycle that includes data collection, model design, training, and deployment, making it adaptable to various applications.
[0098] Figure 3 shows hyperparameters for a YOLO-NAS model, under an embodiment.
[0099] Under an embodiment, these values are optimized using a grid search, which involves training models across a range of hyperparameters.
[0100] Figure 4 describes a machine learning method for monitoring use of a tracheostomy tube in real time, under an embodiment.
[0101] YOLO-NAS models typically process data in a single pass using a convolutional neural network backbone. Under an embodiment, a YOLO-NAS model can adopt grid-based predictions, where the model divides the image into a grid and predicts bounding boxes (with associated confidence scores) for objects in each cell, followed by non-maximum suppression or other post-processing to refine final detections.
[0102] Transformer-based models are also available for fine-tuning. A notable example is DETR (Detection Transformer), which replaces traditional convolutional “heads” for object detection with a transformer encoder-decoder architecture. This approach enables the model to learn relationships between different regions of an image without relying on hand-crafted anchor boxes.
[0103] Vital Vision Wearable Sensing System for Infants
[0104] Vital Vision Pajamas provides continuous vital sign and oxygen saturation measures that can seamlessly integrate with the Eyes-On visual system (described above), to provide a complete Al enabled solution to high risk and technology dependent infant monitoring.
[0105] Garment
[0106] The foundation of Vital Vision Pajamas is a single-piece, footed design that resembles a traditional infant “onesie” with attached feet. The single garment design allows all sensors to remain securely in position — even when an infant kicks, squirms, or rolls around.
[0107] The garment is constructed from soft, hypoallergenic cotton to minimize the risk of skin irritation, ensuring the pajamas feel gentle against a baby’s delicate skin. A full-length zipper along the front, simplifying dressing and diaper changes and allows straightforward access for sensor adjustments.
[0108] Sensors Each foot of the pajamas houses a lightweight photoplethysmography (PPG) sensor — a small LED and photodiode arrangement that measures oxygen saturation (SpOz) and heart rate by reading the blood’s pulsations. Because these sensors are inside the footed portion of the garment, they are shielded from ambient light and have consistent skin contact — factors crucial for accurate readings.
[0109] The respiratory rate is tracked by detecting changes in impedance during chest wall expansion and contraction. The chest portion of the garment incorporates two soft electrodes placed on either side of the upper chest. These conductive patches gently contact the skin to capture the expansion and contraction cycles of each breath. They are non-abrasive, using conductive textiles that balance durability with a baby’s comfort.
[0110] Thin, flexible cables or fabric-based conductors seamlessly run through the garment’s seams and connect the sensors to a small, padded and discreet control module placed away from pressure points on the torso. This streamlined integration ensures that nothing obstructs the baby’s movements or causes discomfort.
[0111] Control Module
[0112] A small, flexible control unit — tucked into a pouch on the side of the garment — processes signals from the sensors and can send continuous readings to the Eyes-On computer system where the information is integrated with the visual monitoring system. The signal is transmitted via low-power Bluetooth to avoid the need for any wires connecting to the patient. The module is positioned to avoid pressure points when the infant is lying down.
[0113] Power management
[0114] A rechargeable battery tucked safely into the garment powers the control module. Careful insulation ensures safe power management. A magnetic inductive charging connector allows for simple recharging. When the sensors aren’t active, the module transitions into a low-power mode to extend battery life and reduce heat generation.
[0115] Figure 5 shows a wearable sensing systems for infants, under an embodiment.
[0116] Integration with the Eyes-On infant monitoring system
[0117] The system integrates with Eyes-On, a separate deep-learning computer vision application designed to detect tracheostomy tube displacement. Alerts from Eyes-On can serve as an additional signal to augment predictions derived from sensor data. The control module feeds data into the Eyes-On software, which continuously monitors the infant’s pulse, respiratory rate, and oxygen saturation. Sensor data is then buffered into a sliding time window (e.g., the past 60 seconds) and preprocessed as needed. A separate deep-learning module works alongside the Eyes-On visual stream to predict and detect abnormal respiratory events, including early indicators of tracheostomy tube displacement. The application is structured as an event-driven, continuously running service, incorporating logging, error handling, and interfaces for integration with external alert systems.
[0118] The deep-learning model is designed to learn temporal patterns from normalized sensor data. Several neural network architectures suitable for streaming time-series data have been evaluated, including recurrent neural networks (e.g., Long Short-Term Memory, or LSTM) and more recent transformer-based approaches (e.g., the Temporal Fusion Transformer, or TFT), which is well-suited for integrating heterogeneous data sources such as sensor data and computer vision streams. The model outputs a probability (via a sigmoid activation) representing the risk of an adverse respiratory event. It is trained offline using historical labeled data and then integrated into the real-time application.
[0119] Figures 6A-6F shows code information that demonstrates this approach using an LSTM architecture which integrates the Eyes-On Model architecture described above, under an embodiment.
[0120] Computer networks suitable for use with the embodiments described herein include local area networks (LAN), wide area networks (WAN), Internet, or other connection services and network variations such as the world wide web, the public internet, a private internet, a private computer network, a public network, a mobile network, a cellular network, a value-added network, and the like. Computing devices coupled or connected to the network may be any microprocessor controlled device that permits access to the network, including terminal devices, such as personal computers, workstations, servers, mini computers, main-frame computers, laptop computers, mobile computers, palm top computers, hand held computers, mobile phones, TV set-top boxes, or combinations thereof. The computer network may include one of more LANs, WANs, Internets, and computers. The computers may serve as servers, clients, or a combination thereof.
[0121] The artificial intelligence assisted monitoring of patients can be a component of a single system, multiple systems, and / or geographically separate systems. The artificial intelligence assisted monitoring of patients can also be a subcomponent or subsystem of a single system, multiple systems, and / or geographically separate systems. The components of the artificial intelligence assisted monitoring of patients can be coupled to one or more other components (not shown) of a host system or a system coupled to the host system.
[0122] One or more components of the artificial intelligence assisted monitoring of patients and / or a corresponding interface, system or application to which the artificial intelligence assisted monitoring of patients is coupled or connected includes and / or runs under and / or in association with a processing system. The processing system includes any collection of processor-based devices or computing devices operating together, or components of processing systems or devices, as is known in the art. For example, the processing system can include one or more of a portable computer, portable communication device operating in a communication network, and / or a network server. The portable computer can be any of a number and / or combination of devices selected from among personal computers, personal digital assistants, portable computing devices, and portable communication devices, but is not so limited. The processing system can include components within a larger computer system.
[0123] The processing system of an embodiment includes at least one processor and at least one memory device or subsystem. The processing system can also include or be coupled to at least one database. The term “processor” as generally used herein refers to any logic processing unit, such as one or more central processing units (CPUs), digital signal processors (DSPs), applicationspecific integrated circuits (ASIC), etc. The processor and memory can be monolithically integrated onto a single chip, distributed among a number of chips or components, and / or provided by some combination of algorithms. The methods described herein can be implemented in one or more of software algorithm(s), programs, firmware, hardware, components, circuitry, in any combination.
[0124] The components of any system that include the artificial intelligence assisted monitoring of patients can be located together or in separate locations. Communication paths couple the components and include any medium for communicating or transferring files among the components. The communication paths include wireless connections, wired connections, and hybrid wireless / wired connections. The communication paths also include couplings or connections to networks including local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), proprietary networks, interoffice or backend networks, and the Internet. Furthermore, the communication paths include removable fixed mediums like floppy disks, hard disk drives, and CD-ROM disks, as well as flash RAM, Universal Serial Bus (USB) connections, RS-232 connections, telephone lines, buses, and electronic mail messages.
[0125] Aspects of the artificial intelligence assisted monitoring of patients and corresponding systems and methods described herein may be implemented as functionality programmed into any of a variety of circuitry, including programmable logic devices (PLDs), such as field programmable gate arrays (FPGAs), programmable array logic (PAL) devices, electrically programmable logic and memory devices and standard cell-based devices, as well as application specific integrated circuits (ASICs). Some other possibilities for implementing aspects of the artificial intelligence assisted monitoring of patients and corresponding systems and methods include: microcontrollers with memory (such as electronically erasable programmable read only memory (EEPROM)), embedded microprocessors, firmware, software, etc. Furthermore, aspects of the artificial intelligence assisted monitoring of patients and corresponding systems and methods may be embodied in microprocessors having software-based circuit emulation, discrete logic (sequential and combinatorial), custom devices, fuzzy (neural) logic, quantum devices, and hybrids of any of the above device types. Of course the underlying device technologies may be provided in a variety of component types, e.g., metal-oxide semiconductor field-effect transistor (MOSFET) technologies like complementary metal-oxide semiconductor (CMOS), bipolar technologies like emitter-coupled logic (ECL), polymer technologies (e.g., silicon-conjugated polymer and metal-conjugated polymer-metal structures), mixed analog and digital, etc.
[0126] It should be noted that any system, method, and / or other components disclosed herein may be described using computer aided design tools and expressed (or represented), as data and / or instructions embodied in various computer-readable media, in terms of their behavioral, register transfer, logic component, transistor, layout geometries, and / or other characteristics. Computer- readable media in which such formatted data and / or instructions may be embodied include, but are not limited to, non-volatile storage media in various forms (e.g., optical, magnetic or semiconductor storage media) and carrier waves that may be used to transfer such formatted data and / or instructions through wireless, optical, or wired signaling media or any combination thereof. Examples of transfers of such formatted data and / or instructions by carrier waves include, but are not limited to, transfers (uploads, downloads, e-mail, etc.) over the Internet and / or other computer networks via one or more data transfer protocols (e.g., HTTP, FTP, SMTP, etc.). When received within a computer system via one or more computer-readable media, such data and / or instruction- based expressions of the above described components may be processed by a processing entity (e g., one or more processors) within the computer system in conjunction with execution of one or more other computer programs.
[0127] Unless the context clearly requires otherwise, throughout the description and the claims, the words “comprise,” “comprising,” and the like are to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense; that is to say, in a sense of “including, but not limited to.” Words using the singular or plural number also include the plural or singular number respectively. Additionally, the words “herein,” “hereunder,” “above,” “below,” and words of similar import, when used in this application, refer to this application as a whole and not to any particular portions of this application. When the word “or” is used in reference to a list of two or more items, that word covers all of the following interpretations of the word: any of the items in the list, all of the items in the list and any combination of the items in the list.
[0128] The above description of embodiments of the artificial intelligence assisted monitoring of patients is not intended to be exhaustive or to limit the systems and methods to the precise forms disclosed. While specific embodiments of, and examples for, the artificial intelligence assisted monitoring of patients and corresponding systems and methods are described herein for illustrative purposes, various equivalent modifications are possible within the scope of the systems and methods, as those skilled in the relevant art will recognize. The teachings of the artificial intelligence assisted monitoring of patients and corresponding systems and methods provided herein can be applied to other systems and methods, not only for the systems and methods described above.
[0129] The elements and acts of the various embodiments described above can be combined to provide further embodiments. These and other changes can be made to the artificial intelligence assisted monitoring of patients and corresponding systems and methods in light of the above detailed description.
Claims
CLAIMS1. A method comprising, receiving first video data of a first plurality of subjects, wherein the first video data tracks a tracheostomy tube position in the first plurality of subjects, wherein the tracheostomy tube position comprises either a first tracheostomy tube position or a second tracheostomy tube position; training a predictive model using information of the first video data, wherein the information of the first video data comprises annotated frames, wherein the annotated frames indicate a first or second tracheostomy tube position, wherein the trained predictive model distinguishes between the tracheostomy tube position states; applying the trained predictive model to second video data of a subject, wherein the second video data tracks a tracheostomy tube position, wherein the trained predictive model detects a transition from the first tracheostomy tube position to the second tracheostomy tube position. la. The method of claim 1, wherein the training comprises constructing bounding boxes around tracheostomy tube positions in frames of the first video data. lb. The method of claim la, wherein the annotated frames comprise annotated bounding boxes. lc. The method of claim lb, wherein the constructing includes implementing a software tool to automatically identify the bounding boxes. ld. The method of claim 1c, wherein the software tool includes at least one of Roboflow. Label Studio, VGG Image Annotator, and COCO Annotator. le. The method of claim 1, wherein the first video data is derived from recordings taken during tracheostomy tube changes.lf. The method of claim le, wherein the plurality of subjects comprises a mannequin with an attached tracheostomy tube. lg. The method of claim le, wherein the plurality of subjects comprises a human subject using a tracheostomy tube. lh. The method of claim 1g, wherein the recordings include additional computer simulation to deidentify the at least one subject. li. The method of claim le, wherein the recordings comprise computer generated imagery lj. The method of claim le, wherein the recordings comprise video captured under a plurality of recording conditions. lk. The method of claim Ij, wherein the plurality of conditions comprises variable lighting. ll. The method of claim Ij, wherein the plurality of conditions comprises variable camera angles. lm. The method of claim 1, wherein the trained predictive model comprises a You Only Look Once Neural Architecture Search. ln. The method of claim 1, comprising providing an alert to a remote monitoring system upon detecting a transition from the first tracheostomy tube position to the second tracheostomy tube position. lp. The method of claim 1, wherein the first tracheostomy tube position comprises an in- place condition. lq. The method of claim 1, wherein the second tracheostomy tube position comprises a dislodged condition.
Citation Information
Patent Citations
Augmented Reality Device for Providing Feedback to an Acute Care Provider
US20190282324A1
Smart endotracheal tube
US20230086031A1
System and method for visualizing placement of a medical tube or line
US20230162352A1
System and Method for Determining a Device Safe Zone
US20240078781A1
Systems and methods for guided airway cannulation
WO2024058849A2