Automated remote medical sensor (ARMS)
The ARMS system addresses the challenge of remote vital sign detection in dynamic environments by using multi-level neural networks for real-time ROI detection and processing, enhancing accuracy and reliability for triage in diverse conditions.
Patent Information
- Application Number
- PCT/US2025/013264
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-08
- Filing Date
- 2025-01-27
- Publication Date
- 2025-09-11
AI Technical Summary
Existing image-based systems for monitoring vital signs are ill-suited for dynamic, unconstrained environments like battlefields due to varying distances, lighting conditions, subject movements, and obstructions, requiring high resolution and close proximity, which complicates accurate remote sensing of medical conditions.
An Automated Remote Medical Sensor (ARMS) system utilizing multi-level neural networks for on-board real-time video processing to detect regions of interest (ROIs) and extract vital signs from a distance, employing embedded GPU systems and asynchronous processing to improve signal-to-noise ratios in diverse scenes.
Enables remote vital sign detection in challenging conditions, allowing medical personnel to provide triage via platforms like drones, with improved accuracy and reliability in dynamic environments.
Smart Images

Figure US2025013264_12092025_PF_FP_ABST
Abstract
Description
AUTOMATED REMOTE MEDICAL SENSOR (ARMS)Cross Reference to Related Applications
[0001] This patent application claims priority to, and thus the benefit of an earlier filing date from, U.S. Provisional Patent Application No. 63 / 562,898 (filed March 8, 2024), the contents of which are hereby incorporated by reference.Background
[0002] Image-based systems designed to monitor vital signs have been developed for use in hospital settings, where patients are typically monitored from their beds. However, these systems generally require fixed camera positions near the patient, with reliable communication bandwidth, stable power sources, and optimal lighting conditions. In contrast, extracting vital sign measurements in dynamic, unconstrained environments, such as battlefields, presents significant challenges. For instance, battlefields feature a range of variables, including varying distances, lighting conditions, platform types, subject movements, and other complicating factors like injuries or obstructions, all of which can obscure accurate measurements, especially from a distance. These challenges are further exacerbated in combat situations.
[0003] Existing techniques for vital sign measurement from video typically rely on facial recognition algorithms, which require high resolution and a close proximity to the subject (e.g., usually less than 5 meters) to identify areas of interest for monitoring. These algorithms also depend on specific conditions, such as minimal motion and well-lit environments. As a result, these methods are ill-suited for remotely sensing the medical conditions of individuals in remote or variable environments.Summary
[0004] Systems and methods herein provide for remote sensing medical conditions of a human subject. The embodiments herein may be useful to provide medical personnel with remote vital sign detection, such as in the cases of military conflict and disaster response. The embodiments herein arc designed for challenging cases and prove that vital sign measurement is possible in previously unattainable conditions. Thus, the embodiments herein may allow medical personnel to remotely provide triage via various platforms such as drones.
[0005] The embodiments herein include an Automated Remote Medical Sensor (ARMS) developed for mobile non-contact sensing of vital signs. One system provides on-board real-time processing of video to determine vital signs from a distance. The system may include multi-level neural networks capable of detecting regions of interest (ROIs) in more diverse scenes and extracting the ROIs at different levels. The system can extend operating range and provide fine resolution segmentation maps that are combined to improve signal to noise ratios of the vitals signals. The system processes video in an asynchronous framework at higher rates onboard a platform without compression.
[0006] In some embodiments, embedded GPU systems, such as the NVidia Jetson, and multi-level neural network segmentation processes are employed. The multi-level neural network segmentation may allow the system to operate in more unconstrained conditions and provide an SNR where vital signs can be measured. The asynchronous processing may alleviate processing which would otherwise be computationally intensive.
[0007] In one embodiment, system for remote sensing medical conditions of a human subject includes a video camera for capturing a video stream of a scene, and a processing system operable to implement at least one trained neural network, said at least one neural network being trained on human biological data. The at least one trained neural network is operable to process one or more video frames of the video stream from the video camera to detect the human subject, and to determine that the human subject is in distress. The at least one trained neural network is operable to process the one or more video frames of the video stream from the video camera to identify a region of interest of the human subject for vital sign extraction. And, the processing system is further operable to process the video stream from the video camera to extract a vital sign of the human subject based on the identified region of interest of the human subject (e.g., a heart rate of the human subject, heart rate variability of the human subject, a blood pressure of the human subject, an oxygen saturation of blood in the human subject, a perfusion of the blood in the human subject, an injury, etc.). In doing so, the processing system may process red and green channels of the video frames to determine the vital sign. In some embodiments, the at least one trained neural network comprises a plurality of trained neural networks operating in parallel. The neural network(s) may be trained from training data comprising at least one of actual human subject vital sign data, actual human subject pose data, generated human subjectvital sign data, or generated human subject pose data. The processing system may also include a video stabilization module operable to compensate for motion of a platform on which the video camera is configured. And, the system may be configured on an airborne platform, such as an unmanned aerial vehicle (UAV), a waterborne platform, a land vehicle, etc.
[0008] In some embodiments, the system includes a satellite navigation device operable to determine a geographic coordinate of the system, and a communication module. In such an embodiment, the processing system is further operable to summarize a medical condition of the human subject based on the extracted vital sign of the human subject, to encapsulate the medical condition of the human subject with the geographic coordinate of the system in an electronic message, and to transmit the electronic message via the communication module.
[0009] In some embodiments, the system may include a lidar module and / or a radar module operable to provide lidar and / or radar data of the human subject to the processing system to improve vital sign extraction of the human. This data may be input to a Kalman filter to process the vital sign of the human subject to improve an estimate of the vital sign of the human subject.
[0010] In some embodiments, the at least one trained neural network is operable to process the one or more video frames of the video stream from the video camera to identify another region of interest of the human subject for an injury. In this regard, the processing system may be further operable to process the video stream from the video camera to determine the injury to the human subject based on the other identified region of interest of the human subject.
[0011] In some embodiments, the at least one trained neural network is operable to process the one or more video frames of the video stream from the video camera to detect another human subject, and to determine that the other human subject is in distress. The at least one trained neural network may be operable to process the one or more video frames of the video stream from the video camera to identify a region of interest of the other human subject for vital sign extraction. And, the processing system may be further operable to process the video stream from the video camera to extract a vital sign of the other human subject based on the identifiedregion of interest of the other human subject. In some embodiments, the processing system may be further operable to extract the vital signs of both human subjects simultaneously.
[0012] The various embodiments disclosed herein may be implemented in a variety of ways as a matter of design choice. For example, some embodiments herein are implemented in hardware whereas other embodiments may include processes that are operable to operate the hardware. Other exemplary embodiments, including software and firmware, are described below.Brief Description of the Drawings
[0013] Some embodiments of the present invention are now described, by way of example only, and with reference to the accompanying drawings. The same reference number represents the same element or the same type of element on all drawings.
[0014] FIG. 1 is a block diagram of a system for remote sensing medical conditions of a human subject, in one exemplary embodiment.
[0015] FIG. 2 is a flowchart of a process of the system of FIG. 1, in one exemplary embodiment.
[0016] FIG. 3 is a block diagram of the system for remote sensing medical conditions of a human subject, in another exemplary embodiment.
[0017] FIGS. 4A and 4B illustrate neural network detection and pose estimation of human subjects in a scene of a video frame (e.g., ROIs), in one exemplary embodiment.
[0018] FIGS. 5A-5F illustrate image extraction in a multi-level neural network to identify layers of segmentation for signal extraction (e.g., ROIs), in one exemplary embodiment.
[0019] FIGS. 6 A and 6B illustrate image extraction in a multi-level neural network to identify layers of segmentation for wound identification (e.g., ROIs), in one exemplary embodiment.
[0020] FIGS. 7 A and 7B illustrate medium wavelength infrared (MWIR) imagery enhancement for neural network detection and pose estimation of human subjects in a scene, in one exemplary embodiment.
[0021] FIGS. 8A-8D illustrate neural network detection of ROIs in a pose of a human subject, in one exemplary embodiment.
[0022] FIGS. 9A-9D illustrate neural network detection of ROIs in a pose of a human subject, in another exemplary embodiment.
[0023] FIG. 10 is a block diagram of a system for remote sensing medical conditions of a human subject being augmented with a frequency modulated continuous wave (FMCW) and / or a FMCW lidar system.
[0024] FIG. 11 illustrates transmit and receive optical frequency or radio frequency (RF) waveforms of an FMCW lidar or an FMCW radar for respiration rate and / or heart rate detection, in one exemplary embodiment.
[0025] FIG. 12 illustrates sequence data of respiration and / or heartbeat measurements, in one exemplary embodiment.
[0026] FIG. 13 illustrates sequence data of respiration and / or heartbeat measurements that has been low pass filtered, in one exemplary embodiment.
[0027] FIG. 14 illustrates UAV platform in which an ARMS system is configured, in one exemplary embodiment.
[0028] FIG. 15 illustrates a boat platform in which an ARMS system is configured, in one exemplary embodiment.
[0029] FIG. 16 is a block diagram of a processing system operable to train and implement a neural network, in one exemplary embodiment.
[0030] FIG. 17 depicts one illustrative cloud computing system operable to perform the above operations by executing programmed instructions tangibly embodied on one or more computer readable storage mediums.Detailed Description of the Drawings
[0031] The figures and the following descriptions illustrate specific exemplary embodiments. It will thus be appreciated that those skilled in the art will be able to devise various arrangements that, although not explicitly described or shown herein, embody certain principles and are included within the scope of the embodiments. Furthermore, any examples described herein are intended to aid in understanding the embodiments and are to be construed as being without limitation to such specifically recited examples and conditions. As a result, the embodiments are not limited to any of the examples described below.
[0032] FIG. 1 is a block diagram of a system 100 for remote sensing medical conditions of a human subject (i.e., an Automated Remote Medical Sensor, or “ARMS”), in one exemplary embodiment. The system 100 is generally configured with a mobile platform, such as a boat, an automobile, a plane, a helicopter, a UAV, or the like. In this embodiment, the system 100 includes a video camera 102 and a processing system 104. The video camera 102 is operable to digitally capture video of a scene, as is traditional with digital video cameras. Thus, in doing so, the video camera 102 captures a plurality of video frames 106-1 - 106-N of the scene (where the reference number “N” is an integer greater than “1” and not necessarily equal to any other “N” reference designated herein). In some embodiments, the video camera 102 is a high-speed video camera capable of frame rates at 60 frames per second or higher. And, as the system 100 is generally configured on a mobile platform, the system 100 may be configured with a video stabilizer 120 that is operable to stabilize the video frames 106 with respect to the motion of the mobile platform. For example, when image frames arc collected from a moving platform, additional processing may be used to compensate for the relative motion between a human subject and the camera. Thus, the system 100 may be configured with the video stabilizer 120 to compensate for the platform’s motion.
[0033] The processing system is operable to implement a plurality of trained neural networks 108-1 - 108-N to process the video frames 106 to detect and identify a human subject in the scene. And, one or more of the neural networks 108 may be operable to detect various biological features of the human subject. For example, each neural network may be trained on human biological data. Examples of human biological data include human shapes, skin tones,hair colors, human poses, heart rates, perfusion, heart rate variability, blood pressure, oxygen saturation, and the like.
[0034] Once trained, a first of the trained neural networks 108-1 may be operable to process one or more of the video frames 106 to identify an ROI where a human subject may be located in a scene of the video frame(s) 106. In this regard, the neural network 108-1 may detect the human subject and determine a pose of the human subject. And, the neural network 108-1 may determine that the human subject is in distress based on the pose and / or motion of the human subject. Then, a second of the trained neural networks 108-2 may process the video frames 106 to identify an ROI where a vital sign of the human subject may be obtained.
[0035] The processing system 104 may also include a plurality of estimators 110-1 - 110-N, each communicatively coupled to a respective neural network 108. The estimators 110 are operable to receive the ROIs from their respective neural networks 108 and to process a live video from the video camera 102 based on the ROIs they receive from their respective neural networks 108. The estimators 110 may do so to track and update the pose of the human subject and extract a vital sign from the live video stream. For example, when a pose of a human subject has been detected by the neural network 108-1, the neural network 108-1 may transfer that ROI to another neural network 108-2 to extract a heart rate of the human subject. As the neural network 108-2 has been trained on locations where heart rates of human subjects can be detected, the neural network 108-2 may search for ROIs on the human subject where heart rates may be extracted. Once these ROIs have been identified, the neural network 108-2 transfers those ROIs to its associated estimator 110-2 for signal extraction of a heart rate. In this regard, the estimator 110 may zoom in on the pixels of the identified ROI in the video stream of the video camera 102 for heart rate extraction. Examples of other vital signs include heart rate variability of the human subject, a blood pressure of the human subject, an oxygen saturation of blood in the human subject, and a perfusion of the blood in the human subject.
[0036] To illustrate, when a blood pulse flows through capillaries, the capillaries expand and exert pressure on the surrounding skin. This pressure changes the reflectance of the skin’s surface ever so slightly. But, the change can be detected with standard red, green, blue (RGB) cameras when averaged over a larger area of skin, such as a forehead. By measuring the averagepixel intensity over a region of skin over several frames, a signal can be extracted. The detected signal may be correlated with a photoplcthysmogram (PPG) to provide information about several relevant vital signs, such as heart rate, oxygen saturation, heart rate variability, blood pressure changes, and perfusion. One of the neural networks 108 may subtract a normalized red channel of the video camera 102 from a normalized green channel of the video camera 102 to isolate the signal from noise.
[0037] The primary frequency of a PPG waveform is the heard rate. Fourier analysis may be used to estimate the peak frequency or heart rate. Ratios of a measured PPG in different wavelengths can be correlated with oxygen saturation. And, phase differences between PPG signals measured at two different locations, such as the face and the hand, may be correlated with pulse transit time. Blood pressure is a function of pulse transit time and individual- specific attributes, such as arterial stiffness. Although the embodiments may not have knowledge of some of these individual- specific attributes that affect absolute blood pressure measurements, they can be assumed to remain constant for an individual. For example, while arterial stiffness may not be known, it will not change in an individual over a matter of minutes. Thus, by monitoring the pulse transit time, the processing system 104 can estimate changes in blood pressure for an individual. Similarly, the processing system 104 may use measurements of the PPG signal at several locations on the body to provide information on perfusion, or blood flow in the extremities with enough pressure.
[0038] The processing system 104 may measure perfusion on a pixel-by-pixel basis by creating a compound score based on two different metrics, oxygenation and color channel correlation. The processing system 104 may measure how oxygenated blood reflects light in the visible band differently than deoxygenated blood by taking the ratio of the AC and DC components of the red and green channels of the video camera 102. This metric may be prone to lighting noise. Accordingly, the processing system 104 may measure how anti-correlated the red and green channels are. Real pulsatile signal from blood causes red content in skin to increase as green content in the skin decreases, and vice versa. The processing system 104 may combine these metrics to yield a measurement that can survive slight changes in lighting and movement.
[0039] The source of pulsatile signal is capillary pressure, and it is typically very weak and made undetectable due to noise, such as changes in illumination, subject motion, and platform motion. In cases where the subject is potentially critically injured, the pulsatile signal may not be consistent or may be even weaker. The processing system 104 may therefore use several hundreds of pixels that are averaged to measure a PPG signal with a frame rate of at least 15 frames per second (fps). To measure differences between PPG signals, the processing system 104 may use faster frame rates of 60-90 fps.
[0040] Other vital signs, such as subject motion, pose, wounds, and respiration rate can be measured more easily because they are larger signals and change at slower rates. The processing system 104 may measure respirations via expansions and contractions of the chest over time using optical flow between frames. The processing system 104 can measure respirations through clothing, including several layers and coats. Since the processing system 104 can be on a moving platform, accurate identification of the chest region in each extraction frame may be used for robust measurement.
[0041] A subject’s pose and their motion can be indicative of life and critical injuries. For example, decorticate or decerebrate poses are indicative of traumatic brain injuries. The processing system 104 monitors pose and subject motion for kinematic irregularities and provides information about other physical injuries. If the processing system 104 detects no motion, along with no measured heart rate or respirations, the processing system 104 can determine death.
[0042] To measure pose, the processing system 104 may employ the neural network 108-1 to identify key positions and joint locations. The processing system 104 may measure several angles between joints over time to assess motion. The processing system 104 may also identify visible wounds and bleeding localized to the body to provide additional assessment information.
[0043] To measure any of the aforementioned vital signs, the processing system 104 identifies pixels in ROIs which contain signal. The neural networks 108 provide region identification in less constrained operational environments, where subjects vary drastically in appearance based on injuries, clothing, camera angles, poses, shadows, and occlusions. Theneural networks 108 perform human subject detection, pose estimation, body region segmentation, wound detection, and skin segmentation. The neural networks 108 contain custom architectures to run in real-time on limited Size, Weight, and Power (SWAP) platforms.
[0044] Coordinating asynchronous processes to get vital signals can be challenging, because temporal agreement between algorithms cannot be assumed. Thus, the processing system 104 may combine outputs from various subprocesses so as to be robust to missing or delayed inputs. And, running these algorithms such that no one algorithm operates too slow may require optimization for low SWAP platforms.
[0045] Extending performance to less constrained conditions can also be challenging. Common assumptions about range, lighting, subject appearance, subject motion, and platform motion do not always hold. For example, training the neural networks 108 to detect ROIs in diverse conditions may require large amounts of training data which can be difficult to obtain due to the lack of technology being applied in this area. Thus, in some embodiments, training data is collected specifically and / or extrapolated from existing datasets for different applications.
[0046] In some embodiments, the processing system 104 may be configured onboard a remotely piloted UAV, another rotor craft, a fixed wing aircraft, a ground vehicle, a ship, etc. On board such a platform, the processing system 104 can measure vital signs from a human subject that may be in any number of environments and light conditions. For example, the subject may be within a forested area, or other natural environment, an urban area, or even within water. A subject’s body may be wet or dry, have varying amounts of skin exposure. A human subject may be in positioned arbitrary lighting, weather conditions, cluttered background, and orientations, and even have dismembered body parts. And, the surrounding surfaces may be wet or dry, concrete, or vegetation.
[0047] The processing system 104 may measure vital signs using real-time algorithms from simultaneous observation using multiple platforms, either unmanned or manned. The processing system 104 provides rapid assessment of vital signs to help prioritize and plan rescue and recovery operations of humans.
[0048] In some embodiments, the system 100 is configured with an output module 1 14 to generate messages pertaining to the vital signs of the detected human subject. For example, the output module 114 may receive outputs from the estimators 110 that the output module 114 may then incorporate into an electronic message (e.g., periodic, continuous, etc.) that can be communicated to a remote base station for search and rescue. The message may be configured in a variety of ways as a matter of design choice. For example, the output module 114 may summarize the vital signs of the human subject and / or provide a live stream of the vital signs of the human subject (e.g., a continuous PPG signal).
[0049] The system 100 may also include a satellite navigation device 112, such as the United States’s Global Positioning System (GPS), Russia’s Global Navigation Satellite System (GLONASS), China’s BeiDou Navigation Satellite System (BDS), and the European Union’s Galileo. With the processing system 104 being configured with a mobile platform, the satellite navigation device 112 can transfer the coordinates to the output module 114 for encapsulation into the message such that the location of the human subject can be quickly known for search and rescue.
[0050] The communication module 116 is operable to transmit the message to the remote base station. The communication module 116 may employ any communication technology as a matter of design choice including any generation of cellular network technology, and various other forms RF communications, including satellite emergency beacons and satellite communication networks, such as the Iridium network.
[0051] In some embodiments, the processing system 104 may be configured to transmit the video data to some sort of base station for processing. However, high bandwidth video would typically need to be compressed, especially for longer range communications from lower power platforms. And, processing algorithms for vital sign measurements are substantially degraded by video compression. Thus, the embodiments herein are generally directed to processing image frames on board the mobile platform so that imagery can be processed without compression, permitting the measurement of higher fidelity vital signs. The use of multi-level neural network algorithms enable operation in more unconstrained environments and diverse operating conditions.
[0052] In some embodiments, the plurality of neural networks may be combined into a single neural network model. In this regard, the model may contain the ability to perform both human detection and pose estimation by having a shared backbone with multiple task heads. Such a model may use the shared backbone between any of human detection, pose estimation, body part segmentation, and ROI selection for injury and vital sign extraction. For example, a single neural network model may be capable of a plurality of tasks, such as the first being human detection and distress determination, the second being injury detection, the third being ROI selection for vitals extraction, etc..
[0053] FIG. 2 is a flowchart of a process 200 of the system of FIG. 1, in one exemplary embodiment. In this embodiment, the process 200 initiates when the video camera 102 begins capturing video of a scene. The processing system 104 receives the video stream from the video camera 102, in the process element 202, and begins processing one or more video frames of the video stream through at least one trained neural network 108 to detect the human subject, and to determine that the human subject is in distress. The neural network 108 is trained on human biological data, as explained above. Then, the processing system 104 processes the one or more video frames through the neural network 108 to identify an ROI of the human subject for vital sign extraction, in the process element 206. As mentioned, the neural networks 108 may operate in parallel or operate as different independent functions within a single neural network.
[0054] With the ROI of the human subject identified, the processing system 104 processes the video stream from the video camera to extract a vital sign of the human subject based on the identified ROI of the human subject, in the process element 208. For example, the processing system 104 may register the video frames being processed through the trained neural networks 108 to the live video stream from the video camera. And, the neural networks 108 may pass the identified ROIs in the video frames to corresponding estimators 110 to identify vital signs signals in the live video stream. In this regal'd, the estimators 110 may focus in on the pixels of the live video stream where the ROIs have been identified to extract vital signs, as described above.
[0055] FIG. 3 is a block diagram of another processing system 300 operable to detect vital signs of a human subject, in one exemplary embodiment. In this embodiment, theprocessing system 300 receives video from the video camera 302 where it passes through a stabilization and registration module 304 which stabilizes the video to compensate for the motion of the mobile platform to which the processing system 300 is configured. The stabilization and registration module 304 selects a portion of video frames from the video camera 302 and registers those video frames to a live video stream from the video camera 302 such that the various estimators and their associated trained neural networks can perform their operations. For example, the detection and pose estimation module 306 includes a trained neural network (e.g., neural network 108-1 of FIG. 1) that has been trained on thousands of images, or more, of human subjects in various positions, motions, etc. The detection and pose estimation module 306, via its trained neural network, first identifies a human subject in the scene from one or more the stabilized video frames and determines a pose of the human subject. With the human subject detected and the human subject’s pose estimated, the pixels in the video frames where the human subject is detected are associated with previous video frames in the module 308 such that subsequent neural networks can focus on the pixels in the video frames where the human subject is located. An example of such as shown in FIGS. A and 4B.
[0056] In FIG. 4A, the detection and pose estimation module 306 detects two human subjects 402-1 and 402-2 in a scene 400 of a video frame. The pixels where the human subjects 402 are in the video frame are selected as a ROI 404, and the remaining pixels are masked out. Then, the neural network of the detection and pose estimation module 306 crops the video frame about the ROI 404, as shown in FIG. 4B. The module 306 may then transfer the ROI 404 to the module 308 such that the module 308 can associate the ROI 404 with previously captured video frames and operate on similar subsequent pixels in the scene 400. Here, the neural network of the module 306 has determined that the poses of the human subjects 402 are either deceased or injured.
[0057] The detection and pose information of the video frames (e.g., the ROIs) of module 308 are then transferred to various other trained neural networks to perform vital sign measurements. For example, the module 308 may transfer the information to the heart rate ROI segmentation module 310. Here, the module 310 uses its trained neural network to identify another ROI from the ROI 404. In doing so, the module 310 may receive the video frames that the detection of pose estimation module 306 received from the stabilization and registrationmodule 304. The trained neural network of the module 310 uses the information from the module 308 to focus on the video frame from the module 304 to identify an ROI on the human subject, such as exposed layers of skin where PPG measurements may be obtained. An example of such is shown in FIGS. 5A-5F.
[0058] FIGS. 5A-5F illustrate the image extraction of a multi-level neural network of the module 310 and how it identifies several layers of segmentation, from coarse to fine, for robust signal extraction. In FIGS. 5 A and 5B, the neural network identifies body pail ROIs 504, 506, and 508 where skin of a human subject 502 may be exposed for potential PPG measurements. Then, in FIGS. 5C and 5D, the neural network identifies skin in a finer level ROI selection for PPG measurement (i.e., ROIs 510, 512, and 514). And finally, in FIGS. 5E and 5F, the neural network identifies ROIs 516 used for signal extraction, with the brighter areas indicating the ROIs 516 in the video frame.
[0059] With the ROIs 516 located, this information is passed to an estimator of the module 310, which locates these regions in the live video stream from the video camera 302. For example, the estimator operates on the ROIs 516 in the live video stream based on the information from the neural network of the module 310. As the video camera 302 is a high resolution video camera with a high frame rate, the estimator can focus on color changes in the ROIs 516 on a frame by frame basis by measuring the average pixel intensity over several frames to extract the PPG signal. With the PPG signal obtained, the heart rate of the human subject can be determined and output to the module 312.
[0060] Similarly, the module 314 may process the detection and pose estimation from the module 308 through its neural network and estimator to identify a ROI for respiration detection in the human subject. For example, the module 314 may identify a chest region of the human subject based on the detection and pose estimation information. Then, the estimator of the module 314 may extract a respiration signal from the live video stream of the video camera 302.
[0061] The module 318 may use its neural network to process the detection pose estimation from the module 308 to identify an injury ROI in the human subject. Then, based onthe live video stream from the video camera 302, the estimator of the module 318 may output an injury characterization 320. An example of such as shown in FIGS. 6A and 6B.
[0062] FIG. 6A is a video frame of a soldier wounded in battle. Based on the detection and pose estimation from the module 308, the module 318 may process the video frame through its trained neural network to segment ROIs in the video frame that may indicate injury, as shown with the ROIs in FIG. 6B. In FIG. 6B, ROI 602 indicates bleeding. The neural network of module 318 is operable to determine whether the blood is fresh or old based on the color of the blood. The ROIs 604 indicate dismemberment. If module 310 and / or module 314 do not detect a pulse or a respiration rate, respectively, the processing system 300 may conclude that the human subject is deceased.
[0063] The module 322 may operate in similar fashion in detecting other vital signs as desired. In any case, the asynchronous processing chain allows for each vital sign to be calculated independently, and thus vital signs can be configured to run on their own or in subsets. Signal measurement requirements, such as frame rate and resolution ultimately depend on the system hardware. For example, memory and processing capability may determine the number of subjects and vital signs that can be measured at one time.
[0064] In this regard, the asynchronous processing chain is operable to separate each process into a background process which can update as fast as desired. These separate processes inform relatively low computational cost estimators (e.g., the estimators 110 of the FIG. 1) which can apply information to new frames at a much higher framerate than their associated neural networks can handle. The asynchronous processing chain allows for multiple vital signs from multiple subjects to be measured simultaneously.
[0065] In some embodiments, the neural networks may be trained on known data using various algorithms. For example, human subject detection, pose estimation, and segmentation of incoming imagery are initial steps for the system 300. One neural network that has been trained on vast amounts of detections and pose estimations includes YOLO (https: / / www.ultralytics.com / yolo), which could be used in the module 308. A module 308 could use Y OLO outputs as a phase correlation between frames to find an X-Y shift between the newest, high framerate feed and an image that YOLO was run on. The combined results couldthen be used on every frame in the high-framerate feed to extract a maximal number of measurements using embedded systems. A maximal number of measurements of a signal generally allows the best estimate of extremely weak signals, such as heart rate and respiration, that are necessary for remote measurements. This method allows the processing to be done onboard, as opposed to transferring the video to a base station where much of the signal would be lost in video compression. In some embodiments, Nvidia Jetsons and / or Google Coral accelerators may be used and combined with standard, modern cameras with high resolution, low noise, and relatively high framerates. The system 300 may thus be tuned for high rate measurements on a relatively low power system.
[0066] The human detection, pose estimation and body region network can identify ROIs at different scales in diverse scenes. Detection of subjects and ROIs at multiple resolutions allows for operation at increased range. Algorithm training to support various poses, occlusions, and injuries further enable performance in search and rescue or otherwise challenging environments.
[0067] OpenPose is another neural network that may be used for pose estimation and human subject detection algorithms. While many of these neural networks can be effective and serve as a good stalling point to build on, many of them tend to fail when used in a combat casualty environments for several reasons. For example, state-of-the-art human subject detection networks tend to fail because of a significant difference between the data they were trained on and data of relevant operating conditions. Typical training data has people well-lit, at large scales, and mostly centered in the image. And, people are easily distinguishable from their surroundings wearing bright clothes. But, battlefield conditions are extremely challenging with military personnel wearing camouflage and differ greatly from the cell phone like images that neural networks are trained on. Typical datasets simply fail to capture extreme postures, relevant scene diversity, and multiple camera angles that may be required for casualty detection from a UAV. The combination of all these discrepancies overwhelms traditional detection neural networks and lead to missed detections in relevant casualty conditions.
[0068] Another reason that these neural networks tend to fail is due to assumptions inherent in both the data and the neural network’s structure. Traditional neural networks aredesigned to have a strong prior for the human shape, which is based on labeling of key features on the body such as joints and facial features. This prior defines the human shape and constrains the poses a body can take, which is often strong enough to overcome new data. Additionally, missing or deformed limbs are important for casualty assessment. And, these neural networks often “hallucinate” a normal limb by extrapolating missing or obscured body parts based on their learned prior. Even worse, subjects that fail to align well with this prior are more likely to be missed completely by the detector. Thus, as a starting point, some embodiments herein train the neural networks on simulated imagery and / or actual imagery from battlefield conditions such as in field exercises.
[0069] Ground truthing this data often takes significant time. Finding and recording the precise location of people and their key points in each and every frame of the data is often intractable. Even with a clean dataset and a good labeling program, this can easily add up to hundreds of hours of manual labor. To leverage data more efficiently from field tests, the embodiments herein may use bootstrapping, a technique which iteratively refines labels by building on neural network predictions. Generally, a neural network is not perfect in creating initial labels. Thus, as a first pass, various open-source networks with permissive licenses, like YOLO, mmPose, or POEM may be used. The task of truthing then becomes much easier, as the truther only needs to correct mistakes made by the previous network. The embodiments herein may also exclude occluded points from the data to reduce the neural network’s tendency to build a human skeleton prior.
[0070] Additionally, a synthetic data generator may be used to produce associated ground-truth labels automatically. The synthetic data generator constructs datasets that focus on under sampled conditions, such as extreme posturing, camouflage, camera angle diversity, and range diversity. Results from hold out datasets and subsequent field tests have shown that performance has improved by incorporating these techniques.
[0071] In some embodiments, to successfully run in real-time from a UAV and to maintain a relatively low power usage to extend fight time and meet system cost requirements, a streamlined neural network is used. The embodiments herein may use shared weights andsmaller feature dimensions to perform well under diverse operating conditions while meeting runtime requirements for SWaP-limitcd platforms, such as UAVs.
[0072] In some embodiments, image frames may be from one or more cameras. For example, the video 302 may come from an RGB visible camera with spectral filtering to improve vital sign measurements. But, the camera may be enhanced with an infrared sensor. An example of how the neural network embodiments process medium wavelength infrared (MWIR) imagery to provide additional information is shown in FIGS. 7A and 7B. In FIG. 7A, a raw MWIR image is shown. After processing, the neural network keys in on human subjects 702 in the scene. The triangles 704 denote key body positions output from the neural network.
[0073] FIGS. 8A-8D and FIGS. 9A-9D illustrate two additional examples of multi-level neural network measurements. FIGS. 8 A and 9 A illustrate raw imagery received by the neural network. Then, as the neural network operates on the imagery, it identifies ROIs until the final ROIs are identified in FIGS. 8D and 9D. The neural network ROI identification may be performed at three different scales to improve performance under unconstrained and challenging imaging conditions (e.g., based on training).
[0074] Some embodiments may include onboard illumination to illuminate subjects. For example, the illuminator may improve nighttime performance and / or enhance performance in the day. In embodiments having illuminators, the spectrum of the illuminator may be designed specifically to enhance detection of blood for variations, for example, a mix of red and green light with matched spectral filters on a receiver.
[0075] In some embodiments, the system 300 may also be configured to process radar data and / or lidar data for making vital sign measurements. FIG. 10 is a block diagram of the system 300 being augmented with a frequency modulated continuous wave (FMCW) and / or a FMCW lidar system 1002. The system 1002 may be adapted to measure velocity and range of surfaces with extremely high accuracy. These technologies measure radial velocity, and can serve as an orthogonal measurement to the RGB camera-based measurements described above. The estimation of either respiration rate and / or heart rate can therefore be made in a joint fashion by using both measurements as an input to a Kalman filter to improve the estimate of therespiration rate and / or the heart rate allowing it to be made faster with more precision and more confidence. Then, the processing system 300 may output the measurements in the module 1008.
[0076] FMCW radar / lidar may measure the respiration rate and / or heart rate by measuring the velocity or displacement of the body to extremely fine detail. For example, a 60 GHz radar is capable of making sub-millimeter measurements throughout a full range. Much of the physics behind FMCW lidar and FMCW radar are the same, as both are electromagnetic radiation.
[0077] FIG. 11 shows an example of the transmit and receive optical frequency or radio frequency (RF) waveforms of an FMCW lidar or an FMCW radar, respectively. The transmitAv frequency is modulated in a sawtooth functional form with linear chirp rate y = — and constantmagnitude, where v is frequency and t is time. The received light has the same chirp but is 2.R delayed by a time (T = — ) that is proportional to the range (R) to a target and inversely proportional to the speed of light (c). The received waveform also has a Doppler shift in the2 v frequency that is proportional to the range rate between the lidar and the subject (5v = — ). During the rising and falling spans of the chirp signals, the received signal and transmitted signal may be coherently mixed where a detector measures the resulting beat frequencies. On the upward chirp, the frequency shift vais measured, and on the downward chirp the frequency shift vbis measured. The range and velocity can then be computed from the two beat frequencies.
[0078] With both FMCW radar and lidar, high repetition rate measurements of relative surface velocities, and ranges can be fused to infer body surface movement associated with respiration or heartbeats. In some embodiments, motion rate measurements from the camera and an inertial measurement unit (IMU) may provide additional information content to separate subject related motions from platform motion.
[0079] In one exemplary embodiment, a lidar that provides both range and range rate may provide a sequence of data as shown in FIG. 12. Here, each data measurement provides a Doppler velocity measurement and a range measurement. The measured velocities have a component representing platform motion as well as a component associated with motion of a subject (e.g., respiration and / or heartbeat). Over a sequence of measurements, oscillatorycontent associated with respiration and heartbeats can be separated from platform motion. For example, a fit of the sequential ranges can be low pass filtered to calculate a range rate associated with platform motion. After subtraction from the doppler measurements, a signal more associated with subject vitals can then be analyzed, as shown in FIG. 13. In additional embodiments, inertial rates and their frequencies may be measured from an IMU sensor and can be used to produce a platform motion power spectrum that can be subtracted from the power spectrum of the measured doppler velocity measurements to obtain a power spectrum associated with subject vitals. Thus, an FMCW lidar can be used to augment vital measurements from an image sensor.
[0080] The embodiments herein are operable to use an asynchronous method of image processing, where processing steps from a single frame of imagery may be performed simultaneously using estimators derived from previous image frames. Those parallel processing steps may include finding subjects within an image frame, determining ROIs of a human subject within an image frame for measurement analysis, and performing the single frame image analysis. Additionally, with each frame, estimators may be calculated to aid in the processing of subsequent frames. Estimators may include information such as pixel shifts, pixel-wise segmentation, pixel-wise classification, rotations for subsequent image frame analysis, filtered histories of previous estimators, and / or information relating to camera exposure times or lighting levels. This method of parallel processing with reliance on previous estimators permits high bandwidth measurements with reduced power requirements and processing hardware. A sequence of measurements in time from the parallel processing may be processed, either in parallel with the measurements or after bursts of measurements, to measure vital signs such as respiration rates or heart beats. Dismemberment, physical injuries, wounds, or subject motion may be determined from image processing related to body pose. Assessments that a subject is deceased may also be made.
[0081] FIGS. 14 and 15 illustrate various platforms in which an ARMS system, such as the system 100 of FIG. 1, may be configured. For example, FIG. 14 illustrates and ARMS system 1402 and a video camera 1404 configured with a UAV 1400 to identify a person 1410 in distress. However, those skilled in the art will readily recognize that the system 1402 and camera 1404 can be configured with other aircraft, such as helicopters and planes. And, FIG. 15illustrates an ARMS system 1502 and a camera 1504 being configured with a boat 1500 to identify a person 1510 in distress in the water.
[0082] FIG. 16 is a block diagram of a processing system 100 operable to train and implement a neural network, in one exemplary embodiment. In this embodiment, a single neural network 108 is shown being communicatively coupled to a single estimator 110 to illustrate the training and detection of a neural network 108 of the processing system 104 of FIG. 1. But, as should be understood by the description of the processing system 104, the processing system 104 is operable to implement a plurality of neural networks 108 and a corresponding plurality of estimators 110. The neural network 108 is operable to receive a plurality of training datasets 1602-1 - 1602-N. Each training data set 1602 may include an image 1606 and a corresponding validation 1604. For example, the data sets 1602 may include imagery pertaining to a human pose. The validation 1604 may indicate the type of pose of the human subject in the imagery. To illustrate, the image 1606-1 may include a human subject laying in a “fencing position” in a battlefield. The validation 1604-1 may indicate that the human subject in the image 1606-1 has a concussion. The image 1606-N may include a human subject limping and in need of rescue. And, the validation 1604-N may indicate such. These datasets are input to the processing system 104 to train the neural network 108. Typically, the number of datasets 1602 used to train the neural network 108 is in the thousands, if not hundreds of thousands, such that the neural network 108 may produce an output that is statistically relevant. Thus, the neural network 108 may be trained on a plurality of validated poses of human subjects.
[0083] Once trained, the neural network 108 receives a video frame 106-1 from the video camera 102 of FIG. 1. The estimator 110 receives a stream of the video frames 106-1 - 106-N from the video camera 102. The trained neural network 108 processes the video frame 106-1 to detect an ROI in the video frame 106-1 to detect and identify a human subject and a pose of the human subject in the video frame 106-1. Once the ROI has been determined, the processing system 104 registers the video frame 106-1 to the stream of video frames 106-1 such that the estimator 110 may “lock in” on the ROI in the video frames 106-1 - 106-N to detect and monitor a pose of the human subject to update or otherwise indicate changes in the pose of the human subject over time. The estimator 110 may then output information pertaining to the poseof the human subject over time to the output module 114, which may then be encapsulated in a message and transmitted to a base station.
[0084] Again, the processing system 104 may include a plurality of neural networks 108. Thus, as with the system 300 of FIG. 3, the neural network 108 may transfer the ROI of the human subject to another neural network to perform a measurement of the human subject. For example, a subsequent neural network 108 may be configured to measure a heart rate of the human subject in the ROI. Accordingly, the neural network 108 may identify an ROI on the subject that is suitable for detecting a heart rate of the human subject. Then, once identified, the ROI may be transferred to its corresponding estimator 110 such that the estimator 110 can extract a heart rate signal from the ROI over time in the video frames 106-1 - 106-N.
[0085] Some nonlimiting examples of the neural network 108 include feedforward neural networks (FNNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory networks (LSTMs), gated recurrent units (GRUs), generative adversarial networks (GANs), autoencoders, transformer networks, radial basis function, networks (RBFNs), modular neural networks (MNNs), spiking neural networks (SNNs), selforganizing maps (SOMs), deep belief networks (DBNs), and reinforcement learning networks. Machine learning may be alternatively or additionally employed in the neural network 108 and / or the estimator 1 10. Some examples of machine learning algorithms that may be used in the neural network 108 and / or the estimator 110 include a supervised learning algorithm, a semisupervised learning algorithm, an unsupervised learning algorithm, a regression analysis algorithm, a reinforcement learning algorithm, a self-learning algorithm, a feature learning algorithm, a sparse dictionary learning algorithm, an anomaly detection algorithm, a generative adversarial network algorithm, a transfer learning algorithm, and an association rules algorithm.
[0086] In FIG. 17, one illustrative cloud computing system 1700 is illustrated and is operable to perform the above operations by executing programmed instructions tangibly embodied on one or more computer readable storage mediums. The cloud computing system 1700 generally includes the use of a network of remote servers hosted on the internet to store, manage, and process data, rather than a local server or a personal computer (e.g., in the computing systems 1702-1 - 1702-N). Cloud computing enables users to use infrastructure and applications via the internet, without installing and maintaining them on-premises. In thisregard, the cloud computing network 1720 may include virtualized information technology (IT) infrastructure (e.g., servers 1724-1 - 1724-N, the data storage module 1722, operating system software, networking, and other infrastructure) that is abstracted so that the infrastructure can be pooled and / or divided irrespective of physical hardware boundaries. In some embodiments, the cloud computing network 1720 can provide users with services in the form of building blocks that can be used to create and deploy various types of applications in the cloud on a metered basis.
[0087] In one embodiment, instructions stored on a computer readable medium direct a computing system of any of the devices and / or servers discussed herein to perform the various operations disclosed herein. In some embodiments, all or portions of these operations may be implemented in a networked computing environment, such as a cloud computing system. Cloud computing often includes on-demand availability of computer system resources, such as data storage (cloud storage) and computing power, without direct active management by a user. Cloud computing relies on the sharing of resources, and generally includes on-demand self-service, broad network access, resource pooling, rapid elasticity, and measured service.
[0088] Various components of the cloud computing system 1700 may be operable to implement the above operations in their entirety or contribute to the operations in part. For example, a computing system 1702-1 may be used to perform all or portions of the operations of the embodiments herein, and then store results in a data storage module 1722 (e.g., a database) of a cloud computing network 1720. Various computer servers 1724-1 - 1724-N of the cloud computing network 1720 may be used to operate on the data and / or transfer analysis of the data and / or the data to another computing system 1702-N.
[0089] Some embodiments disclosed herein may utilize instructions (e.g., code / software) accessible via a computer-readable storage medium for use by various components in the cloud computing system 1700 to implement all or parts of the various operations disclosed hereinabove. Examples of such components include the computing systems 1702-1 - 1702-N.
[0090] Exemplary components of the computing systems 1702-1 - 1702-N may include at least one processor 1704, a computer readable storage medium 1714, program and datamemory 1706, input / output (I / O) devices 1708, a display device interface 1712, and a network interface 1710. For the purposes of this description, the computer readable storage medium 1714 comprises any physical media that is capable of storing a program for use by the computing system 1702. For example, the computer-readable storage medium 1714 may be an electronic, magnetic, optical, electromagnetic, infrared, semiconductor device, or other non-transitory medium. Examples of the computer-readable storage medium 1714 include a solid-state memory, a magnetic tape, a removable computer diskette, a random access memory (RAM), a read-only memory (ROM), a rigid magnetic disk, and an optical disk. Some examples of optical disks include Compact Disk - Read Only Memory (CD-ROM), Compact Disk - Read / Write (CD- R / W), Digital Versatile Disc (DVD), and Blu-Ray Disc.
[0091] The processor 1704 is coupled to the program and data memory 1706 through a system bus 1716. The program and data memory 1706 include local memory employed during actual execution of the program code, bulk storage, and / or cache memories that provide temporary storage of at least some program code and / or data in order to reduce the number of times the code and / or data are retrieved from bulk storage (e.g., a hard disk drive, a solid state drive, or the like) during execution.
[0092] Input / output or VO devices 1708 (including but not limited to keyboards, displays, touchscreens, microphones, pointing devices, etc.) may be coupled either directly or through intervening VO controllers. Network adapter interfaces 1710 may also be integrated with the system to enable the computing system 1702 to become coupled to other computing systems or storage devices through intervening private or public networks. The network adapter interfaces 1710 may be implemented as modems, cable modems, Small Computer System Interface (SCSI) devices, Fibre Channel devices, Ethernet cards, wireless adapters, etc. Display device interface 1712 may be integrated with the system to interface to one or more display devices, such as screens for presentation of data generated by the processor 1704.
[0093] Any of the above embodiments herein may be rearranged and / or combined with other embodiments. Accordingly, the concepts herein are not to be limited to any particular embodiment disclosed herein. Any of the various computing and / or control elements shown in the figures or described herein may be implemented as hardware, as a processor implementingsoftware or firmware, or some combination of these. For example, an element may be implemented as dedicated hardware. Dedicated hardware elements may be referred to as “processors,” “controllers,” or some similar terminology. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, a network processor, application specific integrated circuit (ASIC) or other circuitry, field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), non-volatile storage, logic, or some other physical hardware component or module. Some examples of software include but are not limited to firmware, resident software, microcode, etc.F0094] While the embodiments herein refer to the remote monitoring of human subjects, those skilled in the art should readily recognize that other life forms may be monitored as well. For example, the neural networks described herein may be trained on mammal and / or reptile biological data to detect poses and / or vital signs for remote monitoring purposes.
Claims
ClaimsWhat is claimed is:
1. A system for remote sensing medical conditions of a human subject, the system comprising: a mobile platform comprising: a video camera for capturing a video stream of a scene; and a processing system operable to implement at least one trained neural network, said at least one neural network being trained on human biological data, wherein said at least one trained neural network is operable to process one or more video frames of the video stream from the video camera to detect the human subject, and to determine that the human subject is in distress, wherein said at least one trained neural network is operable to process the one or more video frames of the video stream from the video camera to identify a region of interest of the human subject for vital sign extraction, and wherein the processing system is further operable to process the video stream from the video camera to extract a vital sign of the human subject based on the identified region of interest of the human subject.
2. The system of claim 1, wherein: said at least one trained neural network comprises a plurality of trained neural networks operating in parallel.
3. The system of claim 1 , further comprising: a satellite navigation device operable to determine a geographic coordinate of the system; and a communication module, wherein the processing system is further operable to summarize a medical condition of the human subject based on the extracted vital sign of the human subject, to encapsulate the medical condition of the human subject with the geographic coordinate of the system in an electronic message, and to transmit the electronic message via the communication module.
4. The system of claim 1, wherein: the mobile platform is an airborne platform.
5. The system of claim 1, wherein: the airborne platform is an unmanned aerial vehicle.
6. The system of claim 1, wherein: the mobile platform is a waterborne platform.
7. The system of claim 1, further comprising: a lidar module operable to provide lidar data of the human subject to the processing system to improve vital sign extraction of the human subject.
8. The system of claim 1, further comprising a radar module operable to provide radar data of the human subject to the processing system to improve vital sign extraction of the human subject.
9. The system of claim 1, wherein: the processing system comprises a video stabilization module operable to compensate for motion of a platform on which the video camera is configured.
10. The system of claim 1 , wherein: the vital sign of the human subject comprises at least one of a heart rate of the human subject, heart rate variability of the human subject, a blood pressure of the human subject, an oxygen saturation of blood in the human subject, or a perfusion of the blood in the human subject; and the processing system is further operable to process red and green channels of the video frames to determine the vital sign.
11. The system of claim 1, wherein: said at least one trained neural network is trained from training data comprising at least one of actual human subject vital sign data, actual human subject pose data, generated human subject vital sign data, or generated human subject pose data.
12. The system of claim 1, further comprising: a Kalman filter operable to process the vital sign of the human subject to improve an estimate of the vital sign of the human subject.
13. The system of claim 1, wherein: said at least one trained neural network is operable to process the one or more video frames of the video stream from the video camera to identify another region of interest of the human subject for an injury, and the processing system is further operable to process the video stream from the video camera to determine the injury to the human subject based on the other identified region of interest of the human subject.
14. The system of claim 1 , wherein: said at least one trained neural network is operable to process the one or more video frames of the video stream from the video camera to detect another human subject, and to determine that the other human subject is in distress, said at least one trained neural network is operable to process the one or more video frames of the video stream from the video camera to identify a region of interest of the other human subject for vital sign extraction, and the processing system is further operable to process the video stream from the video camera to extract a vital sign of the other human subject based on the identified region of interest of the other human subject.
15. The system of claim 14, wherein: the processing system is further operable to extract the vital signs of both human subjects simultaneously.
16. A method for remote sensing medical conditions of a human subject from a mobile platform, the method comprising: capturing a video stream of a scene with a video camera; implementing at least one trained neural network in a processing system, said at least one neural network being trained on human biological data; processing one or more video frames of the video stream from the video camera through said at least one trained neural network to detect the human subject, and to determine that the human subject is in distress; processing the one or more video frames of the video stream from the video camera through said at least one trained neural network to identify a region of interest of the human subject for vital sign extraction; and processing the video stream from the video camera through the processing system to extract a vital sign of the human subject based on the identified region of interest of the human subject.
17. The method of claim 16, wherein: the vital sign of the human subject comprises at least one of a heart rate of the human subject, heart rate variability of the human subject, a blood pressure of the human subject, an oxygen saturation of blood in the human subject, or a perfusion of the blood in the human subject; and the processing system is further operable to process red and green channels of the video frames to determine the vital sign.
18. The method of claim 16, further comprising: processing the one or more video frames of the video stream from the video camera through the at least one trained neural network to detect another human subject, and to determine that the other human subject is in distress; processing the one or more video frames of the video stream from the video camera through at least one trained neural network is operable to identify a region of interest of the other human subject for vital sign extraction; and processing the video stream from the video camera through the processing system to extract a vital sign of the other human subject based on the identified region of interest of the other human subject, wherein the processing system is further operable to extract the vital signs of both human subjects simultaneously.
19. The method of claim 16, wherein: said at least one trained neural network comprises a plurality of trained neural networks operating in parallel.
20. A non-transitory computer readable medium comprising instructions that, when executed by a processing system, direct the processing system to remote sense medical conditions of a human subject from a mobile platform, the instructions further directing the processing system to: capture a video stream of a scene with a video camera; implement at least one trained neural network, said at least one neural network being trained on human biological data; process one or more video frames of the video stream from the video camera through said at least one trained neural network to detect the human subject, and to determine that the human subject is in distress; process the one or more video frames of the video stream from the video camera through said at least one trained neural network to identify a region of interest of the human subject for vital sign extraction; and process the video stream from the video camera through the processing system to extract a vital sign of the human subject based on the identified region of interest of the human subject.
Citation Information
Patent Citations
Pre-hospital emergency rescue method and system based on 5G
CN114339148A
911 services and vital sign measurement utilizing mobile phone sensors and applications
US20130072145A1
One-Pass Video Stabilization
US20180041708A1
Use of low frequency electromagnetic signals to detect occluded anomalies by a vehicle
US20230243994A1