Systems and methods for predicting a condition requiring a life-saving intervention including leveraging contextual information and information on reliability of input data streams

The system using UAVs and UGVs with sensors and machine learning algorithms addresses the challenges of triage in mass casualty incidents by automating data processing and predicting life-saving interventions, enhancing efficiency and accuracy in disaster scenarios.

WO2025188356A9PCT designated stage expired Publication Date: 2025-11-06BATTELLE MEMORIAL INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/046386
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-12
Filing Date
2024-09-12
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Triage and delivery of life-saving interventions in mass casualty incidents are challenging due to the unpredictable nature of disaster scenes, overwhelming data streams, frequent false positives from monitoring instruments, and compromised data reliability, leading to inadequate monitoring and delayed interventions.

Method used

A system utilizing unmanned aerial and ground vehicles equipped with sensors to collect and process physiological and contextual data, applying statistical process control and machine learning algorithms to predict the need for life-saving interventions, integrating data streams, and providing timely recommendations.

Benefits of technology

Enables automated, efficient, and accurate prediction of life-saving interventions by reducing the burden on medical personnel, improving decision-making, and ensuring timely delivery of interventions in chaotic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024046386_06112025_PF_FP_ABST
    Figure US2024046386_06112025_PF_FP_ABST
Patent Text Reader

Abstract

To provide a life-saving intervention (LSI) recommendation, physiological data streams are acquired from a patient. Features are extracted from the physiological data streams. Statistical process control (SPC) is applied to the physiological data streams to generate SPC quality metric features for the physiological data streams. The SPC quality metric features are indicative of signal quality of the physiological data streams. The extracted features and the SPC quality metric features for the physiological data streams are combined, and a condition of the patient requiring an LSI is predicted based on the combination of the extracted features and the SPC quality metric features for the physiological data streams.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR PREDICTING A CONDITION REQUIRING A LIFESAVING INTERVENTION INCLUDING LEVERAGING CONTEXTUAL INFORMATION AND INFORMATION ON RELIABILITY OF INPUT DATA STREAMSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 537,991 , filed on September 12, 2023, which is incorporated by reference in its entirety.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0002] This invention was made with government support under contract number HR001124C0334 awarded by the Defense Advanced Research Projects Agency. The government has certain rights in the invention.BACKGROUND

[0003] The present disclosure relates to care, triage, and assessment of injured persons, to improved triage and delivery of life-saving intervention (LSI) in mass casualty incidents (MCls) such as battlefields, natural disasters such as hurricanes or earthquakes, other types of disasters such as high-rise building collapse, and / or so forth.

[0004] Triage and delivery of medical assistance at disaster scenes or military battlefields is challenging, due to factors such as the unpredictable location of such events and the potential for mass casualties with a wide range of possible injury types, such as blunt trauma, gunshot, burns, traumatic brain injuries, inhalation of toxic chemical and / or biohazardous biological agents, and / or so forth. Although triage protocols specific to mass casualty incidents (MCls) have been adopted, such as Sort, Assess, Lifesaving Interventions, Treatment / Transport (SALT), these practices are inadequate for the secondary period since casualties assumed to be stable are often left unmonitored, due to the need to allocate those scarce medical personnel to more immediate priorities. Even in situations where personnel can be allocated to medical monitoring, those that perform this task often can only depend on physiological signatures, such as heart rate, respiration, oxygen saturation, or temperature, to assess the state of each casualty whenattempting to observe many casualties. The sheer volume of information from multiple data streams from each casualty can oversaturate medical personnel attempting to monitor many different individuals, and it is difficult to track trends across different time scales. Alarms on monitoring instruments are employed to alert monitoring personnel once a single value moves outside an acceptable range. However, these signals fail to account for the holistic integration of all data streams and are often ignored in practice by medical personnel due to frequent false positives (i.e., alarm fatigue). These challenges are compounded by the potential chaos of an MCI environment, such as a battlefield or disaster area, or the tasks of monitoring patients during transport to a higher level of care.

[0005] Alarms on monitoring instruments are employed to alert monitoring personnel once a single value moves outside an acceptable range. However, these signals fail to account for the holistic integration of all data streams and are often ignored in practice by medical personnel due to frequent false positives (i.e., alarm fatigue). Thus, there is an unmet need for the automation of this process, using artificial intelligence (Al) technology, which can reduce the burden of collecting and interpreting physiological data streams in real time, provide decision-making support to medical personnel, and facilitate delivery of life-saving interventions (LSIs) in a timely manner.

[0006] For both combat personnel and civilians, triage occurs at every level of care. To maximize health outcomes, the patient should be monitored during initial assessment, evacuation, while in transit, upon arrival at field hospitals or medical centers, during treatment, and in recovery. Continuous monitoring enables rapid response interventions following unexpected events and patient health deterioration. To meet military needs of future conflicts in remote areas, triage support technology should be able to analyze physiological signatures and relay pressing information to field medics or remote care centers to support clinical decisions. In non-military MCls (e.g., natural disasters, building collapse events, et cetera), the same challenge applies; trained first responders may not have the bandwidth to monitor many individuals simultaneously for subtle changes in physiological signatures indicating the need for an LSI. Even at higher levels of care, hospitals can become overwhelmed with casualties, and resources can be misallocated leading to delayed or absent LSIs.

[0007] In certain triage situations, data reliability is a further challenge. For example, in triaging an MCI scene, data communication infrastructure may have been compromised or destroyed by the MCI. In such an MCI scene, casualty data may be transmitted over temporary wired and / or wireless channels with reduced reliability compared with the original infrastructure of the MCI scene. Additionally, sensors may be damaged, dislodged, obstructed or otherwise interfered with during an MCI which may further compromise data integrity. The compromised patient data may therefore be less reliable - but also may be the only data available for triaging casualties to determine whether LSI is needed. Similar reliance on compromised patient data can arise downstream, for example at local hospitals that are inundated with casualties transported to the hospital from the MCI scene.

[0008] The following discloses improved systems and methods for providing LSI recommendations that overcome these problems and others.BRIEF DESCRIPTION

[0009] Multiple different technologies are described herein for addressing such issues in the civilian and military context related to treatment of multiple persons and leveraging technology to provide timely information to improve care. Besides the battlefield, these may be useful for injuries in other hazardous conditions, such as fires, floods, hurricanes, earthquakes, or in a hospital setting, or during transport of a patient in an emergency vehicle, et cetera.

[0010] Disclosed is a triage method for triaging a scene of a mass casualty incident includes acquiring a plurality of data streams pertaining to a patient, where in some embodiments the acquiring of at least one data stream uses one or more sensors of the at least one UAV and / or UGV; extracting features from the data streams; applying statistical process control (SPC) to the data streams to generate SPC quality metric features for the data streams, the SPC quality metric features being indicative of signal quality of the data streams; combining the extracted features and the SPC quality metric features for the data streams; and predicting a requirement for a life-saving intervention (LSI) for the casualty based on the combination of the extracted features and the SPC quality metric features for the data streams.

[0011] Also described herein is a non-transitory computer readable medium storing instructions executable by at least one electronic processor to perform a patient condition assessment method for a patient. The method includes extracting features from patient data streams acquired for the patient; receiving patient contextual data from one or more data records associated with the patient; extracting patient contextual features from the received patient contextual data using a language learning model (LLM); combining the extracted features and the extracted patient contextual features; and predicting a requirement for a LSI for the patient based on the combination of the extracted features and the extracted patient contextual features.

[0012] Some embodiments also relate to an apparatus comprising a computer programmed to perform a life-saving intervention (LSI) recommendation process including: receiving physiological data streams acquired from a patient; extracting features from the physiological data streams; applying statistical process control (SPC) to the physiological data streams to generate SPC quality metric features for the physiological data streams, the SPC quality metric features being indicative of signal quality of the physiological data streams; combining the extracted features and the SPC quality metric features for the physiological data streams; and predicting a condition of the patient requiring an LSI based on the combination of the extracted features and the SPC quality metric features for the physiological data streams.

[0013] Some embodiments also relate to providing a life-saving intervention (LSI) recommendation. To this end, physiological data streams are acquired from a patient. Features are extracted from the physiological data streams. Statistical process control (SPC) is applied to the physiological data streams to generate SPC quality metric features for the physiological data streams. The SPC quality metric features are indicative of signal quality of the physiological data streams. The extracted features and the SPC quality metric features for the physiological data streams are combined, and a condition of the patient requiring an LSI is predicted based on the combination of the extracted features and the SPC quality metric features for the physiological data streams.

[0014] These and other non-limiting aspects and / or objects of the disclosure are more particularly described below.BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Aspects of the present disclosure are best understood from the following detailed description when read with the accompanying figures. It is noted that, in accordance with the standard practice in the industry, various features are not drawn to scale. In fact, the dimensions of the various features may be arbitrarily increased or reduced for clarity of discussion.

[0013] FIG. 1 is a scene 1000 of a mass casualty incident (MCI), in accordance with some embodiments of the present disclosure.

[0014] FIG. 2 is a flow chart showing a patient condition prediction method, in accordance with some embodiments of the present disclosure.

[0015] FIG. 3 is a schematic view of a trained neural network used in the method of FIG. 2, in accordance with some embodiments of the present disclosure.

[0016] FIG. 4 is a framework showing the method of FIG. 2, in accordance with some embodiments of the present disclosure.

[0017] FIGS. 5a-5f show example binary training masks for training transformer models used in the method of FIG. 2, in accordance with some embodiments of the present disclosure.

[0018] FIGS. 6a-6d shows a graph of example interpolated sequences from an unsupervised pretraining task by a transformer models used in the method of FIG. 2, in accordance with some embodiments of the present disclosure.

[0019] FIG. 7 shows another embodiment of the framework of FIG. 4, in accordance with some embodiments of the present disclosure.

[0020] FIG. 8 is a graph showing a comparison of minutes to a need for a life-saving intervention determined by the method of FIG. 2, in accordance with some embodiments of the present disclosure.

[0021] FIGS 9a-9d show an attention mechanism of a BERT language model used by the framework of FIG. 4, in accordance with some embodiments of the present disclosure.

[0022] FIG. 10 is a flow chart showing a patient condition prediction method, in accordance with some further embodiments of the present disclosure.DETAILED DESCRIPTION

[0019] A more complete understanding of the processes and apparatuses disclosed herein can be obtained by reference to the accompanying drawings. These figures are merely schematic representations based on convenience and the ease of demonstrating the existing art and / or the present development, and are, therefore, not intended to indicate relative size and dimensions of the assemblies or components thereof.

[0020] Although specific terms are used in the following description for the sake of clarity, these terms are intended to refer only to the particular structure of the embodiments selected for illustration in the drawings, and are not intended to define or limit the scope of the disclosure. In the drawings and the following description below, it is to be understood that like numeric designations refer to components of like function.

[0021] Spatially relative terms, such as “beneath,” “below,” “lower,” “above,” “upper” and the like, are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. The devices may be otherwise oriented (rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein may likewise be interpreted accordingly.

[0022] Numerical values in the specification and claims of this application should be understood to include numerical values which are the same when reduced to the same number of significant figures and numerical values which differ from the stated value by less than the experimental error of conventional measurement technique of the type described in the present application to determine the value. All ranges disclosed herein are inclusive of the recited endpoint.

[0023] The modifier "about" used in connection with a quantity is inclusive of the stated value and has the meaning dictated by the context (for example, it includes at least the degree of error associated with the measurement of the particular quantity). When used with a specific value, it should also be considered as disclosing that value. For example, the term “about 2” also discloses the value “2” and the range “from about 2 to about 4” also discloses the range “from 2 to 4.” The term “about” may refer to plus or minus 10% of the indicated number.

[0024] FIG. 1 diagrammatically illustrates a scene 1000 of a mass casualty incident (MCI) undergoing initial triage using one or more (illustrative two) unmanned aerialvehicles (UAVs) 1002 and one or more (illustrative one) unmanned ground vehicle (UGV) 1004. The MCI scene 1000 could be the scene of a natural disaster such as an earthquake, hurricane landfall, flood, tsunami landfall, or the like. The MCI scene 1000 could alternatively be the scene of a disaster in an urban environment such as a large- scale fire, a building collapse, or so forth. The MCI scene 1000 could alternatively be a battlefield. These are merely some nonlimiting illustrative examples. Advantageously, the UAVs 1002 can travel over impassible terrain, while the UGVs 1004 may have more limited mobility but can enter close spaces (e.g., inside buildings) that the UAVs 1002 may be unable to access. As diagrammatically indicated in the upper enlarged UAV image, each UAV 1002 includes sensors 1006, 1008, 1010 that provide various sensor signals for locating casualties 1012 in the MCI scene 1000 and for performing triage of such casualties. Although not shown, each UGV 1004 could also include such sensors. By way of some nonlimiting illustrative examples, the sensor 1006 may be an acoustic sensor such as a microphone, a phased acoustic array operating in the audible frequency range, a phased ultrasound (US) array operating in ultrasonic frequencies, various combinations thereof, or so forth. The sensor 1008 may be a radar, such as an ultrawide- band (UWB) radar. The sensor 1010 may be an imaging device such as a video camera, electro-optical (EO) sensor, an infrared (IR) sensor, or combination thereof.

[0025] The UAVs 1002 and / or UGVs 1004 can operate autonomously or semi- autonomously to locate the casualties 1012 using the sensors 1006, 1008, 1010. Some nonlimiting examples of semi-autonomous operation may include an operator designating a search area within which the UAV or UGV performs an autonomous search for casualties 1012, or with an operator taking direct (remote) control of the navigation of the UAV or UGV. A casualty 1012 may be located using the imaging device(s) 1010, e.g., by pattern matching or other object matching image analysis being performed on the acquired images (e.g., video frames), and / or by an IR sensor 1010 detecting the heat signature of a casualty 1012, various combinations thereof, or so forth. The acoustic sensor(s) 1006 may assist in casualty location by detecting sounds (e.g., moans, screams, or the like) from the casualty. The UWB radar may assist in detecting a casualty 1012 by detection of motion, for example. These are merely some nonlimiting illustrative examples. For some MCI scenes, dedicated personnel locator devices may be utilizedfor detecting casualties, e.g., by detecting RFID tags worn by personnel using an RFID tag reader included with the UAV 1002 or UGV 1004.

[0026] When a casualty 1012 is located, the UAV 1002 and / or UGV 1004 can operate autonomously or semi-autonomously to assess or triage the located casualty 1012 using the sensors 1006, 1008, 1010. Such assessment or triage is intended to detect any lifethreatening injuries, so that emergency response personnel 1014 (one of which is illustrated in FIG. 1 as a nonlimiting example) can plan and execute a life-saving intervention (LSI) that is appropriate for the detected life-threatening injury. By way of some nonlimiting illustrative examples, a speaker of the UAV or UGV may be operated to request a vocal response from the casualty, and the acoustic sensor(s) 1006 may then be utilized to assess the responsiveness of the casualty. More directly, if the casualty can speak and state his or her injuries, then the acoustic sensor(s) 1006 can employ speech recognition to derive the semantic meaning of the casualty’s speech. US-based sensors 1006 can detect other physiological signs such as breathing, shivering, and so forth. The UWB radar 1008 can detect physiological signs such as chest movement (respiratory and / or heart rate, e.g., using time-of-flight or Doppler US techniques). The imaging device(s) 1010 can be used to assess pose of the casualty (e.g., lying down, sitting, or standing), blink response, chest movement, body temperature (e.g., using an IR imager), tremors, blood (by color analysis, for example), and / or so forth, by suitable image analysis applied to the images (e.g., video frames) acquired by the imaging device(s) 1010. In making the assessment or triage of the located casualty 1012, a single UAV or UGV may be used, or two (or more) UAVs 1002 and / or UGVs 1004 may be employed, as shown in inset 1016 of FIG. 1 where both a UAV 1002 and a UGV 1004 are being employed to assess or triage a single casualty 1012. Using multiple UAVs and / or UGVs to assess or triage a single casualty can be useful if, for example, no single UAV or UGV has all requisite sensors to do so.

[0027] In addition to assessing or triaging located casualties 1012, the UAVs 1002 and / or UGVs 1004 can operate autonomously or semi-autonomously to identify the located casualties. For example, this can employ facial recognition if the casualty’s face is visible to be imaged by the imaging sensor 1010, or by reading identifying information contained in RFID tag or the like worn by personnel. In another approach, a locatedcasualty may be identified based on location, possibly along with certain identifying physical characteristics, if there is sufficient a priori knowledge to do so. For example, in a battlefield scene 1000 the positioning of soldiers on the battlefield may be known a priori. If the located casualty 1012 can be identified, then an electronic health record of the casualty can be retrieved to provide further information for the assessment or triage. Even if the casualty cannot be identified as a specific individual with a corresponding electronic health record, certain demographic information such as gender, size, approximate body mass index (BMI), race / ethnicity, and / or so forth may be extracted from the imaging acquired by the imaging sensor 1010 and used to provide further information for the assessment or triage.

[0028] The emergency response at the MCI scene 1000 may employ a large number of LIAVs 1002 and / or UGVs 1004, each of which generates a continuous stream of sensor data from its respective sensors 1006, 1008, 1010. This massive and continuous influx of data cannot practically be monitored and analyzed by the emergency response personnel 1014, and moreover, subtle signals indicative of casualty location and physiological signs contained in the video, US data, radar data, and / or other sensor data may not be practically (or at all) detectable by manual monitoring of the data by emergency response personnel 1014. Accordingly, the sensor data streams are wirelessly transmitted to a base station 1020, which in the illustrative example is a pair of mobile ground control stations 1020 deployed in respective vans or other emergency response vehicles. The ground control station(s) 1020 may optionally be connected with (or alternatively, may be) a remote base station such as a remote server computer or computer cluster or cloud computing resource to provide additional data processing capacity. Emergency response personnel 1014 may utilize user interfacing devices 1022 such as an illustrative electronic tablet or cellphone 1022 to interact with the base station 1020. Emergency response personnel 1014 may utilize the user interfacing devices 1022 to receive information raw sensor data (e.g., video acquired by the UAV camera 1010) and / or processed sensor data, such as the location of a casualty determined by the base station 1020 processing sensor data from one or more sensors of one or more UAVs 1002 and / or UGVs 1004, and / or physiological signs of a located casualty determined by the base station 1020 processing sensor data from one or more sensors of one or more UAVs 1002 and / orUGVs 1004, and / or a LSI recommendation for a located casualty determined by the base station 1020 processing sensor data from one or more sensors of one or more UAVs 1002 and / or UGVs 1004.

[0029] The emergency response at the MCI scene 1000 further entails medical stabilization of located casualties 1012 and extraction of the casualties to a field hospital, mobile medical unit, or (if available) a nearby hospital or other medical institution. These operations may be performed by emergency response personnel 1014, and / or by unmanned robotic systems and / or vehicles (not shown). As the emergency response proceeds to the casualty stabilization and extraction phases, the capability to identify the casualty increases - for example, emergency response personnel 1014 may now have access to the casualty’s wallet, purse, or other carried item(s) which may contain personal information, or fingerprinting or DNA testing may be employed to identify the casualty. Again, upon identification of casualties, electronic health records can be accessed to provide further information for the assessment or triage.

[0030] As previously noted, there is an unmet need for the automation of the collection and processing of the data on casualties 1012 from the UAVs 1002 and / or UGVs 1004 to reduce the burden of collecting and interpreting physiological data streams from the in real time, provide decision-making support to medical personnel, and facilitate delivery of LSIs in a timely manner. The following discloses improvements to address this unmet need. The disclosed improvements may also or alternatively find application in other contexts in which large quantities of patient data from a large number of patients are usefully processed to predict and recommend LSIs, such as in hospital emergency rooms or intensive care units.

[0031] FIG. 2 shows an example of a non-transitory computer readable medium implemented as a server computer 10 that includes at least one electronic processor 12 and a memory 14. The memory 14 of the server computer 10 stores instructions executable by the at least one electronic processor 12 to perform a patient condition assessment method 100, e.g., for assessing the condition of the casualties 1012 or other subjects (generally referred to herein as “patients”, although it is to be understood that “patient” should be broadly construed as more generally encompassing any person being monitored for a condition calling for LSI, such as a casualty 1012, e.g., an injured soldieror person injured in a natural disaster or other MCI or so forth. The patient condition assessment method 100 operates on data on the patient (e.g., data on a casualty 1012 collected by one or more of the UAVs 1002 and / or UGVs 1004), and potentially additional data such from electronic health records of the patient (e.g., the casualty 1012 once identified). In other applications, the patient condition assessment method 100 could be employed to detect situations calling for LSI in an emergency room, intensive care unit, general hospital wing, or the like. The method 100 includes operations 101-109. Additional operations may also be performed, and not all illustrated operations must be performed.

[0032] At an operation 101 , patient physiological data is received at the server computer 10. To do so, patient physiological data is collected from one or more sensors (not shown) associated with a patient (not shown). For example, the patient physiological data can include one or more of blood pressure, heart rate, oxygen saturation, respiration, temperature sensor, electroencephalogram (EEG) activity, electrocardiogram (EKG or ECG) activity, weight of the patient, a presence of liquid (i.e. , perspiration), a presence of swelling, contraction of a muscle or skin of the patient, skin conductance, a presence of edema, or any other suitable sensor to measure any other desired parameter or vital sign. The patient physiological data is collected and received by the server computer 10, in real time. In some nonlimiting illustrative examples with MCI context, the patient physiological data may be collected by the sensors 1006, 1008, 1010 of one or more of the UAVs 1002 and / or UGVs 1004, and optionally pre-processed to extract clinically useful information such as heart rate, EEG, ECG, or so forth. After examination by emergency response personnel 1014 and casualty extraction, the patient physiological data may be collected by patient monitors placed by emergency response personnel 1014, such as 12-lead ECG sensors and so forth.

[0033] At an operation 102, the received patient physiological data is filtered using statistical process control (SPC). SPC is a process that comprises monitoring for signal disruptions in real time in the received patient physiological data.

[0034] At an operation 103, patient physiological features from the filtered patient physiological data are extracted. In some embodiments, data-driven features are alsoextracted from the filtered patient physiological data. These data-driven features are derived from one or more machine-learning models (discussed in more detail below).

[0035] To perform the extraction operation 103, a deep learning model 16 such as a neural network (NN), convolutional neural network (CNN), or the like is stored in the memory 14 of the server computer 10 and executed by the processor(s) 12 to extract numerical values from the filtered patient contextual data. The deep-learning model 16 is trained by masking values in training data.

[0036] At an operation 104, patient contextual data is received at the server computer 10. To do so, patient contextual data is retrieved from one or more patient data records of the patient. For example, the patient contextual data can include demographic data of the patient, injury data of the patient, electronic health record (EHR) of the patient, comorbidities data of the patient, lab results data of the patient, and so forth.

[0037] At an operation 105, the received patient contextual data is filtered to detect and remove disrupted or otherwise unreliable signals.

[0038] At an operation 106, patient contextual features from the filtered patient contextual data are extracted. The contextual features comprise information about the patient that may provide context for interpreting the patient physiological data to identify when a life-saving intervention should be performed. The contextual information may include (or be derived from) patient demographic information, information on injuries to the patient, chronic medical conditions of the patient (e.g., cardiac disease, diabetes, et cetera), lifestyle information (e.g., smoker or nonsmoker), and / or so forth. In some embodiments, data-driven features are also extracted from the filtered patient contextual data. These data-driven features are derived from the deep learning model 16 (or another deep learning model dedicated to the patient contextual data). To perform the extraction operation 106, the deep learning model 16 extracts numerical values from the filtered patient contextual data.

[0039] It will be appreciated that the above description, and in view of FIG. 2, that the “physiological data” operations (i.e. , the operations 101, 102, and 103) parallel, for the most part, the “contextual data” operations (i.e., the operations 104, 105, and 106). The patient physiological data operations and the patient contextual data operations each include corresponding receiving operations (101 and 104, respectively), correspondingfiltering operations (102 and 105, respectively), and corresponding extracting operations (103 and 106, respectively). It will also be appreciated that the patient physiological data operations 101, 102, and 103 can be performed concurrently (or sequentially, in either order) with the patient contextual data operations 104, 105, and 106.

[0040] At an operation 107, the extracted numerical values of the patient physiological features and the extracted numerical values of the patient physiological features are combined. To do so, in one nonlimiting illustrative approach the extracted numerical values of the patient physiological features and the extracted numerical values of the patient physiological features are input to a trained neural network (NN) 18 that is stored in the memory 14 of the server computer 10 and executed by the processor(s) 12. The NN 18 is trained using one or more of patient data collected as a time series, textual data from patient contextual data, and patient imaging data. The training data may be labeled with “ground truth” values that the trained NN should output. For example, if the NN 18 is being trained to detect when the LSI of mechanical ventilation should be administered, then the training data may be from previous patients who required such LSI (positive examples) and from previous patients who did not require such LSI (negative examples), with the training data labeled as to whether the LSI was required. The NN training then adjusts weights and / or other parameters of the NN, for example using backpropagation techniques or other known NN training techniques, to cause the NN output indicative of whether the LSI is required to optimally match the training data labels.

[0041] With brief reference to FIG. 3, a nonlimiting illustrative example is shown in which the trained NN 18 includes at least one transformer model 20 configured to extract the features from the patient data, and a fusion model 22 configured to combine the extracted features.

[0042] At an operation 108, a condition of a patient requiring a life-saving intervention (LSI) is predicted based on the combination of extracted features (at the combining operation 107) based on the corresponding need or requirement for an LSI for the patient P. To do so, again with brief reference to FIG. 3, the trained NN 18 includes an output model 24 configured to output a predicted condition of a patient based on the combination of extracted features.

[0043] At an operation 109, the predicted condition of the patient requiring an LSI (at the predicting operation 108) is reviewed by comparing a condition of the patient prior to the prediction and a condition of the patient after a life-saving intervention is provided to the patient based on the predicted condition. In some embodiments, the NN 18 is dynamically trained using the predicted condition of the patient. For example, in some nonlimiting embodiments the review operation 109 is used to generate further training data, e.g., if the patient condition improved after LSI, then this is a positive training example, whereas if the patient condition did not improve (or worsened) after LSI then this is a negative training example. This further training data is used to update-train the NN 18.

[0044] Referring now to FIG. 3, and with continuing reference to FIG. 2, a schematic view of the trained NN 18 is shown. Patient data 26 can be input to the one or more transformer models 20 to extract the features from the patient data 26. As shown in FIG. 3, the patient data 26 can comprise a first time series of patient physiological data collected in real time (i.e., “TimeSeriesT’), a second time series of patient physiological data collected in real time (i.e., “TimeSeries2”), textual data of the patient contextual data, and imaging data of the patient. It will be appreciated that any suitable type of data can be used.

[0045] As shown in FIG. 3, each of these types of patient data 26 can be input into a corresponding transformer model 20. Each transformer model 20 can comprise an individual NN and be dedicated to a specific type of data. For example, TimeSeriesI can comprise heart rate data of the patient, and thus the corresponding transformer model 20 can be a “heart rate” transformer model 20. Similarly, TimeSeries2 can comprise blood oxygen saturation data of the patient, and thus the corresponding transformer model 20 can be an “blood oxygen saturation” transformer model 20. Similarly, the text and image data 26 can be input to dedicated corresponding text and image transformer models 20. There can be any suitable number of transformer models 20 in the NN 18, each of which corresponds to a specific type of patient data 26.

[0046] The transformer models 20 perform the corresponding operation 103. Namely, the transformer models 20 perform the extracting operations 103 and 106.

[0047] Features may be calculated without the transformer models by using clinical metrics or other analysis of patient data, called hand-crafted features. Hand-crafted features are specific to the physiological data type. For example, a heart rate variability feature is calculated from an electrocardiogram by determining the variation in time between heartbeats.

[0048] The fusion model 22 is configured to combine the extracted features. To do so, the fusion model 22 is configured to combine the numerical values of the extracted features (e.g., by summing, averaging, determining a standard deviation, and so forth). The hidden state of the model 28 is passed to the output model 24, which is configured to output a predicted condition 30 of the patient based on the combination of extracted features from the fusion model 22. The output 30 can also be a classification of a state of the patient, the LSI required to treat the patient, and so forth.EXAMPLES

[0049] Disclosed herein is a nonlimiting illustrative example of a machine learning framework, named Continuous Review and Intervention for Timely Care (CRITIC), to predict the need for an LSI. The CRITIC process is an end-to-end solution that predicts the need for an LSI based on raw medical sensor signals and contextual information. The disclosed machine learning framework is a data fusion, machine learning framework that processes medical sensor data and medical documentation to predict the need for an LSI. The disclosed system can handle corrupted signals and various sensor combinations. The disclosed system can be used to improve allocation of medical resources and give greater lead-time on medical emergencies. The illustrative data fusion framework provides advantages such as robustness to damaged or missing sensors.

[0050] CRITIC utilizes the disclosed framework that integrates machine learning methods to form a fully end-to-end solution that predicts LSI based on raw sensor signals and contextual information. The disclosed model is designed to be robust to incoming signal disruptions while remaining flexible to changing sensor arrays and the inclusion of dynamic contextual features.

[0051] FIG. 4 shows another embodiment of the method 100, depicted as a framework. The framework is divided into five levels, each contributing towards the goal of identifying and processing sensitive physiological signatures for accurate and timely LSI predictions. FIG. 4 shows the patient data 26, LEVEL 1 (corresponding to the receiving operations 101 and 104 and the filtering operations 102 and 105), LEVEL 2 (corresponding to the physiological feature extracting operation 103), LEVEL 3 (corresponding to the contextual feature extracting operation 106), FEATURES (corresponding to the combining operation 107), LEVEL 4 (corresponding to the prediction operation 108), and LEVEL 5 (corresponding to the reviewing operation 109).

[0052] LEVEL 1 shows the signal preprocessing and statistical process control (SPC) operations. The SPC is responsible for detecting and cleaning (e.g., removing) disrupted signals from input sensors. Beyond preprocessing such as frequency filtering, this LEVEL 1 shows the SPC process continuously monitoring for signal disruptions in real time.

[0053] In LEVEL 2, physiological signatures are extracted through two synergistic processes to be used as features for LSI predictions. Specifically, both hand-crafted features based on physiological principles and data-driven features identified by pretrained machine learning models are extracted based on the transformer models 20. The transformer models 20 can be pre-trained in an unsupervised manner, allowing for the incorporation of outside datasets that do not pertain to LSI to improve model performance. Alternatively, supervised pre-training can be employed to train the transformer models 20

[0054] LEVEL 3 is focused on integrating contextual information about specific patients and injuries to aid in LSI predictions. LEVEL 3 uses deep learning patient representations (e.g., the deep learning model 16), augmented by open-source and / or other available datasets, to customize LSI predictions to each individual and injury.

[0055] In LEVEL 4, a self-attention feature fusion (SAFF) process (e.g., the trained NN 18) combines physiological and contextual features for LSI prediction. In a typical SAFF process, a multi-headed scaled dot-product attention is used to compute a representation of the input features. The combination of the features (i.e. , the “fusion” aspect) advantageously combines information from the different sources to provide an accurate LSI prediction. The SAFF process is flexible in allowing sensors to be added orremoved on the fly (e.g., due to actual sensor removal or failure, or due to sensor data removal at LEVEL 1 in response to SPC process detection of sensor data unreliability) without changing the model architecture or the need for retraining.

[0056] In LEVEL 5, insights into the decisions by the trained NN 18 can be examined to accelerate closed-loop system development. Through dimensionality reduction and clustering, sensor and contextual features can be visualized to gain insights into physiological states prior to LSI. In the SAFF process, the internal representations of the contextual information can be examined to gain insights into which sensor features the SAFF uses to make decisions about LSI.

[0057] The CRITIC process is designed to overcome the challenges inherent in LSI prediction. In LEVEL 1 , the SPC process can provide automated detection of signal outages. During training of CRITIC, sensor data can be periodically dropped to mimic signal disruptions, to provide simulated signal disruption training examples for training graceful handling of signal disruptions within CRITIC. The LEVEL 2 feature extraction methods in CRITIC have been developed to incorporate outside datasets to extract flexible features from the limited amount of LSI training data. In LEVEL 4, the SAFF is highly flexible and readily adapts to changes in sensor and contextual features, making it capable of accommodating a multitude of sensor configurations. These aspects of CRITIC combine to yield a highly flexible and robust system that is able to predict the need for LSI in a wide range of situations.

[0058] Advantageously, the CRITIC process improves and automates triage procedures and patient care through rapid, accurate, and explainable artificial intelligence (Al)-guided development of injury signatures, systematic data fusion, and predictive modeling. CRITIC is capable of processing sensor data streams with data points in rapid increments (e.g., every second, every minute, et cetera) in real-time to provide timely LSI recommendations. Such real-time analysis and thereby well-informed LSI recommendation cannot be practically performed by medical personnel, especially in an MCI where there may be many casualties. CRITIC improves medical triage and LSI delivery technologies by providing timely LSI recommendation (and, in some contemplated embodiments, automated delivery of the LSI by a robotic or other automated LSI delivery device automatically triggered by the LSI recommendation) basedon such real-time analysis of medical sensor data streams, in some embodiments informed by contextual information such as patient demographics. Incorporation of SPC processing enables real-time removal of unreliable data streams from consideration by the CRITIC LSI recommendation process; by contrast, human medical personnel attempting to interpret the incoming data streams might erroneously accept unreliable data and potentially perform unnecessary and possibly harmful LSI administration based on such unreliable data. Administering triage and life-saving interventions are formidable tasks, especially in chaotic and demanding environments such as active battlefields or mass causality incident sites. In addition to limited responder personnel and extreme levels of stress, decisions can be impaired by inadequate medical diagnostic support. For example, under- and over-triage of patients can each exceed 30%, resulting in insufficient medical care and inefficient use of medical resources, respectively.

[0059] A technical challenge solved by the CRITIC process is the integration of varied and volatile data streams so that reliable predictions and state estimations can be made. For an autonomous system to provide value to first responders and medical personnel, it must be robust to variations in sensor inputs including missing channels or data streams, and differing device manufacturers or models that that affect signal-to-noise, sampling rate, latency, and other properties. CRITIC can span a variety of datasets, preprocessing the underlying time series and contextual data into features that are not dependent on specific hardware or practices.

[0060] Another technical challenge is to create a system robust to varied sources of environmental noise and data corruption. Especially in battlefield or civilian pre-hospital settings, sensors may be damaged, dislodged, or record unpredictable noise due to movement artifacts, electromagnetic noise, and other factors. If not handled properly, aberrant values will propagate throughout the data processing pipeline and interfere with LSI predictions, leading to inadequate medical resource management and increased patient mortality. The processing employed in CRITIC effectively addresses this technological challenge in providing accurate LSI recommendation.

[0061] Another challenge is to account for the differences between subjects given limited training data and complex pathophysiology across a broad range of injuries. For instance, some signatures of hemodynamic compromise are weaker predictors forindividuals with preexisting primary hypertension. Likewise, changes in relevant vitals can be masked by physiological compensatory processes such as vasoconstriction. The identification of sensitive signatures that are robust across patient populations is a complex endeavor that can be accelerated by advanced ML algorithms. In CRITIC, the use of patient contextual information effectively addresses this challenge in providing accurate LSI recommendation.

[0062] Existing triage and vitals monitoring systems lack sensitive physiological signatures and are prone to inaccurate or lagged LSI alerts. The disclosed CRITIC process can incorporate advancements in AI / ML techniques, e.g., self-attention feature fusion, unsupervised pretraining, and statistical process control.

[0063] The disclosed nonlimiting illustrative CRITIC process is augmented with a statistical processes control (SPC) algorithm to monitor and protect from diverse signal disruptions. The self-supervised pretraining of CRITIC allows for incorporation of contextual data not directly related to LSI to improve feature extraction. Furthermore, due to the flexibility of the SAFF, CRITIC can process heterogenous data from many sensors and integrate contextual information to inform LSI predictions. The disclosed approach also leverages explainable ML for closed loop development, process refinement, and interpretability allowing LSIs to be performed with confidence.

[0064] The disclosed CRITIC process has enormous potential to positively impact the secondary triage period, both domestically in the United States and globally for American service personnel on the battlefield and other MCI scenes. This disclosed CRITIC process will grant the ability to effectively identify and treat casualties while they await a higher level of care for LSIs by constantly monitoring available physiological signatures for realtime triage assessments. The use of an autonomous process to perform tasks that may otherwise occupy and / or overwhelm first responders present an additional force multiplier, enabling incident commanders to maximize personnel resources towards extracting casualties and providing direct medical interventions. A data-driven and automated triage decision support system will greatly increase the speed and accuracy of monitoring decisions.

[0065] A secondary benefit is reduction in mental stress, emotional burden and guilt from the person making triage decisions, and accounting for both clinical and non-clinicalfactors unique to each incident. The disclosed system can dynamically adjust as medical resources are made available or are taken offline, and account for unique aspects of each specific type of MCI, which can be continuously improved through outcome assessments following each event.

[0066] With continuing reference to FIG. 4, LEVEL 1 of the CRITIC process can include signal preprocessing. To detect complex and interdependent physiological processes indicating the need for an LSI, a suite of sensors spanning a diverse set of biometrics can be required. Common data streams such as blood pressure, electrocardiograms, photoplethysmography, and respiratory rate may optionally beneficially undergo modality-specific preprocessing. Not only may these sensors capture signals with unique frequency characteristics, but each may be exposed to unique sources of noise, especially during transit and in dynamic MCI environments.

[0067] To do so, closed-loop clinical systems leveraging custom biosignatures across multiple signal modalities have been developed, including electromyography and intracortical neural recordings. Inevitably, these systems experience unexpected signal disruptions that can degrade performance. The CRITIC process can thoroughly characterize these bio-signal disruptions, and the SPC process is used to mitigate their harmful effects. The first stage of processing for the proposed CRITIC process can use the SPC algorithm to automatically detect and neutralize corrupted signals. This method uses statistical analysis to identify a variety of issues including channel drifts and disconnections that will be relevant in austere and uncontrolled environments. This methodology provides identification and integration of statistics relevant to each sensor.

[0068] Individual channels with disruptions will be omitted from the processing pipeline for as long as they are out-of-control. This will be handled flexibly by linking the SPC outputs to the LEVEL 2 feature extractors described in the following section. In LEVEL 2, the transformer models 20 are pretrained by dynamically masking input data, and therefore should be robust to short outages in sensor data. Additionally, if many channels from the same device are out of control indicating severe hardware failure, the feature fusion framework can dynamically omit the entire sensor (see LEVEL 4). This SPC framework detects errant data near the source and prevents it from propagating through later processing stages and interfering with LSI predictions. This level ultimately will helpmitigate false positive and false negative LSI predictions leading to better medical care and efficient resource management.

[0069] LEVEL 2 of the CRITIC process can include feature extraction from the patient data. In one nonlimiting illustrative approach, two complementary approaches can be used for extracting physiological signatures: 1 ) hand-crafted signatures rooted in the understanding of underlying physiological processes, supported by cutting edge biomedical literature and 2) data-driven signatures supported by machine learning methods and training data. Both types of signatures will be used features within the illustrative feature fusion model 22 (LEVEL 4) and will be referred to herein as such.

[0070] For the hand-crafted feature extraction approach, after corrupted channels are cleaned, signatures of impending need for LSI can be identified. Detecting these signatures can be challenging due to physiological compensatory mechanisms, dynamic shifts across time, injury variability, and individual differences. To resolve this, wavelet transforms can be used to extract volitional motor intent from intracortical recordings. This signature is stable and enables decoding of motor intention across several years This signature can be preferred to existing signatures such as neural spike rate and other frequency-based features. This signature may be directly applicable for many other physiological data streams with minor adaptations to account for frequency differences. For example, wavelet-based transforms have many applications in detecting ECG waveform abnormalities in disease states.

[0071] In some examples, electromyography signatures to track neurological recovery from stroke and spinal cord injury survivors can be identified. In other examples, biosignal subject matter expertise can be used to identify and improve state of the art signatures correlated with deterioration / mortality including compensatory reserve index. As shown in FIG. 4, these hand-crafted signatures can be integrated with the ML features described below as well as contextual features (LEVEL 3) during feature fusion (LEVEL 4). These fused features will reflect interdependent physiological processes and will be used to predict the need for an LSI.

[0072] For the data-driven feature extraction approach, learned features can be extracted from sensor inputs through the use of machine learning models. In contrast to the hand derived features described above, the transformer models 20 can be used,which are a class of neural networks that utilize the attention mechanism, to directly learn features from the patient data 26. An attention mechanism of the fusion model 22 was developed in the self-attention feature fusion (SAFF) process to allow the transformer models 20 to utilize information from a sequential input in a flexible way, analogous to information retrieval by creating query, key, and value embeddings of each input. The relationship between these three values captures the contextual information of the input sequence and learns this contextual relationship from the data during training.

[0073] To overcome the size limitations of the LSI prediction datasets, the transformer models 20 can be pre-trained on large amounts of unsupervised data. The transformer models 20 can be scaled well with the amount of training data available. In addition, the transformer models 20 have shown strong results in time series classification and time series regression. In the disclosed CRITIC process, the transformer models 20 are configured to extract features from each of the sensor data streams. As depicted in FIG. 4, the preprocessed data streams 26 from each sensor will be passed to a separate transformer model 20, which learns a set of features for each sensor. These learned features are then combined to make a final prediction in the feature fusion stage of the SAFF process, described in LEVEL 4. This may utilize the end-to-end training of the preprocessed sensor signals and augment hand designed features that may not optimally exploit information from the sensors.

[0074] The pretraining of the transformer models 20 advantageously masks out portions of the time series and trains the transformer to fill in any missing pieces. After the unsupervised pretraining, the pretrained transformer models 20 may optionally be fine-tuned for a down-stream task using a labeled dataset. Transformer models exposed to large amounts of unsupervised data have many interesting properties such as robustness to outliers (i.e., out-of-distribution samples). The transformer models 20 for each sensor can be pre-trained using existing bio-signal datasets. The pretrained transformer models 20 can then be fine-tuned for the prediction of LSI. This will allow the exploitation of existing bio-signal datasets, not used for the prediction of LSI, to improve the performance on the LSI prediction task.

[0075] The model used in the pretraining of the transformer models 20 is a transformer encoder model that will be adapted to the unsupervised pretraining task on ECG data. Inan illustrative approach, a binary mask is created to match the size of the incoming ECG signal. In the binary mask, zeros represent the portions of the input signal that are set to zero and the ones represent portions of the signal that are left unchanged.

[0076] Examples of some suitable binary masks are given in FIGS. 5a-5f, which all display the masked signal (designated with a line) with the masked regions denoted by shading for two channels of ECG data. FIGS. 5c and 5e are masked using the mask in FIG. 5a, while FIGS. 5d and 5f are masked using the mask in FIG. 5b. The mask in FIG. 5a represents a mask generated such that each masked section has a length that follows a geometric distribution followed by an unmasked section that also follows a geometric distribution. The masked sections are chosen independently for each channel. The goal of this process is to create longer temporally masked section compared to a uniformly random mask. The mask in FIG. 5b was also created using masked and unmasked segments that follow a geometric series, however, that sequence is shared across all input channels. This would be done to prevent the transformer model 20 from learning to input the missing values from adjacent channels. During pretraining, these masks will be generated independently for each new batch of data, and multiple mask generation strategies may be employed during training.

[0077] After the input signal (i.e. , patient data 26) has been masked, it is passed to the transformer model 20 which produces a multi-variate time series with the same sequence length and channel count (for example, 12 ECG channels), with the goal of recreating the masked values. The transformer model 20 is then scored by how accurately the masked points were recovered. Specifically, the mean squared error between the imputed points (set to zero in the input signal) and the true input signal is calculated. The transformer model 20 is not scored by the sections of the input signal that have a mask value of one, but since the mask are sampled randomly during training, the transformer model 20 needs to be able reproduce the input at all points (not just masked locations).

[0078] In an implemented test, the transformer model(s) 20, as outlined above, were built in Python using PyTorch. The transformer model(s) 20 were trained using the training set of PTB-XL data and evaluated on the test set of PTB-XL. The transformer mode(s)l 20 were successfully able to learn to impute missing values of ECG time series data in an unsupervised manner, demonstrating a successful pretraining. In FIGS. 6a-6d,examples of input signals, and imputed values are given. Each row in FIGS. 6a-6d represents a separate channel from the 12 lead ECG (4 rows shown for brevity). The lines represent the input signal before masking and dots indicate the locations in the signal that were masked and therefore required the transformer model 20 to impute those values. The transformer model 20 is successfully able to recover the masked portions of the input signal with high fidelity, indicating the unsupervised pretraining was successful, and the model can interpret ECG data. With the transformer model 20 able to learn to understand the structure of the time series data, it opens the possibility for incorporation of large amounts of unsupervised data to improve performance on the limited LSI dataset.

[0079] Turning to LEVEL 3, In addition to the LEVEL 2 physiological state estimations, contextual information can be integrated in the CRITIC process to further enhance LSI predictions. LEVEL 3 focuses on incorporating relevant information such as mechanisms of injury, time since injury, patient demographics, comorbidities, medical procedures, lab results, and diagnoses. These factors may affect physiological signatures or the probability of requiring a future LSI. Incorporating this information is nontrivial because health data often has complex properties including heterogeneity, sparsity, incompleteness, and temporal dependencies. However, the disclosed approaches are specifically designed to address these challenges and to integrate seamlessly with existing infrastructure. LEVEL 3 is implemented to embed contextual information in features to be fused with physiological state estimations for enhanced LSI predictions.

[0080] The foregoing analysis of time series of data from sensors 1006, 1008, 1010 acquired by the UAVs 1002 and / or UVGs 1004, by vital sign sensors placed on MCI casualties when reached and extracted, and so forth can be performed without contextual information about the casualties to provide time-critical LSI recommendations at the MCI scene. This is advantageous since at many MCI scenes the identities of casualties may not be easily determined, and moreover seriously injured casualties may need immediate LSI. However, once a casualty has been identified, patient context information can optionally augment the CRITIC analysis to improve accuracy of LSI recommendations. To create contextual features, one or more deep learning models 16 that extract information from electronic health records (EHRs) can be adapted. To process information with little to no temporal dependencies such as patient demographics,mechanism of injury, and comorbidities, an unsupervised deep learning approach known as Deep Patient Representation. Deep Patient Representation uses a stack of denoising autoencoders to generate representations from EHRs. The autoencoders use a masking noise process during training to blank a percentage of the inputs, thereby simulating missing components in EHRs. During evaluation, if some contextual data streams are unavailable or inaccessible due to situational constraints, these inputs will either be estimated using population values consistent with the MCI or blanked as during training. The latent representation in the final autoencoder will be used in LEVEL 4 to incorporate static context. Experiments have shown that classifiers trained using the Deep Patient Representation outperform alternative feature extraction methods such as principal component analysis.

[0081] If contextual information contains detailed time logs with strong temporal relationships, Temporal Tree Representations (TTR) may also be used. This is relevant when coalescing information from correlated events such as medication administration and resulting changes in physiological signatures. TTR will be particularly useful when encoding a series of LSIs because prior interventions may be indicative of, or may otherwise impact, future interventions. The CRITIC process first constructs hierarchical representations comprising decision trees from temporal sequences of medical events. The tree(s) is constructed using only available medical information and thus handles irregular and sparse data. After the tree is constructed, nodes are traversed, and resulting sequences are treated as a paragraph to be encoded using established embedding techniques. The processes used to construct and traverse the trees are flexible and will be tuned to maximize model performance. Temporal trees may be constructed or updated as new context becomes available and thus are appropriate for handling dynamic context. The embeddings generated from the processed temporal trees will be incorporated in LEVEL 4. Deep Patient Representations and Temporal Tree Representations are only two examples of contextual feature extraction and other methods are possible. In another nonlimiting illustrative embodiment, numerical values from electronic health records such as lab values are standardized, and categorical data such as race, sex, or medical classifications is encoded using one-hot encoding, then passed into the fusion model.

[0082] In LEVEL 4, LSI assessment features learned from physiological data and contextual features (LEVEL 2) are fused using the fusion model 22 of the trained NN 18. The fusion model 22 overcomes two of the technical challenges inherent with the LSI data. First, the fusion model 22 is invariant to input length, allowing for the addition and removal of sensors on the fly without the need for retraining. Second, the fusion model 22 is robust to missing sensor data, such as data flagged by SPC described at LEVEL 1. The resulting fusion model 22 provides a robust and flexible way to fuse features created in LEVEL 2.

[0083] Some common approaches to feature fusion are concatenation and addition. In feature concatenation, the features from each sensor would be concatenated into a single large feature. A drawback of feature concatenation is an expansion in feature size for each new sensor, which could cause computational constraints as the number of sensors increases. In feature addition, the features are simply summed together to create a single feature with the same size as the original, with the drawback that each feature has equal weight, and the system may not learn interactions between features. Although feature addition may allow for flexible addition and removal of sensors from an architecture standpoint, such a system would most likely need to be retrained after the addition or removal of a sensor.

[0084] To overcome the limitations of feature addition and feature concatenation, in the illustrative CRITIC system described herein the fusion model 22 is used to learn contextual information from the sensor features. For example, the fusion model 22 can include a self-attention mechanism. When applied to natural language processing, selfattention extracts meaningful relationships between various words in a sentence. In the CRITIC system described herein, self-attention is utilized for the different purpose of combining features. Each extracted sensor feature can be considered a word, and the collection of sensor features can be considered as a sentence. Therefore, the fusion model 22 learns the contextual relationships between features from the training data. For example, an increase in heart rate by itself may not be a problem, but when combined with other changes in vital signs, could be cause for recommending an LSI. Since selfattention layers of the fusion model 22 handle variable length input sequences, the fusion model 22 can advantageously generalize across different sensor combinations withoutthe need for retraining. To accomplish this, the LEVEL 2 sensor features that get passed to the fusion model 22 during training can be used to simulate various sensor conditions, further improving the robustness of the fusion model 22 under varying sensor configurations.

[0085] FIG. 7 shows another embodiment of the fusion model 22. A classification token 32 can be prepended to a sensor and contextual feature vector. This classification token 32 can give the self-attention feature fusion layer (which corresponds to the fusion model 22) a mechanism to combine the contextual information (i.e. , the patient data 26) from the current set of sensors. The state at the output of the self-attention feature fusion model 22 for this classification token 32 acts as a representation for the need for LSI, similar to the “[class]” token used in a BERT language model. After the self-attention feature fusion layer 22 processes the input features, the output will be a set of embeddings 34 with the same shape as the input. The classification token 32 embedding can be passed through a prediction head (corresponding to the output model 24) to produce a single output variable (i.e., the output 30), the other output embeddings are not used for prediction. The single output value 30 is then used to predict the time until an LSI is required.

[0086] Since the self-attention feature fusion layer 22 does not encode positional information, an additive positional encoding vector may optionally be added to the inputs of the transformer model 20. In natural language processing, positional encoding is normally used to encode information such as word ordering in a sentence. However, in the case of the disclosed CRITIC process, there is no positional information in the sensors, i.e., there is no inherent ordering of the sensors. Hence, in some embodiments, this additive encoding vector is removed. However, in other embodiments as disclosed herein, additional information specific to the prediction of the need for LSI is encoded with this additive encoding. For example, sensor information, such as sensor type or sensor state, can be incorporated into the transformer model 20 in lieu of positional encoding. Another option is for the transformer model 20 to learn the encoding from the data.

[0087] There are also advantages of using SAFF when sensor signals are unreliable. To increase the robustness when there is missing or corrupted sensor data, sensor data can be selectivity dropped out and replaced with a token 32 that represents missing dataduring training. This token 32 can be a learnable parameter of SAFF, allowing the transformer model 20 to optimize the location of this token 32 in LEVEL 2. When the SPC process detects an anomaly, the features for that sensor can be replaced with a token 32 that the self-attention feature fusion layer 24 can recognize as missing information and ignore the missing data. This will allow for a highly adaptable and robust self-attention feature fusion layer 24 when sensor data is unreliable or corrupted.

[0088] In LEVEL 4, to provide useful predictions for the need for LSI, the SAFF process can predict the need for LSI from data in the preceding minutes. There are many ways to formulate this problem, such as predicting LSI at a fixed interval in the future, sequence-to-sequence prediction, and regression. Regressing on the time to LSI has clinical benefits such as informing clinical staff of the approximate time the LSI will occur. However, it is unclear whether LSI indicating biological signatures will present in the data at longer time scales (e.g., 12 hours before an LSI). In these cases, it may be hard for the fusion model 24 to distinguish between a time to LSI of 6 hours vs. 8 hours, potentially leading to issues during training such as overfitting.

[0089] To combat these issues, in some embodiments disclosed herein the time to LSI is mapped to a number between 0 and 1 using a modified sigmoid function. An example of this function of the time to LSI in minutes is given in FIG. 8. This mapping compresses the potentially large time scale (0-24 hours) into the range between 0-1 . Large time values are effectively set to zero, indicating there is a low probability of LSI occurring (indicated in the bottom portion of the graph). At around 15 minutes until the need for LSI, the curve in FIG. 8 begins to increase (indicated in the middle portion of the graph). In this zone, the need for LSI is likely and resources for LSI should be allocated. Finally, at around 8 minutes to LSI, the need for LSI is imminent and resources for LSI should be deployed (indicated in the top portion of the graph). These times are adjustable in the formulation of the mapping function, and the values chosen here are for example only. An actual timeline of events and key points on the timeline can be determined by the need of the clinicians as well as the capabilities of the fusion model 24.

[0090] One advantage of the fusion model 24 is the smoothly increasing nature of the output value 30, allowing the fusion model 24 to gradually increase the output 30 as the need for LSI grows closer. In contrast, if this were a classification problem (e.g., predict ifan LSI is required in the next five minutes), there will always be a discontinuity around the cutoff time. In other words, at 5 minutes 1 second, to LSI the need for LSI would be 0 and at 5 minutes to LSI the need for LSI would be 1 . This could potentially be problematic for the fusion model 24 to learn, if the same biosignatures are present before and after the 5-minute cutoff. The fusion model 24 avoids this problem while also containing adjustable parameters to how and where the output value changes with respect to the time to LSI. In addition, the sigmoid function is differentiable, allowing for easy integration into existing machine learning tools.

[0091] This fusion model 24 can also be extended to classify the type of LSI that is occurring through a process known as multitask learning. In multitask learning, a single machine learning model is developed to perform multiple tasks. In the case of the disclosed CRITIC process, the fusion model 24 would predict the need for LSI metric as well as categorize the type of LSI. The fusion model 24 would share the same architecture for both tasks up until LEVEL 4. The LEVEL 4 architecture would be modified by adding an additional branch to the network dedicated to classifying LSI, while the original branch at LEVEL 4 would remain to predict the need for LSI metric. The shared portions of the fusion model 24 are generalized, to facilitate work on multiple tasks, leading to increased robustness of the fusion model 24.

[0092] LEVEL 5 includes a process refinement process responsible for performance monitoring and enhancement. At LEVEL 5, performance of the fusion model 24 will be used to update and refine LEVELS 1-4. This feedback stage will allow for identification of model weaknesses, exploration of CRITIC generated features, and increased understanding into how the disclosed CRITIC process makes decisions. This framework will allow for rapid development and refinement for seamless translation to field use.

[0093] In some contemplated embodiments, techniques can be employed to increase the explainability of the CRITIC process to identify and refine physiological signatures of impending need for LSI. To understand the features extracted from each sensor, dimensionality reduction can be applied to the extracted feature for each sensor. The low dimensional representation will be used to produce visualizations of the feature spaces for each sensor. These visualizations will be useful for discovery of physiological states based on senor input. In addition to the dimensionality reduction, clustering techniquessuch as hdbscan can be applied to cluster the sensor features, which will allow us to identify similarities in physiological states. These methods facilitate exploration of the learned feature space of each sensor and gain insight into the which physiological states occur prior to an LSI.

[0094] The trained NN 18 can also allow for further explainability of the LSI predictions through examination of so-called query and key pairs for the sensor features. The query and key pairs are generated from the input features and are analogous to information retrieval, such as querying a search. With examination of the product of query and key pairs in the SAFF, contextual insights into the features can be gained. For example, a large query-key pair product value would represent an important contextual relationship between those two sensors at that time. FIGS 9a, 9b, 9c, and 9d show the attention mechanism of a BERT language model examined using BertViz. In each of FIGS 9a, 9b, 9c, and 9d the attention of one token to another is visualized using lines, with shading representing different attention heads and line thickness representing the attention value. FIGS. 9a and 9c display connections between all tokens and FIGS. 9b and 9d display connections between input tokens and the “[CLS]” token (note the “[CLS]” token is used by BERT to make predictions). By visualizing the attention of the “[CLS]” token, insights into what the self-attention layer was focusing on when making predictions can be gained. For instance, in FIG. 9b, the attention head is focusing on tokens dog, cat and rug. In contrast, in FIG. 9d the attention of the “[CLS]” token is focusing on what the dog and cat are doing, e.g., “sat on the mat.” Linking the contextual information with the lowdimensional visualization, clustering information, and LSI events, will increase the interpretability of the proposed model. These analyses will be used to quantify similarities in feature embedding spaces to identify new relationships between injuries, physiological processes, and ultimately, health deterioration and need for LSI, while also offering insight into the inner workings of the SAFF process. This information can be used to identify new signatures and refine existing feature extraction and fusion processes to improve LSI predictions.

[0095] Finally, this formalized process refinement level will allow for clear diagnostics and effective communication to improve data acquisition and model interpretability, effectively closing the loop between model development and refinement. Using thisclosed-loop system, unique physiological states can be identified from the data and use them to gain insights into the need for LSI.

[0096] Referring back to FIG. 2, in the operation 102 the received patient physiological data is filtered using SPC, which monitors for signal disruptions in real time in the received patient physiological data. LEVEL 1 shown in FIG. 4 employs SPC to detect and clean (e.g., remove) disrupted signals from input sensors. The SPC process continuously monitors for signal disruptions in real time. The SAFF process advantageously enables sensors to be added or removed in real time in response to SPC process detection of sensor data unreliability, without retraining or changing the model architecture. The SPC framework detects errant data near the source and prevents it from propagating through later processing stages and interfering with LSI predictions.

[0097] With reference now to FIG. 10, a flow chart is shown diagrammatically depicting a patient condition prediction method, in accordance with some further embodiments of the present disclosure. The flow chart of FIG. 10 is similar to that of FIG. 2. However, the SPC filtering operation 102 of the embodiment of FIG. 2 is omitted in favor of a different approach for incorporating SPC into the LSI recommendation method. As diagrammatically shown in FIG. 10, in this embodiment the SPC is applied to the physiological data streams 101 in an operation 200 to generate SPC quality metric features 202 for the respective physiological data streams.

[0098] Furthermore, the operation 107 of FIG. 2 which combines the extracted numerical values of the patient physiological features and the extracted numerical values of the patient physiological features is replaced in the process of FIG. 10 by a modified operation 207 that combines the extracted numerical values of the patient physiological features, the extracted numerical values of the patient physiological features, and the SPC quality metric features 202 for the respective physiological data streams. The operation 207 is suitably implemented similarly to the operation 107, using a self-attention feature fusion (SAFF) process (e.g., implemented using the trained NN 18 that is modified for the embodiment of FIG. 10 to have further input nodes for the SPC quality metric features 202 for the respective physiological data streams) that combines physiological and contextual features, as well as the SPC quality metric features 202 for the respective physiological data streams, for LSI prediction. In a suitable SAFF process, a multi-headedscaled dot-product attention is used to compute a representation of the input features (here additionally including the SPC quality metric features 202). The combination of the features (i.e. , the “fusion” aspect) advantageously combines information from the different sources (here including the SPC quality metric features 202 produced by the SPC analysis 200) to provide an accurate LSI prediction.

[0099] The prediction model 208 of FIG. 10 corresponds to the operation 108 of FIG. 2, and predicts a condition of a patient requiring LSI based on the combination of extracted features produced by combining operation 207. To do so, again with brief reference to FIG. 3, the trained NN 18 includes the output model 24 configured to output a predicted condition of a patient based on the combination of extracted features.

[0100] An advantage of the patient condition prediction method of FIG. 10 is that signals which may be compromised by signal disruptions or the like are not discarded completely (as per the filtering operation 102 of the embodiment of FIG. 2), but instead are still input to the SAFF LSI recommendation modeling 207, 208 along with input of the SPC quality metric features 202 for the respective physiological data streams. This advantageously allows the SAFF LSI recommendation modeling 207, 208 to continue to utilize the potentially compromised data streams, while accounting for their reduced reliability by way of the SPC quality metric features 202 received in parallel.

[0101] In the illustrative example of FIG. 10, the combining 207 of the extracted features and the SPC quality metric features for the data streams (and optionally also the contextual features) and the predicting 208 of a requirement for an LSI based on the combination of the extracted features and the SPC quality metric features for the data streams (and optionally also the contextual features) is performed using a self-attention feature fusion (SAFF) process. However, the LSI recommendation model 207, 208 could more generally be implemented using another type of predictive model, such as another type of neural network or another type of machine learning (ML) algorithm.

[0102] This approach of FIG. 10 is premised on the recognition herein that usable signal information for LSI prediction may still be extracted during signal disruptions. Physiological events may be correlated or even cause apparent disruptions and thus it may be undesirable to drop or mask features (i.e., remove) data streams during these periods of disruption. A further benefit of the approach of FIG. 10 is that assumptions donot need to be made about the best way to modify individual features during disruptions. Instead, the model learns the relationship between the SPC quality metric features 202 and the signal features 103. This provides an extensible and efficient approach compared to attempting to create complex logic for each type of feature and disruption. The approach of FIG. 10 may be of particular value in scenarios in which pristine physiological data streams may be unavailable or only sporadically available (e.g., at an MCI scene where data streams are coming from UAVs 1002, or downstream at an emergency room of a hospital that is inundated with casualties transported from the MCI scene.

[0103] As previously described, in LEVEL 1 (see FIG. 4), the SPC process of FIG. 2 can be trained by periodically dropping sensor data to mimic signal disruptions, to provide simulated signal disruption training examples for training graceful handling of signal disruptions within CRITIC. For the variant embodiment of FIG. 10, it may be beneficial to implement less drastic signal disruption during the training, such as by randomly introducing noise into training data streams, implementing very short signal drops (e.g., on microsecond or millisecond time frames) to simulate spurious signals, and / or so forth. The signal disruptions introduced during training may suitably simulate the type(s) of signal disruptions expected in a real-world situation (e.g., at a real-world MCI scene).

[0104] With continuing reference now to FIG. 10, a further difference of the patient condition prediction method of FIG. 10 compared with that of FIG. 2 is that the contextual features extraction 106 of FIG. 2 is replaced by an operation 206 in the embodiment of FIG. 10 in which contextual features are extracted from the filtered data more particularly using a large language model (LLM). Hence, in this embodiment the deep learning model 16 that extracts numerical values from the filtered patient contextual data particularly comprises an LLM 216, as diagrammatically shown in FIG. 10. The LLM 216 is suitably pretrained to embed an unstructured text space into an informative feature space useful for predicting the need for LSIs. The LLM 216 may be trained on relevant medical text, for example extracted from an EHR, to enhance the embedding. The text may be open- ended or preprocessed in a standard structure to describe manual logs reported by emergency response medical personnel in transit to the hospital or at the hospital, for example. To ensure a large training dataset, the LLM 216 may be trained on training EHR data from multiple institutions and logs. As new EHR data are recorded, it can be fed intothe LLM 216 to track the patient state. Once the EHR information is embedded, this information may be used (in conjunction with other physiological features or as a standalone feature embedding) with classification or regression models to predict the need for LSIs, e.g. using the SAFF LSI recommendation modeling 207, 208 as previously described.

[0105] An advantage of using the LLM 216 to embed EHR data or other contextual data is that the LLM 216 maps useful / informative, but unstructured and / or inconsistent, data to a consistent feature space. This consistent feature embedding can be input to the SAFF LSI recommendation modeling 207, 208 to make predictions, The LLM 216 advantageously can process incoming contextual data 104 in real-time, and provides standardization across hospitals or other medical institutions.

[0106] A further advantage of the embodiment of FIG. 10 which integrates the contextual features extraction 206 using the LLM 216 with the illustrative SAFF model 207, 208 is that the self-attention aspect of the SAFF model 207, 208 strengthens meaningful connections in the unstructured EHR medical text extracted by operation 206) and the real-time physiological signals whose features are extracted in operation 103 to facilitate making accurate LSI predictions.

[0107] In one nonlimiting illustrative example, the LLM 216 may be a Bidirectional Encoder Representations from Transformers (BERT) large language model employing an encoder-only transformer architecture. However, it is contemplated to employ other types of large language model as the LLM 216 used in the contextual features extraction 206, such as a Text-to-Text Transfer Transformer (T5) model, a Language Model for Dialogue Applications (LaMDA) model, a Pathways Language Model (PaLM), or so forth.

[0108] The present disclosure has been described with reference to several different embodiments. Obviously, modifications and alterations will occur to others upon reading and understanding the preceding detailed description. It is intended that the present disclosure be construed as including all such modifications and alterations insofar as they come within the scope of the appended claims or the equivalents thereof.

Claims

CLAIMS:1 . A triage method for triaging a scene of a mass casualty incident, the triage method comprising: acquiring a plurality of data streams pertaining to a casualty at the scene, the acquiring of at least one data stream using one or more sensors of at least one unmanned aerial vehicle (UAV) and / or unmanned ground vehicle (UGV); extracting features from the data streams; applying statistical process control (SPC) to the data streams to generate SPC quality metric features for the data streams, the SPC quality metric features being indicative of signal quality of the data streams; combining the extracted features and the SPC quality metric features for the data streams; and predicting a requirement for a life-saving intervention (LSI) for the casualty based on the combination of the extracted features and the SPC quality metric features for the data streams.

2. The triage method of claim 1 , wherein the combining and the predicting are performed using a self-attention feature fusion (SAFF) process.

3. The triage method of claim 1 , further comprising: receiving context information for the casualty based on an identify of the casualty; and extracting further features from the context information using a large language model (LLM); wherein the combining combines the extracted features, the SPC quality metric features for the data streams, and the extracted further features.

4. The triage method of claim 1 , wherein receiving context information for the casualty based on an identify of the casualty comprises: receiving patient physiological data from one or more sensors associated with the casualty; andextracting patient physiological features from the received patient physiological data.

5. The triage method of claim 1 , wherein combining the extracted features and the SPC quality metric features for the data streams comprises: inputting the extracted features and the SPC quality metric features for the data streams to a trained neural network (NN); and outputting, with the trained NN, a prediction of the condition of the patient.

6. A non-transitory computer readable medium storing instructions executable by at least one electronic processor to perform a patient condition assessment method for a patient, the method comprising: extracting features from patient data streams acquired for the patient; receiving patient contextual data from one or more data records associated with the patient; extracting patient contextual features from the received patient contextual data using a language learning model (LLM); combining the extracted features and the extracted patient contextual features; and predicting a requirement for a life-saving intervention (LSI) for the patient based on the combination of the extracted features and the extracted patient contextual features.

7. The non-transitory computer readable medium of claim 6, wherein the LLM comprises a Bidirectional Encoder Representations from Transformers (BERT) large language model employing an encoder-only transformer architecture, a Text-to-Text Transfer Transformer (T5) model, a Language Model for Dialogue Applications (LaMDA) model, or a Pathways Language Model (PaLM),8. The non-transitory computer readable medium of claim 6, wherein the combining and the predicting are performed using a self-attention feature fusion (SAFF) process.

9. The non-transitory computer readable medium of claim 6, wherein the combining and the predicting are performed using at least one neural network.

10. The non-transitory computer readable medium of claim 6, wherein the method further comprises: applying statistical process control (SPC) to the patient data streams to generate SPC quality metric features for the patient data streams, the SPC quality metric features being indicative of signal quality of the patient data streams; wherein the combining comprises combining the extracted features and the extracted patient contextual features and the SPC quality metric features for the patient data streams, and the predicting is based on the combination of the extracted features, the extracted patient contextual features, and the SPC quality metric features for the patient data streams.11 . The non-transitory computer readable medium of claim 6, wherein combining the extracted features and the SPC quality metric features for the patient data streams comprises: inputting the extracted features and the SPC quality metric features for the patient data streams to a trained neural network (NN); and the predicting includes outputting, with the trained NN, the prediction of the requirement for an LSI.

12. The non-transitory computer readable medium of claim 6, wherein method further comprises: reviewing the predicted requirement for an LSI by comparing a condition of the patient prior to the prediction and a condition of the patient after the LSI is provided to the patient based on the predicted requirement for the LSI.

13. An apparatus comprising:a computer programmed to perform a life-saving intervention (LSI) recommendation process including: receiving physiological data streams acquired from a patient; extracting features from the physiological data streams; applying statistical process control (SPC) to the physiological data streams to generate SPC quality metric features for the physiological data streams, the SPC quality metric features being indicative of signal quality of the physiological data streams; combining the extracted features and the SPC quality metric features for the physiological data streams; and predicting a condition of the patient requiring an LSI based on the combination of the extracted features and the SPC quality metric features for the physiological data streams.

14. The system of claim 13, wherein the combining and the predicting are performed using a self-attention feature fusion (SAFF) process.

15. The system of claim 13, wherein the combining and the predicting are performed using at least one neural network.

16. The system of claim 15, wherein the computer is further programmed to train the at least one neural network using training data with randomly introduced noise and / or signal drops simulating spurious signals.

17. The system of claim 15, further comprising: at least one unmanned aerial vehicle (UAV) and / or unmanned ground vehicle (UGV) including one or more sensors configured to acquire the physiological data streams from the patient; wherein the receiving comprises receiving the physiological data streams via wireless transmission of the physiological data streams from the at least one UAV and / or UGV to the computer.

18. The system of claim 17, wherein the at least one UAV and / or UGV is further configured to deliver the LSI to the patient in response to the computer predicting the condition of the patient requiring the LSI.

19. The system of claim 13, further comprising: patient sensors operatively connected to acquire the physiological data streams from the patient; wherein the receiving comprises receiving the physiological data streams from the patient sensors.

20. The system of claim 13, wherein the LSI recommendation process further includes: receiving context information for the patient; and extracting further features from the context information using a large language model (LLM); wherein the combining combines the extracted features, the SPC quality metric features for the data streams, and the extracted further features.