Methods and systems for detecting vomit-related events

A machine learning-based method for detecting vomit-related events through acoustic analysis addresses underreporting issues, offering precise symptom tracking and improved treatment strategies.

WO2025238616A1PCT designated stage Publication Date: 2025-11-20TAKEDA PHARMA CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/055141
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-17
Filing Date
2025-05-16
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

Chronic vomiting or retching symptoms, caused by various gastrointestinal and non-gastrointestinal conditions, significantly affect patients' quality of life and are often underreported due to patient recall errors, leading to inaccurate symptom severity assessment and treatment decisions.

Method used

A computer-implemented method using machine learning models to analyze acoustic frames captured by user devices, segmenting and classifying audible sounds into vomit-related or non-vomit-related events, with optional confirmation by patients and integration of non-audio features for enhanced accuracy.

Benefits of technology

Provides objective and accurate detection of vomit-related events, reducing patient self-reporting ambiguity and enabling better treatment adjustments and public health monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025055141_20112025_PF_FP_ABST
    Figure IB2025055141_20112025_PF_FP_ABST
Patent Text Reader

Abstract

A method (300) includes receiving a sequence of acoustic frames (122) and segmenting the sequence of acoustic frames into a plurality of events of interest characterizing one or more audible sounds (106). Each event of interest is associated with a respective subset of sequential acoustic frames (122) segmented from the sequence of acoustic frames (122). For corresponding event of interest, the method also includes: processing, using a machine learning (ML) model (200), the respective subset of sequential acoustic frames (122) to generate a corresponding probability distribution over possible audible sound classifications (202); and labeling, using a labeler (130), based on the corresponding probability distribution over possible audible sound classifications, the particular event of interest as a vomit-related event or a non- vomit-related event. The method also includes determining quantitative bout information indicating a number of the plurality of events of interest that are labeled as vomit-related events.
Need to check novelty before this filing date? Find Prior Art

Description

Methods And Systems For Detecting Vomit- Related EventsTECHNICAL FIELD

[0001] This disclosure relates to methods and systems for detecting vomit-related events.BACKGROUND

[0002] Chronic, persistent, or recurring vomiting or retching can be debilitating and distressing symptoms that can significantly affect a patient’s quality of life. There are several underlying gastrointestinal (GI) causes (e.g., gastroparesis, and cyclic vomiting syndrome (CVS)) and non-GI causes (e.g., medications, motion sickness, pregnancy, vestibular disorders, neurological disorders, hyperemesis gravidarum, chronic nausea and vomiting, and bulimia) that can cause vomiting and retching. Frequent or repeated episodes of vomiting or retching can result in, for example, undernutrition, weight loss, metabolic abnormalities, frequent hospitalization, emergency department visits, and higher healthcare service utilization.SUMMARY

[0003] One aspect of the disclosure provides a computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations that include receiving a sequence of acoustic frames captured by a user device associated with a patient and segmenting the sequence of acoustic frames into a plurality of events of interest characterizing one or more audible sounds, each event of interest associated with a respective subset of sequential acoustic frames segmented from the sequence of acoustic frames. For each particular event of interest of the plurality of events of interest, the operations also include: processing, using a machine learning (ML) model, the respective subset of sequential acoustic frames associated with the particular event of interest to generate a corresponding probability distribution over possible audible sound classifications for the particular event of interest; and labeling, using a labeler, based on the corresponding probability distribution over possible audible sound classifications, the particular event of interest as a vomit-related event or a non-vomit-related event. The operations also include determining, based on the labeled events of interest, quantitative bout information indicating a number of the plurality of events of interest segmented from the sequence of acoustic frames that are labeled as vomit-related events.

[0004] Implementations of the disclosure may include one or more of the following optional features. In some implementations, the quantitative bout information further indicates a respective number of occurrences of each of the possible audible sound classifications. The operations may also include determining, based on the labeled events of interest, quantitative event information indicating occurrences of vomit-related events associated with a health condition of the patient.

[0005] In some examples, labeling the particular event of interest as a vomit-related event or a non- vomit-related event is further based on the corresponding probability distribution over possible audible sound classifications determined for each of one or more events of interest of the plurality of events of interest that are temporally near to the particular event of interest. In these examples, the operations may also include between each temporally near pair of events of interest segmented from the sequence of acoustic frames , determining a corresponding time duration separating a last acoustic frame of the respective subset of sequential acoustic frames associated with an earlier event of interest in the adjacent pair and an initial acoustic frame of the respective subset of sequential acoustic frames associated with a later event of interest in the adjacent pair. Here, labeling the particular event of interest as a vomit-related event or a non-vomit-related event may be further based on the corresponding time duration between the particular event of interest and each of the one or more events of interest of the plurality of events of interest that are temporally near to the particular event of interest.

[0006] In some implementations, the operations further include, when the particular event of interest is labeled as a vomit-related event transmitting the respective subset of sequential acoustic frames associated with the particular event of interest to a remote computing device, wherein the remote computing device is configured to: process, using a second ML model different from the ML model, the respective subset of sequential acoustic frames to generate a vomit-related event classification score for the particularevent of interest; and confirm or reject the labeling of the particular event of interest as a vomit-related event based on the vomit-related event classification score. In these implementations, the operations may further include, when the remote computing device rejects the labeling of the particular event of interest as a vomit-related event, re-labeling the particular event of interest as a non-vomit-related event. Additionally or alternatively, the operations may also include adapting one or more parameters of the second ML model based on the vomit-related classification score.

[0007] In some examples, the operations further include, when the particular event of interest is labeled as a vomit-related event: transmitting the respective subset of sequential acoustic frames associated with the particular event of interest to a remote computing device, wherein the remote computing device is configured to: playback the respective subset of sequential acoustic frames associated with the particular event of interest to a user; and receive, from the user, an indication indicating a confirmation or a rejection of the labeling of the particular event of interest. In these examples, the operations may also include adapting one or more parameters of the ML model based on the indication received from the user. The operations may also include, prior to transmitting the respective subset of sequential acoustic frames to the remote computing device, modifying the respective subset of sequential acoustic frames to at least one of mask an identity of the patient or mask spoken words.

[0008] In some implementations, the operations further include, when the particular event of interest is labeled as a vomit-related event: outputting, from the user device, a prompt prompting the patient to confirm or reject the labeling of the particular event of interest as a vomit-related event; receiving an indication from the patient indicating the patient rejects the labeling of the particular event of interest as a vomit-related event; and based on receiving the indication from the patient indicating the patient rejects the labeling of the particular event of interest as a vomit-related event, re-labeling the particular event of interest as a non-vomit-related event. In these implementations, the operations may also include, based on receiving the indication from the patient indicating the patient rejects the labeling of the particular event of interest as a vomit-related event,adapting one or more parameters of the ML model based on the acoustic frames (122) of the particular event of interest and the indication.

[0009] The data processing hardware may reside on the user device or on a remote computing device in communication with the user device. The possible audible sound classifications may include one or more of a retch, an expulsion, a cough, a spit, a sneeze, speech, laughter, a throat clearing, a splash, water running, an alarm, or a toilet flushing.

[0010] The operations may further include receiving non-audio features associated with the patient and captured by the user device within a threshold period of time from when the user device captured the sequence of acoustic frames. Here, the non-audio features represent at least one of motion of the patient or a skin conductance of the patient. Here, labeling the particular event of interest as a vomit-related event or a nonvomit-related event may be further based on the non-audio features associated with the patient.

[0011] The operations may further include, for each particular event of interest of the plurality of events of interest: processing, using a speaker identification model, the respective subset of sequential acoustic frames associated with the particular event of interest to identify the patient as a source of the particular event of interest. Here, labeling the particular event of interest as a vomit-related event or a non-vomit-related event is further based on identifying the patient as the source of the particular event of interest.

[0012] In some examples, labeling the particular event of interest as a vomit-related event or a non-vomit-related event includes labeling the particular event of interest as a non-vomit-related event or one of multiple different types of vomit-related events. The multiple different types of vomit-related events include at least one of a retch and expulsion event, a retch-only event, an expulsion-only event, a retch and cough event, a spitting event, a coughing event, or a throat-clearing event.

[0013] In some implementations, the ML model is trained by receiving a training dataset including a plurality of training samples that each include a corresponding sequence of training acoustic frames, a corresponding ground-truth audible sound classification for each training event of interest of the training sample, and acorresponding ground-truth label for each training event of interest identifying the training event of interest of the training sample as a vomit-related event or a non-vomit- related event. The corresponding sequence of training acoustic frames includes one or more training events of interest each characterizing one or more audible sounds and associated with a respective subset of sequential training acoustic frames segmented from the sequence of training acoustic frames. For each particular training sample in the training dataset, the ML model is trained by training the ML model on the particular training sample to teach the ML model to learn how to predict the corresponding groundtruth audible sound classifications and the corresponding ground-truth labels. In these implementations, training the ML model on a particular training sample include: segmenting the corresponding sequence of training acoustic frames of the particular training sample into a plurality of training events of interest characterizing one or more audible sounds, each event of interest associated with a respective subset of sequential training acoustic frames segmented from the sequence of training acoustic frames; for each particular training event of interest of the plurality of training events of interest: processing, using the ML model, the respective subset of sequential training acoustic frames associated with the particular training event of interest to generate a corresponding probability distribution over possible audible sound classifications for the particular training event of interest; labeling, using a labeler, based on the corresponding probability distribution over possible audible sound classifications, the particular training event of interest as a vomit-related event or a non-vomit-related event; determining a loss based on one or more of: a probability of the probability distribution corresponding to the corresponding ground-truth audible sound classification; or a difference between the corresponding ground-truth label and the label of the particular training event of interest; and updating a plurality of weights of the ML model based on the determined loss.

[0014] In some examples, the ML model includes a joint audio and speech architecture that includes a pre-trained audio encoder, a pre-trained speech decoder, and a pre-trained large language model (LLM) adapted by a fine-tuned adapter module. In these examples, the joint audio and speech architecture may be trained by: receiving a training dataset including a plurality of training samples that each include: acorresponding sequence of training acoustic frames, the corresponding sequence of training acoustic frames including one or more training events of interest, each training event of interest characterizing one or more audible sounds and associated with a respective subset of sequential acoustic frames segmented from the sequence of training acoustic frames; a corresponding training natural language prompt; a corresponding ground-truth audible sound classification for each training event of interest of the training sample; and a corresponding ground-truth label for each training event of interest identifying the training event of interest of the training sample as a vomit-related event or a non-vomit-related event; and fine-tuning the adapter module on the plurality of training samples by updating a plurality of weights of the adapter module while parameters of the pre-trained audio encoder, the pre-trained speech decoder, and the pre-trained LLM are held fixed to teach the adapter module to learn how to adapt the pre-trained LLM to predict the corresponding ground-truth audible sound classifications and the corresponding ground-truth labels.

[0015] Another aspect of the disclosure provides a system that includes data processing hardware and memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations that include receiving a sequence of acoustic frames captured by a user device associated with a patient and segmenting the sequence of acoustic frames into a plurality of events of interest characterizing one or more audible sounds, each event of interest associated with a respective subset of sequential acoustic frames segmented from the sequence of acoustic frames. For each particular event of interest of the plurality of events of interest, the operations also include: processing, using a machine learning (ML) model, the respective subset of sequential acoustic frames associated with the particular event of interest to generate a corresponding probability distribution over possible audible sound classifications for the particular event of interest; and labeling, using a labeler, based on the corresponding probability distribution over possible audible sound classifications, the particular event of interest as a vomit-related event or a non-vomit-related event. The operations also include determining, based on the labeled events of interest, quantitativebout information indicating a number of the plurality of events of interest segmented from the sequence of acoustic frames that are labeled as vomit-related events.

[0016] This aspect may include one or more of the following optional features. In some implementations, the quantitative bout information further indicates a respective number of occurrences of each of the possible audible sound classifications. The operations may also include determining, based on the labeled events of interest, quantitative event information indicating occurrences of vomit-related events associated with a health condition of the patient.

[0017] In some examples, labeling the particular event of interest as a vomit-related event or a non- vomit-related event is further based on the corresponding probability distribution over possible audible sound classifications determined for each of one or more events of interest of the plurality of events of interest that are temporally near to the particular event of interest. In these examples, the operations may also include between each temporally near pair of events of interest segmented from the sequence of acoustic frames , determining a corresponding time duration separating a last acoustic frame of the respective subset of sequential acoustic frames associated with an earlier event of interest in the adjacent pair and an initial acoustic frame of the respective subset of sequential acoustic frames associated with a later event of interest in the adjacent pair. Here, labeling the particular event of interest as a vomit-related event or a non-vomit-related event may be further based on the corresponding time duration between the particular event of interest and each of the one or more events of interest of the plurality of events of interest that are temporally near to the particular event of interest.

[0018] In some implementations, the operations further include, when the particular event of interest is labeled as a vomit-related event transmitting the respective subset of sequential acoustic frames associated with the particular event of interest to a remote computing device, wherein the remote computing device is configured to: process, using a second ML model different from the ML model, the respective subset of sequential acoustic frames to generate a vomit-related event classification score for the particular event of interest; and confirm or reject the labeling of the particular event of interest as a vomit-related event based on the vomit-related event classification score. In theseimplementations, the operations may further include, when the remote computing device rejects the labeling of the particular event of interest as a vomit-related event, re-labeling the particular event of interest as a non-vomit-related event. Additionally or alternatively, the operations may also include adapting one or more parameters of the second ML model based on the vomit-related classification score.

[0019] In some examples, the operations further include, when the particular event of interest is labeled as a vomit-related event: transmitting the respective subset of sequential acoustic frames associated with the particular event of interest to a remote computing device, wherein the remote computing device is configured to: playback the respective subset of sequential acoustic frames associated with the particular event of interest to a user; and receive, from the user, an indication indicating a confirmation or a rejection of the labeling of the particular event of interest. In these examples, the operations may also include adapting one or more parameters of the ML model based on the indication received from the user. The operations may also include, prior to transmitting the respective subset of sequential acoustic frames to the remote computing device, modifying the respective subset of sequential acoustic frames to at least one of mask an identity of the patient or mask spoken words.

[0020] In some implementations, the operations further include, when the particular event of interest is labeled as a vomit-related event: outputting, from the user device, a prompt prompting the patient to confirm or reject the labeling of the particular event of interest as a vomit-related event; receiving an indication from the patient indicating the patient rejects the labeling of the particular event of interest as a vomit-related event; and based on receiving the indication from the patient indicating the patient rejects the labeling of the particular event of interest as a vomit-related event, re-labeling the particular event of interest as a non-vomit-related event. In these implementations, the operations may also include, based on receiving the indication from the patient indicating the patient rejects the labeling of the particular event of interest as a vomit-related event, adapting one or more parameters of the ML model based on the acoustic frames (122) of the particular event of interest and the indication.

[0021] The data processing hardware may reside on the user device or on a remote computing device in communication with the user device. The possible audible sound classifications may include one or more of a retch, an expulsion, a cough, a spit, a sneeze, speech, laughter, a throat clearing, a splash, water running, an alarm, or a toilet flushing.

[0022] The operations may further include receiving non-audio features associated with the patient and captured by the user device within a threshold period of time from when the user device captured the sequence of acoustic frames. Here, the non-audio features represent at least one of motion of the patient or a skin conductance of the patient. Here, labeling the particular event of interest as a vomit-related event or a nonvomit-related event may be further based on the non-audio features associated with the patient.

[0023] The operations may further include, for each particular event of interest of the plurality of events of interest: processing, using a speaker identification model, the respective subset of sequential acoustic frames associated with the particular event of interest to identify the patient as a source of the particular event of interest. Here, labeling the particular event of interest as a vomit-related event or a non-vomit-related event is further based on identifying the patient as the source of the particular event of interest.

[0024] In some examples, labeling the particular event of interest as a vomit-related event or a non-vomit-related event includes labeling the particular event of interest as a non-vomit-related event or one of multiple different types of vomit-related events. The multiple different types of vomit-related events include at least one of a retch and expulsion event, a retch-only event, an expulsion-only event, a retch and cough event, a spitting event, a coughing event, or a throat-clearing event.

[0025] In some implementations, the ML model is trained by receiving a training dataset including a plurality of training samples that each include a corresponding sequence of training acoustic frames, a corresponding ground-truth audible sound classification for each training event of interest of the training sample, and a corresponding ground-truth label for each training event of interest identifying the training event of interest of the training sample as a vomit-related event or a non-vomit-related event. The corresponding sequence of training acoustic frames includes one or more training events of interest each characterizing one or more audible sounds and associated with a respective subset of sequential training acoustic frames segmented from the sequence of training acoustic frames. For each particular training sample in the training dataset, the ML model is trained by training the ML model on the particular training sample to teach the ML model to learn how to predict the corresponding groundtruth audible sound classifications and the corresponding ground-truth labels. In these implementations, training the ML model on a particular training sample include: segmenting the corresponding sequence of training acoustic frames of the particular training sample into a plurality of training events of interest characterizing one or more audible sounds, each event of interest associated with a respective subset of sequential training acoustic frames segmented from the sequence of training acoustic frames; for each particular training event of interest of the plurality of training events of interest: processing, using the ML model, the respective subset of sequential training acoustic frames associated with the particular training event of interest to generate a corresponding probability distribution over possible audible sound classifications for the particular training event of interest; labeling, using a labeler, based on the corresponding probability distribution over possible audible sound classifications, the particular training event of interest as a vomit-related event or a non-vomit-related event; determining a loss based on one or more of: a probability of the probability distribution corresponding to the corresponding ground-truth audible sound classification; or a difference between the corresponding ground-truth label and the label of the particular training event of interest; and updating a plurality of weights of the ML model based on the determined loss.

[0026] In some examples, the ML model includes a joint audio and speech architecture that includes a pre-trained audio encoder, a pre-trained speech decoder, and a pre-trained large language model (LLM) adapted by a fine-tuned adapter module. In these examples, the joint audio and speech architecture may be trained by: receiving a training dataset including a plurality of training samples that each include: a corresponding sequence of training acoustic frames, the corresponding sequence of training acoustic frames including one or more training events of interest, each trainingevent of interest characterizing one or more audible sounds and associated with a respective subset of sequential acoustic frames segmented from the sequence of training acoustic frames; a corresponding training natural language prompt; a corresponding ground-truth audible sound classification for each training event of interest of the training sample; and a corresponding ground-truth label for each training event of interest identifying the training event of interest of the training sample as a vomit-related event or a non-vomit-related event; and fine-tuning the adapter module on the plurality of training samples by updating a plurality of weights of the adapter module while parameters of the pre-trained audio encoder, the pre-trained speech decoder, and the pre-trained LLM are held fixed to teach the adapter module to learn how to adapt the pre-trained LLM to predict the corresponding ground-truth audible sound classifications and the corresponding ground-truth labels.

[0027] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims.DESCRIPTION OF DRAWINGS

[0028] FIG. 1 is a schematic view of an example system using a vomit-related event detection system for automatically detecting vomit-related events.

[0029] FIG. 2 is a schematic view of an example machine learning (ML) model for classifying audible sounds associated with vomit-related events.

[0030] FIG. 3 is a flowchart of an example arrangement of operations for a computer- implemented method of detecting vomit-related events.

[0031] FIG. 4 is a schematic view of an example training process for training a vomit-related event detection system for automatically detecting vomit-related events.

[0032] FIG. 5 is a confusion matrix representing an example performance of a trained vomit-related event detection system.

[0033] FIG. 6 is a flowchart of an example arrangement of operations for a computer- implemented method of training a vomit-related event detection system for automatically detecting vomit-related events.

[0034] FIG. 7 is a schematic view of an example joint audio and speech architecture for classifying audible sounds associated with vomit-related events.

[0035] FIG. 8 is a schematic view of an example computing device that may be used to implement the systems and methods described herein.

[0036] Like reference symbols in the various drawings indicate like elements.DETAILED DESCRIPTION

[0037] Chronic, persistent, or recurring vomiting or retching can be debilitating and distressing symptoms that can significantly affect a patient’s quality of life. There are several underlying gastrointestinal (GI) causes (e.g., gastroparesis, and cyclic vomiting syndrome (CVS)) and non-GI causes (e.g., medications, motion sickness, pregnancy, vestibular disorders, neurological disorders, hyperemesis gravidarum, chronic nausea and vomiting, and bulimia) that can cause vomiting and retching. Frequent or repeated episodes of vomiting or retching can result in, for example, undernutrition, weight loss, metabolic abnormalities, frequent hospitalization, emergency department visits, and overall healthcare service utilization. Vomiting or retching frequency in terms of number of times, events, bouts (e.g., number of vomit-related events in a period of time), or days is often used as a measure of symptom severity to classify patients and dictate appropriate treatment. Here, a bout may be two or more vomit-related events that occur within a threshold period of time. In one example, a retch plus expulsion (R+C) bout indicaes a sequence of closely timed events, such as five retch and / or expulsion events lasting 6.2 seconds. IN some examples the threshold period time for defining a bout of two or more vomit-related events is equal to 5 seconds. In other examples, the threshold period of time for defining a bout of two or more vomit-related events is equal to 10 seconds. The threshold period of time can be set to values less than 5 seconds, less than 10 seconds, or greater than 10 seconds. Statistically, bouts typically contain up to eight (8) vomit- related events and last 0.81-25.15 seconds. Moreover, vomiting and retching frequencyare often used in clinical trials as an outcome measure to evaluate the efficacy of anti emetic treatment. Vomiting and retching are objective symptomatic events that can be quantified, recalled, and reported by patients. Accordingly, patient reported outcome (PRO) instruments are often used for reporting vomit-related events for GI and non-GI diseases and conditions that are associated with nausea, retching, or vomiting. PRO instruments and their vomiting frequency forms of question items are convenient, suitable for daily diary use, easy to follow, and understandable. However, PRO instruments may have several drawbacks, including being susceptible to recall error. For example, the definition of “vomiting time” or “vomiting episode” may be perceived as ambiguous by patients or clinical trial subjects, generate confusion, and lead to questionable accuracy or inter-subject concordance. For instance, it has been observed that patients may get confused when asked to define the “number of vomiting episodes” independently from the number of bathroom trips. The confusion becomes more evident in diseases with episodic vomiting, such as CVS, in which a bout can last hours or days and consist of many vomit-related events that lead to “brain fog” that may further effect recall. Therefore, there is a need for methods and systems for detecting vomit-related events independent of patient recall and self-reporting. Here, a vomit-related event may be one of multiple different types of vomit-related events, such as a retch and expulsion (i.e., regurgitation) event, a retch-only event, an expulsion-only event, an expulsion with one or more of splashing, flushing, coughing, spitting, or throat-clearing event, or a retch with one or more of coughing, spitting, or throat-clearing event. An expulsion may include expulsion of any type of material from a patient’s stomach, while a retch may include any attempt or effort to expel without an expulsion taking place. A retch (also sometimes referred to as a dry heave) and expulsion may be characterized by an action, such as a muscular contraction of the stomach and / or esophagus to expel some or all of the stomach’s contents.

[0038] Implementations disclosed herein provide objective measurements of vomit- related events over time, address patient self-reporting ambiguity, ensure patient compliance, and / or work well in noisy environments. Obtaining objective measurements of vomit-related events may be useful in objectively assessing a patient’s medicalcondition associated with vomit-related events. Implementations disclosed herein may also be useful in assessing the effectiveness of a therapy for a medical condition to improve treatment or improve quality of life. For example, the ability to objectively and accurately obtain vomit-related events for a subject may assist with adjusting or titrating the timing or dosages of anti-emetic medications, monitoring or tracking the side effects of treatments, or to determine other important correlations related to a subject’s medical condition and / or treatment related thereto. These correlations may include determining the number of vomit-related events in a CSV episode has decreased (e.g., a reduction from 100 events a day to 50 events a day), determining the number of days in a period of time with fewer than a threshold number of vomit-related events, determining the number of vomit-related event free days in a period of time, or identifying specific days when no vomit-related events have occurred. Implementations disclosed herein may also be useful in monitoring the impact of daily activities, such as eating with vomit-related event data, which could help inform decisions to change habits to reduce vomit-related events. Additionally, the ability monitor vomit-related events can provide patient support groups and / or healthcare professionals access to previously unavailable information for quantifying the burdens of such medication conditions, as well as assisting in the design of clinical trials for assessing side-effects associated with treatments, or the effectiveness of treatments. Disclosed implementations may also be useful for public health monitoring. For example, widespread vomit-related event detection could be used to rapidly detect outbreaks of norovirus, flu, widespread food poisoning, or similar diseases involving vomit-related events.

[0039] FIG. 1 is a schematic view of an example of a system 100. In the system 100, a patient 104 may interact with a computing device, such as a user device 10, through audible input (e.g., speech or other audible sounds) or non-audio inputs (e.g., touch input, virtual keyboard, movement, or skin conductance). The user device 10 (also referred to generally as device 10) is configured to capture sounds (e.g., streaming audio) from a patient 104 within an environment 102. The environment 102 may include, for example, a room of a home, a hospital room, a rehabilitation center, etc. Here, the streaming audio may refer to, for example, an audible sound 106 by the patient 104 that functions as aninput to a vomit-related event detection system 110 (also referred to herein as detection system 110), a spoken utterance, or an audible sound captured by the device 10. The detection system 110 processes the captured audio data to detect vomit-related events.

[0040] The device 10 may correspond to any computing device associated with a patient 104 and capable of receiving audio data. Some examples of devices 10 include, but are not limited to, mobile devices (e.g., mobile phones, tablets, laptops, etc.), computers, wearable devices (e.g., smart watches, pendants, rings, headwear, straps, glasses, hearing aids, headphones, etc.), smart appliances, Internet of things (loT) devices, vehicle infotainment systems, smart displays, smart speakers, devices adhered to patients, home security devices, etc. The device 10 includes data processing hardware 12 and memory hardware 14 in communication with the data processing hardware 12. The memory hardware 14 stores instructions that, when executed by the data processing hardware 12, cause the data processing hardware 12 to perform one or more operations, such as those disclosed herein. The device 10 further includes an audio system 16 having an audio capture device (e.g., microphones 16, 16a) for capturing and converting spoken audible sounds 106 within the environment 102 into electrical signals. While the device 10 implements a single audio capture device 16a in the example shown, the device 10 may implement an array of audio capture devices 16a without departing from the scope of the present disclosure, whereby one or more capture devices 16a in the array may not physically reside on the device 10, but be in communication with the audio system 16. The audio system 16 may also include an audio output device (e.g., a speaker 16, 16b) for communicating an audible audio signal (e.g., as output audio data from the device 10).

[0041] In the system 100, the device 10 and / or a remote computing device 70 includes an input subsystem 120 configured to receive audible sounds 106 captured by the audio capture device 16a within the environment 102, and convert the captured audible sounds 106 into a corresponding digital format associated with input acoustic frames 122 (also referred to herein as audio data 122) capable of being processed by the vomit-related event detections system 110. In some implementations, the input subsystem 120 includes a segmenter 124 for segmenting the sequence of acoustic frames 122 into a plurality of events of interest characterizing one or more audible sounds. Here,each event of interest is associated with a respective subset of sequential acoustic frames 122 segmented from the sequence of acoustic frames 122. In some examples, the segmenter 124 identifies events of interest by identifying subsets of sequential acoustic frames 122 that each have a respective magnitude that exceeds a pre- determined threshold. In some implementations, the segmenter 124 excludes events of interest containing speech to help ensure patient privacy. In some implementations, the input subsystem 120 performs audio denoising on the acoustic frames 122, or identifies acoustic frames 122 unsuitable for detecting vomit-related events (e.g., containing too much noise, or background sounds that are too loud).

[0042] The remote computing device 70 may be one or more remote servers of a distributed system executing in a cloud-computing environment in communication with the device 10 via a network 40. The remote computing device 70 includes data processing hardware 72, and memory hardware 74 in communication with the data processing hardware 72. The memory hardware 74 stores instructions that, when executed by the data processing hardware 72, cause the data processing hardware 72 to perform one or more operations, such as those disclosed herein.

[0043] The detection system 110 resides on the device 10 of the patient 104 and / or on the remote computing device 70, and implements a trained machine learning (ML) model 200 and a labeler 130. For each particular event of interest, the machine learning (ML) model 200 of the detection system 110 processes the respective subset of sequential acoustic frames associated with the particular event of interest to generate a corresponding probability distribution over possible audible sound classifications 202 for the particular event of interest (also referred to herein as a classification 202). In some implementations, the possible audible sound classifications 202 include one or more of a retch, an expulsion, a cough, a spit, a sneeze, speech, other, laughter, a throat clearing, a splash, or a toilet flushing. The number and type of possible audible sound classifications the ML model 200 is trained to predict is non-limiting.

[0044] Thereafter, the labeler 130 of the detection system 110 labels, based on the corresponding probability distribution over the possible audible sound classifications 202, each particular event of interest as a vomit-related event or a non-vomit-related event.The labeler 130 outputs, for each event of interest, an event label 132 indicating whether the event of interest is a vomit-related event or a non-vomit-related event. In some implementations, the labeler 130 includes another trained ML model (e.g., an attentionbased model) that is trained to process the corresponding probability distribution over possible audible sound classifications 202 for a particular event of interest to generate or predict the label 132 of the particular event of interest as a vomit-related event or a non- vomit-related event, and / or to highlight the acoustic features and time periods giving rise to event detection. In other implementations, the labeler 130 is a heuristic-based model that uses a set of rules for assigning labels 132 to events of interest. In some examples, the label 132 includes a confidence score indicating a likelihood that the particular event of interest includes a vomit-related event. In other examples, the label 132 includes a binary value of “1” or “0” such that a binary value of “1” indicates a vomit-related event or a non-vomit-related event and a binary value of “0” indicates the other one of the vomit-related event or the non-vomit-related event. Additionally or alternatively, the labeler 130 may use a saliency map, based on a probability distribution over possible audible sound classifications 202 for a particular event of interest, to label 132 the particular event of interest.

[0045] In some examples, the labeling of a particular event of interest as a vomit- related event or a non-vomit-related event is further based on corresponding probability distributions over possible audible sound classifications determined for one or more events of interest that are adjacent to or temporally near the particular event of interest. For example, events of interest that are adjacent to or temporally near the particular event of interest may include events of interest that occur within a threshold period of time from the particular event of interest. For example, a particular event of interest characterizing a cough event, water- related sounds (e.g., a splash resulting from an expulsion into water, running water, etc.) that occurs temporally near another particular event of interest characterizing an expulsion event may be used to confirm, or increase a confidence of, the expulsion event as a vomit-related event. In particular, the labeler 130 may, between each adjacent pair of events of interest segmented from the sequence of acoustic frames 122, determine a corresponding time duration separating a last acousticframe 122 of the respective subset of sequential acoustic frames 122 associated with an earlier event of interest in the adjacent pair and an initial acoustic frame 122 of the respective subset of sequential acoustic frames 122 associated with a later event of interest in the adjacent pair. Here, labeling the particular event of interest as a vomit- related event or a non-vomit-related event may be further based on the corresponding time duration between the particular event of interest and each of the one or more events of interest of the plurality of events of interest that are adjacent to the particular event of interest. In some examples, the device 10 or the remote computing device 70 discards or does not retain the acoustic frames 122 after they are processed by the detection system 110 to ensure the privacy of the patient 104.

[0046] In some implementations, the detection system 110 also includes a speaker identification model 140 that receives and processes acoustic frames 122 associated with a particular event of interest to determine an identity of the patient 104 as a source of the particular event of interest. The speaker identification model 140 may process the acoustic frames 122 associated with the particular event of interest to generate a speaker vector 142 that uniquely identifies the vocal characteristics of the source of the acoustic frames 122 associated with the particular event of interest. Thereafter, the labeler 130 may compare the speaker vector 142 with a reference vector 144 for the patient 104 that represents the vocal characteristics of the patient 104. The labeler 130 may determine that the patient 104 is the source of the acoustic frames 122 associated with the particular event of interest when the speaker vector 142 matches the reference vector 144. Here, the speaker vector 142 may match the reference vector 144 when a metric or distance (e.g., a cosine distance) between the speaker vector 142 and the reference vector 144 in an embedding space satisfies a distance threshold. Scenarios where the speaker vector 142 does not match the reference vector 144 may indicate that while an event of interest may be characteristic of a vomit-related event, a source of the acoustic frames associated with the vomit-related event may include broadcasted media content or otherwise include some person other than the patient 104 that is in proximity of the user device 10 that experienced the vomit-related event.

[0047] The reference vector 144 may be obtained from the patient 104 during a prior enrollment process where the patient 104 provided voice samples and the speaker identification model 140 generated the reference vector 144 based on the voice samples. In some examples, the reference vector 144 for the patient 104 is obtained from prior events of interest that the labeler 130 determined were vomit-related events and the patient 104 confirmed that the patient 104 did in fact experience a vomit-related event. In some examples, the speaker identification model 140 processes acoustic frames 122 corresponding to speech that occur temporally near an event of interest (e.g., within a threshold period of time) to determine the speaker identification 142 for the event of interest. The labeler 130 may label the particular event of interest as a vomit-related event or a non-vomit-related event further based on the patient 104 being identified as the source of the particular event of interest. Notably, by using speaker identification, audible sounds 106 of other persons sharing the environment 102 with the patient 104 can be ignored.

[0048] In some implementations, the detection system 110 also includes a nonacoustic feature model 150 that receives and processes non-acoustic data 126 associated with the patient 104 and captured by the device 10 within a threshold period of time from when the device 10 captured the sequence of acoustic frames 122 to generate a nonacoustic vomiting classification label 152. Here, the non-acoustic data 126 may include image data representing at least one of motion of the patient 104 or a skin conductance of the patient 104 and the non-acoustic vomiting classification label 152 generated by the non-acoustic feature model 150 may indicate whether the motion of the patient 104 is associated with motions when one is retching or expulsing and / or the skin conductance of the patient 104 is associated with sweating. The vomiting classification label 152 may include a score indicating a likelihood that the non-acoustic data 126 is characteristic of a vomit-related event. The labeler 130 may adjust a confidence score for whether a particular event of interest is a vomit-related event based on the non-acoustic vomiting classification label 152 associated with the patient 104. For instance, a vomiting classification label 152 indicating the non-acoustic data 126 is not likely a vomit-related event may cause the labeler 130 to reduce a confidence score for whether the particularevent of interest is a vomit-related event. By contrast, a vomiting classification label 152 having a higher likelihood of a vomit- related event may result in the labeler 130 boosting the confidence score for whether the particular event of interest is the vomit-related event. The non-acoustic data could additionally or alternatively include a heart rate of the patient 104 and / or a temperature of the patient measured by one or more suitable sensors integrated by the user device 10 when the user device is a wearable device (e.g., a watch or pendant). For instance, a rise in one or both of the patient’s body temperature or heart rate may increase confidence that the patient 104 is experiencing a vomit-related event. Moreover, similar to the speaker identification model 140, the non-acoustic feature model 150 may process non-acoustic data 126 that includes image data to generate a facial identifier representing facial characteristics of a source of the event of interest that may be compared to a reference identifier representing facial characteristics of the patient 104 to determine whether a source of the event of interest is in fact the patient 104.

[0049] In some implementations, the detection system 110 is augmented with one or more additional ML models (not shown for clarity of illustration) that also process the respective subset of sequential acoustic frames 122 associated with a particular event of interest to generate additional corresponding probability distributions over possible audible sound classifications for the particular event of interest. Here, the additional ML models may have architectures that are different from the ML model 200, or may be trained to classify different sets of audible sounds. In such implementations, the labeler 130 may combine the different probability distributions using, for example, a weighting function, voting, etc.

[0050] The device 10 and / or the remote computing device 70 also executes a vomit event detection analysis application 20 (also referred to herein as analysis application 20) configured to, in some examples and / or circumstances, display, on a screen 18 of the device 10, a prompt 108 that prompts the patient 104 to confirm or reject the labeling of a particular event of interest as a vomit-related event or a non- vomit-related event, and receives an indication from the patient 104 indicating whether the patient 104 confirms or rejects the labeling of the particular event of interest as a vomit-related event or a nonvomit-related event. Additionally or alternatively, a text-to-speech system (not shown)(e.g., executing on any combination of the device 10 or the remote computing device 70) may convert the prompt 108 into synthesized speech for audible output by the device 10 and / or another device. In the example shown, the patient 104 may provide the indication by simply speaking “Yes” or “No”. However, the analysis application 20 may optionally receive the indication as an activation of either of a “Yes” or “No” virtual button displayed on the screen 18. In some implementations, a “I just vomited” virtual button is displayed on the screen 18 to enable the patient 104 to indicate a vomit-related event occurred when, for example, the detection system 110 falsely labeled an event of interest as a non-vomit-related event (i.e., a false negative) and, thus, didn’t prompt the patient 104 for confirmation. The “I just vomited” virtual button may be displayed by the user device 10 when the label 132 includes a confidence score that satisfies a first threshold, but fails to satisfy a second higher threshold. Additionally or alternately, the user device 10 may implement a hotword detector or speech recognizer configured to detect an utterance of a wake word, such as “vomit.” Such a button could also be used to confirm the correct labeling of an event of interest as a vomit-related event.

[0051] In some examples, the analysis application 20 determines to prompt the patient 104 based on comparing a confidence score associated with the labeling of the particular event of interest with one or more confidence thresholds. In some implementations, the threshold(s) dynamically adjust over time as the vomit- related event detections system 110 learns from past vomit-related events and non-vomit-related events of the patient 104. In some implementations, when a bout of vomit-related events is detected, the analysis application 20 waits to prompt the patient 104 until the bout has ended and then provides the prompt 108 to the patient 104 to confirm that a bout of vomit-related events occurred without confirming each vomit-related event of the bout. In some examples, the analysis application 20 adapts / updates one or more parameters / weights of the ML model 200 based on indications received from the patient 104. In this way, the device 10 can learn, reduce the number of false positives and false negatives, and improve vomit-related event detection accuracy over time.

[0052] The analysis application 20 may also determine, based on the corresponding label 132 output for each event of interest, quantitative bout information indicating anumber of the plurality of events of interest segmented from the sequence of acoustic frames 122 that include labels 132 indicating vomit-related events. Here, the quantitative bout information may indicate a respective number of occurrences of each of the possible audible sound classifications during various time intervals. Additionally or alternatively, the analysis application 20 may determine, based on the labeled events of interest, quantitative event information indicating occurrences of vomit-related events associated with a health condition of the patient 104. Here, the quantitative event information may indicate a respective number of occurrences of vomit-related events during various time intervals. In some examples, the analysis application 20 determines, based on the labeled events of interest, quantitative event information indicating occurrences of vomit-related events associated with particular diseases such as hyperemesis gravidarum, chronic nausea and vomiting, bulimia, or any other particular disease that may be characterized by vomit-related events. That is, the quantitative event information may indicate whether the occurrences of vomit-related events are characteristic of vomit-related events for the particular disease. As such, the application 20 may use the quantitative event information may be used to predict whether or not the patient has a particular disease.

[0053] In some implementations, when a particular event of interest is labeled as a vomit-related event by the detection system 110 on the device 10, the analysis application 20 transmits the respective subset of sequential acoustic frames 122 associated with the particular event of interest to the remote computing device 70. The remote computing device 70 may be configured to process, using a second ML model 76 different from the ML model 200, the respective subset of sequential acoustic frames 122 to generate a vomit-related event classification score for the particular event of interest, and confirm or reject the label 132 (by the detection system 110) of the particular event of interest as a vomit-related event based on the vomit-related event classification score. In some examples, the second ML model 76 is a larger, more accurate, or more sophisticated ML model than can be implemented on the device 10. When the remote computing device 70 rejects the labeling of the particular event of interest as a vomit-related event, the remote computing device 70 re-labels the particular event of interest as a non-vomit-related event. In other words, the second ML model 76 may serve as a second pass model forconfirming or rejecting a vomit-related event based on the classification 202 determined by the first ML model 200 during a first pass. In some implementations, the remote computing device 70 adapts one or more parameters of the second ML model 76 based on the vomit-related classification score. In some examples, prior to transmitting the respective subset of sequential acoustic frames 122 to the remote computing device 70, the analysis application 20 modifies the respective subset of sequential acoustic frames 122 to at least one of mask an identity of the patient 104 or mask spoken words to ensure the privacy of the patient 104. Example modifications of the acoustic frames 122 include, but are not limited to, reversing the order of the acoustic frames 122, random pitch shifting, and audio encoding that preserves acoustic features useful for vomit- related event detection, but masks speech.

[0054] Additionally or alternatively, the remote computing device 70 may playback the respective subset of sequential acoustic frames 122 associated with the particular event of interest for a user (e.g., a healthcare provider) 78 of the remote computing device 70, and receive, from the user 78, an indication indicating a confirmation or a rejection of the labeling of the particular event of interest. In some implementations, the remote computing device 70 adapts one or more parameters of the second ML model 76 based on the indication received from the user 78.

[0055] In some implementations, the analysis application 20 monitors the patient 104 in real time and may, as needed, alert caregivers, healthcare professionals, or emergency responders to the patient’s condition. In some examples, the analysis application 20 reports the event labels 132, quantitative bout information, and / or quantitative event information to the patient 104 and / or a healthcare professional 78 monitoring or treating the patient 104. The event labels 132, quantitative bout information, and / or quantitative event information may be, for example, presented in a cloud- or web-based interface, sent via a short message service (SMS), sent via electronic mail (email), or stored in a data file (e.g., a comma-separated values (CSV) file).

[0056] In the foregoing description, events of interest are analyzed by the detection system 110 in real time (e.g., in a streaming fashion) as they occur. Additionally or alternatively, raw acoustic frames 122 may be stored (e.g., in non-volatile memory 14) asthey are captured (e.g., for 24 hours), and processed in bulk at a later time. In such examples, the device 10 sends the stored acoustic frames 122 in bulk (e.g., each night) to the remote computing device 70 for subsequent analysis. Additionally or alternatively, the device 10 may process the stored acoustic frames 122 whenever the device 10 is plugged in or not being used for another purpose. To reduce the amount of acoustic frames 122 that are stored, the device 10 may only store acoustic frames 122 associated with events of interest (e.g., having a magnitude satisfying a threshold).

[0057] In some implementations, the ML model 200 is trained or fine-tuned using federated learning techniques. For federated learning, the ML model 200 is stored locally on each of a plurality of user devices 10 (e.g., in a home network) associated with respective patients 104, and the second ML model 76 is a global ML model (i.e., a cloudbased counterpart of the local ML model 200) and is stored remotely at the remote computing device 70. A device 10, using its local ML model 200, processes events of interest detected at the device 10 to generate predicted classifications 202 and labels 132. The device 10 compares the predicted classifications 202 and labels 132 to ground-truth outputs (confirmations / rej ections) provided by a patient 104 to generate client gradients that are sent to the remote computing device 70. The remote computing device 70 utilizes received client gradients to update weights of the global ML model 76.Thereafter, the remote computing device 70 transmits the updated global ML model 76, or updated weights of the global ML model 76, to the devices 10. The devices 10 then replace their local ML model 200 with the updated global ML model 76, or replace the weights of their local ML model 200 with the updated weights of the global ML model 76, thereby updating the local ML models 200. Notably, the client gradients do not expose any sensitive data or information associated with the patent 104 to other devices 10 or the remote computing device 70, but allow the global ML model 76 to be updated based on the audible sounds 106 captured and labeled at all of the devices 10.

[0058] FIG. 2 is a schematic view of an example ML model 200 for processing a respective subset of sequential acoustic frames 122 associated with a particular event of interest to generate a probability distribution over possible audible sound classifications 202 for the particular event of interest. The ML model 200 includes a Mel filter bank 210that converts the acoustic frames 122 of the particular event of interest into a sequence of 128-dimensional log Mel filter bank (fbank) features 212. In some examples, the sequence of features 212 is computed every 10 milliseconds (ms) using a 25 ms Hanning window.

[0059] A convolutional neural network 220 processes the features 212 to generate a ([t / 32]x4xl280) dimensional tensor 222. An example convolutional neural network 220 is based on the EfficientNet-BO model architecture, which is described in a paper entitled “EfficientNet: Rethinking Model Scaling For Convolutional Neural Networks” by M. Tan and Q.V. Le and published in the 2019 Proceedings of the International Conference on Machine Learning (ICML), which is hereby incorporated herein by reference in its entirety.

[0060] A frequency mean pooler 230 applies mean pooling over 4 frequency dimensions to the tensor 222 to produce a ([t / 32]xl280) dimensional tensor 232. Thereafter, the tensor 232 is processed by a 1x1 convolutional filter 240 with a sigmoid activation function to generate a ([t / 32] x #classes) dimensional tensor 242, where #classes is the number of possible audible sound classifications. Finally, a temporal mean pooler 250 performs mean pooling on the tensor 242 to generate a (# classes) dimensional tensor 252 representing the probability distribution 202 over the possible audible sound classifications for the particular event of interest.

[0061] FIG. 3 is a flowchart of an example arrangement of operations for a computer- implemented method 300 of detecting vomit-related events. The operations may be performed by data processing hardware 810 (FIG. 8) (e.g., the data processing hardware 12 of the device 10 or the data processing hardware 72 of the remote computing device 70) based on executing instructions stored on memory hardware 820 (e.g., the memory hardware 14 of the device 10 or the memory hardware 74 of the remote computing device 70).

[0062] At operation 302, the method 300 includes receiving a sequence of acoustic frames 122 captured by a user device 10 associated with a patient 104. At operation 304, the method 300 includes segmenting the sequence of acoustic frames 122 into a plurality of events of interest characterizing one or more audible sounds, each event of interestassociated with a respective subset of sequential acoustic frames 122 segmented from the sequence of acoustic frames 122.

[0063] For each particular event of interest of the plurality of events of interest, the method performs operations 306 and 308. At operation 306, the method 300 includes processing, using the ML model 200, the respective subset of sequential acoustic frames 122 associated with the particular event of interest to generate a corresponding probability distribution 202 over possible audible sound classifications for the particular event of interest. At operation 308, the method 300 includes labeling, using the labeler 130, based on the corresponding probability distribution over possible audible sound classifications 202, the particular event of interest as a vomit-related event or a nonvomit-related event.

[0064] At operation 310, the method 300 includes determining, based on the labeled events of interest, quantitative bout information indicating a number of the plurality of events of interest segmented from the sequence of acoustic frames that are labeled as vomit-related events.

[0065] FIG. 4 is a schematic view of an example training process 400 for training the detection system 110. The training process 400 may execute on the remote computing device 70 (i.e., on the data processing hardware 72) or on the device 10 (i.e., on the data processing hardware 12). The training process 400 trains the detection system 110 using a training dataset 401 that includes a plurality of training samples 402, 402a-n. Here, each training sample 402 includes: a corresponding sequence of training acoustic frames 410; a corresponding ground-truth audible sound classification 420 for each training event of interest of the training sample; and a corresponding ground-truth event label 430 for each training event of interest identifying the training event of interest of the training sample 402 as a vomit-related event or a non-vomit-related event. The corresponding sequence of training acoustic frames 410 include one or more training events of interest each characterizing one or more audible sounds. Each training event of interest is associated with a respective subset of sequential training acoustic frames 410 segmented from the sequence of training acoustic frames 410. An example set of ground- truth audible sound classifications 420 includes one or more of a retch, an expulsion, a cough,a spit, a sneeze, speech, other, laughter, a throat clearing, a splash, water running, or a toilet flushing. An example set of ground-truth event labels 430 includes a retch (R), a cough (C), a retch + a cough (R+C), a spit (S), an expulsion (E), a retch + an expulsion (R+E), an expulsion + a splash (E+S), and an expulsion + toilet sound (E+T). However, other ground-truth labels may be used. In some examples, a human, based on listening to the subset of training acoustic frames 122 for a particular event of interest, manually assigns a corresponding ground-truth audible sound classification 420 and a corresponding ground-truth event label 430 to the event of interest. In other examples, training samples 402 are created from acoustic frames 122 captured by a user device 10 together with a confirmed probability distribution over possible classifications 202 and a confirmed event label 132 for each event of interest of the training sample 402. In some implementations, a human labeler is instructed to manually label: all instances of a retch as R; all instances of a cough, a throat clearing, or a spit as R+C if they are within a threshold time period of a retch; a retch event and an expulsion separately as R and E, respectively, when there is a clear time distinction between the retch event and the expulsion event; an instance of a retch event followed by an expulsion event as R+E when there is not a clear time distinction between the two events; and an instance of a retch event followed by a cough as R+C when there is not a clear time distinction between the retch event and the following cough. However, other guidelines for manually labeling vomit-related events may be used.

[0066] It has been advantageously discovered that the sounds associated with vomit- related events are different from other audible sounds. For example, a retch is characterized by a slowly changing pitch structure, a cough has broadband energy with little structure, and speech shows more modulation over time than retch (especially at high frequencies). Therefore, the training process 400 trains the detection system 110 with training samples 402 of vomit-related sounds as well as other audible sounds. In some examples, the training samples 402 are collected from multiple sources of recorded audible sounds. Example sources of audible sound recordings include, but are not limited to, YouTube or other Internet websites, the AudioSet dataset of human-labeled sound clips drawn from YouTube videos, the VocalSound dataset of crowdsourced recordings,the environmental sound classification (ESC) dataset, and recordings captured during clinical trials. In some examples, the training dataset 401 is an IRB-approved dataset of crowd sourced recordings of laughter, sighs, coughs, throat clearings, sneezes, and sniffs that is augmented with vomit, retch, spitting, toilet noises, and environmental sounds captured from the Internet or during clinical trials.

[0067] It has also been advantageously discovered that various patterns of vocal and other sounds that occur within a threshold period of time (i.e., are temporally near to each other) are highly specific to vomit-related events. For example, a splash sound, a cough, splitting, and / or toilet noise often follows a retch or expulsion. Thus, a combination of audible sounds may be used to further increase the likelihood that the labeler 130 correctly labels vomit-related events. That is, the training process leverages the combination of audible sounds high specific to vomit-related events to train the labeler 130 to reduce instances of false-positives and false-negatives. Therefore, the training process 400 trains the detection system 110 with training samples 402 that include sequences of vomit-related sounds paired with other audible sounds. In some examples, some of the training samples 402 are acoustically varied by, for example, adding Gaussian noise, adjusting volume, or adding background sounds or noise.

[0068] Using techniques such as a cross-validation process, the sensitivity, specificity, and precision of the trained ML model 200 for classifying vomit-related event sounds and non- vomit-related event sounds has been demonstrated. For example, FIG. 5 shows a confusion matrix 500 demonstrating an example sensitivity, specificity, and precision of the trained ML model 200. Here, the ML model 200 has a sensitivity of approximately 81% and a specificity of nearly 100%. Thus, as demonstrated, the trained ML model 200 can accurately classify the sounds associated vomit-related events as vomit-related sounds with a low rate of false positives and a low rate of false negatives.

[0069] For each particular training event of interest of each training sample 402 in the training dataset 401, the training process 400 trains the detection system 110 on the respective subset of sequential training acoustic frames 410 associated with the particular training event of interest to teach the detection system 110 to learn how to predict the corresponding ground-truth sound classification 420 and the corresponding ground-truthevent label 430. In particular, for a particular training sample 402, the training process 400 segments the corresponding sequence of training acoustic frames 410 of the particular training sample 402 into a plurality of training events of interest characterizing one or more audible sounds, each event of interest associated with a respective subset of sequential training acoustic frames 410 segmented from the sequence of training acoustic frames 410. For each particular training event of interest, the training process 400 processes, using the ML model 200, the respective subset of sequential training acoustic frames 410 associated with the particular training event of interest to generate a corresponding probability distribution 202 over possible audible sound classifications for the particular training event of interest. The training process 400 then labels, using the labeler 130, based on the corresponding probability distribution 202 over possible audible sound classifications, the particular training event of interest as a vomit-related event or a non- vomit-related event.

[0070] Thereafter, a loss term module 440 determines a first loss term 442 based on a probability of the probability distribution 202 corresponding to the corresponding ground-truth sound classification 420. In some examples, the first loss term 442 is a negative log of the probability of the probability distribution 202 corresponding to the corresponding ground-truth sound classification 420. The loss term module 440 also determines a second loss term 444 based on a difference between the corresponding ground- truth event label 430 and the predicted label 132 of the particular training event of interest. The training process 400 then trains the detection system 110 (e.g., by adjusting one or more coefficients of the ML model 200 and the labeler 130) based on the loss terms 442 and 444. In some examples, training process 400 uses the first loss term 442 to train the ML model 200, and uses the second loss term 444 to train the labeler 130. In other examples, a combined loss term is determined based on the loss terms 442 and 444, and the combined loss term is used to train both the ML model 200 and the labeler 130.

[0071] In some examples, the training process 400 divides the training dataset 401 into K distinct subsets (e.g., 5 or 10), disjoint by patient 104 so that each patient 104 only appears within one subset. The training process 400 then applies standard cross-validation ML procedures to obtain K different detection systems 110, which are then tested on non-overlapping portions of the training dataset 401. The training process 400 may then compute sensitivity, specificity, and precision for each detection system 110. Thereafter, the training process 400 may report mean, minimum, and maximum values of sensitivity, specificity, and precision values to capture average performance as well as best- and worstcase performances. Here, the training process 400 may report the performance of the trained detection system 110 as a confusion matrix, such as the confusion matrix shown in FIG. 5.

[0072] FIG. 6 is a flowchart of an example arrangement of operations for a computer- implemented method 600 of training the vomit-related event detection system 110 for automatically detecting vomit-related events. The operations of method 600 may be performed by data processing hardware 810 (FIG. 8) (e.g., the data processing hardware 12 of the device 10 or the data processing hardware 72 of the remote computing device 70) based on executing instructions stored on memory hardware 820 (e.g., the memory hardware 14 of the device 10 or the memory hardware 74 of the remote computing device 70).

[0073] At operation 602, the method 600 includes receiving a training dataset 401 that includes a plurality of training samples 402. Here, each training sample 402 in the training dataset 401 includes a corresponding sequence of training acoustic frames 410 including on or more training events of interest. Each training event of interest characterizes one or more audible sounds and is associated with a respective subset of sequential training acoustic frames 410 segmented from the sequence of training acoustic frames. Each training sample 402 in the training dataset 401 also includes a corresponding ground-truth audible sound classification 420 for each training event of interest of the training sample, and a corresponding ground-truth event label 430 for each training event of interest identifying the training event of interest of the training sample as a vomit-related event or a non-vomit-related event.

[0074] At operation 604, the method 600 includes, for each particular training sample 402 in the training dataset 401, training the ML model 200 on the particular training sample 402 to teach the ML model 200 to learn how to predict the corresponding ground-truth audible sound classifications 420 and the corresponding ground-truth event labels 430.

[0075] Referring to FIG. 7, in some implementations, the vomit- related event detection system 110 includes a joint audio and speech understanding architecture 700 designed to simultaneously recognize and reason about both speech and non-speech audio events from input audio data. The joint audio and speech understanding architecture (or simply ‘architecture’) 700 is configured to receive, as input, a respective subset of sequential acoustic frames 122 associated with a particular event of interest and a natural language prompt 702, and generate, as output, a predicted sound classification 701 for the particular event of interest. The natural language prompt 702 can include instructions that guide a large language model (LLM) 750 of the architecture to generate predicted sound classifications 701 and / or labels 132 related to vomiting. For instance, non-limiting natural language prompts 702 may include, without limitation, “Identify the sound (500+ class vocabulary)”, “Detect human voice”, “Identify vomit sounds”, and “Infer recording location”. When the natural language prompt 702 instructs the LLM 750 to identify a sound, the LLM 750 may predict a sound classification from a 500+ class vocabulary. A prompt 702 that instructs the LLM 750 to “infer recording location” may cause the LLM 750 to predict a location from where the audio frames were captured, such as a bathroom. In some examples, multiple natural language prompts 702 are paired with the acoustic frames 122 associated with the particular event of interest.

[0076] The respective subset of sequential acoustic frames 122 may be segmented by the segmenter 124 (FIG. 1) when the input subsystem 120 detects acoustic activity. In some examples, the predicted sound classification is a probability distribution over possible sound classifications 202. In some examples, the architecture 700 may correspond to the ML model 200 of FIG. 1 such that the architecture 700 outputs a predicted sound classification 701 (and / or a probability distribution over possible sound classifications 202) for the respective subset of sequential acoustic frames 122 and the labeler 130 of FIG. 1 outputs a predicted audio event label 130 based on the predicted sound classification 701 (and / or the probability distribution over possible sound classifications 202) . In some examples, the architecture 700 corresponds to both the MLmodel 200 and the labeler 130 of FIG. 1 such that the architecture 700 can output both a predicted sound classification 701 for the particular event of interest and a predicted audio event label 132 indicating whether the particular event of interest is a vomit-related event or a non-vomit related event.

[0077] The architecture 700 includes a speech recognizer 710, a Time and Layer- Wise Transformer (TLTR) module 720, a text tokenizer 730, and a large language model (LLM) 750 adapted by a fine-tuned adapter module 740 for predicting sound classifications 701 and / or event labels 132 based on a respective subset of sequential acoustic frames 122 associated with a particular event of interest and a natural language prompt 702. The respective subset of sequential acoustic frames 122 may include a sequence of acoustic frames spanning 10-seconds of audio. The speech recognizer 710 includes an audio encoder 712 and a speech decoder 716. The audio encoder 712 may include a plurality of multi-head attention layers that are pre-trained on a massive labeled speech corpus recorded under diverse conditions. For instance, the massive labeled speech corpus may include 680,000-hours of labeled speech. In some other configurations, the audio encoder is additionally or alternatively pre-trained on unlabeled speech. The multi-head attention layers of the audio encoder 712 may include Transformer layers, Conformer layers, or another type of multi-head attention layers. Similarly, the speech decoder 716 may include a plurality of multi-head attention layers such as Transformer layers, Conformer layers, or another type of multi-head attention layers. In some implementations, the speech recognizer 710 of the architecture 700 includes 32-layer, 1280-dimensional Transformer networks for both the audio encoder 712 and the speech decoder 716.

[0078] In some implementations, the adapter module 740 includes a parallel adapter module that adapts the pre-trained LLM 750 to output sound classifications 701 and / or audio event labels 132 related to vomit detection based on prompt tokens 734 derived from the natural language input prompt 202 and audio and spoken text tokens 724, 732 derived from respective subsets of acoustic frames 722 processed by the pre-trained audio encoder 712.

[0079] For each acoustic frame 122 in the respective subset of sequential acoustic frames 122, each multi-head attention layer of the audio encoder 712 generates a corresponding intermediate encoding 713. The intermediate encoding 713 generated by the final multi-head attention layer of the audio encoder 712 includes an audio encoding 714 generated as output by the audio encoder 712 for the corresponding acoustic frame 122. Notably, the pre-trained audio encoder 712 is capable of encoding not only linguistic information but also rich background sound, paralinguistic information (e.g., emotion, pitch, background sounds, and / or language development), and other non-speech audio events. Described in greater detail below, the intermediate encodings 713 and the corresponding audio encoding 714 generated as output by the audio encoder 712 for each acoustic frame 122 is utilized in two parallel pathways: a speech transcription pathway and an audio event and paralinguistic feature extraction pathway.

[0080] At the speech transcription pathway, the speech decoder 716 receives each audio encoding 714 generated as output by the audio encoder 712 to transcribe the respective subset of sequential acoustic frames 122 into spoken text 718 when the acoustic frames characterizes speech. If the respective subset of sequential acoustic frames does not characterize speech, the output from the speech decoder 716 is empty. The spoken text 718 (if any) output by the speech decoder 718 is tokenized by the text tokenizer 730 into corresponding spoken text tokens 732, while the natural language input prompt 702 is tokenized by the text tokenizer 730 into corresponding prompt tokens 734.

[0081] At the audio event and paralinguistic feature extraction pathway, the TLTR module 720 receives the intermediate audio encodings 713 output by the plurality of multi-head attention layers (e.g., 32 Transformer layers) of the audio encoder 712 for each acoustic frame 122 in the respective subset of acoustic frames 122 and generates a TLTR output 722 that a projection layer 724 projects into a sequence of audio tokens 725 that have an embedding dimension required by the LLM 750. Here, the sequence of audio tokens 732 encode soft audio events and paralinguistic information (e.g., emotion, pitch, background sounds, and / or language development) from the respective subset of acoustic frames 122. In some examples, the TLTR output 722 includes a 1280-dimensional representation and the projection layer 724 projects the TLTR output 722 into a 4096-dimensional space to match an embedding dimension required by the LLM 750. The TLTR module 720 may be pre-trained on an AudioSet of audio event detection. Specifically, the TLTR module 720 applies an attention mechanism over both the temporal and layer dimensions of the intermediate encodings 713, thereby enabling the TLTR module 720 to extract both linguistic and non-linguistic information.

[0082] The LLM 750 adapted by the fine-tuned adapter module 740 functions as a reasoning model that receives a concatenation of the sequence of audio tokens 725 from the projection layer 720, the spoken text tokens 732 tokenized from the spoken text 718 (if any) output by the speech decoder 716, and the prompt tokens 734 tokenized from the input prompt 202 to predict at least one of the sound classification 701 for the respective subset of sequential acoustic frames 122 associated with the particular event of interest, the probability distribution over possible sound classifications 102 for the respective subset of sequential acoustic frames 122 associated with the particular event of interest, or the event label 132 indicating whether the particular event of interest is a vomit- related event or a non-vomit related event. In some examples, the LLM 750 includes seven (7) billion parameters and is fine-tuned on an instruction-following dataset. The architecture 700 may include about 8.5 billion parameters, however, only 49 million parameters may be trainable (e.g., 40 million parameters for the TLTR module 720, 4.2 million parameters for the adapter module 740, and 5 million parameters for the projection layer 730).

[0083] As previously described, the analysis application 20 (FIG. 1) may also determine, based on the corresponding label 132 output by the LLM 750 for each event of interest, quantitative bout information indicating a number of the plurality of events of interest segmented from the sequence of acoustic frames 122 that include labels 132 indicating vomit-related events. Here, the quantitative bout information may indicate a respective number of occurrences of each of the possible audible sound classifications during various time intervals. Additionally or alternatively, the analysis application 20 may determine, based on the labeled events of interest, quantitative event information indicating occurrences of vomit-related events associated with a health condition of thepatient 104. Here, the quantitative event information may indicate a respective number of occurrences of vomit-related events during various time intervals. In some examples, the analysis application 20 determines, based on the labeled events of interest, quantitative event information indicating occurrences of vomit-related events associated with particular diseases such as hyperemesis gravidarum, chronic nausea and vomiting, bulimia, or any other particular disease that may be characterized by vomit-related events. That is, the quantitative event information may indicate whether the occurrences of vomit-related events are characteristic of vomit-related events for the particular disease. As such, the application 20 may use the quantitative event information may be used to predict whether or not the patient has a particular disease.

[0084] A fine-tuning process 780 may train the joint audio and speech understanding architecture 700 as an end-to-end model for audible sound classification and / or predicting vomit-related events from input audio. Specifically, the audio encoder 712, the speech decoder 716, and the LLM 750 may include pre-trained models and the fine-tuning process 780 fine-tunes the adapter module 740 of the LLM 750 while parameters of the audio encoder 712, the speech decoder 716, and the LLM 750 remain fixed. Notably, the freezing of the parameters of the audio encoder 712 and the speech decoder 716 preserves robust speech recognition capabilities of the speech recognizer 710 while also reducing computational costs associated with training / updating parameters of the speech recognizer 710. Similarly, freezing the parameters of the LLM 750 mitigates catastrophic forgetting by the LLM 750 as well as reducing computational overhead that would be required for training / updating parameters of the LLM 750. The adapter module 740 adapts the pre-trained LLM 750 for audible sound classification related to vomiting and / or predicting vomit-related events from input audio. Specifically, the fine-tuning process 780 trains the LLM 750 to learn how to predict ground-truth classifications 420 and / or ground-truth labels 430 from training audio frames 410 by only updating parameters of the adapter module 740 while parameters of the pre- trained LLM 750 are held fixed. In some examples, in addition to updating parameters of the adapter module 740, the fine-tuning process 780 also updates parameters of either or both of the TLTR module 720 and the projection layer 724. In the example shown in FIG. 7, the shading ofthe TLTR module 720, the projection layer 724, and the adapter module 740 indicate trainable parameters by the fine-tuning process 780. Dashed lines and dashed components indicate features only present during the fine-tuning process 780.

[0085] The fine-tuning process 780 trains the joint audio and speech understanding architecture 700 using a training dataset 401 that includes a plurality of training samples 402. For simplicity, only one training sample 402 is depicted. Here, each training sample 402 includes: a corresponding sequence of training acoustic frames 410; a corresponding ground-truth audible sound classification 420 for each training event of interest of the training sample; a corresponding ground- truth event label 430 for each training event of interest identifying the training event of interest of the training sample 402 as a vomit-related event or a non-vomit-related event; and a training natural language prompt 485. The training dataset 401 used by the fine-tuning process 780 may be the same as the training dataset 401 used by the training process 400 of FIG. 4 except that each training sample 402 is modified to further include the training natural language prompt 485. Notably, one or more of the training samples 402 in the training dataset 401 used by the fine-tuning process 780 may include training natural language prompts 485 that are different than the training natural language prompts 485 of the other training samples 402 in the training dataset 401. The corresponding sequence of training acoustic frames 410 include one or more training events of interest each characterizing one or more audible sounds. Each training event of interest is associated with a respective subset of sequential training acoustic frames 410 segmented from the sequence of training acoustic frames 410. An example set of ground-truth audible sound classifications 420 includes one or more of a retch, an expulsion, a cough, a spit, a sneeze, speech, other, laughter, a throat clearing, a splash, water running, an alarm sound, or a toilet flushing. An example set of ground-truth event labels 430 includes a retch (R), a cough (C), a retch + a cough (R+C), a spit (S), an expulsion (E), a retch + an expulsion (R+E), an expulsion + a splash (E+S), and an expulsion + toilet sound (E+T). However, other ground-truth labels may be used. In some examples, a human, based on listening to the subset of training acoustic frames 410 for a particular event of interest, manually assigns a corresponding ground-truth audible sound classification 420, a correspondingground- truth event label 430 to the event of interest, and a corresponding training nautral language prompt 475. In other examples, training samples 402 are created from acoustic frames 122 captured by a user device 10 together with a confirmed probability distribution over possible classifications 202 and a confirmed event label 132 for each event of interest of the training sample 402. In some implementations, a human labeler is instructed to manually label: all instances of a retch as R; all instances of a cough, a throat clearing, or a spit as R+C if they are within a threshold time period of a retch; a retch event and an expulsion separately as R and E, respectively, when there is a clear time distinction between the retch event and the expulsion event; an instance of a retch event followed by an expulsion event as R+E when there is not a clear time distinction between the two events; and an instance of a retch event followed by a cough as R+C when there is not a clear time distinction between the retch event and the following cough. However, other guidelines for manually labeling vomit-related events may be used.

[0086] As previously discussed with reference to the training process 400 of FIG. 4, sounds associated with vomit-related events are different from other audible sounds. For example, a retch is characterized by a slowly changing pitch structure, a cough has broadband energy with little structure, and speech shows more modulation over time than retch (especially at high frequencies). Therefore, the fine-tuning process 780 trains the joint speech and audio architecture 700 with training samples 402 of vomit-related sounds as well as other audible sounds. Additionally, by the architecture 700 implementing the speech recognizer 710 that includes the pre-trained audio encoder 712 and the pre-trained speech decoder 716, the speech recognizer 710 advantageously helps distinguish human-generated audio from environmental or medical sounds. In some examples, the training samples 402 are collected from multiple sources of recorded audible sounds. Example sources of audible sound recordings include, but are not limited to, YouTube or other Internet websites, the AudioSet dataset of human-labeled sound clips drawn from YouTube videos, the VocalSound dataset of crowdsourced recordings, the environmental sound classification (ESC) dataset, and recordings captured during clinical trials. Notably, the TLTR module 720 may be initially pre-trained on samplesfrom the AudioSet dataset of human-labeled sound clips. In some examples, the training dataset 401 is an IRB-approved dataset of crowd sourced recordings of laughter, sighs, coughs, throat clearings, sneezes, and sniffs that are augmented with vomit, retch, spitting, toilet noises, alarms, and environmental sounds captured from the Internet or during clinical trials.

[0087] It has also been advantageously discovered that various patterns of vocal and other sounds that occur within a threshold period of time (i.e., are temporally near to each other) are highly specific to vomit-related events. For example, a splash sound, a cough, splitting, and / or toilet noise often follows a retch or expulsion. Thus, a combination of audible sounds may be used to further increase the likelihood that the LLM 450 (e.g., functioning as the ML network 200 and the labeler 130 of FIG. 1) correctly predicts a sound classification 701 for the particular event of interest and correctly labels vomit- related events. That is, the fine-tuning process 780 leverages the combination of audible sounds highly specific to vomit-related events to train the adapter module 740 to reduce instances of false-positives and false-negatives output by the LLM 750. Therefore, the fine-tuning process 780 fine-tunes the adapter module 740, and optionally the TLDR module 720 and the projection layer 730, with training samples 402 that include sequences of vomit-related sounds paired with other audible sounds. In some examples, some of the training samples 402 are acoustically varied by, for example, adding Gaussian noise, adjusting volume, or adding background sounds or noise.

[0088] For the training event of interest of each corresponding training sample 402 in the training dataset 401, the fine-tuning process 780 trains on the adapter module 740, and optionally the TLDR module 720 and the projection layer 730, on the respective subset of sequential training acoustic frames 410 associated with the particular training event of interest and the training natural language prompt 485 while parameters of the speech recognizer 710 and the LLM 750 are held fixed to adapt the LLM 750 to predict the corresponding ground-truth sound classification 420 and / or the corresponding ground-truth event label 430. In particular, for a corresponding training sample 402, the fine-tuning process 480 segments the corresponding sequence of training acoustic frames 410 of the particular training sample 402 into a plurality of training events of interestcharacterizing one or more audible sounds, each event of interest associated with a respective subset of sequential training acoustic frames 410 segmented from the sequence of training acoustic frames 410. In some examples, each respective subset of sequential training acoustic frames 410 is associated with a 10-second clip of audio. For each training event of interest, the fine-tuning process 480 processes, using pre-trained audio encoder 712, the respective subset of sequential training acoustic frames 410 associated with the corresponding training event of interest to generate a corresponding audio encoding 714 for each training acoustic frame 410 and a plurality of intermediate encodings 713 each generated by a respective one of the plurality of multi-head attention layers of the pre-trained audio encoder 712.

[0089] At the speech transcription pathway, the speech decoder 716 receives each audio encoding 714 generated as output by the audio encoder 712 to transcribe the respective subset of sequential training acoustic frames 410 into spoken text 718 when the acoustic frames characterizes speech. If the respective subset of sequential training acoustic frames 410 does not characterize speech, the output from the speech decoder 716 is empty. The spoken text 718 (if any) output by the speech decoder 718 is tokenized by the text tokenizer 730 into corresponding spoken text tokens 732, while the training natural language prompt 485 is tokenized by the text tokenizer 730 into corresponding prompt tokens 734.

[0090] At the audio event and paralinguistic feature extraction pathway, the TLTR module 720 receives the intermediate audio encodings 713 output by the plurality of multi-head attention layers (e.g., 32 Transformer layers) of the audio encoder 712 for each acoustic frame 410 in the respective subset of training acoustic frames 410 and generates the TLTR output 722 that the projection layer 724 projects into a sequence of audio tokens 725 that have an embedding dimension required by the LLM 750. The sequence of audio tokens 732 may encode soft audio events and paralinguistic information (e.g., emotion, pitch, background sounds, and / or language development) from the respective subset of training acoustic frames 410.

[0091] The fine-tuning process 780 concatenates the sequence of audio tokens 725 from the projection layer 720, the spoken text tokens 732 tokenized from the spoken text718 (if any) output by the speech decoder 716, and the prompt tokens 734 tokenized from the training natural language prompt 485 for processing by the LLM 750 adapted by the adapter module 740 to predict at least one of the sound classification 701 for the respective subset of sequential training acoustic frames 410, the probability distribution over possible sound classifications 202 for the respective subset of sequential training acoustic frames 410, or the event label 132 indicating whether the particular event of interest is a vomit-related event or a non-vomit related event. In some examples, the LLM 750 includes seven (7) billion parameters and is fine-tuned on an instructionfollowing dataset.

[0092] Thereafter, a loss term module 760 determines a first loss term 762 based on the predicted sound classification 701 and the corresponding ground-truth sound classification 420. In some examples, the first loss term 462 is a negative log of a probability of a probability distribution 202 over possible sound classifications corresponding to the corresponding ground-truth sound classification 420. The loss term module 760 also determines a second loss term 764 based on a difference between the corresponding ground-truth event label 430 and the predicted label 132 of the particular training event of interest. The fine-tuning process 780 then fine-tunes the adapter module 740 (e.g., by adjusting one or more coefficients or weights of the adapter module 740) through cross-entropy based on the loss terms 762 and 764. Optionally, the fine-tuning process 780 may also fine-tune the TLTR module 720 and the projection layer 724 through cross-entropy based on the loss terms 762 and 764. In some examples, a combined loss term is determined based on the loss terms 762 and 764, and the combined loss term is used as a cross-entropy loss term to train both the adapter module 740 (and optionally the TLTR module 720 and the projection layer 724).

[0093] Notably, the analysis application 20 may adapt / update one or more parameters / weights of the adapter module 740 (and optionally the TLDR module 720 and the projection layer 730) based on indications received from the patient 104, e.g., in response to issuing a prompt 108 to the patient 104 that prompts the patient to confirm or reject sound classifications 701 output by the LLM 750 and / or event labels 132 output by the LLM 750 or output by the labeler 130 based on the sound classifications 701 outputby the LLM 750. In this way, the device 10 can learn, reduce the number of false positives and false negatives, and improve vomit-related event detection accuracy over time

[0094] FIG. 8 is schematic view of an example computing device 800 that may be used to implement the systems and methods described in this document. The computing device 800 is intended to represent various forms of digital devices and computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The components shown here, their connections and relationships, and their functions, are meant to be exemplary only, and are not meant to limit implementations of the inventions described and / or claimed in this document.

[0095] The computing device 800 includes a processor 810 (i.e., data processing hardware) that can be used to implement the data processing hardware 12 and / or 82, memory 820 (i.e., memory hardware) that can be used to implement the memory hardware 14 and / or 84, a storage device 830 (i.e., memory hardware) that can be used to implement the memory hardware 14 and / or 84, a high-speed interface / controller 840 connecting to the memory 820 and high-speed expansion ports 850, and a low speed interface / controller 860 connecting to a low speed bus 870 and a storage device 830 that can be used to store a conversational training dataset. Each of the components 810, 820, 830, 840, 850, and 860, are interconnected using various busses, and may be mounted on a common motherboard or in other manners as appropriate. The processor 810 can process instructions for execution within the computing device 800, including instructions stored in the memory 820 or on the storage device 830 to display graphical information for a graphical user interface (GUI) on an external input / output device, such as display 880 coupled to high speed interface 840. In other implementations, multiple processors and / or multiple buses may be used, as appropriate, along with multiple memories and types of memory. Also, multiple computing devices 800 may be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multi-processor system).

[0096] The memory 820 stores information non-transitorily within the computing device 800. The memory 820 may be a computer-readable medium, a volatile memoryunit(s), or non-volatile memory unit(s). The non-transitory memory 820 may be physical devices used to store programs (e.g., sequences of instructions) or data (e.g., program state information) on a temporary or permanent basis for use by the computing device 800. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electronically erasable programmable readonly memory (EEPROM) (e.g., typically used for firmware, such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM) as well as disks or tapes.

[0097] The storage device 830 is capable of providing mass storage for the computing device 800. In some implementations, the storage device 830 is a computer- readable medium. In various different implementations, the storage device 830 may be a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid state memory device, or an array of devices, including devices in a storage area network or other configurations. In additional implementations, a computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine- readable medium, such as the memory 820, the storage device 830, or memory on processor 810.

[0098] The high speed controller 840 manages bandwidth-intensive operations for the computing device 800, while the low speed controller 860 manages lower bandwidthintensive operations. Such allocation of duties is exemplary only. In some implementations, the high-speed controller 840 is coupled to the memory 820, the display 880 (e.g., through a graphics processor or accelerator), and to the high-speed expansion ports 850, which may accept various expansion cards (not shown). In some implementations, the low-speed controller 860 is coupled to the storage device 830 and a low-speed expansion port 890. The low-speed expansion port 890, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may becoupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., through a network adapter.

[0099] The computing device 800 may be implemented in a number of different forms, as shown in the figure. For example, it may be implemented as a standard server 800a or multiple times in a group of such servers 800a, as a laptop computer 800b, or as part of a rack server system 800c.

[0100] Various implementations of the systems and techniques described herein can be realized in digital electronic and / or optical circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0101] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an “application,” an “app,” or a “program.” Example applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.

[0102] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor, and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms “machine- readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer readable medium, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used toprovide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0103] The processes and logic flows described in this specification can be performed by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for performing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Computer readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0104] To provide for interaction with a user, one or more aspects of the disclosure can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen for displaying information to the user and optionally a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used toprovide interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user's client device in response to requests received from the web browser.

[0105] Unless expressly stated to the contrary, the phrase “at least one of A, B, or C” is intended to refer to any combination or subset of A, B, C such as: (1) at least one A alone; (2) at least one B alone; (3) at least one C alone; (4) at least one A with at least one B; (5) at least one A with at least one C; (6) at least one B with at least C; and (7) at least one A with at least one B and at least one C. Moreover, unless expressly stated to the contrary, the phrase “at least one of A, B, and C” is intended to refer to any combination or subset of A, B, C such as: (1) at least one A alone; (2) at least one B alone; (3) at least one C alone; (4) at least one A with at least one B; (5) at least one A with at least one C; (6) at least one B with at least one C; and (7) at least one A with at least one B and at least one C. Furthermore, unless expressly stated to the contrary, “A or B” is intended to refer to any combination of A and B, such as: (1) A alone; (2) B alone; and (3) A and B.

[0106] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the disclosure. Accordingly, other implementations are within the scope of the following claims.

Claims

WHAT IS CLAIMED IS:

1. A computer-implemented method (300) executed on data processing hardware (810) that causes the data processing hardware (810) to perform operations comprising: receiving a sequence of acoustic frames (122) captured by a user device (10) associated with a patient (104); segmenting the sequence of acoustic frames (122) into a plurality of events of interest characterizing one or more audible sounds (106), each event of interest associated with a respective subset of sequential acoustic frames (122) segmented from the sequence of acoustic frames (122); for each particular event of interest of the plurality of events of interest: processing, using a machine learning (ML) model (200), the respective subset of sequential acoustic frames (122) associated with the particular event of interest to generate a corresponding probability distribution over possible audible sound classifications (202) for the particular event of interest; and labeling, using a labeler (130), based on the corresponding probability distribution over possible audible sound classifications (202), the particular event of interest as a vomit-related event or a non-vomit-related event; and determining, based on the labeled events of interest, quantitative bout information indicating a number of the plurality of events of interest segmented from the sequence of acoustic frames (122) that are labeled as vomit-related events.

2. The method (300) of claim 1, wherein the quantitative bout information further indicates a respective number of occurrences of each of the possible audible sound classifications (202).

3. The method (300) of claim 1 or 2, wherein the operations further comprise determining, based on the labeled events of interest, quantitative event information indicating occurrences of vomit-related events associated with a health condition of the patient (104).

4. The method (300) of any of claims 1-3, wherein labeling the particular event of interest as a vomit-related event or a non-vomit-related event is further based on the corresponding probability distribution over possible audible sound classifications (202) determined for each of one or more events of interest of the plurality of events of interest that are temporally near to the particular event of interest.

5. The method (300) of claim 4, wherein the operations further comprise: between each temporally near pair of events of interest segmented from the sequence of acoustic frames (122), determining a corresponding time duration separating a last acoustic frame (122) of the respective subset of sequential acoustic frames (122) associated with an earlier event of interest in the adjacent pair and an initial acoustic frame (122) of the respective subset of sequential acoustic frames (122) associated with a later event of interest in the adjacent pair, wherein labeling the particular event of interest as a vomit-related event or a non- vomit-related event is further based on the corresponding time duration between the particular event of interest and each of the one or more events of interest of the plurality of events of interest that are temporally near to the particular event of interest.

6. The method (300) of any of claims 1-5, wherein the operations further comprise, when the particular event of interest is labeled as a vomit-related event: transmitting the respective subset of sequential acoustic frames (122) associated with the particular event of interest to a remote computing device (70), wherein the remote computing device (70) is configured to: process, using a second ML model (76) different from the ML model (200), the respective subset of sequential acoustic frames (122) to generate a vomit- related event classification score for the particular event of interest; and confirm or reject the labeling of the particular event of interest as a vomit- related event based on the vomit-related event classification score.

7. The method (300) of claim 6, wherein the operations further comprise, when the remote computing device (70) rejects the labeling of the particular event of interest as a vomit-related event, re-labeling the particular event of interest as a non-vomit-related event.

8. The method (300) of claim 6 or 7, wherein the operations further comprise adapting one or more parameters of the second ML model (76) based on the vomit- related classification score.

9. The method (300) of any of claims 1-8, wherein the operations further comprise, when the particular event of interest is labeled as a vomit-related event: transmitting the respective subset of sequential acoustic frames (122) associated with the particular event of interest to a remote computing device (70), wherein the remote computing device (70) is configured to: playback the respective subset of sequential acoustic frames (122) associated with the particular event of interest to a user; and receive, from the user, an indication indicating a confirmation or a rejection of the labeling of the particular event of interest.

10. The method (300) of claim 9, wherein the operations further comprise adapting one or more parameters of the ML model (200) based on the indication received from the user.

11. The method (300) of claim 9 or 10, wherein the operations further comprise, prior to transmitting the respective subset of sequential acoustic frames (122) to the remote computing device (70), modifying the respective subset of sequential acoustic frames (122) to at least one of mask an identity of the patient (104) or mask spoken words.

12. The method (300) of any of claims 1-11, wherein the operations further comprise, when the particular event of interest is labeled as a vomit-related event:outputting, from the user device (10), a prompt (108) prompting the patient (104) to confirm or reject the labeling of the particular event of interest as a vomit-related event; receiving an indication from the patient (104) indicating the patient (104) rejects the labeling of the particular event of interest as a vomit-related event; and based on receiving the indication from the patient (104) indicating the patient (104) rejects the labeling of the particular event of interest as a vomit-related event, relabeling the particular event of interest as a non-vomit-related event.

13. The method (300) of claim 12, wherein the operations further comprise, based on receiving the indication from the patient (104) indicating the patient (104) rejects the labeling of the particular event of interest as a vomit-related event, adapting one or more parameters of the ML model (200) based on the acoustic frames (122) of the particular event of interest and the indication.

14. The method (300) of any of claims 1-13, wherein the data processing hardware (810) resides on the user device (10).

15. The method (300) of any of claims 1-14, wherein the data processing hardware (810) resides on a remote computing device (70) in communication with the user device (10).

16. The method (300) of any of claims 1-15, wherein the possible audible sound classifications (202) comprise one or more of a retch, an expulsion, a cough, a spit, a sneeze, speech, laughter, a throat clearing, a splash, water running, an alarm, or a toilet flushing.

17. The method (300) of any of claims 1-16, wherein the operations further comprise: receiving non-audio features (126) associated with the patient (104) and captured by the user device (10) within a threshold period of time from when the user device (10)captured the sequence of acoustic frames (122), the non-audio features representing at least one of motion of the patient (104) or a skin conductance of the patient (104), wherein labeling the particular event of interest as a vomit-related event or a nonvomit-related event is further based on the non-audio features (126) associated with the patient (104).

18. The method (300) of any of claims 1-17, wherein the operations further comprise, for each particular event of interest of the plurality of events of interest: processing, using a speaker identification model (140), the respective subset of sequential acoustic frames (122) associated with the particular event of interest to identify the patient (104) as a source of the particular event of interest, wherein labeling the particular event of interest as a vomit-related event or a nonvomit-related event is further based on identifying the patient (104) as the source of the particular event of interest.

19. The method (300) of any of claims 1-18, wherein labeling the particular event of interest as a vomit-related event or a non-vomit-related event comprises labeling the particular event of interest as a non-vomit-related event or one of multiple different types of vomit-related events, the multiple different types of vomit-related events comprising at least one of a retch and expulsion event, a retch-only event, an expulsion-only event, a retch and cough event, a spitting event, a coughing event, or a throat-clearing event.

20. The method (300) of any of claims 1-19, wherein the ML model (200) is trained by: receiving a training dataset (401) comprising a plurality of training samples (402), each training sample (402) in the training dataset (401) comprising: a corresponding sequence of training acoustic frames (410), the corresponding sequence of training acoustic frames (410) including one or more training events of interest, each training event of interest characterizing one or more audible sounds (106), each training event of interest associated with a respective subset ofsequential acoustic frames (122) segmented from the sequence of training acoustic frames (410); a corresponding ground-truth audible sound classification (420) for each training event of interest of the training sample (402); and a corresponding ground-truth label (430) for each training event of interest identifying the training event of interest of the training sample (402) as a vomit-related event or a non-vomit-related event; and for each particular training sample (402) in the training dataset (401), training the ML model (200) on the particular training sample (402) to teach the ML model (200) to learn how to predict the corresponding ground-truth audible sound classifications (420) and the corresponding ground-truth labels (430).

21. The method (300) of claim 20, wherein training the ML model (200) on a particular training sample (402) comprises: segmenting the corresponding sequence of training acoustic frames (410) of the particular training sample (402) into a plurality of training events of interest characterizing one or more audible sounds (106), each event of interest associated with a respective subset of sequential training acoustic frames (410) segmented from the sequence of training acoustic frames (410); for each particular training event of interest of the plurality of training events of interest: processing, using the ML model (200), the respective subset of sequential training acoustic frames (410) associated with the particular training event of interest to generate a corresponding probability distribution over possible audible sound classifications (202) for the particular training event of interest; labeling, using a labeler (130), based on the corresponding probability distribution over possible audible sound classifications (202), the particular training event of interest as a vomit-related event or a non-vomit-related event; determining a loss based on one or more of:a probability of the probability distribution corresponding to the corresponding ground-truth audible sound classification (420); or a difference between the corresponding ground-truth label (430) and the label (132) of the particular training event of interest; and updating a plurality of weights of the ML model (200) based on the determined loss.

22. The method (300) of any of claims 1-21, wherein the ML model (200) comprises a joint audio and speech architecture (700) that comprises a pre-trained audio encoder (712), a pre-trained speech decoder (716), and a pre-trained large language model (LLM) 750 adapted by a fine-tuned adapter module (740).

23. The method (300) of claim 22, wherein joint audio and speech architecture (700) is trained by: receiving a training dataset (401) comprising a plurality of training samples (402), each training sample (402) in the training dataset (401) comprising: a corresponding sequence of training acoustic frames (410), the corresponding sequence of training acoustic frames (410) including one or more training events of interest, each training event of interest characterizing one or more audible sounds (106), each training event of interest associated with a respective subset of sequential acoustic frames (122) segmented from the sequence of training acoustic frames (410); a corresponding training natural language prompt (485); a corresponding ground-truth audible sound classification (420) for each training event of interest of the training sample (402); and a corresponding ground-truth label (430) for each training event of interest identifying the training event of interest of the training sample (402) as a vomit-related event or a non-vomit-related event; and fine-tuning the adapter module (740) on the plurality of training samples (402) by updating a plurality of weights of the adapter module (740) while parameters of thepre-trained audio encoder (712), the pre- trained speech decoder (716), and the pre- trained LLM (750) are held fixed to teach the adapter module (740) to learn how to adapt the pre-trained LLM (750) to predict the corresponding ground-truth audible sound classifications (420) and the corresponding ground-truth labels (430).

24. A system (100) comprising: data processing hardware (810); and memory hardware (820) in communication with the data processing hardware (810), the memory hardware (820) storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations comprising: receiving a sequence of acoustic frames (122) captured by a user device (10) associated with a patient (104); segmenting the sequence of acoustic frames (122) into a plurality of events of interest characterizing one or more audible sounds (106), each event of interest associated with a respective subset of sequential acoustic frames (122) segmented from the sequence of acoustic frames (122); for each particular event of interest of the plurality of events of interest: processing, using a machine learning (ML) model (200), the respective subset of sequential acoustic frames (122) associated with the particular event of interest to generate a corresponding probability distribution over possible audible sound classifications (202) for the particular event of interest; and labeling, using a labeler (130), based on the corresponding probability distribution over possible audible sound classifications (202), the particular event of interest as a vomit-related event or a non-vomit-related event; and determining, based on the labeled events of interest, quantitative bout information indicating a number of the plurality of events of interest segmented from the sequence of acoustic frames (122) that are labeled as vomit-related events.

25. The system (100) of claim 1, wherein the quantitative bout information further indicates a respective number of occurrences of each of the possible audible sound classifications (202).

26. The system (100) of claim 24 or 25, wherein the operations further comprise determining, based on the labeled events of interest, quantitative event information indicating occurrences of vomit-related events associated with a health condition of the patient (104).

27. The system (100) of any of claims 24-26, wherein labeling the particular event of interest as a vomit-related event or a non-vomit-related event is further based on the corresponding probability distribution over possible audible sound classifications (202) determined for each of one or more events of interest of the plurality of events of interest that are temporally near to the particular event of interest.

28. The system (100) of claim 27, wherein the operations further comprise: between each temporally near pair of events of interest segmented from the sequence of acoustic frames (122), determining a corresponding time duration separating a last acoustic frame (122) of the respective subset of sequential acoustic frames (122) associated with an earlier event of interest in the adjacent pair and an initial acoustic frame (122) of the respective subset of sequential acoustic frames (122) associated with a later event of interest in the adjacent pair, wherein labeling the particular event of interest as a vomit-related event or a non- vomit-related event is further based on the corresponding time duration between the particular event of interest and each of the one or more events of interest of the plurality of events of interest that are temporally near to the particular event of interest.

29. The system (100) of any of claims 24-28, wherein the operations further comprise, when the particular event of interest is labeled as a vomit-related event:transmitting the respective subset of sequential acoustic frames (122) associated with the particular event of interest to a remote computing device (70), wherein the remote computing device (70) is configured to: process, using a second ML model (76) different from the ML model (200), the respective subset of sequential acoustic frames (122) to generate a vomit- related event classification score for the particular event of interest; and confirm or reject the labeling of the particular event of interest as a vomit- related event based on the vomit-related event classification score.

30. The system (100) of claim 29, wherein the operations further comprise, when the remote computing device (70) rejects the labeling of the particular event of interest as a vomit-related event, re-labeling the particular event of interest as a non-vomit-related event.

31. The system (100) of claim 29 or 30, wherein the operations further comprise adapting one or more parameters of the second ML model (76) based on the vomit- related classification score.

32. The system (100) of any of claims 24-31, wherein the operations further comprise, when the particular event of interest is labeled as a vomit-related event: transmitting the respective subset of sequential acoustic frames (122) associated with the particular event of interest to a remote computing device (70), wherein the remote computing device (70) is configured to: playback the respective subset of sequential acoustic frames (122) associated with the particular event of interest to a user; and receive, from the user, an indication indicating a confirmation or a rejection of the labeling of the particular event of interest.

33. The system (100) of claim 32, wherein the operations further comprise adapting one or more parameters of the ML model (200) based on the indication received from the user.

34. The system (100) of claim 32 or 33, wherein the operations further comprise, prior to transmitting the respective subset of sequential acoustic frames (122) to the remote computing device (70), modifying the respective subset of sequential acoustic frames (122) to at least one of mask an identity of the patient (104) or mask spoken words.

35. The system (100) of any of claims 24-34, wherein the operations further comprise, when the particular event of interest is labeled as a vomit-related event: outputting, from the user device (10), a prompt (108) prompting the patient (104) to confirm or reject the labeling of the particular event of interest as a vomit-related event; receiving an indication from the patient (104) indicating the patient (104) rejects the labeling of the particular event of interest as a vomit-related event; and based on receiving the indication from the patient (104) indicating the patient (104) rejects the labeling of the particular event of interest as a vomit-related event, relabeling the particular event of interest as a non-vomit-related event.

36. The system (100) of claim 35, wherein the operations further comprise, based on receiving the indication from the patient (104) indicating the patient (104) rejects the labeling of the particular event of interest as a vomit-related event, adapting one or more parameters of the ML model (200) based on the acoustic frames (122) of the particular event of interest and the indication.

37. The system (100) of any of claims 24-36, wherein the data processing hardware (810) resides on the user device (10).

38. The system (100) of any of claims 24-37, wherein the data processing hardware (810) resides on a remote computing device (70) in communication with the user device (10).

39. The system (100) of any of claims 24-38, wherein the possible audible sound classifications (202) comprise one or more of a retch, an expulsion, a cough, a spit, a sneeze, speech, laughter, a throat clearing, a splash, water running, an alarm, or a toilet flushing.

40. The system (100) of any of claims 24-39, wherein the operations further comprise: receiving non-audio features (126) associated with the patient (104) and captured by the user device (10) within a threshold period of time from when the user device (10) captured the sequence of acoustic frames (122), the non-audio features representing at least one of motion of the patient (104) or a skin conductance of the patient (104), wherein labeling the particular event of interest as a vomit-related event or a nonvomit-related event is further based on the non-audio features (126) associated with the patient (104).

41. The system (100) of any of claims 24-40, wherein the operations further comprise, for each particular event of interest of the plurality of events of interest: processing, using a speaker identification model (140), the respective subset of sequential acoustic frames (122) associated with the particular event of interest to identify the patient (104) as a source of the particular event of interest, wherein labeling the particular event of interest as a vomit-related event or a nonvomit-related event is further based on identifying the patient (104) as the source of the particular event of interest.

42. The system (100) of any of claims 24-41, wherein labeling the particular event of interest as a vomit-related event or a non-vomit-related event comprises labeling theparticular event of interest as a non-vomit-related event or one of multiple different types of vomit-related events, the multiple different types of vomit-related events comprising at least one of a retch and expulsion event, a retch-only event, an expulsion-only event, a retch and cough event, a spitting event, a coughing event, or a throat-clearing event.

43. The system (100) of any of claims 24-42, wherein the ML model (200) is trained by: receiving a training dataset (401) comprising a plurality of training samples (402), each training sample (402) in the training dataset (401) comprising: a corresponding sequence of training acoustic frames (410), the corresponding sequence of training acoustic frames (410) including one or more training events of interest, each training event of interest characterizing one or more audible sounds (106), each training event of interest associated with a respective subset of sequential acoustic frames (122) segmented from the sequence of training acoustic frames (410); a corresponding ground-truth audible sound classification (420) for each training event of interest of the training sample (402); and a corresponding ground-truth label (430) for each training event of interest identifying the training event of interest of the training sample (402) as a vomit-related event or a non-vomit-related event; and for each particular training sample (402) in the training dataset (401), training the ML model (200) on the particular training sample (402) to teach the ML model (200) to learn how to predict the corresponding ground-truth audible sound classifications (420) and the corresponding ground-truth labels (430).

44. The system (100) of claim 43, wherein training the ML model (200) on a particular training sample (402) comprises: segmenting the corresponding sequence of training acoustic frames (410) of the particular training sample (402) into a plurality of training events of interest characterizing one or more audible sounds (106), each event of interest associated with arespective subset of sequential training acoustic frames (410) segmented from the sequence of training acoustic frames (410); for each particular training event of interest of the plurality of training events of interest: processing, using the ML model (200), the respective subset of sequential training acoustic frames (410) associated with the particular training event of interest to generate a corresponding probability distribution over possible audible sound classifications (202) for the particular training event of interest; labeling, using a labeler (130), based on the corresponding probability distribution over possible audible sound classifications (202), the particular training event of interest as a vomit-related event or a non-vomit-related event; determining a loss based on one or more of: a probability of the probability distribution corresponding to the corresponding ground-truth audible sound classification (420); or a difference between the corresponding ground-truth label (430) and the label (132) of the particular training event of interest; and updating a plurality of weights of the ML model (200) based on the determined loss.

45. The system (100) of any of claims 24-44, wherein the ML model (200) comprises a joint audio and speech architecture (700) that comprises a pre-trained audio encoder (712), a pre-trained speech decoder (716), and a pre-trained large language model (LLM) 750 adapted by a fine-tuned adapter module (740).

46. The system (100) of claim 45, wherein joint audio and speech architecture (700) is trained by: receiving a training dataset (401) comprising a plurality of training samples (402), each training sample (402) in the training dataset (401) comprising: a corresponding sequence of training acoustic frames (410), the corresponding sequence of training acoustic frames (410) including one or more trainingevents of interest, each training event of interest characterizing one or more audible sounds (106), each training event of interest associated with a respective subset of sequential acoustic frames (122) segmented from the sequence of training acoustic frames (410); a corresponding training natural language prompt (485); a corresponding ground-truth audible sound classification (420) for each training event of interest of the training sample (402); and a corresponding ground-truth label (430) for each training event of interest identifying the training event of interest of the training sample (402) as a vomit-related event or a non-vomit-related event; and fine-tuning the adapter module (740) on the plurality of training samples (402) by updating a plurality of weights of the adapter module (740) while parameters of the pre-trained audio encoder (712), the pre- trained speech decoder (716), and the pre- trained LLM (750) are held fixed to teach the adapter module (740) to learn how to adapt the pre-trained LLM (750) to predict the corresponding ground-truth audible sound classifications (420) and the corresponding ground-truth labels (430).

Citation Information

Patent Citations

  • Systems and methods for treatment of a patient by automated patient care

    US10682288B1

  • Sensor fusion to validate sound-producing behaviors

    US20220071588A1

  • Nausea and Vomiting Management System

    US20230395254A1