Systems and methods for clinical artificial intelligence with patient-reported multi-modal audio data

The system addresses the challenges of healthcare system inefficiencies and AI model limitations by using a processor to generate voice EHR data from patient-reported multimodal data, enabling accurate and timely clinical predictions and improving healthcare efficiency.

WO2025128765A1PCT designated stage expired Publication Date: 2025-06-19THE GOVERNMENT OF THE UNITED STATES OF AMERICA AS REPRESENTED BY THE SECRETARY DEPARTMENT OF HEALTH & HUMAN SERVICES

Patent Information

Application Number
PCT/US2024/059680
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-11
Filing Date
2024-12-11
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current healthcare systems face challenges such as long waitlists, nursing and physician shortages, and increased provider burnout, which are exacerbated by the COVID-19 pandemic. Additionally, existing AI models in healthcare are limited in their ability to process patient-reported data effectively, particularly in high-volume settings like emergency departments.

Method used

A system and method that utilize a processor to access semi-structured multimodal data from patients, including speech and language data, to generate voice EHR data. This system transcribes the voice EHR data, extracts key phrases relevant for health-related tasks, and uses a transformer AI model to generate task-specific probabilistic predictions regarding health-related treatments.

Benefits of technology

The system enables efficient collection and processing of patient-reported data, improving the accuracy and detail of health information, particularly in high-volume clinical settings. It facilitates timely and effective clinical decisions, reducing the burden on healthcare providers and improving patient outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024059680_19062025_PF_FP_ABST
    Figure US2024059680_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Various embodiments of a system for an artificial intelligence transformer model for modelling voice EHR data collected through a mobile application and then completing clinical tasks as disclosed herein.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR CLINICAL ARTIFICIAL INTELLIGENCE WITH PATIENT-REPORTED MULTI-MODAL AUDIO DATAFIELD

[0001] The present disclosure generally relates to systems and methods to facilitate the collection of voice EHR (Electronic Health Records), and particularly to an application for data acquisition of voice EHR and training a pretransformer model to perform clinical tasks based on multimodal data inputs.BACKGROUND

[0002] The COVID-19 pandemic underscored the limitations of healthcare systems and highlighted the need for data innovations to support both care providers and patients. The high volume of patients seeking medical care has caused extraordinary challenges, including long waitlists, limited time for each patient, increased testing costs, exposure risks for healthcare workers, and documentation burdens - particularly in high-volume settings like emergency departments. Adding to the problem, the world is facing nursing and physician shortages which are expected to rise dramatically over the next 10 years. This contributes to the increasing rates of provider burnout and a loss of trust in the healthcare system, both of which have been particularly severe since the onset of the COVID-19 pandemic. To address these problems, artificial intelligence (Al) has been proposed as a mechanism to rapidly perform key clinical tasks such as diagnostics, triage, and patient monitoring, improving the efficiency of the healthcare system. This has become particularly true with the advent of GPT and other multimodal large language models (LLMs), which have advanced capabilities in question answering, image interpretation, programming, and other complex tasks. As a result, technology companies have begun to develop foundation Al models for the healthcare space. However, these are mainly designed for processing and diagnostic tasks with privileged data (e.g., images) or as a Chatbot tool for general question-answering rather than providing specific recommendations from patient data. While future LLMs may add value to the healthcare system, serious data challenges remain for the widespread, equitable deployment of Al models in healthcare.

[0003] It is with these observations in mind, among others, that various aspects of the present disclosure were conceived and developed.SUMMARY

[0004] In some embodiments, a system includes a processor in communication with memory, the memory including instructions executable by the processor to access, at the processor, semi-structured multimodal data from an individual comprising speech and language data to generate voice EHR data; transcribe, at the processor, the voice EHR data related to speech and language; extract key phases, by an Al model, from the transcribed voice EHR data related to speech and language of the individual relevant for health-related tasks; extract key phrases, by the Al model, from the transcribed voice EHR data related to speech and language of the individual relevant for observed changes in voice and / or speech of the individual; extract key phrases, by the Al model, from the transcribed voice EHR data related to speech and language of the individual relevant for a health related issue; generate representation vectors, by the Al model, for extracted phrases from the transcribed voice EHR data; and generate, by the Al model, a taskspecific probabilistic prediction regarding a health-related treatment of the individual using the representation vectors of the extracted phrases from the transcribed voice EHR data.

[0005] In some embodiments, the system further includes the processor accessing semi-structured multimodal data from an individual comprising sound data from the patient to generate voice EHR data; generate a spectrogram from the voice EHR data; and extract sound features from the spectrogram.

[0006] In some embodiments, the system also generates, by an Al model, representation vectors from the extracted sound features from the spectrogram.

[0007] The system may further align the extracted phrases of the voice EHR with the extracted sound features from the spectrogram.

[0008] In some embodiments, the sound data comprises breathing tasks, prolonged phonation of vowels, and scripted words spoken by the user.

[0009] In some embodiment, an application is in operative communication with a smartphone or a web-based computer, wherein the application is operable for collecting the multimodal audio data from the individual.

[0010] In one aspect, the multimodal data includes demographic data, location data, background health data about the health of the patient, longitudinal information about the current complaint / illness, a baseline for contextualizingchanges in voice, speech, conventional acoustic data, and information from patients or providers about engagements with the healthcare systems, like physical exams of imaging studies.

[0011] In some embodiments, the system adjusts the extracted key phrases based on global and local positional information of the extracted key phrases in the sequence of extracted key phrases

[0012] In one aspect, the probabilistic prediction relates to detection of a disease in the individual and / or the likelihood of a hospital admission for the individual.

[0013] In another embodiment, the system includes a processor in communication with memory, the memory including instructions executable by the processor to access, at the processor, semi-structured multimodal data from sound data of a patient to generate voice EHR data; generate a spectrogram from the voice EHR data; extract sound features from the spectrogram; and generate representation vectors, by an Al model, for the extracted sound features; and generate, by the Al model, a task-specific probabilistic prediction regarding a health- related treatment of the individual using the representation vectors extracted from the sound features, wherein the probabilistic prediction relates to detection of a disease in the individual and / or the likelihood of a hospital admission for the individual.BRIEF DESCRIPTION OF THE DRAWINGS

[0014] FIG. 1 is a simplified illustration showing a model workflow for an Al system designed for triage / disease severity assessments in high-volume clinical settings like the emergency department.

[0015] FIG. 2 is a simplified illustration showing the workflow for the Al system of FIG. 1.

[0016] FIG. 3 is a simplified diagram showing an example computing system for implementation of the Al system of FIG. 2.

[0017] FIGS. 4A-4C are related flow diagrams showing the process flow of the Al system described in FIG. 2.

[0018] Corresponding reference characters indicate corresponding elements among the view of the drawings. The headings used in the figures do not limit the scope of the claims.DETAILED DESCRIPTION

[0019] The present system is directed to an accessible, artificial intelligence-driven tool for clinical tasks using correlated multimodal data (text, sound). The aim is that any patient or healthcare provider (including in telehealth settings) with a smartphone or a tablet can immediately receive Al-generated predictions based on sound features and patient-reported text which approximates longitudinal EHR, with a particular focus on volume management and quality of care in high-volume settings like emergency departments. To do so, an Al system, designated 100, in FIGS. 1 , 2, 3 and 4A-4C is disclosed herein that will accept semi-structured voice data in various languages as input. The present Al system 100 allows the patient to anonymously report their own time-series “voice EHR” as well as the sound data (changes in voice, speech, breathing) present in the recording. The data collection process is designed to mimic an in-person or virtual interaction with a clinician, wherein the patient provides an account of their health based on prompts / questions and the healthcare provider is also able to assess any signs or symptoms based on the voice EHR provided by the patient. The Al system 100 may be used for key clinical tasks in high volume health settings like the emergency department. Specific tasks include the assessment of patients entering the emergency department to determine likelihood of deterioration / future hospital admission and characterization of respiratory disease cases to identify indicators of infectiousness such as viral illness and variant status as shown in FIG. 1.

[0020] Smartphone and web-based application for data acquisition. In one aspect, patients can provide their own voice EHR using a smartphone or webbased application implemented on a smartphone, tablet, or computer for data acquisition by uploading recorded descriptions from the patient of (1 ) background health data and other information the patient feels may be relevant; and (2) the longitudinal trajectory of their illness such as the progression of symptoms / interventions over the course of day 1 , repeat for day 2, up until the time of submission. The smartphone and web-based application prompts the patient to describe any newly observed changes in their speech / voice, which generates an acoustic baseline. For example, a patient who smokes often may have chronic voice changes, but this would be unrelated to a recent viral infection. The patient can be prompted by the smartphone and web-based application to report any known diagnosis and input relevant demographic information (non-PII). Finally, the patient orphysician may offer a brief overview of the physical exam, diagnosis, labs, treatments, and / or other next steps. The process for collecting voice EHR will take approximately 10-12 minutes. Voice EHR data may be collected from multiple sites, including low- and middle- income countries (LMICs) due to the accessibility of the smartphone, tablet, or computer to such patients.

[0021] Data Processing with Large Language Models In some embodiments, voice EHR may include key background information, a longitudinal description of illness up to point of data collection (as a proxy for traditional timeseries EHR), descriptions of recent speech, voice, changes from the patient, and acoustic tasks to capture the sounds of the voice. In another aspect shown in the process flow of FIGS. 2 and 4A-4C, once the recorded language components of the voice EHR have been obtained the Al system 100 includes a module for transcribing these patient audio inputs at step 202. Once the voice EHR has been transcribed, a large language model with tunable, customized prompts extracts key phrases (in longitudinal order) from the transcribed voice EHR which may be relevant for health- related Al tasks at step 204 as well as other metadata if necessary (relating to the quality and nature of symptoms. The key phrase extraction operation reduces noise in the data, enhances interoperability with data from diverse sources by introducing consistency, and increases the likelihood of successful, accurate task completion in later stages. Key phrases are then extracted from the patient assessment of health history in the transcribed voice EHR at step 206, the observed change in voice / speech in the transcribed voice EHR at step 208, and the account of the main complaint / illness in the transcribed voice EHR at step 210. As shown in FIGS. 1 and 2, these key phrases / sentences 106, 108, and 110 extracted in steps 206-210 may then be embedded with a natural language processing model to capture local relationships between words in the entity at step 212, thereby enabling more detailed information to be input into the transformer Al model 101 (FIG. 1 ). For example, a qualifier like “very” which is an indicator of severity rather than the physiological nature of the symptom. The resultant embedding representation vectors may be used to form embedding matrices representing all information from the voice EHR, such as background health, current challenges, voice changes, and other relevant circumstances.

[0022] Automated Sound Processing with Deep Learning Models.As further shown in FIGS. 2 and 4C, at step 222, audio data from the patient’s voiceEHR is converted into a Mel spectrogram 112, which is a representation of sound that contains information on the intensity of different frequencies for each time-point in the original waveform. Deep learning models are used to extract sound features from the spectrogram 112 of sound data at step 224. These methods are used to encode the audio data as representation vectors for the sound features at step 226 which can be aligned with text embedding representation vectors at step 228 and input into the customized transformer Al model 101 described below at steps 216, 218 and 220. At this stage, the data is now in the correct format for input into a trained transformer Al model 101 which outputs an identification of complex health biomarkers which involve both acoustic and semantic insights.

[0023] Training of a transformer Al model. The transformer Al model 101 is used to perform a task such as a “severity assessment of general diseases” based on voice, speech, and language information contained in the voice EHR of a patient. Within the training process of the customized transformer Al model 101 , the representations of text and sound as noted above are merged for the purpose of extracting higher-order biomarkers via attention mechanisms at step 228. At step 214, specific encodings are used to adjust the extracted key phrase / sentence embeddings based on global and local positional information such as the position of the phrases in the sequence of phrases and the overall position of the information within voice EHR (i.e., a history of chest pain has a different contextual relevance than current, ongoing chest pain at step 212. At step 218, a multilayer perceptron (MLP) neural network is used to convert the key information encoded by the transformer Al model 101 into a task-specific probabilistic prediction at step 220 which can be interpreted to determine, for example, whether the patient’s condition is mild or severe which would likely require hospital admission. A different multilayer perceptron is used depending on the task (each MLP neural network is taskspecific). A generative component of the Al system 100 may be used to better interpret and communicate transformer Al model 101 results, thereby supporting patient and physician users.Use Cases for Al Predictions

[0024] As shown in FIG. 1 , the Al system 100 can be implemented in settings such as emergency departments to triage patients, determining which patients are likely to need admission, which patients may have a contagiousinfectious disease, and / or which can be referred to outpatient care (health management and resource allocation). In one example, the Al system 100 may detect that, based on reported symptoms and voice changes, that the patient has new-onset atrial fibrillation and should be admitted for evaluation / monitoring as shown at block 116 of FIG. 1 . In another case, the Al system 100 may assess that, based on the reported symptoms and distinct laryngitis, that the patient has a contagious COVID-19 variant (i.e., Omicron) and should self-quarantine rather than be admitted to the hospital as shown in block 118 of FIG. 1.Strategy

[0025] In one aspect, the smartphone and web-based application for voice EHR data acquisition and training the transformer Al model 101 may include the following features:1 ) Preprocess data to ensure suitability for input into a transformer Al model 101.2) Training of a customized transformer Al model 101 for severity assessment of general diseases from a high-volume setting like an emergency department. Predictions generated by the transformer Al model 101 may be based on sound data and voice EHR.3) Validation of the transformer Al model 101 on held-out test data, including data collected from resource-constrained settings.Hypotheses

[0026] The following hypotheses were considered when developing the smartphone and web-based application disclosed herein for data acquisition and the pre-trained transformer model as described herein:

[0027] Hypothesis #1: “Voice EHR” will contain more detailed and accurate information compared to conventional data types, particularly in low- resource settings.

[0028] Hypothesis #2 Sound data with clinical context and longitudinal depth will be significantly more useful than unimodal sound data. Beyond binary screening for illnesses (the main use case of current voice Al tools), this data will facilitate specific classifications and assessments of diseases in high-volume settings like emergency departments, outperforming existing methods. Moreover, thediverse dataset will reduce model biases against groups who are often excluded from consideration in Al studies.Preliminary Results

[0029] In initial pilot studies involving the present data and Al system 100, the following results were obtained: (1 ) a voice EHR dataset collected from multiple sites, including the emergency department at Tampa General Hospital (TGH) was found to be significant more informative than data collected from patients via conventional means (e.g., short answer, multiple choice questions). Through the use of an automated rating pipeline involving multimodal large language models, voice EHR was found to be equally or more informative (a score of 3 / 5 or higher) in the context of a clinical assessment (i.e., such as those performed in the emergency department) in over 80% of cases. In nearly 60% of considered cases, voice EHR was found to be significantly more informative than conventional data (score of 5 / 5). In a second pilot study, the keyword extraction pipeline paired with a customized transformer model led to an AUROC score of approx. 0.80 at a task involving infectious disease characterization based on viral strain, similar to the concept of triaging diseases in a high-volume setting based on factors like contagion potential and severity / likelihood of hospital admission. These results were achieved using noisy, fully unstructured data from online sources which was similar to voice EHR (speakers describing their illness). Such outcomes implied that the use of semistructured voice EHR, rather than fully uniform or fully free-form data, will not only offer valuable context through the longitudinal health information but will also improve the voice / speech analysis by capturing diverse acoustic features for each patient.Methods and Experimental Plan

[0030] The data collected from the smartphone and web-based application was quality checked and preprocessed to prepare for using within the transformer-based systems. Following additional validation studies, high-performing transformer models may be connected to an interface for real-time use in clinical settings like emergency departments and tested by data collection partners (i.e., Tampa General Hospital).Data Collection

[0031] The present smartphone and web-based application disclosed herein for data collection may be operable through either a web-based application using on a computer or a mobile application through a smartphone or tablet. The smartphone and web-based application may be operable for collecting the following data, such as semi-structured voice EHR and conventional voice data (elongated vowels, scripts) for comparison purposes. No personal identifiers of the patients were included in the research data.Demographic and Clinical Data

[0032] Examples of demographic and clinical data may include one or more of the following: i. Age ii. Sex at birth iii. Racial / ethnic identification iv. Insurance status v. Highest level of education vi. Zip code or other approximate indicator of location vii. Probable diagnosis and / or laboratory-confirmed diagnosis viii. Selected co-morbidities (from a multiple-choice section) ix. Selected symptoms (from a multiple-choice section) x. Smoking history xi. Duration of illness xii. Progression of illness (improving / stable / deteriorating)Voice-EHR: Patient-submitted voice recordings

[0033] In some embodiments, a subject provides a recorded response to background information about their health that the subject feels may be relevant in response to the following prompts. This establishes a baseline from the health history of the patient which may be used to identify key changes due to illness.Prompt: Please tell us background information about your health before your current illness, including Chronic conditions (such as high blood pressure or diabetes), Recent illnesses (for example, COVID-19), Other physical health problems, mental health problems, such as anxiety, medications you currently take, and any recent changes to your medication which made you feel differently.Prompt: Is there anything else you would like us to know about your health or circumstances that you feel we have missed?For example, you can tell us about: your employment, your lifestyle habits, and / or any challenges you have had with the healthcare system, including delays with receiving care or problems with quality of care that may have impacted your health.

[0034] In some embodiments, the subject describes signs, symptoms, and interventions (longitudinally) in detail in response to the following prompt. This prompt may capture the complaint of the patient by approximating a longitudinal record of illness progression.Prompt: In as much detail as possible, please tell us how your illness has developed from the time when you first noticed symptoms until now. Include any medications you took (like Tylenol) or steps you use to reduce your symptoms. Please use words / phrases like "on the first day", "in the morning", "then", "after that" and use descriptive words like "mild", "severe". No detail is too small.

[0035] In some embodiments, a voice-specific baseline of the subject is established, wherein the subject reports if they or anyone else have noticed recent voice or speech changes in response to the following prompt. This may establish an audio baseline to potentially identify longer-term changes in the voice or speech - for example, stemming from lifestyle factors like smoking behaviors.Prompt: Please tell us if you or anyone else has noticed any recent changes in your voice (like hoarse, raspy, or lost voice) speech (like difficulty getting words out or slurring words), or breathing. If so, describe these changes. These should be changes that started around the same time as this illness episode, not any chronic long-term changes.

[0036] In some embodiments, conventional acoustic data is collected to characterize changes in breathing, the voice (sound produced by vibration of the vocal cords) and speech (the physical, neuromuscular production of sound to form words).Prompt: Say each of these vowels for as long as you can: aaaaa (as in made); eeeee (beet); ooooo (cool)Prompt: Read these sentences: “When the sunlight strikes raindrops in the air, they act as a prism and form a rainbow. The rainbow is a division of whitelight into many beautiful colors. These take the shape of a long round arch, with its path high above, and its two ends apparently beyond the horizon.”Prompt: Hold the device near your nose and record yourself breathing normally for 30 seconds with your mouth closed.Prompt: Hold the device near your mouth and record yourself taking 3 deep breaths through your mouth.

[0037] In some embodiments, the subject or physician describes the physical exams, lab results, diagnosis, treatment plan, and / or any next steps in response to the following prompt. This prompt may provide an audio approximation of other multimodal data types which may be key context for patient-reported data collected by the application.Prompt: Your physician or other provider should briefly describe the physical exam (given to you by the physician), any available lab results, imaging studies, the diagnosis, hospital admission / discharge status, and other next steps related to testing, treatment, or monitoring the illness. If the healthcare provider is not available or you are at home, you can record this information yourself.

[0038] If necessary, physicians can ask additional questions during the recordings, to guide the participant towards providing more relevant information. These will be included as inputs to the model.

[0039] In one aspect, voice EHR data is structured to mirror a visit to a medical professional or other point of care setting. The use of voice EHR introduces numerous potential benefits, including those which may improve the overall performance of Al models and reduce bias towards underrepresented groups.

[0040] In one aspect, the smartphone and web-based application described herein facilitates the rapid collection of “voice EHR” data in a user-friendly way, without 1 ) requiring time-consuming and error-prone text data entry on the part of the individual, and 2) enforcing a rigid, pre-defined data schema found in traditional EHR, which may limit the incorporation of information which the patient considers to be important. Furthermore, the process of creating a “voice EHR” may be useful to healthcare workers.

[0041] With the introduction of text-sound correlates, voice EHR may additionally compensate for sources of confusion that are often found in clinical datathrough “biomarker reinforcement.” Even if participants provide incoherent and / or incomplete data in terms of semantic meaning, the smartphone and web-based application still captures voice and breathing data which may independently contribute to the robustness of the data. For example, lapses in patient memory, incomplete notes from healthcare workers, or information reported in colloquial terminology may compromise the value of language data, but acoustic features from the voice may be unaffected in these scenarios. The converse may also be true, in which transcripts of patient-reported health information still provide usable data despite background noise or recording errors (e.g., the device was held too far from the mouth).

[0042] For cases in which both modalities are viable, the use of voice / sound data in combination with transcribed health information can capture a more comprehensive composite of diseases with diverse phenotypes, particularly at the time of presentation. Sound data may contribute biomarkers for certain diseases which would not currently be captured in clinician notes. Ultimately, self-reported multimodal audio data expands upon basic health information traditionally used for digital health systems and may allow Al models to better consider chronic conditions, voice changes, speech patterns, word choices indicating mood / sentiment, potential exposures, behavioral influences, and specific disease progression. Compared to similar methods like ambient listening, semi-structured voice EHR may also reduce the variability of multimodal audio data, potentially enabling machine learning modelling from a smaller sample size. This methodology may also reduce Al biases against clinics / healthcare environments which do not engage in conventional workflows or styles of patient interaction (which could reduce the value of ambient listening in these settings).

[0043] In under-resourced settings, the EHR is often incomplete, incorrect, or “low-tech”, which disadvantages patients who may rely on EHR-driven Al technologies developed in high-income settings. The smartphone and web-based application will facilitate the collection of time-series “EHR” data in a fast and accessible way, rather than requiring the patient to type large amounts of text into a lengthy form, which would likely decrease patient participation. The smartphone and web-based application collects key components of a health assessment, including background information, a baseline for contextualizing changes in voice / speech, and a longitudinal account of the illness.

[0044] The use of a transformer Al model 101 to parse patient inputs and perform clinical tasks can address various common challenges which have been shown to obstruct physician-patient communication, particularly in understaffed settings. These include overwhelmingly long and complex lists of symptoms, a perceived lack of empathy from the physician, and lack of physician insight into the condition despite receiving detailed information from the patient.

[0045] The data collection process is deployed through two primary channels: 1) public use of the smartphone and web-based application; and 2) partnerships with healthcare professionals working at pharmacies, walk-in / primary care clinics, ICUs, emergency departments, and other point-of care settings. Collaboration with healthcare professionals will likely improve the reliability of data and expand access to participants with diagnoses (labels for the training data). The smartphone and web-based application can be used broadly within a healthcare ecosystem, including physicians, physician’s assistants, nurses, technicians, clinical researchers, medical students, and patients.Data Preprocessing

[0046] In one aspect, the first step in the preprocessing pipeline will be the removal of data that does not meet the criteria for inclusion in the study. For example, submissions with significant amounts of missing data will be excluded from downstream analyses. To optimize participant compliance and ensure access to data with corresponding diagnoses, partnerships have been established with healthcare providers working at point-of-care settings. Healthcare professionals supporting this effort will assist participants with data collection and may input descriptive data on diagnoses / treatment plans.

[0047] Audio data will be preprocessed, acoustic signals normalized, converted into spectrograms, and the embedded via Al models pre-trained on large sound datasets. Transcripts of the recordings will then be extracted using speech-to- text Al models such as Whisper. Once the transcriptions have been extracted, a large language model will be used to remove invalid data. For example, a submission would be excluded for containing a description of obviously irrelevant symptoms, a description of a topic clearly unrelated to health, or an incomprehensible transcription due to low-quality audio. The feasibility of this step was assessed in a prior study, and the task was completed accurately with the largelanguage model. Data will then undergo key phrase / sentence extraction and embedding prior to input into the transformer Al model 101.Model Fine-tuning and Evaluation

[0048] During the training process, transformer Al models 101 are be adjusted based on the error of the output using standard loss and optimization functions.

[0049] To assess the reliability of the transformer Al model 101 , multiple validation and test sets will be generated using random sampling without replacement. Standard evaluation metrics will be used, including sensitivity, specificity, and AUG. Metrics will be reported for each data collection site, to ensure a balanced performance.

[0050] As further shown, the sample model workflow for the Al system 100 may perform a severity assessment for triage / volume management of the patient in an emergency department. Spectrograms 112 containing acoustic features from recorded sound made by the patient, such as the phonation of an elongated vowel, are aligned with the transcribed voice EHR, including key phrases / sentences extracted from patient-reported background health information 106 at step 206, key phrases extracted from the observed changes in speech / voice 110 (“acoustic baseline”) at step 210, and key phrases / sentences extracted from the time-series voice EHR data on symptom progression 108 at step 208 are all used to ensure a robust prediction by the Al system 100 related to the task at step 218. The final output of the Al system 100 is an interpreted probability value at step 220, which is provided to a generative Al model given the patient context and is an output of a multilayer perception (MLP) neural network at step 218. The MLP learned insights and derived a probability value from the full scope of the patient-reported data and the objective audio of the patient. This reduces the likelihood of different possible biases, including a lack of consideration for information which the patient considers to be important. The nature of the probabilistic outputs varies based on the task but may include: (1) indications of the severity of disease and the subsequent likelihood of hospital admission / deterioration shown at block 116 of FIG. 1 and (2) indication of the infectiousness of the illness and the subsequent type of infection (e.g., Omicron COVID-19) shown at block 118 of FIG. 1.Computer-Implemented System

[0051] FIG. 3 is a schematic block diagram of an example device 202 that may be used with one or more embodiments described herein, e.g., implementing the Al system 100 shown in FIGS 1 and 2. In some embodiments, the example device 202 may be either a smartphone, tablet, or computer in which the smartphone and web-based application discussed above is implemented.

[0052] Device 202 comprises one or more network interfaces 210 (e.g., wired, wireless, PLC, etc.), at least one processor 220, and a memory 240 interconnected by a system bus 250, as well as a power supply 260 (e.g., battery, plug-in, etc.) in which the device 202 is in operable association with the transformer Al model 101 shown in FIG. 1.

[0053] Network interface(s) 210 include the mechanical, electrical, and signaling circuitry for communicating data over the communication links coupled to a communication network. Network interfaces 210 are configured to transmit and / or receive data using a variety of different communication protocols. As illustrated, the box representing network interfaces 210 is shown for simplicity, and it is appreciated that such interfaces may represent different types of network connections such as wireless and wired (physical) connections. Network interfaces 210 are shown separately from power supply 260, however it is appreciated that the interfaces that support PLC protocols may communicate through power supply 260 and / or may be an integral component coupled to power supply 260.

[0054] Memory 240 includes a plurality of storage locations that are addressable by processor 220 and network interfaces 210 for storing software programs and data structures associated with the embodiments described herein. In some embodiments, device 202 may have limited memory or no memory (e.g., no memory for storage other than for programs / processes operating on the device and associated caches). Memory 240 can include instructions executable by the processor 220 that, when executed by the processor 220, cause the processor 220 to implement aspects of the system and the methods outlined herein.

[0055] Processor 220 comprises hardware elements or logic adapted to execute the software programs (e.g., instructions) and manipulate data structures 245. An operating system 242, portions of which are typically resident in memory 240 and executed by the processor, functionally organizes device 202 by, inter alia,invoking operations in support of software processes and / or services 290 executing on the device 202. These software processes and / or services may include taskspecific prompts and descriptive prediction processes / services 290, which can include aspects of system 100 having the Transformer Al model 102 illustrated in FIG. 1 and / or implementations of various modules described herein. Note that while task-specific prompts and Al-driven prediction processes / services 290 are illustrated in centralized memory 240, the Al-driven prediction processes / services 290 may be in operable association with the transformer Al model 101 for providing the functions disclosed herein. Alternative embodiments provide for the process to be operated within the network interfaces 210, such as a component of a MAC layer, and / or as part of a distributed computing network environment.

[0056] It will be apparent to those skilled in the art that other processor and memory types, including various computer-readable media, may be used to store and execute program instructions pertaining to the techniques described herein. Also, while the description illustrates various processes, it is expressly contemplated that various processes may be embodied as modules or engines configured to operate in accordance with the techniques herein (e.g., according to the functionality of a similar process). In this context, the term module and engine may be interchangeable. In general, the term module or engine refers to model or an organization of interrelated software components / functions. Further, while the task-specific prompts and Al-driven prediction processes / services 290 is shown as a standalone process of an Al system 100, those skilled in the art will appreciate that this process may be executed as a routine or module within other processes.

[0057] It should be understood from the foregoing that, while particular embodiments have been illustrated and described, various modifications can be made thereto without departing from the spirit and scope of the invention as will be apparent to those skilled in the art. Such changes and modifications are within the scope and teachings of this invention as defined in the claims appended hereto.

Claims

CLAIMSWhat is claimed is:

1. A system comprising: a processor in communication with memory, the memory including instructions executable by the processor to: access, at the processor, semi-structured multimodal data from an individual comprising speech and language data to generate voice EHR data; transcribe, at the processor, the voice EHR data related to speech and language; extract key phases, by an Al model, from the transcribed voice EHR data related to speech and language of the individual relevant for health-related tasks; extract key phrases, by the Al model, from the transcribed voice EHR data related to speech and language of the individual relevant for observed changes in voice and / or speech of the individual; extract key phrases, by the Al model, from the transcribed voice EHR data related to speech and language of the individual relevant for a health related issue; generate representation vectors, by the Al model, for extracted phrases from the transcribed voice EHR data; and generate, by the Al model, a task-specific probabilistic prediction regarding a health-related treatment of the individual using the representation vectors of the extracted phrases from the transcribed voice EHR data.

2. The system of claim 1 , further comprising: access, at the processor, semi-structured multimodal data from an individual comprising sound data from the patient to generate voice EHR data; generate a spectrogram from the voice EHR data; and extract sound features from the spectrogram.

3. The system of claim 2, further comprising: generate, by an Al model, representation vectors from the extracted sound features from the spectrogram.

4. The system of claim 3, further comprising: aligning the extracted phrases of the voice EHR with the extracted sound features from the spectrogram.

5. The system of claim 2, wherein the sound data comprises breathing tasks, prolonged phonation of vowels, and scripted words spoken by the user.

6. The system of claim 1 , further comprising: an application in operative communication with a smartphone or a webbased computer, wherein the application is operable for collecting the multimodal audio data from the individual.

7. The system of claim 6, wherein multimodal data comprises demographic data, location data, background health data about the health of the patient, longitudinal information about the current complaint / illness, a baseline for contextualizing changes in voice, speech, conventional acoustic data, and information from patients or providers about engagements with the healthcare systems, like physical exams of imaging studies.

8. The system of claim 1 , further comprising: adjust the extracted key phrases based on global and local positional information of the extracted key phrases in the sequence of extracted key phrases9. The system of claim 1 , wherein the probabilistic prediction relates to detection of a disease in the individual.

10. The system of claim 1 , wherein the probabilistic prediction relates to the likelihood of a hospital admission for the individual.

11. A system comprising: a processor in communication with memory, the memory including instructions executable by the processor to: access, at the processor, semi-structured multimodal data from sound data of a patient to generate voice EHR data; generate a spectrogram from the voice EHR data; extract sound features from the spectrogram; generate representation vectors, by an Al model, for the extracted sound features; and generate, by the Al model, a task-specific probabilistic prediction regarding a health-related treatment of the individual using the representation vectors extracted from the sound features.

12. The system of claim 11 , wherein the probabilistic prediction relates to detection of a disease in the individual.

13. The system of claim 11 , wherein the probabilistic prediction relates to the likelihood of a hospital admission for the individual.

Citation Information

Patent Citations

  • Systems and methods for mental health assessment

    US20210110894A1

  • Speech analysis for monitoring or diagnosis of a health condition

    US20230255553A1

Cited By

  • Medical dialogue context integrity guarantee and application method and system and storage medium

    CN122494200A