A system and method for predicting mental health status based on speech / text and language processing.
A system that analyzes patient-provider conversations using acoustic and linguistic models addresses the limitations of self-reported surveys by providing real-time, objective mental health assessments, improving accuracy and efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ELLIPSIS HEALTH INC
- Filing Date
- 2024-06-13
- Publication Date
- 2026-07-24
AI Technical Summary
Current mental health screening tools, such as the PHQ-9 and GAD-7, rely on self-reported surveys that are repetitive, lack individualization, and are often incomplete or biased, leading to inaccurate assessments and increased healthcare costs.
A system that passively analyzes conversations between patients and healthcare providers using acoustic and linguistic models to predict mental health symptoms, identifying relevant topics and speaker roles, and generating electronic reports for accurate symptom severity assessment.
Provides real-time, objective, and precise mental health assessments without additional effort, reducing workflow disruption and improving patient outcomes by enhancing the accuracy and efficiency of mental health evaluations.
Smart Images

Figure 2026524831000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 507,973, filed on June 13, 2023, entitled "Systems and Methods for Predicting States Based on Conversations", and U.S. Provisional Patent Application No. 63 / 571,398, filed on March 28, 2024, entitled "Systems and Methods for Predicting Mental Health States Based on Passive Processing of Speech and Language in Conversations", the contents of which are incorporated herein by reference.
[0002] This disclosure generally relates to the field of health assessment, and more specifically, embodiments relate to devices, systems, and methods that use artificial intelligence to predict the severity of symptoms of mental health states in patients based on conversations with a healthcare provider.
Background Art
[0003] Introduction Behavioral health is a serious concern. The most widely used tools for screening behavioral health status can rely on accurate self - reporting by the screened patients. For example, the current "gold standard" for screening based on a depression questionnaire is the Patient Health Questionnaire - 9 (PHQ - 9), a written depression health survey that includes nine multiple - choice questions. Other similar health surveys include the Patient Health Questionnaire - 2 (PHQ - 2) and the Generalized Anxiety Disorder 7 (GAD - 7).
[0004] These and other screening or monitoring surveys may be unmotivating due to their repetitive nature and lack of individualization. They may also lack an element of objectivity because they are self-reported. Assessing a patient's behavioral health can be difficult if no assessment questionnaire is provided or if the survey questions are incomplete. Patients may become mechanical in their responses after receiving the same questions over multiple sessions, making it difficult to assess their progress. Patients may also respond to surveys in a biased manner, either to influence the therapist or to conceal their condition due to fear of unemployment or discrimination.
[0005] Finally, completing these surveys requires effort from both the person conducting them (e.g., healthcare team members, case managers, etc.) and the patient (e.g., some patients may require assistance), which disrupts the workflow for both. Often, these surveys are conducted verbally and may constitute a significant portion of the clinical consultation. Furthermore, when these surveys are conducted verbally, they are often conducted improperly (e.g., not all questions are asked, questions are not asked verbatim, correct answer choices are not provided, etc.), rendering their results essentially invalid. This can lead to increased healthcare costs for patients, providers, and payers. In some cases, it can also lead to worse patient outcomes (e.g., if inappropriate stratification due to incorrect scores prevents or delays patients from seeking appropriate care).
[0006] Improvements are needed in the area of patient case management / care management. [Overview of the project]
[0007] Screening or monitoring surveys may not be well-received due to their repetitive nature and lack of individualization, and they are not always objective because they are self-reported. This problem can be further exacerbated if these questions are delivered in written survey format or by automated questionnaire distribution systems. Patients may be more willing to participate if the questionnaire questions are delivered by a human (e.g., a staff member) based on the patient's situation. It could be even more beneficial if answers to the questions could be elicited from existing conversations with the patient.
[0008] The caregiver may be too busy talking to the patient to simultaneously extract relevant points from the conversation. Furthermore, the caregiver may not be able to elicit relevant answers in a timely manner, or in some cases, not at all. A robust and automated system or method is needed to passively listen to the conversation between the caregiver and the patient in order to extract relevant answers or other relevant information.
[0009] If assessment questionnaires are not provided, or if the survey questions are incomplete, it may be difficult to assess a patient's behavioral health. Responsible personnel may be able to elicit relevant responses from patients using conversational approaches that can make progress assessment easier. The systems and methods described herein may further aid in predicting a patient's condition by eliciting relevant topics and relevant portions of conversations. The systems and methods described herein can utilize existing conversations to conduct assessments, eliminating the need for verbatim assessments.
[0010] Finally, completing the study requires effort from both the clinical team members and the patient (for example, some patients may need assistance to complete it), which disrupts the workflow for both the practitioner and the patient. Providing a system or method for performing session analysis, transcription, and annotation work helps practitioners provide rapid (e.g., real-time or just-in-time), accurate, and precise inquiries and treatments. This not only leads to better patient outcomes but can also further free up practitioners' ability to evaluate more patients in a shorter period of time or to spend the saved time addressing the concerns or needs of other patients.
[0011] In one embodiment, a system is provided for analyzing conversation (voice and / or text) to predict the severity of symptoms. The system comprises at least one input device for receiving conversation data from at least one user, at least one output device for outputting an electronic report, and at least one computing device for communicating with at least one input device and at least one output device. The at least one computing device is configured to receive conversation data from at least one input device, process the conversation data to generate a language model output and / or an acoustic model output, optionally apply weights to the language model output and the acoustic model output, and merge the language model output and the acoustic model output by generating a composite output from the weighted output, generate an electronic report, and transmit the electronic report to the output device.
[0012] In one embodiment, a system is provided for identifying the roles of speakers in a conversation. The system comprises at least one input device for receiving conversational data from at least one user, at least one output device for outputting an electronic report, and at least one computing device for communicating with the at least one input device and the at least one output device. The at least one computing device is configured to receive conversational data from the at least one input device, determine at least one role of at least one speaker, process the conversational data to generate a language model output and / or acoustic model output, apply weights to the language model output and / or acoustic model output, each of which includes multiple outputs corresponding to multiple time segments of conversational data, weights which are optionally time-based and partially based on at least one role of at least one speaker in each time segment, generate an electronic report, and transmit the electronic report to the output device.
[0013] In some embodiments, at least one computing device is further configured to merge weighted language model outputs and acoustic model outputs to produce a composite output. The composite output may represent a fused output from the fusion of the language model outputs and acoustic model outputs.
[0014] In some embodiments, the electronic report may, based on the combined output, identify the severity of at least one symptom of the condition.
[0015] In some embodiments, the state may include a mental health state.
[0016] In some embodiments, the electronic report includes annotations of language model output that indicate the prominence of at least one of one or more time segments.
[0017] In some embodiments, at least one speaker includes an agent, and applying weights to the language model output and acoustic model output includes applying a zero weight to the acoustic model output corresponding to the agent.
[0018] In some embodiments, the weights are partially based on the topics in each time segment.
[0019] In some embodiments, the language model output includes the identification of at least one query based on conversational data from the person in charge, and at least one response to at least one query based on conversational data from the patient, wherein at least one query is mapped to a predetermined query, and at least one response to at least one query is mapped to a predetermined response to the predetermined query.
[0020] In some embodiments, processing conversational data to generate language model outputs and acoustic model outputs involves using language model neural networks and acoustic neural networks trained on labeled conversational data collected from one or more other subjects, wherein the labeled conversational data is labeled as (i) having a state at some level and (ii) not having that state.
[0021] In some embodiments, at least one role of at least one speaker includes at least one of patient, agent, interactive voice response, and bot speaker.
[0022] In some embodiments, the weights applied to the language model output and the acoustic model output are partially based on determining whether the number of speakers matches the expected number of speakers.
[0023] In some embodiments, conversation data is processed based in part on at least one of a patient profile and a practitioner profile, the profile including at least one of historical data, career data, demographic data, and longitudinal data.
[0024] In some embodiments, the conversation data includes at least one of utterance data and text-based data.
[0025] In some embodiments, at least one computing device is configured to run a model based on human-interpretable features.
[0026] In one embodiment, a system is provided for identifying topics in a conversation. The system comprises at least one input device for receiving conversation data from at least one user, at least one output device for outputting an electronic report, and at least one computing device for communicating with the at least one input device and the at least one output device. The at least one computing device is configured to receive conversation data from the at least one input device, process the conversation data to generate a language model output, the language model output including one or more topics corresponding to one or more time ranges, to generate a weighted output, the output including multiple outputs corresponding to multiple time segments of the conversation data, to which weights, optionally time-based and partially based on one or more topics in each time segment are applied, to the output, to generate a weighted output, to generate an electronic report, and transmit the electronic report to the output device.
[0027] In some embodiments, at least one computing device is configured to fuse a language model output and an acoustic model output by processing conversation data to generate an acoustic model output, applying weights to the language model output and the acoustic model output, and generating a composite output from the weighted outputs, where the acoustic model output includes a plurality of outputs corresponding to a plurality of time segments of the conversation data, the weights are optionally time-based, and are partially based on the topic within each time segment.
[0028] In some embodiments, the time range of conversation data corresponding to a given topic is processed to generate a language model output using a more computationally robust model than that used for the time range of conversation data not corresponding to the given topic.
[0029] In some embodiments, the electronic report includes a transcription of the language model output annotated partially based on weights that are partially based on the topic within each time segment.
[0030] In some embodiments, the electronic report identifies the severity of at least one symptom of a condition based on the language model output.
[0031] [[ID=H16]]In some embodiments, the condition includes a mental health condition.
[0032] In some embodiments, the electronic report includes an annotation of the language model output indicating the salience of at least one of the one or more time segments.
[0033] In some embodiments, the computing device is further configured to determine at least one role of at least one speaker, and the weights are partially based on at least one role of at least one speaker within each time segment. <00001H7>
[0034] In some embodiments, the language model output includes the identification of at least one query based on conversational data from the person in charge, and at least one response to at least one query based on conversational data from the patient, wherein at least one query is mapped to a predetermined query, and at least one response to at least one query is mapped to a predetermined response to the predetermined query.
[0035] In some embodiments, processing conversational data to generate language model output involves using a language model neural network trained on labeled conversational data collected from one or more other subjects, where the labeled conversational data is labeled as (i) having a state at some level and (ii) not having that state.
[0036] In some embodiments, at least one query belongs to a set of queries, and the computing device is configured to predict an overall score based on the set of queries based on the response to each of the queries in the set.
[0037] In some embodiments, conversation data is processed based in part on at least one of a patient profile and a practitioner profile, the profile including at least one of historical data, career data, demographic data, and longitudinal data.
[0038] In some embodiments, the conversation data includes at least one of utterance data and text-based data.
[0039] In some embodiments, at least one computing device is configured to run a model based on human-interpretable features.
[0040] In one embodiment, a system is provided for scoring surveys based on conversations. The system comprises at least one input device for receiving conversation data from at least one user, at least one output device for outputting an electronic report, and at least one computing device for communicating with the at least one input device and the at least one output device. The at least one computing device is configured to receive conversation data from the at least one input device, process the conversation data, and generate a language model output, the language model output comprising: an identification of at least one query based on conversation data from a person in charge; at least one query based on conversation data from a patient, the at least one query being mapped to a predetermined query, the at least one response to the at least one query being mapped to a predetermined response to the predetermined query; generate an electronic report; and transmit the electronic report to the output device.
[0041] In some embodiments, at least one computing device is further configured to process conversational data to generate an acoustic model output, apply weights to the language model output and the acoustic model output, and merge the language model output and the acoustic model output by generating a composite output from the weighted output, wherein the language model output and the acoustic model output each include multiple outputs corresponding to multiple time segments of the conversational data, and the weights are optionally time-based.
[0042] In some embodiments, the weights are partially based on at least one query or topic within each time segment.
[0043] In some embodiments, the system is configured to output at least one unresolved query to an output device. An unresolved query may be one for which the user has not provided a response.
[0044] In some embodiments, the system is configured to output a flag indicating that at least one response to at least one query does not map to a given response to a given query with a confidence level above a threshold.
[0045] In some embodiments, the system is configured to prompt the system to repeat a given query if it is determined with a confidence level above a threshold that at least one response to at least one query does not map to a given response to a given query.
[0046] In some embodiments, the electronic report identifies the severity of at least one symptom of the condition based on the language model output.
[0047] In some embodiments, the state includes a mental health state.
[0048] In some embodiments, the electronic report includes annotations of language model output that indicate the prominence of at least one of one or more time segments.
[0049] In some embodiments, the computing device is further configured to determine at least one role of at least one speaker, and the weights are partially based on at least one role of at least one speaker during each time segment.
[0050] In some embodiments, processing conversational data to generate language model output involves using a language model neural network trained on labeled conversational data collected from one or more other subjects, where the labeled conversational data is labeled as (i) having a state at some level and (ii) not having that state.
[0051] In some embodiments, conversation data is processed based in part on at least one of a patient profile and a practitioner profile, the profile including at least one of historical data, career data, demographic data, and longitudinal data.
[0052] In some embodiments, the conversation data includes at least one of utterance data and text-based data.
[0053] In some embodiments, at least one computing device is configured to run a model based on human-interpretable features.
[0054] In one embodiment, a method is provided for identifying the roles of speakers in a conversation. The method includes receiving conversational data from at least one input device; determining at least one role of at least one speaker; processing the conversational data to generate language model outputs and / or acoustic model outputs; applying weights to the language model outputs and / or acoustic model outputs, wherein each language model output and / or acoustic model output includes multiple outputs corresponding to multiple time segments of the conversational data, and the weights are optionally time-based and partially based on at least one role of at least one speaker in each time segment; generating an electronic report; and transmitting the electronic report to an output device.
[0055] In some embodiments, the process further includes fusing weighted language model outputs with acoustic model outputs to generate a composite output. The composite output may represent a fused output from the fusing of the language model outputs and acoustic model outputs.
[0056] In some embodiments, the electronic report identifies the severity of at least one symptom of the condition based on the combined output.
[0057] In some embodiments, the state includes a mental health state.
[0058] In some embodiments, the electronic report includes annotations of language model output that indicate the prominence of at least one of one or more time segments.
[0059] In some embodiments, at least one speaker includes an agent, and applying weights to the language model output and acoustic model output includes applying a zero weight to the acoustic model output corresponding to the agent.
[0060] In some embodiments, the weights are partially based on the topics in each time segment.
[0061] In some embodiments, the language model output includes the identification of at least one query based on conversational data from the person in charge, and at least one response to at least one query based on conversational data from the patient, wherein at least one query is mapped to a predetermined query, and at least one response to at least one query is mapped to a predetermined response to the predetermined query.
[0062] In some embodiments, processing conversational data to generate language model outputs and acoustic model outputs involves using language model neural networks and acoustic neural networks trained on labeled conversational data collected from one or more other subjects, wherein the labeled conversational data is labeled as (i) having a state at some level and (ii) not having that state.
[0063] In some embodiments, at least one role of at least one speaker includes at least one of patient, agent, interactive voice response speaker, and bot speaker.
[0064] In some embodiments, the weights applied to the language model output and the acoustic model output are partially based on determining whether the number of speakers matches the expected number of speakers.
[0065] In some embodiments, conversation data is processed based in part on at least one of a patient profile and an agent profile, the profile including at least one of historical data, career data, demographic data, and / or longitudinal data.
[0066] In some embodiments, the conversation data includes at least one of utterance data and text-based data.
[0067] In some embodiments, at least one computing device is configured to run a model based on human-interpretable features.
[0068] In one embodiment, a method is provided for identifying topics in a conversation. The method includes receiving conversation data from at least one input device; processing the conversation data to generate a language model output, wherein the language model output includes one or more topics corresponding to one or more time ranges; applying weights to the output to generate a weighted output, wherein the output includes multiple outputs corresponding to multiple time segments of the conversation data, and the weights are optionally time-based and partially based on one or more topics in each time segment; generating an electronic report; and transmitting the electronic report to an output device.
[0069] In some embodiments, the method further includes processing conversational data to generate an acoustic model output, and fusing the language model output and the acoustic model output by applying weights to the language model output and the acoustic model output and generating a composite output from the weighted output, wherein the acoustic model output includes multiple outputs corresponding to multiple time segments of the conversational data, and the weights are optionally time-based and partially based on the topics in each time segment.
[0070] In some embodiments, the time range of conversational data corresponding to a given topic is processed using a computationally more robust model than the one used for the time range of conversational data not corresponding to the given topic, in order to generate language model output.
[0071] In some embodiments, the electronic report includes a transcript of language model output annotated with weights partially based on the topics in each time segment.
[0072] In some embodiments, the electronic report identifies the severity of at least one symptom of the condition based on the language model output.
[0073] In some embodiments, the state includes a mental health state.
[0074] In some embodiments, the electronic report includes annotations of language model output that indicate the prominence of at least one of one or more time segments.
[0075] In some embodiments, this further includes determining at least one role of at least one speaker, and the weights are based in part on at least one role of at least one speaker during each time segment.
[0076] In some embodiments, the language model output includes the identification of at least one query based on conversational data from the person in charge, and at least one response to at least one query based on conversational data from the patient, wherein at least one query is mapped to a predetermined query, and at least one response to at least one query is mapped to a predetermined response to the predetermined query.
[0077] In some embodiments, processing conversational data to generate language model output involves using a language model neural network trained on labeled conversational data collected from one or more other subjects, where the labeled conversational data is labeled as (i) having a state at some level and (ii) not having that state.
[0078] In some embodiments, at least one query belongs to a set of queries, and the computing device is configured to predict an overall score based on the set of queries based on the response to each of the queries in the set.
[0079] In some embodiments, conversation data is processed based in part on at least one of a patient profile and a practitioner profile, the profile including at least one of historical data, career data, demographic data, and longitudinal data.
[0080] In some embodiments, the conversation data includes at least one of utterance data and text-based data.
[0081] In some embodiments, the method is configured to run a model based on human-interpretable features.
[0082] In one embodiment, a method is provided for scoring a survey based on conversations. The method includes receiving conversation data from at least one input device; processing the conversation data to generate a language model output, the language model output including the identification of at least one query based on conversation data from a person in charge and at least one response to at least one query based on conversation data from a patient, wherein at least one query is mapped to a predetermined query and at least one response to at least one query is mapped to a predetermined response to the predetermined query; generating an electronic report; and transmitting the electronic report to an output device.
[0083] In some embodiments, the method further includes processing conversational data to generate an acoustic model output, and fusing the language model output and the acoustic model output by applying weights to the language model output and the acoustic model output and generating a composite output from the weighted output, wherein the language model output and the acoustic model output each include multiple outputs corresponding to multiple time segments of the conversational data, and the weights are optionally time-based.
[0084] In some embodiments, the weights are partially based on at least one query or topic within each time segment.
[0085] In some embodiments, the method further includes outputting at least one unresolved query to an output device. An unresolved query may be one for which the user has not provided a response.
[0086] In some embodiments, the method further includes outputting a flag indicating that at least one response to at least one query does not map to a given response to a given query with a confidence level above a threshold.
[0087] In some embodiments, the method further includes prompting the user to repeat a given query if it is determined that at least one response to at least one query does not map to a given response to a given query with a confidence level exceeding a threshold.
[0088] In some embodiments, the electronic report identifies the severity of at least one symptom of the condition based on the language model output.
[0089] In some embodiments, the state includes a mental health state.
[0090] In some embodiments, the electronic report includes annotations of language model output that indicate the prominence of at least one of one or more time segments.
[0091] In some embodiments, the method further includes determining at least one role of at least one speaker, and the weights are based in part on at least one role of at least one speaker in each time segment.
[0092] In some embodiments, processing conversational data to generate language model output involves using a language model neural network trained on labeled conversational data collected from one or more other subjects, where the labeled conversational data is labeled as (i) having a state at some level and (ii) not having that state.
[0093] In some embodiments, conversation data is processed based in part on at least one of a patient profile and a practitioner profile, the profile including at least one of historical data, career data, demographic data, and longitudinal data.
[0094] In some embodiments, the conversation data includes at least one of utterance data and text-based data.
[0095] In some embodiments, the method is configured to run a model based on human-interpretable features.
[0096] In one embodiment, a non-temporary computer-readable medium is provided that includes program instructions for causing a computer to perform one of the above methods.
[0097] In one embodiment, a system is provided for training one or more baselines of language models, acoustic models, and fusion models (models) to directly or indirectly detect behavioral or mental health states using machine learning. Training includes using the model to predict behavioral or mental health states in training data and updating the model based on the accuracy of the predictions.
[0098] In some embodiments, the labels in the training data are evaluated for reliability before or during training, and each training data is reweighted according to the reliability of its label.
[0099] In some embodiments, the training data is augmented with at least one of paraphrases or synthesized utterances to generate additional training data for use with the training data.
[0100] In some embodiments, training includes predicting speaker IDs and penalizing the model for accurately identifying speaker IDs. In some embodiments, the model is penalized for learning information about speaker IDs rather than information about the speaker's state.
[0101] In one embodiment, a system is provided for predicting the severity of at least one symptom of a subject's behavioral or mental health condition. The system comprises at least one input device for receiving conversational data from the subject, and at least one computing device for communicating with the at least one input device. The at least one computing device is configured to receive in-context learning, including explanations related to one or more questions on a questionnaire, receive conversational data from the at least one input device, and predict the severity of at least one symptom of the subject's behavioral or mental health condition based on the in-context learning and the conversational data.
[0102] In some embodiments, a computing device accesses a large-scale language model to predict the severity of at least one symptom of a subject's behavioral or mental health condition based on in-context learning and conversational data.
[0103] In some embodiments, predicting the severity of at least one symptom of a behavioral or mental health condition includes predicting the results of a questionnaire.
[0104] In some embodiments, the prediction of the severity of at least one symptom of a behavioral or mental health condition includes the prediction of the outcome of at least one of one or more questions on a questionnaire.
[0105] In one embodiment, a system is provided for preprocessing transcript data for use in predicting the severity of at least one symptom of a subject's behavioral or mental health condition. The system comprises at least one input device for receiving conversational data from a subject, and at least one computing device for communicating with the at least one input device. The at least one computing device is configured to receive in-context learning including a description of a behavioral or mental health condition, receive conversational data from the at least one input device, preprocess the conversational data by performing at least one of the following based on the relationship between at least one segment of the conversational data and a behavioral or mental health condition: weighting at least one segment of the conversational data, summarizing at least one segment of the conversational data, providing an analysis on at least one segment of the conversational data, summarizing at least one aspect of the behavioral or mental health condition, and providing an analysis on at least one aspect of the behavioral or mental health condition, and to send the preprocessed conversational data to one or more models to predict the subject's behavioral or mental health condition.
[0106] In some embodiments, computing devices access large-scale language models to preprocess conversational data.
[0107] Many further features relating to the embodiments described herein, and combinations thereof, will become apparent to those skilled in the art by reading this disclosure. [Brief explanation of the drawing]
[0108] [Figure 1] This figure shows an exemplary system for predicting a patient's condition based on passive listening, according to several embodiments. [Figure 2] Several embodiments of methods for preprocessing speech data are described. [Figure 3] Several embodiments of methods for detecting the role of a speaker are described. [Figure 4] Several embodiments of methods for detecting topics within a conversation are presented. [Figure 5] Several embodiments describe a method for scoring a survey based on passive listening of conversations. [Figure 6] This document presents exemplary UIs for a case manager in several embodiments. [Figure 7] This document presents exemplary UIs for a case manager during a call with a patient, based on several embodiments. [Figure 8] This document presents an exemplary UI for a case manager while the system delivers behavioral health scores, based on several embodiments. [Figure 9] The following are exemplary call summaries of care meetings in several embodiments. [Figure 10] Schematic diagrams of computing devices in several embodiments are shown. [Modes for carrying out the invention]
[0109] This specification provides systems and methods that can passively listen to conversations involving a patient and analyze the conversation using acoustic and linguistic models to estimate the patient's potential mental health status. Conversations can take place between a patient and a representative, such as a healthcare team member or case manager. Conversations can also take place between a patient and a bot (e.g., a chatbot or AI representative). The systems and methods described herein may work with different forms of conversational data, such as oral conversations (e.g., voice conversations) or written conversations (e.g., conducted via text messaging). The systems and methods described herein may work with utterances or text information provided by a single user (e.g., analysis of a patient's spontaneous utterances or journal entries, which may be referred to as "conversational data"). The systems and methods described herein may also work with asynchronous conversational data (e.g., conversations that take place over time with intervals between each turn, e.g., text conversations). The system and method may include role detection, diarization, topic detection to support evaluation and computational analysis, and survey scoring to support evaluation and speaker classification by the person in charge. The system and method described herein may analyze conversations using natural language processing (NLP) and / or acoustic models. The system and method described herein may be useful, for example, for non-experts in mental health to perform assessments and provide support for clinical judgment.
[0110] Embodiments of the systems and methods described herein generate severity scores for mental health symptoms from existing person-to-person conversations. The systems and methods described herein can address over- or under-detection of mental health conditions and the inadequacy of relying on self-report measures such as the PHQ-9 for this purpose. Populations with chronic illnesses may be at high risk of depression and anxiety, and these individuals may participate in regular calls with caregivers, who may not be trained in mental health. Implementing self-report screening tools may not be appropriate in such situations for a number of reasons, including a lack of consistency in use and implementation, lack of engagement, extra time, and repetition. Evaluating these informal conversations may be beneficial.
[0111] In some embodiments, the systems and methods described herein may be even more advantageous because they may be able to better identify and account for overreporting or underreporting. Some patient populations may overreport or underreport in self-report mental health surveys. For example, individuals who experienced World War II may be underreporters, and individuals under 25 years of age may be overreporters. The systems and methods described herein may be able to provide more accurate scoring for these individuals because the systems can score patients by benefiting from training data from the entire population of individuals. Furthermore, the systems and methods described herein may be used, for example, to predict the mental health status of a patient population (or subpopulation) based on patient demographics.
[0112] Furthermore, the systems and methods described herein may be configured to use data (e.g., demographic data, metadata, etc.) to assess the likelihood of patients overreporting or underreporting and to adjust the model's weighting accordingly. For example, the model may be configured to give higher weight to signs of depression in underreporters to account for a tendency to underreport.
[0113] For example, it may be beneficial to assess a patient's condition not only based on their responses to survey questions, but also on their free-form conversation and the acoustic characteristics of that conversation. Such analyses may include evaluating relevant topics in the conversation based on language recognition, and evaluating the acoustic properties of the conversation. Furthermore, by reserving sophisticated analysis of the conversation for segments of the conversation relevant to relevant topics, computational resources may be used more economically and efficiently. Such features may be particularly useful when delivering real-time feedback or when using third-party paid models (e.g., when using more sophisticated and expensive models only for the most relevant parts of the conversation).
[0114] In some embodiments, the systems and methods described herein can be used to passively listen to conversations between patients and healthcare professionals. Unlike agents, healthcare professionals may have more impromptu utterances. Questions asked may suggest a particular condition but may not map entirely to conventional questionnaires. Speaker segmentation and text summarization help healthcare professionals meet record-keeping requirements without requiring additional work from them.
[0115] Furthermore, the actual implementation of PROs (Patient-Reported Outcomes) may be carried out by personnel with varying levels of competence, who may not ask questions verbatim, may not ask all questions, or may input incorrect patient responses using one or a different set of options for which the PRO was validated.
[0116] The systems and methods described herein utilize regular conversations between the practitioner and the patient to provide a professional "ear" trained to listen for evidence of mental health symptoms in speech and language. The present invention may require no additional effort or time from either party.
[0117] Some of the technical advantages of the systems and methods described herein include improved prediction of mental health symptoms compared to conventional systems and methods, reduced overall prediction time, improved trust between patients and caregivers, more patient-centered content, more accurate scoring of oral surveys, improved compliance with a set of questions for patient assessment, and improved patient triage and referral (leading to improved outcomes).
[0118] In some embodiments, the systems and methods described herein may analyze a conversation in real time and, while the conversation is in progress, provide the agent with clues / hints on how to improve it (e.g., what to ask next, what to discuss in more detail, what topics to cover, etc.). Such implementations may further guide the agent to provide advice to the patient (e.g., how to cope with difficult times).
[0119] In some embodiments, the systems and methods described herein can be implemented without interrupting or extending the conversation between the patient and the caregiver. Furthermore, the caregiver may require little to no additional training. The system's output may be usable "out of the box," for example, to predict the severity of mental health symptoms. In some embodiments, the systems and methods may be able to determine which speaker is the caregiver and which is the patient, based on the absence of training by the caregiver or patient, and analyze the conversation accordingly. In some embodiments, the system may perform just-in-time analysis to generate a score for the severity of mental health symptoms at or before the end of a naturally occurring case management / care management conversation. In some embodiments, the output of the system and methods may be a symptom score, a survey score from an oral survey, or a summary via automated summarization. Automated summarization may be useful because it reduces the amount of downtime required for each caregiver to meet documentation requirements. The mental health score can be associated with the organization's existing clinical pathways to provide appropriate referrals or recommendations commensurate with the mental health score.
[0120] Surveys may be more engaging when distributed by a person rather than through a written survey or automated distribution system. To make the conversation flow more freely, it may be beneficial to allow the person in charge to guide the survey questions so that the conversation becomes relevant within that free flow. It may also be beneficial to allow the person in charge to rephrase the questions based on their natural presentation and / or the patient's situation.
[0121] Furthermore, it would be beneficial if the system were configured to elicit relevant answers to survey questions regardless of differences in the order or content of question distribution. Such a system would free up staff time, allowing them to see more patients on a working day. Moreover, it could enable staff to provide accurate and precise inquiries and treatments with low latency (e.g., real-time or just-in-time).
[0122] Multiple conditions may coexist. The systems and methods described herein may determine the presence of another condition by taking a previous diagnosis or prediction by the system of a condition as input. Different subtypes may exist, and different treatments may exist for different subtypes (e.g., patients who respond well to therapy, SSRIs, or TMS, or anorexia in depression). When treating the whole person, the more information available, the better the system will be at recommending appropriate treatment for different subtypes.
[0123] In some embodiments, the system may be configured to score patients in relation to the severity of their condition. For example, this score may be associated with a confidence metric. This score may be used to determine triage among patients, for example, if medical resources are low throughout the system, the system may ensure that patients with more severe cases are seen before those with milder or more stable cases, while ensuring that all individuals exhibiting depression are referred to therapy. Furthermore, the system may be configured to recommend more aggressive and urgent treatment for patients whom the system predicts with high confidence that they are suffering from a serious and urgent condition.
[0124] In some embodiments, a patient's response to a drug can provide information about the progression of their condition (for example, the system can monitor many patients taking the drug to identify the expected progression of a patient's condition and apply that progression to make further predictions for the patient).
[0125] In some embodiments, the system may be configured to store longitudinal information about a patient in a patient profile, for example. In some embodiments, the system may be able to measure a patient's risk under different drugs or treatments, dosages, and timelines. The system may be configured to predict whether a drug is working for a particular patient over time (rather than periodically rotating patients). By measuring inter-session changes in conversation data, the system may be able to confirm the effectiveness of a drug earlier and more accurately than through surveys. The system may also use inter-session changes to support drug administration and other adaptations. The system may also be configured to extract adverse event data from sessions.
[0126] Patient-completed questionnaires are formalized, and patient responses may be based on habitual responses rather than genuine self-reflection. Therefore, questionnaires may require more significant results than those measured herein to demonstrate their validity through questionnaire measurements. There can be numerous reasons why self-reported responses are unreliable. In some experiments, only a small number of patients can reliably reproduce their responses. Questionnaires may be completed based on the benefits perceived by the patient, or they may be completed to elicit specific responses from those interpreting the results (e.g., to please the therapist by suggesting the patient is likely getting better).
[0127] Although the above usage is primarily described in the context of drugs, it can be similarly applied to any treatment scheme or other compliance regime. For example, in some embodiments, the treatment may be an exercise routine, a nutrition plan, cognitive behavioral therapy, or mindfulness exercise. In some embodiments, the treatment plan may be a combination of two or more elements (e.g., drugs and a nutrition plan). Any type of treatment or therapy may be compatible with this system.
[0128] General system Figure 1 shows an exemplary system 100 for predicting a patient's condition based on passive listening, according to several embodiments. System 100 includes an input device 102, a computing device 104, and an output device 126.
[0129] Input 102 The input device 102 may be configured to receive and relay data, for example, including audio data of a conversation between a patient and a healthcare provider (e.g., conversation data). The input device 102 may capture audio data from the conversation. For example, the input device may be a voice recorder present in the room, a passive listening device configured to listen to a telephone (or other remote) call, or another capture device methodology. In some embodiments, the input device 102 may include one or more of the following: a microphone or other audio capture device that passively listens to a one-on-one face-to-face session (e.g., between a doctor and a patient) or a face-to-face session in a room where other people may be present (e.g., an ambient microphone in a hospital environment or home healthcare visit), a landline telephone, a mobile phone, a meeting call service, a video meeting call service, a call center, etc. In some embodiments, the input device 102 may be any device configured to listen to a patient having a conversation with a healthcare provider.
[0130] In some embodiments, the input device 102 may capture a two-way conversation using a single channel. In some embodiments, the input device 102 may capture a conversation using multiple channels. In some embodiments, the input device 102 may be configured to include a keyboard, mouse, camera, touchscreen, and / or microphone for receiving input from the person in charge or patient. It may be beneficial for the system to be speaker-independent (i.e., it may not have a baseline of patient utterances).
[0131] In some embodiments, the input device 102 may include a channel from a care manager at a call center and a channel from a patient with whom the care manager is speaking. In some embodiments, the care manager and the patient may be on the same channel. In some embodiments, the care manager may be located at the call center, and its input device may include further functionality for reviewing patient data and / or past conversations (including call notes).
[0132] In some embodiments, the systems and methods described herein may be implemented (or supported) using a large-scale language model-based chatbot (e.g., ChatGPT) or other generated speech / text scheme instead of a live person. In such embodiments, the system may be configured to prompt a specific response from the patient, but may be trained to prompt responses that elicit longer responses from the user.
[0133] In some embodiments, these large-scale language model-based chatbots can be used to collect training data. For example, the models may be implemented to simulate actual personnel and / or actual patients to augment the training data for one or more aspects of the system.
[0134] In some embodiments, large-scale language model-based chatbots may be supported by virtual avatars. In addition to prompting patient responses, the virtual avatars may express emotions or provide backchannel communication. This can be advantageous for automating the role of the agent while still providing the patient with a virtual human connection. Avatars may be further configured for digital, augmented, and / or virtual reality systems.
[0135] In some embodiments, the input device 102 may further receive additional input from the conversation. For example, in some embodiments, the input device 102 may be configured to capture visual data from one or more of the conversation participants (e.g., patients). In such embodiments, the visual data may be used to detect visual indicators of mental state in order to further enhance the reliability of the system.
[0136] In some embodiments, the input device 102 may include video input. For example, a session may be conducted via a video meeting call service, and video data may be supplied to the system instead of, or in addition to, audio data. In such a situation, the model may be trained to evaluate the video data for cues that may indicate the patient's mental state or areas of the session that may have higher predictive values (e.g., topic detection, which will be described in more detail below). For example, the system may evaluate the patient's eye position, posture, gaze fixation, or head movement (headmode) and make an assessment of the patient based on this information. In some embodiments, it may evaluate the importance of areas of the session based on the video data (for example, the system may weight a portion of the session more highly if the patient exhibits spontaneously withdrawn body language or increased self-sedating behavior, as this may indicate a portion of the session in which the patient is experiencing a significant emotional response). In some embodiments, the video data may be used to assist the system in, for example, conversational data processing. For example, the system may track whose mouth is moving during turns of conversation in order to better segment the session as speaker. For example, if a patient is on a video call with a caregiver, the system can better distinguish between the patient's speech and the caregiver's speech by detecting whose mouth is moving. As a further example, if the patient is lying down, the conversation data may change slightly, but using video, the system may be able to take this change into account.
[0137] In some embodiments, system 100 may be configured to track conversations conducted via SMS services or other text-based interactions (e.g., using such a device as input device 102). In such embodiments, the model may be specifically trained to evaluate the patient's texting habits. The communication style exhibited by the patient via text may differ substantially from the communication style exhibited via voice chat. Therefore, the model may need to be specifically trained to evaluate the state in a text environment. Furthermore, conversations via text messaging may be less synchronized (e.g., the patient is doing something else) or not synchronized at all (e.g., the agent and patient respond at random intervals throughout the day), which can dramatically affect the manner in which the data is analyzed. For example, since the patient's situation may differ during the time between interactions due to many factors (e.g., related to their mental state, such as BPD, or not related to their mental state, such as, for example, getting burned while cooking), it may be more appropriate to treat each asynchronous interaction in a more independent manner.
[0138] Conversations conducted via text-based communication may be more important to individuals within certain demographic groups. For example, older individuals may be less inclined to text and tend to treat it like formal correspondence, while younger individuals may be more comfortable sharing cues about their mental state through text rather than face-to-face interaction. Furthermore, different subsets of a group may have different etiquette rules surrounding text-based communication (e.g., the use of punctuation marks may be interpreted as a sign of hostility among younger individuals). Adding further complexity to text-based communication, the use of emojis within text can have significantly different meanings depending on how they are used in the message and the demographics of the person using them (e.g., a "blowing a kiss" emoji may be used affectionately by an older individual but ironically by a younger individual or depending on the context).
[0139] The use of text-based communication may be used in place of speech data. The use of text-based communication may be used in addition to speech data. For example, a patient may send text messages to their caregiver intermittently between formal voice sessions, either at the caregiver's request or at the patient's discretion. Topics discussed during text conversations may be used to inform the analysis of speech data. For example, a patient or caregiver may refer to topics discussed via text, and the system may be configured to analyze the text message conversation to retrieve the topics discussed.
[0140] In some embodiments, the system may be configured to assess the mental state of one or more patients in a group text.
[0141] In some embodiments, the input device 102 may extract relevant information from other resources. For example, the input device 102 may be configured to retrieve history data or other patient data from an external source (e.g., contained in an electronic health record (EHR)) and provide it to the computing device 104 to enhance the predictive capabilities of the system 100. In addition, it may be beneficial to incorporate patient personality or profile data (e.g., whether the patient tends to share things, whether the patient tends to complain, past personality survey results, etc.).
[0142] Processor 106 The computing device 104 may be configured, for example, to process audio data of a conversation between a patient and a caregiver to predict the patient's mental state. The computing device 104 may optionally be configured to detect the role of each speaker in the conversation (e.g., patient or caregiver), to detect the topic of the conversation, and to score the investigation. The computing device 104 may be a server, a network appliance, a computer expansion module, a personal computer, a laptop, a personal data assistant, a mobile phone, a smartphone device, a UMPC tablet, a video display terminal, an e-reader device, a wireless hypermedia device, or any other computing device that can be configured to perform the methods described herein. The computing device 104 includes at least one processor 106, at least one memory 108, and at least one network interface 110.
[0143] The processor 106 may be a microprocessor or microcontroller, a digital signal processing (DSP) processor, an integrated circuit, a field-programmable gate array (FPGA), a reconfigurable processor, a programmable read-only memory (PROM), or any combination thereof. The processor 106 may include a speech preprocessor 112, a role detector 114, a language processor 116, an acoustic processor 118, a fusion model 120, and a topic detector 122, a survey scorer 124.
[0144] Memory 108 Memory 108 may include instructions for performing any one of the following functions: a voice preprocessor 112, a role detector 114, a language processor 116, an acoustic processor 118, a fusion model 120, a topic detector 122, and a survey scorer 124. Memory 108 may include a suitable combination of various types of computer memory located either internally or externally, such as random access memory (RAM), read-only memory (ROM), compact disk read-only memory (CDROM), electro-optical memory, magneto-optical memory, erasable programmable read-only memory (EPROM), and electro-erasable programmable read-only memory (EEPROM), and ferroelectric RAM (FRAM).
[0145] Network interface 110 The network interface 110 enables the computing device 104 to communicate with other components (e.g., input device 102 and output device 126), exchange data with other components, access and connect to network resources, provide applications, and run other computing applications by connecting to networks (or multiple networks) capable of carrying data, including the Internet, Ethernet, Plain Old Telephone Service (POTS) lines, Public Switched Telephone Networks (PSTN), Integrated Services Digital Network (ISDN), Digital Subscriber Line (DSL), coaxial cable, optical fiber, satellite, mobile, wireless (e.g., Wi-Fi, WiMAX), SS7 signaling communications networks, fixed lines, local area networks, wide area networks, and others, including any combination thereof.
[0146] Output 126 The output device 126 may be configured to provide output. For example, in some embodiments, the output device 126 may include a display. In some embodiments, the display may be configured to provide real-time or just-in-time results to the person in charge. In some embodiments, the person in charge may be able to access previously processed results using the output device 126. In some embodiments, the output device 126 may be configured to provide some or all of the results of the computing device 104's determination to a storage database and / or further computing device, for example, further processing or other output reasons. In some embodiments, the output device 126 includes a display screen and a speaker.
[0147] The delivery of results may include documenting the conversation and providing the results to the person in charge to determine the next steps for the patient regarding mental health screening or management. For example, the delivery of results may include summarizing the conversation, analyzing the conversation, providing predictions and confidence levels regarding the predictions and data, and comparing the results with previous data of the conversation.
[0148] In some embodiments, results can be delivered in real time (i.e., real-time mode) and dynamically updated. In some embodiments, results are delivered at the end of the conversation, but while the agent is still on the phone with the patient (i.e., just-in-time mode). In some embodiments, conversations can be processed in batches after the conversation has ended (i.e., offline mode). In offline mode, results are delivered centrally or directly to the agent.
[0149] In real-time dialogue mode, the system's configuration can allow the agent to dynamically modify the course of the conversation with the patient to achieve better results. For example, if sufficient information about the patient's mental health has not yet been addressed, the system can prompt the agent to insert mental health-related questions into the existing conversation. The system can also dynamically inform the agent whether the system's confidence in its predictions is sufficient, allowing the agent to save time and conclude the conversation. The system can also inform the agent when to ask for more details or when to encourage a longer turn.
[0150] In just-in-time mode, preprocessing is performed strictly sequentially (i.e., preprocessing is performed, roles are identified, and then information is de-identified). In just-in-time mode, it may be beneficial to end passive listening a little before the end of the conversation to give the system time to analyze the conversation. In some embodiments, the system may be configured to predict when a conversation will end (e.g., based on changes in pitch) so that the just-in-time analysis ends at the same time the conversation ends. Such embodiments may provide automated just-in-time analysis in that the analysis is performed without further input from the participants.
[0151] Offline mode can offer more economical analysis (i.e., conversations can be run offline in batches when the demand for processing power is lower, and dynamically adjusted without affecting the final delivery time).
[0152] As part of the output provided to the user, the system may provide analysis. Analysis can provide end users with information about the content, quality, and dialogue of conversations. Analysis can include an unlimited range of measurements derived from the output of processing steps, particularly using word-based information. For example, the system may provide the total time spoken by each speaker, the number of words spoken by each speaker, the percentage of utterances from the patient, the length of turns, the distribution of topics, etc. Analysis can capture interpretable and useful information about conversations. For example, the number of response words can be used as a surrogate indicator of engagement level (e.g., a conversation with a low percentage of patient words may indicate a low level of trust). Analysis can be used by case management companies to see overall trends, understand the performance of different agents, and know which questions were asked in which conversations.
[0153] The analysis can be based on text and audio (i.e., acoustic information). The specific form of patient responses may indicate that the person in charge is skilled at eliciting sufficient responses from the patient. The objective of the systems and methods described herein (particularly machine learning predictions) is to extract mental health-related cues. In some embodiments (e.g., for case management), a more general interest may be directed to knowing whether other information (not just mental health-related cues) is obtained, as well as whether there is patient engagement and trust.
[0154] In some embodiments, the analysis may be provided in real time during the conversation with the patient. The system may be configured to flag if the conversation is not eliciting enough information from the patient, or if there are unresolved questions that need to be answered so that the agent can correct the course during the conversation. The model requires patient utterances to perform, and the model may perform better if the agent is eliciting many personal utterances. The system may perform best if the patient is giving longer responses that are inherently deeper and more introspective.
[0155] In some embodiments, the output may further provide a text summary. The text summary may provide a summary of the discussions that took place during the conversation. It may provide a final report by the agent or help to recall the agent's memory later. Creating this summary without input from the agent may save the agent time (it gives the agent an opportunity to verify accuracy, but eliminates the need for the agent to take detailed notes or create a summary). The text summary may represent a function that can be used independently of the prediction of mental health. The text summary may use a Large Language Model (LLM) to summarize ASR-based output from a case management call.
[0156] In some embodiments, conversation topics may be extracted from the conversation to be provided as part of the analysis. By providing the topics discussed during the conversation, the metrics can become more interpretable (end users can see both what the patient's assessment was and which topics were discussed and flagged by the system).
[0157] Some embodiments may further provide survey scoring as part of the output (as described above). The system may, for example, track the oral administration of the survey and maintain the score via a graphical user interface. For example, in conversational mode, the person conducting the survey can use the information to better conduct and score such surveys. It should be noted that these improvements may be due to the person conducting the survey using the information (allowing the conversation to proceed naturally for the patient), as the purpose in conversational mode may be to avoid distracting the patient.
[0158] UI examples The systems and methods described herein can provide call intelligence tools. These call intelligence tools can use AI to quantify behavioral health issues (e.g., assess the severity of depression and anxiety), effectively triage and monitor, minimize administrative tasks with memo summaries, analyze various aspects of a call (e.g., patient engagement, topics discussed, treatment alliances), and provide real-time guidance to the person in charge.
[0159] Figure 6 shows an exemplary UI600 for a case manager according to several embodiments.
[0160] In some embodiments, a care manager can open the software and launch an exemplary UI. The launch screen can provide the care manager with information such as the patient's name and information, a care summary (e.g., including any barriers to medical care, social determinants of health, and care gaps). UI600 may also include information such as any recent assignments (e.g., discharge planning or assessment scheduling). UI600 may also include the patient's health risk score.
[0161] Figure 7 shows an exemplary UI700 for a case manager while communicating with a patient, according to several embodiments.
[0162] From this UI portal, care managers can initiate a workflow by calling a patient (see icon 702). While on the call with the patient, or at the beginning of the conversation, the system may display a screen for loading behavioral health scores (704).
[0163] Figure 8 shows an exemplary UI800 for a case manager while the system delivers behavioral health scores, according to several embodiments.
[0164] When a care manager speaks to a patient, the system can analyze what and how the patient is saying and deliver a behavioral health score. For example, it can display behavioral health scores of 802 for depression and 804 for anxiety. As a further example, this information may be provided along with historical data to show whether the patient's score has increased or decreased (see increases in depression 806 and anxiety 808). Other health scores could include stress scores, etc.
[0165] Using existing care pathways, the system can triage care based on severity and recommend the next best action. For example, a high score might prompt a care manager to send a behavioral health referral (see Recommendation 812 under Intelligent Guidance). Care managers can perform all of these actions during a call to ensure timely care. The behavioral health score may also take into account the health risk score 810 to provide information about the patient's risk of readmission.
[0166] Care managers can also access call summaries from the calls for future reference.
[0167] Several systems and methods can streamline and expedite the screening, monitoring, and triage of behavioral health problems, saving time while ensuring a better patient experience and quality of care.
[0168] Figure 9 shows exemplary call summaries 900 of care meetings in several embodiments.
[0169] In some embodiments, the system may be further configured to generate note summaries based on care meetings between the patient and the care manager. The call summaries may include logistical information about the meeting (e.g., the date and time the meeting was held, the duration, etc.). The call summaries may include summaries of the topics discussed during the meeting. These summary notes may be extracted from the output of a language model based on a machine learning algorithm. The call summaries may also include relevant details related to the interventions discussed during the meeting, along with the interventions themselves. Interventions may be extracted, for example, using a machine learning algorithm applied to a language model or an acoustic model.
[0170] General system operation During runtime, the systems and methods described herein may perform data preprocessing, model inference (which may include analysis, text summarization, and / or survey scoring, prediction of the severity of symptoms of mental health conditions), and result delivery.
[0171] The systems and methods described herein may use either or both acoustic and / or linguistic models, and optionally a fusion thereof. In some embodiments, all available models for health conditions may be used (e.g., depression, anxiety, etc.). In some embodiments, historical information about the patient may be used to analyze longitudinal trends. In some embodiments, metadata about the patient may be used to condition model predictions. In some embodiments, the analysis may generate information about the topics discussed, the length of utterance regions, the overall utterance share by participants, and many other analyses. Summarization may be used to create an automated summary of the conversation. The systems and methods described herein may be configured to handle multiple languages (and potentially include code-switching). Topic regions may be determined in the conversation, and these regions may be dynamically weighted for more efficient processing and higher accuracy of the model. Such weighting may be used to explain where cues are located. Data from agents (in addition to the patient) may be used for linguistic processing. Linguistic data may be used to find utterance regions on which to run the acoustic model. The roles of speakers (e.g., patient or agent) may be detected automatically. Confidence scoring can be used, for example, to measure signal quality (e.g., SNR, clipping), confidence from third-party ASR, length, and model-based estimates.
[0172] Pre-treatment In some embodiments, input data may be preprocessed. For example, audio data may be processed by the audio preprocessor 112. Other preprocessing modalities are also conceivable (e.g., text preprocessing, visual preprocessing).
[0173] The audio preprocessor 112 may optionally be configured to process any audio data received from the input device 102. In some embodiments, the audio signal may be preprocessed to filter out irrelevant noise (e.g., through the use of a bandpass filter) or to render the data into a format usable by the computing device 104.
[0174] De-identification can be performed before speech or language processing. The recording can be subjected to automatic speech recognition (ASR) to mark protected health information (PHI) / personally identifiable information (PII), and then in speech processing, (white noise or some other modification to mask or remove speech content) the PHI / PII regions are placed, and in language processing, the PHI / PII regions are removed.
[0175] Figure 2 shows a method 200 for preprocessing conversation data according to several embodiments.
[0176] In some embodiments, metadata and audio data (e.g., conversation data) can be processed in parallel.
[0177] For example, some systems can determine demographic information about patients related to their mental health status. For instance, a postal code can be mapped to an SVI class (Social Vulnerability Index class). This information can then be used to map individual responses within a PHQ-8 label, for example. In the case of audio data, the audio quality can be determined. All of this information can be included as final metadata for the meeting.
[0178] ASR may also be applied to the audio data to extract words from the speech. In some embodiments, the system may also provide speaker segmentation such as meeting, timing, language ID (LID), and PII detection. The data may be anonymized to detect and anonymize any PHI. This processed information may then be split into transcript information and audio data. The transcript information may be evaluated for transcription, for example, role detection (described in more detail below) and manual survey removal (which may be done so that the model can be trained on non-survey conversations rather than simply relying on patient responses to survey questions). The audio data may be anonymized by providing masked speech. In some embodiments, ASR may be combined with other services such as speaker segmentation, PHI and PII detection, or LID detection.
[0179] Speech masking may include applying white noise or other forms of speech masking. In some embodiments, PII and / or PHI may be masked with silence or other speech distortions. It may be important to train the system to handle such silence or speech distortions. In some embodiments, PII and / or PHI detection may require specific trusted PII / PHI detection software during use (e.g., to comply with regulations). Such software may be computationally expensive and / or only accessible on a paid basis. In some embodiments, it may be advantageous to train the system with PII / PHI detection software that can mimic specific trusted software, but it may be computationally simpler and / or less expensive. Such software may be used during training, ensuring that PII and / or PHI are not viewed by other entities (e.g., data is used only within a company with restricted access), but not during actual use (session information may be delivered to external entities). Such embodiments may be trained in a configuration that more or less matches the use of PII / PHI detection and masking, but are trained with less expensive software (financially and / or computationally). Such embodiments may have the advantage of training a practical model without incurring the potentially higher overhead required for a reliable model.
[0180] Next, the transcript data of the conversation (with roles) and the processed speech data can be input into acoustic data preparation (for subsequent model analysis, e.g., acoustic model analysis). The transcript (with roles) can be input into NLP (or any other language model) data preparation (for subsequent model analysis, e.g., language model analysis). The final metadata can be passed to the system for analysis and metrics.
[0181] At runtime, the trained model (e.g., housed in the language processor 116 and the acoustic processor 118) and an optional model fusion 120 can process the data and assign a severity score for each mental health condition symptom. The model output may optionally include a confidence scale (in particular, data sample audio recording quality, data length, topic content, ASR quality, and model quality) to allow the end user to determine how much confidence they have in the model's predictions for a particular data sample. The confidence scale can also allow the end user to determine how much confidence they have in the predictions and analysis. The ASR confidence scale is also important for the reliability of the summarization and analysis output (primarily word-based).
[0182] In some embodiments, inputs required by one component of the system may also be required by another component of the system. The system may improve its processing efficiency by processing the information in a way that reuses these inputs rather than regenerating them. For example, a model using ASR output (words or word time marks) and other types of speech processing may be configured such that the speech recognition process first passes the data to determine, for example, word time marks. Other speech processing components may also use those time marks for their processes rather than passing the information another way to regenerate them. For example, a speech recognition system may record the time of each word used by the patient, and a speech processor may use the time of the first word used by the patient to perform speech processing on the string of the patient's utterances (which may include a buffer beforehand).
[0183] In one embodiment, a system 100 is provided for preprocessing transcript data for use in predicting the severity of at least one symptom of a subject's behavioral or mental health condition. The system 100 comprises at least one input device 102 for receiving conversational data from a subject, and at least one computing device 104 for communicating with at least one input device. The at least one computing device 104 is configured to receive in-context learning including a description of a behavioral or mental health condition, receive conversational data from at least one input device 102, preprocess the conversational data by performing at least one of the following based on the relationship between at least one segment of the conversational data and a behavioral or mental health condition: weighting at least one segment of the conversational data, summarizing at least one segment of the conversational data, providing an analysis on at least one segment of the conversational data, summarizing at least one aspect of the behavioral or mental health condition, and providing an analysis on at least one aspect of the behavioral or mental health condition, and to send the preprocessed conversational data to one or more models to predict the subject's behavioral or mental health condition.
[0184] In some embodiments, the computing device 104 accesses a large-scale language model to preprocess conversational data.
[0185] Role detection The role detector 114 may be configured to process the data at will to determine the role of each speaker. In some embodiments, it may not be clear what role each speaker is participating in (e.g., agent or patient). Not all setups for passive listening provide information about which speaker is the patient and which is the agent (e.g., setups that use only one channel for both patient and agent). If necessary, role detection can be performed to detect which speaker is the patient. The system may then attribute the speech data to the patient (for acoustic analysis) and the NLP data to both patient and agent. In some embodiments, the agent's acoustics may be retained and / or used. Furthermore, some systems may not have a reliable means of distinguishing between the two roles.
[0186] The relevant roles may generally be those of the patient or member, so that the systems and methods described herein can be used to assess the mental health of the patient or member.
[0187] In some embodiments where there are two or more speakers on the same channel, it may be necessary to identify which speaker is the patient and which speaker is the agent. Doing so helps map NLP data to the correct speaker (a process called speaker segmentation). Once roles are identified, the system can then use turn-taking to determine which speaker is currently speaking.
[0188] Role detection may occur within the first few seconds or minutes of a conversation, allowing the system to provide a real-time assessment of the patient's mental state based on the conversation. In some embodiments, roles may be detected by identifying the voice of a representative or patient compared to previous recordings. For example, a representative may provide enough utterance samples to identify their voice stored in the system. In some embodiments, a representative may register with the system as a speaker, allowing the system to recognize the presence of their voice. In these embodiments, the system may then have a known speaker model for that person. Furthermore, a patient may be identifiable if they participate in the system multiple times (and the system takes in the patient's utterance data for identification). In some embodiments, roles may be detected by applying a machine learning system to the conversation that is configured to recognize utterance patterns that generally correspond to a representative (e.g., specific introductory comments such as "How are you?" or intonation), and assigning that speaker as the representative.
[0189] In some embodiments (for example, in the case of a monaural audio channel), role detection and speaker classification may be performed using ASR. For example, ASR may be implemented to determine the words and phrases used by each speaker and to assign roles, for example, agent and patient, using the words and phrases used. In some embodiments, role detection analysis may be performed using algorithms and / or machine learning methods. For example, the words and phrases used may be classified using, for example, term frequency-inverse document frequency analysis with hyperparameters tuned with Bayesian optimization. The role detector 114 may be trained to identify a speaker's role by the use of a particular word or phrase. In some embodiments, the role detector 114 may also consider other factors. In some embodiments, the role detector 114 may be able to retrieve historical data from one or more of the patient and agent to assist in its determination (for example, previous utterance samples from the agent stored in an agent profile, etc.).
[0190] In some embodiments, there may be two (or more) channels so that role detection can be simplified. For example, if the patient speaks through one channel and the caregiver speaks through another, then role detection may be assigned based on the channels of communication, for example. However, even in such embodiments, role detection may nevertheless be advantageous when, for example, multiple people (e.g., 2+) are present on one line. Such a situation may occur, for example, when a patient is being cared for by a caregiver. In such a situation, the patient and caregiver may share a communication channel (e.g., both are heard through the same channel, or both speak alternately, for example). In such embodiments, role detection can be used to distinguish the patient's role from the caregiver's role. Other roles such as translator / interpreter, background speaker (non-entity of the conversation) may also be distinguished.
[0191] An exemplary method of role detection may include performing speaker segmentation to obtain the number of speakers and the number of words assigned to each speaker. If two or more speakers are detected by speaker segmentation, the system may use role detection to assess which speaker is playing which role if there is no other way to determine the role (e.g., if the speakers are not on separate known channels). For example, the system may use a role detection model to score each speaker's words. For example, for each speaker model, it may provide a log probability of being a patient or a person in charge. For example, Speaker 1: (-0.1797,'[s1]','P'),(-1.8046,'[s1]','A'), Speaker 2:(-0.3669,'[s2]','P'),(-1.1804,'[s2]','A'), Speaker 3: (-1.6420,'[s3]','P'),(-0.2151,'[s3]','A')
[0192] All detected roles can be sorted in order of probability. (-0.1797,'[s1]','P'), (-0.2151,'[s3]','A'), (-0.3669,'[s2]','P'), (-1.1804,'[s2]','A'), (-1.6420,'[s3]','P'), (-1.8046,'[s1]','A')
[0193] Next, the system can assign all roles, starting with the one with the highest probability. Patient:s1 Person in charge: s3 Others: s2
[0194] If the system detects only one speaker, this may represent a speaker identification failure. The system may use a fallback procedure to extract usable information from the session regardless of the speaker identification failure.
[0195] In some embodiments, the systems described herein may generally be able to identify calls that should be rejected (e.g., IVR—as described in more detail below, responses from family members and patient non-participation). Some such systems may, for example, be able to identify when an IVR system has been activated. Some systems may be able to identify conversation data that matches a situation where a family member answered the phone and the patient was not participating (e.g., shorter conversations, negative responses to whether the patient could participate, etc.). Such systems may improve the reliability and performance of the model because inappropriate content is not scored.
[0196] In some embodiments, it may be necessary to identify the respective roles of speakers (e.g., personnel or patients) before or during the analysis. Identification may be important to ensure that the correct conversation data is analyzed (i.e., focusing on patient conversation data in the analysis).
[0197] Figure 3 shows a method 300 for detecting the role of a speaker, according to several embodiments.
[0198] In one embodiment, a method is provided for identifying the roles of speakers in a conversation. The method 300 includes receiving conversation data from at least one input device (302), determining at least one role of at least one speaker (304), processing the conversation data to generate a language model output and / or an acoustic model output (306), applying weights to the language model output and / or acoustic model output (308), wherein the language model output and / or acoustic model output each include a plurality of outputs corresponding to a plurality of time segments of the conversation data, and the weights are optionally time-based and partially based on at least one role of at least one speaker in each time segment, generating an electronic report (310), and transmitting the electronic report to an output device (312).
[0199] In some embodiments, Method 300 includes fusing weighted language model outputs and acoustic model outputs to generate a composite model. The composite output may represent a fused output from the fusing of the language model outputs and acoustic model outputs.
[0200] In some embodiments, the electronic report identifies the severity of at least one symptom of the condition based on the combined output.
[0201] In some embodiments, the state includes a mental health state.
[0202] In some embodiments, the electronic report includes annotations of language model output that indicate the prominence of at least one of one or more time segments.
[0203] In some embodiments, at least one speaker includes a representative, and applying weights to the language model output and acoustic model output includes applying a zero weight to the acoustic model output corresponding to the representative. In such embodiments, the language model may use only the patient's language information, or the language information of both the patient and the representative.
[0204] In some embodiments, the weights are partially based on the topics in each time segment.
[0205] In some embodiments, the language model output includes the identification of at least one query based on conversational data from the person in charge, and at least one response to at least one query based on conversational data from the patient, wherein at least one query is mapped to a predetermined query, and at least one response to at least one query is mapped to a predetermined response to the predetermined query.
[0206] In some embodiments, processing conversational data to generate language model outputs and acoustic model outputs involves using language model neural networks and acoustic neural networks trained on labeled conversational data collected from one or more other subjects, wherein the labeled conversational data is labeled as (i) having a state at some level and (ii) not having that state.
[0207] In some embodiments, at least one query belongs to a set of queries, and the computing device is configured to predict an overall score based on the set of queries based on the response to each of the queries in the set.
[0208] In some embodiments, conversation data is processed based in part on at least one of a patient profile and a practitioner profile, the profile including at least one of historical data, career data, demographic data, and longitudinal data.
[0209] In some embodiments, the conversation data includes at least one of utterance data and text-based data.
[0210] In some embodiments, method 300 is configured to run a model based on human-interpretable features.
[0211] In one embodiment, a non-temporary computer-readable medium is provided that includes program instructions for causing a computer to perform one of the above methods.
[0212] Interactive voice response In some embodiments, the system may be configured to manage an interactive voice response (IVR) system.
[0213] Interactive Voice Response (IVR) technology can enable a user to interact with a computer-operated telephone system via voice or dual-tone multi-frequency input. IVR systems often utilize pre-recorded or dynamically generated voices for communication with the user. The systems described herein may be encountered, for example, for voicemail services, to navigate call routing systems, or when a wrong number is dialed. IVR systems can influence the model. For example, the system may attempt to segment and catalog correspondences with the IVR system, or attempt to score a patient partially based on the voice from the IVR system.
[0214] In some embodiments, the system is configured to recognize and remove or omit speech originating from an IVR system. In some embodiments, the system is trained to recognize an IVR system and appropriately exclude such speech. In some embodiments, the system may exclude speech from an IVR system from speaker classification, or it may include it as a special "IVR" speaker (different from human speakers). In some embodiments, the system is configured to exclude IVR speech from patient scoring.
[0215] In some embodiments, the system may be configured to identify the IVR speaker within the speaker segment. In some embodiments, the content of the IVR speaker may be used to aid in prediction. For example, the IVR voice may be used to score the patient's condition if the IVR system is, for example, a call triage system for a mental health clinic. In such embodiments, the system may be configured to evaluate the language used by the patient (or dual-tone multi-frequency input from the patient's phone) to score the user (e.g., to extract the context of the call) using, for example, the language used by the IVR system within NLP.
[0216] In some embodiments, IVR systems may be recognized based on words and phrases commonly used in IVR systems (e.g., "Press 1 for...", "Hang up and call again"). In some embodiments, IVR systems may be recognized by specific voice patterns they exhibit (e.g., if there are patient speakers and IVR speakers, the system may identify a monotonous, professional, and somewhat cheerful speaker as the IVR speaker). In some embodiments, the system may be trained against sample IVR systems to quickly identify IVR systems and manage their data accordingly.
[0217] In some embodiments, the systems described herein may generally be able to identify calls that should be rejected (e.g., IVR, responses from family members, and patient non-participation). Some such systems may, for example, identify when an IVR system has been activated. Some systems may identify conversation data that matches a situation where a family member answered the phone and the patient was not participating (e.g., shorter conversation, negative response to whether the patient could participate). Such systems may improve the reliability and performance of the model because inappropriate content is not scored.
[0218] Once roles are assigned at the beginning of the conversation, the system can use turn-taking to analyze data from each speaker in different ways. A language model may analyze the patient's utterances in more detail. An acoustic model may, for example, weight the patient's voice more highly and analyze it with a more sophisticated model.
[0219] NLP models (described in more detail below) may or may not use turns, but turns can be beneficial to acoustic models in identifying who is speaking and when (and therefore giving greater weight to patient utterances). The person in charge is generally used as a starting point (e.g., after the role of the person in charge has been identified by the system), and the system then annotates that the data is between the person in charge and the patient. The system may also annotate other speakers (e.g., caregivers and interpreters) if they are present.
[0220] Topic detection The topic detector 122 may be configured to process data at will to determine the topic of conversation. The topic detector 122 may have access to common topics of conversation segments for processing or evaluation purposes. For example, topics may be evaluated to determine which segments of conversation deserve more detailed examination (i.e., more robust analysis) or higher importance (i.e., higher weighting).
[0221] In some embodiments, the system may use information related to the detected topics to estimate the scoring confidence. In some embodiments, heuristics may be used based on the number and types of topics, as well as their duration (or word count) or region associated with each topic. For example, the system may be configured to score the confidence of the output higher if the topics mentioned in the scored session are the same as the topics encountered during training.
[0222] These topic statistics can also be used to improve model training and scoring. In some embodiments, the results of topic detection can be provided as input features for the model. In some embodiments, regions corresponding to particular topics may not be very useful for scoring the session and can therefore be lightly weighted or completely excluded from scoring.
[0223] Some embodiments may provide some explainability of the prognosis provided by the system by highlighting and flagging relevant sentences, expressions, or words from conversations that were important to the evaluation proposed by the system.
[0224] Conversations between patients and therapists can be lengthy (30 minutes to an hour or more is not uncommon) and may be filled with information less relevant to predicting the severity of symptoms of mental health conditions. The systems and methods described herein can be used to identify the most prominent areas of conversation and to predict cost-effectiveness and accuracy (particularly in language models). For this purpose, automated topic detection can be used. Topics can be detected in therapist utterances, patient utterances, or both (e.g., patient responses may require contextual understanding of therapist queries). Topics related to mental health and related conditions (e.g., social determinants of health (SDOH)) can be constructed using multiple methods (empirical, clinical, rule-based). Areas containing estimated relevant topics are given arbitrary higher weighting in the modeling to improve performance. For example, areas of conversation related to medicine, exercise, living situation (e.g., living alone), and interpersonal relationships (e.g., family and friends) may be more important than other parts of the conversation (and therefore can be weighted higher). Furthermore, other prominent content that correlates with predictive accuracy may be included and weighted higher. You may also evaluate areas related to the same topic to determine if there are many negative, positive, or neutral words (for example, negative words related to family may indicate poor family support and could be helpful in predicting mental health).
[0225] In some embodiments, the positive / negative value of a word may be determined from modeling rather than from the sentiment of the word. For example, "tennis" may be a predictor of patients who are not depressed, while "bed" may be a predictor of patients who are depressed.
[0226] Prominent topic areas can be made available to the system to optimize processing efficiency. Once a portion of a conversation is identified as relating to a prominent topic, it may be beneficial to perform more expensive speech recognition on that portion. Third-party automatic speech recognition (ASR) can be expensive, and reducing the amount of conversation that would be processed using this ASR can be very beneficial to the cost of these systems. While ASR may be necessary to identify topic areas themselves, once identified, these areas can be prioritized for more expensive and accurate ASR for better predictive modeling.
[0227] Evaluating a topic may require analyzing linguistic data from both the patient and the caregiver. While the patient's responses are the data of interest, the caregiver's queries (or statements) may also be evaluated to confirm the context of the patient's responses. For example, the answer "yes" has entirely different meanings depending on whether it is in response to the question "Can you hear me?" or "Do you have suicidal thoughts?"
[0228] In some embodiments, the system may use computationally lighter models to find clues that may contain relevant topics. For example, the system may analyze a session for clusters of words or phrases related to a topic of interest, or it may identify points in the patient's voice where there are significant differences in emotional expression. The model may be trained to further analyze important areas of the session in a different way than the rest of the session (e.g., by using a more robust analytical model). The advantages of such a system include increased efficiency and reduced data noise. Specifically, because language models use the session history, the computational complexity can increase exponentially with the length of the session. Computational resources can be saved by omitting some parts of the session. Furthermore, these models can be expensive to use, and therefore, using them on a subset of the session (rather than the entire session) can save financial costs while improving the efficiency of the model.
[0229] In some embodiments, identifying the topic of conversation can be beneficial. This allows the system to focus on or further process segments of conversation that deal with important topics such as medicine, exercise, or interpersonal relationships (i.e., using a computationally more robust model).
[0230] Figure 4 shows a method 400 for detecting topics in a conversation, according to several embodiments.
[0231] In one embodiment, a method 400 for identifying topics in a conversation is provided. The method includes receiving conversation data from at least one input device (402); processing the conversation data to generate a language model output (404) wherein the language model output includes one or more topics corresponding to one or more time ranges; applying weights to the output to generate a weighted output (406) wherein the output includes multiple outputs corresponding to multiple time segments of the conversation data, and the weights are optionally time-based and partially based on one or more topics in each time segment; generating an electronic report (408); and transmitting the electronic report to an output device (410).
[0232] In some embodiments, the method further includes processing conversational data to generate an acoustic model output, and fusing the language model output and the acoustic model output by applying weights to the language model output and the acoustic model output and generating a composite output from the weighted output, wherein the acoustic model output includes multiple outputs corresponding to multiple time segments of the conversational data, and the weights are optionally time-based and partially based on the topics in each time segment.
[0233] In some embodiments, the time range of conversational data corresponding to a given topic is processed using a computationally more robust model than the one used for the time range of conversational data not corresponding to the given topic, in order to generate language model output.
[0234] In some embodiments, the electronic report includes a transcript of language model output annotated with weights partially based on the topics in each time segment.
[0235] In some embodiments, the electronic report identifies the severity of at least one symptom of the condition based on the language model output.
[0236] In some embodiments, the state includes a mental health state.
[0237] In some embodiments, the electronic report includes annotations of language model output that indicate the prominence of at least one of one or more time segments.
[0238] In some embodiments, this further includes determining at least one role of at least one speaker, and the weights are based in part on at least one role of at least one speaker during each time segment.
[0239] In some embodiments, the language model output includes the identification of at least one query based on conversational data from the person in charge, and at least one response to at least one query based on conversational data from the patient, wherein at least one query is mapped to a predetermined query, and at least one response to at least one query is mapped to a predetermined response to the predetermined query.
[0240] In some embodiments, processing conversational data to generate language model output involves using a language model neural network trained on labeled conversational data collected from one or more other subjects, where the labeled conversational data is labeled as (i) having a state at some level and (ii) not having that state.
[0241] In some embodiments, at least one query belongs to a set of queries, and the computing device is configured to predict an overall score based on the set of queries based on the response to each of the queries in the set.
[0242] In some embodiments, conversation data is processed based in part on at least one of a patient profile and a practitioner profile, the profile including at least one of historical data, career data, demographic data, and longitudinal data.
[0243] In some embodiments, the conversation data includes at least one of utterance data and text-based data.
[0244] In some embodiments, method 400 is configured to run a model based on human-interpretable features.
[0245] In one embodiment, a non-temporary computer-readable medium is provided that includes program instructions for causing a computer to perform one of the above methods.
[0246] LLM / Chatbot / ChatGPT Implementation In some embodiments, predictive models can incorporate large-scale language models (LLMs). LLMs have the advantage of using less labeled mental health data (which may be rare or expensive to generate). They can also include emotional and sentimental content. Thus, LLMs can serve as a source of information for prediction. These LLMs can be used in various combinations and structures. Mental health predictions can be fine-tuned or adapted. LLMs can be used as generative models to provide real-time guidance and suggestions to healthcare professionals during person-to-person (H2H) conversations, such as what or tone to use in the next turn.
[0247] In some embodiments, results from large-scale language models can be used directly in fusion with other models. In some embodiments, results from large-scale language models can be used in heuristic or rule-based setups, cascade-based setups, confidence estimators, or any type of trained model, either alone or in combination with the models described herein (e.g., one or more acoustic and / or language models), to generate more accurate final estimates.
[0248] In some embodiments, when used as a confidence estimator, the actions taken by the system to provide the user with final information may be modified to take the confidence estimator into account. For example, if a large language model produces a response that deviates significantly more than a more well-trained model, the system may take actions such as indicating that the result may be less confident, troubleshooting data processing, or requesting more input.
[0249] In some embodiments, the chatbot may be a prompt-based model that, for example, asks a large language model about the severity level of PHQ (or other mental health condition) based on transcription. This answer may provide confidence and validity. The model may also ask for estimated answers to individual questions based on transcription. These individual answers may have confidence and validity. Such individual answers may be aggregated into a final score.
[0250] In some embodiments, the transcript may be processed, for example, with a large-scale language model (e.g., ChatGPT) to identify content regions relevant to the patient's mental health (perhaps weighted according to their relevance). This can then be used to highlight these regions during inference or training. In some embodiments, the transcript may be processed, for example, with a large-scale language model (e.g., ChatGPT) to identify content regions unrelated to the patient's mental health (perhaps weighted according to their unrelevance). This can then be used to reduce the importance of these regions or to skip them during inference or training.
[0251] In some embodiments, the transcripts may be processed by a large-scale language model (e.g., ChatGPT) to summarize aspects of several PHQs and / or GADs (which can be extended to other broader behavioral aspects such as stress, life satisfaction, and wellness). For example, each individual PHQ question may include aspects such as: "Do you have little interest or pleasure in doing things?", "Do you feel depressed, gloomy, or hopeless?", "Do you have trouble falling asleep, wake up in the middle of the night, or oversleep?", "Do you feel tired or lacking energy?", "Do you have a poor appetite or overeat?", etc. Alternatively, each individual PHQ topic may include: "Interest or pleasure in doing things", "Depression, gloom, or hopelessness", "Sleep", "Energy levels, fatigue", "Eating habits", etc. Thus, a large-scale language model (e.g., ChatGPT) can extract and summarize individual aspects, including confidence levels (the amount of evidence found for each individual aspect) and validity.
[0252] Such data can be used as input features for the inventors' model (for training and inference). This can be used alone or in conjunction with existing transcriptions. Such implementations can be used alone or in conjunction with existing transcriptions.
[0253] Data augmentation during inference time In some embodiments, the system may augment the data from the patient (for example, modifying both the NLP and acoustic model signals by paraphrasing, masking, cutting, substituting with random words, introducing ASR errors (substitution for similar pronunciation), phase shifting, adding white noise, etc.). Such techniques can enable several inferences on the same data points and yield a distribution of predictions (or a distribution of distributions, if the model can generate probability scores) instead of a single prediction.
[0254] At runtime, each model of interest may have a component that instructs it to generate N different input samples using (for example, a given) augmentation approach for each model. Each model is then run on each augmentation test sample, and the results can be combined to provide a better performance estimate. For example, each model may apply score fusion to the results, create a distribution, and use the distribution to derive a better performance estimate. This may involve presenting the distribution itself or using point estimates from the distribution with confidence.
[0255] In some embodiments, language data can be augmented through the use of large-scale language models. Such models may be used to generate different ways in which a patient or agent might say the same thing (e.g., paraphrasing).
[0256] In some embodiments, acoustic data may be augmented through the use of synthesized speech. Synthesized speech (or other methods) may be used to generate different intonations when speaking the same or different content.
[0257] These can provide data augmentation for the purpose of training models. In fields such as mental health assessment, obtaining sufficient raw data for training can be difficult (it is even more difficult because the data has privacy constraints and needs to be labeled soon after collection).
[0258] Natural Language Processing Models (NLP) The language processor 116 may be configured to process speech data and extract transcript data and language data from the speech data received from the input device 102. The language processor 116 may use an algorithm to evaluate the language data. The language processor 116 may use a natural language processor (NLP) model, or a combination of different language models or language models parameterized in different ways. The NLP model may use words as input and may not use acoustic signals directly.
[0259] The purpose of including NLP models is generally to help predict trends in NLP models within this domain. In some embodiments, for NLP tasks, the speech signal may first be transcribed using a publicly available ASR service. A high ASR error rate may be acceptable for the NLP model (possibly due to good cue redundancy). The NLP model may be based on a transformer architecture and may utilize transfer learning from language modeling tasks. A DeBERTa pre-trained model can be used. Alternatively, a RoBERTA, ALBERT, or BERT model can be used.
[0260] DeBERTa may have the advantage of having fewer (435M) parameters than other comparable models while still offering good performance. This can be useful when a large number of experiments need to be performed. In such embodiments, the DeBERTa model can be pre-trained on over 80GB of text data from the following common corpora: Wiki, Books, and OpenWebtext. The input context window length can be set to 512, and the tokenizer can be trained on 128K tokens.
[0261] In such an embodiment for fine-tuning, a predictive head can be attached to a language model, and a binary classifier can be trained. All hyperparameters except the learning rate may be fixed. The learning rate can be set proportionally to the amount of training data for each experiment (grid search approach). Early termination can be used to avoid excessive use of execution time.
[0262] In some embodiments, the language model may be selected from a group consisting of mood models, statistical language models, topic models, syntactic models, embedding models, dialogue or discourse models, emotion or affect models, and speaker personality models. Examples of NLP algorithms include semantic analysis, mood analysis, vector space semantics, and relation extraction. In some embodiments, the methods described herein may be able to generate assessments without requiring the presence or intervention of a person in charge. In other embodiments, the methods described herein may be used to extend or enhance assessments provided by a person in charge, or to assist a person in providing assessments. Assessments may include queries containing subjects adapted or modified from screening or monitoring methods such as PHQ-9 and GAD-7 assessments. Assessments described herein may not only use questions from such surveys verbatim, but may also adaptively modify queries based at least in part on responses from the patient in question.
[0263] In some embodiments, the language model may be trained using a person-to-device corpus. In some embodiments, the language model may be trained using a person-to-person corpus.
[0264] In some embodiments, the language model may not use any information other than speaker metadata or information obtained from processing recorded conversations. The language model may use a modern deep learning architecture. The language model may use large amounts of data, including out-of-domain data, for pre-training the model. The language model may be based on a transformer architecture and may utilize transfer learning from language modeling tasks.
[0265] The systems and methods disclosed herein may use natural language processing (NLP) to perform semantic analysis on patient utterances. As disclosed herein, semantic analysis may refer to the analysis of spoken language from a patient's responses to assessment questions or from captured conversations in order to determine the meaning of spoken language for the purpose of conducting screening or monitoring of a patient's mental health. The analysis may be of words or phrases and may be configured to consider primary or follow-up queries. The analysis may also be applied to the utterances of an operator. Where used herein, the terms “semantic analysis” and “natural language processing (NLP)” may be used interchangeably. Semantic analysis may be used to determine the meaning of a patient’s utterance in context. It may also be used to determine the topic the patient is talking about.
[0266] With respect to weighted word scores, the semantic model of the language processor 116 can estimate the patient's health status from the positive and / or negative content of the patient's utterances. The semantic model correlates individual words and phrases with specific health conditions designed for the semantic model to detect. System 100 can retrieve all responses to a given question from the collected patient data and use the semantic model to determine the correlation of each word in each response to one or more health conditions. The weighted word score of an individual response is the statistical mean of the correlation of the weighted word scores. System 100 can quantify the quality of the question as a statistical measure of the weighted word scores of the responses, for example, as its statistical mean.
[0267] The language processor 116 may include several text-based machine learning models to (i) predict depression, anxiety, and possibly other health conditions directly from words spoken by the patient, and (ii) model factors that correlate with such health conditions. Examples of machine learning that directly model health conditions include mood analysis, semantic analysis, language modeling, word / document embedding and clustering, topic modeling, discourse analysis, syntactic analysis, and dialogue analysis. The model does not need to be constrained to one type of information. The model may include information from both mood and topic-based features, for example. The language information may include score outputs from specific modules, e.g., scores from a mood detector trained on mood rather than mental health conditions. The language information may include information acquired via a transfer learning-based system.
[0268] The language processor 116 may store text metadata and modeling dynamics, and may share this data with the acoustic model 118. The text metadata may include, for example, data that identifies the speech portion (syntactic analysis), sentiment analysis, semantic analysis, topic analysis, etc., for each word or phrase. The modeling dynamics include data that represents the components of the constitutive model of the language processor 116. Such components include the machine learning features of the language processor 116, and other components such as long short-term memory (LSTM) units, gated recurrent units (GRUs), hidden Markov models (HMMs), and sequence-to-seq (seq2seq) translation information.
[0269] In some embodiments, the navigator can receive language model output and semantically analyze the linguistic results of the command language in near real time. Such commands may include phrases such as "Could you say that again?", "Speak louder," and "I don't want to talk about that." These types of "command" phrases indicate to the system that an immediate action is being requested by the user.
[0270] Language models can utilize words modeled through natural language processing. NLP models for native and second-language speakers can also differ significantly. Even across generations, NLP models vary considerably to address nuances in slang and other forms of speech. Making models of different granularities available to individuals allows for the application of the most appropriate model, thereby significantly improving the classification accuracy of these models.
[0271] Many approaches can be used for model inference on data in languages other than the target language. First, the incoming data can be processed via a language ID to determine the language. Then, one or more of the many approaches can be used. In some embodiments, the non-target language portion may be skipped or omitted separately from language scoring. In some embodiments, the system may translate the incoming data into the target language, for example, using a third-party translation. The original target language model can then be used. In some embodiments, a model for a new language can be created by translating training data into the new language and training a new model. In some embodiments, a model specific to the language used in the incoming data (for example, created from training data for that language) can be selected. At runtime, the appropriate model for the incoming language can be selected. In some embodiments, a model trained to work in multiple languages (a multilingual model) can be used. Such approaches are described in PCT Patent Application No. PCT / US2022 / 015147 (published as WO2022169995A1) (incorporated herein by reference).
[0272] The reliability of the model may be affected by languages not present in the training data. For example, foreign languages may not be adequately represented in the training data, making it difficult to provide high confidence ratings. In a system that translates patient utterances from a foreign language to the target language, the analysis of word string selections made by the patient may be hindered.
[0273] In some embodiments, the speaker (e.g., patient or caregiver) may switch languages during a session. In such embodiments, it may be desirable for the system to continuously monitor the language for changes in language ID. In embodiments where there is little discussion in the non-target language, it may be easy (e.g., computationally convenient) to omit the analysis of that section of utterance without negatively impacting the model. However, especially when longer sessions are conducted in the non-target language, it may be advantageous to use one of the other approaches described above (e.g., translating incoming data, using a language-specific model, or using a multilingual model). The approach taken by the system may differ during a session (e.g., the system may initially ignore the non-target language until the volume of utterances reaches, for example, a certain duration threshold, and / or until, for example, the analysis value of the non-target language portion exceeds the computational overhead of analyzing it). Embodiments that can confirm and adapt to language switching may be advantageous when speakers switch languages (for example, when a patient can express certain concepts in the target language and other concepts in the non-target language, or when a patient and caregiver speak with a person in the target language and with each other in the non-target language), or when speakers are using different languages (for example, when a patient is speaking in the non-target language, which may be their native language, but an interpreter or caregiver is speaking with a person in the target language).
[0274] The systems and methods described herein may function without prior knowledge of the speaker's voice, language, or metadata (i.e., the only input that may be required to estimate the severity of the symptoms of the speaker's mental health condition is speech from a conversation).
[0275] In the context of modeling a single speaker over time, longitudinal information can be used to better calibrate predictions (to determine whether the patient's severity is increasing or decreasing).
[0276] Acoustic model The acoustic processor 118 may be configured to process speech data and extract conversational data from speech data received from the input device 102. The acoustic processor 118 may be configured to analyze specific conversational data of a patient to determine whether the patient has a health condition. The speech model can use acoustic information in the signal. The acoustic processor 118 can analyze the speech portion of the audiovisual signal to find patterns associated with various health conditions (e.g., depression). The association between acoustic patterns in speech and health is, in some cases, applicable to different languages without retraining. They may also be retrained with data from that language. Thus, the acoustic processor 118 can analyze the signal in a language-independent manner. The acoustic processor 118 can use machine learning approaches such as encoder-decoder architectures, convolutional neural networks (CNNs), long short-term memory (LSTM) units, and hidden Markov models (HMMs) to learn high-level representations and model the temporal dynamics of the audiovisual signal. The acoustic model may use the speech signal as input, but not directly use words.
[0277] The purpose of including acoustic models is to illustrate how acoustic models and signal-based models generally behave in this domain. In some embodiments, the acoustic models used may be based on encoder-decoder architectures. Alternatively, models may be used that utilize other deep learning architectures, including long-shorter-term memory (LSTM) and convolutional neural networks (CNNs). Acoustic models can function in two stages. Utterances can initially be segmented every 25 seconds. The model can learn latent representations at the segment level using filter bank coefficients as input. The representations can then be fused to make predictions at the response or session level. Model training can utilize transfer learning from automatic speech recognition (ASR) tasks. The ASR decoder can be discarded after the pre-training stage, and the encoder can be fine-tuned along with the prediction layer.
[0278] In such an embodiment, the encoder can consist of a CNN, followed by a layer of LSTM. The prediction layer can use a Recursive CNN (RCNN). The last layer of the RCNN module can be used as a vector representation of a given audio segment. It can be passed to a fusion module that makes the final prediction for the session, along with other representations of the same session. The fusion model can use a max operation over all the vectors to obtain an aggregated representation of the session and then use a multi-layer perceptron to make the final prediction.
[0279] In some embodiments, the acoustic model can include one or more of an acoustic embedding model, a spectral-temporal model, a supervector model, an acoustic emotion model, a speaker personality model, an intonation model, a speaking rate model, a pronunciation model, a non-verbal model, or a fluency model.
[0280] In some embodiments, the acoustic model can be trained using a human-to-device corpus. In some embodiments, the acoustic model can be trained using a human-to-human corpus.
[0281] In some embodiments, the acoustic model may not use any other information other than the speaker's metadata or information obtained from the processing of the recorded conversation. The acoustic model may not use pre-extracted features such as pitch or energy, and instead the input may use the raw utterance signal. Raw input and advanced modeling may produce better results than an approach that uses pre-extracted features. The acoustic model may use state-of-the-art deep learning architectures. The acoustic model may use a large amount of data, including out-of-domain data, for pre-training of the model. The acoustic model may use transfer learning from an automatic speech recognition task, followed by fine-tuning.
[0282] In some embodiments, the acoustic model can include a wav2vec style model. In some embodiments, the acoustic model can include a multilingual acoustic model.
[0283] In some embodiments, the system may utilize a speech recognizer. In some embodiments, the system may utilize hesitancy, latency, absence of laughter, or other information in the conversation data. In some embodiments, the system may utilize the sound of words precisely. In some embodiments, filler words may be used as input. For example, a person presenting depression may use more filler words, and thus, the analysis of filler words can improve the accuracy of the model.
[0284] The acoustic processor 118 can store data representing the attributes of the audio signal and the machine learning features of the acoustic processor 118 as acoustic model metadata and share that data with the language processor 116. The acoustic model metadata can include, for example, data representing the spectrogram of the visual-auditory signal of the patient's response. Further, the acoustic model metadata can include both basic features and high-level feature representations of the machine learning features. More basic features can include, for example, the mel-frequency cepstral coefficients (MFCC) of the acoustic processor 118, and various log filter banks. High-level feature representations can include, for example, the convolutional neural network (CNN), autoencoder, variational autoencoder, deep neural network, and support vector machine of the acoustic processor 118. The acoustic model metadata enables the language processor 116 to improve the sentiment analysis of words and phrases, for example, using the acoustic analysis of visual-auditory signals.
[0285] In some embodiments, specific models of the acoustic processor 118 can be generated to classify acoustic signals as either representing individuals in a depressed state or not. The voice quality, pitch, and speech rate of the voice input can vary significantly between young and elderly individuals. Therefore, specific models can be developed based on whether the patient being screened is young or elderly. Similarly, women generally exhibit greater variability in acoustic signals compared to men, suggesting that yet another set of acoustic models may be advantageous. It is also clear that combined models for young and elderly women, versus young and elderly men, would be desirable. Clearly, the number of applicable models increases exponentially as further individualization groups are generated.
[0286] In some embodiments, conversation data is provided to a high-level feature expressor that works in conjunction with a temporal dynamics modeler. A model conditioner influences the operation of these components. The high-level feature expressor and temporal dynamics modeler also receive raw feature extractor outputs and higher-level feature extractor outputs that identify features within the input acoustic signal, and feed them to the model. The high-level feature expressor and temporal dynamics modeler generate results for an acoustic model, which may be fused into a final result classifying the health status of an individual, or consumed by other models for conditioning purposes.
[0287] High-level feature expressors involve leveraging existing models for frequency, pitch, amplitude, and other acoustic features, which provide valuable insights into feature classification. Several off-the-shelf "black-box" algorithms accept acoustic signal inputs and provide classifications of emotional states to a commensurate degree of accuracy. For example, emotions such as sadness, happiness, anger, and surprise can already be identified in acoustic samples using existing solutions. Additional emotions such as jealousy, tension, excitement, ecstasy, fear, disgust, trust, and anticipation will also be leveraged as development progresses. However, this system and method goes further by matching these emotions, their intensity, and confidence in the emotions to patterns of emotional profiles that indicate specific mental health states. For example, pattern recognition may be trained to identify the emotional states of depressed respondents based on patients known to suffer from depression.
[0288] The acoustic processor 118 may take into account, for example, pitch / energy, timbre / vocalization, speech fluency, and articulatory coordination. In some embodiments, the acoustic model uses phonemes. In some embodiments, the acoustic model directly uses signal patterns (i.e., a 2vec acoustic model).
[0289] In normal conversation, interjections refer to interjections made by one speaker in response to another, indicating attention, understanding, sympathy, or agreement. These can include both verbal responses ("I understand") and nonverbal responses (nodding).
[0290] In some embodiments, feature extraction may be performed on acoustic data. In some embodiments, acoustic modeling may be performed directly on the energy distribution or spectral / cepstral representation across frequency bands.
[0291] In some embodiments, there may be N labels representing different mental states and / or conditions (e.g., stress, anxiety, and depression). In some embodiments, there may be three labels (stress, anxiety, and depression), and models may be trained separately. In some embodiments, the system may output results for all three models. The same model can be trained and optimized for different metrics. Some models may have different properties (e.g., different types of acoustic models).
[0292] When translating an input language to a target language, acoustic data may not be translated. If the language in which the system is trained is not the same as the input language, it is possible to run the acoustic system "as is" or run a system that has been adjusted for the language closest to that language. Such embodiments can still provide useful information from both the linguistic and acoustic data. In a preferred embodiment, the new model is generated based on training data from the input language. In such embodiments, subtle nuances in the acoustic data can be better analyzed with training data originating from the same language as the input language.
[0293] In some embodiments, the model may be penalized for learning speaker information. The training dataset may contain multiple sessions from the same patient. Therefore, it may be advantageous to actively penalize the model for learning speaker information. During training, the model may learn to predict mental health status by predicting speaker identity rather than relying on generalizable information from conversational data. In some embodiments, during training, the model may be asked to predict speaker ID in addition to mental health status, and may be penalized if it succeeds in doing so. For example, a term may be added to the loss function so that the loss is larger when the speaker ID is correct. Other methods of penalizing the model are also conceivable.
[0294] Fusion of short-time segment outputs In some embodiments, the input may be segmented into shorter time domains. For example, an acoustic model may need to perform segmentation due to the large amount of information in each segment. As a further example, an NLP model may also need, or benefit from, reduction from long strings into regions of interest containing important information, based on the results of light training (e.g., skipping noise regions (which contain no information for the language model) but preserving long-term dependencies).
[0295] In such embodiments, scoring by the model may be performed for each region. These scored segments may then need to be fused. Different methods for fusion (e.g., maxpool, mean, std, var, etc.) may yield different answers, and combinations can be used. Segment fusion can be performed in such a way that the segment-based model learns representations that are less correlated with the output of the NLP model, potentially resulting in better combined performance.
[0296] Fusion of outputs from different models The fused model 120 may optionally be configured to combine the outputs from the language processor 116 and the acoustic processor 118. The estimations of both types of models can optionally be fused to provide better performance and robustness. The fused model 120 may be configured to assign human-readable labels to scores. The fused model 120 may combine the model outputs into a single, unified classification of health states. Model weighting may be performed using static weighting, such as weighting the outputs (e.g., weighting the language processor output more than the acoustic processor output). However, more robust and dynamic weighting methodologies can also be applied. For example, the weights of a given model output may, in some embodiments, be modified based on the confidence level of the model's classification. For instance, if the language model's output renders a person as not depressed with a confidence level of 0.56 (out of 0.00 to 1.00), while the acoustic model's output classifies them as depressed with a confidence level of 0.97, in some cases the model output weights may be weighted such that the acoustic model's output receives a greater weight. In some embodiments, the weights of a given model can be linearly scaled by multiplying the model's confidence level by a baseline weight. In yet another embodiment, the output weights are time-based. For example, generally, a language model output may be weighted more heavily than other outputs, but if there is a video output, the video output may be weighted more heavily for its time domain when the user is not speaking. Similarly, if the video and acoustic model outputs independently suggest that the person is nervous and dishonest (e.g., frequent eye shifts, increased sweating, increased pitch modulation, increased speech rate), the weight of the language model output may be minimized because it is more likely that the individual is not answering the question honestly.
[0297] In some embodiments, the fusion model may be configured to modify the fusion weights using confidence levels (e.g., for the entire conversation or by region of the conversation). In some embodiments, the system is configured to calculate confidence levels for segments or turns of conversation and combine them in an acoustic model, for example. For example, the system may evaluate segments or turns for confidence levels and allow only the analysis segments or turns that produce a particular confidence level.
[0298] In some embodiments, the system may verify the performance of the personnel and use it to adjust the weights of the fusion model (or other models).
[0299] In some embodiments, the system may be configured to handle speaker segmentation failures, for example, by adjusting the fusion model. In some speaker segmentation failures, the system fails to recognize the patient and the person in charge, for example, by misidentifying both the patient and the person in charge as the same person. In such situations, the system may be configured to identify the failure due to detecting only one speaker (i.e., in this example, a speaker segmentation failure, where at least two speakers are required) and apply custom weights to the fusion model.
[0300] Custom weights can be 0 for acoustic model output and 1 for language model output. Even if speaker classification fails, the language of both the patient and the caregiver can still indicate the patient's mental state. Acoustic models may be relatively unhelpful because they mix acoustic data from both the patient and the caregiver. Speaker classification failures can also occur if it is not possible to distinguish between the patient and, for example, the role of a caregiver or interpreter.
[0301] In some embodiments, the fusion weights may be automatically adjusted based on the extracted information. For example, information such as signal-to-noise, duration, number of words per role, detected topic, estimated gender or age, estimated distance to the microphone, reverberation level, detected background noise type (e.g., car engine, TV / radio conversation, TV / radio music) may be used to automatically adjust the fusion weights. For example, based on the detection of information, the system may identify that there is a negative impact on the acoustic data and that it is reducing its probative value (e.g., detection of a reverberation level that interferes with the acoustic data), and the fusion model may adjust the weight of the language model output higher and the weight of the acoustic model output lower than in a situation without this impact. The magnitude of the correction to the fusion weights may be partially based on the magnitude of the detected impact in the extracted information (e.g., a quiet car engine sound may cause little or no adjustment of the fusion weights, while a noisy car engine sound may cause a decrease in the weighting of the acoustic model output compared to the language model output). In some embodiments, the characteristics of the extracted information may cause a specific weight change (e.g., if the number of words spoken by a patient increases beyond a threshold, the system may focus more sharply on the acoustic data by weighting the acoustic data more compared to when the threshold is not exceeded).
[0302] In some embodiments, the fusion weights may partially depend on metadata (e.g., gender, age, patient category, call category, etc.). For example, based on some metadata, the fusion weights may be adjusted (e.g., in the form of a dynamic modifier adjusted based on a modifier over the entire session or additional variables evaluated within the session itself). For example, the metadata may indicate that the patient is elderly, and the system may adjust the fusion weights to weight the language model output higher than when the patient is younger (e.g., an elderly patient may speak more quietly, and thus the acoustic model may be likely to produce less reliable results).
[0303] After merging and weighting the model outputs, the resulting classifications can be combined with features and other user information to generate the final result. These results are provided to the user for storage and potentially as future training material, and are also provided to the output device 126 for display.
[0304] Detection of multiple health conditions In some embodiments, the systems and methods described herein are trained and operated as models capable of evaluating multiple health conditions simultaneously. In some embodiments, the system may be configured to evaluate mental health conditions when they are correlated within a population.
[0305] Such models may have an advantage because they can perform better by capturing the correlation between states in a population. Such models may also be able to use assessments of certain states to assist in determining the state of interest (for example, when the model does not have enough training data to accurately label the state of interest). For instance, even if there is no or limited training data for depression, the system may be configured to assess both anxiety and depression. Such models may also be more efficient when multiple state outputs are desired.
[0306] Such systems may be more useful in populations where the correlation between two states is similar to or the same as the correlation in the training data. For example, anxiety and depression may be highly correlated in younger populations, but the correlation between anxiety and depression may be weak or nonexistent in older populations. Therefore, a model trained to have a high correlation between anxiety and depression may perform better in younger populations.
[0307] Interpretability In some embodiments, the system may be configured to run the model based on human-interpretable features (e.g., pitch in the case of acoustic modeling, pronoun usage in the case of linguistic modeling) in addition to the systems and methods described herein. In some cases, the models described herein perform better than models based on human-interpretable features, but by running these human-interpretable models in addition to the models described herein, the system may be able to provide some explainability for the output of the model. In some embodiments, an interpretable feature in one type of model (e.g., an acoustic model, a linguistic model) may be used to justify a different type of model.
[0308] In some applications, explainability may be desirable. For example, certain regulatory regimes require that predictions from models regarding mental health status be justified in some way. By using human-interpretable base models in addition to the models described herein, the system may be able to satisfy such requirements. For example, pitch variation (e.g., monotone) and pronoun use (e.g., frequent use of personal pronouns) may be used to justify the finding that a patient has depression, but the deep learning model itself is not improved by including such features.
[0309] In some embodiments, explainability may also incorporate longitudinal information from, for example, patient profiles. In such embodiments, the system may not only perform an analysis on patient information in a single session, but also analyze how it compares with conversational data from past sessions. Such an analysis may reveal, for example, that the patient uses more personal pronouns, which may provide a justification for a higher score on depression metrics, for example.
[0310] Conditional estimation In some embodiments, metadata may be used to customize and / or modify model predictions. For example, medical records, drug treatments, obesity, family history, and other information may be supplied to the model. The model may or may not use this information. For example, the system may use this metadata to modify predictions or select a different model to analyze the patient. In some embodiments, the system may use metadata only if the model predictions exhibit low confidence (e.g., are used only as conditional estimates), and in such situations, the system may use metadata to improve the confidence of the model predictions.
[0311] This conditional estimation can be used as a screening tool. For example, based on metadata from patients, the model can be modified based on those conditions. In this way, the model can treat different patients differently, just as it can treat patients differently. Furthermore, it can provide a convenient way to enable comparisons between different patients (for example, the system can attribute part of the adjustments to this metadata).
[0312] In some embodiments, conditional estimation may be used to cancel out differences arising from changes in circumstances between sessions. For example, patient conversation data may be influenced by various factors that are not causally related to mental state (e.g., time of day, people around during the session, whether the patient was in a hurry, or whether they had a near-miss accident immediately before the session). The system can be configured to take these inputs and adjust the predictions made for the patient in that session (e.g., if the patient had a near-miss car accident before the session, panic or agitation in the patient's conversation data may be attributable to a reaction to a near-miss, and the model prediction can be adjusted to take that attribution into account).
[0313] In some embodiments, these conditional estimations can be used longitudinally. For example, the system tracks metadata associated with each session with a patient and develops a model to understand how the patient behaves at different times of the day. For instance, the model might understand that the patient is a night owl and therefore less attentive or willing to respond in sessions conducted early in the morning. In this situation, the system could modify the model for early morning sessions to attribute at least some of the patient's possible inattention or irritability to the time of day, as opposed to the patient's state. In further embodiments, the system may attempt to schedule sessions with the patient at times that are historically more useful (e.g., if the patient is a night owl, schedule meetings at the same time to ensure a consistent patient state; or schedule meetings at later times when the patient is more likely to provide sufficient responses for the system to assess).
[0314] Survey scoring The survey scorer 124 can optionally be configured to process data and score surveys (e.g., PHQ-9, PHQ-2, GAD-7).
[0315] Survey scoring can be based on conversation. Surveys may include PHQ-9, PHQ-2, and GAD-7. Survey scoring may be used independently of predicting mental health. The systems and devices described herein may include dedicated NLP and rule-based detectors trained to find agent questions that map to survey questions. The system may then observe the patient's responses and map them to one of a set of forced-choice options (e.g., “never,” “sometimes,” “frequently,” “always”). In either case, the issue is not obvious as it may differ from fixed survey words. There may be interjections from the patient or agent between questions, or the questions may not be complete, out of order, or significantly separated by time. Tracking scores on orally administered surveys may be something that automation can provide to support the accuracy and efficiency of case managers. In real-time conversational modes, case managers may also be prompted to ask questions that may have been forgotten during the survey.
[0316] The systems and methods described herein can dynamically evaluate casual conversations and automatically input the results into a survey scorecard as the conversation progresses, allowing the person in charge to review the data at a glance.
[0317] In some embodiments, the system may be configured not only to score the survey and provide a final score for the patient, but also to predict the patient's responses to one or more questions. In some embodiments, predicting the patient's responses to individual questions can improve the accuracy of the overall scoring. In some embodiments, measuring individual questions may help identify subtypes of mental health conditions the user may be suffering from.
[0318] In some embodiments, the labels used in training may be based on the weighted contribution of each question to a set of questions. For example, the system may weight each question differently when scoring the survey, depending on how it contributed to the overall score of the survey. As mentioned above, the labels used in training do not have to be based on complete dates (e.g., some questions may be skipped or asked in a misleading way), and therefore, weighting questions in a way different from what would be expected in the raw test may lead to better performance or a more accurate confidence level by the model.
[0319] In some embodiments, responses to different individual questions can be used to assess and / or estimate subtypes of mental health conditions. For example, this can be used to identify subtypes of treatment or to estimate subcategories of severity types. Such analysis may lead to different treatment suggestions for these patients.
[0320] In some embodiments, it may be beneficial to extract survey questions from the operator's utterances and survey responses from the patient's utterances to quickly and efficiently score the patient on various diagnoses or other surveys. In some embodiments, the system may be configured to assess whether there are any unresolved questions that need to be queried, or to prompt the operator to clarify responses if they do not map to a predetermined response. In some embodiments, survey scoring may be performed in real time to provide the operator with progress on the patient's score in the survey.
[0321] In some embodiments, the model can be trained using paraphrased data generated using a Large-Scale Language Model (LLM; e.g., a GPT-type model). In some embodiments, the system utilizes transformers in its architecture.
[0322] Figure 5 shows a method 500 for scoring a survey based on passively listening to a conversation, according to several embodiments.
[0323] In one embodiment, a method is provided for scoring a survey based on conversation. The method includes receiving conversation data from at least one input device (502), processing the conversation data to generate a language model output (504), the language model output including the identification of at least one query based on conversation data from a person in charge and at least one response to at least one query based on conversation data from a patient, wherein at least one query is mapped to a predetermined query and at least one response to at least one query is mapped to a predetermined response to the predetermined query, generating an electronic report (506), and sending the electronic report to an output device (508).
[0324] In some embodiments, the method further includes processing conversational data to generate an acoustic model output, and fusing the language model output and the acoustic model output by applying weights to the language model output and the acoustic model output and generating a composite output from the weighted output, wherein the language model output and the acoustic model output each include multiple outputs corresponding to multiple time segments of the conversational data, and the weights are optionally time-based.
[0325] In some embodiments, the weights are partially based on at least one query or topic within each time segment.
[0326] In some embodiments, the method further includes outputting at least one unresolved query to an output device. An unresolved query may be one for which the user has not provided a response.
[0327] In some embodiments, the method further includes outputting a flag indicating that at least one response to at least one query does not map to a given response to a given query with a confidence level above a threshold.
[0328] In some embodiments, the method further includes prompting the user to repeat a given query if it is determined that at least one response to at least one query does not map to a given response to a given query with a confidence level exceeding a threshold.
[0329] In some embodiments, the electronic report identifies the severity of at least one symptom of the condition based on the language model output.
[0330] In some embodiments, the state includes a mental health state.
[0331] In some embodiments, the electronic report includes annotations of language model output that indicate the prominence of at least one of one or more time segments.
[0332] In some embodiments, the method further includes determining at least one role of at least one speaker, and the weights are based in part on at least one role of at least one speaker in each time segment.
[0333] In some embodiments, processing conversational data to generate language model output involves using a language model neural network trained on labeled conversational data collected from one or more other subjects, where the labeled conversational data is labeled as (i) having a state at some level and (ii) not having that state.
[0334] In some embodiments, conversation data is processed based in part on at least one of a patient profile and a practitioner profile, the profile including at least one of historical data, career data, demographic data, and longitudinal data.
[0335] In some embodiments, the conversation data includes at least one of utterance data and text-based data.
[0336] In some embodiments, method 500 is configured to run a model based on human-interpretable features.
[0337] In one embodiment, a non-temporary computer-readable medium is provided that includes program instructions for causing a computer to perform one of the above methods.
[0338] Speaker Profile In some embodiments, profiles (or other structures for capturing speaker history, demographic, historical, longitudinal data, or other metadata) can be implemented to help the model adapt to those speakers (e.g., patients and / or agents) over time to become more accurate. Such embodiments may be configured to assess a patient in the first session and track changes in the patient over time to assess their improvement or deterioration. Patients will naturally differ from one another in terms of utterance patterns and language patterns. Individuals may exhibit certain characteristics that suggest a mental health condition, but in reality do not possess that mental health condition. For example, some patients may naturally have poor emotional expression and therefore be misclassified as having more severe depression than they actually do. By tracking user differences between sessions, the system can better assess whether a particular patient's mental state is improving or worsening. In this way, profiles can store longitudinal data about speakers. Such information can be stored in the profile. Furthermore, profiles may be able to provide patients, agents, or other parties with speaker trends over time (e.g., whether the patient is improving or worsening). Longitudinal data may be useful in predicting future risks to a user's mental health based on past patterns of symptom severity over time.
[0339] In some embodiments, a patient profile can be constructed from patterns of speech and conversation with the patient. The profile can then be used to predict which therapies and interventions will work best for the patient.
[0340] In some embodiments, the system may predict differences between sessions. For example, the system may expect the patient's condition to improve with each session while the patient is receiving treatment. This allows for longitudinal monitoring of the patient regarding the effectiveness of the treatment (e.g., type of treatment / titration or therapy). Such an approach may be advantageous because it can use relative scores for a particular patient rather than absolute levels to better assess whether the patient should be seen by a healthcare provider or seek other attention. Patients who deviate significantly from these predictions may prompt the system to alert the caregiver that the treatment choice may be ineffective (and may suggest further alternatives), that the patient is not adhering to the treatment regimen (e.g., not taking medication), or something else. As a further example, if a patient initiates a new therapy that is unlikely to be effective in several sessions, the system may predict that the patient's condition will remain relatively stable until the therapy is expected to be effective. In further embodiments, the systems and methods described herein may be used to longitudinally assess the patient's condition and adjust the dosage as needed.
[0341] In some embodiments, the systems and methods described herein may be used to predict future patient conditions, for example, in the case of conditions that may involve relapses (e.g., addiction), manic episodes (e.g., bipolar disorder), the effectiveness of treatment, and the effectiveness of treatment (e.g., for future treatment). Such embodiments may be useful in predicting changes in the patient condition before they occur and potentially suggesting interventions to avoid or mitigate future conditions. For example, in the case of relapse, the system may suggest a higher level of counseling. As a further example, it may suggest that treatment be modified or otherwise suggest resources for the patient. As a final example, in the case of treatments that may last for an unpredictable amount of time (e.g., when the patient is prescribed a drug for the first time and it is unknown how long it will take to metabolize), the system may monitor the patient to see when the effectiveness of the treatment is lost and suggest retreatment. This may be efficient in ensuring that the patient's condition is treated without overtreatment. Furthermore, the system may be configured to monitor treatments that lose their effectiveness with retreatment and adjust the treatment (e.g., dosage) over time.
[0342] Treatment may include any treatment scheme or other compliance framework. For example, in some embodiments, treatment may be an exercise routine, a nutrition plan, cognitive behavioral therapy, or mindfulness exercise. In some embodiments, a treatment plan may be a combination of two or more elements (e.g., medication and a nutrition plan). Any type of treatment or therapy may be compatible with this system.
[0343] In some embodiments, the system may detect when the patient is in a much worse condition than in the previous session and / or much worse than expected. The system may then alert the person in charge. This alert can be given early in the conversation so that the person in charge can steer the conversation towards topics that may be more helpful in eliciting more information or to offer the patient interventions to improve the patient's condition. This alert may also be used so that the person in charge or the system can alert another party, such as the patient's physician or emergency medical professional. The patient may be referred to another party by the system or the person in charge.
[0344] In some embodiments, the system may be configured to track users who have received treatments that wear off. The system may be configured to predict when the effects of the treatment may wear off (and the user may require further treatment sessions). The system may be configured to monitor sessions for clues that the effects of the treatment are wearing off and to further adjust the patient's expected trajectory. As the patient repeats these expected trajectories of treatment, the system may be able to predict and recommend when the patient will next require treatment. Furthermore, by using longitudinal data, the system can assess whether its accuracy for the same patient improves over time or whether the effects of the treatment decline over time.
[0345] In some embodiments, the speaker may first be asked to participate in an action that results in an individualized profile (e.g., calibrating the profile). For example, the system may ask them to respond to a series of prompts to confirm the speaker's baseline. The system may also be configured to update the profile based on the speaker session to further tailor the profile to the user.
[0346] In some embodiments, the patient profile may store data about the patient. For example, the profile may store information about the patient's introversion and extroversion, speaking and / or listening tendencies, agreeableness, etc. Such metrics may further be used by the system to match the patient with an agent the system can predict and to better build a trusting relationship with the patient.
[0347] In some embodiments, the system may be configured to assess a patient's "speaker type." The speaker type may include, for example, an index of cues that may be presented during speech, which could indicate a specific severity of symptoms in a mental health condition. Such types can help the system reduce variability in severity predictions.
[0348] In some embodiments, the system may access metadata information during a session with a patient. For example, the provider may ask the patient questions during the session, and the responses may be incorporated as metadata, or the system may make inferences based on conversational data from the patient (e.g., it may determine that the patient has a smoking history based on the patient's voice signals). In some embodiments, the system may attempt to verify such inferences (e.g., prompt the provider to ask a question, or prompt the user to respond). In some embodiments, conversational history (including patient conversational history, provider conversational history, and patient / provider conversational history) may be used to analyze how to spend time most efficiently in future conversations, triage patients, and adjust patient-provider matches. Current and past conversations can be used to discover subcategories of mental health status for each patient. These subcategories can be used to design treatments more effectively.
[0349] End users include not only the staff but also managers, supervisors, and executives. In some embodiments, staff may also have profiles. These profiles can be used to evaluate the staff's performance over time. These profiles can also be used to evaluate the staff's relationship with a particular patient over time. For example, the system may find that if the staff-patient relationship is a good match, it will follow a certain expected course of trust building. The system may preferentially match staff with patients with whom they are likely to build trust (e.g., patients similar to other patients with whom the staff has successfully built trust in the past).
[0350] Not all staff members can conduct patient interviews equally effectively. For example, staff members may influence a patient's responses with leading comments. For instance, a staff member might ask a patient, "Oh, are you thinking about suicide?" This could lead the patient to answer "no" in line with a clear expectation set by the staff member's wording. Such leading questions can cause certain conditions or symptoms to be over- or under-reported.
[0351] The systems and methods described herein may be useful for evaluating and monitoring the effectiveness of the personnel themselves. For example, a manager of personnel may verify whether personnel are complying with policies, regulations, or standards when conducting surveys or asking required questions. The manager may also evaluate whether some personnel are performing well or whether sessions are being well-labeled.
[0352] The systems and methods described herein may also help improve the clarity of communication, enhance trust between caregivers and patients, and encourage compliance with questionnaires and policies.
[0353] Trust can help attract new patients to a case management company. Trust also encourages patients to provide complete and honest answers, thereby potentially identifying patient problems early and eliminating the need for follow-up visits (and thus saving money). The systems and methods described herein benefit from conversations in which patients share a lot of information, especially information about their condition. Real-time monitoring of trust can help staff determine how much trust has been built and whether it needs to be built further. Trust can be mapped to features that correlate with trust (e.g., short pauses in speech, speech rate, use of interjections like "yes").
[0354] Improving the trust relationship between the patient and the caregiver can make the patient more proactive. The more proactive the patient, the more input the model will have available to accurately predict any mental state the user may be experiencing.
[0355] In some embodiments, the system may evaluate the staff member based on the content and number of mental health-related questions they ask. In some embodiments, the system may evaluate the staff member based on how deeply the patient answers the questions they are asked (patients who are more willing to speak freely may have a better relationship of trust with their staff member). In some embodiments, the system may automatically monitor the staff member's performance. In some embodiments, the system may identify areas for improvement in the staff member (e.g., instructions to avoid interrupting the patient, instructions to recognize the patient, or instructions to ask for clarification).
[0356] In some embodiments, the parallel speaker segmentation model evaluates agent data. For example, acoustic and linguistic data may be used to determine agent performance. In such embodiments, agents may provide voice samples so that the system can more easily recognize the agent (and perhaps as a result, the system can recognize the patient as another speaker through a filtering process). Such embodiments may store agent information in an agent profile, for example, which can track the agent's longitudinal performance and / or other metrics.
[0357] In some embodiments, the agent's performance metrics may be used to supply confidence for a session or a portion of a session. In person-to-device implementations, factors that may influence confidence include, for example, voice quality, ASR, length, and model estimation (all of which may also influence person-to-person implementations). In agent-based implementations, the agent's performance may also influence the confidence of the model's results. Some further factors that may influence the model's confidence may include, for example, topics related to mental health and the length of those topics. The agent's performance may influence the confidence of results for the entire session (e.g., their coverage) and / or for a specific patient (e.g., based on trust). The system can use this performance to verify the reliability of any labels output.
[0358] In some embodiments, the agent may be scored for their ability to accurately label patients based on a standardized test (e.g., PHQ-9, PHQ-2, GAD-7). The system may be configured to semantically analyze a session and score that session against a set of questions of interest. The system may be configured to detect questions in a session that closely correspond to questions in the set. The closer the agent asks the questions to the format provided in the set, the higher the agent's score will be. This can be used to assess the reliability of the agent's scoring of patients based on the standardized test, or the reliability of the system's own score for patients (because the system may not be able to properly score patients who have not been asked or have not been asked all of the questions in the set).
[0359] Exemplary aspects According to embodiments of this specification, a system 100 for identifying the roles of speakers in a conversation is provided. The system comprises at least one input device 102 for receiving conversational data from at least one user, at least one output device 126 for outputting an electronic report, and at least one computing device 104 for communicating with the at least one input device 102 and the at least one output device 126. At least one computing device 104 is configured to receive conversation data from at least one input device 102, determine at least one role of at least one speaker using a role detector 114, process the conversation data to generate language model outputs and / or acoustic model outputs using a language processor 116 and an acoustic processor 118, respectively, apply weights, optionally time-based and partially based on at least one role of at least one speaker in each time segment, to the language model outputs and / or acoustic model outputs, each including multiple outputs corresponding to multiple time segments of conversation data, generate an electronic report, and transmit the electronic report to an output device 126.
[0360] In some embodiments, at least one computing device is further configured to use a fusion model 120 to fuse weighted language model outputs and acoustic model outputs to produce a composite output. The composite output may represent a fusion output from the fusion of the language model output and the acoustic model output.
[0361] In some embodiments, the electronic report may identify at least one symptom of a condition based on the combined output.
[0362] In some embodiments, the state may include a mental health state.
[0363] In some embodiments, the electronic report includes annotations of language model output that indicate the prominence of at least one of one or more time segments.
[0364] In some embodiments, at least one speaker includes a representative, and applying weights to the language model output and acoustic model output includes applying a zero weight to the acoustic model output corresponding to the representative. In such embodiments, the language model may use only the patient's language information, or the language information of both the patient and the representative.
[0365] In some embodiments, the weights are partially based on the topics in each time segment.
[0366] In some embodiments, the language model output includes the identification of at least one query based on conversational data from the person in charge, and at least one response to at least one query based on conversational data from the patient, wherein at least one query is mapped to a predetermined query, and at least one response to at least one query is mapped to a predetermined response to the predetermined query.
[0367] In some embodiments, processing conversational data to generate language model outputs and acoustic model outputs involves using language model neural networks and acoustic neural networks trained on labeled conversational data collected from one or more other subjects, wherein the labeled conversational data is labeled as (i) having a state at some level and (ii) not having that state.
[0368] In some embodiments, at least one role of at least one speaker includes at least one of patient, agent, interactive voice response, and bot speaker.
[0369] In some embodiments, the weights applied to the language model output and the acoustic model output are partially based on determining whether the number of speakers matches the expected number of speakers.
[0370] In some embodiments, conversation data is processed based in part on at least one of a patient profile and a practitioner profile, the profile including at least one of historical data, career data, demographic data, and longitudinal data.
[0371] In some embodiments, the conversation data includes at least one of utterance data and text-based data.
[0372] In some embodiments, at least one computing device 104 is configured to run a model based on human-interpretable features.
[0373] In one embodiment, a system 100 is provided for identifying topics in a conversation. The system 100 includes at least one input device 102 for receiving conversation data from at least one user, at least one output device 126 for outputting an electronic report, and at least one computing device 104 for communicating with the at least one input device 102 and the at least one output device 126. The at least one computing device 94 is configured to receive conversation data from the at least one input device 102, process the conversation data, and use a language processor 116 to generate a language model output, the language model output including one or more topics corresponding to one or more time ranges, and to generate weighted outputs, the outputs including multiple outputs corresponding to multiple time segments of conversation data, by applying weights, optionally time-based and partially based on one or more topics in each time segment, to the outputs, generate an electronic report, and transmit the electronic report to the output device 126.
[0374] In some embodiments, at least one computing device 94 is configured to use an acoustic processor 118 to process conversational data to generate an acoustic model output, apply weights to the language model output and the acoustic model output, and generate a composite output from the weighted output, thereby fusing the language model output and the acoustic model output using a fusion model 120, wherein the acoustic model output includes multiple outputs corresponding to multiple time segments of conversational data, and the weights are optionally time-based and partially based on the topic in each time segment.
[0375] In some embodiments, the time range of conversational data corresponding to a given topic is processed using a computationally more robust model than the one used for the time range of conversational data not corresponding to the given topic, in order to generate language model output.
[0376] In some embodiments, the electronic report includes a transcript of language model output annotated with weights partially based on the topics in each time segment.
[0377] In some embodiments, the electronic report identifies the severity of at least one symptom of the condition based on the language model output.
[0378] In some embodiments, the state includes a mental health state.
[0379] In some embodiments, the electronic report includes annotations of language model output that indicate the prominence of at least one of one or more time segments.
[0380] In some embodiments, the computing device is further configured to determine at least one role of at least one speaker, and the weights are partially based on at least one role of at least one speaker during each time segment.
[0381] In some embodiments, the language model output includes the identification of at least one query based on conversational data from the person in charge, and at least one response to at least one query based on conversational data from the patient, wherein at least one query is mapped to a predetermined query, and at least one response to at least one query is mapped to a predetermined response to the predetermined query.
[0382] In some embodiments, processing conversational data to generate language model output involves using a language model neural network trained on labeled conversational data collected from one or more other subjects, where the labeled conversational data is labeled as (i) having a state at some level and (ii) not having that state.
[0383] In some embodiments, at least one query belongs to a set of queries, and the computing device is configured to predict an overall score based on the set of queries based on the response to each of the queries in the set.
[0384] In some embodiments, conversation data is processed based in part on at least one of a patient profile and a practitioner profile, the profile including at least one of historical data, career data, demographic data, and longitudinal data.
[0385] In some embodiments, the conversation data includes at least one of utterance data and text-based data.
[0386] In some embodiments, at least one computing device 104 is configured to run a model based on human-interpretable features.
[0387] In one embodiment, a system 100 is provided for scoring a survey based on conversations. The system 100 includes at least one input device 102 for receiving conversation data from at least one user, at least one output device 126 for outputting an electronic report, and at least one computing device 104 for communicating with the at least one input device 102 and the at least one output device 126. At least one computing device 104, configured to receive conversation data from at least one input device 102, process the conversation data using a language processor 116, generate a language model output which includes the identification of at least one query based on conversation data from a person in charge, at least one query based on conversation data from a patient, at least one response to at least one query which is mapped to a predetermined query, and at least one response to at least one query which is mapped to a predetermined response to a predetermined query, generate an electronic report, and send the electronic report to an output device 126.
[0388] In some embodiments, at least one computing device 104 is further configured to use an acoustic processor 118 to process conversational data to generate an acoustic model output, apply weights to the language model output and the acoustic model output, and generate a composite output from the weighted output, thereby fusing the language model output and the acoustic model output using a fusion model 120, wherein the language model output and the acoustic model output each include multiple outputs corresponding to multiple time segments of conversational data, and the weights are optionally time-based.
[0389] In some embodiments, the weights are partially based on at least one query or topic within each time segment.
[0390] In some embodiments, the system is configured to output at least one unresolved query to an output device. An unresolved query may be one for which the user has not provided a response.
[0391] In some embodiments, the system is configured to output a flag indicating that at least one response to at least one query does not map to a given response to a given query with a confidence level above a threshold.
[0392] In some embodiments, the system is configured to prompt the system to repeat a given query if it is determined with a confidence level above a threshold that at least one response to at least one query does not map to a given response to a given query.
[0393] In some embodiments, the electronic report identifies the severity of at least one symptom of the condition based on the language model output.
[0394] In some embodiments, the state includes a mental health state.
[0395] In some embodiments, the electronic report includes annotations of language model output that indicate the prominence of at least one of one or more time segments.
[0396] In some embodiments, the computing device 104 is further configured to determine at least one role of at least one speaker, and the weights are partially based on at least one role of at least one speaker during each time segment.
[0397] In some embodiments, processing conversational data to generate language model output involves using a language model neural network trained on labeled conversational data collected from one or more other subjects, where the labeled conversational data is labeled as (i) having a state at some level and (ii) not having that state.
[0398] In some embodiments, conversation data is processed based in part on at least one of a patient profile and a practitioner profile, the profile including at least one of historical data, career data, demographic data, and longitudinal data.
[0399] In some embodiments, the conversation data includes at least one of utterance data and text-based data.
[0400] In some embodiments, at least one computing device 104 is configured to run a model based on human-interpretable features.
[0401] In one embodiment, a non-temporary computer-readable medium is provided that includes program instructions for causing a computer to perform one of the above methods.
[0402] training The model can be trained using deep learning. Its performance can be validated with labeled data from unseen speakers. The language processor, acoustic processor, and fusion model can each be trained separately, together, or initially separately and then fine-tuned together.
[0403] The model can be trained using large amounts of labeled data about patients' mental health. In these datasets, the handling of areas that may contain strong clues to responses depends on the operational data in question. If the data is not expected to include, for example, survey areas (if any), these areas can be removed from the training data. Therefore, when standard health surveys are not conducted, the model can be encouraged to learn about cues for mental health, making the software more universally applicable and reducing its reliance on survey implementation.
[0404] In some embodiments, the model may be trained before or after the removal of any oral surveys. If the model is trained before the removal of oral surveys, it may not perform well if the operational data does not include surveys. If the operational data includes oral surveys, it may be best to train the model with the survey areas included in the training data. Otherwise, the surveys themselves may make it too easy for the model to predict mental health states, and the goal is to perform without requiring the presence of surveys.
[0405] As part of training in implementing survey scoring, it may be necessary to provide the model with multiple ways of representing different questions and different answers.
[0406] The training may utilize further methods to detect and remove portions of training conversations related to any survey. These survey portions can then be removed from the conversations to train a model used in the system to identify and evaluate patients based on non-survey responses.
[0407] Training models for older populations may need to overcome several unique challenges. Since older patients tend to have lower technological literacy, it may be important to make the device as easy to use as possible. Older populations tend to speak more slowly and softer, which may require special training for certain modes. Furthermore, older populations may have more health issues that could further impact the model's speech recognition.
[0408] Confidence of training data Training data from person-to-person care sessions are not always reliably labeled. Generally, for training purposes, the final scores of standardized tests entered by the caregiver may be used as part or all of the labels for the training data. Therefore, it is crucial that such standardized tests are administered completely and accurately in order to generate meaningful labels. In some situations, caregivers may ask questions from standardized tests in an inappropriate order, omit all questions, or fail to properly score the patient (for example, asking a question as a binary choice and then attempting to assess the patient on a graded scale).
[0409] In some embodiments, the system may be able to assess the reliability of labels in the training data and, for example, weight the data accordingly. For example, the system may search the training data semantically for questions similar to or identical to questions from a set of interest (e.g., from a standardized test). A session that has questions closely corresponding to all questions in the set may be more reliable than a session that is missing some questions, or a session where the closest question does not closely correspond to a question from the set (or both). Furthermore, the more questions there are, and the more the agent and patient discuss them, the more reliable the data becomes. Such confidence levels can be used to filter out bad / good data or to weight the data during training. Confidence levels can be used to filter out good data from bad data during training. For example, each session may be scored for reliability based on the semantic closeness of the questions asked in the session to the questions in the set. Confidence scores may be used to weight the training data and / or to include / exclude the training data.
[0410] Furthermore, such considerations can be equally applied during use. For example, an employee may evaluate the reliability of a label generated by the system while they are asking questions. Such a system may then prompt the employee to ask further questions or to do so appropriately.
[0411] Data expansion during training In some embodiments, training data may be augmented to enhance pre-operational model training. In some embodiments, linguistic and acoustic data may be augmented. Data augmentation may be particularly advantageous for providing training data for demographics where available training data (e.g., data from speakers of a particular dialect or accent, or data from patients with a particular speaking style) is relatively limited. Data augmentation may also be useful in adapting a model to a particular population by familiarizing it with specific types of words and phrases, as well as the speaking style or voice type expected from the target population.
[0412] In some embodiments, language data can be augmented through the use of large-scale language models. Such models may be used to generate different ways in which a patient or agent might say the same thing (e.g., paraphrasing).
[0413] In some embodiments, acoustic data can be augmented through the use of synthesized speech. Synthesized speech (or other methods) can be used to generate different voices that say the same or different content. Furthermore, conversational data from an unmodified session can be modified to capture a specific speaking style.
[0414] Zero-shot learning Zero-shot learning can describe a learning scenario in which a model is asked to classify objects from classes that were not observed during training. In some embodiments, a zero-shot learning model may associate auxiliary information with observed and unobserved classes so that the model can distinguish between classes based on several distinguishing features of the objects.
[0415] In some embodiments, the model may be trained using some form of zero-shot learning (e.g., if one or more classes are not observed during training). Such embodiments may be advantageous because they can eliminate the need for at least some labeled training data.
[0416] In some embodiments, the systems and methods described herein may use a language model (e.g., a large-scale language model) as a form of zero-shot model. A large-scale language model is a model (e.g., pre-trained, self-supervised, semi-supervised) trained to predict the next token or word based on input text. In some embodiments, such a large-scale language model may be used to predict the mental health assessment and / or symptom severity of a subject.
[0417] In some embodiments, in-context learning (e.g., prompt engineering) may be used to instruct a model to predict behavioral or mental health conditions based on zero-shot or future-shot learning. For example, a large-scale language model may be provided with a description of depression or other conditions, from which it may be possible to predict the severity (and severity level) of the subject's depression, for example, based on a description and transcription of the subject's conversations with an agent.
[0418] In some embodiments, zero-shot learning may be used in conjunction with, for example, a language model. In some embodiments, zero-shot learning may be used directly or indirectly. In some embodiments, a large language model may be asked about the severity level of a subject (e.g., PHQ risk level or other mental health severity level) based on a transcript of a conversation. In some embodiments, the large language model may be asked to estimate questionnaires (e.g., PHQ or other mental health questionnaires) for the answers to individual questions based on a transcript of a conversation. In such embodiments, individual answers may have confidence levels associated with them. The answers may be aggregated into a final score. As a further example, individual answers may further help in subtyping the classes of conditions. In some embodiments, any of the above outputs may be further used as input to further models, which may not be zero-shot learning models (e.g., non-zero-shot fusion models, language models, acoustic models, etc.).
[0419] In some embodiments, the output from the zero-shot model can be fused with other models (e.g., acoustic models).
[0420] In some embodiments, zero-shot learning models may be used to preprocess data (e.g., indirectly). For example, a large-scale language model can be used to extract and summarize topics and characteristics of interest, including confidence levels (e.g., how much evidence is found for each individual aspect). Such data can then be used as input features for further models (e.g., for training and / or inference). This can be used alone or in conjunction with existing transcripts. In some embodiments, a transcript of a conversation can be preprocessed with a large-scale language model to identify content regions related to a patient's mental health (and potentially weighted according to how relevant they are). This can then be used to emphasize the conversation regions during inference or training (e.g., by applying higher weights). In some embodiments, a transcript of a conversation can be preprocessed with a large-scale language model to identify content regions not related to a patient's mental health (and potentially weighted according to how irrelevant they are). This can then be used to avoid emphasizing the conversation regions during inference or training (e.g., by applying lower weights). In some embodiments, a large-scale language model can be used to identify topics in a transcript, provide analysis, and / or summarize them. In some embodiments, large-scale language models can be used to identify, analyze, and / or summarize certain aspects of a questionnaire (e.g., PHQ or GAD). For example, it may be possible to provide individual responses to questions such as: "Do you have little interest or pleasure in doing things?", "Do you feel depressed, depressed or hopeless?", "Do you have trouble falling asleep, wake up in the middle of the night, or oversleep?", "Do you feel tired or lacking energy?", and "Do you have a poor appetite or overeat?". In some embodiments, large-scale language models can be used to identify, analyze, and / or summarize behavioral aspects (e.g., stress, life satisfaction, or wellness).
[0421] In some embodiments, preprocessed conversational data may be fed into further models. For example, weighting by a zero-shot model may be used by a language model or an acoustic model to analyze the conversational data. As a further example, topic summaries may be analyzed by a language model to predict the severity of symptoms of behavioral or mental health conditions. The results of the preprocessing may also be provided to the end user (e.g., the person in charge), for example, to support explainability.
[0422] Optimization based on metrics and distributions Different mental health applications may emphasize different metrics (e.g., binary screening, class-based, regression-based, correlation coefficient matching). Optimal performance for one metric does not guarantee optimal performance for another. Data distribution can also have different effects on metrics.
[0423] In some embodiments, a metric of interest and an expected approximate target distribution may be selected, and the model and / or fusion model may be optimized with respect to them. The fusion weighting can be trained to optimize for the metric of interest. Optimizing the fusion model, as well as the language and acoustic models, may improve the performance of the models.
[0424] Exemplary training implementation In one embodiment, a system 100 is provided that trains one or more baselines of language models, acoustic models, and fusion models (models) to directly or indirectly detect behavioral or mental health states using machine learning. Training includes using the models to predict behavioral or mental health states in training data and updating the models based on the accuracy of the predictions.
[0425] In some embodiments, the labels in the training data are evaluated for reliability before or during training, and each training data is reweighted according to the reliability of its label.
[0426] In some embodiments, the training data is augmented using at least one of paraphrasing or synthesized speech to generate additional training data for use with the training data.
[0427] In some embodiments, training includes predicting speaker IDs and penalizing the model for accurately identifying speaker IDs. In some embodiments, the model is penalized for learning information about speaker IDs rather than information about the speaker's state.
[0428] In one embodiment, a system 100 is provided for predicting the severity of at least one symptom of a subject's behavioral or mental health condition. The system 100 comprises at least one input device 102 for receiving conversational data from the subject, and at least one computing device 104 for communicating with the at least one input device 102. The at least one computing device 104 is configured to receive in-context learning, including explanations related to one or more questions on a questionnaire, receive conversational data from the at least one input device, and predict the severity of at least one symptom of the subject's behavioral or mental health condition based on the in-context learning and the conversational data.
[0429] In some embodiments, the computing device 104 accesses a large-scale language model to predict the severity of at least one symptom of a subject's behavioral or mental health condition based on in-context learning and conversational data.
[0430] In some embodiments, predicting the severity of at least one symptom of a behavioral or mental health condition includes predicting the results of a questionnaire.
[0431] In some embodiments, the prediction of the severity of at least one symptom of a behavioral or mental health condition includes the prediction of the outcome of at least one of one or more questions on a questionnaire.
[0432] compatibility Optionally, the devices, systems, and methods described herein may be used to implement embodiments of the devices, systems, and methods described in U.S. Patent No. 1,0748644, filed September 4, 2019, entitled “Systems and Methods for Mental Health Assessment” (which is incorporated herein by reference in its entirety). Accordingly, the devices, systems, or methods described herein may be interoperable with methods for determining whether a subject is at risk of having a mental or physiological condition, which include acquiring data from such subject, including conversational data and optionally associated visual data, and processing such data using a plurality of machine learning models, including natural language processing (NLP) models and acoustic models, to generate NLP and acoustic outputs, wherein the plurality of machine learning models include neural networks trained on labeled conversational data collected from one or more other subjects, and the labeled conversational data from each of the one or more other subjects indicates that (i) the subject has the mental or physiological condition at some level, or (ii) ) generating, labeled as not having the mental or physiological condition; merging the NLP output and the acoustic output by (1) applying weights to the NLP output and the acoustic output to generate a weighted output; and (2) generating a composite output from the weighted output, wherein the NLP output and the acoustic output each include multiple outputs corresponding to multiple time segments of the conversation data, and the weights in (1) are time-based; and outputting an electronic report, based at least on the composite output, identifying whether the subject is at risk of having the mental or physiological condition, wherein the risk is quantified in the form of a score having a confidence level provided in the report.
[0433] Optionally, the devices, systems, and methods described herein may be used to implement embodiments of the devices, systems, and methods described in U.S. Patent Application No. 17 / 493687, filed October 4, 2021, entitled “Reliability Assessment for Measuring the Reliability of Behavioral Health Survey Results” (which is incorporated herein by reference in its entirety). Accordingly, the devices, systems, or methods described herein may be interoperable with methods for measuring the reliability of responses received from human subjects in health surveys for assessing the health status of subjects, the methods comprising: acquiring response data generated by a subject in response to prompts presented to the subject during the conduct of a health survey of the subject, wherein the response data includes a plurality of conditioned events and a plurality of conditioned events; and determining a first probability that a first conditioned event is present in the response data, based in part on the presence of a first conditioned event in the response data, wherein the plurality of conditioned events include the first conditioned event and a plurality The process includes determining a conditioned event which includes a first conditioned event and a first probability which is used to determine a first pair of events which includes the first conditioned event and the first conditioned event; repeating the first two steps for two or more other conditioned events and other conditioned events to generate multiple additional probabilities for multiple additional pair of events; and generating confidence vector data which combines (i) the first probability and the multiple additional probabilities or (ii) two or more probabilities selected from the multiple additional probabilities, the confidence vector data which generates response data in response to a health survey and represents a measure of confidence in the reliability of the subject.
[0434] Optionally, the devices, systems, and methods described herein may be used to implement embodiments of the devices, systems, and methods described in U.S. Patent Application No. 17 / 726999, filed April 22, 2022, entitled “Acoustic and Natural Language Processing Models for Voice-Based Behavioral Health Screening and Monitoring,” which is incorporated herein by reference in its entirety. Accordingly, the devices, systems, or methods described herein may be interoperable with methods for detecting behavioral or mental health conditions in a subject, the methods comprising: obtaining an utterance sample from the subject comprising one or more utterance segments; and performing at least one of (i) or (ii), wherein (i) includes processing the utterance sample with at least one acoustic model comprising an encoder to produce an acoustic model output comprising an abstract feature representation of the utterance sample, the encoder being pre-trained to perform a first task other than detecting the behavioral or mental health condition in the subject; and (ii) includes processing the utterance sample, its derivatives, and / or the utterance sample transcribed into a text sequence with at least one natural language processing (NLP) model to produce a language model output; and individually or jointly generating an output indicating whether the subject has the behavioral or mental health condition using at least one of the acoustic model output or the language model output.
[0435] Optionally, the devices, systems, and methods described herein may be used to implement embodiments of the devices, systems, and methods described in PCT Patent Application No. PCT / US2022 / 015147, filed on 3 February 2022, entitled “System and Method for Multilingual Adaptive Mental Health Risk Assessment from Spoken and Written Languages” (which is incorporated herein by reference in its entirety). Accordingly, the devices, systems, or methods described herein may be interoperable with methods for detecting behavioral or mental health conditions, the methods comprising: receiving an input signal comprising a plurality of speech or lexical characteristics of an utterance in question, wherein at least one of the plurality of speech or lexical characteristics of the utterance is associated with at least one language based at least in part on the plurality of speech or lexical characteristics of the input signal; selecting one or more speech or natural language processing (NLP) models, wherein at least one of the speech or NLP models is a multilingual or language-independent model; and processing the input signal with a fused or combined model derived from one or more speech or NLP models to detect results indicating the presence or absence of a behavioral or mental health condition.
[0436] Details of the implementation The above description provides many exemplary embodiments. Each embodiment represents a single combination of the elements of the present invention, but other examples may include all possible combinations of the disclosed elements. Thus, if one embodiment includes elements A, B, and C, and a second embodiment includes elements B and D, other remaining combinations of A, B, C, or D may also be used.
[0437] The terms "connected" or "joined" can include both direct joining (two elements joined together are in contact with each other) and indirect joining (at least one additional element is located between the two elements).
[0438] Throughout the preceding description, numerous references have been made to servers, services, interfaces, portals, platforms, or other systems formed from computing devices. It should be understood that the use of such terms is considered to represent one or more computing devices having at least one processor configured to execute software instructions stored on a computer-readable, tangible, non-temporary medium. For example, a server may include one or more computers operating as a web server, database server, or other type of computer server in a manner that fulfills the described roles, responsibilities, or functions.
[0439] Embodiments of devices, systems, and methods described herein may be implemented in combination of both hardware and software. These embodiments may be implemented on a programmable computer, each computer comprising at least one processor, a data storage system (including volatile memory or non-volatile memory, or other data storage elements, or a combination thereof), and at least one communication interface. Embodiments of devices, systems, and methods described herein may be implemented, for example, using cloud computing, services, and / or edge computing.
[0440] Program code is applied to input data to perform the functions described herein and generate output information. The output information is applied to one or more output devices. In some embodiments, the communication interface may be a network communication interface. In embodiments, the communication interface may be a software communication interface, such as one for inter-process communication, where elements can be combined. In yet other embodiments, there may be a combination of communication interfaces implemented as hardware, software, and a combination thereof.
[0441] The technical solution of the embodiment may take the form of a software product. The software product may be stored on a non-volatile or non-temporary storage medium, which may be a compact disc read-only memory (CD-ROM), a USB flash disk, or a removable hard disk. The software product includes several instructions that enable a computer device (personal computer, server, or network device) to perform the method provided by the embodiment.
[0442] The embodiments described herein are implemented by physical computer hardware, including computing devices, servers, receivers, transmitters, processors, memory, displays, and networks. The embodiments described herein provide useful physical machines and particularly configured computer hardware configurations. The embodiments described herein cover electromachines adapted for processing and converting electromagnetic signals representing various types of information, as well as methods implemented by electromachines. The embodiments described herein are broad and integrated in nature relating to machines and their uses, and the embodiments described herein have no meaning or practical applicability other than their use with computer hardware, machines, and various hardware components. For example, replacing physical hardware particularly configured to implement various operations with non-physical hardware, for example, by using mental steps, may substantially affect how the embodiments function. Such limitations of computer hardware are clearly essential elements of the embodiments described herein and cannot be omitted or replaced by mental means without materially affecting the operation and structure of the embodiments described herein. Computer hardware is essential for implementing the various embodiments described herein and is not merely used to perform steps quickly and efficiently.
[0443] Figure 10 shows a schematic diagram of a computing device 1000 according to several embodiments. As depicted, the computing device 1000 includes at least one processor 1002, memory 1004, at least one I / O interface 1006, and at least one network interface 1008. The computing device 1000 may be implemented in system 100 as computing device 104.
[0444] For simplicity, only one computing device 1000 is shown, but the system may include more computing devices 1000 that can be operated by the user to access remote network resources and exchange data. The computing devices 1000 may be the same or different types of devices. The computing device 1000 comprises at least one processor, a data storage device (including volatile memory or non-volatile memory, or other data storage elements, or a combination thereof), and at least one communication interface. The computing device components may be connected in a variety of ways, including being directly coupled, indirectly coupled via a network, distributed across a wide geographical area, and connected via a network (which may be referred to as “cloud computing”).
[0445] For example, but not limited to, computing devices may include servers, network appliances, set-top boxes, embedded devices, computer expansion modules, personal computers, laptops, personal data assistants, mobile phones, smartphone devices, UMPC tablets, video display terminals, game consoles, e-reading devices, and wireless hypermedia devices, or any other computing devices that can be configured to perform the methods described herein.
[0446] Each processor 1002 could be, for example, any kind of general-purpose microprocessor or microcontroller, a digital signal processing (DSP) processor, an integrated circuit, a field-programmable gate array (FPGA), a reconfigurable processor, a programmable read-only memory (PROM), or any combination thereof.
[0447] The memory 1004 may include, for example, any combination of computer memories located either internally or externally, such as random access memory (RAM), read-only memory (ROM), compact disk read-only memory (CDROM), electro-optical memory, magneto-optical memory, erasable programmable read-only memory (EPROM), and electro-erasable programmable read-only memory (EEPROM), and ferroelectric RAM (FRAM).
[0448] Each I / O interface 1006 enables the computing device 1000 to interconnect with one or more input devices such as a keyboard, mouse, camera, touchscreen, and microphone, or with one or more output devices such as a display screen and speakers.
[0449] Each network interface 1008 enables the computing device 1000 to communicate with other components, exchange data with other components, access and connect to network resources, provide applications, and run other computing applications by connecting to networks (or multiple networks) capable of carrying data, including the Internet, Ethernet, Plain Old Telephone Service (POTS) lines, Public Switched Telephone Networks (PSTN), Integrated Services Digital Network (ISDN), Digital Subscriber Line (DSL), Coaxial Cable, Optical Fiber, Satellite, Mobile, Wireless (e.g., Wi-Fi, WiMAX), SS7 Signaling Communications Networks, Fixed Lines, Local Area Networks, Wide Area Networks, and others, including any combination thereof.
[0450] The computing device 1000 may operate to register and authenticate users (for example, using logins, unique identifiers, and passwords) before providing access to applications, local networks, network resources, other networks, and network security devices. The computing device 1000 may serve one or more users.
[0451] Although this embodiment has been described in detail, it should be understood that various changes, substitutions, and modifications can be made herein without departing from the scope defined by the appended claims.
[0452] Furthermore, the scope of this application is not intended to be limited to specific embodiments of the processes, machines, manufactures, material compositions, means, methods, and steps described herein. As those skilled in the art will readily understand from the disclosure of the present invention, existing or subsequently developed materials, means, methods, or steps, processes, machines, manufactures, and compositions, which perform substantially the same functions or achieve substantially the same results as the corresponding embodiments described herein, may be utilized. Accordingly, the appended claims are intended to include within their scope such processes, machines, manufactures, material compositions, means, methods, or steps.
[0453] For the purposes of this invention, the embodiments described and illustrated above are intended to be illustrative only. The scope is defined by the attached claims.
Claims
1. A system for identifying the role of a speaker in a conversation, wherein the system At least one input device for receiving conversation data from at least one user, At least one output device for outputting electronic reports, A computing device that communicates with the at least one input device and the at least one output device, The conversation data is received from at least one input device. Determine at least one role of at least one speaker, The aforementioned conversation data is processed to generate language model output and / or acoustic model output, The language model output and / or the acoustic model output, each including a plurality of outputs corresponding to a plurality of time segments of the conversation data, is to be weighted, optionally time-based, and partially based on the at least one role of the at least one speaker in each time segment. Generate an electronic report, A system comprising at least one computing device configured to transmit the electronic report to the output device.
2. The at least one computing device, The system according to claim 1, further configured to fuse the weighted language model output and the acoustic model output to generate a composite output.
3. The system according to claim 1, wherein the electronic report identifies the severity of at least one symptom of a condition based on the composite output.
4. The system according to claim 3, wherein the aforementioned state includes a mental health state.
5. The system according to claim 1, wherein the electronic report includes annotations of the language model output indicating at least one of the time segments.
6. The system according to claim 1, wherein the at least one speaker includes a person in charge, and applying the weights to the language model output and the acoustic model output includes applying a zero weight to the acoustic model output corresponding to the person in charge.
7. The system according to claim 1, wherein the weights are partially based on the topics in each time segment.
8. The system according to claim 1, wherein the language model output includes the identification of at least one query based on conversation data from a person in charge, and at least one response to the at least one query based on conversation data from a patient, wherein the at least one query is mapped to a predetermined query, and the at least one response to the at least one query is mapped to a predetermined response to the predetermined query.
9. The system according to claim 1, wherein processing the conversation data to generate the language model output and the acoustic model output includes using a language model neural network and an acoustic neural network trained on labeled conversation data collected from one or more other subjects, wherein the labeled conversation data is labeled as (i) having a state at a certain level and (ii) not having the state.
10. The system according to claim 1, wherein the at least one role of the at least one speaker includes at least one of patient, staff member, interactive voice response, and bot speaker.
11. The system according to claim 1, wherein the weights applied to the language model output and the acoustic model output are in part based on a determination that the number of at least one speaker matches the expected number of speakers.
12. The system according to claim 1, wherein the conversation data is processed in part based on at least one of a patient profile and a caregiver profile, and the profile includes at least one of historical data, career data, demographic data, and longitudinal data.
13. The system according to claim 1, wherein the conversation data includes at least one of utterance data and text-based data.
14. The system according to claim 1, wherein the at least one computing device is configured to run a model based on human-interpretable features.
15. A system for identifying the topic of conversation, wherein the system At least one input device for receiving conversation data from at least one user, At least one output device for outputting electronic reports, A computing device that communicates with the at least one input device and the at least one output device, The conversation data is received from at least one input device. The aforementioned conversation data is processed to generate a language model output which includes one or more topics corresponding to one or more time ranges, Outputs, which include multiple outputs corresponding to multiple time segments of the conversation data, are weighted by applying weights, optionally time-based and partially based on one or more topics in each time segment, to the outputs to generate weighted outputs. Generate an electronic report, A system comprising at least one computing device configured to transmit the electronic report to the output device.
16. The at least one computing device, The aforementioned conversation data is processed to generate an acoustic model output, The system according to claim 15, configured to merge the language model output and the acoustic model output by applying weights to the language model output and the acoustic model output and generating a composite output from the weighted output, wherein the acoustic model output includes a plurality of outputs corresponding to a plurality of time segments of the conversation data, and the weights are optionally time-based and partially based on the topic in each time segment.
17. The system according to claim 15, wherein the time range of the conversation data corresponding to a predetermined topic is processed to generate the language model output using a model that is more computationally robust than the model used for the time range of the conversation data not corresponding to the predetermined topic.
18. The system according to claim 15, wherein the electronic report includes a transcription of the language model output annotated partly on the topic in each time segment and partly on the weights.
19. The system according to claim 15, wherein the electronic report identifies the severity of at least one symptom of a condition based on the language model output.
20. The system according to claim 19, wherein the aforementioned state includes a mental health state.
21. The system according to claim 15, wherein the electronic report includes annotations of the language model output indicating at least one of the time segments.
22. The system according to claim 15, wherein the computing device is further configured to determine at least one role of at least one speaker, and the weights are partially based on the at least one role of the at least one speaker in each time segment.
23. The system according to claim 15, wherein the language model output includes the identification of at least one query based on conversation data from a person in charge, and at least one response to the at least one query based on conversation data from a patient, wherein the at least one query is mapped to a predetermined query, and the at least one response to the at least one query is mapped to a predetermined response to the predetermined query.
24. The system according to claim 15, wherein processing the conversation data to generate the language model output includes using a language model neural network trained on labeled conversation data collected from one or more other subjects, wherein the labeled conversation data is labeled as (i) having a state at a certain level and (ii) not having the state.
25. The aforementioned at least one query belongs to the set of queries, The system according to claim 15, wherein the computing device is configured to predict an overall score based on the set of queries based on the response to each of the queries in the set of queries.
26. The system according to claim 15, wherein the conversation data is processed in part based on at least one of a patient profile and a caretaker profile, and the profile includes at least one of historical data, career data, demographic data, and longitudinal data.
27. The system according to claim 15, wherein the conversation data includes at least one of utterance data and text-based data.
28. The system according to claim 15, wherein the at least one computing device is configured to run a model based on human-interpretable features.
30. A system for scoring surveys based on conversations, At least one input device for receiving conversation data from at least one user, At least one output device for outputting electronic reports, A computing device that communicates with the at least one input device and the at least one output device, The conversation data is received from at least one input device. The conversation data is processed to generate a language model output which includes the identification of at least one query based on the conversation data from the person in charge, and the at least one query based on the conversation data from the patient, wherein the at least one query is mapped to a predetermined query, and the at least one response to the at least one query is mapped to a predetermined response to the predetermined query. Generate an electronic report, A system comprising at least one computing device configured to transmit the electronic report to the output device.
31. The at least one computing device, The aforementioned conversation data is processed to generate an acoustic model output, The system according to claim 30, further configured to merge the language model output and the acoustic model output by applying weights to the language model output and the acoustic model output and generating a composite output from the weighted outputs, wherein the language model output and the acoustic model output each include a plurality of outputs corresponding to a plurality of time segments of the conversation data, and the weights are optionally time-based.
32. The system according to claim 31, wherein the weights are partially based on the at least one query or topic in each time segment.
33. The aforementioned system The system according to claim 30, configured to output at least one unresolved query to the output device.
34. The aforementioned system The system according to claim 30, wherein the system is configured to output a flag indicating that the at least one response to the at least one query does not map to a predetermined response to the predetermined query with a confidence level above a threshold.
35. The aforementioned system The system according to claim 30, wherein if it is determined with confidence that the at least one response to the at least one query is not mapped to a predetermined response to the predetermined query, the system is configured to output a prompt to repeat the predetermined query.
36. The system according to claim 30, wherein the electronic report identifies the severity of at least one symptom of a condition based on the language model output.
37. The system according to claim 36, wherein the aforementioned state includes a mental health state.
38. The system according to claim 30, wherein the electronic report includes annotations of the language model output indicating at least one of the time segments.
39. The system according to claim 30, wherein the computing device is further configured to determine at least one role of at least one speaker, and the weights are partially based on the at least one role of the at least one speaker in each time segment.
40. The system according to claim 30, wherein processing the conversation data to generate the language model output includes using a language model neural network trained on labeled conversation data collected from one or more other subjects, wherein the labeled conversation data is labeled as (i) having a state at a certain level and (ii) not having the state.
41. The system according to claim 30, wherein the conversation data is processed in part based on at least one of a patient profile and a caretaker profile, and the profile includes at least one of historical data, career data, demographic data, and longitudinal data.
42. The system according to claim 30, wherein the conversation data includes text-based data.
43. The system according to claim 30, wherein the at least one computing device is configured to run a model based on human-interpretable features.
44. A method for identifying the role of a speaker in a conversation, wherein the method is Receiving conversational data from at least one input device, To determine at least one role of at least one speaker, The process involves processing the aforementioned conversation data to generate language model output and acoustic model output, The language model output and / or the acoustic model output, each including a plurality of outputs corresponding to a plurality of time segments of the conversation data, wherein the language model output and / or the acoustic model output are weights, optionally time-based, and partially based on the at least one role of the at least one speaker in each time segment, Generating electronic reports, A method comprising transmitting the aforementioned electronic report to an output device.
45. The method according to claim 44, further comprising fusing the weighted language model output and the acoustic model output to generate a composite output.
46. The system according to claim 44, wherein the electronic report identifies the severity of at least one symptom of a condition based on the composite output.
47. The system according to claim 46, wherein the aforementioned state includes a mental health state.
48. The system according to claim 44, wherein the electronic report includes annotations of the language model output indicating at least one of the time segments.
49. The system according to claim 44, wherein the at least one speaker includes a person in charge, and applying the weights to the language model output and the acoustic model output includes applying zero weight to the acoustic model output corresponding to the person in charge.
50. The system according to claim 44, wherein the weights are partially based on the topics in each time segment.
51. The system according to claim 44, wherein the language model output includes the identification of at least one query based on conversation data from a person in charge, and at least one response to the at least one query based on conversation data from a patient, wherein the at least one query is mapped to a predetermined query, and the at least one response to the at least one query is mapped to a predetermined response to the predetermined query.
52. The method according to claim 44, wherein processing the conversation data to generate the language model output and the acoustic model output includes using a language model neural network and an acoustic neural network trained on labeled conversation data collected from one or more other subjects, wherein the labeled conversation data is labeled as (i) having a state at a certain level and (ii) not having the state.
53. The method according to claim 44, wherein the at least one role of the at least one speaker includes at least one of patient, staff member, interactive voice response speaker, and bot speaker.
54. The method according to claim 44, wherein the weights applied to the language model output and the acoustic model output are in part based on a determination that the number of at least one speaker matches the expected number of speakers.
55. The method according to claim 44, wherein the conversation data is processed in part based on at least one of a patient profile and a practitioner profile, the profile includes at least one of historical data, career data, demographic data, and longitudinal data.
56. The method according to claim 44, wherein the conversation data includes at least one of utterance data and text-based data.
57. The method according to claim 44, wherein the method is configured to run a model based on human-interpretable features.
58. A method for identifying a topic in a conversation, wherein the method is Receiving conversational data from at least one input device, The process involves processing the aforementioned conversation data to generate a language model output, wherein the language model output includes one or more topics corresponding to one or more time ranges. To generate a weighted output by applying weights to the output, wherein the output includes a plurality of outputs corresponding to a plurality of time segments of the conversation data, and the weights are optionally time-based and partially based on one or more topics in each time segment. Generating electronic reports, A method comprising transmitting the aforementioned electronic report to an output device.
59. The process involves processing the aforementioned conversation data to generate an acoustic model output, The method of claim 58, further comprising fusing the language model output and the acoustic model output by applying weights to the language model output and the acoustic model output and generating a composite output from the weighted output, wherein the acoustic model output includes a plurality of outputs corresponding to a plurality of time segments of the conversation data, and the weights are optionally time-based and partially based on the topic in each time segment.
60. The method according to claim 58, wherein the time range of the conversation data corresponding to a predetermined topic is processed to generate the language model output using a model that is more computationally robust than the model used for the time range of the conversation data not corresponding to the predetermined topic.
61. The method according to claim 58, wherein the electronic report includes a transcript of the language model output annotated partly on the topic in each time segment and partly on the weights.
62. The method according to claim 58, wherein the electronic report identifies the severity of at least one symptom of a condition based on the language model output.
63. The method according to claim 62, wherein the state includes a mental health state.
64. The method according to claim 58, wherein the electronic report includes annotations of the language model output indicating at least one of the time segments.
65. The method according to claim 58, further comprising determining at least one role of at least one speaker, wherein the weights are in part based on the at least one role of the at least one speaker in each time segment.
66. The method according to claim 58, wherein the language model output includes the identification of at least one query based on conversation data from a person in charge, and at least one response to the at least one query based on conversation data from a patient, wherein the at least one query is mapped to a predetermined query, and the at least one response to the at least one query is mapped to a predetermined response to the predetermined query.
67. The method according to claim 58, wherein processing the conversation data to generate the language model output includes using a language model neural network trained on labeled conversation data collected from one or more other subjects, wherein the labeled conversation data is labeled as (i) having a state at a certain level and (ii) not having the state.
68. The aforementioned at least one query belongs to the set of queries, The method according to claim 58, wherein the computing device is configured to predict an overall score based on the set of queries based on the response to each of the queries in the set of queries.
69. The method according to claim 58, wherein the conversation data is processed in part based on at least one of a patient profile and a practitioner profile, the profile includes at least one of historical data, career data, demographic data, and longitudinal data.
70. The method according to claim 58, wherein the conversation data includes at least one of utterance data and text-based data.
71. The method according to claim 58, wherein the method is configured to run a model based on human-interpretable features.
72. A method for scoring surveys based on conversations, Receiving conversational data from at least one input device, The process involves processing the aforementioned conversation data to generate a language model output, wherein the language model output includes the identification of at least one query based on the conversation data from the person in charge, and at least one response to the at least one query based on the conversation data from the patient, wherein the at least one query is mapped to a predetermined query, and the at least one response to the at least one query is mapped to a predetermined response to the predetermined query. Generating electronic reports, A method comprising transmitting the aforementioned electronic report to an output device.
73. The process involves processing the aforementioned conversation data to generate an acoustic model output, The method of claim 72, further comprising merging the language model output and the acoustic model output by applying weights to the language model output and the acoustic model output and generating a composite output from the weighted outputs, wherein the language model output and the acoustic model output each include a plurality of outputs corresponding to a plurality of time segments of the conversation data, and the weights are optionally time-based.
74. The system according to claim 73, wherein the weights are partially based on the at least one query or topic in each time segment.
75. The system according to claim 72, further comprising outputting at least one unresolved query to the output device.
76. The system according to claim 72, further comprising outputting a flag indicating that the at least one response to the at least one query does not map to a predetermined response to the predetermined query with confidence exceeding a threshold.
77. The system according to claim 72, further comprising outputting a prompt to repeat the predetermined query if it is determined with confidence above a threshold that the at least one response to the at least one query does not map to a predetermined response to the predetermined query.
78. The system according to claim 72, wherein the electronic report identifies the severity of at least one symptom of a condition based on the language model output.
79. The system according to claim 78, wherein the aforementioned state includes a mental health state.
80. The system according to claim 72, wherein the electronic report includes annotations of the language model output indicating at least one of the time segments.
81. The system according to claim 72, wherein the computing device is further configured to determine at least one role of at least one speaker, and the weights are partially based on the at least one role of the at least one speaker in each time segment.
82. The system according to claim 72, wherein processing the conversation data to generate the language model output includes using a language model neural network trained on labeled conversation data collected from one or more other subjects, wherein the labeled conversation data is labeled as (i) having a state at a certain level and (ii) not having the state.
83. The method according to claim 72, wherein the conversation data is processed in part based on at least one of a patient profile and a practitioner profile, the profile includes at least one of historical data, career data, demographic data, and longitudinal data.
84. The method according to claim 72, wherein the conversation data includes text-based data.
85. The method according to claim 72, wherein the at least one computing device is configured to run a model based on human-interpretable features.
86. A non-temporary computer-readable medium comprising program instructions for causing a computer to perform the method according to any one of claims 44 to 85.
87. A system for training one or more baselines of language models, acoustic models, and fusion models (models) to directly or indirectly detect behavioral or mental health conditions using machine learning, wherein the training is Using the aforementioned model, predict the behavioral or mental health status in the training data. A system including updating the model based on the accuracy of the prediction.
88. The system according to claim 87, wherein the labels in the training data are evaluated for reliability before or during training, and each training data is reweighted according to the reliability of its label.
89. The system according to claim 87, wherein the training data is augmented using at least one of paraphrasing or synthesized speech to generate additional training data for use with the training data.
90. The aforementioned training, Predicting speaker ID and The system according to claim 87, further comprising imposing a penalty on the model for identifying the speaker ID.
91. A system for predicting the severity of at least one symptom of a subject's behavioral or mental health condition, wherein the system At least one input device for receiving conversation data from the aforementioned object, At least one computing device that communicates with the aforementioned at least one input device, Receive in-context learning that includes explanations related to one or more questions in the questionnaire. The conversation data is received from at least one input device. A system comprising at least one computing device configured to predict the severity of at least one symptom of the behavioral or mental health condition of the subject, based on the in-context learning and the conversational data.
92. The system according to claim 91, wherein the computing device accesses a large-scale language model to predict the severity of at least one symptom of the behavioral or mental health condition of the subject based on the in-context learning and the conversational data.
93. The system according to claim 91, wherein the prediction of the severity of the at least one symptom of the behavioral or mental health condition includes a prediction of the results of the questionnaire.
94. The system according to claim 91, wherein the prediction of the severity of the at least one symptom of the behavioral or mental health condition includes a prediction of the result of at least one of the one or more questions on the questionnaire.
95. A system for preprocessing transcription data for use in predicting the severity of at least one symptom of a subject's behavioral or mental health condition, wherein the system is At least one input device for receiving conversation data from the aforementioned object, At least one computing device that communicates with the aforementioned at least one input device, Upon receiving in-context learning including a description of the behavioral or mental health state, The conversation data is received from at least one input device. Weighting the at least one segment of the conversation data based on the relationship between the at least one segment of the conversation data and the behavioral or mental health state, Summarizing at least one segment of the conversation data, To provide an analysis of at least one segment of the aforementioned conversation data, To summarize at least one aspect of the aforementioned behavioral or mental health condition, and Preprocessing the conversation data by performing at least one of the following: providing an analysis relating to the at least one aspect of the behavioral or mental health state. A system comprising at least one computing device configured to transmit the pre-processed conversation data to one or more models to predict the behavioral or mental health state of the subject.
96. The system according to claim 95, wherein the computing device accesses a large-scale language model and preprocesses the conversation data.