Systems and methods for generating exposure hierarchies

LLMs facilitate the generation of personalized ERP hierarchies for mental health disorders by addressing unique patient symptoms, enhancing treatment efficacy and accessibility.

WO2026050569A1PCT designated stage Publication Date: 2026-03-05THE GENERAL HOSPITAL CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/044054
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-28
Filing Date
2025-08-28
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing systems struggle to generate personalized and effective exposure and response prevention (ERP) hierarchies for mental health disorders like OCD, as they often fail to address a patient's unique symptoms and require significant clinical expertise, leading to suboptimal treatment and low response rates.

Method used

Utilizing large language models (LLMs) to generate ERP hierarchies by receiving user inputs, extrapolating missing symptoms, referencing a database of ERP tasks, and organizing them into a personalized hierarchy, with expert oversight to ensure safety and relevance.

Benefits of technology

LLMs can effectively generate tailored ERP hierarchies that are safe, specific, and useful, reducing the treatment gap by enabling more clinicians and patients to deliver personalized ERP, despite challenges in clinical expertise and time constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025044054_05032026_PF_FP_ABST
    Figure US2025044054_05032026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods of generating an ERP hierarchy for a subject implement a multi-agent model to generate a profile for a subject; generate a set of ERP task recommendations; generate a preliminary ERP hierarchy; and present the preliminary ERP hierarchy to the user. The user may edit or approve the agent outputs at several stages in the workflow.
Need to check novelty before this filing date? Find Prior Art

Description

Client No: MGH 2024-165-02 Quarles: 125141.04855 PCT PATENT APPLICATION FOR SYSTEMS AND METHODS FOR GENERATING EXPOSURE HIERARCHIES Sabine Wilhelm Adam C. Jaroszewski Ryan J. Jacoby QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) SYSTEMS AND METHODS FOR GENERATING EXPOSURE HIERARCHIES CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of and priority to U.S. Provisional Patent Application No.63 / 688,245, filed on August 28, 2024, titled “AI-Driven Tool for Generating Exposure Hierarchies,” the entire contents of which are each herein incorporated by reference for all purposes. STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT

[0002] Not applicable. BACKGROUND

[0003] Mental health disorders affect a large number of people worldwide. Obsessive compulsive disorder (OCD), for example, is characterized by persistent, intrusive thoughts coupled with compulsions, or repetitive actions or rituals intended to neutralize or rid oneself of the obsessions. OCD occurs in approximately 2% of the population and was named the 10th leading cause of impairment among all health conditions by the World Health Organization given disproportionately high rates of poor outcomes like unemployment (up to 41%) and suicidal behaviors (up to 27%). Without treatment, OCD has a chronic and severe course, underscoring the critical importance of access to effective treatment for OCD sufferers. As another example, body dysmorphic disorder (BDD) is characterized by persistent, intrusive concerns regarding one’s own appearance. BDD also occurs approximately 2% of the population, and a large majority (up to 77%) of persons with BDD report that their systems interfere with their occupational or academic progress.

[0004] Cognitive behavioral therapy (CBT), and specifically exposure and response prevention (ERP), has been used as a treatment for OCD and other mental health disorders. When used to treat OCD, ERP involves systematically confronting the thoughts, images, objects, and situations that make the individual anxious or provoke their obsessions while 1 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) limiting engagement in compulsive responses. Exposures can occur both in real life (e.g., touching surfaces perceived to be contaminated) and in one’s imagination by reading or listening to an exposure script (e.g., imagining contracting a deadly virus). During ERP patients learn that their feared outcomes are less likely to come true than they initially thought (e.g., that they do not show signs of illness despite not washing their hands after touching something that they considered to contaminated) and that they can tolerate distress without engaging in compulsive or ritualized behaviors. Despite robust supporting evidence for ERP, the majority (57.3%) of individuals with OCD cannot access any treatment, let alone ERP with a qualified clinician.

[0005] In comparative examples, technological advances have been used in an attempt to reduce the treatment gap. These examples attempt to use, for example, virtual provider training to increase the number of providers equipped to deliver ERP by lowering barriers like cost and the logistics of traveling to in-person programs. Program evaluations show that clinicians new to treating OCD can be effectively trained in ERP through virtual, light touch instruction. These examples may also involve developing smartphone- and internet-based platforms to deliver or support ERP (typically with the help of clinicians or coaches), which can reduce the cost, associated stigma, and other barriers to care as users can access content whenever and wherever is convenient. Trials have shown that guided digital CBT for OCD is feasible, safe, and effective.

[0006] However, even for those who are able to access ERP for mental health disorders such as OCD, response and remission rates across in-person and virtual treatment remain too low. Personalizing treatment, particularly developing a hierarchy of exposures that are considered to be useful and safe, is difficult; this difficulty contributes to providers delivering suboptimal ERP, as well as to patients and coaches struggling with digital and self-guided treatments. Treatment personalization requires creativity and clinical acumen, as effective exposures must directly address a patient’s specific obsessions and compulsions, be appropriately graded (i.e., not too difficult or too easy), and be both feasible and safe within a patient’s context. This can also be time consuming, further straining already burdened clinicians and discouraging self-help or digital treatment users who may struggle to develop their own hierarchies. If an automated (e.g., computerized) system is used to generate ERPs 2 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) that are ineffective, as would be the case in comparative examples, scarce computing resources may be wasted. Moreover, given the high heterogeneity of obsessions and compulsions seen in OCD, example hierarchies and exposures provided in a smartphone app, treatment manual, or training may not be directly relevant to a given patient’s unique symptoms. For example, an individual who worries that having an unacceptable taboo thought while getting dressed means that they need to change their clothes to prevent something bad from happening may struggle to apply ERP principles to their specific obsessions and compulsions if the examples provided to them include more traditional contamination and harm themes.

[0007] Accordingly, there exists a need for systems and methods that provide effective guidance in planning exposures to realize the promise of scaling trainings and treatments for OCD and other mental health disorders. SUMMARY

[0008] The present disclosure addresses these and other needs by providing systems, devices, methods, algorithms, and / or media for generating an ERP hierarchy. In examples, the techniques set forth herein leverage large language models (LLMs) or other generative artificial intelligence (AI) based platforms.

[0009] According to one aspect of the present disclosure, a method of generating an exposure and response prevention (ERP) hierarchy for a subject is provided. The method comprises generating an ERP hierarchy for a subject perform or include receiving an input from a user via a first interface, the input including a plurality of input symptoms regarding the subject, the plurality of input symptoms corresponding to a behavioral health condition; generating a profile for the subject by a first agent, wherein the first agent is configured to receive the input, extrapolate at least one missing symptom from the input, and generate the profile based on the input and the at least one missing symptom; generating a set of ERP task recommendations by a second agent, wherein the second agent is configured to receive the input, reference a database of preexisting ERP tasks, and select the set of ERP task recommendations from among the preexisting ERP tasks based on the input; generating a preliminary ERP hierarchy by a third agent, wherein the third agent is configured to receive the profile and the set of ERP task recommendations, generate a plurality of ERP task 3 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) candidates, and organize at least one of the plurality of ERP task candidates into the preliminary ERP hierarchy; and presenting the preliminary ERP hierarchy to the user via a second interface.

[0010] According to another aspect of the present disclosure, a system for generating an ERP hierarchy for a subject is provided. The system comprises a memory; a user interface configured to interact with a user; and at least one processor operatively connected to the memory and the user interface, the at least one processor configured to receive an input from the user via a first component of the user interface, the input including a plurality of input symptoms regarding the subject, the plurality of input symptoms corresponding to a behavioral health condition, generate a profile for the subject by a first agent, wherein the first agent is configured to receive the input, extrapolate at least one missing symptom from the input, and generate the profile based on the input and the at least one missing symptom, generate a set of ERP task recommendations by a second agent, wherein the second agent is configured to receive the input, reference a database of preexisting ERP tasks, and select the set of ERP task recommendations from among the preexisting ERP tasks based on the input, generate a preliminary ERP hierarchy by a third agent, wherein the third agent is configured to receive the profile and the set of ERP task recommendations, generate a plurality of ERP task candidates, and organize at least one of the plurality of ERP task candidates into the preliminary ERP hierarchy, and present the preliminary ERP hierarchy to the user via a second component of the user interface. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Some examples of the disclosure are described herein with reference to the accompanying figures. The description, together with the figures, makes apparent to a person having ordinary skill in the art how some implementations of the disclosure may be practiced. The figures are for the purpose of illustrative discussion and no attempt is made to show structural details of an example in more detail than is necessary for a fundamental understanding of the teachings of the disclosures. In the drawings:

[0012] FIG. 1 illustrates an example workflow for a system according to various aspects of the present disclosure. 4 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758)

[0013] FIG.2 illustrates an example workflow for generating an initial ERP hierarchy according to various aspects of the present disclosure.

[0014] FIG. 3A illustrates an example interface according to various aspects of the present disclosure.

[0015] FIG. 3B illustrates an example interface according to various aspects of the present disclosure.

[0016] FIG. 3C illustrates an example interface according to various aspects of the present disclosure.

[0017] FIG.4 illustrates an example system according to various aspects of the present disclosure.

[0018] FIG.5 illustrates an example method according to various aspects of the present disclosure. DETAILED DESCRIPTION

[0019] In the following detailed description, reference is made to the accompanying drawings in which specific examples are shown by way of illustration. These examples are described in sufficient detail to enable those of ordinary skill in the art to practice the disclosure. It should be understood, however, that the detailed description and the specific examples, while indicating examples of embodiments of the disclosure, are given by way of illustration only and not by way of limitation. From this disclosure, various substitutions, modifications, additions rearrangements, or combinations thereof within the scope of the disclosure may be made and will become apparent to those of ordinary skill in the art.

[0020] For example, while the following description provides examples related primarily to OCD, the principles set forth herein may be applied to other disorders to generate effective ERPs or other aspects of CBT. For example, the systems and methods of the present disclosure may be applied to BDD or other disorders that are treatable using CBT plans. Thus, any reference to a particular disorder (e.g., OCD) should be understood as being made by way of example and not limitation. Moreover, while the following examples may be described with the use of particular LLMs such as ChatGPT, in other examples the present disclosure may be 5 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) realized through the use of other LLMs, including custom LLMs, either alone or in combination with ChatGPT. Thus, any reference to a particular LLM herein should be understood as being made by way of example and not limitation.

[0021] Unless otherwise indicated, the various features illustrated in the drawings may not be drawn to scale. The illustrations presented herein are not necessarily intended to be actual views of any particular method, device, or system, but are merely idealized representations that are employed to describe various embodiments of the disclosure. Accordingly, the dimensions of the various features as illustrated may be arbitrarily expanded or reduced for clarity. In addition, some of the drawings may be simplified for clarity. Thus, the drawings may not depict all of the components of a given apparatus (e.g., device) or method. In addition, like reference numerals may be used to denote like features throughout the specification and figures.

[0022] It should be understood that any reference to an element herein using a designation such as “first,” “second,” and so forth does not limit the quantity or order of those elements, unless such limitation is explicitly stated. Rather, these designations may be used herein as a convenient method of distinguishing between two or more elements or instances of an element. Thus, a reference to first and second elements does not mean that only two elements may be employed there or that the first element must precede the second element in some manner. Also, unless stated otherwise a set of elements may comprise one or more elements.

[0023] Unless otherwise specified or indicated by context, the terms “a,” “an,” and “the” mean “one or more.” As used herein, unless otherwise limited or defined, “or” indicates a non-exclusive list of components or operations that can be present in any variety of combinations, rather than an exclusive list of components that can be present only as alternatives to each other. For example, a list of “A, B, or C” indicates options of: A; B; C; A and B; A and C; B and C; and A, B, and C. Correspondingly, the term “or” as used herein is intended to indicate exclusive alternatives only when preceded by terms of exclusivity, such as “only one of,” or “exactly one of.” For example, a list of “only one of A, B, or C” indicates options of: A, but not B and C; B, but not A and C; and C, but not A and B. In contrast, a list preceded by “one or more” (and variations thereon) and including “or” to separate listed elements indicates options of one or more of any or all of the listed elements. For example, the 6 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) phrases “one or more of A, B, or C” and “at least one of A, B, or C” indicate options of: one or more A; one or more B; one or more C; one or more A and one or more B; one or more B and one or more C; one or more A and one or more C; and one or more A, one or more B, and one or more C. Similarly, a list preceded by “a plurality of” (and variations thereon) and including “or” to separate listed elements indicates options of one or more of each of multiple of the listed elements. For example, the phrases “a plurality of A, B, or C” and “two or more of A, B, or C” indicate options of: one or more A and one or more B; one or more B and one or more C; one or more A and one or more C; and one or more A, one or more B, and one or more C.

[0024] As used herein, “about,” “approximately,” “substantially,” and “significantly” will be understood by persons of ordinary skill in the art and will vary to some extent on the context in which they are used. If there are uses of these terms which are not clear to persons of ordinary skill in the art given the context in which they are used, “about” and “approximately” will mean plus or minus ≤10% of the particular term and “substantially” and “significantly” will mean plus or minus >10% of the particular term.

[0025] As used herein, the terms “include” and “including” have the same meaning as the terms “comprise” and “comprising” in that these latter terms are “open” transitional terms that do not limit claims only to the recited elements succeeding these transitional terms. The term “consisting of,” while encompassed by the term “comprising,” should be interpreted as a “closed” transitional term that limits claims only to the recited elements succeeding this transitional term. The term “consisting essentially of,” while encompassed by the term “comprising,” should be interpreted as a “partially closed” transitional term which permits additional elements succeeding this transitional term, but only if those additional elements do not materially affect the basic and novel characteristics of the claim.

[0026] In some examples, the systems and methods set forth herein may be implemented on or using one or more computing devices, each of which includes a processor and a memory. As used herein, a “processor” may include one or more individual electronic processors, each of which may include one or more processing cores, and / or one or more programmable hardware elements. The processor may be or include any type of electronic processing device, including but not limited to central processing units (CPUs), graphics 7 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) processing units (GPUs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), microcontrollers, digital signal processors (DSPs), or other devices capable of executing software instructions. When a device is referred to as “including a processor,” one or all of the individual electronic processors may be external to the device (e.g., to implement cloud or distributed computing). In implementations where a device has multiple processors and / or multiple processing cores, individual operations described herein may be performed by any one or more of the microprocessors or processing cores, in series or parallel, in any combination. In some implementations, one or more of the processing units or processing cores may be remote (e.g., cloud-based).

[0027] As used herein, a “memory” may be any storage medium, including a non- volatile medium, e.g., a magnetic media or hard disk, optical storage, or flash memory; a volatile medium, such as system memory, e.g., random access memory (RAM) such as dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM), static RAM (SRAM), extended data out (EDO) DRAM, extreme data rate dynamic (XDR) RAM, double data rate (DDR) SDRAM, etc.; on-chip memory; and / or an installation medium where appropriate, such as software media, e.g., a CD-ROM, or floppy disks, on which programs may be stored and / or data communications may be buffered. The term “memory” may also include other types of memory or combinations thereof. For the avoidance of doubt, cloud storage is contemplated in the definition of memory. A memory is an example of a non-transitory computer-readable medium which stores instructions that are executable by a processor (or processors), the execution of which causes the executing device (e.g., a computer) to perform certain operations, such as those operations described herein.

[0028] As noted above, comparative examples of generating ERPs suffer from several challenges. The present disclosure sets forth systems and methods that address these challenges. In one example, the systems and methods set forth herein leverage LLMs. Such generative AI-based platforms, like OpenAI’s ChatGPT, use natural language processing algorithms to engage in human-like conversation. For example, ChatGPT can support clinical decision making, pass medical exams, answer patient queries, and assist in scientific writing and literature reviews with largely “passing” performance. Within psychiatry specifically, LLMs have shown preliminary promise in supporting diagnosis and clinical decision making 8 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) and directly supporting patients through psychoeducation and therapeutic conversations. Overall, LLMs, and ChatGPT in particular, appear more adept in following clinical guidelines and answering simple questions and less adept when presented with more severe or complex cases or asked to provide personalized clinical guidance. Overall, publicly available platforms like ChatGPT perform increasingly well on clinical tasks, while also continuing to exhibit concerning weaknesses, including inaccurate references and sometimes even dangerous suggestions. Thus, to safely and optimally integrate these technologies into clinical workflows, it may be preferable that models are evaluated by experts and further trained using specialized datasets.

[0029] In support of the present disclosure, studies were performed to demonstrate the current ability of a widely used and publicly available LLM (GPT-4) to produce exposure hierarchies for OCD treatment. The impact of various clinical (e.g., subtype and complexity of OCD symptoms) and demographic (e.g., patient age) features of simulated patient case descriptions on the ChatGPT-generated output was evaluated. ChatGPT performance was also benchmarked against expert, doctoral-level OCD therapists. Given the low rate of clinicians offering, let alone specializing, in OCD treatment, such implementations may effectively augment face-to-face and digital care.

[0030] In the study, ChatGPT (GPT-4) was used to generate 10-item exposure hierarchies for a series of simulated patient cases described further below. All model testing was performed in new ChatGPT sessions to limit influence from previous prompts. Prompts were written by experts in the diagnosis and treatment of OCD and providers in an OCD specialty clinic, and manually entered into the ChatGPT website. Responses were copied verbatim into a database for review, as will be described in more detail below.

[0031] Prompts included (1) context for the request, (2) how the model was meant to respond, (3) the specific request, and (4) the output format. Specifically, prompts used the following template: “I am a mental health provider creating a treatment plan. Please respond as if you are an expert in Exposure and Response Prevention (ERP) treatment for Obsessive Compulsive Disorder (OCD). Please create a 10-item graded exposure hierarchy for a patient whose primary obsession(s) is (are) __. The corresponding compulsion(s) is (are) __. The patient is __years old and identifies as a __.” Given the known impact of prompt engineering 9 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) on generative LLM output, prompt formats were standardized and the following dimensions were systematically varied: OCD subtype (contamination, accidental harm, sexual / physical violence), symptom complexity or number (low [1 obsession and 1 compulsion], moderate [2 obsessions and 4 compulsions], high [3 obsessions and 6 compulsions]), level of detail (low, high), patient age (15 years old, 40 years old), patient gender (woman, man).

[0032] An example of a low detail prompt is: “I am a mental health provider creating a treatment plan. Please respond as if you are an expert in Exposure and Response Prevention (ERP) treatment for Obsessive Compulsive Disorder (OCD). Please create a 10-item graded exposure hierarchy for a patient whose primary obsession is fear and disgust of being contaminated by bodily excretions (urine and feces). The corresponding compulsion is excessive handwashing. The patient is 15 years old and identifies as a man.” An example of a high detail prompt is: “I am a mental health provider creating a treatment plan. Please respond as if you are an expert in Exposure and Response Prevention (ERP) treatment for Obsessive Compulsive Disorder (OCD). Please create a 10-item graded exposure hierarchy for a patient whose primary obsession is fear of being contaminated by bodily excretions (urine and feces). He is afraid that urine and feces will get on his skin after using the bathroom or in public restrooms, unknowingly enter his body, and get him sick and that he will then unknowingly spread this sickness to others. He also reports feeling disgusted after using the restroom at the thought of fecal matter being on his skin. The corresponding compulsion is washing his hands with very hot water for more than 20 minutes every time he uses the bathroom or has a thought about being contaminated. He also washes his hands in a ritualized order by going in between each finger in order and then washing up the arm to the elbow. The patient is 15 years old and identifies as a man.”

[0033] In total, 72 prompts were submitted to ChatGPT, capturing all possible combinations of the aforementioned dimensions. In the event that ChatGPT failed to produce a hierarchy (i.e., responded with an error), the response was still pasted into the database for review. Subsequently, investigators were to prompt ChatGPT a second time using: “What activities might help a patient with these symptoms?”. This output was used for blinded reviews (described below). In order to compare human-vs ChatGPT- generated hierarchies, expert clinicians (EB, AJ, RJ) generated hierarchies for 18 of the 72 case description prompts (short 10 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) and long version of each subtype / number of symptom combinations; gender and age were counterbalanced).

[0034] To analyze the data, an initial review was performed. All ChatGPT responses were reviewed and evaluated by two psychologists along the following dimensions: (1) task completion and (2) degree to which input information was incorporated in the output. Task completion was rated categorically: not at all, partial, or complete. A ChatGPT response was marked as “complete” if it included a list of 10 scenarios that fit the definition of an exposure with response prevention. A response was marked as “not at all” complete if ChatGPT failed to generate any suggestions after the initial prompt. All other responses were marked “partial”. A third psychologist reviewed responses to resolve any discrepancies. The degree to which input information (i.e., prompt content) was incorporated was rated on a 5-point scale from 1=not at all to 5=completely. A response received a 1 (not at all) if, for example, the prompt referred to contamination and violence obsessions for an adult and the response provided symmetry exposures in a school environment. A response received a 5 (completely) if all suggestions fit the prompt and all prompt content was incorporated. Note that this rating was not an evaluation of the quality of the suggestions. Interrater reliability was assessed by calculating intraclass correlation coefficients (ICC) for each rating dimension, based on a mean rating (k = 2), absolute agreement, two-way mixed-effects model. Raters had good agreement on both ‘task completion’ (ICC = .76) and ‘input information incorporated’ (ICC = .74).

[0035] Descriptives are presented and quantitatively evaluated as to whether any variable (e.g., prompt length, OCD subtype) performed worse than others on task completion (chi-squares) and input information incorporation (t-tests, ANOVAs).

[0036] Thereafter, a blinded review was performed. Three blinded psychologists with expertise in OCD and ERP served as blinded raters. Raters all split their time practicing in both hospital (M=56.33%, SD=14.15) and private practice (M=40.00%, SD=18.03%) settings and were on average 13.33 years (SD=9.07) post-terminal degree. Raters have each treated over 100 patients with OCD and reported using ERP in the treatment of nearly all of them (M=99.67% of patients with OCD, SD=0.58). They were presented with all ChatGPT- generated (n=72) and human-generated (n=18) hierarchies but were blinded to the generation source and aims of the project. They were told that team members were developing new 11 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) programs to train and support people in delivering ERP for OCD and looking for outside experts to assess how the pilot program went. They rated each hierarchy along the following dimensions using a 5-point scale (1-strongly disagree to 5-strongly agree): (1) appropriateness (how appropriate the suggestions are given patient’s symptoms), (2) specificity (how actionable the suggestions are), (3) variability (to what degree suggestions fit the expected range of exposures needed for a full hierarchy; they were reminded that instructions were always to generate 10-item exposures, so hierarchies should not be penalized for length), (4) safety / ethics (to what degree suggestions are safe and ethical for the patient to conduct), and (5) overall usefulness or quality. ICCs for rating dimensions were based on a mean rating (k = 3) absolute agreement two-way mixed-effects model Raters had low agreement on hierarchy, generated, matched by prompt), asked to identify which was AI-generated, and rated their confidence in this judgment (0-100% confident). Below, descriptives of ratings are presented and quantitatively evaluated as to whether any variable (e.g., prompt length, OCD subtype) biased the above ratings (t-tests, ANOVAs).

[0037] ChatGPT generated partial (n=15, 20.8%) or complete (n=55, 76.4%) responses to 70 of 72 initial prompts. The two prompts (2.8%) that resulted in errors (i.e., flagged by OpenAI’s content filtering system, with no output generated) were both cases describing sexual / violent obsessions with high number of symptoms and high detail. With a single follow- up prompt, ChatGPT produced hierarchy suggestions for these as well. Although ChatGPT occasionally failed to produce all 10 exposure suggestions, the cause for a hierarchy being labeled as ‘partially’ completed was more often (13 out of 15, 86.7%) due to the inclusion of activities that did not meet the definition of an exposure (e.g., discussing one’s feelings or QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758)Gender^^^^2=3.05, p=.22, V=.21t(59.64)=1.39, p=.17, d=.33

[0039] Blinded ratings are summarized in Table 2. Overall, ChatGPT-generated hierarchies were viewed as largely appropriate (M=4.47, SD=0.58), specific (M=4.17, SD=0.65), variable (M=3.96, SD=0.79), safe (M=4.89, SD=0.24), and overall useful (M=3.99, SD=0.82). However, expert human-generated hierarchies were still rated as significantly more appropriate, specific, variable, and useful, ps<.05, though, importantly, just as safe and ethical, p=.24. Whereas ratings of the expert human-generated hierarchies never fell below a 3 out of 5, there was more variability in the ChatGPT-generated hierarchies, and some did receive unacceptable scores. Table 2 – Blinded Ratings of Hierarchies 13 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) ChatGPT- Human- Generated Generated (N=72) (N=18)

[0040] In the ChatGPT-generated hierarchies specifically, OCD subtype, number of symptoms, patient age, and patient gender were all unrelated to hierarchy ratings, ps > .05 (See Table 3). Symptom detail was a significant predictor of appropriateness, specificity, variability, and overall utility ratings, ps<.05 (but not safety ratings, p=.15), such that hierarchies were rated significantly better on these dimensions when prompts were more detailed. Examining qualitative feedback provided by the raters, specificity and overall usefulness ratings were typically lower when raters noted insufficient detail regarding ritual prevention. Variability ratings were typically lower in cases where raters felt there were not enough challenging exposures or where suggestions were focused on limited contexts (e.g., exposure only to public toilets, not including imaginal exposures).

[0041] After unblinding, raters correctly identified 94.44% (Rater 1: 17 / 18), 66.67% (Rater 2: 12 / 18), and 66.67% (Rater 3: 12 / 18) pairs of ChatGPT- versus human-generated 14 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) hierarchies and reported varying levels of confidence M(SD) = 86.72% (15.32%), M(SD)=46.72% (16.09%), and M(SD)=55.11% (14.70%), respectively. Table 3 – Predictors of Ratings for ChatGPT-generated Hierarchies Appropriate- Specificity Variability Safety Overall ness Usefulness OCD Sub- F(2,69)=.73, F(2,69)=2.52, F(2,69)=.54, F(2,69)=1.29, F(2,69)=1.26, Type p=.49 p=.09 p=.59 p=.28 p=.29 Symptom F(2,69)=.33, F(2,69)=1.71, F(2,69)=.09, F(2,69)=.41, F(2,69)=.03, Number p=.72 p=.19 p=.92 p=.66 p=.97 DetailF(1,70)=6.92,F(1,70)=11.51, F(1,70)=8.31, F(1,70)=2.16, F(1,70)=6.93, p=.01 p=.001 p=.005 p=.15 p=.01 AgeF(1,70)=1.05,F(1,70)=.13, F(1,70)=.72, F(1,70)=.65, F(1,70)=.18, p=.31 p=.72 p=.40 p=.42 p=.67 GenderF(1,70)=1.36,F(1,70)=.02, F(1,70)=.72, F(1,70)=.03, F(1,70)=2.86, p=.25 p=.90 p=.40 p=.87 p=.10

[0042] ChatGPT was able to generate graded exposure hierarchies for a range of OCD presentations and levels of complexity. Suggested exposures were, on average, rated as being fairly appropriate, specific or actionable, variable in difficulty, safe, ethical, and useful. Where novice clinicians may struggle with designing exposures for certain OCD subtypes, such as violent or sexual obsessions, or more complex symptom presentations, ChatGPT showed no such bias (at least once a follow-up prompt was used to bypass original error flags with sexual / violent obsessions). Additionally, no biases were detected for the gender or age of the participant. Encouragingly, all suggestions provided by ChatGPT were rated as safe and ethical and safety / ethics ratings were not statistically different than those for hierarchies created by OCD experts. This feasibility study demonstrates the ability to use LLMs to support ERP hierarchy generation.

[0043] In view of the above, certain safeguards may be applied to ensure that no harm is done. First, although appropriateness, specificity, variability, and overall utility of 15 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) hierarchies benefited when prompts contained more clinical detail, such prompts were also less likely to fully integrate all symptoms. However, the fact that some symptoms were occasionally omitted in the output, particularly where the prompt included a high number of symptoms (3 obsessions and 6 compulsions), is plausibly an artifact of the prompt constraining the output to just 10 hierarchy items. The larger issue regarding task completion scores was the inclusion of suggestions that were not exposures; although sometimes these alternative activities were benign or otherwise therapeutic (e.g., recognizing unhelpful thinking patterns), other times they would be contraindicated (e.g., providing reassurance about the safety of home appliances). Relatedly, ChatGPT suggested using relaxation strategies during exposures in ways which would typically not be recommended.

[0044] Additionally, although exposure suggestions were largely reviewed positively, a meaningful number were rated as inadequate. Specifically, raters indicated that 2.78% of generated hierarchies were not appropriate, 5.56% were not specific, 15.28% were not variable, and 18.06% were not useful (average rating ≤ 3). In contrast, no human-generated hierarchies received these scores. Unsurprisingly then, ChatGPT- and human-generated hierarchies were distinguishable more often than not, and ChatGPT-generated hierarchies received significantly lower ratings (as well as more variable ratings) across these dimensions compared to human- generated hierarchies, in this initial study. Some drivers of this difference include ChatGPT providing insufficient ritual prevention (e.g., limiting handwashing in addition to touching contaminated surfaces), omitting imaginal exposures where they would be appropriate, omitting difficult exposures (e.g., using a sharp knife close to a loved one), and not deviating from the specific examples provided in the prompt (e.g., all contamination exposures focused on toilets, no exposures combined multiple feared stimuli). Such weaknesses underscore the importance of leveraging clinical expertise to fine-tune models for this task. Continuous monitoring and updating would also be important; as models suggest increasingly challenging exposures, risk of suggesting dangerous or unethical exposures may increase.

[0045] Implementations of the above study may also implement thoughtful prompt engineering as well as a user interface to ensure sufficient detail in the input. Additionally, to fully realize the promise of this approach and limit errors, many more OCD sub-types and 16 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) presentations as well as potentially relevant demographic and medical factors (e.g., psychiatric or medical comorbidities) may be tested.

[0046] Overall, the above-described feasibility study illustrates how LLMs can empower more providers to deliver high quality, personalized ERP. Implementation strategies may be adjusted to guard against inherent risks of these tools. For example, users may be reminded to confirm the safety and feasibility of the suggestions for a given patient’s context and not enact any suggestions outside their competence. Involving provider and patient stakeholders throughout this process is also beneficial. User-centered design and usability testing ensures that the tool fits the intended users’ needs, context, and preferences, integrates into clinical workflows, and ultimately improves outcomes. Outstanding ethical and privacy questions should also be considered. For example, such applications must uphold data security and privacy standards and patients should be informed about how their data would be used.

[0047] Certain conclusions can be drawn from the above-described study. The first- line treatment for OCD and a range of other disorders (e.g., anxiety disorders, body dysmorphic disorder) involves exposure and response prevention. To be effective, exposure and response prevention exercises should be tailored to the patient and graded (ranging from easy to very challenging). Generating appropriately targeted exercises is difficult for many patients and clinicians, as it typically involves considerable session time, creativity, expertise in treating a given population, and clinical experience. This is a barrier to clinicians using exposures and to patients completing self-guided treatments (e.g., app-based; sustaining gains after treatment). Despite its strong evidence base, many patients never receive adequate exposure therapy. LLMs have the potential to allow more professional and paraprofessional support persons to deliver personalized, high-quality ERP for OCD, to enable patients to use self-help apps more effectively, and thus to reduce the treatment gap and associated disease burden.

[0048] FIG.1 illustrates an example workflow 100 for a system of generating exposure hierarchies, according to the present disclosure. In FIG. 1, solid lines show determined flow and dashed lines show conditional flow. The workflow 100 implements an agentic model, through which one or more AI models perform operations in concert. The AI models (or “agents”) are capable of independently operating to perform tasks (e.g., to analyze data, to generate decisions, etc.) in an automated manner. An agent may be or include an LLM or other 17 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) AI model (e.g., a neural network), and may include or be in communication with one or more databases. In operation, an agent may receive an input (e.g., from a user or from another agent) and, based on the training of the AI model and in some cases by referencing the associated database, generate an output. Where an agent implements an LLM, the LLM may act as a reasoning engine to implement one or more techniques, such as retrieval-augmented generation (RAG), to generate the output.

[0049] The workflow 100 begins with an interface 102 through which a user may interact with the system. In some examples, the interface 102 may be a web page or application presented to the user, for example via a computing device (e.g., a laptop computer, a workstation, a smartphone, a tablet computer, etc.) associated with the user. One such example will be described in more detail below with regard to FIGS. 3A-3C. In other examples, the interface 102 may be a chatbot-style interface (e.g., text-based, voice-based, etc.) with which the user may interact using natural language. The user may be a clinician or a patient. In still other implementations, the interface 102 may be configured to receive its input from another agent (e.g., a next-generation model).

[0050] The user enters information regarding a subject (e.g., a patient) via the interface 102, which is parsed to determine a set of input symptoms 104. The input symptoms 104 are passed to both a missing symptom agent 106 and a domain analysis agent 118. In some implementations, the full set of input symptoms 104 are passed to both the missing symptom agent 106 and the domain analysis agent 118; however, in other implementations a partial set may be passed to the missing symptom agent 106, the domain analysis agent 118, or both (in which case the partial set may be different for each recipient agent). The role of the missing symptom agent 106 is to check for possible missing symptoms in the set of input symptoms 104 (e.g., symptoms that the user may have failed to enter or may have failed to realize are present). These missing symptoms are passed to a symptom extrapolator agent 108, which is an agent that interprets the symptoms to make labeled inferences and generate extrapolated symptoms. To perform these tasks, the symptom extrapolator agent 108 may consult a symptom list 110 (e.g., as stored in a database) that may include external triggers, internal triggers, obsessions, etc. The extrapolated symptoms are then passed (e.g., in conjunction with 18 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) the input symptoms 104) to a profile generator 112, which serves to generate a subject profile based on its received input.

[0051] The subject profile generated by the profile generator agent 112 may then be subjected to user approval. That is, the subject profile, in a preliminary or initial form, may be presented to a user via an interface 114. The interface 114 may be a web page or application presented to the user, a chatbot-style interface, and the like. Via the interface 114, the user is presented with the option to approve the initial subject profile or edit the initial subject profile. If the user desires to edit the initial subject profile, the edited profile may be treated as a new set of input symptoms 104 and the workflow of FIG.1 (implementing elements 106-114) may be performed anew. This procedure may proceed in an iterative manner until the user approves the subject profile, in which case the approved profile is provided to an ERP designer agent 116. In some implementations, however, the interface 114 may be omitted or replaced with an additional agent that checks the subject profile.

[0052] The parsed input symptoms 104 (or portion thereof) are also passed to the domain analysis agent 118, which may analyze the input symptoms 104 on a domain-by- domain basis to determine symptoms and associated ERP tasks. A more detailed description of this operation is provided below with regard to FIG. 2. In general, however, the domain analysis agent 118 determines a set of symptom domains for the input symptoms 104 and, working in conjunction with an ERP retrieval tool 120, determines a set of similar ERP cases to pass to the ERP designer agent 116. In so doing, the ERP retrieval tool 120 may consult an ERP database 122, which will also be described in more detail below with regard to FIG.2.

[0053] The ERP designer agent 116 generates ERPs based on the approved subject profile (received from the interface 114) and the retrieved ERP cases (received from the ERP retrieval tool 120). The ERP communicates with an imaginal reviewer agent 124, an ERP scorer agent 126, an ERP re-ranker tool 128 (which may also be agentic), a difficulty reviewer agent 130, and an ERP-10 selector agent 132. The imaginal reviewer agent 124 may be an agent that generates or selects forms of imaginal exposure therapy that may form the basis for or be included in the ERP. The ERP scorer agent 126 may review the imaginal exposure therapy or other elements of the generated ERP and assigns them scores corresponding to various parameters (difficulty, intensity, etc.). These scores are output to the ERP re-ranker tool 128 19 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) which orders them according to the parameter(s). For example, the ERP scorer agent may assign difficulty scores to various aspects of the ERP and, in coordination with the difficulty reviewer agent 130, sorts the aspects in the correct order so that the resulting ERP may have increased effectiveness (e.g., so that the tasks are most digestible by the subject). The ERP-10 selector agent 132 may select ten elements to be included in the ERP hierarchy. Of course, in other implementations, the selector agent 132 may be configured to select a different number of elements.

[0054] In either case, the selected hierarchy may be output to the interface 134. The interface 134 may be a web page or application presented to the user, a chatbot-style interface, and the like. Via the interface 134, the user is presented with the option to approve the proposed ERP hierarchy or edit the proposed ERP hierarchy. While FIG.1 illustrates the interface 102 and the interface 134 as separate elements, the interface 102, the interface 114, and the interface 134 may be the same interface or may be separate elements (e.g., separate screens or tabs) of the same interface. Moreover, the user that interacts with the interface 102, the user that interacts with the interface 114, and the user that interacts with the interface 134 may be the same user (e.g., the same clinician) or different users. If the user desires to edit the proposed ERP hierarchy, the edited hierarchy may be treated as a new set of ERP parameters and the workflow of FIG.1 (implementing elements 116-132) may be performed anew. This procedure may proceed in an iterative manner until the user approves the ERP hierarchy, in which case the approved hierarchy is presented in the form of a final report 136. In some implementations the interface 134 may be omitted or replaced with an additional agent that checks the ERP hierarchy, however, it may be preferred that final approval be performed by a user.

[0055] While FIG. 1 illustrates each agent as a separate entity, in practical implementations subsets of the illustrated agents may be implemented as the same agent. In one particular example, the symptom extrapolator agent 108, the profile generator agent 112, the ERP designer agent 116, the domain analysis agent 118, and the ERP-10 selector agent 132 may be implemented by a first agent; the missing symptom agent 106 may be implemented by a second agent; and the imaginal reviewer agent 124, the ERP scorer agent 126, and the difficulty reviewer agent 130 may be implemented by a third agent. This grouping of agents is merely exemplary. In implementations, the agents of FIG. 1 may be grouped based on, for 20 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) example, similarities in the operations performed, the reasoning applied, the format of inputs / outputs, the databases or other sources consulted, and so on.

[0056] FIG.2 illustrates an example workflow 200 for a portion of the workflow 100. In particular, FIG.2 illustrates an example of how data input via the interface 102 may be used to generate an initial ERP hierarchy by the ERP designer 116. As noted above with regard to FIG.1, the input symptoms 104 are used to generate a profile (via profile generator agent 112) and are also output to the domain analysis agent 118. The profile may include information regarding triggers for the patient, maladaptive cognitive reactions of the patient, maladaptive behavioral responses of the patient, and the like. The profile may also be provided to the domain analysis agent 118, where it and the symptom domain information are used in combination with an ERP database 122. For example, aspects of the generated patient profile may be parsed or sorted into a series of symptom domains.

[0057] The ERP database 122 represents a corpus of ERP hierarchies created by experts. The entries may be used by agents to perform RAG; that is, to refer back to a source in order to match the subject’s characteristics with predetermined knowledge. This may be especially beneficial in the case of rare, complicated, or confusing manifestations of OCD and / or other disorders. By using the ERP database 122 in such instances, the outputs of the underlying agents (e.g., the LLM(s)) may be improved, for example by reducing hallucinations. In the illustrated example, the ERP database 122 may include a plurality of entries organized into a plurality of symptom domains (that is, categories of symptom types). The symptom domains may include, without limitation, domains such as a perceived need to know or remember something, a checking of compulsions, concern with environmental contaminants, excessive existential concerns or doubts, excessive concern or doubt about interpersonal relationships, mental compulsions, concerns with what is right / wrong or moral, sexual obsessions, and so on. For each domain, one or a plurality of symptoms (that is, according to a clinician’s original input or interpretation). For example, the domain concerned with what is right / wrong or moral may include symptoms associated with reading news about events that are morally questionable, symptoms that are associated with completing governmental forms (such as tax returns), etc. These symptoms may be associated with external triggers, internal 21 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) triggers, and obsessions. The ERP database 122 may further include one or more ERP tasks for each symptom.

[0058] In an example, the generated patient profile may be used by the domain analysis agent 118 to determine the symptom domain for each symptom. For example, the domain analysis agent 118 may perform a similarity search on the “symptoms” portion of the ERP database 122 to generate the list of symptom domains. These symptom domains may then be the subject of a domain match by the domain analysis agent 118, referencing the ERP database 122. The top three cases for each matched domain may then be selected (although in other implementations, more or fewer cases may be selected per domain) and provided to the ERP designer agent 116 to generate the initial ERP hierarchy. Each case may include an example, including a symptom and an associated ERP task. Thus, the initial ERP hierarchy generated by the ERP designer agent may include a set of exposure levels, a set of triggers and maladaptive behavioral responses, a set of exposure tasks, and a set of response preventions. By referencing the ERP database 122, the domain analysis agent 118 and ERP designer agent 116 may be able to perform with a reduced computational load (e.g., by avoiding the need for repetitively generating outputs from scratch).

[0059] As noted above, where an interface is present, it may take any one or more of a variety of forms, including web-based, application-based, chatbot-based, and the like. FIGS. 3A-3C illustrate example UI screens for implementations where a web- or application-based interface is used. In particular, FIG.3A shows a first screen 300A corresponding to an example of the interface 102, FIG.3B shows a second screen 300B corresponding to an example of the interface 114, and FIG. 3C shows a third screen 300C corresponding to an example of the interface 134.

[0060] The first screen 300A includes a set of fields, some of which are prepopulated and some of which are fillable by the user. In the illustrated example, the first screen 300A includes a field for a clinician to enter a clinician ID. This field may be fillable in the form of free text, or may include a prepopulated list of clinicians from which the user may select her or his ID. In some implementations, the clinician ID may instead be determined by user log-in credentials or another method, such that this field may be omitted. Following the instructions, the user may then proceed in order and enter various information regarding the subject. In one 22 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) example, this information is presented in a sequence of external triggers, followed by obsessive thoughts, followed by feared consequences, followed by compulsive behaviors, followed by avoidance strategies. Information regarding internal triggers may be grouped with obsessive thoughts in this example. By asking this information up-front, the information gathering process may be improved compared to alternative implementations. For each category, the user may first select a subfield from a dropdown menu of prepopulated templates. When a template is selected from the dropdown menu, the corresponding template is added to a text box that the user may manipulate to provide details (e.g., filling in the blank). The list of prepopulated templates may be extensive, to provide highly personalized selectable templates based on expert input.

[0061] Once the user has completed entering the subject information, she or he may operate a button to submit the entered information. Upon submission, the various agents of the system (e.g., as shown in FIG.1) operate to generate an initial subject profile (also referred to as a “preliminary subject profile”). The second screen 300B of FIG.3B shows an example of how the initial subject profile is presented to the user. In the illustrated example, the initial profile includes triggers, maladaptive cognitive reactions, and maladaptive behavioral responses. For each category, additional information is provided, including information entered by the user as well as any further information (e.g., additional symptoms) that have been extrapolated by one or more of the agents. For example, under the category of maladaptive cognitive reactions, the second screen 300B presents obsessions and distorted beliefs, as well as a fear symptom that has been extrapolated by the system. The user is presented with the option to edit the initial subject profile as described above. The user may then request the system to generate a new profile, or may elect to proceed with the profile as directly edited by the user. In either case, when the user is satisfied with the profile she or he may confirm it.

[0062] Upon confirmation, the various agents of the system (e.g., as shown in FIG.1) operate to generate an initial ERP hierarchy, which may be presented to the user by a screen such as the third screen 300C of FIG. 3C. The third screen 300C shows a set of exposure therapy steps having a predetermined number (e.g., ten) and being arranged in order of increasing difficulty. Each step may include a trigger or maladaptive behavioral response that the step seeks to address, along with an exposure task and any instructions to assist in 23 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) maladaptive response prevention. Again, the user is presented with the option to edit the initial ERP hierarchy as described above. The user may then request the system to generate a new hierarchy, or may elect to proceed with the hierarchy as directly edited by the user. In either case, when the user is satisfied with the hierarchy she or he may confirm it, print it, or both. If the user chooses to print the hierarchy, it may be printed in physical form or saved as a file (e.g., a JSON file).

[0063] The above techniques may be implemented by an ERP hierarchy generating system. FIG.4 illustrates one such example system 400, including at least one processor 402, at least one memory 404, and at least one user interface 406, operatively connected to one another. While FIG.4 illustrates the processor 402, the memory 404, and the user interface 406 as being part of a single device, the system 400 may alternatively be implemented according to a distributed computing architecture. The at least one processor 402 is configured to execute instructions (e.g., in the form of software instructions) stored in the memory 404. By executing the instructions, the system 400 may then implement any of the techniques discussed above, such as the workflow 100 shown in FIG.1.

[0064] Thus, the processor 402, by executing the instructions, may perform operations including receiving an input from the user via a first component of the user interface 406 (e.g., the interface 104 shown in FIG.1), the input including a plurality of input symptoms regarding the subject, the plurality of input symptoms corresponding to a behavioral health condition; generating a profile for the subject by a first agent, wherein the first agent is configured to receive the input, extrapolate at least one missing symptom from the input, and generate the preliminary profile based on the input and the at least one missing symptom; generating a set of ERP task recommendations by a second agent, wherein the second agent is configured to receive the input, reference a database of preexisting ERP tasks, and select the set of ERP task recommendations from among the preexisting ERP tasks based on the input; generating a preliminary ERP hierarchy by a third agent, wherein the third agent is configured to receive the profile and the set of ERP task recommendations, generate a plurality of ERP task candidates, and organize at least one of the plurality of ERP task candidates into the preliminary ERP hierarchy; and presenting the preliminary ERP hierarchy to the user via a second component of the user interface 406 (e.g., the interface 134 shown in FIG.1). These operations 24 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) are shown in FIG.5. It should be understood, as noted above, that in some implementations the user may be presented with the option to approve or edit the preliminary profile. In such implementations, the processor 402 may further perform operations including presenting the preliminary profile to the user via a third component of the user interface 406 (e.g., the interface 114 shown in FIG.1).

[0065] In particular, FIG.5 shows a method 500 for generating an ERP hierarchy for the subject. The method begins with operation 502 of receiving an input from a user. The input may be received via a first interface (e.g., the first component of the user interface 406, which may be the interface 104 of FIG. 1), and may include an indication of a plurality of input symptoms regarding the subject. The plurality of input symptoms may correspond to a behavioral health condition, such as OCD or BDD.

[0066] At operation 504, the method 500 generates a subject profile using an agent (e.g., the first agent). Operation 504 may include receiving the input at the first agent, extrapolating one or more missing symptoms from the input, and generating the subject profile based on the input and the at least one missing symptom. Operation 504 may include a feedback component, for example by generating a preliminary profile by the agent, presenting it to the user (e.g., via a third interface, which may be the interface 114 of FIG. 1), and providing a manner in which the user may approve or request edits to the preliminary profile. If the user requests to edit the preliminary profile, operation 504 may be repeated to modify the preliminary profile based on the request. If the user approves the preliminary profile, the preliminary profile may be set as the subject profile and the method 500 may proceed to operation 506.

[0067] At operation 506, a set of ERP task recommendations is generated by an agent (e.g., the second agent). Operation 506 may include receiving the input, referencing a database of preexisting ERP tasks, and selecting one or more ERP task recommendation components from among the preexisting ERP tasks based on the input. For example, operation 506 may proceed as described above with regard to FIG.2.

[0068] Operation 508 includes generating a preliminary ERP hierarchy by an agent (e.g., the third agent). Thus, operation 508 may include receiving the profile and the set of ERP task recommendations by the third agent, generating a plurality of ERP task candidates, and 25 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) organizing at least one of the plurality of ERP task candidates into the preliminary ERP hierarchy. The preliminary ERP hierarchy may be presented to the user at operation 510 (e.g., via a second interface, which may be the interface 134 of FIG. 1). Operations 508 and 510 may also include a feedback component, for example by presenting the preliminary ERP hierarchy to the user and providing a manner in which the user may approve or request edits to the preliminary ERP hierarchy. If the user requests to edit the preliminary ERP hierarchy, operation 508 may be repeated to modify the preliminary ERP hierarchy and operation 510 may be repeated to present the modified preliminary ERP hierarchy to the user for approval. If the user approves the preliminary ERP hierarchy (or modified preliminary ERP hierarchy), the preliminary ERP hierarchy may be saved, printed, or presented in the form of a report.

[0069] The first, second, and third agents may implement one or more LLMs, which may be the same LLM or a plurality of different LLMs. The agents may be local to the system 400 or may be remotely located (e.g., on a cloud server), in any combination. In some examples, the agents may be distributed such that they include a frontend portion that is displayed on the user interface 406 and a backend portion that operates on a remote device. In such implementations, the system 400 may transmit various information for the operation of the LLMs to the remote device, and receive responses (e.g., outputs) from the remote device. Moreover, where multiple different LLMs are implemented, they may reside (wholly or partly) on the same remote device or across multiple different remote devices.

[0070] The user interface 406 is configured to interact with a user of the system 400, including by receiving inputs from the user and providing outputs to the user. The user interface 406 may include physical or software interfaces for one or more input or output devices, including but not limited to a keyboard, a mouse, a display, a touch screen, a microphone, a speaker, a haptic feedback device, and the like. In one example, the user interface 406 includes a graphical user interface (GUI) displayed on a display screen, thereby to implement at least one of a web page interface, an application interface, or a chatbot interface. The user interface 406 may have multiple components (e.g., multiple screens).

[0071] Other examples and uses of the disclosed technology will be apparent to those having ordinary skill in the art upon consideration of the specification and practice of the invention disclosed herein. The specification and examples given should be considered QB\125141.04855\98133476.3

Claims

Docket No. MGH 2023-216-02 (125141.04758) CLAIMS What is claimed is:

1. A method of generating an exposure and response prevention (ERP) hierarchy for a subject, the method comprising: receiving an input from a user via a first interface, the input including a plurality of input symptoms regarding the subject, the plurality of input symptoms corresponding to a behavioral health condition; generating a profile for the subject by a first agent, wherein the first agent is configured to receive the input, extrapolate at least one missing symptom from the input, and generate the profile based on the input and the at least one missing symptom; generating a set of ERP task recommendations by a second agent, wherein the second agent is configured to receive the input, reference a database of preexisting ERP tasks, and select the set of ERP task recommendations from among the preexisting ERP tasks based on the input; generating a preliminary ERP hierarchy by a third agent, wherein the third agent is configured to receive the profile and the set of ERP task recommendations, generate a plurality of ERP task candidates, and organize at least one of the plurality of ERP task candidates into the preliminary ERP hierarchy; and presenting the preliminary ERP hierarchy to the user via a second interface.

2. The method of claim 1, wherein at least one of the first agent, the second agent, or the third agent is configured to implement a large language model (LLM).

3. The method of claim 1, wherein the first interface includes at least one of a web page interface, an application interface, or a chatbot interface.

4. The method of claim 1, wherein the first interface is the same as the second interface. 28 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) 5. The method of claim 1, wherein generating the profile for the subject includes: generating a preliminary profile by the first agent; presenting the preliminary profile to the user via a third interface; in response to receiving a request to edit the preliminary profile from the user via the third interface, modifying the preliminary profile by the first agent based on the request; and in response to receiving an approval of the preliminary profile from the user via the third interface, setting the preliminary profile as the profile for the subject.

6. The method of claim 1, further comprising: in response to receiving a request to edit the preliminary ERP hierarchy from the user via the second interface, modifying the preliminary ERP hierarchy by the third agent based on the request; and in response to receiving an approval of the preliminary ERP hierarchy from the user via the second interface, setting the preliminary ERP hierarchy as the ERP hierarchy for the subject.

7. The method of claim 1, wherein the first agent includes: a missing symptom agent configured to check the input for missing symptoms; a symptom extrapolator agent configured to make inferences from the input and extrapolate the at least one missing symptom; and a profile generator agent to generate the profile based on the input and the extrapolated at least one missing symptom.

8. The method of claim 1, wherein the third agent includes: an ERP designer agent configured to generate the plurality of ERP task candidates; an imaginal reviewer agent configured to generate an imaginal exposure therapy task; an ERP scorer agent configured to assign respective scores to the plurality of ERP task candidates and the imaginal exposure therapy task; a difficulty reviewer agent configured to rank the plurality of ERP task candidates and the imaginal exposure therapy task based on the respective scores; and 29 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) an ERP selector agent configured to select a predetermined number of the plurality of ERP task candidates and the imaginal exposure therapy task based on the ranking.

9. The method of claim 1, wherein the behavioral health condition is obsessive compulsive disorder (OCD).

10. The method of claim 1, wherein the behavioral health condition is body dysmorphic disorder (BDD).

11. A system for generating an exposure and response prevention (ERP) hierarchy for a subject, the system comprising: a memory; a user interface configured to interact with a user; and at least one processor operatively connected to the memory and the user interface, the at least one processor configured to: receive an input from the user via a first component of the user interface, the input including a plurality of input symptoms regarding the subject, the plurality of input symptoms corresponding to a behavioral health condition, generate a profile for the subject by a first agent, wherein the first agent is configured to receive the input, extrapolate at least one missing symptom from the input, and generate the profile based on the input and the at least one missing symptom, generate a set of ERP task recommendations by a second agent, wherein the second agent is configured to receive the input, reference a database of preexisting ERP tasks, and select the set of ERP task recommendations from among the preexisting ERP tasks based on the input, generate a preliminary ERP hierarchy by a third agent, wherein the third agent is configured to receive the profile and the set of ERP task recommendations, generate a plurality of ERP task candidates, and organize at least one of the plurality of ERP task candidates into the preliminary ERP hierarchy, and present the preliminary ERP hierarchy to the user via a second component of the user interface. 30 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) 12. The system of claim 11, wherein at least one of the first agent, the second agent, or the third agent is configured to implement a large language model (LLM).

13. The system of claim 12, wherein the at least one processor is configured to at least one of: generate the profile for the subject by transmitting the input to a first device implementing the first agent and receiving the profile from the first device; generate the set of ERP task recommendations by transmitting the input to a second device implementing the second agent and receiving the set of ERP task recommendations from the second device; or generate the preliminary ERP hierarchy by transmitting the profile and the set of ERP task recommendations to a third device implementing the third agent and receiving the preliminary ERP hierarchy from the third device.

14. The system of claim 13, wherein the first device, the second device, and the third device are the same device.

15. The system of claim 11, wherein the first component of the user interface includes at least one of a web page interface, an application interface, or a chatbot interface.

16. The system of claim 11, wherein the second component of the user interface includes at least one of a web page interface, an application interface, or a chatbot interface.

17. The system of claim 11, wherein the at least one processor is configured to generate the profile by: generating a preliminary profile by the first agent; presenting the preliminary profile to the user via a third component of the user interface; in response to receiving a request to edit the preliminary profile from the user via the third component of the user interface, modifying the preliminary profile by the first agent based on the request; and 31 QB\125141.04855\98133476.3Docket No. MGH 2023-216-02 (125141.04758) in response to receiving an approval of the preliminary profile from the user via the third component of the user interface, setting the preliminary profile as the profile for the subject.

18. The system of claim 11, wherein the at least one processor is further configured to: in response to receiving a request to edit the preliminary ERP hierarchy from the user via the second component of the user interface, modify the preliminary ERP hierarchy by the third agent based on the request; and in response to receiving an approval of the preliminary ERP hierarchy from the user via the third interface, set the preliminary ERP hierarchy as the ERP hierarchy for the subject.

19. The system of claim 11, wherein the first agent includes: a missing symptom agent configured to check the input for missing symptoms; a symptom extrapolator agent configured to make inferences from the input and extrapolate the at least one missing symptom; and a profile generator agent to generate the profile based on the input and the extrapolated at least one missing symptom.

20. The system of claim 11, wherein the third agent includes: an ERP designer agent configured to generate the plurality of ERP task candidates; an imaginal reviewer agent configured to generate an imaginal exposure therapy task; an ERP scorer agent configured to assign respective scores to the plurality of ERP task candidates and the imaginal exposure therapy task; a difficulty reviewer agent configured to rank the plurality of ERP task candidates and the imaginal exposure therapy task based on the respective scores; and an ERP selector agent configured to select a predetermined number of the plurality of ERP task candidates and the imaginal exposure therapy task based on the ranking. 32 QB\125141.04855\98133476.3

Citation Information

Patent Citations

  • Medical risk factors evaluation

    US20180101657A1

  • Mental health response system and method

    US20230128090A1

  • Artificial intelligence and machine learning techniques using input from mobile computing devices to diagnose medical issues

    US20230411008A1