Follow-up question generation for enhanced patient-provider conversations

WO2026183197A1PCT designated stage Publication Date: 2026-09-03TRUSTEES OF DARTMOUTH COLLEGE THE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/016615
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-26
Filing Date
2026-02-25
Publication Date
2026-09-03

Smart Images

  • Figure US2026016615_03092026_PF_FP_ABST
    Figure US2026016615_03092026_PF_FP_ABST
Patent Text Reader

Abstract

A method of automatic follow-up question generation includes accessing, using a computerized system, a message received from a patient using a medical dialogue interface. The method further includes obtaining, using the computerized system, electronic health record (EHR) information corresponding to the patient. The method includes providing, using the computerized system, the message and the EHR information to a multi-agent model comprising an EHR reasoning agent, a differential diagnostic agent, and a message clarification agent. The method includes creating, using the multi-agent model and the computerized system, a set of questions, wherein each question in the set of questions responds to the message. The method includes assigning, using the computerized system, a quality value to each question in the set of questions. The method includes outputting, via the medical dialogue interface, a question from the set of questions assigned a highest quality value.
Need to check novelty before this filing date? Find Prior Art

Description

FOLLOW-UP QUESTION GENERATION FOR ENHANCED PATIENT-PROVIDER CONVERSATIONSCROSS-REFERENCE TO RELATED APPLICATION(S)

[0001] This application claims priority to U. S. Provisional Patent Application Serial No.63 / 763,845, filed February 26, 2025, the content of which is hereby incorporated by reference in its entirety.BACKGROUND

[0003] Asking relevant, useful follow-up questions while conversing fosters deeper understanding, and ensures meaningful and productive conversations. Generating relevant and meaningful follow-up questions can involve collecting relevant information that is fragmented across multiple sources. For example, in patient provider communication, providers may consider patient utterances while attending to information scattered throughout the patient’s electronic health record (EHR), as well as consider numerous different thought processes while formulating a patient’s diagnosis. Similarly, in customer service interactions, support agents may integrate information from a customer’s message while referencing their past interactions, purchase history or account details, as well as simultaneously consider multiple resolution strategies, such as troubleshooting steps, redirection to other support or refund policies.SUMMARY

[0004] This summary' is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter.

[0005] Aspects of the present disclosure relate to automatic follow-up question generation in medical dialogue systems. In one aspect, a method includes accessing, using a computerized system, a message received from a patient using a medical dialogue interface. The method further includes obtaining, using the computerized system, electronic health record (EHR) information corresponding to the patient. The method includes providing, using the computerized system, the message and the EHR information to a multi-agent model comprising an EHR reasoning agent, a differential diagnostic agent, and a message clarification agent. The method includes creating, using the multi-agent model and the computerized system, a set ofQB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)2questions, wherein each question in the set of questions responds to the message The method includes assigning, using the computerized system, a quality value to each question in the set of questions. The method includes outputting, via the medical dialogue interface, a question from the set of questions assigned a highest quality value.

[0006] In another aspect, a non-transitory computer-readable storage medium stores instructions that, when executed by a processor, cause the processor to perform a method of automatic follow-up question generation. The method includes accessing, using a computerized system, a message received from a patient using a medical dialogue interface. The method further includes obtaining, using the computerized system, electronic health record (EHR) information corresponding to the patient. The method includes providing, using the computerized system, the message and the EHR information to a multi-agent model comprising an EHR reasoning agent, a differential diagnostic agent, and a message clarification agent. The method includes creating, using the multi-agent model and the computerized system, a set of questions, wherein each question in the set of questions responds to the message. The method includes assigning, using the computerized system, a quality value to each question in the set of questions. The method includes outputting, via the medical dialogue interface, a question from the set of questions assigned a highest quality value.

[0007] The foregoing general description of the illustrative embodiments and the following detailed description thereof are merely exemplary aspects of the teachings of this disclosure and are not restrictive.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] FIG 1 is a block diagram conceptually illustrating a system for automatic medical question generation and evaluation, according to some embodiments.

[0009] FIG. 2 is a flow diagram illustrating an example process for follow-up question generation and selection, according to some embodiments.

[0010] FIGS. 3A and 3B illustrate an example overview of a follow-up question generation system, according to some embodiments.

[0011] FIG. 4 is a flow diagram illustrating an example process for training a medical question enhancing agent based on clinician-annotated data, according to some embodiments.[00121 FIG. 5 is an example diagram of a patient message response drafted by a multi-agent model and edited by a clinician, according to some embodiments.QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)3

[0013] FIGS 6A and 6B illustrate an example overview of an LLM-as-judge frame for evaluating clinician feedback, according to some embodiments.

[0014] FIG. 7 is a flow diagram illustrating an example process for sorting patient messages, according to some embodiments.

[0015] FIG. 8 is a table comparing a default inbox with an urgency-aware inbox, according to some embodiments.

[0016] FIG. 9 is a graph illustrating example metric results for agent follow-up question generation, according to some embodiments,

[0017] FIG. 10 illustrates example message data formats from two data resources, according to some embodiments.DETAILED DESCRIPTION

[0018] The detailed description set forth below in connection with the appended drawings is intended as a description of various configurations and is not intended to represent the only configurations in which the subject matter described herein may be practiced. The detailed description includes specific details to provide a thorough understanding of various embodiments of the present disclosure. However, it will be apparent to those skilled in the art that the various features, concepts, and embodiments described herein may be implemented and practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid obscuring such concepts.

[0019] The systems and methods described herein streamline asynchronous medical conversations between clinicians and patients by reducing time spent by clinicians reviewing client, information and previous messages in order to formulate follow-up questions. In particular, the systems and methods described herein are directed to multiagent framework designed to automatically generate follow-up questions in asynchronous medical conversations. Moreover, the systems and methods provided herein may triage patient messages received to reduce time spent by clinicians sorting through multiple messages.

[0020] FIG. 1 shows a block diagram illustrating an example system for automatic medical question generation (i.e., question generation system 100), according to some embodiments. The question generation system 100 may further be used to determine an urgency associated with one or more messages received from patients. In some examples, a computing device 110 can obtain or receive electronic health record (EHR) materialsQB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)4102 and / or messages 104 via the communication network 130 to generate items, such as follow-up question designed to elicit relevant information from an individual to make diagnostic hypotheses. For example, the EHR materials 102 may include test results, vaccination records, family medical history, past and present medical conditions, demographics, medications, or the like. Moreover, in some examples, the messages 104 may correspond to conversations between one or more medical practitioners and a patient using a dialogue system. In further examples, the messages 104 may further correspond to automatically generated messages (or questions) produced by the computing device 110

[0021] In some examples, the computing device 110 can include a processor 112. In some embodiments, the processor 112 can be any suitable hardware processor or combination of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), a microcontroller (MCU), etc.

[0022] In further examples, the computing device 110 can further include a memory 114. The memory 114 can include any suitable storage device or devices that can be used to store suitable data (e.g., EHR materials 102, one or more generated questions, etc.) and instructions that can be used, for example, by the processor 112 to receive documents, files, or materials. The memory 114 can include any suitable volatile memory, non-volatile memory, storage, or any suitable combination thereof. For example, memory 114 can include random access memory (RAM), read-only memory (ROM), electronically- erasable programmable read-only memory (EEPROM), other forms of volatile memory, other forms of non-volatile memoiy, one or more forms of semi-volatile memoiy, one or more flash drives, one or more hard disks, one or more solid state drives, one or more optical drives, etc. In some embodiments, the processor 112 can execute at least a portion of process 200 described below in connection with FIG. 2.

[0023] In further examples, computing device 110 can further include communications system 118. Communications system 118 can include any suitable hardware, firmware, and / or software for communicating information over communication network 130 and / or any other suitable communication networks. For example, communications system 118 can include one or more transceivers, one or more communication chips and / or chip sets, etc. In a more particular example, communications system 118 can include hardware, firmware and / or software that can be used to establish a Wi-Fi connection, a Bluetooth connection, a cellular connection, an Ethernet connection, etc.QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)5

[0024] In further examples, computing device 110 can receive or transmit information (e.g., EHR materials 102, one or more generated messages 104, etc.) and / or any other suitable system over a communication network 130. In some examples, the communication network 130 can be any suitable communication network or combination of communication networks. For example, the communication network 130 can include a Wi¬ Fi network (which can include one or more wireless routers, one or more switches, etc.), a peer-to-peer network (e.g., a Bluetooth network), a cellular network (e.g., a 3G network, a 4G network, a 5G network, etc, complying with any suitable standard, such as CDMA, GSM, LTE, LTE Advanced, NR, etc.), a wired network, etc. In some embodiments, communication network 130 can be a local area network, a wide area network, a public network (e.g., the Internet), a private or semi-private network (e.g., a corporate or university intranet), any other suitable type of network, or any suitable combination of networks. Communications links shown in FIG. 1 can each be any suitable communications link or combination of communications links, such as wired links, fiber optic links, Wi-Fi links, Bluetooth links, cellular links, etc.

[0025] In further examples, computing device 110 can further include a display 116 and / or one or more inputs 120. In some embodiments, the display 116 can include any suitable display devices, such as a computer monitor, a touchscreen, a television, an infotainment screen, etc. to display messages 104 or diagnostic hypothesis to a medical practitioner. In further embodiments, and / or the input(s) 120 can include any suitable input devices (e.g., a keyboard, a mouse, a touchscreen, a microphone, etc.).

[0026] In an example, the computing device 110 can operate as a standalone device or the computing device 110 can be connected (e.g., networked) to other machines (e.g, via the communication network 130). In a networked deployment, the computing device 110 can operate in the capacity of either a server or a client machine in server-client network environments. In an example, computing device 110 can act as a peer machine in peer-to-peer (or other distributed) network environments. The computing device 110 can be a personal computer (PC). a tablet PC, a set-top box ( STB), a Personal Digital Assistant (PDA), a mobile telephone, a web appliance, a network router, switch or bridge, or any machine capable of executing instructions (sequential or otherwise) specifying actions to be taken (e.g., performed) by the computing device 110. Further, while only a single computing device 110 is illustrated, the computing device 110 may also include any collection of computing devices thatQB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)6individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods or processes described in the present disclosure.

[0027] Moreover, the question generation system 100 may include one or more agents used to produce follow-up questions and / or sort messages received from patients, such as EHR reasoning agent(s) 152, differential diagnostic agent(s) 154, message clarification agent(s) 156, a response enhancing agent 158, and a patient message ranking (PMR) agent 160. These agents may work together to create one or more questions for a comprehensive dialogue with a patient, which can enable efficient diagnoses and evaluation of a patient's concerns.

[0028] In some examples, the EHR reasoning agent(s) 152 may include one or more agents used to mitigate challenges associated with generating questions from fragmented data sources. As described above, EHR materials 102 may include information related to a patient’s current inquiry can be implicit or fragmented across different data tables, and fields. For a given EHR record C = {A. H, M}, the EHR reasoning agent 152 may include a medical history reasoning agent and a medication list reasoning agent. Each agent can first extracts relevant pieces of EHR information concerning the patient’s current inquiry, T. In some examples, these portions of a patient’s medical history' and medication list can contain information that is not relevant to their current message. This may be defined using the following equations: lhist=f{A, H, PextractH, T) and / med= f(M, PextractM, '0, where lhistand lmedare the elements from the patient's medical history and medication list most relevant to the patient’s message. PextractHanc' PextractMaretopic-specific information extraction prompts for history and medications. The EHR reasoning agent(s) 152 may then generate EHR-specific follow-up questions QEHR= QhistU Qmedas follows: Qhist= f(T, lhist, Phist, k) and Qmed= f(T med> Pmed, )- Here,and Pmedguide f on generating questions related to lhistand lmed, respectively.

[0029] In some examples, providers who read and respond to patient messages often mentally perform a differential diagnosis before coming up with the follow-up questions. Specifically, providers can decide what could be wrong with the patient and ask questions to rule out various diagnostic hypotheses Inspired by this mental framework, the different diagnostic agent(s) 154 may generate a set of possible patient diagnoses Dd^^, then generate follow-up questions based on these potential diagnoses, i.e. follow-up questions to rule out each diagnosis dte Ddtff- Differential diagnostic agent(s) 154 may first compute a pseudo-differential diagnosis by identifying the k best and worst-case diagnoses for a given patient, using prompts Pbestand Pworstrespectively: Ddiff — f ( ’ Pbest, k) f T, Pworst, k). Thus Ddtff1Sl^e unio of theQB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)7potential diagnoses produced by funder the assumptions of the best and worst case scenarios. This strategy is motivated by the following observations: providers cast a wide net for gathering relevant information.

[0030] Next differential diagnostic agents 154 iteratively build the question set Qndiff~ Qd... U.... which has targeted questions to rule out each diagnosisfy G F°r agiven possible diagnosis d{, Qd. is computed as follows. Qdi= f (T, di, Prule..mit, k where f outputs a question set of size k using the prompt Pruie-out to guide questions to see if the patient is suffering from potential diagnosis cty By taking the union of question sets about each potential diagnosis, differential diagnostic agents produce the set of questions Qodiff-

[0031] In some examples, the message clarification agent(s) 156 may be used to increase the clarity' of different aspects of the patient’s message. For example, the message clarification agent(s) 156 may include a symptom inquiry' agent that extract symptoms from the message and ask clarifying questions about each symptom as needed (e.g., location of abdominal pain). The message clarification agent(s) 156 can further include self-treatment agents that ask patients to elaborate on how they are treating their symptoms (e.g. patients may' self-treat with over-the-counter pain medications or herbal supplements). The message clarification agent(s) 1.56 can also include a temporal reasoning agent that generates questions to increase clarity in the timeline of presented symptoms, (e.g., duration and frequency of pain). Message ambiguity agents may target reducing overall ambiguity of the message (e.g. “tell me more about what you mean by you are feeling off"). For each message clarification agent 156. clar^ Qctari~ f(T> clarf k), where Pctariis a prompt specific to clarification agent clat. Taking the union over one or more clarification agents generates a question set ''Qciar-

[0032] A final question pool Qpmay include questions generated by the differential diagnostic agents 154 and message clarification agents 156: Qp— {Q. >diff, Q.£HR, Q.ciar}- The value k, which controls the number of questions generated, is specific to each agent. As described herein, the number of questions generated by each agent is controlled to retain granular control of the output size. However, it should be noted that flexible variations, such as generating up-to-k questions (or no constraints) may also be used.

[0033] In some examples, the response enhancing agent 158 may' use drafts produced by agents 152, 154, and 156 in supporting clinician responses to patient messages, by evaluating alignment of the agents 152, 154, and 156 to responses generated by real clinicians. Specifically, the response enhancing agent 158 may evaluate content-level and theme-levelQB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)8alignments between clinician-written and the multi-agent-generated responses, to inform responsible use of NLP in patient message response drafting. In particular, the response enhancing agent 158 may evaluate the following questions: What constitutes a high-quality clinician response to a patient message9How might the evaluation of multi -agent-generated responses draft quality be automated, with respect to clinician editing workload? How can agents 152, 154, and 156 be adapted to support clinicians in generating quality responses to patient messages?

[0034] In answering these research questions, a clinically relevant set of “themes” and frames is created to systematically characterize clinician responses to patient messages. Moreover, a two-level evaluation framework for assessing clinician editing load given multi-agent-drafted responses to patient messages. In some examples, an expert-clinician-annotated dataset may be used to evaluate performance on the patient message response drafting task, as described below with respect to FIG. 4.

[0035] FIG 2 is a flow diagram illustrating an example process 200 for follow-up question generation and selection, according to some embodiments. As described below, a particular implementation can omit some or all illustrated features / steps, may be implemented in some embodiments in a different order, and may not require some illustrated features to implement all embodiments. In some examples, an apparatus (e.g., computing device 110, processor 112 with memory 114, etc.) in connection with FIG. 1 (described above) can be used to perform example process 200. However, it should be appreciated that any suitable apparatus or means for carrying out the operations or features described below may perform process 200.

[0036] At step 202, the process 200 accesses one or more messages received from a patient. In some examples, the one or more messages may be received using a medical dialogue interface. The medical dialogue interface may be a portal, a mobile application, a web interface, or the like that allows a patient to interact with a clinician. The messages may describe symptoms, concerns, or question, as well as attach photos, test results, and other test information.

[0037] At step 204, the process 200 obtains EHR information corresponding to the patient. In some examples, a patient's EHR record may be defined as C ------ {A -I. M}. This includes a patient’s demographics A (e.g., age and gender), medical history H (e.g,, problem list and recent medical encounters), and medication list M. Each component of C is represented as a string in the framework.

[0038] At step 206, the process 200 generates a set of questions corresponding to the one or more received messages. In some examples, for a patient message T, corresponding EHR C,QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)9and a text generator f (e.g., a multi-agent-based framework) may be represented as: f(T, C) = Q — ■■■ > n}- The text generation system may produce a set Q, where each q, e Q is a follow-up question to the patient's message. In some examples, f may exist in an asynchronous environment without real-time access to the patient and may generate all pertinent follow-up questions as a list. An example summary of this framework is illustrated in FIGS, 3A and 3B, In some examples, three thought processes were considered for the asynchronous follow-up question generation process: (i) EHR Reasoning: providers obtaining contextual knowledge from fragmented EHR data to guide their inquiry, (11) Differential Diagnostics: providers developing a mental list of potential diagnoses or explanations that describe the patient’s symptoms, guiding their question formulation, and (iii) Message Clarifications: providers asking a series of questions to fill in gaps in the patient's reported symptoms.

[0039] At step 208, the process 200 assigns a quality value to each question in the set of questions. For example, in the context of use cases of asynchronous medical dialogue, the quality of the set of questions Q — [q,..-. qn produced by / may be determined by the reduction in dialogue turns used by the provider to make a diagnosis or recommendation. To make a decision, providers may first collect information from the patient To reduce provider workload, patient responses to questions in Q may contain at least the information requested in the ground truth question set Q. Thus, metrics for comparing generated questions Q against ground truth questions Q are designed to identify how well the information requested in Q is covered by the information requested in Q. Moreover, a generated questionE Q is considered as matching a ground truth question qj E Q when the same information is required (i.e., invokes similar responses), m addition to when they match exactly. For example, an agent asking “Do you have a cough or fever?” may elicit a similar response to the doctor’s question “Have you been coughing?” and thus should be considered a match. In some examples, the response enhancing agent 158 may perform this semantic matching, as described below with respect to FIG. 4.

[0040] In some examples, the quality value may correspond to a calculated Requested Information Match (RIM) metric represented by: RIM(Q, Q) =““• RIM measures thesample-wise percentage of provider-generated questions that are also generated by system f. RIM is a task-specific case of the T versky Index. RIM does not penalize based on the size of the set of generated questions, |Q|. For example, if only three questions are asked to a patient,QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)10it does not mean there are no other useful inquiries to generate. A different list of questions may be generated based on a provider’s experience, preferences, and current mental models, resulting in subjectivity in the ground truth. Thus, requesting additional information helps a provider who may have forgotten to consider a certain outlying issue the patient may be facing, and address the subjectivity of their thought process.

[0041] Subsequently, to generate asynchronous follow-up questions, coverage of ground truth questions is optimized by maximizing RIM and by controlling the size of the list of generated questions, Q|. A RIM score of 1.0 may denote a case where a system (i.e.. system 100) requested information also requested by the patient’s provider. For each output with an RIM score of 1.0, / has reduces the number of clarification requests a provider needs to send by 1, assuming the patient responds to all questions. Conversely, while RIM scores below 1.0 are still suggestive of message improvement, they may not suggest any reduction in outgoing messages to be sent by a provider.

[0042] In some examples, each question in the set of questions may further be assigned a second metric - Message Reduction%(MR%), which measures the percentage of samples where the RIM score = 1.0. Models with a higher MR% correspond to greater workload reductions.

[0043] At step 210. a question assigned the highest quality value is output. As described above, Q questions may be generated based on hyperparameter choices. In some examples, Qpmay produce a Qptoo large for a patient to answer. Therefore, the size of Qpmay be reduced to a target size k based on the assigned quality7values. Moreover, question de-duplication and top-k question selection may be performed. For example, the question de-duplication framework may use LLMs or agents to filter non-unique inquiries fromIn some examples, different agents may request the same information in the context of different thought processes. The top-k selection may take the de-dupli cated question list and select the k highest quality7questions to output.

[0044] FIG. 4 is a flow diagram illustrating an example process 400 for training a medical question enhancing agent based on clinician-annotated data, according to some embodiments. As described below, a particular implementation can omit some or all illustrated features / steps, may be implemented in some embodiments in a different order, and may not require some illustrated features to implement all embodiments. In some examples, an apparatus (e.g., computing device 110, processor 112 with memory 114, etc.) in connection with FIG. 1 (described above) can be used to perform example process 400. However, it should beQB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)11appreciated that any suitable apparatus or means for cany mg out the operations or features described below may perform process 400.[0045 At operation 402, the process 400 receives an annotated question corresponding to a generated follow-up question. In some examples, the generated follow-up questions are questions generation in process 200 using the electronic health record reasoning agent(s) 152, the differential diagnostic agent(s) 154. and the message clarification agent(s) 156. In particular, the generated follow-up questions may be provided to one or more clinicians for review,

[0046] FIG. 5 illustrates an example diagram of patient message response drafting by an LLM or multi-agent model and editing of responses by a healthcare provider. The multi-agent model may draft responses to patient messages, then clinicians edit the draft by deleting and adding content as needed,

[0047] At operation 404, the process 400 identifies deletions, additions, and matches between the annotated question and the generated follow-up question. Referring again to FIG. 5, the provider response deletions are indicated as strikethroughs, while the italicized text indicates additions made by the healthcare provider.

[0048] At operation 406, the process 400 identifies a theme of the annotated question and the generated follow-up question. In some examples, the theme may be used to characterize a quality of clinician responses to patient messages. Table 1 below lists several themes that may be used.

[0049] Table 1: Themes derived from clinician responses to patient portal messages, alongside representati e frames and example response elements / utterances. For example, "‘explanation of test result"’ is a frame within the medical assessment theme, and '‘your iron levels look normal"’ is a clinician response component that falls under this frame. In one example, 8 clinician response themes comprised of 67 unique frames were derived.Theme Example Frame Example Response Element Empathy Encouragement of You've been doing a patient treatment effort great job with your tapering.Symptom Question Asking about location Has your pain only been of symptoms in your lower back?QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)12Medication Question Asking about intake of Have you been taking medications your Amoxicillin regularly? Medical Assessment Explanation of test Your iron levels look result normal.Medical Planning Confirmation of Let’s get you in for a required testing bloodwork test.Logistics Confirmation of clinic You can only offer policy telehealth in the state.Care Coordination Promise of future We’ll reach out after we patient contact receive the results.Contingency Planning Symptom-related If you’re feeling dizzy,backup plan please call triage.

[0050] At operation 408, the process 400 trains the response enhancing agent based on the identified deletions, additions, matches, and themes. In some examples, the response enhancing agent may first be provided with one or more prompts that describe each theme to provide associated LLM(s) of the agent with context. Moreover, the enhancing agent may be instructed identify predicted annotations of newly multi-agent generated questions, based on the identified deletions, additions, matches, and themes provided to the enhancing agent of clinician-annotated questions. For example, the agent may be instructed to answer questions such as: 1) how much content would the clinician add to the multi-agent drafted question? and 2) how much content would the clinician remove from the LLM draft? Hence, by providing the response enhancing agent 158 the identified deletions, additions, matches, and themes from operations 402, 404, and 406, the response enhancing agent may use a reference-based approach to directly compare the multi-agent generated questions with a question by an expert clinician. Comparing what needs to be removed from and added to an LLM-drafted response to achieve an expert- written response, is analogous to measuring 1) recall (i.e., how much of the expert-written response is covered by the multi-agent-drafted question), and 2) precision (i.e., how much of the LLM-drafted response is matched in the clinician's response). An overview' of this “EditJudge Evaluation Framework” is illustrated in FIGS. 6A and 6B. This framework is a human-AI collaborative, task-specific, reference- based, LLM-as-judge evaluation framework.QB\178981.0003S\101030782.1Docket No. (78981.00038 (2025-020, 2025-011-001)13

[0051] In some examples, two measures of editing load are determined by the response enhancing agent to capture complementary aspects of alignment between generated and reference responses. The content-level edit- Fi score assesses whether a response drafting LLM reproduces specific clinical facts, instructions, or action items present in the reference. However, clinically appropriate drafts may win. substantially in wording or level of detail while addressing the same underlying intent. The theme-level edit- Fi score captures higher-level alignment by measuring whether the response addresses thematically similar clinical goals, concerns, and communicative functions (e.g.. reassurance, triage guidance, or follow-up planning), even when the granular content differs. Using both metrics distinguishes incomplete response drafts from those that are semantically (content-level) and thematically aligned but phrased differently, providing a more reliable evaluation of response draft quality.

[0052] Given an expert-written clinician response re and an LLM response draft rd, the contentlevel edit- Fi score aims to identify how many expected additions (EA and expected deletions (ED) are needed from the clinician, in order to unify rd with re. Matching content m the response draft rd is referred to as an expected match (EM), meaning it would not be expected for the clinician to have to rewrite that content in order to achieve their desired response re, saving the clinician time and achieving reliability via multi-agent response drafting.

[0053] In some examples, the response enhancing agent may be framed using an algorithm for counting EA, ED, and EM. This algorithm splits an expert-writen response reinto atomic elements (e.g., sentences), then for each element uses a fine-tuned judge LLM (e.g., a content¬ level edit-judge of the response enhancing agent) to either identify expected matches EMin the response draft rd, or expected additions EA to the response draft to achieve that element. The content-level edit judge takes as input a sentence from the clinician-written response seand the multi-agent drafted response (i.e., questions) rd, and outputs either the matching content from the LLM-drafted response Sd, or “NO MATCH” if there is no matching content. Finally, this algorithm identifies expected deletions ED in the response draft by quantifying the remaining amount of unmatched content By treating expected matches, expected additions, and expected deletions as true positives, false negatives, and false positives respectively, recall is calculated (i.e., the percentage of the expert-written response rewhich does not need to be added to rd, and precision is calculated (i.e., the percentage of the response draft rd, which does not need to be removed). In one examples, the harmonic mean of the content-level 261 recall and precision scores (i.e. Fi) were calculated and called the content-level edit- Fi score AssumingQB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)14additions and deletions are evenly-weighted, content-level edit- Fi go es the expected reduction in editing load for the clinician by using the LEM response draft.[0054 Moreover, given a clinician response re, and a multi-agent response draft ra, the themelevel edit- Fi score can identity the higher-level themes in the clinician response rewhich are correctly matched by the themes in the multi-agent response draft ra. To identify themes in each response, the theme-level editJudge model of the response enhancing agent is used. Given a sentence from either the clinician response se<= reor the multi-agent drafted response sa e ra, the theme-level editJudge model assigns a theme label Is. Predicting clinician response themes is a 9-class multi-label classification task, as there are 8 high-level themes (see Table 1) and an ‘'Other” class to capture miscellaneous themes not captured in the mam 8 classes. Using the theme labels ha assigned to sentences sa from the multi-agent drafted response ra as predictions for the theme labels he assigned to sentences sefrom the clinician response re, the theme-level edit- Fi score is the micro average Fi of theme predictions.

[0055] FIG. 7 is a flow diagram illustrating an example process 700 for sorting patient messages, according to some embodiments. As described below, a particular implementation can omit some or all illustrated features / steps. may be implemented in some embodiments in a different order, and may not require some illustrated features to implement all embodiments. In some examples, an apparatus (e.g., computing device 110, processor 112 with memory 114, etc.) in connection with FIG. 1 (described above) can be used to perform example process 700. However, it should be appreciated that any suitable apparatus or means for carrying out the operations or features described below may' perform process 700.

[0056] At operation 702, the process 700 receives a set of patient messages. In some examples, the patient messages may be received from one or more data sources. As described further below, m one example, the patient messages may be obtained from an open forum, a repository of messages from regional hospitals, and / or example messages provided by' one or more clinicians.

[0057] For a set of patient messages P = {p1;..., pn, the patient message ranking (PMR) agent 160 may learn the optimal ordering P ' such that messages with a higher degree of medical urgency are ranked higher in P The resulting sort enables patients with greater medical needs to receive clinicians’ attention sooner. As illustrated in FIG. 8, the PMR agent 160 may' produce an ‘Urgency Aware’ inbox that is sorted by' medical urgency.

[0058] At operation 704, the process 700 obtains electronic health record (EHR) information corresponding to the patient. Each p, e P may contain an associated structured EHR record, et.QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)15As described above, the EHR records may contain relevant patient details such as age, gender, medication list, diagnosis history, and active problem list. Such context is often taken into consideration when patient messages are reviewed, making PMR a multi-modal inference problem.

[0059] At operation 706, the process 700 assigns each message an urgency score based on a plurality of pre-defined urgency levels and the EHR information. In some examples, the responses may be placed into an ordinal set of urgency categories. For example, a 6-tier ordinal scale for labeling urgency in patient portal messages may be used. The scale may begin at Level 1 (Most Urgent - Emergency Attention Needed) and go to Level 6 (Least Urgent - No Medical Attention Needed).

[0060] Moreover, for a sample (q, r) where q is a patient message and r is a response from a clinical expert, an LLM g (associated with the PMR agent 160) may classify the response r into the above scale to determine the degree of urgency of the message. For example, patients instructed to go to the ED are classified as '‘Level 1" while patients given self-care strategies are classified as " Level 5".

[0061] At operation 708, the process 700 creates a plurality of message pairs from the set of patient messages based on the urgency scores. In some examples, pairwise annotation is created for two messages (<,-, ty) based on the rel tive ranking of their respective responses (g(r;). g(i ) In some examples, sample quality' filtering and judge models are utilized to validate the label accuracy of each pair of messages. Moreover, P may be sorted via pairwise comparisons across patients.

[0062] At operation 710, the process 700 compares messages contained in each message pair to determine what message is more urgent. In some examples, any two patient messages (ty, p) may be fed into a model f whose job is to determine which of the tw o patient messages should be attended to first. In particular, the model / may utilized urgency labels that can be converted from the ordinal urgency scale into relevancy scores to map the problem into an information retrieval setting. For example. Level 1 samples may have the highest relevancy scores, indicating that the corresponding patient messages should be attended to before low?er urgency patient messages.

[0063] At operation 712, the process 700 increases an urgency score of the more urgent 2 j pairwisecomparisons are computed. Each time a sample is deemed more urgent than another, its scoreQB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)16is incremented by (1 +; / ). where / / is the difference in normalized probabilities and reward scores.[00641 At operation 714. the process 700 sorts the set of patient messages based on the urgency score. In some examples, the set of patient messages in a clinician's inbox may be sorted based on the total score of each sample, as described above with respect to operation 712. In further examples, the clinician’s inbox may also be re-sorted using / as the comparator in a sorting / n\algorithm or to compute one or more [ J comparisons and sort messages based on their "winrate."Follow-Up Question Generation Examples and Experiments

[0065] The inventors evaluated the methods described herein using a benchmark FollowupBench (FB) containing two asynchronous Portal Message datasets: FB-Real and FB-Synth, which are described in Table 1.

[0066] FB-Real consists of real messages and EHR records sent from adult patients to their providers between January 2020 and June 2024 at a large university' medical center in the United States. From a corpus of over 500k messages, the inventors performed a multistep filtering process, including extensive human evaluation, to select messages that ensure each patient in the dataset is symptomatic and received a provider response containing follow-up questions. Human reviewers ensure that the clarification questions included in the ground truth are in scope for an Al system to generate (i.e. are grounded to the message or chart and not to prior m-person patient-provider interaction). Additionally, human reviewers ensure the questions are specific to symptom clarifications (i.e. not aimed to generate logistical questions related to scheduling, insurance coverage, or medication refills). The extracted ground truth questions 'ere broken down into single topic questions to promote granular evaluation of NLP systems (e.g " Do you have any fever or cough?" is converted to the following two questions: "1. Do you har e any fever? 2. Do you have any cough?"). As the human review process is expensive and extremely time-consuming, FB-Real was limited to 150 unique patient messages and the corresponding EHR data of those 150 patients. As FB-Real contains protected health information, it cannot be shared publicly.

[0067] Table 1: Dataset statistics for FollowupBench (FB). FB-Real has fewer questions on average as they were composted in alive, time-constrained work environment.Dataset # of Messages Total / Mean # of Mean # ofQuestions SentencesQB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)17FB-Real 150 514 / 3.4 5.3FB-Synth 250 2,336 / 9.3 6.5

[0068] FB-Synth is a semi-synthetic dataset consisting of 250 (medical chart, patient message) pairs with over 2,300 follow-up questions written by a team of 9 physicians, nurses, and medical residents at a large university medical center. This patient data was created by first sampling a random, deidentified medical chart from the corpus of patients used to create FB-Real. Then, a synthetic message was created using a grounded message generation technique. FB-Synth has a higher mean number of questions per sample compared to FB-Real (9.3 vs 3.4). This may be due to various factors, including (a) research subjects often act differently under observation, known commonly as the Hawthorne Effect (b) providers may request more information when they have more time to reflect on a given patient's needs. The latter point is suggestive of FB-Synth potentially being closer to the target set of questions an Al system should strive to generate.

[0069] Unbounded-Generation: An LLM with both patient message and EHR data was provided, and prompted to write as many questions as necessary to clarify ambiguity and seek missing details. Both unbounded generation in the context of both O-shot and few-shot prompting were evaluated.

[0070] L-Questi on-Generation: Tire same prompt as unbounded-generation, but with specific guidance to output k questions. This allowed for the exploration of performance correlation with the scale of the generated set. For large values of k, this experiment explores the intrinsic limits of LLMs to reach longtail questions, k = 40 was used in this study as the mean number of questions output by the FollowupQ framework before filtration is ~ 35.

[0071] Long-Thought Generation: The performance of models that generate follow-up questions was evaluated after significant Chain-of-Thought output. This baseline is explicitly instructed to generate as many questions as it sees fit.

[0072] Due to the sensitive nature of FB-Real, experiments were performed on a secure computing cluster with no access to the internet and a single NVIDIA A40 GPU. In one examples, the experiments were performed using 4-bit quantized versions of the following models due to computational restraints: (i) Llama3-8b, (li) Llama3-8b Aloe, a Llama3 variant trained on healthcare data, (iii) Qwen2.5-32b-instruct, and (iv) Qwen2.5-32b distillation of DeepSeek Rl, for long-thought baseline.QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)

[0073] Table 2: The Mean RIM / Mean Number of Questions Asked for each experiment on FB-Real.Llama3-8b Llama3-8b-Aloe Q wen-32bZero-Shot-U 0.35 / 11 0.35 / 17 0.35 / 10Few-Shot-U 0.29 / 9 0.33 / 19 0.36 / 10Zero-Shot-40 0.40 / 30 0.40 / 30 0.41 / 35Few- Shot-40 037 / 35 0.36 / 35 045 / 37Long-Thought - - 0.31 / 10Follow-up Q 0.62 / 36 0.64 / 58 0.54 / 35

[0074] Table 3: The Mean RIM / Mean Number of Questions Asked for each experiment on FB-Synth.Llama3-8b Llama3-8b-Aloe Qwen-32bZero-Shot-U 0.30 / 11 0.30 / 17 0.30 / 11Few-Shot-U 0.24 / 8 0.27 / 18 0.27 / 10Zero-Shot-40 0.34 / 30 0.35 / 31 0.43 / 36Few7- Shot-40 0.34 / 35 0.35 / 35 0.41 / 38Long-Thought - 0.27 / 11Follow -up Q 0.48 / 35 0.51 / 59 0.44 / 34

[0075] In some examples, each metric may depend on determining whether a provider and multi-agent-generated question are requesting the same information. To detect matching questions, a fine-tuned PHI -4- 14b based LLM-as-Judge framework is employed to do a pairwise comparison between the true and generated follow-up questions. A test set (n::T00 question pairs) hand-labeled by a family medicine physician with over twenty years of experience is also used. The judge model can detect matching question pairs that elicit the same information with a macro Fl-score of 0.87.

[0076] In Table 2, it is shown that FollowupQ (Llama3-8b) achieves a mean RIM score of 0.62 on FB-Real while generating 36 questions. This is a 22-pomt increase in RIM compared to comparable zero and few-shot baselines using Llama3-8b - showing the effectiveness of FollowupQ. Although Llatna3- 8b- Aloe achieves an even higher score of 0.64, it also generates additional questions. However, it can be noted that Llama3-8b-Aloe uniquely struggled to follow7instructions in ways that support the ability to control output set size. This can result from its training data.QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)19

[0077] While baseline solutions do see an increase in performance when encouraged to generate more questions (i.e. Unbounded 40-question generation), they may struggle to match the performance of FollowupQ. In other words, while FB-Real averages 3.4 follow-up questions per sample, the results demonstrate that even encouraging the LLM to generate over lOx the number of questions written by a real doctor does not solve this problem - motivating a more intricate solution. In summary', FollowupQ’s ability^ to generate diverse questions provides significant improvements over baseline LLMs in both zero and few-shot settings.[00781 On the FB-Synth dataset, it can be seen that FollowupQ with Llama3-8b provides a 5-point improvement over the closest baseline. When employing Qwen-32b, it was found that the increase in performance is more subtle, but that FollowupQ still outperforms zero and fewshot baselines while generating fewer questions on average.

[0079] FollowupQ reduces the number of information-seeking messages providers send by 34%. In Table 4, each model’s ability to achieve a RIM score of 1.0 on FB-Real is compared, indicating that the provider follow-up questions were captured by the model. It was found that FollowupQ provides a significant 19% increase over the nearest baseline. This suggests that when patients respond to the FollowupQ questions, providers request additional information in 34% fewer symptomatic inquiries, effectively reducing their workload. FollowupQ’s providercentric question-generation strategy produces a broader range of question types, improving performance over baseline LLMs, which often overlook less common concerns in a provider’s differential diagnosis.[0080| Table 4: Number of samples fully matched by each method for the Llama3-8b experiments.Method Message Reduction %Zero-Shot-40 0.15Few-Shot-40 0.13FollowupQ 0.34

[0081] FollowupQ Surfaces Patterns in Providers’ Thought Process: In FIG. 9 the RIM score of each specialist agent in FollowupQ on FBReal is illustrated. It was found that most performance comes from the Differential (worst-case) agent, displaying how follow-up questions are often concerned with ruling out worst-case scenarios as they would require urgent care. However, anon-trivial amount of inquiries come from other specialist agents focusing on medications, timeline clarification, and ambiguity clarification. Notably 10% of performance came from inquiries concerning the patient’s EHR, highlighting the importance of EHR dataQB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)20in personalized follow-up question generation. Thus, FIG. 9 not only demonstrates the interpretability7of the framework, but also surfaces the underlying thought processes of the providers.

[0082] Performance After Filtration Step: In addition to the primary evaluations, question filtration on FB-Real was explored using outputs from the best performing LLM, Llama-3-8b. In this experiment, | Qp| was filtered to a size of 10. The choice of &=10 is motivated by both conversations with clinical experts and the mean number of questions asked in the synthetic dataset (see Table 1). It was found that the top-performing model, on average, generates 36 questions. After de-duplication, the number of questions went from 36 to 22 questions per sample with a mean RIM score of 0.62 from 0.57.Response Enhancement Examples and Experiments

[0083] Locally-hosted LLMs may be stored locally in clinical settings due to the sensitive nature of protected health information (PHI) and the frequency with which PHI occurs in patient portal messages. Token throughput and hosting memory constraints are also considerations. As such, 7-8b parameter LLMs on the response drafting task were evaluated in an experiment. Three models were used: (i) the instruction-tuned L]ama3-8B model, (ii) a healthcare-specific version of the same model Aloe-8B, and (hi) Qwen3-8B from a different model family. Three commercial models were also tested on the Sy PPM dataset and a public dataset: (i) Claude 4.5 Sonnet, (ii) Gemini 2.5 Pro, and (iii) GPT-OSS. Moreover, several avenues for aligning LLMs with expert clinicians were evaluated to improve reliability and responsibility.

[0084] O-Shot: Minimally-guided responses from each model are evaluated to identify how closely-aligned the LLM is with expert clinicians.

[0085] Thematic: In some examples, prompting techniques can improve LLM performance on patient messaging tasks. The experiments described herein determine whether the themes derived in process 400 can align LLMs more closely with expert clinicians. The thematic prompt includes a brief explanation of each of the 8 themes, to guide the LLM w ith context.

[0086] RAG: Retrieval augmented generation can be used in patient messaging tasks to improve style and content of LLM responses 5-shot RAG prompting was performed, as described hereinQB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)21

[0087] SFT: Supervised fine-tuning on prior patient-clinician conversations can be used to adapt LLM for patient message response drafting. SFT is performed using all 144k training messages.

[0088] TADPOLE: Thematic Agentic Direct Preference Optimization for Learning Enhancement strategy was developed for creating theme-driven preference training data for DPO. TADPOLE uses response enhancement agents designed for each theme derived in process 400. Several preference pair creation strategies are developed, and the results of models trained using the best-performing strategy were reported.

[0089] A consideration when evaluating the LLM-clinician alignment is how closely-aligned clinicians are with each other. Clinician alignment may vary based on experience factors (e.g. role, years of experience, specialty), personality factors (e.g. writing style), and interpersonal factors (e.g. relationship with the patient). 3 expert responses to 40 samples from the SyPPM dataset were collected to quantify inter-annotator predictability (IAP). IAP was calculated using the editJudge framework to compare inter-human alignment on patient message response drafting. IAP provides a measure of how useful a different clinician’s response might be when used as a response draft.

[0090] As manually annotating sentence themes in the responses would be inefficient, the empirically-validated sentence-level theme classifier was used to classify themes in the clinician responses (i.e., ground truth) and the LLM response drafts to estimate thematic tendencies (see Table 7).

[0091] Six LLMs and five adaptation techniques were evaluated on the IPPM and SyPPM response drafting evaluation datasets. Table 5 contains both content-level and theme-level edit- Fl scores, averaged across the three local LLMs described above, alongside standard deviation. Table 6 contains content- and theme-level edit-Fl scores for Claude 4.5 Sonnet, Gemini 2.5 Pro, and GPT-OSS reasoning models, using both O-shot and thematic prompting adaptation. In Tables 5 and 6. micro average precision, recall, and edit-Fl at the content and theme levels are reported Table 7 contains theme frequencies for clinician responses and adapted LLM drafts, averaged across the evaluation datasets.

[0092] Table 5: Edit-Fl scores for LLM adaptations on the IPPM and SyPPM patient message response drafting datasets.Content-Level Theme-Level Dataset Model Precision Recall Edit-Fl Precision Recall Edit-Fl IPPM O-Shot 0.07±0.02 0.26±0.04 0.10±0.02 0.49±0.03 0.74±0.03 0.58±0.02Theme 0.06±0.01 0.30±0.05 0.09±0.01 0.47±0.01 0803=0.02 0.583=0.01RAG 0.11±0.03 0.030±0.17 0.13±0.01 0.48±0.20 0.66±0.09 0.56±0.02QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)22SFT 0.15±0.01 0.16±0.00 0.14±0.01 0.64±0.01 0.57±0.01 0.60±0.01 TADPOLE 0.13±0.01 0.18±0.01 0.14+0.01 0.54±0.00 0.65±0.02 0.59±0.01 SyPPM O-Shot 0.12±0.04 0.31±0.03 0.16±0.04 0.47±0.02 0.46±0.03 0.47±0.02Theme 0.11±0.00 0.33±0.10 0.15±0.02 0.50±0.01 0.58±0.01 0.54±0.00 RAG 0.17±0.08 0.28±0.07 0.18±0.05 0.47±0.03 0.43±0.02 0.45±0.02 SFT 0.22±0.01 0.17±0.01 0.18±0.0 0.64±0.01 0.41±0.01 0.50±0.01 TADPOLE 0.21±0.01 0.20±0.02 0.20±0.01 0.62±0.01 0.54±0.02 0.58±0.01 Gemini 0.20 0.43 0.26 0.58 0.69 0.64IAP 0.26 0.25 0.24 0.61 0.63 0.62

[0093] Usefulness of Thematic Context: It was found that fine-tuned models achieve highest precision, theme prompted models achieve highest recall, and the TADPOLE adaptation strategy7offers the best blend of precision and recall with the highest average content-level edit- Fl scores. Moreover, it was found that that added context improves LLM alignment with individual clinicians, and that edit-Fl performance generally scales with the amount of added context. Examining theme-specific content-level recall. TADPOLE-adapted models blend precision with empathetic communication content (0.30 average recall vs 0.28 average among other adaptations) and contingency planning content (0,27 vs 0,21) - two themes which tend to appear more in “ideal” response drafts. Among commercial models, thematic prompting adaptation improves performance of the three LLMs. It was found that the best frontier-level model in the evaluation is Gemini 2.5 Pro adapted with thematic prompting, achieving 0.26 content-level and 0.64 theme-level edit-Fl. The single best-performing TADPOLE model (Qwen3-8B trained on the “corrupted” preference pairs) achieves comparable performance (0.25 content-level edit-Fl score) to the best-performing frontier model (Gemini 2.5 Pro + theme prompt. 0.26). This evaluation suggests that using one of these models in patient message response drafting would lead to a 25-26)% reduction in clinician edits.

[0094] Individual variation stemming from epistemic uncertainty is often observed in medicine, including patient message response drafting. When one clinician’s responses are used as drafts for another clinician, an average content-level edit-Fl score of 0.24 was found, meaning that using another clinicians response as a draft only reduces clinician edits by 24%. This indicates substantial epistemic uncertainty at the content level of clinician responses (i.e., LLMs specialized at the task level are subject to performance loss due to inter-clinician variation in judgment and preferences).

[0095] Evaluating at the theme level shows that LLMs are capable of generating some themes accurately, while other themes are more challenging. For example, LLMs tend to generate the empathetic communication theme frequently (Table 7), and they perform well overall at generating this theme (e.g., TADPOLE-adapted models achieve an average theme-level edit- Fl score of 0.99 on the empathetic communication theme in SyPPM). On the contrast, Table 5QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)23shows that unaligned LLMs will rarely ask follow-up questions. Unaligned LLMs tend to be misaligned with clinicians on question asking themes (e g., O-shot models achieve only 0.17 and 0.08 average theme-level edit-Fl scores on SyPPM symptom and medication questionasking themes). Contextual adaptation improves LLM performance at question asking, with TADPOLE-adapted LLMs improving to 0.79 and 0.49 average theme-level edit-Fl scores on SyPPM symptom and medication question-asking themes.

[0096] Table 6: Edit-Fl results for Claude 3.5 Sonnet, Gemini 2,5 Pro, and GPT-OSS reasoning models on the publicly-available SyPPM evaluation dataset.Content-Level Theme-LevelPrompt Model Pr Re Edit-Fl Pr Re Edit-Fl O-Shot GPT 0.03 0.21 0.05 0.45 0.64 0.53Gemini 0.17 0.40 0.23 0.52 0.56 0.54 Claude 0.20 0.38 0.25 0.52 0.54 0.53 Avg 0.13 0.33 0.18 0.50 0.58 0.53 Theme GPT 0.06 0.30 0.09 0.49 0.77 0.60Gemini 0.20 0.43 0.26 0.56 0.69 0.64 Claude 0.16 0.37 0.22 0.58 0.69 0.63 Avg 0.14 0.37 0.19 0.54 0.72 0.62 IAP 0.26 0.25 0.24 0.61 0.63 0.62

[0097] In general, IAP is much higher at the theme level than at the content level, indicating that theme-level alignment is a more achievable goal when drafting clinician responses. However, some individual themes have very low IAP. Discussions with various clinicians, including the annotators, highlight that different clinicians tend to think differently about how content will be perceived by patients (e.g.. some clinicians indicate that the benefits of providing contingency plans do not outweigh the burden it places on patients).

[0098] It was found that unadapted LLMs tend to generate medical assessment themes more successfully than contextually-adapted LLMs. This is supported by the estimate of theme proportions (Table 7), which finds that unadapted LLMs generate far more medical assessment and treatment planning themes than clinicians and contextually-adapted LLMs. These themes cover utterances related to medical decision making and communication (i.e., explaining test results, symptoms, and potential diagnoses: and recommending various forms of treatment). Unadapted LLMs generate these themes more frequently as they relate to general LLM alignment principles.QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)24

[0099] Table 7: Proportion of responses containing different thematic content, found in responses written by clinicians and various model adaptations.Response Emp Sym Q Med Q Assess Plan Logis Coord Cont Oth Clinicians 0.85 0.36 0.30 0.34 0.19 0.56 0.45 0.22 0.02 O-Shot 0.94 0.02 0.05 0.89 0.82 0.59 0.78 0.18 0.14 Theme 0.95 0.26 0.13 0.94 0.79 0.64 0.82 0.18 0.20 RAG 0.77 0.01 0.05 0.79 0.65 0.56 0.69 0.11 0.19 SFT 0.97 0.02 0.02 0.23 0.26 0.38 0.69 0.02 0.02 TADPOLE 0.99 0.29 0.20 0.28 0.31 0.36 0.83 0.25 0.01

[0100] The evaluation measures how many edits a clinician would make to the LLM-generated draft before sending the response. This is different from the goal of measuring response quality along pre-defined axes, and influences the decision to define a ground truth as a single clinician response, rather than a strategy such as rubric-based evaluation or surveying expert feedback on a generated response. Results from the targeted evaluation highlight the challenge of aligning models with individual clinicians’ judgment, tone, and preferences when responding to patients.Urgency Ranking Examples and Experiments

[0101] Medical triage is the task of allocating medical resources and prioritizing patients based on medical need. The inventors introduce a large-scale public dataset for studying medical triage in the context of asynchronous outpatient portal messages. The task formulation views patient message triage as a pairwise inference problem, where LLMs are trained to choose “which message is more medically urgent" in a head-to-head tournament-style resort of a physician’s inbox. The benchmark PMR-Bench contains 1569 unique messages and 2.000+ high-quality test pairs for pairwise medical urgency assessment alongside a scalable training data generation pipeline. PMR-Bench includes samples that contain both unstructured patient-written messages alongside real electronic health record (EHR) data, emulating a real-world medical tri ge scenario.

[0102] An automated data annotation strategy to provide LLMs with in-domain guidance on this task was designed. The resulting data was used to train two model classes, UrgentReward and UrgentSFT, leveraging Bradley -Terry and next token prediction objective, respectively to perform pairwise urgency classification. It was found that UrgentSFT achieves top performance on PMR-Bench, with UrgentReward showing distinct advantages in low-resourceQB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)25settings. For example, UrgentSFT-8B and UrgentReward-8B provide a 15- and 16-point boost, respectively, on inbox sorting metrics over off-the shelf 8B models.[01031 In this study, a benchmark, Patient Message Ranking (PMR)-Bench, is introduced, a pairwise text classification task covering a diverse array of medical conditions in primary care where the goal is to decide which of the two messages is more medically urgent. Unlike classification, this task formulation is more directly connected to the real problem of deciding which messages should be treated as having higher priority. In addition, intuitively, a binary higher- versus lower-urgency comparison involves simpler comparison semantics than an ordinal set of three or more urgency categories. The ability to compute these comparisons inherently solves the ranking problem, as a PMR model can be deployed as the comparator m a sorting algorithm (e.g., bubble sort / quick sort) to rank a doctor's inbox based on medical urgency.

[0104] PMR-Bench contains 1,569 unique patient messages with clinicians’ ordinal annotations This enables large-scale generation of data pairs for pairwise urgency detection (i.e. which of two patients is more medically urgent) from two different medical communication platforms. First, PMR-Reddit is a publicly available, curated set of patient messages from r / AskDocs -- an online forum where medical experts respond to patient queries. Second, PMR-Synth is a publicly available set of pairwise message comparisons using high-quality, synthetic patient portal messages, paired with real EHR data to emulate a real patient-portal environment in which urgency is determined using both patient message and structured EHR data. Finally. PMR-Real is a proprietary set of real-patient messages and corresponding EHR data sourced from a large regional hospital in the US.

[0105] Two fine-tuning strategies for LLMs were explored. The first is UrgentSFT, which uses Supevised Fine-Tuning (SFT) to adapt LLMs to the novel task. Furthermore, UrgentReward is introduced, which frames pairwise inference training in a re ard modeling context. It was shown that UrgentReward achieves high performance with only an 8B parameter LLM. UrgentReward-8B outperforms GPT-OSS (120B) on this task and achieves 95% of the performance of larger finetuned LLMs. A set of task-specific metrics to evaluate model performance in this task were also defined.

[0106] PMR-Bench is a first large-scale benchmark for pairwise medical urgency assessment. PMR-Bench includes patient messages paired with real structured EHR data-emulating a realistic patient message triage environment. 8 LLMs were benchmarked.QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)26

[0107] Two models were developed, UrgentReward and UrgentSFT, which are pairwise inference approaches for determining which of two patient messages should be attended to first. The methods described herein optimize for accuracy and efficiency with strong performance across 4B, 8B, 27B, and 32B parameter model variants, making the methods more suitable for low-resource settings. For example, UrgentSFT-8B and UrgentReward-8B provide a 15- and 16-point boost, respectively, on inbox sorting metrics over off-the-shelf 8B models.

[0108] Table 8 provides an overview of PMR Bench, which contains 1,569 unique healthcare messages from multiple data sources, (i) PMR Reddit contains messages sourced from the subreddit r / AskDocs, (li) PMR-Real leverages a proprietary corpus of patient messages from a large regional hospital in the US, and (iii) PMR-Synth uses expert-written messages which aim to emulate the style and prose of PMR-Real while enabling us to share high-quality data publicly for reproducibility. Example data from each source is shown in FIG, 10. Due to the high cost and challenges of collecting reliable annotation at scale, a reproducible annotation method that infers labels from clinician responses to messages that are readily available in real responses to patient messages on r / AskDocs and from real physicians in PMR-Real was developed. For PMR-Synth, clinical experts directly annotated pairs of messages.

[0109] Table 8: PMR-Bench dataset statistics.Reddit Synth RealAvg Tokens 228 ± 153 457 ± 100 511 ± 147 Unique Msg 1121 60 388Level 1 126 10 34Level 2 45 10 68Level 3 59 10 79Level 4 384 10 100Level 5 283 10 100Level 6 324 10 7Has EHR No Yes Yes

[0110] PMR-Reddit: Messages were sourced from r / AskDocs, a forum where patients can request feedback from verified clinical experts on Reddit. Using an LLM, those patients who are looking for feedback on acutely onset symptoms were filtered Each comment from a verified expert clinician was classified using GPT-5. Note that only posts with expert comments which (i) received at least 5 upvotes, or (ii) had comment sections with agreement across multiple expert comments were considered. Each comment was classified into the 6-Ievel urgency scale.QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)27The resulting dataset has 1,121 unique posts with the label distribution. The ordinal label assigned to the expert comment is used to sample posts for pairwise comparisons. The PMR- Reddit test set has a total of 1502 pairwise comparisons derived from 362 posts.

[0111] PMR-Real: A set of real patient portal messages were sourced from a large regional medical center in the United States. Similar to the preprocessing of PMR -Reddit, patient messages describing acutely onset symptoms were filtered and the responses to the resulting messages were classified into the ordinal urgency scale. A difference between PMR-Real and PMR-Reddit is the inclusion of structured EHR data Each PMR-Real message has as its linked EHR data, including the patient’s (i) active problem list, (li) recent diagnoses, (in) active medications, and (iv) demographic information.

[0112] PMR-Synth: Messages in PMR-Synth were written and curated by expert members of the study team and subsequently reviewed by additional clinical experts to ensure high-quality, realistic samples covering a wide array of medical topics. The EHR record associated with each message is the de-identified EHR of a real patient from the same pool of patients used to create PMR-Real. Unlike PMR-Reddit and PMR-Synth, a team of medical experts directly classified the message pairs for two separate inboxes of size 30, annotating which of two patients should receive priority medical care. One inbox was treated as a training and the other as a testing set.

[0113] Sample Difficulty Quantification: Each message now has a label 1-6 describing how urgent the message is. The difference between the two labels was considered to be a proxy for pairwise sample difficulty. Easy: Samples with a difficulty level of at least 4 (e.g. Level 1 vs Level 5 / 6). Medium: Samples which are 2-3 difficulty levels apart (e g. Level 1 vs Level 3 / 4). Hard: Samples which are less than 2 difficulty levels apart (e.g. Level 2 and Level 3).

[0114] To solve the pairwise urgency task, UrgentSFT was developed: a tw o-step inference procedure for classifying pairwise medical urgency. Consider an LLM / which processes a pair of texts fla, b) where a and b denote two different patient messages. UrgentSFT asks the model to output a probability (via the ‘YES’ token as in other 1R tasks) deciding if message b should be attended to before message a based on their respective medical urgency. For a given pair of patient messages (a, b) where b was annotated as the more urgent message, model correctness was evaluated using the following probability difference η = P(YES|f(a, b)) − P(YES|f(b, a)). If r] > 0, then the model is correct, as it has successfully attributed a higher probability to the more urgent patient message. UrgentSFT was formulated as a two-step inference leveraging probability scores to help prevent ties and improve models’QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)28sensitivity to input order. Using patient messages which have been classified into the ordinal urgency scale, training pairs for UrgentSFT fine-tuning were created.[01151 UrgentReward frames pairwise classification using a Bradley-Terry (BT) loss to help the LLM better internalize a relative urgency ranking among patient messages. It can be noted that this is in contrast to UrgentSFT, which is trained using the standard next token prediction objective.

[0116] For a given patient message t. UrgentReward training triplets (I, tm, ti) are created, where tmand ti are patient messages that are more and less urgent than t. respectively. Let tPdenote a prompt which contains within it the message t. Specifically, the UrgentReward prompt template asks the model to write a message more medically urgent than the one provided. Thus, the task of quantifying how much more urgent one message is framed when compared to another as the scoring of a prompt completion - aligning the work with prior studies on reward modeling. A reward model is fine-tuned using Urusing the TRL package, which uses a Bradley-Terry objective to maximize Ur(tp, tm) — Ur(tp, tl).

[0117] At test time, UrgentReward was similarly applied to UrgentSFT. For a given pair of patients (a, b), two inferences were run: f(a, b) = s1 and f(b, a) = s2. If s1 > s2, this means that b is more urgent than It can be noted that the pairwise application of the BT model is different than IR that apply BT models in a pointwise re-rank setting. Two BT inferences are performed per pair, and use those scores only to produce a pairwise classification label. The UrgentReward models were fine-tuned on Qwen-based SkyWork-Reward-v2 models, which are LLM-based sequence classifiers pre-trained on 26 million preference pairs.

[0118] The classification accuracy is reported in Table 9, as the evaluation setup is structurally similar to that of a reward modeling evaluation, The overall accuracy is reported, as well as per-difficulty accuracy, where difficulty is defined by the difference in ordinal triage rankings.

[0119] Table 9: Pairwise classification accuracy on each dataset, reported by each difficulty level (easy, medium (med), hard).PMR-Reddit PMR-Synth PMR-RealEasy Med Hard Total Easy Med Hard Total Easy Med Hard Total # of Test Pairs 318 736 448 1502 75 175 185 435 51 274 241 566 Instruct ModelsQwen3-4B 0.85 0.72 0.66 0.73 0.76 0.66 0.50 0.61 0.74 0.51 0.51 0.66 Qwen3-8B 0.85 0.73 0.68 0.74 0.80 0.74 0.55 0.67 0.72 0.60 0.60 0.68 Qwen3-32B 0.89 0.76 0.69 0.77 0.77 0.70 0.51 0.63 0.76 0.58 0.58 0.70 MedGemma-27b 0.89 0.74 0.70 0.76 0.95 0.73 0.63 0.72 0.74 0.60 0.60 0.68 Deep ReasoningModelsQwen3-32B-R 0.81 0.64 0.56 0.65 0.68 0.58 0.38 0.51 0.78 0.58 0.41 0.53QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)29GPT-OSS 0.93 0.76 0.67 0.77 0.79 0.60 0.50 0.59 0.86 0.64 0.56 0.63 UrgentSFTQwen3-4BSFT0.84 0.69 0.63 0.70 0.81 0.75 0.50 0.65 0.90 0.79 0.67 0.75 Qwen3-8BSFT0.90 0.75 0.69 0.76 0.83 0.68 0.52 0.64 0.92 0.80 0.64 0.74 Qwen3-32BSFT0.96 0.86 0.84 0.87 0.92 0.70 0.60 0.69 †MedGemma-27BSFT0.98 0.85 0.87 0.88 0.93 0.77 0.60 0.73 0.92 0.78 0.72 0.77UrgentRewardReward-4Bbase0.79 0.67 055 066 0.75 0.64 0.53 0.61 0.82 0.75 061 070 Reward-4Burgent0.91 0.80 0.80 0.82 0.80 0.66 0.56 0.64 0.92 0.82 0.63 0.75 Reward-8Bbase 0.80 0.67 0.55 0.66 0.84 0.75 0.54 0.68 0.86 0.72 0.57 0.67Reward-8Burgent0.93 0.82 0.85 0.85 0.91 0.73 0.62 0.71 0.92 0.81 0.63 0.74

[0120] Messages of 508 varying urgency levels were sampled to create a diverse inbox of approximately 30 messages per corpus. Upon converting the urgency labels from the ordinal urgency scale into relevancy scores, classic IR metrics are computed directed from the relevancy -mapped samples.

[0121] Instruct Models: Four non-reasoning models were explored, Qwen3-4 / 8 / 32B with thinking disabled, and a medical LLM, Medgemma-27b text-it. Reasoning Models: Two reasoning models were explored, Qwen3-32B and GPT-OSS. Unlike instruct models, which use the probability of the “YES" token, a probability of 1.0 is attributed when the model predicts “YES" as the reasoning process makes use of token probabilities less meaningful.

[0122] Training Data: The training data for UrgentSFT and UrgentReward are the same for each of the three datasets. Training triplets are curated from the pool of samples classified into the 6-label scale. For UrgentReward, this translates into an anchor sample and then a chosen and rejected completion used for training e.g. (Anchor, More Urgent, Less Urgent). The same triplet is converted into multiple SFT samples (e.g. (Anchor, More Urgent, YES) and ('Anchor. Less Urgent, NO)).

[0123] Multi-Class Baseline: For the extrinsic evaluation, LLM capacity was also evaluated to predict the class label directly. GPT-OSS was explored, as well as MedGemma-27B and Qwen3-32B with and without SFT on this task.

[0124] Table 10: Results of the extrinsic evaluation where each model is tasked with re-ranking a clinician's inbox.Reddit Synth Real@10 @30 @10 @30 @10 @30 Multi-ClassMedGem27B 0.49 0.25 0.54 0.64 0.64 0.32 MedGEM27Bsm 0.62 0.32 0.27 0.66 0.66 0.35QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)30Qwen32B 0.59 0.28 0.50 0.62 0.62 0.34 Qwen32BSFT0.50 0.24 0.57 † † GPT-OSS 0.40 0.18 0.48 0.48 0.48 0.23 O-Shot PairsQwen8B 0.41 0.20 0.52 027 052 0.18 Qwen32B 0.54 0.25 0.62 0.28 0.42 0.17 MedGem27B 0.38 0.16 0.66 0.32 0.70 0.31 GPT-OSS 0.44 0.24 0.58 0.28 0.59 0.32 UrgentSFT PairsQwen8BSFT0.51 0.23 0.62 0.30 0.77 0.39 Qwen32BSFT0.77 0.36 0.73 0.34 † MedGEM27BSFT0.61 0.29 0.66 0.35 0.75 0.39 UrgentRewardPairsReward-4BBase0.22 0.10 0.37 0.18 0.48 0.18 Reward-3Burgent0.69 0.30 0.55 0.26 0.65 0.37 Reward-8BBase 0.33 0.14 0.57 0.26 0.54 0.20 Reward-8Burgent 0.58 0.26 0.64 0.29 0.71 0.38

[0125] Table 10 displays the pairwise classification results, it was found that on PMR-Reddit and PMR-Synth, UrgentSFT with MedGemma-27b achieves the highest overall performance. In general, PMR-Reddit scores are higher than PMR-Synth and PMR-Real. This result is intuitive as models do not need to process structured EHR information in PMR-Reddit. Also noteworthy is that PMR-Reddit has a much larger training set, likely contributing to higher performance. However, the ablation study shows that UrgentReward has the capacity to get more out of less training data when applied to smaller models, making it a viable option when fewer resources are accessible

[0126] The comparison between baseline instruct models and reasoning models can be noted. It was found that that Qwen3-32B without reasoning out-performs Qwen3-32B with reasoning. This can be due to input order biases being exaggerated by reasoning as well as the Instruct Models having the advantage of tie-breaking via token probabilities Overall, the methods substantially reduce pairwise triage error on the real datasets, demonstrating that UrgentSFT and UrgentReward can deliver real-world triage improvements while remaining lightweight and deployable with smaller language models.QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)31

[0127] The results of extrinsic evaluation are presented in Table 10. The T-NDCG@30 metric considers the full inbox, reflecting broad-scale sorting quality for a given model. In contrast, T-NDCG@10 metrics emphasizes the top and bottom of the list, more-heavily rewarding correct placement - and penalizing misplacement - of highly urgent messages. The larger performance gap at@10 suggests the models are particularly effective at handling more urgent samples. On PMR-Reddit it was found that UrgentSFT with Qwen3-32B is the highest performing model. Notably, this model outperforms the multi-class baseline, supporting the hypothesis that pairwise inference is more effective. On PMR Synth it was similarly found that the top-performing model is UrgentSFT with Qwen3-32B. Finally, on PMR-Real it was observed that UrgentSFT-8B is the top performing models, with a T-NDCG of 0.77 and 0.39, @ 10 and @ 30, respectively.QB\178981.00038\ 101030782.1

Claims

Docket No. 178981.00038 (2025-020, 2025-011-001)32CLAIMSWHAT IS CLAIMED IS:

1. A method of automatic follow-up questions generation, the method comprising:accessing, using a computerized system, a message received from a patient using a medical dialogue interface;obtaining, using the computerized system, electronic health record (EHR) information corresponding to the patient;providing, using the computerized system, the message and the EHR information to a multi-agent model, the multi-agent model comprising an EHR reasoning agent, a differential diagnostic agent, and a message clarification agent;creating, using the multi-agent model and the computerized system, a set of questions, wherein each question in the set of questions responds to the message;assigning, using the computerized system, a quality value to each question in the set of questions;outputting, via the medical dialogue interface, a question from the set of questions assigned a highest quality value.

2. The method of claim 1, wherein the EHR information includes: a patient age, a patient gender, a medical history of the patient, and a list of medications prescribed to the patient.3 The method of claim 1. wherein the EHR information is represented as a string.

4. The method of claim 1, wherein the multi-agent model was trained using a list of provider-generation questions.

5. The method of claim 4, further comprising:comparing, using the computerized system, each question of the set of questions to the list of provider-generation questions; anddetermining, using the computerized system, a request information match (RIM) metric based on the comparison, wherein the RIM metric is a percentage of questions in the set of questions that match questions in the list of provider-generated questions.wherein the quality value is assigned based on the RIM metric.QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001).

336. The method of claim 1, wherein the multi-agent model further comprises a response enhancing agent.

7. The method of claim 6, further comprising:identifying, via the response enhancing agent, at least one of an expected addition or an expected deletion in the question, andmodifying, using the computerized system, the question based on the expected addition or the expected deletion.

8. The method of claim 7, wherein the response enhancing agent was trained by:receiving, using the computerized system, an annotated question corresponding to a previously -generated follow-up question;identifying, using the computerized system, at least one of a deletion, an addition, or a match between the annotated question and the previously-generated follow-up question; and identifying, using the computerized system, a theme of the annotated question and a theme of the previously -generated follow-up question.

9. The method of claim 8, wherein the theme is selected from list comprising: empathy, symptom, medication, medical assessment, medical planning, logistic, care coordination, and contingency planning.

10. The method of claim 1, wherein the multi-agent model further comprises a patient message ranking agent, wherein the patient message ranking agent determines an urgency associated with the message received from the patient based on the EHR information.1 1. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform a method of automatic follow-up question generation, the method comprising:accessing, using a computerized system, a message received from a patient using a medical dialogue interface;obtaining, using the computerized system, electronic health record (EHR) information corresponding to the patient;QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001)34providing, using the computerized system, the message and the EHR information to a multi-agent model, the multi-agent model comprising an EHR reasoning agent, a differential diagnostic agent, and a message clarification agent;creating, using the multi-agent model and the computerized system, a set of questions, wherein each question in the set of questions responds to the message;assigning, using the computerized system, a quality value to each question in the set of questions;outputting, via the medical dialogue interface, a question from the set of questions assigned a highest quality value.

12. The non-transitory computer-readable storage of claim 11, wherein the EHR information includes: a patient age, a patient gender, a medical history of the patient, and a list of medications prescribed to the patient.

13. The non-transitory computer-readable storage of claim 11, wherein the EHR information is represented as a string.

14. The non-transitory computer-readable storage of claim 11, wherein the multi -agent model was trained using a list of provider-generation questions.

15. The non-transitory computer-readable storage of claim 14, wherein the method further comprises:comparing, using the computerized system, each question of the set of questions to the list of provider-generation questions; anddetermining, using the computerized system, a request information match (RIM) metric based on the comparison, wherein the RIM metric is a percentage of questions in the set of questions that match questions in the list of provider-generated questions,wherein the quality value is assigned based on the RIM metric.

16. The non-transitory computer-readable storage of claim 11, wherein the multi -agent model further comprises a response enhancing agent.QB\178981.00038\ 101030782.1Docket No. 178981.00038 (2025-020, 2025-011-001).3517. The non-transitory computer-readable storage of claim 16, wherein the method further comprises:identifying, via the response enhancing agent, at least one of an expected addition or an expected deletion in the question; andmodifying, using the computerized system, the question based on the expected addition or the expected deletion.

18. The non-transitory computer-readable storage of claim 17, wherein the response enhancing agent was trained by:receiving, using the computerized system, an annotated question corresponding to a previously-generated follow-up question;identifying, using the computerized system, at least one of a deletion, an addition, or a match between the annotated question and the previously-generated follow-up question; and identifying, using the computerized system, a theme of the annotated question and a theme of the previously -generated follow-up question.

19. The non-transitory computer-readable storage of claim 18, wherein the theme is selected from list comprising: empathy, symptom, medication, medical assessment, medical planning, logistic, care coordination, and contingency planning.

20. The non-transitory computer-readable storage medium of claim 11, wherein the multi¬ agent model further comprises a patient message ranking agent, wherein the patient message ranking agent determines an urgency associated with the message received from the patient based on the EHR information.QB\178981.00038\ 101030782.1