Guided report interpretation using large language model
A two-module pipeline system with a deterministic rule-based engine and LLM generates accurate, non-invasive gastric function assessments, addressing the limitations of current diagnostic methods by providing actionable biomarkers for personalized therapy.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ALIMETRY LTD
- Filing Date
- 2026-01-16
- Publication Date
- 2026-07-23
AI Technical Summary
Current diagnostic methods for chronic gastro-duodenal disorders are invasive, costly, and lack objective biomarkers, leading to prolonged and uncertain diagnostic processes and ineffective treatment plans.
A two-module pipeline system using a deterministic rule-based interpretation engine and a large language model (LLM) to generate a report interpretation summary from gastric test data, incorporating an electrode array patch for non-invasive electrophysiological analysis and digital symptom profiling, with guardrails to ensure clinical safety and accuracy.
Provides actionable biomarkers for personalized therapy by accurately stratifying patients, reducing invasive testing, and improving diagnostic and treatment efficacy while minimizing patient harm and healthcare costs.
Smart Images

Figure IB2026050414_23072026_PF_FP_ABST
Abstract
Description
PATENT Atorney Docket No. 112142-001310PC-1532135GUIDED REPORT INTERPRETATION USING EARGE LANGUAGE MODELCROSS REFERENCE TO RELATED APPLICATION DATA
[0001] The present application claims the benefit of priority to U.S. Provisional Patent Application. No. 63 / 746,581 filed January 17, 2025, the full disclosure of which is incorporated herein by reference in its entirety for all purposes.BACKGROUND OF THE INVENTION
[0002] Chronic gastro-duodenal symptoms affect more than 10% of the global population and have a significant healthcare burden, resulting in a significant economic impact.Functional gastrointestinal (GI) disorders are among the most prominent causes of chronic ill-health in both adults and children. Chronic gastroduodenal diagnosis paradigms rely on symptom-based criteria which group nausea, vomiting, abdominal pain, early satiety, and / or excessive fullness into disorders such as chronic nausea and vomiting syndromes (CNVS), functional dyspepsia (FD), and when gastric emptying is delayed, gastroparesis. However, these classifications substantially overlap, limiting their clinical utility and ability to effectively inform individual patient management.
[0003] Functional gastrointestinal disorders (FGIDs, or disorders of gut-brain interaction) place an economic burden on healthcare systems and reduce patient quality of life. Functional gastrointestinal disorders generally affect 35% to 70% of people at some point in life, women more often than men. For example, more than 70% of patients indicate that their symptoms interfere with everyday life and 46% report missing work or school. A recent review of 26 studies found that between 10-29% of school children reported symptoms consistent with a functional GI disorder. The symptoms are frequently distressing and may be severe and debilitating, encompassing chronic abdominal pain, abdominal distension, anorexia, and chronic nausea and vomiting. These disorders collectively extract a major illness burden, including a significantly reduced quality of life, and are common reasons for adults and children missing work or school.
[0004] Current GI diagnosis include gastroparesis and functional dyspepsia disorders. Gastroparesis is defined by symptoms of nausea and vomiting, typically with other symptoms e.g., abdominal pain, bloating, burning, excessive fullness, early satiation, and / or documented presence of delayed gastric emptying. Functional dyspepsia is defined by chronic symptoms such as distress after eating, indigestion, abdominal pain, bloating, burning, excessive fullness, and / or early satiation. Gastric emptying may also be delayed in up to 25% of patients identified with functional dyspepsia, and therefore overlaps with gastroparesis, however nausea and vomiting are not considered the dominant feature. Because these disorders overlap significantly, or at least many patients are on the same disease spectrum, there is a state of confusion in the clinical field. For example, healthcare professionals are often unsure how to define, distinguish and diagnose such patients, and therefore are unable to provide appropriate patient specific management plans, typically reverting to trial and error type therapies.
[0005] Objectively evaluating and treating adults and children with chronic upper GI symptoms is a major clinical challenge, owing to a lack of routine tests that may reliably and safely distinguish specific underlying disorders. Relying on symptom-based diagnoses often results in less than ideal and potentially hazardous attempts at trial-and-error treatments. Currently, both adult and pediatric patients with chronic GI symptoms frequently undergo a protracted diagnostic process that may include endoscopies, biopsies, lab tests, nuclear medicine studies, manometry and radiology exams, often over numerous years. Many of these tests are invasive and involve radiation, yet the diagnostic results are often inconclusive. For example, gastric scintigraphy and antroduodenal manometry are two tests that are commonly performed in adult and pediatric gastroenterology, as they may distinguish myopathic or neuropathic functional disorders, and may impact diagnosis and treatment in 15-20% of patients with chronic upper GI symptoms. However, the interpretation of these tests may be uncertain, especially in pediatric applications due to a lack of diagnostic norms in children. Furthermore, these tests typically involve long wait times and high cost as these tests are generally only available in specialist referral centers.
[0006] There is a pressing need for improved and less invasive diagnostic tests that have clinical utility, offer actionable and objective biomarkers that improve both adult and pediatric diagnostic and treatment efficacy, reduce patient harm from negative invasive or unnecessary testing, and directly impact clinical care decisions and treatment. The advent of a less invasive and technically safer diagnostic test for adults and children would broadenavailability and access and reduce the high healthcare expenses of motility testing. An optimal diagnostic solution would be non-invasive, user-friendly, easy to apply and interpret, and provide meaningful results that correlate with symptoms and inform clinical care.SUMMARY OF THE INVENTION
[0007] According to one embodiment, a system and method for generating a report interpretation summary utilize a two-module pipeline architecture designed to ensure clinical safety and accuracy. The first module comprises a deterministic rule-based interpretation engine configured to receive gastric test data measured with an electrode array patch and transform the data into a plurality of structured interpretation points using predefined clinical rules and normative thresholds. These interpretation points span multiple clinical domains, including test quality, spectral analysis, patient-reported symptoms, and gut-brain wellbeing. This module also identifies applicable phenotypes and retrieves relevant clinical information for the applicable phenotype for a given patient.
[0008] The second module includes a summarization engine configured to strictly utilize the structured interpretation points as input prompts. This module constructs a dedicated prompt for each of a plurality of sections, provides the prompt to a large language model (LLM), and generates a narrative summary synthesized from the clinical facts provided in the interpretation points. To mitigate the risk of hallucinations, the system implements guardrails. These include a confidence analysis subsystem that calculates token-level log probabilities for the generated narrative summary, screening for information generated by the model that lacks a basis in the input, regular expression-based content validation, and a final clinician oversight step including manual review. When any guardrail criteria are met, the system may trigger a validation flag for human review and / or initiates corrective actions, such as prompt resubmission or modification.
[0009] According to one embodiment, a method for generating a report interpretation summary includes receiving a request to generate a summary for test data associated with gastric activity of a patient over a predetermined time period, where the test data includes electrical signals that are measured with an electrode array patch disposed over a skin surface of the patient and identifying a plurality of sections to be included in the summary. The method further includes processing the test data using a deterministic rule-based interpretation engine to generate a plurality of structured interpretation points including data. For each section in the plurality of sections, the method includes constructing a prompt forthe section based upon a prompt template configured for the section and the structured interpretation points, providing the prompt as input to an LLM, and responsive to the providing, generating by the LLM a summary for the section. The method further includes generating an overall summary that includes the summaries generated for the plurality of sections and performing a set of one or more validation checks to check contents of the overall summary.
[0010] The method may include various optional embodiments. The method may further include for each validation check in the set of one or more validation checks, upon identifying one or more errors in the overall summary as a result of performing the validation check, performing one or more actions to rectify the one or more errors, where performing the one or more actions causes at least a portion of the contents of the overall summary to be changed resulting in an updated overall summary. The method may further include generating a final overall summary using the updated overall summary and providing the final overall summary as a response to the request. Constructing the prompt may further include identifying an LLM to be used for generating a summary for the section, providing the constructed prompt as an input to the LLM, and outputting, by the LLM, a summary for the section. Constructing the prompt may further include identifying at least one required segment in the identified prompt template, identifying zero or more conditional segments in the identified prompt template, and determining, based upon the test data, which, if any, conditional segments from the conditional segments are to be included for the prompt construction. Constructing the prompt may further include for each variable portion in the constructed prompt, determining, based upon the identified test data, content to be inserted in the variable portion, and for each variable portion in the constructed prompt, replace the variable portion with the content determined for the variable portion. The method may further include determining whether an identified error in the overall summary could lead to patient harm; and in response to determining that the identified error would lead to patient harm, updating the overall summary to indicate a warning. The method may further include determining whether an identified error in the overall summary could lead to patient harm and in response to determining that the identified error would lead to patient harm, requesting healthcare professional review of the identified error. The one or more actions may include prompt resubmission, prompt resubmission with changed LLM settings, or modifying the prompt. The method may further include refining the results of the overall summary where refining the results comprises manually editing the overall summary or regenerating theoverall summary with prompt refinements. The skin surface of the patient may include at least one of an abdomen, a torso, or a flank of the patient. The set of one or more validation checks may include checking key interpretation points, checking numerical accuracy, checking allowed words, checking LLM confidence, or additional LLM queries.
[0011] According to another embodiment, a system for processing gastric activity data includes an electrode array patch disposed over a skin surface of a patient for measuring electrical signals associated with gastric activity of the patient over a predetermined time period and a processor configured to receive a request to generate a summary for test data associated with gastric activity of a patient over a predetermined time period, where the electrical signals are measured with an electrode array patch disposed over a skin surface of the patient and identify a plurality, of sections to be included in the summary. The processor is further configured to process the test data using a deterministic rule-based interpretation engine to generate a plurality of structured interpretation points. The processor is further configured to, for each section in the plurality of sections, construct a prompt for the section based upon a prompt template and the structured interpretation points, provide the prompt as input to an LLM, and generate by the LLM a summary for the section. The processor is further configured to generate an overall summary that includes the summaries generated for the plurality of sections and perform a set of one or more validation checks to check contents of the overall summary.
[0012] The system may include various optional embodiments. The processor may be further configured to, for each validation check in the set of one or more validation checks, upon identifying one or more errors in the overall summary as a result of performing the validation check, perform one or more actions to rectify the one or more errors, where performing the one or more actions causes at least a portion of the contents of the overall summary to be changed resulting in an updated overall summary, generate a final overall summary using the updated overall summary, and provide the final overall summary as a response to the request. Constructing the prompt may further include identifying an LLM to be used for generating a summary for the section, providing the constructed prompt as an input to the LLM; and outputting, by the LLM, a summary for the section. Constructing the prompt may further include identifying at least one required segment in the identified prompt template, identifying zero or more conditional segments in the identified prompt template, and determining, based upon the test data, which, if any, conditional segments from the conditional segments are to be included for the prompt construction. Constructing the promptmay further include, for each variable portion in the constructed prompt, determining, based upon the identified test data, content to be inserted in the variable portion, and, for each variable portion in the constructed prompt, replace the variable portion with the content determined for the variable portion. The processor may be further configured to determine whether an identified error in the overall summary could lead to patient harm and in response to determining that the identified error would lead to patient harm, update the overall summary to indicate a warning. The processor may be further configured to determine whether an identified error in the overall summary could lead to patient harm, and in response to determining that the identified error would lead to patient harm, request healthcare professional review of the identified error. The one or more actions may include prompt resubmission, prompt resubmission with changed LLM settings, or modifying the prompt.
[0013] According to another embodiment, a non-transitory computer-readable medium include storing instructions executable by one or more processors for causing the one or more processors to perform operations including receiving a request to generate a summary for test data associated with gastric activity of a patient over a predetermined time period, where the test data includes electrical signals that are measured with an electrode array patch disposed over a skin surface of the patient, and identifying a plurality of sections to be included in the summary. The operations further include processing the test data using a deterministic rulebased interpretation engine to generate a plurality of structured interpretation points. For each section in the plurality of sections, the operations further include constructing a prompt for the section based upon a prompt template configured for the section and the structured interpretation points, providing the prompt as input to an LLM, and responsive to the providing, generating by the LLM a summary for the section. The operations further include generating an overall summary that includes the summaries generated for the plurality of sections, and performing a set of one or more validation checks to check contents of the overall summary.
[0014] Other embodiments and variations thereof will become apparent from the following description which is given by way of example only and with reference to the accompanying drawings.
[0015] It is acknowledged that the term ‘comprise’ may, under varying jurisdictions, be attributed with either an exclusive or an inclusive meaning. For the purpose of this specification, and unless otherwise noted, the term ‘comprise’ shall have an inclusivemeaning, allowing for inclusion of not only the listed components or elements, but also other non-specified components or elements. The terms ‘comprises’ or ‘comprised’ or ‘comprising’ have a similar meaning when used in relation to the system or to one or more steps in a method or process.
[0016] As used hereinbefore and hereinafter, the term “and / or” means “and” or “or”, or both. As used hereinbefore and hereinafter, “(s)” following a noun means the plural and / or singular forms of the noun. As used hereinbefore and hereinafter, the term “continuous” or “semi-continuous” with respect to the test period is to be interpreted as ongoing throughout the entire or nearly entire test period.BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The invention will now be described by way of example only and with reference to the drawings in which:
[0018] FIG. l is a perspective view of a flexible electrode patch, according to various embodiments of the present disclosure.
[0019] FIG. 2 illustrates a network for a test data interpretations system, in accordance with various embodiments of the present disclosure.
[0020] FIG. 3 illustrates a detailed view of the test data interpretations system of FIG. 2, in accordance with various embodiments of the present disclosure.
[0021] FIG. 4 is a flowchart of a method for generating a report interpretation, in accordance with various embodiments of the present disclosure.
[0022] FIG. 5 is a flowchart of a sub-method for generating a report interpretation, in accordance with various embodiments of the present disclosure.
[0023] FIG. 6 is a flowchart of a sub-method for generating a report interpretation, in accordance with various embodiments of the present disclosure.
[0024] FIG. 7 is a flowchart of a sub-method for generating a report interpretation, in accordance with various embodiments of the present disclosure.
[0025] FIG. 8 is a flowchart of a method for risk assessment, in accordance with various embodiments of the present disclosure.
[0026] FIGS. 9A-9E illustrate various guidelines for interpreting test results, in accordance with various embodiments of the present disclosure.
[0027] FIGS. 10A-10D illustrate various guidelines for interpreting test results, in accordance with various embodiments of the present disclosure.
[0028] FIGS. 11 A-l IE illustrate various guidelines for interpreting test results, in accordance with various embodiments of the present disclosure.
[0029] FIG. 12 illustrates various components of the gastrointestinal tract, in accordance with various embodiments of the present disclosure.
[0030] FIG. 13 illustrates various components of the gastrointestinal tract, in accordance with various embodiments of the present disclosure.
[0031] For purposes of the description hereinafter, the terms “upper”, “lower”, “right”, “left”, “vertical”, “horizontal”, “top”, “bottom”, “lateral”, “longitudinal” and derivatives thereof shall relate to the teachings herein as it is oriented in the drawing figures. However, it is to be understood that the variations of the teachings herein may assume various alternative variations, except where expressly specified to the contrary. It is also to be understood that the specific devices illustrated in the attached drawings and described in the following description are simply exemplary embodiments. Hence, specific dimensions and other physical characteristics related to the embodiments disclosed herein are not to be considered as limiting.DETAILED DESCRIPTION OF THE INVENTION
[0032] The present invention provides non-invasive assessment of gastric function using electrophysiological analysis and digital symptom profiling of the gastric conduction system to provide actionable biomarkers that stratify patients into therapeutic groups (e.g., such as groups where gastric dysfunction is present versus absent) to provide a roadmap for personalized (e.g., patient specific) therapy. Various embodiments of the present disclosure utilize large language model (LLM) based report interpretation systems and methods for analyzing gathered data and providing guided interpretation information to healthcare professionals for making informed decisions regarding a patient’s condition and potential treatment options.
[0033] FIG. l is a perspective view of a flexible electrode patch. A flexible electrode patch 10 includes a flexible substrate 12 and a plurality of electrodes 14 disposed thereon. The flexible substrate 12 may include a sensing region 16, a connector region 18, and a tail region 20. In various embodiments and as shown in FIG. 1, the plurality of electrodes 14 is disposed on the sensing region 16 of the flexible substrate 12.
[0034] In various embodiments, the flexible electrode patch 10 is larger than conventional ECG patches or the like. For example, ECGs are typically recorded using a maximum of 10 electrodes that are individually applied to the patient. In another example, ECGs may be recorded via a wearable monitor that typically consists of a maximum of 4 electrodes. In contrast, the flexible electrode patch 10 of the present disclosure is relatively large to ensure that the gastric and / or colonic regions are fully covered by the plurality of electrodes 14 (e.g., including, but not limited to 64 electrodes). Furthermore, the stomach and other organs within the abdomen have variable sizes and configurations between patients and a relatively larger flexible electrode patch 10 may be used for with a wide range of patient sizes. According to at least some embodiments, the area of the sensing region 16 of the flexible electrode patch 10 is a 225 cm2(e.g., 15 cm by 15 cm). Additionally, gastric or colonic signals are relatively weak compared to signals of the heart that are measure by ECGs. Conventional ECGs would not be capable of reliably and accurately gathering the signal data compared to the flexible electrode patch 10 as described herein.
[0035] As shown in the exemplary embodiment of FIG. 1, there are total of 66 electrodes out of which 64 electrodes are arranged in an array of 8 rows and 8 columns, and the remaining two electrodes are the ground and reference electrodes. In use, electrical potentials may be measured as the difference between each of the 64 electrodes and the reference electrode. The ground electrode may be the “driven right leg” or “bias” electrode. The purpose of the ground electrode in some embodiments is to keep voltage level of the subject’s body within an acceptable range and to minimize any common-mode in the subject’s body (e.g., 50 / 60 Hz power-line noise). The driven right leg may act as a source or sink. However, the flexible electrode patch 10 may comprise more than 66 electrodes or less than 66 electrodes. The ground and reference electrodes may be different than what is shown in FIG.1. In an embodiment, the patch may comprise less than, greater than, or equal to 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 125, 150, 175, 200, 225, 250, 275 or 300 electrodes any value or range of values therebetween in 1 increments (e.g., 33, 94, 44 to 192, etc.). According to at least some embodiments, each ofthe 64 electrodes (e.g., any electrodes other than the ground electrode and reference electrodes) may be equally spaced from one another. The plurality of electrodes 14 may be arranged in parallel lines (e.g., other than the ground electrode and reference electrode) or including the ground electrode and reference electrode.
[0036] According to some embodiments, for example, for a flexible electrode patch 10 having 64 electrodes as shown in FIG. 1, each electrode of the plurality of electrodes 14 may have a diameter between 10 mm and 13 mm, inclusive. Furthermore, the spacing between each of the plurality of electrodes 14 (having any number of electrodes) may be a center to center spacing between 18 mm to 22 mm, inclusive. In some embodiments, for example, for a flexible electrode patch 10 having 32 electrodes, each electrode of the plurality of electrodes 14 may have a diameter between 10 mm and 25 mm, inclusive, and a center to center spacing between 18 mm to 45 mm, inclusive.
[0037] According to various embodiments, one or more of the plurality of electrodes 14 may be deactivated during use of the flexible electrode patch 10. For example, 8 to 10 electrodes (in addition to or including the ground electrode and reference electrode) may be used for mapping in response to a determination that the 8 to 10 electrodes receive the strongest signals or the like. In various embodiments, any number of electrodes may be activated or deactivated. For example, any number of electrodes (e.g., up to and including the total number of electrodes) may be deactivated for various reasons such for saving power consumption, extending battery life, etc.
[0038] According to at least some embodiments, the flexible substrate 12 (and / or any other layers to be described herein) may be pre-formed in a convex shape such that the sensing region 16 is the first portion to contact the skin of the patient, thereby ensuring full contact between at least some of the plurality of electrodes 14 and the skin of the patient.
[0039] The flexible electrode patch 10 may further include a connector assembly 22 disposed at least partially on the connector region 18 of the flexible substrate 12. For example, the connector assembly 22 may originate at the connector region 18 and extend into the sensing region 16 and / or the tail region 20. In various embodiments, the cutout 29 is positioned between at least two connector assemblies or between at least two portions of the connector assembly 22, as shown at least in FIG. 1. A split connector assembly reduces the mating force required for the data acquisition device 100 to couple with the connector assembly 22, thereby providing a more reliable (e.g., less likely to be damaged) connection.Furthermore, the split connector assembly increases the flexibility of the flexible electrode patch 10. The combination of the cutout 29 and the split connector assembly creates a free-floating portion of the flexible electrode patch 10 that improves the conformability of the flexible electrode patch 10 to the skin of the patient, thereby increasing the reliability of the plurality of electrodes 14. The split connector assembly further simplifies manufacturing of the flexible electrode patch 10 and enables manufacturing of the relatively small plurality of electrically conductive tracks 24 of the connector assembly 22.
[0040] According to various embodiments, the flexible electrode patch 10 further includes a plurality of electrically conductive tracks 24 disposed on the flexible substrate for electrically coupling each of the plurality of electrodes 14 in the sensing region 16 and a plurality of electrically conductive contact pads 26 disposed on the connector region 18. The plurality of electrically conductive tracks 24 may be disposed on the flexible substrate 12 and run between the each of the plurality of electrodes 14 and each of the plurality of electrically conductive contact pads 26. In various embodiments, the number of the plurality of electrically conductive tracks 24 is the same as the number of the plurality of electrodes 14. Similarly, the number of the plurality of electrically conductive contact pads 26 may be the same as the number of the plurality of electrodes 14.
[0041] According to some embodiments, the plurality of electrodes 14 and / or the plurality of electrically conductive tracks 24 are screen-printed on to the flexible substrate 12 in a manner which would be appreciated by one having ordinary skill in the art. According to various embodiments, the connector assembly 22 does not include any wires (e.g., cables or the like) coupled to the flexible electrode patch 10. The omission of wires not only improves the wearability of the flexible electrode patch 10 but also simplifies manufacturing and use as wires are often difficult to make and maintain (e.g., the wires must be cleaned between each patient, etc.)
[0042] According to various embodiments, the plurality of electrodes 14 of the flexible electrode patch 10 contact the skin of a patient for use as electrophysiological sensors.Signals from the plurality of electrodes 14 may be read by connection of appropriate electronic hardware in electrical communication with the plurality of electrically conductive contact pads 26. In exemplary embodiments, the flexible electrode patch 10 is configured for use in monitoring gastro-intestinal electrical activity and / or colonic electrical of a patient, in part by an appropriate spatial arrangement of the plurality of electrodes 14 in an array, suchas according to the sensor array and various embodiments as is described in WO 2021 / 130683, which is herein incorporated in its entirety and for all purposes.
[0043] In at least some embodiments, the connector region 18 of the flexible substrate 12 includes one or more alignment holes 28 for aligning and electrically connecting a data acquisition device 100 to the plurality of electrically conductive contact pads 26 of the connector region 18 during use. The one or more alignment holes 28 may correspond to projections 102 of the data acquisition device 100 for ensuring that the data acquisition device 100 is properly connected to the flexible electrode patch 10. In exemplary embodiments, one or more alignment holes 28 are offset for further emphasizing the correct orientation of a data acquisition device 100 relative to the flexible electrode patch 10. For example, alignment hole 28a and alignment hole 28b may be offset from each other along an axis perpendicular to a longitudinal axis of the flexible electrode patch 10, to be described in further detail below.
[0044] According to some embodiments, the data acquisition device 100 may include a first clamping member 106 and a second clamping member 108 that are configured to move between an open position as shown in FIG. 1 and a closed position having the first clamping member 106 and the second clamping member 108 adjacent to surface 110. As shown, in the open position the first clamping member 106 and the second clamping member 108 are both configured to pivotally move away from the surface 110 and reveal the surface 110 and the connector region 18 of the flexible electrode patch 10. Similarly, in the closed position the first clamping member 106 and the second clamping member 108 are both configured to move pivotally towards the surface 110 and conceal the surface 110 and the connector region 18 of the flexible electrode patch 10. The first clamping member 106 and / or the second clamping member 108 may include at least one connector 112 that is configured to be physically and operatively connected with the flexible electrode patch 10 receiving the electrical signals from plurality of electrodes 14 of the flexible electrode patch 10 to allow monitoring the electrical activity generated by the gastric or colonic activity of the patient. Therefore, no cable is required to connection between the connector assembly 22 and the flexible electrode patch 10.
[0045] In at least some embodiments, the data acquisition device 100 is free floating at the connector region 18 of the flexible substrate 12. For example, there is no adhesive coupling the flexible electrode patch 10 to the data acquisition device 100 and the flexible electrodepatch 10 may conform freely to the patient's skin even with the data acquisition device 100 coupled to the flexible electrode patch 10.
[0046] In various embodiments, the flexible electrode patch 10 may include a cutout 29 for a display 104 of the data acquisition device 100 to protrude through when the data acquisition device 100 is coupled to the flexible electrode patch 10. The display 104 may include information associated with the status of the flexible electrode patch 10 and / or the data acquisition device 100 including power status, charging status, a mapping mode, etc.
[0047] Various embodiments of the present disclosure takes as input a structured set of results from an electrode patch array as described with respect to FIG. 1, in addition to any other results from the system such as from other measurement devices, and outputs a report interpretation summary. The results are generated based at least in part on test metadata including test data quality, meal information, duration, etc., and test outputs including spectral metrics, symptom values, survey results, etc. Various systems and methods as described herein transform the structured results into plaintext prompts, pass each of the plaintext prompts as queries into an LLM, and filter the LLM outputs using various automated guardrails.
[0048] In various embodiments, the test data interpretation system 200 is configured as a two-module pipeline to ensure the clinical safety, accuracy, and reproducibility of the generated summaries. This architecture prevents the large language model (LLM) from directly interpreting raw, unstructured sensor data, thereby mitigating the risk of "black box" hallucinations or mathematical errors.
[0049] According to various embodiments, the first module may include a deterministic rule-based interpretation engine. This engine ingests raw test outputs, including spectral metrics, patient-reported symptoms, and test quality metadata. Using a predefined set of clinical rules and normative thresholds derived from report interpretation guidelines, the engine translates raw quantitative metrics into a structured set of human-readable "interpretation points". For example, a raw Principal Gastric Frequency (PGF) value is not passed directly to the LLM. The rule-based interpretation engine evaluates the value against a normative range (e.g., 2.65-3.35 cpm) and generates a categorical interpretation point such as "PGF overall normal". This ensures that the foundational clinical facts are verified by deterministic logic before any narrative synthesis occurs.
[0050] According to various embodiments, the second module may include an LLM summarization engine. This engine takes the structured interpretation points from the first module and synthesizes them into a concise, narrative summary. The behavior of the LLM is tightly controlled by "locked prompts" which include role-setting directives and strict output templates. By utilizing the LLM solely for linguistic synthesis rather than clinical calculation, the system ensures that the final report remains grounded in the verified interpretation points. The handoff between these two modules is further protected by automated guardrails that verify the LLM's output against the original rule-based interpretation points.
[0051] To ensure the narrative summaries remain reliable, in some embodiments, the system includes a confidence analysis subsystem that serves as an intrinsic mathematical check of the LLM’s output. This subsystem analyzes the internal confidence of the summarization model for every generated sentence by calculating token-level log probabilities. The log probability is a mathematical measure of the model’s certainty in its prediction of each specific token (word or sub-word) within the clinical summary. The subsystem aggregates these values to produce average per-word, per-sentence, and persection confidence scores. Experimental data has demonstrated that log probabilities scores serve as a valid proxy for the accuracy of a clinical summary. For instance, in validation testing, cases where the LLM failed to mention a required clinical finding (such as a borderline BMI) or incorrectly categorized a meal response exhibited significantly lower log probabilities compared to accurate summaries. Consequently, if a summary score falls below a predetermined "sufficiently confident" threshold identified through manual review, the system is configured to automatically flag the section for human clinician review or trigger a summary regeneration action.
[0052] The LLM summarization engine is optimized for clinical accuracy through a multistage supervised fine-tuning process. This process aligns the linguistic output of the model with expert clinical preferences and standardized reporting styles. In a first stage, referred to as model distillation, a large foundational model is used to generate a distillation dataset consisting of approximately 1,000 report summaries. A smaller, more computationally efficient model is then tuned using this dataset to produce outputs that mimic the performance and complexity of the larger model for the specific summarization task. In a second stage, referred to as expert alignment, the distilled model is further refined using an "Expert-Aligned Dataset." This dataset may include summaries created through manual review and revision by expert clinicians. By training the model on these clinician-revised targets, thesystem ensures that the final LLM output adheres to clinical best practices, prioritizes the most relevant clinical phenotypes, and maintains a tone appropriate for medical documentation.
[0053] FIG. 2 illustrates a network for a test data interpretations system 200. FIG. 2 illustrates an exemplary test data interpretations system 200 having various inputs to be described in further detail below. The test data interpretations system 200 may include more or less inputs than those shown in FIG. 2. Any of the inputs may be transferred across a network including to and from the test data interpretations system 200 in a manner known in the art. For example, the test data interpretations system 200 may be provided as a cloud service.
[0054] Test data sources 202 may be input into the test data interpretations system 200. Test data sources 202 may include one or more sensor-enabled devices data (such as the electrode patch array as described with respect to FIG. 1, a smart watch or smart phone, any other tracking device such as a sleep tracker, heart rate monitor, glucose monitor, etc.), lab tests, patient symptom surveys, etc. Each of these test data sources 202 provides at least one type of test data 204 including sensor data (such as sensor data from the electrode patch array as described with respect to FIG. 1) gathered from one or more sensor-enabled devices, lab test data, patient symptoms information, or any other patient data such as from testing including any one of or any combination of breath testing, stool analysis, motility and function testing (e.g., gastric emptying, manometry), imaging and visual diagnostics (e.g., endoscopy, CT, dynamic MRI), biochemical and serological testing, pH and reflux testing, food sensitivity and intolerance testing, etc.
[0055] The test data interpretations system 200 further receives or otherwise includes configuration information for generating various interpretation reports 206. The test data interpretations system 200 may generate various types of reports including various combinations of sections, as desired by the healthcare professionals. For example, a healthcare professional may request a report regarding gastric activity that emphasizes results associated with a patient consuming a predetermined meal. Accordingly, the test data interpretations system 200 may be configured to generate a “gastric activity report” having sections for each of “pre meal activity,” “during meal activity,” and “post meal activity.” In some embodiments, the healthcare professional may only request the “gastric activity report” having the “post meal activity” section. Accordingly, the configuration information forgenerating various interpretation reports 206 may include information associated with each of the reports that may be generated and each of the sections of each of the reports that may be generated. Said another way, the configuration information for generating various interpretation reports 206 includes templates for generating one or more reports where each of the one or more reports includes one or more sections, to be described in further detail below. The configuration information for generating various interpretation reports 206 may be updated continuously or periodically as new reports are generated and / or requested from the healthcare professional or other users.
[0056] The test data interpretations system 200 may further have one or more LLMs 208. The LLMs 208 may include various LLMs known in the art. In various embodiments, the test data interpretations system 200 may use different LLMs for different reports and / or for different sections of the reports.
[0057] The test data interpretations system 200 may further receive expert inputs 210. For example, expert inputs 210 may include feedback on previously generated reports, updates reports and / or sections of reports, analysis of the reports and / or sections of the reports that is input back into the test data interpretations system 200 to improve generated reports and / or sections of the reports.
[0058] According to various embodiments, the test data interpretations system 200 receives a request to interpret test data from a user 212 via a user system 214. For example, a healthcare professional may use a user system within the healthcare environment to request a report based on a meal consumption test and gastric activity gathered from an electrode patch array applied to the patient’s abdomen throughout the duration of the meal consumption test. In response to the request to interpret test data, the test data interpretations system 200 provides the interpretation report based at least in part on the inputs described in detail above. One or more users 212 may request test data interpretation via one or more user systems 214, as shown in FIG. 2.
[0059] FIG. 3 illustrates a detailed view of the test data interpretations system of FIG. 1. According to various embodiments, the test data interpretations system 200 include various subsystems for performing subprocesses of the interpretation report generation. The test data interpretations system 200 may include more or fewer subsystems than those shown in FIG.3. The test data interpretations system 200 includes a prompt generation subsystem 300, summary generation subsystem 301, validation checks subsystem 302, a risk assessmentsubsystem 304, an actions subsystem 306, and a refinement subsystem 308. Each of these subsystems may perform any combination of embodiments described herein. For example, the subsystems may perform operations as described at least in method 400, method 500, method 600, method 700, method, 800, etc.
[0060] FIG. 4 is a flowchart of a method 400 for generating a report interpretation. Method 400 may include more or less operations than those shown in FIG. 4 and various operations may be performed in alternative configurations unless otherwise noted herein. Various embodiments of method 400 may be performed by the test data interpretations system 200 as described with respect to FIG. 2. Operation 402 includes receiving a request requesting a particular type of interpretation summary and identifying the test data to be used for the interpretation. For example, a healthcare professional may request a particular type of report based on the type of test that was performed on the patient. The healthcare professional may further specify that the report should include test data from the same type of test but taken a predetermined period of time removed from when the present test data was collected.
[0061] Method 400 may further include operation 404 including determining a set of sections to be included in the summary to be generated for the particular type of interpretation identified in the request received in operation 402. Each interpretation report may include one or more sections, and a healthcare professional may request that certain sections be included in the interpretation summary. For example, for a meal consumption interpretation summary, a healthcare professional may request only gastric activity sections be included in the interpretation summary and that sections relating to blood sugar or heart rate, for example, need not be included in the interpretation summary.
[0062] Operation 406 includes constructing a prompt for each section determined in operation 404. Operation 408 includes generating a summary for each section determined in operation 404 using an LLM and the prompt generated for the sections in operation 406. According to various embodiments, each section of the report is generated by inputting a dedicated prompt into the LLM, which includes style and formatting instructions, and programmatically-determined key points summarizing test result interpretations. The LLM then generates a concise sentence or paragraph. For each section, detailed guidelines on interpretation rules and prompt structure may be provided, along with examples of both an interpretation and a corresponding prompt. One LLM may be used to generate each summary for each section, according to some embodiments. In other embodiments, different LLMsmay be used to generate the summaries. For example, for an interpretation summary having three sections, a first LLM may be used to generate two summaries out of three summaries and a second LLM may be used to generate the third summary.
[0063] Rule-based interpretations may be used as part of the prompt construction as described in operation 406. Detailed guidelines on interpretation rules and prompt structure may be provided, along with examples of both an interpretation and a corresponding prompt. Exemplary rule-based interpretations are described below.
[0064] According to some embodiments, test quality may be evaluated with criteria for indicating “interpret with caution” precautions or other terminology to describe certain results. In some embodiments, if any of the quality metrics indicate “interpret with caution”, then the quality interpretation points start with this phrase, followed by an indented list of each result warranting caution. Else the quality interpretation starts with “Pass”. In both cases, the non-cautionary variables are then reported one per line if remarkable. The quality template may be dependent on the quality metrics as well. If the quality was pass, the template opens with "Test Quality: Pass. ", otherwise “Test Quality: Interpret with Caution [why?] ”
[0065] As other results are interpreted, if they are not cause for caution, but they are notable, they are included by appending the quality template for high non-caution interpretations. For example, if the percent of the session with artifacts result is between 15 and 50%, the quality template is appended with "[if artifacts mild or moderate, [Mild / Moderate] artifacts]."
[0066] In some embodiments, for spectral interpretation, the interpretation includes characterizing spectral metric averages across the entire session by comparing them to normative ranges, then evaluating hourly averages for the metrics and noting “transient” spectral abnormalities. For example, the exemplary spectral metrics may be labeled as “low”, “normal”, “borderline”, or “high” based on their values relative to published reference intervals.
[0067] In various embodiments, the post-meal hourly averages for Principal Gastric Frequency (PGF), Gastric Alimetry Rhythm Index (GA-RI), and BMLAdjusted Amplitude are translated to a categorical value and filters may be applied including removing normal hourly spectral metrics, removing hourly spectral metrics beyond the post-prandial duration, removing, hourly abnormal spectral metrics that match their overall interpretation category,etc. The overall values may be filled into a template defined in the reporting guidelines, followed by bullet points describing remarkable transient abnormalities.
[0068] In various embodiments, the symptom interpretation may include several broad assessments and then characterizes individual symptoms according to patient reported and observed events and interpretation maps from severity scores and event counts to interpretation categories.
[0069] According to various embodiments, the interpretation of the gut-brain survey may involve score inversion for certain questions and grouping into psychiatric domain. The total score and subscore for each domain is reported, along marked symptoms. For example, the interpretation may include computing total and section scores (do not mention sections with score 0) and noting pertinent questions / symptoms (ratings of 3 or 4 out of 4).
[0070] A conclusion interpretation may involve an overall assessment of whether the spectral analysis was normal or abnormal, followed by a synthesis of the previously described sections and adding clinical associations when phenotypes are observed. Phenotypes may be interpreted according to embodiments described in PCT Application No.PCT / IB2024 / 061366 entitled “Gastrointestinal Diagnostic Aid” which is incorporated by reference herein in its entirety and for all purposes. Further embodiments of phenotypes of results interpreted by embodiments described herein may include embodiments described in US Patent Application No. 18 / 933,661 entitled " Systems And Methods For Body Surface Colonic Mapping” which is incorporated by reference herein in its entirety and for all purposes.
[0071] Operation 410 includes generating an overall summary that includes the summaries generated in operation 408. For example, each summary of each section may include text portions and / or graphical information (such as a graph or the like) and the overall summary may include each of the generated text portions and / or graphical information. The overall summary may be generated by the LLM (or one of the LLMs) used to generate the summaries for the sections or the overall summary may be generated by a new LLM. In further embodiments, operation 410 may include generating a new text portion and / or graphical portion based on all of the summaries.
[0072] Operation 411 may include performing a set of one or more validation checks on the contents of the overall summary and take corrections actions, as needed, to correct any errors identified by the validation checks where each corrective action changes a portion of thecontents of the overall summary. Performing the set of one or more validations checks may be performed by a subsystem such as the validation checks subsystem 302 as shown in FIG.3.
[0073] Operation 412 may include taking corrections actions, as needed, to correct any errors identified by the validation checks where each corrective action changes a portion of the contents of the overall summary. Corrective actions may include regenerating the summary of the section with the same LLM that was previously used to generate the summary of the section, generating a new summary for the section using a different LLM that was previously used to generate the summary of the section, deleting the section, modifying the section (e.g., by modifying the prompt and / or the test data used), combining one or more sections, Any actions may be performed by a subsystem such as the actions subsystem 306 as shown in FIG. 3.
[0074] Operation 414 may further include performing risk assessment on the overall summary resulting from the processing performed in operation 412. Operation 414 may further include updating the overall summary, as needed, based upon the risk assessment. Risk assessments may be performed by a subsystem such as the risk assessment subsystem 304 as shown in FIG. 3. Operation 416 may further include enabling the overall summary resulting from operation 414 to be refined. Refining the overall summary may be performed by a subsystem such as the refinement subsystem 308 as shown in FIG. 3.
[0075] Method 400 may further include operation 418 including designating the overall summary resulting from the processing performed in operations 412, 414, and 416 as the final interpretation summary (alternatively referred to herein as the final interpretation report or the final guided interpretation report). Operation 420 may further include providing the final interpretation summary as a response to the request received in operation 402. The final interpretation summary may be provided to the healthcare professional who requested the interpretation summary. In various embodiments, the final interpretation summary may be provided to a downstream consumer such as the patient or the like.
[0076] FIG. 5 is a flowchart of a sub-method for generating a report interpretation. Method 500 is a sub-method of performing operation 412 in FIG. 4. Method 500 includes operation 502 for identifying a set of validation checks to be performed. The validation checks are further described in detail with respect to at least FIG. 6. Operation 504 includes, for each validation check identified in 502, a series of operations. Operation 506 includes performingthe validation check and operation 508 includes determining whether the validation check passed. If yes, method 500 proceed to operation 512 of marking the validation check as passed and then performing the next validation check in operation 514, as needed. If no, method 500 proceeds to determining whether to perform a corrective action in operation 510.
[0077] In various embodiments, a corrective action may be available to rectify the validation check failure. Operation 516 includes identifying one or more actions to be performed for corrective one or more errors identified by the validation check and operation 518 includes performing the one or more actions identified in operation 516. The action may include reperforming the validation check, reverting back to method 400 and modifying the prompt, the test data, the LLM used to generate the summary of the section and / or the overall summary, etc. Method 500 may proceed until all of the validation checks identified in operation 502 are marked as passed or otherwise bypassed. A failed validation check may be bypassed either manually (e.g., via a healthcare professional input) or due to a predefined exception defined by one or more conditions. For example, a failed validation check may be bypassed upon the condition that it is marked as failed.
[0078] FIG. 6 is a flowchart of a sub-method for generating a report interpretation. Method 600 illustrates a sub-method for performing operation 412 of FIG. 4 for performing a set of one or more validation checks on the contents of the overall summary. Validation checks, interchangeably referred to herein as guardrails, are programmatic tests for checking the LLM-generated summaries. These guardrails serve to directly mitigate the risks above. As described here, these function as a step in the summary generation, but they can also be used for automated evaluation. The guardrails can be implemented using a variety of deterministic approaches to checking the content of plain text, including string search, pattern matching and regular expressions, keyword and phrase presence with position tracking, unique word and phrase detection, structured result extraction, approximate string matching, etc.
[0079] The foregoing operations may be performed in any order and in any combination other than that shown in FIG. 6. Method 600 includes operation 602 includes checking key interpretation points. Operation 602 includes confirming that rule-based interpretations are accurately captured in the LLM generated summaries. For rule-based interpretation points and prompt templates that describe a result with 2 (binary) or several options (categorical) (e.g. “overall [ab]normal spectral analysis”), various tests using regular expressions areimplemented in operation 602 to extract the summary category and ensure it matches the rule-based category.
[0080] Table 1 includes exemplary interpretation points with categorical mappings for the quality, spectral, and conclusion report sections. For each row, the guardrail extracts the relevant value using a regular expression and ensures that it matches the rule-based interpretation point from the Interpretation Point column.
[0081] Table 1. Exemplary Interpretation Points.
[0082] Operation 604 includes checking numerical accuracy. For example, to avoid the LLM “filling in” a missing value or mutating a numeric test result, this guardrail checks for any numeric values that are mentioned in the LLM summaries but not in interpretation points (de novo numbers). This addresses the accuracy risk. An extension of this guardrail may include ensuring the accuracy of the context in which the numbers are used. For example, by searching for keywords or phrases around the number to ensure that the number appearing in the summary is being associated with the same metric that it is associated with in the prompt.
[0083] Operation 606 includes checking allowed words. A list of all words that are present in the summary but not included in the prompt may be generated. This list of words is compared against a predetermined list of acceptable words. The list of acceptable words may be generated from extensive experimentation and manual validation across a large historical database. Alternatively, the final implementation may instead use a list of excluded words, or some combination of these two lists.
[0084] Operation 608 includes checking the LLM confidence. According to some embodiments, when generating summaries, the LLM produces a per-token confidence score (calculated using the model’s predicted probability that the selected token is the correct next token). The average per-word, per-sentence, and per-section confidence scores can be compared to independent thresholds to identify portions of the summary that may have issues.
[0085] Operation 610 may include querying the LLM again to check the overall summary against approved interpretation logic and score. Operation 610 may include reperforming at least operation 602. Operation 612 may include checking the LLM generated in operation 610. In some embodiments, operation 610 may confirm phenotype associations and conclusions. For example, wherever phenotypes are mentioned in the LLM summaries, this guardrail ensures that the phenotype was included in the rule-based interpretations.
[0086] Method 600 may further include an operation including healthcare professional review. In some embodiments, before proceeding to a subsequent section, the healthcare professional may review and approve the LLM-generated summary. The healthcare professional is given the opportunity to edit the contents of the summary. Additional checks for unsafe phrases can be implemented via extensive review by healthcare professionals. For example, a cohort of tests from a historical database that represent a wide variety of phenotype combinations may be used to generate summaries and healthcare professionals may mark any phrases / sentences that are not desirable responses. Further keyphrase detection guardrails may be developed to ensure that likely undesired summaries are not permitted. This guardrail may rely on approximate string-matching techniques, such as word embeddings, N-grams with similarity measures, and / or fuzzy matching.
[0087] In some embodiments, for sections where the prompt includes vague instructions for the exact summary structure, an LLM may be used to further verify that the summary does not misinterpret or make any conclusions that are not directly implied by the information in the prompt. Specifically, for the final clinical summary section, which synthesizes background information on multiple phenotypes in some cases, an LLM would be prompted to assess whether or not the conclusions in the output summary are consistent with the information provided in the prompt. The prompt would include specific instructions to only output a “pass” or “fail.” This guardrail would then be validated (and possibly fine-tuned) on a collection of summaries that were manually assessed for correctness or manually modified to be incorrect.
[0088] For various operations of method 600, in the case of a failure, method 600 proceeds to operation 614 for performing corrective actions. A collection of actions may be used to handle rare instances in which the guardrails identify an error in a summary. Actions may include prompt resubmission, prompt resubmission with changed LLM settings, modifying the prompt, modifying the prompt and the original output, implementing few-shot learning, etc.
[0089] According to some embodiments, if a particular LLM-generated summary fails one of the guardrail tests, the system will first try one or more methods to regenerate a summary that is compliant with the guardrails. Given the randomness in LLMs, resubmitting the same prompt again may produce a different (and guardrail compliant) result. In some embodiments, changing the LLM parameters that determine the number of tokens to selectfrom, the distribution from which they are selected, the randomness in the output, etc., may enable the model to generate a summary that is different from the previous one and compliant with the guardrails. Example parameters may include temperature (e.g., how much randomness to introduce, 0 will yield the same result every time, while higher values (e.g., 0.8) will increase creativity and variance), repeat penalty (e.g., how much to discourage repeating the same token (e.g., 1.1)), min P sampling (e.g., minimum base probability for a token to be selected for output (e.g., 0.05)), top P sampling (e.g., minimum cumulative probability for the possible next tokens (e.g., 0.95)), top K sampling (e.g., limits the next token to one of the top-k most probable tokens (e.g., 40)), etc.
[0090] In further embodiments, the prompt may be modified depending on the guardrail that flagged the initial summary. For example, a numerical hallucination may trigger the addition of “don’t make up any numbers” in the prompt, or use of an unacceptable word may trigger “do not use words like Unacceptable word>.” The prompt may be modified to include any combination of the original prompt, additional context on the issues with the original output, the original output, and an instruction to make corrections to the original output.
[0091] In some embodiments, a modified prompt can be generated with fixed or dynamically determined (e.g., nearest neighbours based on test in question). For example, input / output pairs may be included in the prompt as examples. These input / output pairs can be stored in a database with verified outputs so as to ensure that they are good examples to use.
[0092] FIG. 7 is a flowchart of a sub-method for generating a report interpretation. Method 700 describes constructing the prompts for each section of the interpretation summary. In various embodiments, a prompt template may be identified for each section identified in operation 404 of FIG. 4. Method 700 describes further details of constructing the prompt (as in operation 406 of FIG. 4) using the prompt template. Operation 702 includes identifying a required segment in the prompt template identified in operation 404 of FIG. 4. A required segment may refer to a segment that is required based on the request for the interpretation summary. A required segment may also refer to a segment that cannot be removed from the prompt template. For example, if the healthcare professional requests an interpretation summary for a gastrointestinal phenotype associated with sensor data from an array applied to the abdomen or other skin surface of the patient, a required segment of the prompttemplate may include a phenotype analysis (e.g., text or visual data representing whether the test data falls within one or more phenotypes). In some embodiments, each prompt template may include at least one required segment. In other embodiments, the prompt template includes only required segments. In yet further embodiments, a prompt template may only include conditional segments, to be described in detail below.
[0093] Operation 704 include identifying zero or more conditional segments in the prompt template identified in operation 404 of FIG. 4. A conditional segment may refer to a segment that a healthcare professional may choose to include or not to include in the prompt template. For example, a healthcare professional may choose to delete a segment relating to heart rate if they are only interested in a segment related to gastric activity. Again, a prompt template may not include any conditional segments and only includes required segments, according to various embodiments. Operation 706 includes determining, based upon the test data, which, if any, conditional segments from the conditional segments identified in operation 704 are to be included for the prompt. For example, the test data may only be relevant to certain conditional segments and not others. Accordingly, the irrelevant conditional segments may be removed from the prompt template. Further, a healthcare professional may further decide to remove conditional segments as desired. Operation 708 includes constructing a prompt that includes the required segment identified in operation 702 and each conditional segment determined in operation 706.
[0094] Operation 710 further includes, for each variable portion in the prompt constructed in operation 708, determining based upon the test data, content to be inserted into the variable portion. Operation 712 may further include, for each variable portion in the prompt constructed in operation 708, replace the variable portion with the content determined for the variable portion in operation 710. Operation 714 may further include identifying the prompt resulting from the processing in operation 712 as the prompt to be used for the section.
[0095] FIG. 8 is a flowchart of a method for risk assessment. In the event that the summary regeneration techniques fail to generate a guardrail-compliant summary, one or more of the operations of method 800 may be used to enable the healthcare professional to safely use the system. For example, method 800 may include highlighting the values or sections, removing the values or sections, requiring manual review, requiring external expert review, etc., to be described in further detail below. Operation 802 includes evaluating whether an incorrect output could lead to severe patient harm. If yes, method 800 may proceed to operation 704including highlighting the section and displaying a “warning” statement to the user. In some embodiments, the values or sections in question can be highlighted to show the healthcare professional that there may be an issue with a portion of the summary. The indication can provide context for the risk, for example, saying that a value may be inaccurate, a conclusion may not be consistent with approved uses, etc. Furthermore, the severity of the highlighting (e.g., whether there is a simple note, a warning, or a caution) may be context-specific, depending on the exact nature of the guardrail that initiates the need for highlighting and that guardrail’s relationship to identified risks to patients.
[0096] Operation 806 may include requiring expert review. In addition to highlighting or removing a value or section, a healthcare professional may approve or edit the values in question before proceeding. Furthermore, before presenting the summary to the healthcare professional who prescribed the test, the summary may be approved by an external healthcare professional who is trained in assessing the accuracy of the interpretation. Operation 808 includes reevaluating whether an incorrect output could lead to moderate patient harm. If yes, highlight the section may be upgraded to display a “caution” statement to a user which indicates a higher level of risk. Operation 812 may include further expert review and approval based on the increased level of risk. In further embodiments, values or sections containing issues can be removed. The summary may include an explanation as to why they were removed.
[0097] According to some embodiments, the test quality section may produce an “interpret with caution” result based on one or more measures of test quality, each of which has different implications for how the test results should be interpreted. The table below contains exemplary sources of an interpret with caution warning, the warning that is displayed on the report, and the corresponding summary sections that would be implicated:<><>>>
[0098] In each of these instances, the relevant text in the summary can be formatted to direct the healthcare professional ’s interpretation. For example, by highlighting the relevant text (as in the right column) and displaying the warning (as in the middle column) when the healthcare professional hovers over the highlighted text or requiring the clinician to acknowledge the warning on each highlighted section.
[0099] According to various embodiments, the summaries should contain all information needed to properly interpret the report. This includes, for example, noting that a test should be interpreted with caution due to poor quality or that a particular metric falls outside of a normative reference interval. The summaries should contain accurate information and should be checked for unwanted generalizations and / or hallucinations. The scope of the information contained in the summaries should be consistent with any provided instructions for use and the claims approved by relevant regulatory authorities. For example, the summary should not contain diagnoses unless the device is approved as a diagnostic device.
[0100] By implementing multiple, independent layers of defence, the likelihood of a failure slipping through all layers is greatly reduced. In the context of the medical-grade LLM-based system for generating report summaries, the following layers of risk mitigation are applied structured, controlled input prompts which ensure input consistency guiding the LLM toward predictable and accurate outputs, deterministic, testable guardrails including automated checks that catch specific errors, such as missing key interpretations or incorrect numerical data, and ensuring adherence to expected behaviour where it can directly be assessed. Further layers of risk mitigation as described herein include validated LLM-based guardrails which provide additional validation by a separate LLM to verify the correctness of the summary and prevent misinterpretation or hallucinations, integrated alerts / cautions / warnings where the user interface flags potential issues, allowing healthcare professionals to review or correct content before use, and healthcare professional oversight via final review and approval by the healthcare professional.
[0101] According to various embodiments, the overall summary may be further refined. A healthcare professional may want to refine the summaries even when it is guardrail compliant. In some embodiments, the summary may be presented to the healthcare professional in an editable text box, allowing the healthcare professional to directly edit the text before accepting the summary and proceeding to the next section. The healthcare professional may further have the option to regenerate the summary from a selection of prompt refinements that may or may not be section specific. General refinements may include “include more detail,” “explain these results,” “be more concise,” “only list abnormalities,” etc. Specific refinements may include “comment on all symptoms,” “add comment on heartbum,” “exclude phenotypes,” etc. These options may also be paired with the ability to invoke them on a specific portion of text, for example via highlighting the text. With this feature, there will still be a limited number of options available to ensure that the system can still be comprehensively tested and validated for its safety as a medical device.
[0102] According to various embodiments, healthcare professionals may submit feedback, for example via preference / satisfaction logging (e.g., thumbs up / thumbs down) or manual guardrail triggering (e.g., “this value is incorrect”, “this word should be excluded”, “this conclusion was missing”). Furthermore, healthcare professionals may manually add words, phenotypes, conclusions, etc., that should be included or excluded and can be implemented using the same guardrail functionalities that are already in place. The automatic refinements based on user selection may be manually or automatically added to a healthcare professionalprofile. For example, a healthcare professional could set their preferences such that their summaries always include the “be more concise” refinement. Additionally, the system could estimate the probability that a healthcare professional will select a particular refinement given the test data (e.g., that they always select “be more concise” when all metrics are normal) and automatically include these refinements in the prompt.
[0103] Example 1.Example of a test where all spectral metrics were within the reference intervals, but the patient exhibited symptoms when the stomach was active.Symptoms PromptGut-Brain Wellbeing PromptConclusion PromptExample 1 Summary
[0104] Example 2.Example of a patient with normal overall spectral metrics, but transient abnormalities and a delayed response of the gastric activity to the mealQuality PromptSpectral PromptSymptom PromptGut-Brain Wellbeing PromptConclusion PromptExample 2 Summary
[0105] Results
[0106] To verify the reliability of the pipeline architecture described herein, the system was evaluated against a cohort of 1,133 historical test samples that were excluded from the initial development and LLM distillation phases. The performance of the LLM summarization engine was measured by its ability to accurately reflect the structured interpretation points generated by the rule-based interpretation engine.
[0107] In total, 19 interpretation point tests were conducted across the quality, spectral, and conclusion sections of the report to extract the summary category and ensure it matched the rule-based category. The summary accuracy for most quality-related interpretation points was 100%, including overall pass / fail, impedance category, artifact category, meal type, meal completion binary, standard duration, short preprandial duration, long preprandial duration, short postprandial duration, and long postprandial duration. For BMLrelated quality points, accuracy was 99.7% (1,128 of 1,131), with the system correctly mentioning BMI when high or borderline and appending "within limits" for borderline BMI.
[0108] In the spectral analysis domain, interpretation points for Principal Gastric Frequency (PGF) overall, Gastric Alimetry Rhythm Index (GA-RI) overall, and BMI-adjusted amplitude all achieved 100% accuracy. The meal response category showed 99.9% accuracy, and meal response lag was 99.8% accurate. The overall spectral assessment in the conclusion section was 100% accurate.
[0109] The effectiveness of the LLM for summary evaluation was evaluated based on its ability to identify semantic inconsistencies or hallucinations (e.g., symptom phenotype mix-ups) that might bypass deterministic checks. Across 4,111 report sections, the evaluation LLM maintained an overall alarm rate of approximately 0.75%. Human review of the evaluation LLM’s flags confirmed a positive predictive value (PPV) greater than 60% across all clinical sections, with the spectral analysis section achieving a peak PPV of 75%.Furthermore, the system identified a strong correlation between secondary LLM flags and lower token-level log probability scores, confirming that model uncertainty often coincides with semantic errors.
[0110] FIGS. 9A-9E illustrate various guidelines for interpreting test results. Test interpretation may include checking test technical quality, as shown in FIG. 9 including (A) Checking impedance for electrode signal quality is ‘good’ for at least half the electrodes; (B) Checking meal completion is above 50%; (C) Checking proportion of artifacts is less than 50%; (D) Checking app usage was at least every 15 min; (E) Checking raw signal traces for uncertainties in artifacts.[OHl] For (a), the impedance of the skin-electrode interface is a key determinant of signal quality. If signal quality (good / marginal / poor) is ‘good’ for at least half of the electrodes, this is considered a pass, with marginal electrodes considered acceptable. However, if signal quality is marginal or poor across a majority of channels, then the test should be interpreted with caution. The key risk in this context is that motion artifacts may be accentuated in the presence of poor impedance. For (b), if meal completion is <50%, the test should be interpreted with caution. This determination was based on a sensitivity analysis revealing that a half-sized portion was sufficient to trigger meal responses and reliably detect dysrhythmic. For (c), a mobile application notifies the patient to update their symptoms every 15 min. If the patient interacts with the application infrequently, the symptom data may be compromised and should be interpreted with caution. For (d), artifacts are automatically detected and corrected using the onboard accelerometer and validated algorithms. Time periods where artifacts were detected are shown by the ‘Artifact Detected’ bar. Excessive artifacts occur when the patient moves, tenses their abdominal muscles, talks and / or laughs,leading to poor data quality or data loss. If artifacts are present in >50% of the study period, the test should be interpreted with caution. When artifacts are severe, the data may not be plotted. For (e), the signal traces are consulted when there is uncertainty about whether artifacts have significantly affected the signal.
[0112] FIGS. 10A-10C illustrate various guidelines for interpreting test results. The spectral analysis produces a spectrogram (graphical representation of the signal amplitude at different frequencies across time) and associated metric tables. FIG. 10 illustrates (A) Normal reference intervals for as generated from a large database of healthy adults from diverse demographics (n = 110). Four independent spectral metrics are defined with reference to the standardized 4.5 h test protocol: Gastric Alimetry Rhythm Index (GA-RI), Principal Gastric Frequency, Fed:Fasted Amplitude Ratio and Average Amplitude; (B) Assess amplitude curves for meal response: note that the high fasting baseline is a common normal variant; (C) Assess for transient abnormalities that may not have been detected in the overall summary metrics.
[0113] A first metric includes Principal Gastric Frequency (cpm) [Reference interval 2.65-3.35 cpm]. The intrinsic gastric frequency is the dominant feature of the spectrogram. It is observed in normal tests as a distinct horizontal yellow band in the spectrogram and reported in cycles per minute (cpm). Legacy EGG methodologies defined the normal gastric frequency range as 2-4 cpm. The Principal Gastric Frequency is more refined than previous approaches, with normative reference intervals lying within a narrow range of 2.65-3.35 in healthy adults. Small deviations outside this range may be normal, and while females show a slightly higher frequency than males, they are currently assessed using the same range. In legacy EGG, dysrhythmias were defined by frequency abnormalities, with ‘bradygastric’ and ‘tachy gastric’ frequencies found in association with diverse gastric disorders. However, with the robust separation of frequency and rhythm parameters in BSGM, together with signal-processing advances, isolated deviations in frequency are much less commonly identified in Gastric Alimetry reporting. However, frequency elevation (rarely observed to >4 cpm) may be seen in long-term diabetes, hypothesized to reflect autonomic neuropathy, and also in vagal injury. Low frequencies (rarely observed to <2.2 cpm) may be associated with intrinsic gastric pacemaker dysfunction or surgical resections. Abnormalities may not be sustained throughout the entire meal response and can exist transiently. A Principal Gastric Frequency is not reported when the rhythm stability is low or falls below a critical threshold, indicated by a (-) in the metric table.
[0114] A second metric includes BMI-Adjusted Amplitude (pV) [Reference interval 22-70 pV], The amplitude of the gastric signal is corrected for BMI in the Gastric Alimetry system and is reported as microvolts (pV). Based on classical EGG data, it is plausible that sustained high amplitudes (or sustained activity of normal amplitude in the presence of delayed gastric emptying) could be associated with gastric outlet resistance. Low amplitudes may be associated with hypomotility and / or neuromuscular dysfunction. It is also important to note that opiates could reduce gastric amplitudes or induce transient dysrhythmias, meaning that these drugs should be ceased at least 24 h prior to testing when possible.
[0115] A third metric includes Gastric Alimetry Rhythm Index (GA-RI): [Reference interval > 0.25], GA-RI is a measure of stability (between 0-1) of gastric activity and quantifies the extent to which activity is concentrated within a normal principal frequency band over time, relative to the residual spectrum. Higher values indicate greater stability, whereas lower values indicate greater spectral scatter. GA-RI is not reported when the amplitude falls below a threshold of <10 pV (indicated by a (-)). A low GA-RI is the biomarker for dysrhythmia and is currently considered to be a key feature indicative of a gastric neuromuscular disorder, which likely reflects impaired slow-wave generation and coordination in the presence of underlying ICC network impairment.
[0116] A fourth metric includes Fed:Fasted Amplitude Ratio (ff-AR): >1.08. A meal response is indicated by the increase in signal power after the test meal compared to before the meal, which is calculated as a ratio of the maximum amplitude in any single 1-h postprandial period to the amplitude in the pre-prandial period (ff-AR). During reference range development, it was found that approximately 30% of patients showed a ‘high fasting baseline’ amplitude, such that the reference range cut-off was low (>1.08). The ff-AR metric is therefore not considered a reliable indicator of gastric dysfunction in isolation and is used solely as a supporting metric for an abnormal test in combination with other metrics.
[0117] It should be noted that transient abnormalities in the spectral metrics can also occur as shown in portion (c). Such abnormalities will be captured in the hourly reported metrics but may be associated with normal metrics for the overall time period. Assessment of transient abnormalities may be performed on a case-by-case basis. For example, low amplitude or GA-RI before a meal is expected, whereas an hour of high or low frequency activity or low GA-RI immediately after the meal may be indicative of gastric dysfunction, even if it is followed by normal activity.
[0118] Meal response curves that show a delayed rise and / or do not return to baseline may be suspicious for gastric dysfunction; however, dedicated studies addressing meal response curves are still awaited before diagnostic utility can be ascertained. In the initial classification scheme proposed by the BSGM working group, five spectral phenotypes have been described: dysrhythmic (GA-RI < 0.25), low-amplitude (BMI-adjusted amplitude < 22 pV), high-amplitude (BMI-adjusted amplitude > 70 pV), high-frequency (frequency > 3.35 cpm); and low-frequency (frequency < 2.65 cpm). In the context of the present disclosure, a 'phenotype' may refer to a distinct classification of gastric function derived from a cluster of spectral metrics, symptom metrics, and / or other test outputs.
[0119] FIGS. 11 A-l IE illustrate various guidelines for interpreting test results. When spectral analysis is abnormal, the symptom analysis provides complementary data. When the spectral analysis is normal, specific symptom phenotypes may be identifiable in over half of cases which link to gastric activity patterns. Symptom analysis includes both the pattern and severity of individual symptoms. FIG. 11 illustrates (A) Assess for symptom baseline; (B) Assess whether symptoms are meal-responsive or meal non-responsive; (C) Assess the symptom curve pattern: declining curve, continuous curve or late up trending curve; (D) Assess for correlation between symptom curves and gastric amplitude; (E) Assess the timing, type and number of discrete symptom events.
[0120] It should be noted whether symptoms are present before the meal (including type and severity), followed by an assessment of how the symptoms changed in relation to the meal. The presence of early satiation should be noted as a marker of post-prandial distress, which is assessed as a single time-point symptom immediately after the meal (scored out of 10). Meal-responsive symptoms either increase after the meal and decline over time or increase with the meal and then remain constant. A symptom curve that increases then decreases in profile has been described in association with gastric emptying decay curves, with symptoms abating as food transitions to the small intestine, therefore being a strong indicator that the relevant symptoms have a gastric origin. Alternatively, symptoms may remain relatively continuous throughout the test, which has been associated with a higher frequency of gut-brain axis (centrally mediated) disorders and vagal neuropathy in published series.
[0121] If symptoms trend upwards late into the test, this may suggest a ‘post-gastric’ (small intestine) symptom origin, with symptom burden progressively increasing as a greatervolume of contents progress beyond the pylorus. Symptom curves can also present as mixed profiles, and work is ongoing to further characterize these symptom profiles (refer Tips and Pitfalls). It should be noted whether symptoms are present before the meal (including type and severity), followed by an assessment of how the symptoms changed in relation to the meal. The presence of early satiation should be noted as a marker of post-prandial distress, which is assessed as a single time-point symptom immediately after the meal (scored out of 10).
[0122] Symptom and gastric amplitude curves can be assessed together, to determine whether they are correlated, which may indicate visceral hypersensitivity. This assessment can be aided by the total symptom burden bar, which is shown directly under the spectral map in the Gastric Alimetry report. Symptom curves may also show correlations with transient spectral abnormalities. Timing, type and number of symptom ‘events’ (vomiting, reflux and / or belching) should be assessed. The timing of these events can also be correlated with the gastric amplitude.
[0123] Various embodiments of the present disclosure may be applied to test data collected from an electrode array patch disposed over a skin surface of the patient. The skin surface may include at least one of an abdomen, a torso, or a flank of the patient. Accordingly, the electrode patch array may gather data from any portion of the gastrointestinal tract of the patient and any region / quadrant of the abdomen of the patient.
[0124] FIG. 12 illustrates a gastrointestinal tract of a patient. A gastrointestinal tract 1100 is made up of organs that food and liquids travel through when they are swallowed, digested, absorbed, and leave the body as feces. The gastrointestinal tract 1100 may include the mouth 1102, the pharynx (throat) 1104, the esophagus 1106, the stomach 1108, the small intestine 1110, the large intestine 1112, the rectum 1114, and the anus 1116. Furthermore, the large intestine 1112, referred to interchangeably as the colon in various embodiments, includes the cecum 1118, the ascending colon 1120, the transverse colon 1122, the descending colon 1124, the sigmoid colon 1126, and the sigmoid colon 1128. The gastrointestinal tract 1100 may include the jejunum 1130. A gastrointestinal tract of a patient according to various embodiments of the present disclosure may include more components than those shown in FIG. 12.
[0125] FIG. 13 illustrates regions and quadrants of an abdomen of a patient. The abdomen 1200 of a patient may be divided into regions 1201 and / or quadrants 1203 to describe thelocation of an organ or structure. The regions 1201 of the abdomen 1200 may include the right hypochondriac / hypochondrium region 1202, the epigastric / epigastrium region 1204, the left hypochondriac / hypochondrium region 1206, the right lumbar / flank / latus / lateral region 1208, the umbilical region 1210, the left lumbar / flank / latus / lateral region 1212, the right inguinal / iliac region 1214, the hypogastric / suprapubic region 1216, and the left inguinal / iliac region 1218. The quadrants 1203 of the abdomen 1200 may include the right upper quadrant 1220, the left upper quadrant 1222, the right lower quadrant 1224, and the left lower quadrant 1226. As shown in FIG. 13, various components of the gastrointestinal tract 1100 described with respect to FIG. 12 are considered to be at least partially disposed in the abdomen 1200 of the patient.
[0126] To those skilled in the art to which the invention relates, many changes in construction and widely differing embodiments and applications of the invention will suggest themselves without departing from the scope of the invention as defined in the appended claims.
[0127] This invention may also be said broadly to consist in the parts, elements and features referred to or indicated in the specification of the application, individually or collectively, and any or all combinations of any two or more of said parts, elements or features, and where specific integers are mentioned herein which have known equivalents in the art to which this invention relates, such known equivalents are deemed to be incorporated herein as if individually set forth.
Claims
WHAT IS CLAIMED IS:
1. A method for generating a report interpretation summary, the method comprising:receiving a request to generate a summary for test data associated with gastric activity of a patient over a predetermined time period, wherein the test data comprises electrical signals that are measured with an electrode array patch disposed over a skin surface of the patient;identifying a plurality of sections to be included in the summary; andfor each section in the plurality of sections:constructing a prompt for the section based upon a prompt template configured for the section;providing the prompt as input to an LLM; andresponsive to the providing, generating by the LLM a summary for the section;generating an overall summary that includes the summaries generated for the plurality of sections; andperforming a set of one or more validation checks to check contents of the overall summary.
2. The method of claim 1, further comprising:for each validation check in the set of one or more validation checks, upon identifying one or more errors in the overall summary as a result of performing the validation check, performing one or more actions to rectify the one or more errors, wherein performing the one or more actions causes at least a portion of the contents of the overall summary to be changed resulting in an updated overall summary;generating a final overall summary using the updated overall summary; and providing the final overall summary as a response to the request.
3. The method of claim 1, wherein constructing the prompt further comprises: identifying an LLM to be used for generating a summary for the section; providing the constructed prompt as an input to the LLM; andoutputting, by the LLM, a summary for the section.
4. The method of claim 1, wherein constructing the prompt further comprises: identifying at least one required segment in the prompt template;identifying zero or more conditional segments in the prompt template; and determining, based upon the test data, which, if any, conditional segments from the conditional segments are to be included for the prompt construction.
5. The method of claim 1, wherein constructing the prompt further comprises: for each variable portion in the constructed prompt, determining, based upon the test data, content to be inserted in the variable portion; andfor each variable portion in the constructed prompt, replacing the variable portion with the content determined for the variable portion.
6. The method of claim 1, wherein constructing the prompt further comprises: processing the test data using a deterministic rule-based interpretation engine to generate a plurality of structured interpretation points, wherein the structured interpretation points comprise data derived from applying predefined clinical rules, logical conditions, or normative thresholds to the test data; andpopulating the prompt template with the structured interpretation points rather than raw test data.
7. The method of claim 2, further comprising:determining whether an identified error in the overall summary could lead to patient harm; andin response to determining that the identified error would lead to patient harm, updating the overall summary to indicate a warning.
8. The method of claim 2, further comprising:determining whether an identified error in the overall summary could lead to patient harm; andin response to determining that the identified error would lead to patient harm, requesting healthcare professional review of the identified error.
9. The method of claim 2, wherein the one or more actions comprise prompt resubmission, prompt resubmission with changed LLM settings, or modifying the prompt.
10. The method of claim 1, further comprising refining the overall summary, wherein refining the overall summary comprises manually editing the overall summary or regenerating the overall summary with prompt refinements.
11. The method of claim 1, wherein the skin surface of the patient comprises at least one of an abdomen, a torso, or a flank of the patient.
12. The method of claim 1, wherein the set of one or more validation checks comprises checking key interpretation points, checking numerical accuracy or checking allowed words.
13. The method of claim 1, wherein performing the set of one or more validation checks further comprises:constructing a verification prompt containing the generated summary and a set of approved interpretation logic;providing the verification prompt as input to the LLM or a second LLM to assess consistency between the generated summary and the approved interpretation logic; and flagging the summary for correction if the LLM outputs a failure indicator.
14. The method of claim 1, wherein performing the set of one or more validation checks further comprises:calculating a token-level log probability for each token in the generated summary; aggregating the token-level log probabilities to compute a confidence score for the generated summary; andautomatically triggering a corrective action if the confidence score falls below a predetermined confidence threshold.
15. The method of claim 6, wherein performing the set of one or more validation checks further comprises:extracting a first set of numerical values from the plurality of structured interpretation points;extracting a second set of numerical values from the generated overall summary; and identifying an error if the second set of numerical values contains a value not present in the first set of numerical values.
16. The method of claim 1, further comprising:determining a risk level associated with the generated summary; and in response to determining the risk level indicates potential patient harm, automatically modifying the overall summary to include a warning label associated with a specific section of the summary.
17. The method of claim 1, wherein the LLM comprises a distilled model, wherein the distilled model is trained using a training dataset generated by a foundational model having a parameter count larger than the distilled model, and wherein the distilled model is fine-tuned using a dataset of clinician-validated summaries.
18. The method of claim 1, further comprising:receiving a user input identifying a refinement to the generated overall summary; and updating a user profile associated with the request to automatically apply the identified refinement to subsequent prompt constructions associated with that user profile.
19. A system for processing gastric activity data comprising:an electrode array patch disposed over a skin surface of a patient for measuring electrical signals associated with gastric activity of the patient over a predetermined time period; anda processor configured to:receive a request to generate a summary for test data associated with gastric activity of a patient over a predetermined time period, wherein the electrical signals are measured with an electrode array patch disposed over a skin surface of the patient;identify a plurality of sections to be included in the summary; andfor each section in the plurality of sections:construct a prompt for the section based upon a prompt template configured for the section;provide the prompt as input to an LLM; andresponsive to the providing, generate by the LLM a summary for the section;generate an overall summary that includes the summaries generated for the plurality of sections; andperform a set of one or more validation checks to check contents of the overall summary.
20. The system of claim 19, wherein the processor is further configured to: for each validation check in the set of one or more validation checks, upon identifying one or more errors in the overall summary as a result of performing the validation check, perform one or more actions to rectify the one or more errors, wherein performing the one or more actions causes at least a portion of the contents of the overall summary to be changed resulting in an updated overall summary;generate a final overall summary using the updated overall summary; and provide the final overall summary as a response to the request.
21. The system of claim 19, wherein constructing the prompt further comprises: identifying an LLM to be used for generating a summary for the section; providing the constructed prompt as an input to the LLM; andoutputting, by the LLM, a summary for the section.
22. The system of claim 19, wherein constructing the prompt further comprises: identifying at least one required segment in the prompt template;identifying zero or more conditional segments in the prompt template; and determining, based upon the test data, which, if any, conditional segments from the conditional segments are to be included for the prompt construction.
23. The system of claim 19, wherein constructing the prompt further comprises: for each variable portion in the constructed prompt, determining, based upon the test data, content to be inserted in the variable portion; andfor each variable portion in the constructed prompt, replacing the variable portion with the content determined for the variable portion.
24. The system of claim 20, wherein the processor is further configured to: determine whether an identified error in the overall summary could lead to patient harm; andin response to determining that the identified error would lead to patient harm, update the overall summary to indicate a warning.
25. The system of claim 20, wherein the processor is further configured to: determine whether an identified error in the overall summary could lead to patient harm; andin response to determining that the identified error would lead to patient harm, request healthcare professional review of the identified error.
26. The system of claim 20, wherein the one or more actions comprise prompt resubmission, prompt resubmission with changed LLM settings, or modifying the prompt.
27. A non-transitory computer-readable medium storing instructions executable by one or more processors for causing the one or more processors to perform operations comprising:receiving a request to generate a summary for test data associated with gastric activity of a patient over a predetermined time period, wherein the test data comprises electrical signals that are measured with an electrode array patch disposed over a skin surface of the patient;identifying a plurality of sections to be included in the summary; andfor each section in the plurality of sections:constructing a prompt for the section based upon a prompt template configured for the section;providing the prompt as input to an LLM; andresponsive to the providing, generating by the LLM a summary for the section;generating an overall summary that includes the summaries generated for the plurality of sections; andperforming a set of one or more validation checks to check contents of the overall summary.