Systems and methods for genetic test selection through artificial intelligence and / or large language models
The described system addresses the inefficiencies in genetic test selection by using AI and large language models to extract and analyze phenotypic data, resulting in improved accuracy and efficiency of genetic test recommendations and laboratory operations.
Patent Information
- Application Number
- PCT/US2024/057333
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-24
- Filing Date
- 2024-11-25
- Publication Date
- 2025-05-30
AI Technical Summary
The current process of selecting genetic tests is inefficient due to the limitations of test requisition forms and the burden of reviewing extensive clinical histories to extract pertinent phenotypic information, leading to incomplete knowledge for laboratory teams and challenges in filtering relevant genetic variants.
A computing system that uses artificial intelligence and large language models to recommend genetic tests by extracting relevant data from multiple databases, generating a condensed phenotypic description, and displaying pertinent patient information through an expandable user interface, facilitating accurate genetic test selection and result interpretation.
This solution streamlines the genetic test selection process, enhances the accuracy of test recommendations, and improves the efficiency of laboratory operations by providing detailed phenotypic information and facilitating comprehensive result interpretation.
Smart Images

Figure US2024057333_30052025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR GENETIC TEST SELECTION THROUGH ARTIFICIAL INTELLIGENCE AND / OR LARGE LANGUAGE MODELSBACKGROUND
[0001] In its current state, genetic test selection is done during clinical consultation where a summary of the clinical presentation is reviewed to assess if there is a potential genetic etiology of the phenotypic presentation (diagnosis / prognosis) or if there are therapeutic options for treatment that would be informed by the results of genetic testing. As the understanding of genetic disease and phenotypic expression continues to grow, the involvement of genetic testing to rule in or rule out diagnoses and inform patient management will continue to grow. With a genetic testing landscape that spans targeted gene panel testing incorporating genes that are known to contribute to a specific disease state all the way to exploratory genome testing (e.g. chromosomal microarray, exome sequencing, genome sequencing, etc), the information needed to select the appropriate test given the constellation of phenotypic presentations a patient may exhibit is currently not an easy task. The current process is driven primarily by a provider searching the genetic test registry (GTR) website or laboratory specific testing catalogs to find the right test(s) for the patient’s presentation.
[0002] Patient phenotypic information currently is provided to the laboratory on test requisition forms that have inherent limitations for providing thorough and complete information. In lieu of these test requisition forms, the laboratory will also receive pages of patient clinical history / notes that require personnel in the laboratory to review and extract pertinent phenotypic terms to facilitate appropriateness of genetic test selection as well as accurate and relevant analysis of genetic results and results interpretation. As this process can be burdensome, a significant proportion of the testing performed in the lab is completed without receipt of this information which creates a significant gap in knowledge for the laboratory team and the tools utilized to filter to the most relevant reportable variants specific to the patient’s phenotype.
[0003] The systems and methods disclosed herein provide solutions to these problems and may provide solutions to the ineffectiveness, insecurities, difficulties, inefficiencies, encumbrances, and / or other drawbacks of conventional techniques.SUMMARY
[0004] The present embodiments may relate to, inter alia, determining a recommendation for a genetic test for a patient.
[0005] In one aspect, a computing system for determining a recommendation for a genetic test for a patient may be provided. The system may include one or more processors, and one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: (1) receive, via the one or more processors, a patient identifier (ID) of the patient and test code identifier of the genetic test ordered for the patient; (2) extract, via the one or more processors, from multiple databases, previous testing histories, relevant concurrent tests, family history, internal notes, reason for testing, as well as other data elements associated with the ordered test to be evaluated; (3) extract, via the one or more processors, from a patient database, patient features based on the patient ID; (4) generate, via the one or more processors, a condensed phenotypic description of the patient using contextualized targeted clinical documents; (5) generate, via the one or more processors, a list of relevant phenotype ontology terms based on the condensed phenotypic description of the patient; (6) display, via the one or more processors, the ordered test information, condensed phenotypic description, and phenotype ontology terms of the patient; genetic test; (7) generate, a contextualized database to provide a chatbot service to facilitate further exploration of the medical record by laboratory experts; and (8) enable, genetic test validation of the existing test order(s) or inform new genetic test recommendations by laboratory personnel by displaying all of the pertinent patient information via an expandable and comprehensive user interface (UI). The computing system may include additional, less, or alternate functionality, including that discussed elsewhere herein.
[0006] In another aspect, a non-transitory computer-readable medium for determining a recommendation for a genetic test for a patient may be provided. The non-transitory computer- readable medium may have stored thereon instructions that when executed, cause a computer to: (1) receive, via the one or more processors, a patient identifier (ID) of the patient and test code identifier of the genetic test ordered for the patient; (2) extract, via the one or more processors, from multiple databases, previous testing histories, relevant concurrent tests, family history, internal notes, reason for testing, as well as other data elements associated with the ordered testto be evaluated; (3) extract, via the one or more processors, from a patient database, patient features based on the patient ID; (4) generate, via the one or more processors, a condensed phenotypic description of the patient using contextualized targeted clinical documents; (5) generate, via the one or more processors, a list of relevant phenotype ontology terms based on the condensed phenotypic description of the patient; (6) display, via the one or more processors, the ordered test information, condensed phenotypic description, and phenotype ontology terms of the patient; genetic test; (7) generate, a contextualized database to provide a chatbot service to facilitate further exploration of the medical record by laboratory experts; and (8) enable, genetic test validation of the existing test order(s) or inform new genetic test recommendations by laboratory personnel by displaying all of the pertinent patient information via an expandable and comprehensive user interface (UI). The non-transitory computer-readable medium may include additional, less, or alternate functionality, including that discussed elsewhere herein.
[0007] In yet another aspect, a computer- implemented method for determining a recommendation for a genetic test for a patient may be provided. In one example, the method may include: (1) receive, via the one or more processors, a patient identifier (ID) of the patient and test code identifier of the genetic test ordered for the patient; (2) extract, via the one or more processors, from multiple databases, previous testing histories, relevant concurrent tests, family history, internal notes, reason for testing, as well as other data elements associated with the ordered test to be evaluated; (3) extract, via the one or more processors, from a patient database, patient features based on the patient ID; (4) generate, via the one or more processors, a condensed phenotypic description of the patient using contextualized targeted clinical documents; (5) generate, via the one or more processors, a list of relevant phenotype ontology terms based on the condensed phenotypic description of the patient; (6) display, via the one or more processors, the ordered test information, condensed phenotypic description, and phenotype ontology terms of the patient; genetic test; (7) generate, a contextualized database to provide a chatbot service to facilitate further exploration of the medical record by laboratory experts; and (8) enable, genetic test validation of the existing test order(s) or inform new genetic test recommendations by laboratory personnel by displaying all of the pertinent patient information via an expandable and comprehensive user interface (UI). The method may include additional, fewer, or alternate actions, including those discussed elsewhere herein.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Advantages will become more apparent to those skilled in the ail from the following description of the preferred embodiments which have been shown and described by way of illustration. As will be realized, the present embodiments may be capable of other and different embodiments, and their details are capable of modification in various respects. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive.
[0009] The figures described below depict various aspects of the applications, methods, and systems disclosed herein. It should be understood that each figure depicts an embodiment of a particular aspect of the disclosed applications, systems and methods, and that each of the figures is intended to accord with a possible embodiment thereof. Furthermore, wherever possible, the following description refers to the reference numerals included in the following figures, in which features depicted in multiple figures are designated with consistent reference numerals.
[0010] Figure 1 depicts an example computing environment in which the techniques disclosed herein may be implemented, according to some aspects.
[0011] Figure 2 illustrates an example computer-implemented method for determining a recommendation for a genetic test.
[0012] Figure 3 depicts a combined block and logic diagram in which exemplary computer- implemented methods and systems for training a machine learning (ML) chatbot or an artificial intelligence (Al) model are implemented according to some embodiments.
[0013] Figure 4 depicts an example screen including example order information.
[0014] Figure 5 depicts a diagram of an example extraction process.
[0015] Figure 6 depicts an example screen including an example chat session.
[0016] Figure 7 depicts an example diagram of an example implementation of generating a phenotypic description.
[0017] Figure 8 depicts an example screen displaying an example phenotypic description.
[0018] Figure 9 depicts an example phenotypic terms extraction process.
[0019] Figure 10 depicts an example screen including phenotypic terms.DETAILED DESCRIPTION
[0020] The present embodiments relate to, inter alia, determining a recommendation for a genetic test.Overview
[0021] Genetic data is becoming a central component to the management of patients and the portfolio of genetic tests is expanding daily as genetic science and technologies advance within the healthcare system. An unintended consequence of the fast-paced evolution of genetic testing is the confusion that can exist for ordering providers and laboratory clinicians regarding what test(s) to order and the increasing complexity of test result interpretation. Based on current estimates by World Economic Forum, the global genomic sequencing market is expected to grow from its previous $10.7 billion valuation in 2018 to $37.7 billion by 2026. A large medical institution’s pathology laboratory, such as a Department of Laboratory Medicine and Pathology (DLMP), test volumes and offerings are expected to reflect this growth as an increasing number of clinical sub-specialties utilize genetic testing for diagnosis, prognosis, and informing newer targeted therapies. Increased volumes are pivotal to the long-term success of the laboratory but create challenges with test selection in the growing landscape of available tests and will exacerbate the existing challenges of complex result interpretation. This reality creates a unique opportunity for the DLMP genomic lab to develop infrastructure using generative artificial intelligence (Al) and large language model (LLM), readily available to plug in to the disclosed LLM backbone infrastructure, to extract pertinent information efficiently and accurately from a patient’s electronic health records (EHR) data to generate decision support models to facilitate determination of the appropriateness of genetic test selection and identify and present detailed phenotypic information to incorporate for result interpretation. Some embodiments disclosed herein focus on patients for initial model generation and solution implementation. Some embodiments advantageously will allow for and facilitate secure access to external patients’ EHR data. Techniques disclosed herein will revolutionize the way genetic testing is performed, driving efficiencies in the laboratory and distinguishing the laboratory within a highly competitive genetic testing landscape.
[0022] Genetic testing is a highly competitive market with economic and performance metrics driving provider selection. A key barrier to the current state of molecular testing is the limitedclinical information supplied by a provider when ordering these tests that is (i) critical information utilized in assessing the ‘appropriateness of testing’ to optimize informed comprehensive molecular test recommendations based on DLMP’s evolving genetics testing portfolio and (ii) essential information needed for comprehensive analysis and result interpretation of complex genomic tests.
[0023] Solving these and / or other limitations would revolutionize the way genetic testing is performed by providing novel infrastructure for personalized genomic test selection (“the right test”) and driving efficiencies in the laboratory by having detailed patient- specific phenotypic information to facilitate analysis and results interpretation. Currently, there are no solutions that are readily available to perform the scope of work and provide the deliverables as in this disclosure.
[0024] To examine effectiveness of example implementations of the present techniques, we accessed clinical data from approximately 20,500+ unique internal molecular genetic test orders received in a large medical institution’s genomics lab. Access to these clinic records for internal patients allowed for comprehensive clinical data extraction.
[0025] Further, following receipt of a molecular test order by the genetic testing laboratory (DLMP) for a patient, we configured our model to extract identified key features from Mayo Clinic Cloud (MCC)-based identifiable clinical data resources, such as the unified data platform (UDP), etc. We validate the existing test order(s) or make informed molecular test recommendations based on the DLMP genetic testing portfolio. This data was reviewed by the laboratory relative to the test(s) ordered and recommended changes were made where appropriate. To facilitate appropriate laboratory results analysis and interpretation, the implemented models automated the creation of a detailed summarization of phenotypic descriptions of the patient and laboratory values with longitudinal context accessible to the lab staff using a chatbot.
[0026] Further, some embodiments included a combination of standard data query with a Retrieval Augmented Generation (RAG) pipeline built with LangChain and NeMo Guardrails.Example System
[0027] To this end, Figure 1 illustrates an exemplary computer system 100 for determining a recommendation for a genetic test for a patient in which the exemplary computer-implemented methods described herein may be applied. The high-level architecture includes both hardware and software applications, as well as various data communications channels for communicating data between the various hardware and software components.
[0028] Broadly speaking, the genetic test determining computing device 102 may determine a genetic test to be administered to a patient 199.
[0029] The genetic test determining computing device 102 may include one or more processors 120 such as one or more microprocessors, controllers, and / or any other suitable type of processor. The genetic test determining computing device 102 may further include a memory 122 (e.g., volatile memory, non-volatile memory) accessible by the one or more processors 120, (e.g., via a memory controller). The one or more processors 120 may interact with the memory 122 to obtain and execute, for example, computer-readable instructions stored in the memory 122. Additionally or alternatively, computer-readable instructions may be stored on one or more removable media (e.g., a compact disc, a digital versatile disc, removable flash memory, etc.) that may be coupled to the genetic test determining computing device 102 to provide access to the computer-readable instructions stored thereon. In particular, the computer-readable instructions stored on the memory 122 may include instructions for executing various applications, such as chatbot 124, genetic test determiner 125 (e.g., including phenotype fusion model 126), and / or machine learning (ML) training module 128.
[0030] The genetic test determining computing device 102 may further include display device 135. The genetic test determining computing device 102 may be operated by a genetic laboratory expert 139.
[0031] In operation, the chatbot 124 may generate a phenotypic description of the patient 199 (e.g., using a machine learning model). The chatbot 124 may include a large language model (LLM). The chatbot 124 may further participate in or initiate conversations (e.g., with the genetic laboratory expert 139 on display 135). For example, the chatbot 124 may answer questions about the genetic test. In some embodiments, the chatbot 124 is trained as describedwith respect to Figure 3 (e.g., via the ML training module 128). In some embodiments, the chatbot 124 comprises a LLM pre-trained, fine-tuned and / or trained using information from the patient database 180, the genetic test database 182, the historical patient phenotype database 184, and / or other sources.
[0032] For example, in some instances, the genetic test determining computing device 102 may: (i) identify, in the information from the historical patient phenotype database 184, historical patient phenotypes data 185a, and / or historical genetic tests data 185b; and / or (ii) train the chatbot 124 by fine-tuning the chatbot 124 using the historical patient phenotypes data 185a, and / or the historical genetic tests data 185b. This will be described in more detail elsewhere herein (e.g., with respect to Figure 3, etc.).
[0033] The genetic test determiner 125 may display information to genetic laboratory expert 139 on display 135, to enable appropriateness of genetic test order determination for genetic test patient 199. For example the genetic test determiner 125 may receive the phenotypic description of the patient 199 from the chatbot 124, or the clinician computing device 142, and output this information to genetic laboratory expert 139, who confirms genet test order or recommends a new genetic test.
[0034] In some embodiments, the genetic test determiner 125 comprises an Al or ML model or algorithm trained with historical patient phenotypes 185a (e.g., historical patient phenotypic descriptions) as inputs (e.g., also referred to as independent variables, or explanatory variables), and with historical genetic tests 185b (e.g., genetic tests given to historical patients of the historical patient phenotypes 185a) as outputs (e.g., also referred to as a dependent variables, or response variables). The Al or ML model or algorithm may be trained according to any suitable technique, such as supervised learning, unsupervised learning, or semi- supervised learning.
[0035] The genetic test determiner 125 may include phenotype fusion model 126. In some embodiments, the phenotype fusion model 126 may extract, from an output of the chatbot 124 (e.g., the phenotypic description of the patient, etc.), phenotypic terms. In turn, the genetic test determiner 125 may display the extracted phenotypic terms to allow genetic laboratory expert 139 to determine the appropriate genetic test. The phenotypic terms will be described elsewhere herein.
[0036] Such use of phenotype fusion model 126 improves technical functioning. For example, the phenotype fusion model 126 may communicate with the chatbot 124 (e.g., with an LLM of the chatbot 124, etc.). In some instances, advantageously, the LLM’s communication and / or formatting needs are defined, and the phenotype fusion model 126 provides the LLM with information accordingly. For example, regarding the communication needs, the LLM may have a predefined list of phenotypic terms, and the phenotype fusion model 126 provides the LLM with information according to the list (e.g., tags on conversations according to the list).Regarding the formatting, the phenotype fusion model 126 may provide the LLM with the phenotypic terms according to formatting specified by the LLM, thereby advantageously saving computing resources by avoiding a formatting conversion step.
[0037] Further advantageously, the chatbot 124 (or an LLM thereof) may have an application programming interface (API) endpoint to interface with the phenotype fusion model 126. This advantageously streamlines and improves communication between the chatbot 124 and the phenotype fusion model 126. In some embodiments, the phenotype fusion model 126 sends requests containing specific queries or instructions to the chatbot 124 via the API endpoint. The chatbot 124 may process the request, generate a response, and send it back. In some examples, the genetic test determiner 125, the phenotype fusion model 126, the clinician computing device 142, and / or the patient computing device 199 may call the API endpoint.
[0038] Additionally or alternatively, a middleware layer may be provided to process and / or format the data going to and from the chatbot 124 (or LLM thereof). This may help manage session data, handle errors, parse the output for specific information, and translate it into a format usable by phenotype fusion model 126 and / or chatbot 124. In some embodiments, the middleware implements batch processing of communications between the phenotype fusion model 126 and the chatbot 124, which advantageously reduces latency. In some implementations, the middleware applies asynchronous calls (e.g., programming operations that allow a program to initiate a task and continue executing other tasks without waiting for the initial task to complete), which advantageously also reduces latency.
[0039] Further regarding the middleware layer, the middleware layer may advantageously anonymize and / or encrypt data (e.g., from the clinician computing device 142 and / or patientcomputing device 199) sent to the chatbot 124. The data may be anonymized by removing patient names, addresses, etc.
[0040] In addition, API requests to the chatbot 124 may sometimes fail due to network issues, server downtimes, or model limitations. Advantageously, the middleware may implement errorhandling mechanisms, such as retry logic or fallbacks. The middleware may further help by managing resilience by catching errors, logging them, and retrying or escalating when needed.
[0041] Further advantageously, if the chatbot’s 124 output requires heavy processing, the middleware may apply transformations to minimize the load on the phenotype fusion model 126.
[0042] Still further advantageously, the middleware may improve performance of the chatbot 124. In particular, to maintain context (e.g., across a conversation between the genetic laboratory expert 139 and the chatbot 124, etc.), the middleware may store session data, manage user-specific or request-specific context, thereby improving continuity across requests.
[0043] Additionally or alternatively, a client may be implemented that connects to the chatbot’s 124 API endpoint. In some embodiments, this is an HTTP client implemented in languages, such as Python, or Java. Furthermore, in some embodiments, the API call is optimized, which advantageously reduces latency.
[0044] In some embodiments, when an output is sent from the chatbot 124 to the phenotype fusion model 126, the output is parsed and / or processed (e.g., by the middleware or other component). The format of the output may be further modified to suit the phenotype fusion model 126, which advantageously makes it easier for the phenotype fusion model 126 to apply tags corresponding to phenotypic terms to the output.
[0045] Advantageously, to improve communication between the chatbot 124 and the phenotype fusion model 126, a library common to both the chatbot 124 and the phenotype fusion model 126 may be implemented.
[0046] The clinician computing device 142 may include one or more processors 170 such as one or more microprocessors, controllers, and / or any other suitable type of processor. The clinician computing device 142 may further include a memory 172 (e.g., volatile memory, nonvolatile memory) accessible by the one or more processors 170, (e.g., via a memory controller). The one or more processors 170 may interact with the memory 172 to obtain and execute, forexample, computer-readable instructions stored in the memory 172. Additionally or alternatively, computer-readable instructions may be stored on one or more removable media (e.g., a compact disc, a digital versatile disc, removable flash memory, etc.) that may be coupled to the clinician computing device 142 to provide access to the computer-readable instructions stored thereon. In addition, although the example of Figure 1 illustrates the components of the memory 122, such as chatbot 124, genetic test determiner 125, and / or ML training module 128, as part of the genetic test determining computing device 102, these components may additionally or alternatively be implemented via the memory 172.
[0047] The clinician computing device 142 may further include display device 175.
[0048] The clinician computing device 142 may be operated by a human clinician 179. In some examples, the clinician 179 is a person seeking a recommendation for a genetic test. Examples of the clinician 179 include doctors, nurses, and / or other healthcare workers.
[0049] The patient database 180 may store any data. For example, the patient database 180 may store electronic health records (EHR) 181a (e.g., including an EHR of the patient 199), and / or patient feature data 181b. In some embodiments, the patient feature data 181b is derived from the EHR 181a.
[0050] The genetic test database 182 may store any data. For example, the genetic test database 182 may store a genetic test catalog 183a. The genetic test catalog 183a may include names of and information of any genetic tests. Examples of the genetic tests includes tests for: cystic fibrosis; Huntington’s disease; congenital hypothyroidism; sickle cell disease; phenylketonuria (PKU); epilepsy; low muscle tone; short stature; etc. Additional examples include genetic tests for a risk of a particular type of cancer (e.g., colon cancer, breast cancer, etc.).
[0051] The historical patient phenotype database 184 may store any data. For example, the historical patient phenotype database 184 may store a historical patient phenotypes 185a, and / or historical genetic tests 185b. In some embodiments, the information stored in the historical patient phenotype database 184 is used to train the chatbot 125 and / or the genetic test determiner 125.
[0052] The patient computing device 198 may be any suitable computing device, such as a laptop, a smartphone, a tablet, a phablet, etc. The patient computing device 198 may include one or more processors, one or more display devices, etc.
[0053] In addition, further regarding the example system 100, the illustrated exemplary components may be configured to communicate, e.g., via a network 104 (which may be a wired or wireless network, such as the internet), with any other component. Furthermore, although the example system 100 illustrates only one of each of the components, any number of the example components are contemplated (e.g., any number of genetic test determining computing devices, patient computing devices, clinician computing devices, patient databases, genetic test databases, historical patient phenotype databases, etc.).Example Method
[0054] Figure 2 illustrates an example computer-implemented method 200 for determining a recommendation for a genetic test. However, it should be understood that the example method 200 applies equally to determining a hereditary test. In some embodiments, the example method 200 is implemented by the example computing environment 100. Although the description below refers to many of the blocks as performed by specific components, it should be understood that any of the blocks may be additionally or alternatively performed by any other suitable component, such as any component illustrated in Figure 1.
[0055] The example method may begin at block 205 when clinician computing device 142 prompts the clinician 179 for entry of a patient identifier (ID). For example, in some embodiments, an HTML form (written in Python Dash) is presented to the user asking for a clinic number and test catalog ID (e.g., the patient ID comprises the clinic number and / or catalog ID). In some examples, the prompt is displayed on the display 175. Additionally or alternatively, the clinician computing device 142 may prompt the clinician 179 for a test code identifier of a genetic test ordered for the patient 199 (e.g., a test code identifier of a genetic test to be evaluated).
[0056] At block 210, the clinician 179 enters the patient ID (e.g., the clinic number, the catalog ID, or any other suitable patient ID). At this point, clinician 179 may additionally or alternatively submit a genetic test order for patient 199 (e.g., submit a test code identifier), whichis transmitted through network 104, to optionally amongst other systems, genetic test determining compute device 102. The clinician 179 may also optionally input a date range (e.g., a date range of medical records to be extracted, etc.).
[0057] To this end, genetic laboratory expert 139 interacts with system components (chatbot 124, genetic test determiner 125, phenotype fusion model 126) as illustrated in Figure 4 depicts example screen 400 (e.g., displayed on the display 135) including entry box 410 allowing entry of the patient ID, which, in the illustrated example, is a clinic number.
[0058] Example screen 400 further includes example order information, including: soft ID of an order, softlab order number, medical record number (MRN), patient date of birth, clinician (e.g., in the illustrated example, a physician), client name, and client number. It should be appreciated that the example area 430 allows a user (e.g., the genetic laboratory expert 139) to toggle between order information, clinical summary (e.g., generated phenotypic description), and human phenotypic ontology (HPO) terms.
[0059] At block 215, the genetic test determining computing device 102 receives the patient ID.
[0060] At block 220, the genetic test determining computing device 102 extracts, from the patient database 180, patient features 181b based on the patient ID. In some examples, the patient features 181b are comprised in a Mayo Clinic Cloud (MCC) fast healthcare interoperability resources (FHIR) store. For example, block 220 may include a series of database extractions from identified key features from the MCC FHIR store. For instance, predefined identifiable clinical data resources may be collected for the patient identified by the patient ID. Optionally, the patient features 181b may also be extracted according to a date range input at block 210.
[0061] Figure 5 depicts a diagram of an example extraction process 500. With reference thereto, the MCC FHIR store 510 may send SOFT replication data 520 (e.g., data that provides eventual consistency rather than immediate, strict consistency, etc.) and Advanced Data Lake (ADL) FHIR tables / views 530 to Appropriateness of Testing (AOT) staging 540. AOT staging 540 may then produce the patient features 181b, which may be held at the AOT datastore.Furthermore, the ADL document store 550 may also produce patient features 181b to be held at the staging storage.
[0062] Further at block 220, any of the following may be extracted (e.g., from the patient database 180 or any other database): previous testing histories, relevant concurrent tests, family history, internal notes, reason for testing, as well as other data elements associated with the ordered test to be evaluated. Optionally, any of these may also be extracted according to a date range input at block 210.
[0063] At optional block 223, the genetic test determining computing device 102, via the chatbot 124, converses with the genetic laboratory expert 139. Figure 6 depicts an example screen 600 (e.g., displayed on the display 135, 175, etc.) including an example chat session e.g., occurring in example chat session area 610). Here, in some examples, the chatbot 124 may generate a question asking for additional information of the patient 199. The question may be generated or determined based on the patient features 181b extracted at block 220. For example, the chatbot 124 may determine that it could produce an improved phenotypic description of the patient if it had a particular piece of information. To this end, example, questions asking for additional information include:• “What is the patient’s eye color?”• “Does the patient have a family history of pancreatic cancer?”• “Has the patient ever had a seizure?”
[0064] At block 225, the genetic test determining computing device 102 generates a phenotypic description of the patient 199 using chatbot 124 (e.g., by inputting the extracted features, such as any of those extracted at block 220, into the chatbot 124). In some examples, the chatbot 124 is trained as described with respect to Figure 3. Optionally, the chatbot 124 may also determine the recommendation for the genetic test at block 225 as well.
[0065] In some embodiments, the chatbot 124 comprises an LLM (such as MedPal2, etc.) that creates detailed phenotypic descriptions of the patient and laboratory values with longitudinal context using simple prompt engineering (rather than fine-tuning). The phenotypic descriptions will include both a tabular representation and a clinical narrative. Medical ontologies, such as Systematized Nomenclature of Medicine (SNOMED), human phenotypic ontology (HPO),Logical Observation Identifiers, Names, and Codes (LOINC), may be used to ensure consistent representation in a standardized way (defined by Lab Director) to ensure human-in-the-loop validation remains an option.
[0066] Figure 7 depicts an example diagram 700 of an example implementation of generating a phenotypic description. With reference thereto, processed documents 710 (e.g., the patient features 181b) may be input into the chatbot 124 where they may be received by the retrieval augmentation generation (RAG) Q&A framework 720. The machine learning (ML) platform 730 may process the patient features 181b into the vector database 740. RAG summaries 745 may be generated, and then further processed by the chatbot 124 to generate the phenotypic description 750.
[0067] However, in some embodiments, information received at optional block 223 may be used to update an existing phenotypic description 750 rather than to generate a new one. For example, the genetic test determining computing device 102 may receive information of patient’s 199 eye color, and then use the eye color to update an initial phenotypic description 750 to produce an improved phenotypic description 750. In some such examples, optional block 223 may occur subsequent to any or all of blocks 225, 230, 235, and / or 240, and the update occur once the information is received. The phenotypic description is sometimes referred to a “condensed” phenotypic description because it may condense multiple documents into the phenotypic description.
[0068] At optional block 230, the generated phenotypic description is displayed (e.g., at the display device 135) for genetic laboratory expert 139 approval, rejection, or modification. Figure 8 depicts example screen 800 displaying an example phenotypic description 750. In the illustrated example, buttons 810 allow the genetic laboratory expert 139 (or other user) to approve, reject, or modify the phenotypic description 750.
[0069] At block 235, the phenotypic description 750 (either as generated by the chatbot 124, or as modified by the genetic laboratory expert 139) is input into the genetic test determiner 125 to facilitate the laboratory experts review to determine the recommended genetic test. The information generated by the genetic test determiner 125 may be utilized to rank order the most appropriate test for the patient based on available test lists. In some embodiments, this includes the clinical narrative and tabular- representation being assessed for appropriateness of the specific ordered service. In some embodiments, this is accomplished by a retrieval augmentationgeneration (RAG)-based search (e.g., of the genetic test database 182). Additionally or alternatively, a Vector Search Index of all genetic tests in the test catalog (TEST ID, Specimen, clinical interpretive, performance, fees and codes) may be made. Then, the tabular result and or the clinical narrative will retrieve the most similar catalog items and submit to the LLM to determine if the correct test was selected or why another test might be more suitable (again with prompt engineering).
[0070] In some embodiments, determining the recommended genetic test includes a two step process. First, the phenotype fusion model 126 may extract, from an output of the chatbot 124 (e.g., the phenotypic description 750, etc.), phenotypic terms. Second, the genetic test determiner 125 may use the extracted phenotypic terms to determine the genetic test.
[0071] Figure 9 depicts an example phenotypic terms extraction process 900. In some implementations, the phenotypic terms correspond, wholly or partially, to an ontology (e.g., HPO, SNOMED, LOINC, etc.). At block 910, the phenotype fusion model 126 (phenotagger and custom built LLM ) may pre-process the phenotypic description 750. At block 920, the phenotype fusion model 126 and the LLM+Negation 925 may extract the phenotypic terms (e.g., a list of relevant phenotype ontology terms). One or both of the phenotype fusion model 126 and / or the LLM+Negation 925 may make a confidence calculation that the extracted phenotypic term(s) is correct. At block 930, the results are post-processed (e.g., combined and / or merged, etc.).
[0072] As mentioned above, the phenotypic terms may correspond, wholly or partially, to an ontology (e.g., HPO, SNOMED, LOINC, etc.). As such, the phenotypic terms may correspond to genetic tests. In some embodiments, the genetic test is determined by a mapping included in the ontology mapping phenotypic terms to genetic tests.
[0073] Example screen 1000 (e.g., displayed on the display 135.) of Figure 10 depicts example phenotypic terms 1010 corresponding to HPO IDs. Examples of the phenotypic terms include ataxia, psychosis, suicidal ideation, urinary retention, atypical sorting, cognitive impairment, specific leaning disability, and falls. Examples of the clinical text 1020 from which the HPO terms were extracted from the clinical summary in the illustrated example screen 1000 include text for: cerebellar ataxia, psychosis, suicidal ideation, urinary retention, order, cognitive disorder, learning disorder, and falling.
[0074] The example screen 1000 further illustrates original clinical text 1020 and their corresponding confidences 1030 (e.g., calculated at block 920 of Figure 9, etc.) in the HPO term 1010 extracted by phenotype fusion model 126.
[0075] At block 240, the genetic test is displayed on a display device, such as display 135.
[0076] In some example implementations, at optional blocks 245-255, the genetic test determining computing device 102 further converses, via the chatbot 124, with the genetic laboratory expert 139. NeMo guardrails may be used to direct any follow-up questions to the appropriate system - either the FHIR query results or the test catalog. For example, the guardrails may direct follow-up questions to an FHIR store that houses patient features 181b or the genetic test catalog 183a.
[0077] At optional block 245, the genetic test determining computing device 102 may receive a question (e.g., from the genetic laboratory expert 139). The question may be, for example, about the genetic test. In another example, the question may be about the patient 199, such as question 620 in the example of Figure 6.
[0078] At optional block 250, the genetic test determining computing device 102 generates answer to question e.g., via the chatbot 124). Additionally, the genetic test determining computing device 102 may generate, based on the extracted phenotype ontology terms e.g., by the example process 900 of Figure 9), a contextualized database configured to be used by the chatbot 124 to generate the answer to the question.
[0079] At optional block 255, the answer is displayed on a display device, such as display 135. Such an example is illustrated by example answer 630 of Figure 6. Furthermore, any other relevant patient information may be displayed, such as the: previous testing histories, relevant concurrent tests, family history, internal notes, reason for testing, and / or other data elements associated with the ordered test to be evaluated (e.g., the test associated with the test code identifier).
[0080] At block 260, the determined genetic test is administered to the patient 199.
[0081] At block 265, treatment is administered to the patient based on the results of the administered genetic test. For example, the patient 199 may take particular medications, or eat a particular diet based on the results of the administered genetic test.
[0082] It should be understood that not all blocks and / or events of the exemplary signal diagrams and / or flowcharts are required to be performed. Moreover, the exemplary signal diagrams and / or flowcharts are not mutually exclusive (e.g., block(s) / events from each example signal diagram and / or flowchart may be performed in any other signal diagram and / or flowchart). The exemplary signal diagrams and / or flowcharts may include additional, less, or alternate functionality, including that discussed elsewhere herein.Exemplary Training of the Chatbot
[0083] Techniques described herein may use programmable chatbots, such as the chatbot 124 and / or an ML chatbot (e.g.. ChatGPT), to provide tailored, conversational-like customer service, generate phenotypic descriptions, determine recommendations for genetic tests, etc. The chatbot may be capable of understanding customer requests, providing relevant information, escalating issues, any of which may assist and / or replace the need for customer service assets of an enterprise. Additionally, the chatbot may generate data from customer interactions which the enterprise may use to personalize future support and / or improve the chatbot’s functionality, e.g.. when retraining and / or fine-tuning the chatbot.
[0084] In certain embodiments, the machine learning chatbot may be configured to utilize artificial intelligence and / or machine learning techniques. For instance, the machine learning chatbot or voice bot may be a ChatGPT chatbot. The machine learning chatbot may employ supervised or unsupervised machine learning techniques, which may be followed by, and / or used in conjunction with, reinforced or reinforcement learning techniques. The machine learning chatbot may employ the techniques utilized for ChatGPT. The machine learning chatbot may be configured to generate verbal, audible, visual, graphic, text, or textual output for cither human or other bot / machine consumption or dialogue.
[0085] The ML chatbot may provide advanced features as compared to a non-ML chatbot. For example, the ML chatbot may include and / or derive functionality from a large language model (LLM). The ML chatbot may be trained on a server, such as the genetic test determining computing device 102, using large training datasets of text which may provide sophisticated capability for natural-language tasks, such as answering questions and / or holding conversations. The ML chatbot may include a general-purpose pretrained LLM which, when provided with a starting set of words (prompt) as an input, may attempt to provide an output (response) of themost likely set of words that follow from the input. In one aspect, the prompt may be provided to, and / or the response received from, the ML chatbot and / or any other ML model, via a display (e.g., display 135). This may include a user interface device operably connected to the server via an I / O module, such as an I / O module of the genetic test determining computing device 102. Exemplary user interface devices may include a touchscreen, a keyboard, a mouse, a microphone, a speaker, a display, and / or any other suitable user interface devices.
[0086] Multi-tum (z'.e., back-and-forth) conversations may require LLMs to maintain context and coherence across multiple user prompts and / or utterances, which may require the ML chatbot to keep track of an entire conversation history as well as the current state of the conversation. The ML chatbot may rely on various techniques to engage in conversations with users, which may include the use of short-term and long-term memory. Short-term memory may temporarily store information that may be required for immediate use and may keep track of the current state of the conversation and / or to understand the user’s latest input in order to generate an appropriate response. Long-term memory may include persistent storage of information which may be accessed over an extended period of time. The long-term memory may be used by the ML chatbot to store information about the user (e.g., preferences, chat history, etc.) and may be useful for improving an overall user experience by enabling the ML chatbot to personalize and / or provide more informed responses.
[0087] The system and methods to generate and / or train an ML chatbot model e.g., via the ML training module 128 of the genetic test determining computing device 102) which may be used to train the an ML chatbot, may include three steps: (1) a supervised fine-tuning (SFT) step where a pretrained language model (e.g., an LLM) may be fine-tuned on a relatively small amount of demonstration data curated by human labelers to learn a supervised policy (SFT ML model) which may generate responses / outputs from a selected list of prompts / inputs. The SFT ML model may represent a cursory model for what may be later developed and / or configured as the ML chatbot model; (2) a reward model step where human labelers may rank numerous SFT ML model responses to evaluate the responses which best mimic preferred human responses, thereby generating comparison data. The reward model may be trained on the comparison data; and / or (3) a policy optimization step in which the reward model may further fine-tune and improve the SFT ML model. The outcome of this step may be the ML chatbot model using anoptimized policy. In one aspect, step one may take place only once, while steps two and three may be iterated continuously, e.g., more comparison data is collected on the current ML chatbot model, which may be used to optimize / update the reward model and / or further optimize / update the policy.
[0088] In some embodiments, the language model may be pre-trained by a set of vectors associated with a set of training data. The set of training data may include documents. Creating the set of vectors may include (1) extracting text from documents, (2) splitting the text into semantic clusters, and (3) encoding the semantic clusters as the set of vectors. The semantic clusters may be one or more words, a portion of a word, or a character. A distance between the vectors (e.g., a cosine distance, a Euclidean distance) may depend on a relevance between the semantic clusters corresponding to the vectors.
[0089] In some embodiments, the genetic test determining computing device 102 or an external computing device may encode the vectors using a trained machine learning model (e.g., via the ML training module 128). The trained machine learning model may include a plurality of parameters. When training the machine learning model, the plurality of parameters may be updated iteratively. In other embodiments, the genetic test determining computing device 102 may encode the vectors using existing encoding tables and / or libraries.Supervised Fine-Tuning ML Model
[0090] Figure 3 depicts a combined block and logic diagram 300 for training an ML chatbot model, in which the techniques described herein may be implemented, according to some embodiments. Some of the blocks in Figure 3 may represent hardware and / or software components, other blocks may represent data structures or memory storing these data structures, registers, or state variables, and other blocks may represent output data (e.g., 325). Input and / or output signals may be represented by arrows labeled with corresponding signal names and / or other identifiers. The methods and systems may include one or more servers 302, 304, 306, such as the genetic test determining computing device 102 or an external computing device.
[0091] In one aspect, the server 302 may fine-tune a pretrained language model 310. The pretrained language model 310 may be obtained by the server 302 and be stored in a memory, such as memory 122. The pretrained language model 310 may be loaded into the ML trainingmodule 128 by the server 302 for retraining / fine-tuning. A supervised training dataset 312 may be used to fine-tune the pretrained language model 310 wherein each data input prompt to the pretrained language model 310 may have a known output response for the pretrained language model 310 to learn from. The supervised training dataset 312 may be stored in a memory of the server 302, e.g., the memory 122. In one aspect, the data labelers may create the supervised training dataset 312 prompts and appropriate responses. The pretrained language model 310 may be fine-tuned using the supervised training dataset 312 resulting in the SFT ML model 315 which may provide appropriate responses to user prompts once trained. The trained SFT ML model 315 may be stored in a memory of the genetic test determining computing device 102, e.g., memory 122.
[0092] In some embodiments, the server 302 may fine-tune the pretrained language model 310 using a set of vectors associated with a set of training data. In some instances, the set of training data may include prompts associated with questions and documents, and responses associated with the prompts. Creating the set of vectors may include (1) splitting the text of the prompts, associated questions and / or associated documents into semantic clusters, and (2) encoding the semantic clusters as the set of vectors. The semantic clusters may be one or more words, a portion of a word, or a character. A distance between the vectors (e.g., a cosine distance, a Euclidean distance) may depend on a relevance between the semantic clusters corresponding to the vectors.Training the Reward Model
[0093] In one aspect, training the ML chatbot model 350 may include the server 304 training a reward model 320 to provide as an output a scaler value / reward 325. The reward model 320 may be required to leverage Reinforcement Learning with Human Feedback (RLHF) in which a model (e.g., ML chatbot model 350) learns to produce outputs which maximize its reward 325, and in doing so may provide responses which are better aligned to user prompts.
[0094] Training the reward model 320 may include the server 304 providing a single prompt 322 to the SFT ML model 315 as an input. The input prompt 322 may be provided via an input device (e.g., a keyboard) via the I / O module. The prompt 322 may be previously unknown to the SFT ML model 315, e.g., the labelers may generate new prompt data, the prompt 322 may include testing data stored on any of the databases 180, 182, 184, and / or any other suitable prompt data. The SFT ML model 315 may generate multiple, different output responses 324A,324B, 324C, 324D to the single prompt 322. The server 304 may output the responses 324A, 324B, 324C, 324D via an I / O module to a user interface device, such as a display (e.g., as text responses), a speaker (e.g., as audio / voice responses), and / or any other suitable manner of output of the responses 324A, 324B, 324C, 324D for review by the data labelers.
[0095] The data labelers may provide feedback via the server 304 on the responses 324A, 324B, 324C, 324D when ranking 326 them from best to worst based upon the prompt-response pairs. The data labelers may rank 326 the responses 324A, 324B, 324C, 324D by labeling the associated data. The ranked prompt-response pairs 328 may be used to train the reward model 320. In one aspect, the server 304 may load the reward model 320 via the ML module (e.g., the ML training module 128) and train the reward model 320 using the ranked response pairs 328 as input. The reward model 320 may provide as an output the scalar reward 325.
[0096] In one aspect, the scalar reward 325 may include a value numerically representing a human preference for the best and / or most expected response to a prompt, i. e. , a higher scaler reward value may indicate the user is more likely to prefer that response, and a lower scalar reward may indicate that the user is less likely to prefer that response. For example, inputting the “winning” prompt-response (i.e., input-output) pair data to the reward model 320 may generate a winning reward. Inputting a “losing” prompt-response pair data to the same reward model 320 may generate a losing reward. The reward model 320 and / or scalar reward 325 may be updated based upon labelers ranking 326 additional prompt-response pairs generated in response to additional prompts 322.
[0097] In one example, a data labeler may provide to the SFT ML model 315 as an input prompt 322, “Describe the sky.” The input may be provided by the labeler via the server 304 running a chatbot application utilizing the SFT ML model 315. The SFT ML model 315 may provide as output responses to the labeler via the genetic test determining computing device 102: (i) “the sky is above” 324A; (ii) “the sky includes the atmosphere and may be considered a place between the ground and outer space” 324B; and (iii) “the sky is heavenly” 324C. The data labeler may rank 326, via labeling the prompt-response pairs, prompt-response pair 322 / 324B as the most preferred answer; prompt-response pair 322 / 324A as a less preferred answer; and prompt-response 322 / 324C as the least preferred answer. The labeler may rank 326 the prompt-response pair data in any suitable manner. The ranked prompt-response pairs 328 may be provided to the reward model 320 to generate the scalar reward 325.
[0098] While the reward model 320 may provide the scalar reward 325 as an output, the reward model 320 may not generate a response (e.g., text). Rather, the scalar reward 325 may be used by a version of the SFT ML model 315 to generate more accurate responses to prompts, i.e., the SFT model 315 may generate the response such as text to the prompt, and the reward model 320 may receive the response to generate a scalar reward 325 of how well humans perceive it. Reinforcement learning may optimize the SFT model 315 with respect to the reward model 320 which may realize the configured ML chatbot model 350.RLHF to Train the ML Chatbot Model
[0099] In one aspect, the server 306 may train the ML chatbot model 350 (e.g., via the ML training module 128) to generate a response 334 to a random, new and / or previously unknown user prompt 332. To generate the response 334, the ML chatbot model 350 may use a policy 335 (e.g., algorithm) which it learns during training of the reward model 320, and in doing so may advance from the SFT model 315 to the ML chatbot model 350. The policy 335 may represent a strategy that the ML chatbot model 350 learns to maximize the reward 325. As discussed herein, based upon prompt-response pairs, a human labeler may continuously provide feedback to assist in determining how well the ML chatbot’s 350 responses match expected responses to determine the rewards 325. The rewards 325 may feed back into the ML chatbot model 350 to evolve the policy 335. Thus, the policy 335 may adjust the parameters of the ML chatbot model 350 based upon the rewards 325 it receives for generating good responses. The policy 335 may update as the ML chatbot model 350 provides responses 334 to additional prompts 332.
[0100] In one aspect, the response 334 of the ML chatbot model 350 using the policy 335 based upon the reward 325 may be compared using a cost function 338 to the SFT ML model 315 (which may not use a policy) response 336 of the same prompt 332. The cost function 338 may be trained in a similar manner and / or contemporaneous with the reward model 320. The server 306 may compute a cost 340 based upon the cost function 338 of the responses 334, 336. The cost 340 may reduce the distance between the responses 334, 336, i.e., a statistical distance measuring how one probability distribution is different from a second, in one aspect the response 334 of the ML chatbot model 350 versus the response 336 of the SFT model 315. Using the cost340 to reduce the distance between the responses 334, 336 may avoid a server over-optimizing the reward model 320 and deviating too drastically from the human-intended / preferred response. Without the cost 340, the ML chatbot model 350 optimizations may result in generating responses 334 which are unreasonable but may still result in the reward model 320 outputting a high reward 325.
[0101] In one aspect, the responses 334 of the ML chatbot model 350 using the current policy 335 may be passed by the server 306 to the rewards model 320, which may return the scalar reward 325. The ML chatbot model 350 response 334 may be compared via the cost function 338 to the SFT ML model 315 response 336 by the server 306 to compute the cost 340. The server 306 may generate a final reward 342 which may include the scalar reward 325 offset and / or restricted by the cost 340. The final reward 342 may be provided by the server 306 to the ML chatbot model 350 and may update the policy 335, which in turn may improve the functionality of the ML chatbot model 350.
[0102] To optimize the ML chatbot 350 over time, RLHF via the human labeler feedback may continue ranking 326 responses of the ML chatbot model 350 versus outputs of earlier / other versions of the SFT ML model 315, i.e., providing positive or negative rewards 325. The RLHF may allow the servers (e.g., servers 304, 306) to continue iteratively update the reward model 320 and / or the policy 335. As a result, the ML chatbot model 350 may be retrained and / or finetuned based upon the human feedback via the RLHF process, and throughout continuing conversations may become increasingly efficient.
[0103] Although multiple servers 302, 304, 306 are depicted in the exemplary block and logic diagram 300, each providing one of the three steps of the overall ML chatbot model 350 training, fewer and / or additional servers may be utilized and / or may provide the one or more steps of the ML chatbot model 350 training. In one aspect, one server may provide the entire ML chatbot model 350 training.Additional Exemplary Aspects - 1
[0104] Aspect 1. A computing system for determining a recommendation for a genetic test for a patient, comprising: one or more processors, andone or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: receive, via the one or more processors, a patient identifier (ID) of the patient; extract, via the one or more processors, from a patient database, patient features based on the patient ID; generate, via the one or more processors, a phenotypic description of the patient using a chatbot; determine, via the one or more processors, the genetic test based on the generated phenotypic description of the patient; and cause, via the one or more processors, the determined genetic test to be displayed on a display device.
[0105] Aspect 2. The computing system of aspect 1, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to determine, via the one or more processors, the genetic test by: using a phenotype fusion model to extract one or more phenotypic terms from the phenotypic description; and determining the genetic test based on the extracted one or more phenotypic terms.
[0106] Aspect 3. The computing system of any one of aspects 1-2, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to, subsequent to the extraction of the patient features: generate, via the one or more processors, using the chatbot, a question asking for additional information of the patient; present, via the one or more processors, the question asking for the additional information of the patient; receive, via the one or more processors, an answer to the question asking for the additional information of the patient; and generate, via the one or more processors, the phenotypic description of the patient further based on the answer to the question asking for the additional information of the patient.
[0107] Aspect 3. The computing system of any one of aspects 1-2, wherein the genetic test tests for at least one of: cystic fibrosis;Huntington’s disease; congenital hypothyroidism; sickle cell disease; phenylketonuria (PKU); epilepsy; low muscle tone; short stature; risk of colon cancer; or risk of breast cancer.
[0108] Aspect 5. The computing system of any one of aspects 1-4, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to, prior to the determination of the genetic test: cause, via the one or more processors, the generated phenotypic description of the patient to be displayed on the display device; and receive, via the one or more processors, approval, rejection, or modification of the generated phenotypic description.
[0109] Aspect 6. The computing system of one of aspects 1-5, wherein the chatbot comprises a large language model (LLM) pre-trained, fine-tuned and / or trained using information from a historical patient phenotype database.
[0110] Aspect 7. The computing system of any one of aspects 1-6, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: identify, in the information from the historical patient phenotype database, historical patient phenotypes, and historical genetic tests; and train, via the one or more processors, the chatbot by fine-tuning the chatbot using the historical patient phenotypes, and the historical genetic tests.
[0111] Aspect 8. The computing system of any one of aspects 1-7, wherein the patient ID comprises a clinic number and a test catalog ID number, and wherein the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to display, on the display, in hypertext markup language form, a prompt asking a clinician to input the clinic number and test catalog ID number.
[0112] Aspect 8a. The computing system any of one of aspects 1-8, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to, subsequent to the extraction of the patient features: generate, via the one or more processors, using the chatbot, a question asking for additional information of the patient; present, via the one or more processors, the question asking for the additional information of the patient; receive, via the one or more processors, an answer to the question asking for the additional information of the patient; and update, via the one or more processors, the generated phenotypic description of the patient based on the answer to the question asking for the additional information of the patient.
[0113] Aspect 8b. The computing system of any one of aspects l-8a, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to determine the genetic test via a retrieval augmentation generation (RAG)-based search of a genetic test database.
[0114] Aspect 8c. The computing system of any one of aspects l-8b, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: receive, via the one or more processors, a follow up question; and direct, via the one or more processors, using a guardrail, the follow up question to a fast healthcare interoperability resources (FHIR) module or a genetic test catalog.
[0115] Aspect 8d. The computing system of any one of aspects l-8c, wherein the genetic test is a hereditary test.
[0116] Aspect 8e. The computing system of any one of aspects l-8d, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: extract, via the one or more processors, from a patient database, based on the patient ID: a previous testing history, a relevant concurrent test, family history information, internal notes, a reason for testing, and / or another data element associated with an ordered genetic test to be evaluated; and generate, via the one or more processors, the phenotypic description further based on the: previous testing history, relevant concurrent test, family history information, internal notes, reason for testing, and / or another data element associated with the ordered genetic test to be evaluated.
[0117] Aspect 8f. The computing system of any one of aspects l-8e, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: receive, via the one or more processors, a test code identifier; and determine, via the one or more processors, the genetic test further by evaluating a genetic test identified by the test code identifier.
[0118] Aspect 8g. The computing system of any one of aspects l-8f, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: generate, via the one or more processors, a phenotype ontology terms based on the generated phenotypic description of the patient; and generate, via the one or more processors, based on the phenotype ontology terms, a contextualized database configured to be used by the chatbot to generate the answer to the question asking for the additional information of the patient.
[0119] Aspect 9. A non-transitory computer-readable medium, having stored thereon instructions that when executed, cause a computer to: receive, via one or more processors, a patient identifier (ID) of the patient; extract, via the one or more processors, from a patient database, patient features based on the patient ID;generate, via the one or more processors, a phenotypic description of the patient using a chatbot; determine, via the one or more processors, the genetic test based on the generated phenotypic description of the patient; and cause, via the one or more processors, the information to determine appropriateness of the genetic test ordered to be displayed on a display device.
[0120] Aspect 10. The non-transitory computer-readable medium of aspect 9, having stored thereon instructions that when executed, cause a computer to determine the genetic test via a retrieval augmentation generation (RAG)-based search of a genetic test database.
[0121] Aspect 11. The non-transitory computer-readable medium of any one of aspects 9-10, wherein the genetic test tests for at least one of: cystic fibrosis;Huntington’s disease; congenital hypothyroidism; sickle cell disease; phenylketonuria (PKU); epilepsy; low muscle tone; short stature; risk of colon cancer; or risk of breast cancer.
[0122] Aspect 12. The non-transitory computer-readable medium of any one of aspects 9-11, having stored thereon instructions that when executed, cause a computer to: receive, via the one or more processors, a follow up question; and direct, via the one or more processors, using a guardrail, the follow up question to a fast healthcare interoperability resources (FHIR) module or a genetic test catalog.
[0123] Aspect 13. The non-transitory computer- readable medium of any one of aspects 9-12, having stored thereon instructions that when executed, cause a computer to: cause, via the one or more processors, the generated phenotypic description of the patient to be displayed on the display device; andreceive, via the one or more processors, approval, rejection, or modification of the generated phenotypic description.
[0124] Aspect 14. The non-transitory computer-readable medium of any one of aspects 9-13, wherein the chatbot comprises a large language model (LLM) pre-trained, fine-tuned and / or trained using information from a historical patient phenotype database.
[0125] Aspect 15. A computer-implemented method for determining a recommendation for a genetic test for a patient, the method comprising: receiving, via one or more processors, a patient identifier (ID) of the patient; extracting, via the one or more processors, from a patient database, patient features based on the patient ID; generating, via the one or more processors, a phenotypic description of the patient using a chatbot; determining, via the one or more processors, the genetic test based on the generated phenotypic description of the patient; and causing, via the one or more processors, the determined genetic test to be displayed on a display device.
[0126] Aspect 16. The computer- implemented method of aspect 15, wherein the determining the genetic test comprises determining the genetic test via a retrieval augmentation generation (RAG)-based search of a genetic test database.
[0127] Aspect 17. The computer-implemented method of any one of aspects 15-16, wherein the genetic test tests for at least one of: cystic fibrosis;Huntington’s disease; congenital hypothyroidism; sickle cell disease; phenylketonuria (PKU); epilepsy; low muscle tone; short stature;risk of colon cancer; or risk of breast cancer.
[0128] Aspect 18. The computer- implemented method of any one of aspects 15-17, further comprising: receiving, via the one or more processors, a follow up question; and directing, via the one or more processors, using a guardrail, the follow up question to a fast healthcare interoperability resources (FHIR) module or a genetic test catalog.
[0129] Aspect 19. The computer- implemented method of any one of aspects 15-18, further comprising: causing, via the one or more processors, the generated phenotypic description of the patient to be displayed on the display device; and receiving, via the one or more processors, approval, rejection, or modification of the generated phenotypic description.
[0130] Aspect 20. The computer-implemented method of any one of aspects 15-19, wherein the chatbot comprises a large language model (LLM) pre-trained, fine-tuned and / or trained using information from a historical patient phenotype database.Additional Exemplary Aspects - II
[0131] Aspect 1. A computing system for determining a recommendation for a genetic test for a patient, comprising: one or more processors, and one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: retrieve via the one or more processors, from a patient database, a collection of records identified by a specific set of test codes and date ranges, the retrieved data includes patient features associated with the corresponding genetic tests and the corresponding clinical documents.
[0132] Aspect 2. The computing system of aspect 1, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to utilize the relevant clinical documents to generate phenotypic description of the patient using a Retrieval Augmented Generation; and cause, via the one ormore processors, to store the generated phenotypic description of the patient in a permanent data store for further display.
[0133] Aspect 3. The computing system of any one of the preceding aspects, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to extract from the generated phenotypic description, via the one or more processors, a list of phenotypic terms by: Using a custom-built fused model that combines Natural Language Processing, Large Language Models and utilizes the Human Phenotypic Ontology to provide medical context helping determine the correct phenotypic terms to be extracted.
[0134] Aspect 4. The computing system of any one of the preceding aspects, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to, prior to determining the appropriateness of the ordered genetic test: cause, via the one or more processors, the generated phenotypic description of the patient to be displayed on the display device; and allow users to, via the one or more processors, review and modify the generated phenotypic description.
[0135] Aspect 5. The computing system of any one of the preceding aspects, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to, after the extraction of the patient features and clinical documents: access, via the one or more processors, a comprehensive vector database that can be used by a chatbot, allowing users to ask for additional information of the patient; present, via the one or more processors, a mechanism for asking questions for the additional information of the patient; receive, via the one or more processors, an answer to the question asking for the additional information of the patient; and display, via the one or more processors, the answer to the question asking for the additional information of the patient.
[0136] Aspect 6. The computing system of any one of the preceding aspects, wherein the chatbot comprises a vector database and a Retrieval Augmented Generation (RAG) based search to provide patient document contextualized answers using information from predefined set of patient specific clinical documents.Other Maters
[0137] Although the text herein sets forth a detailed description of numerous different embodiments, it should be understood that the legal scope of the invention is defined by the words of the claims set forth at the end of this patent. The detailed description is to be construed as exemplary only and does not describe every possible embodiment, as describing every possible embodiment would be impractical, if not impossible. One could implement numerous alternate embodiments, using either current technology or technology developed after the filing date of this patent, which would still fall within the scope of the claims.
[0138] It should also be understood that, unless a term is expressly defined in this patent using the sentence “As used herein, the term ‘ ’ is hereby defined to mean...” or a similar sentence, there is no intent to limit the meaning of that term, either expressly or by implication, beyond its plain or ordinary meaning, and such term should not be interpreted to be limited in scope based upon any statement made in any section of this patent (other than the language of the claims). To the extent that any term recited in the claims at the end of this disclosure is referred to in this disclosure in a manner consistent with a single meaning, that is done for sake of clarity only so as to not confuse the reader, and it is not intended that such claim term be limited, by implication or otherwise, to that single meaning.
[0139] Throughout this specification, plural instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed concurrently, and nothing requires that the operations be performed in the order illustrated. Structures and functionality presented as separate components in example configurations may be implemented as a combined structure or component.Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.
[0140] Additionally, certain embodiments are described herein as including logic or a number of routines, subroutines, applications, or instructions. These may constitute either software (code embodied on a non-transitory, tangible machine-readable medium) or hardware. In hardware, the routines, etc., are tangible units capable of performing certain operations and may beconfigured or arranged in a certain manner. In example embodiments, one or more computer systems (e.g., a standalone, client or server computer system) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or application portion) as a hardware module that operates to perform certain operations as described herein.
[0141] In various embodiments, a hardware module may be implemented mechanically or electronically. For example, a hardware module may comprise dedicated circuitry or logic that is permanently configured (e.g., as a special-purpose processor, such as a field programmable gate array (FPGA) or an application- specific integrated circuit (ASIC) to perform certain operations). A hardware module may also comprise programmable logic or circuitry e.g., as encompassed within a general -purpose processor or other programmable processor) that is temporarily configured by software to perform certain operations. It will be appreciated that the decision to implement a hardware module mechanically, in dedicated and permanently configured circuitry, or in temporarily configured circuitry (e.g., configured by software) may be driven by cost and time considerations.
[0142] Accordingly, the term “hardware module” should be understood to encompass a tangible entity, be that an entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain manner or to perform certain operations described herein. Considering embodiments in which hardware modules are temporarily configured (e.g., programmed), each of the hardware modules need not be configured or instantiated at any one instance in time. For example, where the hardware modules comprise a general-purpose processor configured using software, the general-purpose processor may be configured as respective different hardware modules at different times. Software may accordingly configure a processor, for example, to constitute a particular hardware module at one instance of time and to constitute a different hardware module at a different instance of time.
[0143] Hardware modules can provide information to, and receive information from, other hardware modules. Accordingly, the described hardware modules may be regarded as being communicatively coupled. Where multiple of such hardware modules exist contemporaneously, communications may be achieved through signal transmission (e.g., over appropriate circuits andbuses) that connect the hardware modules. In embodiments in which multiple hardware modules are configured or instantiated at different times, communications between such hardware modules may be achieved, for example, through the storage and retrieval of information in memory structures to which the multiple hardware modules have access. For example, one hardware module may perform an operation and store the output of that operation in a memory device to which it is communicatively coupled. A further hardware module may then, at a later time, access the memory device to retrieve and process the stored output. Hardware modules may also initiate communications with input or output devices, and can operate on a resource (e.g., a collection of information).
[0144] The various operations of example methods described herein may be performed, at least partially, by one or more processors that are temporarily configured e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented modules that operate to perform one or more operations or functions. The modules referred to herein may, in some example embodiments, comprise processor- implemented modules.
[0145] Similarly, the methods or routines described herein may be at least partially processor- implemented. For example, at least some of the operations of a method may be performed by one or more processors or processor-implemented hardware modules. The performance of certain of the operations may be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In some example embodiments, the processor or processors may be located in a single location e.g., within a home environment, an office environment or as a server farm), while in other embodiments the processors may be distributed across a number of geographic locations.
[0146] Unless specifically stated otherwise, discussions herein using words such as “processing,” “computing,” “calculating,” “determining,” “presenting,” “displaying,” or the like may refer to actions or processes of a machine (e.g., a computer) that manipulates or transforms data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.
[0147] As used herein any reference to “one embodiment” or “an embodiment” means that a particular element, feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment.
[0148] Some embodiments may be described using the expression “coupled” and “connected” along with their derivatives. For example, some embodiments may be described using the term “coupled” to indicate that two or more elements are in direct physical or electrical contact. The term “coupled,” however, may also mean that two or more elements are not in direct contact with each other, but yet still co-operate or interact with each other. The embodiments are not limited in this context.
[0149] As used herein, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having” or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, “or” refers to an inclusive or and not to an exclusive or. For example, a condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or not present), A is false (or not present) and B is true (or present), and both A and B are true (or present).
[0150] In addition, use of the “a” or “an” are employed to describe elements and components of the embodiments herein. This is done merely for convenience and to give a general sense of the description. This description, and the claims that follow, should be read to include one or at least one and the singular also includes the plural unless it is obvious that it is meant otherwise.
[0151] Upon reading this disclosure, those of skill in the art will appreciate still additional alternative structural and functional designs for the approaches described herein. Thus, while particular embodiments and applications have been illustrated and described, it is to be understood that the disclosed embodiments are not limited to the precise construction and components disclosed herein. Various modifications, changes and variations, which will be apparent to those skilled in the art, may be made in the arrangement, operation and details of themethod and apparatus disclosed herein without departing from the spirit and scope defined in the appended claims.
[0152] The particular features, structures, or characteristics of any specific embodiment may be combined in any suitable manner and in any suitable combination with one or more other embodiments, including the use of selected features without corresponding use of other features. In addition, many modifications may be made to adapt a particular application, situation or material to the essential scope and spirit of the present invention. It is to be understood that other variations and modifications of the embodiments of the present invention described and illustrated herein are possible in light of the teachings herein and are to be considered part of the spirit and scope of the present invention.
[0153] While the preferred embodiments of the invention have been described, it should be understood that the invention is not so limited and modifications may be made without departing from the invention. The scope of the invention is defined by the appended claims, and all devices that come within the meaning of the claims, either literally or by equivalence, are intended to be embraced therein.
[0154] It is therefore intended that the foregoing detailed description be regarded as illustrative rather than limiting, and that it be understood that it is the following claims, including all equivalents, that are intended to define the spirit and scope of this invention.
[0155] Furthermore, the patent claims at the end of this patent application are not intended to be construed under 35 U.S.C. § 112(f) unless traditional means-plus-function language is expressly recited, such as “means for” or “step for” language being explicitly recited in the claim(s). The systems and methods described herein are directed to an improvement to computer functionality, and improve the functioning of conventional computers.
Claims
WHAT IS CLAIMED:
1. A computing system for determining a recommendation for a genetic test for a patient, comprising: one or more processors, and one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: receive, via the one or more processors, a patient identifier (ID) of the patient; extract, via the one or more processors, from a patient database, patient features based on the patient ID; generate, via the one or more processors, a phenotypic description of the patient using a chatbot; determine, via the one or more processors, the genetic test based on the generated phenotypic description of the patient; and cause, via the one or more processors, the determined genetic test to be displayed on a display device.
2. The computing system of claim 1, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to determine, via the one or more processors, the genetic test by: using a phenotype fusion model to extract one or more phenotypic terms from the phenotypic description; and determining the genetic test based on the extracted one or more phenotypic terms.
3. The computing system of claim 1, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to, subsequent to the extraction of the patient features: generate, via the one or more processors, using the chatbot, a question asking for additional information of the patient; present, via the one or more processors, the question asking for the additional information of the patient;receive, via the one or more processors, an answer to the question asking for the additional information of the patient; and generate, via the one or more processors, the phenotypic description of the patient further based on the answer to the question asking for the additional information of the patient.
4. The computing system of claim 1, wherein the genetic test is a hereditary test.
5. The computing system of claim 1, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to, prior to the determination of the genetic test: cause, via the one or more processors, the generated phenotypic description of the patient to be displayed on the display device; and receive, via the one or more processors, approval, rejection, or modification of the generated phenotypic description.
6. The computing system of claim 1, wherein the chatbot comprises a large language model (LLM) pre-trained, fine-tuned and / or trained using information from a historical patient phenotype database.
7. The computing system of claim 6, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: identify, in the information from the historical patient phenotype database, historical patient phenotypes, and historical genetic tests; and train, via the one or more processors, the chatbot by fine-tuning the chatbot using the historical patient phenotypes, and the historical genetic tests.
8. The computing system of claim 1, wherein the patient ID comprises a clinic number and a test catalog ID number, and wherein the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors,cause the computing system to display, on the display, in hypertext markup language form, a prompt asking a clinician to input the clinic number and test catalog ID number.
9. A non-transitory computer-readable medium for determining a recommendation for a genetic test for a patient, the non-transitory computer-readable medium having stored thereon instructions that when executed, cause a computer to: receive, via one or more processors, a patient identifier (ID) of the patient; extract, via the one or more processors, from a patient database, patient features based on the patient ID; generate, via the one or more processors, a phenotypic description of the patient using a chatbot; determine, via the one or more processors, the genetic test based on the generated phenotypic description of the patient; and cause, via the one or more processors, the determined genetic test to be displayed on a display device.
10. The non-transitory computer-readable medium of claim 9, having stored thereon instructions that when executed, cause a computer to determine the genetic test via a retrieval augmentation generation (RAG)-based search of a genetic test database.
11. The non-transitory computer-readable medium of claim 9, wherein the genetic test tests for at least one of: cystic fibrosis;Huntington’s disease; congenital hypothyroidism; sickle cell disease; phenylketonuria (PKU); epilepsy; low muscle tone; short stature; risk of colon cancer; orrisk of breast cancer.
12. The non-transitory computer-readable medium of claim 9, having stored thereon instructions that when executed, cause a computer to: receive, via the one or more processors, a follow up question; and direct, via the one or more processors, using a guardrail, the follow up question to a fast healthcare interoperability resources (FHIR) module or a genetic test catalog.
13. The non-transitory computer-readable medium of claim 9, having stored thereon instructions that when executed, cause a computer to: cause, via the one or more processors, the generated phenotypic description of the patient to be displayed on the display device; and receive, via the one or more processors, approval, rejection, or modification of the generated phenotypic description.
14. The non-transitory computer-readable medium of claim 9, wherein the chatbot comprises a large language model (LLM) pre-trained, fine-tuned and / or trained using information from a historical patient phenotype database.
15. A computer-implemented method for determining a recommendation for a genetic test for a patient, the method comprising: receiving, via one or more processors, a patient identifier (ID) of the patient; extracting, via the one or more processors, from a patient database, patient features based on the patient ID; generating, via the one or more processors, a phenotypic description of the patient using a chatbot; determining, via the one or more processors, the genetic test based on the generated phenotypic description of the patient; and causing, via the one or more processors, the determined genetic test to be displayed on a display device.
16. The computer-implemented method of claim 15, wherein the determining the genetic test comprises determining the genetic test via a retrieval augmentation generation (RAG)-based search of a genetic test database.
17. The computer-implemented method of claim 15, wherein the genetic test tests for at least one of: cystic fibrosis;Huntington’s disease; congenital hypothyroidism; sickle cell disease; phenylketonuria (PKU); epilepsy; low muscle tone; short stature; risk of colon cancer; or risk of breast cancer.
18. The computer-implemented method of claim 15, further comprising: receiving, via the one or more processors, a follow up question; and directing, via the one or more processors, using a guardrail, the follow up question to a fast healthcare interoperability resources (FHIR) module or a genetic test catalog.
19. The computer-implemented method of claim 15, further comprising: causing, via the one or more processors, the generated phenotypic description of the patient to be displayed on the display device; and receiving, via the one or more processors, approval, rejection, or modification of the generated phenotypic description.
20. The computer-implemented method of claim 15, wherein the chatbot comprises a large language model (LLM) pre-trained, fine-tuned and / or trained using information from a historical patient phenotype database.