Predicting a therapeutic response to a disease using a large language model

By using a genomic transformer to extract exomic data and a multimodal disease prediction model to integrate various patient data types, the system addresses the limitations of conventional models in predicting medication response for rheumatoid arthritis, achieving more accurate and comprehensive predictions.

WO2025111531A1PCT designated stage expired Publication Date: 2025-05-30MAYO FOUNDATION FOR MEDICAL EDUCATION & RESEARCH

Patent Information

Application Number
PCT/US2024/057034
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-24
Filing Date
2024-11-22
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Conventional pharmacogenomic models struggle to accurately predict a patient's response to medication for rheumatoid arthritis due to reliance on single nucleotide polymorphisms and inability to process entire exomes, while also neglecting clinical notes and medical images.

Method used

A computer-implemented method and system that utilize a genomic transformer to extract exomic data from genomic data and a multimodal disease prediction model to generate predictions of a patient's response to medication by integrating exomic data with medical image data and clinical data.

Benefits of technology

This approach enables more accurate and comprehensive predictions of a patient's response to medication by analyzing full exomes and multiple types of patient data, improving the efficiency and effectiveness of treatment decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024057034_30052025_PF_FP_ABST
    Figure US2024057034_30052025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for generating a prediction of a response of a patient to a medication to treat a disease are provided. The system may obtain genomic data, medical image data, and clinical data of the patient. The system may provide the genomic data to a genomic transformer trained to extract exomic data of the patient associated with the disease. The system may provide the exomic data, the medical image data, and the clinical data to a disease prediction model trained to generate the prediction of the response of the patient to the medication. The system may provide the prediction to a user device.
Need to check novelty before this filing date? Find Prior Art

Description

PREDICTING A THERAPEUTIC RESPONSE TO A DISEASE USING A EARGELANGUAGE MODELCROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims priority to and the benefit of the filing date of provisional U.S. Patent Application No. 63 / 602,539, entitled “THE USE OF LLM IN GENOMICS TO PREDICT THERAPEUTIC RESPONSE” and filed on November 24, 2023, the entire contents of which is hereby expressly incorporated herein by reference.FIELD OF THE INVENTION

[0002] The present disclosure generally relates to disease prediction, and more particularly, to predicting a therapeutic response to a disease using a language model.BACKGROUND

[0003] Methotrexate is a first-line treatment of rheumatoid arthritis (RA), however, it can take months to determine its efficacy, if the patient responds at all to the treatment, and causes significant side effects. The patient’s genetics, such as exome variations, can be a significant predictor of RA development and responsiveness to different treatment medications.Conventional pharmacogenomic models that rely on specific single nucleotide polymorphisms of the patient’s genetics to predict their medication response can provide poor replicability in independent patient populations, and are also generally unable to digest entire exomes of the patient’s genetic information. Moreover, other types of patient information such as clinical notes and medical images that may be relevant to predicting the disease progression, treatment response, or other disease predictions, is generally not considered. Conventional disease and treatment response prediction techniques may include additional ineffectiveness, inefficiencies, encumbrances, and / or other drawbacks.SUMMARY

[0004] In one aspect, a computer-implemented method for generating a prediction of a response of a patient to a medication to treat a disease is provided. The computer-implemented method may include obtaining, by one or more processors, genomic data, medical image data, and clinical data of the patient; providing, by the one or more processors, the genomic data to agenomic transformer trained to extract exomic data of the patient associated with the disease; providing, by the one or more processors, the exomic data, the medical image data, and the clinical data to a disease prediction model trained to generate the prediction of the response of the patient to the medication; and providing, by the one or more processors, the prediction to a user device. The computer-implemented method may include additional, less, or alternate functionality or actions, including those discussed elsewhere herein.

[0005] In another aspect, a system for generating a prediction of a response of a patient to a medication to treat a disease is provided. The system may include one or more processors; and one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, may cause the system to: obtain genomic data, medical image data, and clinical data of the patient, provide the genomic data to a genomic transformer trained to extract exomic data of the patient associated with the disease, provide the exomic data, the medical image data, and the clinical data to a disease prediction model trained to generate the prediction of the response of the patient to the medication, and provide the prediction to a user device. The system may include additional, less, or alternate functionality, including that discussed elsewhere herein.

[0006] In another aspect, a non-transitory computer-readable medium is disclosed. The non- transitory computer-readable medium may have stored thereon instructions that, when executed by one or more processors, may cause the one or more processors to at least: obtain genomic data, medical image data, and clinical data of a patient; provide the genomic data to a genomic transformer trained to extract exomic data of the patient associated with a disease; provide the exomic data, the medical image data, and the clinical data to a disease prediction model trained to generate a prediction of a response of the patient to a medication; and provide the prediction to a user device. The instructions may direct additional, less, or alternate functionality, including that discussed elsewhere herein.

[0007] Additional, alternate and / or fewer actions, steps, features and / or functionality may be included in an aspect and / or embodiments, including those described elsewhere herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The figures described below depict various aspects of the applications, methods, and systems disclosed herein. It should be understood that each figure depicts one embodiment of a particular aspect of the disclosed applications, systems and methods, and that each of the figures is intended to accord with a possible embodiment thereof. Furthermore, wherever possible, the following description refers to the reference numerals included in the following figures, in which features depicted in multiple figures are designated with consistent reference numerals.

[0009] FIG. 1 depicts a block diagram of an exemplary computing environment 100 for generating a prediction of a response of a patient to a medication to treat a disease, according to embodiments.

[0010] FIG. 2A depicts a block diagram of an example training data preprocessing of genomic transformer training data, according to embodiments.

[0011] FIG. 2B depicts a block diagram of the file conversion process for converting the BAM data to BED data, FASTQ data, and log files, according to embodiments.

[0012] FIG. 2C depicts a block diagram of a training data filtering and compression process, according to embodiments.

[0013] FIG. 2D depicts a block diagram of an example tokenization process, according to embodiments.

[0014] FIG. 2E depicts a block diagram of an example soft clipping process, according to embodiments.

[0015] FIG. 2F depicts a block diagram of an example NPZ to H5 statistical analysis and data preparation process, according to embodiments.

[0016] FIG. 2G depicts a block diagram of different packing techniques, according to embodiments.

[0017] FIG. 3 depicts a block diagram of an example training process of a model, according to embodiments.

[0018] FIG. 4 depicts a combined block and logic diagram for training an LLM, according to embodiments.

[0019] FIG. 5 depicts a How diagram of an exemplary computer-implemented method for generating a prediction of a response of a patient to a medication to treat a disease, according to embodiments.

[0020] Advantages will become more apparent to those skilled in the ail from the following description of the preferred embodiments which have been shown and described by way of illustration. As will be realized, the present embodiments may be capable of other and different embodiments, and their details are capable of modification in various respects. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive.DETAILED DESCRIPTIONOVERVIEW

[0021] The systems and methods disclosed herein provide a support tool for rheumatoid arthritis (RA) or other chronic diseases, such as autoimmune diseases. An ensemble of models including a genomic transformer and a multimodal model may predict a patient’ s response to a treatment of the disease, such as predicting the patient’ s response to Methotrexate for treating RA. The genomic transformer (e.g., a fine-tuned GPT-based LLM) may receive the patient’s genomic data and extract exomic information associated with the disease, such as exomic variations of the patient respective to a human reference genome. A multimodal model may receive the exomic data of the patient, as well as the patient’s medical image data (e.g., x-rays of one or more hands and / or feet) and clinical data (e.g., clinical notes associated with the disease) to generate the prediction of the response of the patient to the medication. A computing device may receive the prediction, for example to allow a medical professional to determine whether to treat the patient using Methotrexate or an alternate medication or treatment modality.

[0022] The disclosed systems and methods provide an innovative approach to using genomic data for predicting a patient’s response to treatment of a disease, and improve the technical field of disease treatment response modeling. The systems and methods may process genetic information (e.g., exomic data) in a textual format, allowing the exomic information to become interpretable to a language model. Genetic associations indicated by the exomic data may be identified using models incorporating a language model, such as disease progression, patientresponse to treatment, and otherwise patterns encoded in the patient’s genetics that may impact disease treatment. For example, the genetic transformer employing a language model may be trained on genomic data including the human reference genome, diverse human genomes, and genomes from a variety of species. The trained genetic transformer may identify variations in the patient’ s exomes respective to the human reference genome that may be indicative of treatment efficacy. A multimodal disease prediction model may receive the patient’s exomic data, along with medical images and clinical data associated with the disease (e.g., x-rays of feet and / or hands indicating the extent of joint damage from RA and relevant laboratory results, clinical notes medications, and medical history). Analyzing a diverse, comprehensive set of information may allow the disease prediction to make more accurate predictions than analyzing exomic information alone. Moreover, the genomic transformer and multimodal disease prediction model may be fine-tuned (e.g., modifying a classifier head) to make predictions associated with other diseases, such as cancer (e.g., blood cancer), irritable bowel syndrome, autoimmune diseases, as well as others.

[0023] The disclosed model architecture provides advanced technological capabilities compared to conventional models, such as the analysis of full exomes rather than single nucleotide polymorphisms, the capturing and analysis exomic interactions across different chromosomes and parts of the genome (e.g., short neighborhood and long neighborhood), and implementation into different clinical applications, systems (e.g., EPIC medical record systems) and workflows with minimal effort. As most RA patients continue to experience moderate disease activity and disease flares rather than achieving sustained remission, disease insights and predictions via the disclosed techniques can be an invaluable patient treatment decision support tool for medical providers when determining the medication best suited to treat the patient.

[0024] In at least some embodiments the disclosed system and methods provide novel techniques for preprocessing exomic data (e.g., for training and / or fine-tuning a model). Raw genetic sequencing data is converted to other formats, tokenized and / or compressed to reduced file size, represent entire exomes, pack 151 base-pair reads together to create tokens having an appropriate length for model processing, and identify exomic regions and / or variations of interest. Preprocessing the training data improves the time required to train / fine-tune thegenomic model, reduces computing resources for model training such as processing cycles, memory, and / or network bandwidth required as compared to non-preprocessed training data (e.g., substantially larger in size). Accordingly, the training data preprocessing improves the operation of the computer / computing environment and advances the technology of preprocessing exomic data for model training.

[0025] The present disclosure generally refers to predicting a patient’s response to Methotrexate treatment of RA, however, it should be understood that the techniques disclosed herein may be easily applied to other diseases (e.g., cancer, autoimmune) and / or other predictions (e.g., probability of developing a disease, disease progression, disease mortality). Moreover, as used herein, genomic information / data or genetic information / data may generally refer to the entire genome (e.g., all DNA sequences, coding regions, non-coding regions), whereas exomic information / data may generally be limited to the exome (e.g., the coding regions of genes, variations in the protein-coding regions).EXEMPLARY COMPUTER ENVIRONMENT

[0026] FIG. 1 depicts a block diagram of an exemplary computing environment 100 for generating a prediction of a response of a patient to a medication to treat a disease, according to embodiments. The computing environment 100 may include a server 105 communicatively coupled, via a network 110, to one or more databases 108 and a user device 115. Although FIG.1 depicts certain entities, components, equipment, and devices, it should be appreciated that additional or alternate entities, components, equipment, and devices are envisioned.

[0027] The server 105 may perform the functionalities associated with disease prediction, such obtaining patient data, preprocessing model training data, model training, model execution, and generating disease predictions. The server 105 may be part of a cloud network or may otherwise communicate with other hardware or software components within one or more cloud computing environments to send, retrieve, or otherwise analyze data or information described herein. For example, in certain aspects of the present techniques, the computing environment 100 may comprise an on-premise computing environment, a multi-cloud computing environment, a publiccloud computing environment, a private cloud computing environment, and / or a hybrid cloud computing environment. For example, an entity (e.g., a medical institution) may host one or more services in a public cloud computing environment (e.g., Alibaba Cloud, Amazon Web Services (AWS), Google Cloud, IBM Cloud, Microsoft Azure, etc.). The public cloud computing environment may be a traditional off-premise cloud (i.e., not physically hosted at a location owned / controlled by the medical institution). Alternatively, or in addition, aspects of the public cloud may be hosted on-premises at a location owned / controlled by the medical institution. The public cloud may be partitioned using visualization and multi-tenancy techniques and may include one or more inl'rastructurc-as-a-scrvicc (laaS) and / or platform-as-a- service (PaaS) services.

[0028] A network 110 may comprise any suitable network or networks, including a local area network (LAN), wide area network (WAN), Internet, or combination thereof. For example, the network 110 may include a wireless cellular service (e.g., 4G, 5G, 6G, etc.). Generally, the network 110 enables bidirectional communication between the servers 105 and a user device 115. In one aspect, the network 110 may comprise a cellular base station, such as cell tower(s), communicating to the one or more components of the computing environment 100 via wired / wireless communications based upon any one or more of various mobile phone standards, including NMT, GSM, CDMA, UMTS, LTE, 5G, 6G, and / or the like. Additionally, or alternatively, the network 110 may comprise one or more routers, wireless switches, or other such wireless connection points communicating to the components of the computing environment 100 via wireless communications based upon any one or more of various wireless standards, including by non-limiting example, IEEE 402.11a / ac / ax / b / c / g / n (Wi-Fi), Bluetooth, and / or the like.

[0029] The server 105 may include one or more processors 102. The processors 102 may include one or more suitable processors (e.g., central processing units (CPUs) and / or graphics processing units (GPUs)). The processors 102 may be connected to a memory 104 via a computer bus (not depicted) responsible for transmitting electronic data, data packets, or otherwise electronic signals to and from the processors 102 and memory 104 in order to implement or perform the machine-readable instructions, methods, processes, elements, orlimitations, as illustrated, depicted, or described for the various flowcharts, illustrations, diagrams, figures, and / or other disclosure herein. The processors 102 may interface with the memory 104 via a computer bus to execute an operating system (OS) and / or computing instructions contained therein, and / or to access other services / aspects. For example, the processors 102 may interface with the memory 104 via the computer bus to create, read, update, delete, or otherwise access or interact with the data stored in the memory 104 and / or a database 108.

[0030] The memory 104 may include one or more forms of volatile and / or non-volatile, fixed and / or removable memory, such as read-only memory (ROM), electronic programmable readonly memory (EPROM), random access memory (RAM), erasable electronic programmable read-only memory (EEPROM), and / or other hard drives, flash memory, MicroSD cards, and others. The memory 104 may store an operating system (OS) (e.g., Microsoft Windows, Linux, UNIX, etc.) capable of facilitating the functionalities, apps, methods, or other software as discussed herein.

[0031] In general, a computer program or computer based product, application, or code (e.g., the model(s), such as ML models, or other computing instructions described herein) may be stored on a computer usable storage medium, or tangible, non-transitory computer-readable medium (e.g., standard random access memory (RAM), an optical disc, a universal serial bus (USB) drive, or the like) having such computer-readable program code or computer instructions embodied therein, wherein the computer-readable program code or computer instructions may be installed on or otherwise adapted to be executed by the processor(s) 102 (e.g., working in connection with the respective operating system in memory 104) to facilitate, implement, or perform the machine readable instructions, methods, processes, elements or limitations, as illustrated, depicted, or described for the various flowcharts, illustrations, diagrams, figures, and / or other disclosure herein. In this regard, the program code may be implemented in any desired program language, and may be implemented as machine code, assembly code, byte code, interpretable source code or the like (e.g., via Golang, Python, C, C++, C#, Objective-C, Java, Scala, ActionScript, JavaScript, HTML, CSS, XML, etc.).

[0032] The server 105 may include, and / or have access to (e.g., via network 110) one or more databases 108. The database 108 may be a relational database, such as Oracle, DB2, MySQL, a NoSQL based database, such as MongoDB, or another suitable database. The database 108 may store data and / or datasets include one or more types of data, records, files, etc., however the term data and dataset may be used interchangeably herein. Some types of data and / or datasets stored in the database 108, memory 104, and / or otherwise available to the computing environment 100 may be publicly available datasets (e.g., from a laboratory), private datasets (e.g., of patients), or any other suitable type of data.

[0033] The computing environment 100 may include a model training database 108A storing datasets to train and / or operate one or more models, as further described below with respect to FIG. 3. The model training database 108A may store, electronic health record data, clinical data, medical image data, human reference genome data, exomic data, multispecies genomic data, historical predictions, and / or any other suitable training data for a model disclosed herein.

[0034] The computing environment 100 may include an EHR database 108B storing electronic health records (EHRs), also referred to as electronic medical records (EMRs), of patients. The EHR database 108B may include genomic data from sequencing the genes of the patient or biological information (e.g., proteomic data, transcriptomic data, exomic data, metabolomic data, lipidomic data, epigenomic data). The 108B may store clinical data (e.g., blood count, medical history information, prescription information, demographic information, bloodwork results, clinical notes, reports associated with medical images, longitudinal patient records), and / or other suitable data. The 108B may store x-rays, computer tomography (CT) images, ultrasound images, magnetic resonance imaging (MRI) data, positron emission tomography (PET) image data, or other suitable medical images of patients. For example, X- rays of feet and hands of a patient may indicate the extent of joint damage which may be a predictor for RA disease progression, morbidity, etc. The EHR database 108B may store complete blood count (CBC) datasets including CBC data of one or more patients. The CBC data may include and / or indicate one or more of a blood count, blood indices (e.g., MCV, MCHC, RDW, MPV), a blood indices determination, a morphology, genetic sequencing, or other suitable data associated with the blood of the patient. The CBC data may indicate relationshipsand / or associations with a patient developing blood-based cancers / malignancies (e.g., acute myeloid leukemia, myelodysplastic syndromes, CLL, etc.), and / or overall mortality, due to cardiovascular diseases or other diseases. For example, the CBC data may be obtained from a medical record of a patient from blood tests results. The EHR database 108B may include longitudinal patient record (LPR) data, e.g., a patient record combining data from a variety of sources, such as one or more healthcare systems over the course of a patient’s life. The LPR data may include retrospective patient data indicating patient outcomes from a disease.

[0035] The computing environment 100 may include genomic database 108C storing genetic data from a plurality of humans (e.g., participants in the Mayo Clinic Tapestry DNA Sequencing Research Study), multispecies genomic data (e.g., of non-human species form the National Institute of Health repository), genetic information of the human reference genome, and / or other suitable genetic, genomic, and / or exomic data.

[0036] The memory 104 may store a plurality of computing modules 1 12, implemented as respective sets of computer-executable instructions (e.g., one or more source code libraries) as described herein. The computing modules 112 may include a data preprocessing module 122. The data preprocessing module 122 may perform preprocessing of one or more types of data, such as genomic data, before providing the data to a model (e.g., the 132 / / ), as further described in FIGs. 2A-2F.

[0037] The computing modules 112 may include an ML module 114. The ML module 114 may include ML training module (MLTM) 116 and / or ML operation module (MLOM) 118. In some embodiments, at least one of a plurality of ML methods and algorithms may be applied by the ML module 114, which may include, but are not limited to: linear or logistic regression, instance-based algorithms, regularization algorithms, decision trees, Bayesian networks, cluster analysis, association rule learning, artificial neural networks, deep learning, combined learning, reinforced learning, dimensionality reduction, and support vector machines. In various embodiments, the implemented ML methods and algorithms are directed toward at least one of a plurality of categorizations of ML, such as supervised learning, unsupervised learning, and reinforcement learning. In one aspect, the ML based algorithms may be included as a library orpackage executed on server(s) 105. For example, libraries may include the TensorFlow based library, the PyTorch library, and / or the scikit-learn Python library.

[0038] In one embodiment, the ML module 114 employs supervised learning, which involves identifying patterns in existing data to make predictions about subsequently received data. Specifically, the ML module is “trained” (e.g., via MLTM 116) using training data, which includes exemplary inputs and associated exemplary outputs. Based upon the training data, the ML module 114 may generate a predictive function which maps outputs to inputs and may utilize the predictive function to generate ML outputs based upon data inputs. The exemplary inputs and exemplary outputs of the training data may include any of the data inputs or ML outputs described above. In the exemplary embodiments, a processing element may be trained by providing it with a large sample of data with known characteristics or features.

[0039] In another embodiment, the ML module 114 may employ unsupervised learning, which involves finding meaningful relationships in unorganized data. Unlike supervised learning, unsupervised learning does not involve user- initiated training based upon exemplary inputs with associated outputs. Rather, in unsupervised learning, the ML module 114 may organize unlabeled data according to a relationship determined by at least one ML method / algorithm employed by the ML module 114. Unorganized data may include any combination of data inputs and / or ML outputs as described above.

[0040] In yet another embodiment, the ML module 114 may employ reinforcement learning, which involves optimizing outputs based upon feedback from a reward signal. Specifically, the ML module 114 may receive a user-defined reward signal definition, receive a data input, utilize a decision-making model to generate the ML output based upon the data input, receive a reward signal based upon the reward signal definition and the ML output, and alter the decision-making model so as to receive a stronger reward signal for subsequently generated ML outputs. Other types of ML may also be employed, including deep or combined learning techniques.

[0041] The MLTM 116 may comprise a set of computer-executable instructions implementing ML training (e.g., model creation, fine-tuning, retraining, etc.). The MLTM 116 may access one or more databases 108 (e.g., the model training database 108 A) or any other data source for training data suitable to generate and / or otherwise train one or more ML models. The trainingdata may be sample data with assigned relevant and comprehensive labels (classes or tags) used to fit the parameters (weights) of an ML model with the goal of training it by example. In one aspect, once an appropriate ML model is trained and validated to provide accurate predictions and / or responses, the trained model may be loaded into MLOM 118 at runtime to process input data and generate output data.

[0042] The MLTM 116 may receive labeled data at an input layer of a model having a networked layer architecture (e.g., an artificial neural network, a convolutional neural network, etc.) for training the one or more ML models. The received data may be propagated through one or more connected deep layers of the ML model to establish weights of one or more nodes, or neurons, of the respective layers. Initially, the weights may be initialized to random values, and one or more suitable activation functions may be chosen for the training process. The present techniques may include training a respective output layer of the one or more ML models. The output layer may be trained to output a prediction, for example.

[0043] The MLOM 118 may comprise a set of computer-executable instructions implementing ML loading, configuration, initialization and / or operation functionality. The MLOM 118 may include instructions for storing trained models (e.g., in the electronic database 108). As discussed, once trained, the one or more trained ML models may be operated in inference mode, whereupon when provided with de novo input that the model has not previously been provided, the model may output one or more predictions, classifications, etc., as described herein.

[0044] While various embodiments, examples, and / or aspects disclosed herein may include training and generating one or more ML models for the server 105 to load at runtime, it is also contemplated that one or more appropriately trained ML models may already exist (e.g., in database 108) such that the server 105 may load an existing trained ML model at runtime. It is further contemplated that the server 105 may retrain, fine-tune, update and / or otherwise alter an existing ML model before and / or after loading the model at runtime.

[0045] In one aspect, the computing modules 112 may include an input / output (I / O) module 120, comprising a set of computer-executable instructions implementing communication functions. The I / O module 120 may include a communication component configured to communicate (e.g., send and receive) data via one or more external / network port(s) to one ormore networks or local terminals, such as the computer network 110 and / or the user device 115 (for rendering or visualizing) described herein. In one aspect, the servers 105 may include a client-server platform technology such as ASP.NET, Java J2EE, Ruby on Rails, Node.js, a web service or online API, responsive for receiving and responding to electronic requests. VO module 120 may further include or implement an operator interface configured to present information to an administrator or operator and / or receive inputs from the administrator and / or operator. An operator interface may provide a display screen. The VO module 120 may facilitate VO components (e.g., ports, capacitive or resistive touch sensitive input panels, keys, buttons, lights, LEDs), which may be directly accessible via, or attached to, servers 105 or may be indirectly accessible via or attached to the user device 115. According to one aspect, an administrator or operator may access the servers 105 via the user device 115 to review information, make changes, input training data, initiate training via the MLTM 116, and / or perform other functions (e.g., operation of one or more trained models via the MLOM 1 18).

[0046] The computing modules 112 may include a data preprocessing module 122. The data preprocessing module 122 may preprocess model training data, such as exomic training data for training the genomic transformer 126, and / or training data for any other model. For example, the data preprocessing module 122 may preprocess binary alignment map (BAM) training data indicating patient sequence alignment data (e.g., aligned reads from a genomic sequencing platform). The preprocessing may include generating BED and FASTQ data files from the BAM data / files, filtering out human reference genome matches or noisy data, tokenizing the FASTQ and BED data, and / or compressing the data into an NPZ data binary file, as further described below. The NPZ data may include array data stored using in stored in a compressed format (e.g., NumPy array stored using gzip compression).

[0047] The computing modules 112 may include one or more transformers 124. One or more of the transformers 124 may be based upon a decoder-only generative pre-trained transformer (GPT) architecture. The transformers 124 may include a genomic transformer 126 for transforming genomic data (e.g., whole genomes, whole exomes), a vision transformer 128 for transforming image data (e.g., medical images, x-rays, etc.), and language transformer 130 for transforming clinical data (e.g., unstructured or structured text such as medical records,laboratory results, clinical notes medications, medical history). Each of the transformers 124 may generate embeddings or other outputs based upon receiving a corresponding modality of data as an input. For example, the vision transformer 128 may receive medical image data (e.g., x-rays) and generate associated embeddings that are input into the disease prediction model 132. Similarly, the language transformer 130 may receive clinical data (e.g., structed and unstructured text) and generate associated embeddings that are input into the disease prediction model 132.

[0048] In at least some embodiments, the genomic transformer 126 may receive genomic information of a patient (e.g., form the EHR database 108B, the genomic database 108C) and generate embedding of exomic data as an output, as further described below. The genomic transformer 126 may include one or more of a short context model trained to understand a short neighborhood of exomic information, or a long context model trained to understand exomic information from at least one entire chromosome, as further described below.

[0049] The memory 104 and / or the computing modules 1 12 may store one or more models, including the disease prediction model 132. The disease prediction model 132 may be trained to receive medical image data (e.g., x-rays), exomic data (e.g., generated by the genomic transformer 126), and clinical data (e.g., laboratory results, clinical notes medications, medical history, etc.), and in response generate one or more predictions associated with a disease. The prediction may be associated with one or more of: exomic predictions (e.g. pathogenic or non- pathogenic exomic variations), treatment response and / or outcome predictions such as predicting patient responsiveness to medication (which medication, patient toxicity to medication, etc.), disease progression, disease morbidity, patient hospitalization, joint damage, surgeries (e.g., short and long-term), etc. For example, the disease prediction model 132 may generate a prediction of a patient’s response to methotrexate to treat RA based upon receiving exomic data indicating the patient’s exomic variation associated with RA, medical images of the patient’s hands and feet depicting RA progression, and clinical notes associated with the patient’s RA. The disease prediction model 132 may include one or more transformers, multimodal models, classifiers, language models (e.g., large language, small language, tiny language, hybrid language), and / or other suitable models.

[0050] The computing environment 100 may include at least one user device 115. The user device 115 may comprise one or more computers, such as multiple, redundant, or replicated client computers accessed by one or more users. The user device 115 may access devices, services, and / or other components of the computing environment 100 via the network 110. The user device 115 may include a desktop computer, laptop, mobile computing device, augmented, virtual, mixed, and / or extended reality glasses / headsets, chatbots, and / or other electronic or electrical components. The user device 115 may include a processor 142 (e.g., the processor 102), a memory 144 (e.g., the memory 104), and a NIC 146 (e.g., the NIC 106). The user device 115 may include a user interface 148, such as one or more components to receive an input and / or generate an output. The user interface 148 may include one or more of a keyboard, a mouse, a touchscreen, a microphone, a speaker, an imaging device, etc.

[0051] The user device 115 may store a disease prediction application 150 in the memory 144. The disease prediction application 150 may cause the server 105 or other suitable computing device to generate the one or more predictions associated with a disease. In one example, the disease prediction application 150 may generate a prediction request which the user device 115 transmits to the server 105 via the network 110. In response, the server 105 may generate and transmit the prediction to the user device 115. In other embodiments, the disease prediction application 150 may generate the prediction at the user device 115, e.g., by obtaining (e.g., from the memory 144, from a database 108, the server 105, etc.) the genomic transformer 126, the disease prediction model 132, and patient data provided as inputs to the genomic transformer 126 and the disease prediction model 132, to generate the prediction.

[0052] In operation, the server 105 may obtain and preprocess model training data via the data preprocessing module 122, for example preprocessing raw genomic information obtained from the model training database 108 A and / or the genomic database 108C, as further described with respect to FIGs. 2A-2F. The data preprocessing module 122 may generate model training data for the genomic transformer 126, and store the model training data the model training database 108A. The server 105 (e.g., via the MLTM 116) may also train one or more of the transformers 124 and / or the disease prediction model 132. Training the genomic transformer 126 may include evaluating the performance of the genomic transformer 126 to make exomic predictions (e.g.which changes are pathogenic or non-pathogenic), identify patients with a specific disease (e.g., RA), and / or predict which medication the patient will respond to for treating the disease. Training the disease prediction model 132 may include assessing the accuracy of predictions of patient response to Methotrexate (e.g., many patients were accurately identified as Methotrexate responders).

[0053] The trained the genomic transformer 126 and / or disease prediction model 132 may provide one or more predictions associated with a disease. For example, the user device 115 may be located at a medical facility and integrated with a patient medical record system (e.g., EPIC®). A user of the user device 115 may request a prediction of treatment response for a patient via the disease prediction application 150 integrated with the patient medical record system. In response, the disease prediction application 150 may generate a request for the prediction, and transmit the request to the server the server 105 via the network 110. Responsive to receiving the request, the server 105 may obtain genomic data, medical image data, and clinical data of the patient from the EHR database 108B. The server 105 load the genomic transformer 126 via the MLOM 118. The MLOM 118 may provide the genomic data to the genomic transformer 126. In response, the genomic transformer 126 may generate omic data (e.g., via extraction from the genomic data). The MLOM 118 may load the disease prediction model 132, and provide the omic data, medical image data, and clinical data to the disease prediction model 132. In response, the disease prediction model 132 may generate a disease prediction, such as the response of the patient suffering from RA to Methotrexate as an RA treatment.

[0054] The computing environment 100 may include additional, fewer, and / or alternate components, and may be configured to perform additional, fewer, or alternate actions, including components / actions described herein. Although the computing environment 100 is shown in FIG. 1 as including one instance of various components such as user device 115, server 105, network 110, etc., various aspects include the computing environment 100 implementing any suitable number of any of the components shown in FIG. 1 and / or omitting any suitable ones of the components shown in FIG. 1. For instance, information described as being stored in the model training database 108A may be stored in the memory 104, and therefore the modeltraining database 108A may be omitted. Moreover, various aspects include the computing environment 100 including any suitable additional component(s) not shown in FIG. 1, such as but not limited to the exemplary components described above. Furthermore, it should be appreciated that additional and / or alternative connections between components shown in FIG. 1 may be implemented. As just one example, server 105 and the model training database 108A may be connected via a direct communication link (not shown in FIG. 1) instead of, or in addition to, via the network 110.TRAINING DATA PREPARATION

[0055] The genomic transformer 126 may be trained to receive genomic data (e.g., BAM data) and generate associated exomic embeddings as an output. The genomic transformer 126 may be trained using human reference genome data associated with the human reference genome, multispecies genomic data of multiple species, and exomic data of a plurality of humans (e.g., from the Tapestry project), which may be stored in one or more databases (e.g., the model training database 108A, the genomic database 108C). The human reference genome data may inform the genomic transformer 126 of sequences or otherwise genomic information associated with the human reference genome. The plurality human exomic data may inform the genomic transformer 126 of naturally occurring genetic diversity (e.g., genetic variants, their frequency) across the human species. The multispecies genomic data may allow the genomic transformer 126 to identify the species of subject genomic data and / or increase the understanding of the genomic transformer 126 (e.g., for transfer learning).

[0056] The training data for the genomic transformer 126 may undergo preprocessing to allow the genomic transformer 126 to digest entire exomes (e.g., for capturing long range interactions), identifying clinically important variants of the exome associated with the disease respective to the human reference genome, among other things. Preprocessing the genomic transformer training data may include (i) transforming the BAM files into BED and FASTQ files; (ii) filtering and compressing the BED and FASTQ files into NPZ files; and (iii) statistical analysis of the compressed NPZ files to generate an H5 data for ingestion by the genomic transformer 126.

[0057] FIG. 2A depicts a block diagram of an example training data preprocessing process 200 of genomic transformer training data (e.g., via the data preprocessing module 122), according to embodiments. One or more steps of the preprocessing process 200 may be performed by the server 105 (e.g., a high-performance computing environment such as a Google cloud platform). Preprocessing the training data may exemplify structural variances in the exomes of the training data that are associated with a disease. The preprocessing process 200 may include obtaining 210 a patient’s BAM data (e.g., a BAM file from the genomic database 108C). The BAM data may include one or more binary files that stores sequence alignment data from aligned reads obtained by a genomic sequencing platform. The BAM data may include a header section and an alignment section, and may contain information such as read name, sequence, quality, and alignment position.

[0058] The preprocessing process 200 may include a conversion process 220 for converting the BAM data to BED data, FASTQ data, and log files. The BED data may include one or more text files that store genomic regions of interest, such as genes, exons, or variants. The BED data may contain information arranged one or more rows (e.g., of a matrix or table), each row indicating a chromosome name, chromosome start coordinate, chromosome end coordinate, a name of the line in the BED file, a score, and a DNA strand orientation. The rows of a BED file may include a header section, a body section, and use whitespace delimiters.

[0059] FIG. 2B depicts a block diagram of the file conversion process 220 for converting the BAM data to browser extensible (BED) data, FASTQ data, and log files, according to embodiments. The file conversion process 220 may include the data preprocessing module 122 receiving BAM data 222 including one or more BAM files, and converting 224 a BAM file to generate 226 one or more BED files using BED conversion parameters such as minimum mapping quality, concise idiosyncratic gapped alignment report (CIGAR) score, minimum read length, and minimum coverage. The data preprocessing module 122 may perform the file conversion process 220 for each BAM file in the BAM data.

[0060] The file conversion process 220 may include the data preprocessing module 122 merging, sorting, parsing, and / or filtering one or more of the BED files using parameters such as minimum region length, maximum region length, and minimum region score. The tileconversion process 220 may include the data preprocessing module 122 converting the BAM data into two FASTQ data using FASTQ parameters such as reference genome, variant calling method, and variant filtering criteria. The FASTQ data may include at least two FASTQ files, including a first FASTQ fde containing forward reads of the genomic data, and a second FASTQ file containing the reverse reads of the genomic data. The file conversion process 220 may include generating one or more log files that log the conversion processes of the BAM data to BED data and FASTQ data. The file conversion process 220 may include verifying the integrity and completeness of the transformed BED data and FASTQ data files using checksums and file sizes. The file conversion process 220 may include storing the BED data and / or FASTQ data (e.g., in the model training database 108A).

[0061] The preprocessing process 200 may include a BED and FASTQ file filtering and compression process 230. The data compression may cause the genomic transformer 126 to digest entire exomes. FIG. 2C depicts a block diagram of a training data filtering and compression process 230, according to embodiments. The filtering and compression process 230 may include processing each row of a BED data file to retrieve, trim, and / or align each genomic sequence indicated by the BED data. The following rules may be applied to each BED row processed to retrieve, trim, and align each genomic sequence:

[0062] Confirm a final CIGAR Score has a matching score (M) of greater than or equal to 131;

[0063] Use the Read Name string from the BED file row to lookup the 151-character (A, C, T, G, N values) sequence from the FASTQ file;

[0064] Apply a substring trimming method to the 151-character long sequence;

[0065] If soft clip / insertion / deletion, pad the sequence to 170 characters, filtering reads to only one CIGAR change;

[0066] Check the strand score and apply the reverse complement for alignment purposes if the strand is denoted as a negative (-) strand;

[0067] All alternate contigs tracked by the contig index and are tokenized. The contig index may be added to the BED file so a separate token for the contig index and name is generated; and

[0068] Performing HLA typing (e.g., to train the genomic transformer 126 for diseases such as autoimmune diseases and cancer.Tokenization

[0069] FIG. 2D depicts a block diagram of an example tokenization process 240, according to embodiments. The tokenization process 240 may convert an alpha value for each nucleotide into an integer value enabling proper processing downstream by the genomic transformer 126.Match Case

[0070] All 15 IM sequences may be organized in their own BED file. The CIGAR score may be confirmed as 15 IM match, and then used to read name to find the proper sequence from the FASTQ file and then generate the tokenization. Once the sequence is found, the preprocessing process 200 may include checking the strand value, which can be negative (-) or positive (+). The strand value may indicate when a reverse complement of the sequence is required to enable proper alignment across all sequences before finally tokenizing each individual nucleotide. The start location of every sequence may be tracked from the BED file with a separate token for the start position.

[0071] FIG. 2E depicts a block diagram of an example soft clipping process 250, according to embodiments. The soft clipping process 250 may be performed to highlight structural variances. Soft clipping may include retaining the clipped bases in the read data, and include deletions that indicate missing bases and do not retain any data for the deleted segment. Soft clipping may be used to deal with low quality reads or partially aligned reads. The soft clipping process 250 may include receiving 252 a patient data / files including at least one BED file, and performing clipping 254 via matching soft clipping (MS) and / or soft clipping matching (SM) on one or more BED files. Matching soft clipping may include situations where the initial portion of the sequence is perfectly aligned, and before the 151 mark is cleared, a portion of the sequence is clipped. Soft clipping matching may be considered the opposite of matching soft clipping, where the initial portion of the sequence is absent, and the rest of the read is intact and aligned. Both types of clipping may leverage the nucleotide token values for the matching and filler values tocreate a sequence of 151 tokenized values. The substring value identified within the BED file row may be kept when the MS or SM pattern is detected for the sequence. The data preprocessing module 122 may perform trimming or removing of the soft clipping part using the substring value identified during the bio-informatics process. The strand value is used to determine an application of the reverse complement in cases where the strand value is a negative (-) value. The data preprocessing module 122 may tokenize the sequence using the matching values and for any substring characters that are removed, and add a filler token (value of 127) to retain the 151-length of the tokenized sequence. Fillers tokens may be placed at the end of the sequence, such that the stall position is always at the same location in the tokenized sequence.Deletion

[0072] In the case of deletion, there will be two sequence rows in the BED file. Processing the first BED row for a deletion case may include finding the first substring sequence using the Read Name from the FASTQ file. When the strand value is negative (-), the soft clipping process 250 may include applying the reverse complement before tokenizing the remaining parts of the sequence. Once the substring is applied and the strand logic is used for alignment, the soft clipping process 250 may include adding the deleted nucleotide token value (-5) for each identified deleted token in the original CIGAR score (example 140M1D11M has a ID, requiring appending one deleted token value during the tokenization process). Filler tokens may be added for a maximum token length of 170 tokens for the entire sequence.Insertion Case

[0073] In the case of insertion, there will be two rows in the BED file. The soft clipping process 250 may include identifying in the BED row the sequence substring from the FASTQ file matching the read name. When applying the strand reverse complement, if the strand value in the row is negative (-) or in a negative strand case, the insertion may be moved to the front of the sequence after applying the reverse complement. For the positive case the insertion may remain at the end of the sequence. Any insertion added to the sequence may use the nucleotide token values denoted in an inserted values table (A=5, T=6, G=7, C=8), to indicate the tokens areinsertions. The soft clipping process 250 may include filling the rest of the row with the filler token (127) until the entire sequence is a 170 tokenized nucleotide length.

[0074] The soft clipping process 250 for the second row of the insertion case may include applying a substring method and strand method, ands using the filler token to make the length of the token sequence equal to 170 tokens.

[0075] The above process is iterated over every row in the BED file and all unneeded characters may be stripped. After tokenization is complete, the data preprocessing module 122 may generate NPZ files representing the compressed state of the reads (e.g., a compressed binary file format used to store NumPy arrays in a).Genomic Model Statistical Analysis and H5 Data Preparation

[0076] FIG. 2F depicts a block diagram of an example NPZ to H5 statistical analysis and data preparation process 260, according to embodiments. The data preprocessing module 122 (e.g., via Python scripts) may generate H5 files based upon a statistical analysis of NPZ files.

[0077] The H5 file type may store a large amount of exomic data created after the data preprocessing process. The statistical analyses may include ensuring the data preprocessing conserves all of the information from the original BAM files. The NPZ to H5 statistical analysis and data preparation process 260 may include quantifying the number of matching the human reference genome to the token value of the nucleotide (e.g., a true or false determination). The NPZ to H5 statistical analysis and data preparation process 260 may include quantifying the per position read pile up score that may indicate how many times a position occurs in the exomic (e.g., Tapestry) data.

[0078] In at least some embodiments, the genomic transformer 322 may be trained using transplant learning / training using different approaches including different compression techniques to identify best construction of tokens, and different packing techniques of the 151 base pairs reads into a token input vector (e.g., a 1,024 token input vector).Packing Techniques

[0079] The preprocessing process 200 may include data packing to generate the input vectors for the genomic transformer 126. FIG. 2G depicts a block diagram 270 of different packing techniques, according to embodiments. The packing techniques may include one or more of: (i) random packing 272 by randomly concatenating reads; (ii) consecutive spacing packing 274 by concatenating ordered reads from same chromosome, spaced by read number; (iii) consecutive packing 276 by concatenate consecutive reads; and / or (iv) consecutive overlap packing 278 by create the longest possible sequence using overlapping data.MODEL TRAINING

[0080] FIG. 3 depicts a block diagram 300 of an example training process of a model (e.g., the genomic transformer 126, the disease prediction model 132), according to embodiments.Generally, a machine learning (ML) engine 310 (e.g., the MLTM 116) trains a model 320 using training data 330. The ML engine 310 may train the ML model 320 via regression, k-nearest neighbor, support vector regression, and / or random forest algorithms and / or models, although any type of applicable ML algorithm and / or model may be used, including training using one or more of supervised learning, unsupervised learning, semi-supervised learning, and / or reinforcement learning. Once trained, the ML model 320 may perform operations on one or more data inputs 340 to produce a desired data output 350.

[0081] Generally, the training data 330 may include historical training data / datasets, such as information from biological studies such as proteomics, transcriptomics, genomics, metabolomics, lipidomics, exomics, and epigenomics; data generated (e.g., via an LLM) from unstructured datasets / sources (e.g., medical record data, genomic test reports, etc.); medical record data, genomic sequencing data, rheumatology notes, demographic information, clinical information, disease characteristics, laboratory values, medications, disease- specific markers, x- ray data, computer tomography (CT) image data, ultrasound image data, magnetic resonance imaging (MRI) data, positron emission tomography (PET) image data, or other suitable training data 330, as further described below.

[0082] The server 105 and / or other suitable device(s) or components(s) may update the training data 330, e.g., to include new data indicative of predictions associated with a disease,new medical images, new clinical notes, new human genomic / exomic data, etc. Subsequently, the ML model 320 may be retrained and / or fine-tuned using the updated training data 330. Retraining and / or fine-tuning using updated training data 330 may improve operation of the model 320, cause the model 320 to have additional capabilities, etc.Genomic Transformer

[0083] The ML model 320 may include a genomic transformer 322. In at least some embodiments, the genomic transformer 322 may include a genomics foundation model including an autoregressive, decoder-only transformer architecture. The genomics foundation model may be pretrained on the next token prediction task in a self- supervised manner.

[0084] The genomic transformer 322 may be trained using genomic transformer training data 332. The genomic transformer training data 332 may include historical data such as HRC data of the HRC, multispecies genomic data of a plurality of species, genomic data of a plurality of humans, exomic data of a plurality of humans, and / or any other suitable data to train the genomic transformer 322. The genomic transformer 322 may be configured to process the genomic transformer training data 332 to learn associations and relationships in the genomic transformer training data 332, e.g., relationships indicating exomic variations of humans or other genetic indicators associated with a disease (e.g., RA); genomic and exomic similarities and / or variations between the human reference genome, genomes and exomes of different species, and / or genomes and exomes of humans; exomic variations that are pathogenic and / or non-pathogenic; human response to a medications, and / or any other suitable associations or relationships.Short Context Model

[0085] In at least some embodiments, the genomic transformer 322 may include a short context model 322A trained to understand a short neighborhood of exomic information, e.g., when the regions of exome information relevant to a prediction (e.g., a prediction of medication response) may be known, so the prediction is based upon known correlations in the exome regions. The short context model may be trained to distinguish human DNA from bacterial DNA. The short context model 322A may include one or more of a decoder-only GPT styletransformer, up to 20billion parameters, and / or up to 100 layers. In at least one embodiment, the NPZ to H5 statistical analysis and data preparation process 260 may include 1 billion parameters and 20 layers.

[0086] The short context model 322A may be trained to identify unique exomes respective to the human reference genome. Accordingly, the sequence information around exomes of interest (e.g., variants), including the nucleotides shared with the human reference genome, allow the short context model 322A to understands both how and where the patient’s exomes are different.

[0087] Training the short context model 322A may include genomic transformer training data 332 with sequences from the human reference genome and sequences from a plurality of different subjects, allowing the short context model 322A to understand the base reference genome, what variations are seen across subjects, and at what location. The short context model 322A may be trained with genomes from other organisms (e.g., up to 768 species) with multiple genomes per species.

[0088] The short context model 322A may be trained to filter out or otherwise ignore reads from a subject where the complete read perfectly matches the human reference genome with no differing nucleotides, as no new information is provided by these reads as compared to the human reference genome, and avoids the short context model 322A paying disproportionate attention to the human reference genome instead of the different variations across subjects. The short context model 322A may be trained to filter out noisy data (e.g., low coverage and / or low confidence positions).

[0089] To train the short context model 322A to prioritize learning from the per-subject variations in the genomic transformer training data 332 even though they may be encountered less frequently compared to human reference genome sequences, a loss weighting may be applied to each base pair. The loss weighting may be proportional to how many times the variations are identified by the short context model 322A. In this manner, we can break the link between how often the model sees a sequence vs the importance it places on learning from it.

[0090] The loss weighting may include the formula lossseq= ^i=Osequence_lengthlossi*base_weighti*match_HRG(i)Where:loss_i represents the degree of error in the prediction of the model at a given location in the sequence; base_weighti=l I frequencyi and is the inverse of the frequency of a nucleotide’s position in the dataset; and match_HRG(i)= {a if nucleotide at position for a given subject matches nucleotide in ref. genome f> if nucleotide at position for a given subject does not match ref. genome and is used to assign different multipliers for sequences that originate from cxomic data depending on whether the nucleotide at a given position matches the nucleotide at the reference genome at the same position.

[0091] Training the short context model 322A may occur in multiple stages. The first stage may include training the model on the human reference genome and multispecies sequences. The second stage may include per-subject sequences within the training data. Such training allows the short context model 322A to understand the human reference genome such that when presented with per-subject sequences, the short context model 322A understand what is unique about a per-subject sequence and how it compares to the corresponding human reference genome sequence. The short context model 322A training may include concatenating multiple independent sequences of a subject having variations from the human reference genome, separated by a token (e.g., <bos> token) indicating to the short context model 322A the variations co-occurring with each other in a neighborhood. Training the short context model 322A may include concatenating multiple variations relevant to a downstream prediction task.Long Context Model

[0092] In at least some embodiments, the genomic transformer 322 may include a long context model 322B trained to understand exomic information from a whole chromosome and even multiple chromosomes to identify long range interactions and correlations across the variants of a subject. In at least some embodiments, the short context model 322A may include one or more of a decoder-only GPT style transformer, 5 billion parameters or more, and / or 40-55 layers or more in total. The long context model 322B model may serve as a starting point to generatepatient specific embeddings capturing all of the unique exomic signatures of patients, e.g., for integration into other models.

[0093] The genomic transformer training data 332 may be used to train the long context model 322B. The genomic transformer training data 332 may include a genome- wide position embedding. The position embedding for each position in the human reference genome. Such embeddings allow the long context model 322B to understand a “global view” of the whole genome information of a subject even though the long context model 322B may receive the genome information a piecewise manner. The genomic transformer training data 332 for the long context model 322B may include compressed data indicating a subject’s exome information across the whole chromosome or multiple chromosomes into a small number (e.g., 128 to 1,024) of high-dimensional latent vectors (e.g., 1,280 to 2,048 dimensions). The compression may allow the long context model 322B to understand what is unique about each subject. The compressed representation may also be used for downstream classifiers and multimodal models (e.g., the disease prediction model 132).

[0094] The position embedding and compressed data may represent a subject’s whole genome in latent space. The compression may generate a unique set of vectors per patient, that may capture the genetic fingerprint of a patient in a compressed learned format effectively creating a subject-specific embedding. Such embeddings may be used in a classifier head to be disease specific.

[0095] Once trained, the genomic transformer 322 (e.g., the short context model 322A and / or the long context model 322B) may receive genomic information of a patient (e.g., whole genome information from a sequencing platform) as an input 342, and generate exomic data of the patient (e.g., exomic variants associated with a disease such as RA) as an output 352. In at least some embodiments, the ML model 320 may include multiple genomic transformers 322, for example each transformer fine-tuned to perform operations associated with a specific disease such as a first genomic transformer 322 trained to detect exomic variations associated with RA, a second genomic transformer 322 trained to detect exomic variations associated with irritable bowel syndrome, etc.Disease Prediction Model

[0096] In at least some embodiments, the models 320 may include a disease prediction model 324 (e.g., the disease prediction model 132). The disease prediction model 324 may be trained using disease prediction model training data 334. The disease prediction model training data 334 may include historical data such as exomic data of a plurality of subjects, clinical data of a plurality of subjects (e.g., laboratory results, clinical notes medications, medical history, etc.), medical image data of a plurality of subjects (e.g., x-rays of hands and feet), historical disease predictions, and / or any other suitable training data for training the disease prediction model 324. The disease prediction model 324 may be configured to process the disease prediction model training data 334 to learn associations and relationships in the disease prediction model training data 334, e.g., relationships indicating, based upon one or more of a patient’s exomic data, clinical data and / or medical image data, a response to medication to treat the disease, mortality rate from the disease, identification of a disease and / or its progression, and / or any other suitable associations or relationships.

[0097] Once trained, the disease prediction model 324 may receive one or more of exomic data (e.g., the exomic data 352), medical image data (e.g., x-rays) and / or clinical data (e.g., laboratory results, clinical notes medications, medical history) of a patient as an input 344, and generate a predictions associated with a disease as an output 354. The output predictions 354 may include treatment response and / or outcome predictions such as predicting patient responsiveness to medication (e.g., responsiveness to Methotrexate, patient toxicity to medication, etc.), disease progression, disease morbidity, patient hospitalization, joint damage, surgeries (e.g., short and long-term), etc. In at least some embodiments, the ML model 320 may include multiple disease prediction models 324, for example each disease prediction model 324 fine-tuned to perform operations associated with a specific disease such as a first disease prediction model 324 trained to make predictions associated with RA, a second disease prediction model 324 trained to make predictions associated with irritable bowel syndrome, etc.

[0098] In at least some embodiments, the disease prediction model 324 may be trained to generate the disease prediction 354 if one or more of the inputs 344 is not provided. For example, the disease prediction model 324 may generate the prediction output 354 if only receiving exomic data and clinical data of the patient.LLM

[0099] The present techniques may include language modeling via one or more language models, such as an LLM, wherein one or more deep learning models are trained by processing token sequences using a large language model architecture. For example, in some aspects, a transformer architecture (e.g., the genomic transformer 322) may be used to process a sequence of tokens. Such a transformer model may include a plurality of layers including self- attention and feedforward neural networks. This architecture may enable the model to learn contextual relationships between the tokens, and to predict the next token in a sequence, based upon the preceding tokens. During training, the model is provided with the sequence of tokens and it learns to predict a probability distribution over the next token in the sequence. This training process may include updating one or more model parameters (e.g., weights or biases) using an objective function that minimizes the difference between the predicted distribution and a true next token in the training data. Alternatives to the transformer architecture may include recurrent neural networks, long short-term memory networks, gated recurrent networks, convolutional neural networks, recursive neural networks, and other modeling architectures.

[0100] In at least some embodiments, an application programming interface (API) may allow software components to send requests and / or receive responses from the LLM. For example, the software component may send a request typically structured as text input or a prompt via an API. In at least some embodiments, a middleware layer may process or format the data provided to and / or received from the LLM. Middleware can help manage session data, handle errors, parse the output for specific information, and translate it into a format usable by the software component. Determining a software component to communicate with an LLM may involve one or more technical considerations, such as latency, scalability, load balancing, security, data sensitivity, data processing and / or formatting requirements, technical compatibility, compliance and regulatory requirements, etc.

[0101] In some aspects, the ML engine 310 may perform pretraining of a language model, which as used herein generally refers to a process that may span pre-processing of training data and initialization of an as-yet untrained language model. In general, a pre-trained model is one that has no prior training of specific tasks. For example, the model pretraining module mayinclude instructions that initialize one more model weights. In some aspects, model pretraining module may initialize the weights to have random values. The pretraining may train one or more models using unsupervised learning, wherein the one or more models process one or more tokens (e.g., preprocessed data) to learn to predict one or more elements (e.g., tokens). The pretraining may include one or more optimizing objective functions that the model pretraining module applies to the one or more models, to cause the one or more models to predict one or more most- likely next tokens, based on the likelihood of tokens in the training data. In general, the model pretraining causes the one or more models to learn linguistic features such as grammar and syntax. The pretraining may include additional steps, including training, data batching, hyperparameter tuning and / or model checkpointing.

[0102] The model pretraining may include instructions for generating a model that is pretrained for a general purpose, such as general text processing / understanding. The pretrained model may be referred to as a base model or foundational model, in some aspects. The foundational model may be further trained by downstream training process(es), for example via the ML engine 310. The model pretraining generally trains foundational models that have general understanding of language and / or knowledge. Pretraining may be a distinct stage of model training in which training data of a general and diverse nature (i.e., not specific to any particular task or subset of knowledge) is used to train the one or more models. In some aspects, a single model may be trained and copied. Copies of this model may serve as respective base models for a plurality of fine-tuned models.

[0103] In some aspects, foundational models may be trained to have specific levels of knowledge. For example, the model pretraining may train an omic foundation model language model to understand omic “language” or otherwise information. The engine 310 may fine-tune the omic foundational model to detect exomic variants associated with a specific disease. In this way, the omic foundational model can start from a relatively advanced stage, without requiring pretraining of each more advanced model individually. This strategy represents an advantageous improvement, because pretraining can take a long time (many days).LLM Training

[0104] The system and methods to generate and / or train an LLM may consist of three steps: (1) a supervised fine-tuning (SFT) step where a pretrained language model (e.g., a base LLM) may be fine-tuned on a relatively small amount of demonstration data curated by human labelers to learn a supervised policy (SFT ML model) which may generate responses / outputs from a selected list of prompts / inputs. The SFT ML model may represent a cursory model for what may be later developed and / or configured as the LLM model; (2) a reward model step where human labelers may rank numerous SFT ML model responses to evaluate the responses which best mimic preferred human responses, thereby generating comparison data. The reward model may be trained on the comparison data; and / or (3) a policy optimization step in which the reward model may further fine-tune and improve the SFT ML model. The outcome of this step may be the LLM model using an optimized policy. In one aspect, step one may take place only once, while steps two and three may be iterated continuously, e.g., more comparison data is collected on the current ML model, which may be used to optimize / update the reward model and / or further optimize / update the policy.Supervised Fine-Tuning ML Model

[0105] FIG. 4 depicts a combined block and logic diagram 400 for training an LLM, in which the techniques described herein may be implemented, according to embodiments. Some of the blocks in FIG. 4 may represent hardware and / or software components, other blocks may represent data structures or memory storing these data structures, registers, or state variables (e.g., data structures for training data 412), and other blocks may represent output data (e.g., 425). Input and / or output signals may be represented by arrows labeled with corresponding signal names and / or other identifiers. The methods and systems may include one or more servers 402, 404, 406, such as the server 105 of FIG. 1.

[0106] In one aspect, the server 402 may fine-tune a pretrained language model 410. The pretrained language model 410 may be obtained by the server 402 and be stored in a memory, such as memory 104 and / or database 108. The pretrained language model 410 may be loaded by the server 402 (e.g., the MLTM 116, the ML engine 310) for retraining / fine-tuning. A supervised training dataset 412 may be used to fine-tune the pretrained language model 410 wherein each data input prompt to the pretrained language model 410 may have a known output response forthe pretrained language model 410 to learn from. The supervised training dataset 412 may be stored in a memory of the server 402, e.g., the memory 104 or the model training database 108A. In one aspect, the data labelers may create the supervised training dataset 412 prompts and appropriate responses. The pretrained language model 410 may be fine-tuned using the supervised training dataset 412 resulting in the SFT ML model 415 which may provide appropriate responses to user prompts once trained. The trained SFT ML model 415 may be stored in a memory of the server 402, e.g., memory 104 and / or a database 108.

[0107] In one embodiment, the SFT ML model 415 may be fine-tuned based upon one or more profiles, e.g., to understand specific types of data associated with a disease profile to understand exomic variants associated with the disease of the profile.Training the Reward Model

[0108] In one aspect, training the LLM 450 may include the server 404 training a reward model 420 to provide as an output a scaler value / reward 425. The reward model 420 may be required to leverage reinforcement learning with human feedback (RLHF) in which a model (e.g., model 320) learns to produce outputs which maximize its reward 425, and in doing so may provide responses which are better aligned to user prompts.

[0109] Training the reward model 420 may include the server 404 providing a single prompt 422 to the SFT ML model 415 as an input. The input prompt 422 may be provided via an input device (e.g., a keyboard) via the I / O module of the server, such as I / O module 120. The prompt 422 may be previously unknown to the SFT ML model 415, e.g., the labelers may generate new prompt data, the prompt 422 may include testing data stored on database 108, and / or any other suitable prompt data. The SFT ML model 415 may generate multiple, different output responses 424A, 424B, 424C, 424D to the single prompt 422. The server 404 may output the responses 424A, 424B, 424C, 424D via an I / O module (e.g., I / O module 120) to a user interface device, such as a display (e.g., as text responses), a speaker (e.g., as audio / voice responses), and / or any other suitable manner of output of the responses 424A, 424B, 424C, 424D for review by the data labelers.

[0110] The data labelers may provide feedback via the server 404 on the responses 424A, 424B, 424C, 424D when ranking 426 them from best to worst based upon the prompt-responsepairs. The data labelers may rank 426 the responses 424A, 424B, 424C, 424D by labeling the associated data. The ranked prompt-response pairs 428 may be used to train the reward model 420. In one aspect, the server 404 may load the reward model 420 via the ML module (e.g., the ML module 114) and train the reward model 420 using the ranked response pairs 428 as input. The reward model 420 may provide as an output the scalar reward 425.

[0111] In one aspect, the scalar reward 425 may include a value numerically representing a human preference for the best and / or most expected response to a prompt, i.e., a higher scaler reward value may indicate the user is more likely to prefer that response, and a lower scalar reward may indicate that the user is less likely to prefer that response. For example, inputting the “winning” prompt-response (i.e., input-output) pair data to the reward model 420 may generate a winning reward. Inputting a “losing” prompt-response pair data to the same reward model 420 may generate a losing reward. The reward model 420 and / or scalar reward 425 may be updated based upon labelers ranking 426 additional prompt-response pairs generated in response to additional prompts 422.

[0112] In one example, a data labeler may provide to the SFT ML model 415 as an input prompt 422, “Describe the sky.” The input may be provided by the labeler via the user device 115 over network 110 to the server 404 running a chatbot application utilizing the SFT ML model 415. The SFT ML model 415 may provide as output responses to the labeler via the user device 115: (i) “the sky is above” 424A; (ii) “the sky includes the atmosphere and may be considered a place between the ground and outer space” 424B; and (iii) “the sky is heavenly” 424C. The data labeler may rank 426, via labeling the prompt-response pairs, prompt-response pair 422 / 424B as the most preferred answer; prompt-response pair 422 / 424A as a less preferred answer; and prompt-response 422 / 424C as the least preferred answer. The labeler may rank 426 the prompt-response pair data in any suitable manner. The ranked prompt-response pairs 428 may be provided to the reward model 420 to generate the scalar reward 425.

[0113] While the reward model 420 may provide the scalar reward 425 as an output, the reward model 420 may not generate a response (e.g., text). Rather, the scalar reward 425 may be used by a version of the SFT ML model 415 to generate more accurate responses to prompts, i.e., the SFT model 415 may generate the response such as text to the prompt, and the reward model420 may receive the response to generate a scalar reward 425 of how well humans perceive it. Reinforcement learning may optimize the SFT model 415 with respect to the reward model 420 which may realize the configured LLM 450.RLHF to Train the LLM

[0114] In one aspect, the server 406 may train the LLM 450 (e.g., via the ML module 114) to generate a response 434 to a random, new and / or previously unknown user prompt 432. To generate the response 434, the LLM 450 may use a policy 435 (e.g., algorithm) which it learns during training of the reward model 420, and in doing so may advance from the SFT model 415 to the LLM 450. The policy 435 may represent a strategy that the LLM 450 learns to maximize its reward 425. As discussed herein, based upon prompt-response pairs, a human labeler may continuously provide feedback to assist in determining how well the LLM’s 450 responses match expected responses to determine rewards 425. The rewards 425 may feed back into the LLM 450 to evolve the policy 435. Therefore, the policy 435 may adjust the parameters of the LLM 450 based upon the rewards 425 it receives for generating good responses. The policy 435 may update as the LLM 450 provides responses 434 to additional prompts 432.

[0115] In one aspect, the response 434 of the LLM 450 using the policy 435 based upon the reward 425 may be compared using a cost function 438 to the SFT ML model 415 (which may not use a policy) response 436 of the same prompt 432. The server 406 may compute a cost 440 based upon the cost function 438 of the responses 434, 436. The cost 440 may reduce the distance between the responses 434, 436, i.e., a statistical distance measuring how one probability distribution is different from a second, in one aspect the response 434 of the LLM 450 versus the response 436 of the SFT model 415. Using the cost 440 to reduce the distance between the responses 434, 436 may avoid a server over-optimizing the reward model 420 and deviating too drastically from the human-intended / preferred response. Without the cost 440, the LLM 450 optimizations may result in generating responses 434 which are unreasonable but may still result in the reward model 420 outputting a high reward 425.

[0116] In one aspect, the responses 434 of the LLM 450 using the current policy 435 may be passed by the server 406 to the rewards model 420, which may return the scalar reward 425. The LLM 450 response 434 may be compared via cost function 438 to the SFT ML model 415response 436 by the server 406 to compute the cost 440. The server 406 may generate a final reward 442 which may include the scalar reward 425 offset and / or restricted by the cost 440. The final reward 442 may be provided by the server 406 to the LLM 450 and may update the policy 435, which in turn may improve the functionality of the LLM 450.

[0117] To optimize the LLM 450 over time, RLHF via the human labeler feedback may continue ranking 426 responses of the LLM 450 versus outputs of earlier / other versions of the SFT ML model 415, i.e., providing positive or negative rewards 425. The RLHF may allow the servers (e.g., servers 404, 406) to continue iteratively updating the reward model 420 and / or the policy 435. As a result, the LLM 450 may be retrained and / or fine-tuned based upon the human feedback via the RLHF process, and throughout continuing conversations may become increasingly efficient.

[0118] Although multiple servers 402, 404, 406 are depicted in the exemplary block and logic diagram 400, each providing one of the three steps of the overall LLM 450 training, fewer and / or additional servers may be utilized and / or may provide the one or more steps of the LLM 450 training. In one aspect, one server may provide the entire LLM 450 training.EXEMPLARY METHOD FOR GENERATING A PREDICTION OF A RESPONSE OF A PATIENT TO A MEDICATION TO TREAT A DISEASE

[0119] FIG. 5 depicts a flow diagram of an exemplary computer-implemented method 500 for generating a prediction of a response of a patient to a medication to treat a disease, according to embodiments. One or more steps of the method 500 may be implemented as a set of instructions stored on a computer-readable memory and executable via one or more local or remote processors (e.g., the processor 102), servers (e.g., the server 105), and / or other electronic or electrical components, which may be in wired or wireless communication with one another.

[0120] The computer-implemented method 500 may include obtaining genomic data, medical image data, and clinical data of the patient (block 510). The genomic data may include the full genomic information of a patient. The medical image data may include x-rays of one or more of a hand or a foot of the patient. The clinical data may include clinical notes, lab test results, demographic data, medical record data, and / or other structured and / or unstructured text.

[0121] The computer-implemented method 500 may include providing the genomic data to a genomic transformer trained to extract exomic data of the patient associated with the disease (block 520). The exomic data indicates exomic variations of the patient respective to a human reference genome. The genomic transformer may include one or more of a short context model (e.g., the short context model 322A) trained to understand a short neighborhood of exomic information, or a long context model (e.g., the long context model 322B) trained to understand exomic information from at least one entire chromosome. The genomic transformer includes a classifier head fine-tuned to make predictions associated with the disease.

[0122] The computer-implemented method 500 may include providing the medical image data, and the clinical data to a disease prediction model trained to generate the prediction of the response of the patient to the medication (block 530). The disease may be rheumatoid arthritis and the prediction may be a response of the patient to Methotrexate.

[0123] The computer-implemented method 500 may include providing the prediction to a user device (block 540) (e.g., the user device 115).

[0124] In at least some embodiments, the computer-implemented method 500 may include (i) obtaining binary alignment map (BAM) data of a plurality of patients; (ii) preprocessing the BAM data including (a) converting the BAM data to browser extensible (BED) data using BED conversion parameters, (b) converting the BAM data to FASTQ data using FASTQ conversion parameters, (c) generating one or more tokens representing the BED data and the FASTQ data, (d) generating NPZ data from the tokenized BED data and the tokenized FASTQ data, wherein the NPZ data includes array data stored in a compressed format, and (e) generating H5 data based upon analyzing the NPZ data; and (iii) training the genomic transformer using the H5 data.

[0125] In some such embodiments, the BED conversion parameters may include one or more of a minimum mapping quality, a concise idiosyncratic gapped alignment report (CIGAR) score, a minimum read length, or a minimum coverage. In some such embodiments, the FASTQ conversion parameters may include one or more of a reference genome, a variant calling method, or a variant filtering criteria. In some such embodiments, training the genomic transformer may include fine-tuning a genomic transformer foundational model trained using human referencegenome data of a human reference genome, multispecies genomic data of multiple species, and exomic data of a plurality of humans.

[0126] It should be understood that not all blocks of the exemplary flow diagram of FIG. 5 are required to be performed.Additional Exemplary Aspects

[0127] Aspect 1. A computer-implemented method for generating a prediction of a response of a patient to a medication to treat a disease, the computer-implemented method comprising: obtaining, by one or more processors, genomic data, medical image data, and clinical data of the patient; providing, by the one or more processors, the genomic data to a genomic transformer trained to extract exomic data of the patient associated with the disease; providing, by the one or more processors, the exomic data, the medical image data, and the clinical data to a disease prediction model trained to generate the prediction of the response of the patient to the medication; and providing, by the one or more processors, the prediction to a user device.

[0128] Aspect 2. The computer- implemented method of aspect 1, wherein the exomic data indicates exomic variations of the patient respective to a human reference genome.

[0129] Aspect 3. The computer- implemented method of aspect 1 or 2, wherein the medical image data includes x-rays of one or more of a hand or a foot of the patient.

[0130] Aspect 4. The computer-implemented method of any one of aspects 1-3, wherein the clinical data includes clinical notes associated with the disease.

[0131] Aspect 5. The computer-implemented method of any one of aspects 1-4, wherein the disease is rheumatoid arthritis; and the prediction is a response of the patient to Methotrexate.

[0132] Aspect 6. The computer-implemented method of any one of aspects 1-5, further comprising: obtaining, by the one or more processors, binary alignment map (BAM) data of a plurality of patients; preprocessing, by the one or more processors, the BAM data comprising: converting the BAM data to browser extensible (BED) data using BED conversion parameters, converting the BAM data to FASTQ data using FASTQ conversion parameters, generating one or more tokens representing the BED data and the FASTQ data, generating NPZ data from the tokenized BED data and the tokenized FASTQ data, wherein the NPZ data includes array datastored in a compressed format, and generating H5 data based upon analyzing the NPZ data; and training, by the one or more processors, the genomic transformer using the H5 data.

[0133] Aspect 7. The computer-implemented method of aspect 6, wherein the BED conversion parameters include one or more of a minimum mapping quality, a concise idiosyncratic gapped alignment report (CIGAR) score, a minimum read length, or a minimum coverage.

[0134] Aspect 8. The computer- implemented method of aspect 6 or 7, wherein the FASTQ conversion parameters include one or more of a reference genome, a variant calling method, or a variant filtering criteria.

[0135] Aspect 9. The computer-implemented method of claim 6, wherein training the genomic transformer includes fine-tuning a genomic transformer foundational model trained using human reference genome data of a human reference genome, multispecies genomic data of multiple species, and exomic data of a plurality of humans.

[0136] Aspect 10. The computer-implemented method of any one of aspects 1-9, wherein the genomic transformer includes one or more of: a short context model trained to understand a short neighborhood of exomic information; or a long context model trained to understand exomic information from at least one entire chromosome.

[0137] Aspect 11. The computer-implemented method of any one of aspects 1-10, wherein the genomic transformer includes a classifier head fine-tuned to make predictions associated with the disease.

[0138] Aspect 12. A system for generating a prediction of a response of a patient to a medication to treat a disease, the system comprising: one or more processors; and one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the system to: obtain genomic data, medical image data, and clinical data of the patient, provide the genomic data to a genomic transformer trained to extract exomic data of the patient associated with the disease, provide the exomic data, the medical image data, and the clinical data to a disease prediction model trained to generate the prediction of the response of the patient to the medication, and provide the prediction to a user device.

[0139] Aspect 13. The system of aspect 12, wherein one or more of: the exomic data indicates exomic variations of the patient respective to a human reference genome; the medical image data includes x-rays of one or more of a hand or a foot of the patient; or the clinical data includes clinical notes associated with the disease.

[0140] Aspect 14. The system of aspect 12 or 13, wherein: the disease is rheumatoid arthritis; and the prediction is a response of the patient to Methotrexate.

[0141] Aspect 15. The system of any one of aspects 12-14, further comprising instructions that, when executed by the one or more processors, cause the system to: obtain binary alignment map (BAM) data of a plurality of patients; preprocess the BAM data comprising: converting the BAM data to browser extensible (BED) data using BED conversion parameters, converting the BAM data to FASTQ data using FASTQ conversion parameters, generating one or more tokens representing the BED data and the FASTQ data, wherein the NPZ data includes array data stored in a compressed format, generating NPZ data from the tokenized BED data and the tokenized FASTQ data, and generating H5 data based upon analyzing the NPZ data; and train the genomic transformer using the H5 data.

[0142] Aspect 16. The system of aspect 15, wherein the BED conversion parameters include one or more of a minimum mapping quality, a concise idiosyncratic gapped alignment report (CIGAR) score, a minimum read length, or a minimum coverage.

[0143] Aspect 17. The system of aspect 15 or 16, wherein the FASTQ conversion parameters include one or more of a reference genome, a variant calling method, or a variant filtering criteria.

[0144] Aspect 18. The system of any one of aspects 15-17, wherein training the genomic transformer includes fine-tuning a genomic transformer foundational model trained using human reference genome data of a human reference genome, multispecies genomic data of multiple species, and exomic data of a plurality of humans.

[0145] Aspect 19. The system of any one of aspects 12-18, wherein the genomic transformer includes one or more of: a short context model trained to understand a short neighborhood of exomic information; or a long context model trained to understand exomic information from at least one entire chromosome.

[0146] Aspect 20. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to at least: obtain genomic data, medical image data, and clinical data of a patient; provide the genomic data to a genomic transformer trained to extract exomic data of the patient associated with a disease; provide the exomic data, the medical image data, and the clinical data to a disease prediction model trained to generate a prediction of a response of the patient to a medication; and provide the prediction to a user device.ADDITIONAL CONSIDERATIONS

[0147] Although the text herein sets forth a detailed description of numerous different embodiments, it should be understood that the legal scope of the invention is defined by the words of the claims set forth at the end of this patent. The detailed description is to be construed as exemplary only and does not describe every possible embodiment, as describing every possible embodiment would be impractical, if not impossible. One could implement numerous alternate embodiments, using either current technology or technology developed after the filing date of this patent, which would still fall within the scope of the claims.

[0148] It should also be understood that, unless a term is expressly defined in this patent using the sentence “As used herein, the term ‘ ’ is hereby defined to mean...” or a similar sentence, there is no intent to limit the meaning of that term, either expressly or by implication, beyond its plain or ordinary meaning, and such term should not be interpreted to be limited in scope based upon any statement made in any section of this patent (other than the language of the claims). To the extent that any term recited in the claims at the end of this disclosure is referred to in this disclosure in a manner consistent with a single meaning, that is done for sake of clarity only so as to not confuse the reader, and it is not intended that such claim term be limited, by implication or otherwise, to that single meaning. Finally, unless a claim element is defined by reciting the word “means” and a function without the recital of any structure, it is not intended that the scope of any claim element be interpreted based upon the application of 35 U.S.C. 112(f).

[0149] Throughout this specification, plural instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed concurrently, and nothing requires that the operations be performed in the order illustrated. Structures and functionality presented as separate components in exemplary configurations may be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.

[0150] Additionally, certain embodiments are described herein as including logic or a number of routines, subroutines, applications, or instructions. These may constitute either software (code embodied on a non-transitory, tangible machine-readable medium) or hardware. In hardware, the routines, etc., are tangible units capable of performing certain operations and may be configured or arranged in a certain manner. In exemplary embodiments, one or more computer systems (e.g., a standalone, client or server computer system) or one or more hardware modules of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or application portion) as a hardware module that operates to perform certain operations as described herein.

[0151] In various embodiments, a hardware module may be implemented mechanically or electronically. For example, a hardware module may comprise dedicated circuitry or logic that is permanently configured (e.g., as a special-purpose processor, such as a field programmable gate array (FPGA) or an application- specific integrated circuit (ASIC) to perform certain operations). A hardware module may also comprise programmable logic or circuitry (e.g., as encompassed within a general -purpose processor or other programmable processor) that is temporarily configured by software to perform certain operations. It will be appreciated that the decision to implement a hardware module mechanically, in dedicated and permanently configured circuitry, or in temporarily configured circuitry (e.g., configured by software) may be driven by cost and time considerations.

[0152] Accordingly, the term “hardware module” should be understood to encompass a tangible entity, be that an entity that is physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a certain manner or to perform certain operations described herein. Considering embodiments in which hardware modules are temporarily configured (e.g., programmed), each of the hardware modules need not be configured or instantiated at any one instance in time. For example, where the hardware modules comprise a general-purpose processor configured using software, the general-purpose processor may be configured as respective different hardware modules at different times. Software may accordingly configure a processor, for example, to constitute a particular’ hardware module at one instance of time and to constitute a different hardware module at a different instance of time.

[0153] Hardware modules can provide information to, and receive information from, other hardware modules. Accordingly, the described hardware modules may be regarded as being communicatively coupled. Where multiple of such hardware modules exist contemporaneously, communications may be achieved through signal transmission (e.g., over appropriate circuits and buses) that connect the hardware modules. In embodiments in which multiple hardware modules are configured or instantiated at different times, communications between such hardware modules may be achieved, for example, through the storage and retrieval of information in memory structures to which the multiple hardware modules have access. For example, one hardware module may perform an operation and store the output of that operation in a memory device to which it is communicatively coupled. A further hardware module may then, at a later time, access the memory device to retrieve and process the stored output. Hardware modules may also initiate communications with input or output devices, and can operate on a resource (e.g., a collection of information).

[0154] The various operations of exemplary methods described herein may be performed, at least partially, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented modules that operate toperform one or more operations or functions. The modules referred to herein may, in some exemplary embodiments, comprise processor-implemented modules.

[0155] Similarly, the methods or routines described herein may be at least partially processor- implemented. For example, at least some of the operations of a method may be performed by one or more processors or processor-implemented hardware modules. The performance of certain of the operations may be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In some exemplary embodiments, the processor or processors may be located in a single location (e.g., within a home environment, an office environment or as a server farm), while in other embodiments the processors may be distributed across a number of geographic locations.

[0156] Unless specifically stated otherwise, discussions herein using words such as processing,” “computing,” “calculating,” “determining,” “presenting,” “displaying,” or the like may refer to actions or processes of a machine (e.g., a computer) that manipulates or transforms data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.

[0157] As used herein any reference to “one embodiment” or “an embodiment” means that a particular element, feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment.

[0158] Some embodiments may be described using the expression “coupled” and “connected” along with their derivatives. For example, some embodiments may be described using the term “coupled” to indicate that two or more elements are in direct physical or electrical contact. The term “coupled,” however, may also mean that two or more elements are not in direct contact with each other, but yet still co-operate or interact with each other. The embodiments are not limited in this context.

[0159] As used herein, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having” or any other variation thereof, are intended to cover a non-exclusive inclusion. Forexample, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, “or” refers to an inclusive or and not to an exclusive or. For example, a condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or not present), A is false (or not present) and B is true (or present), and both A and B are true (or present). In addition, use of the “a” or “an” are employed to describe elements and components of the embodiments herein. This is done merely for convenience and to give a general sense of the description. This description, and the claims that follow, should be read to include one or at least one and the singular also includes the plural unless it is obvious that it is meant otherwise.

[0160] Upon reading this disclosure, those of skill in the art will appreciate still additional alternative structural and functional designs for the approaches described herein. Therefore, while particular embodiments and applications have been illustrated and described, it is to be understood that the disclosed embodiments are not limited to the precise construction and components disclosed herein. Various modifications, changes and variations, which will be apparent to those skilled in the art, may be made in the arrangement, operation and details of the method and apparatus disclosed herein without departing from the spirit and scope defined in the appended claims.

[0161] The particular features, structures, or characteristics of any specific embodiment may be combined in any suitable manner and in any suitable combination with one or more other embodiments, including the use of selected features without corresponding use of other features. In addition, many modifications may be made to adapt a particular application, situation or material to the essential scope and spirit of the present invention. It is to be understood that other variations and modifications of the embodiments of the present invention described and illustrated herein are possible in light of the teachings herein and are to be considered part of the spirit and scope of the present invention.

[0162] While the preferred embodiments of the invention have been described, it should be understood that the invention is not so limited and modifications may be made without departing from the invention. The scope of the invention is defined by the appended claims, and alldevices that come within the meaning of the claims, either literally or by equivalence, are intended to be embraced therein. It is therefore intended that the foregoing detailed description be regarded as illustrative rather than limiting, and that it be understood that it is the following claims, including all equivalents, that are intended to define the spirit and scope of this invention.

Claims

WHAT IS CLAIMED:

1. A computer-implemented method for generating a prediction of a response of a patient to a medication to treat a disease, the computer-implemented method comprising: obtaining, by one or more processors, genomic data, medical image data, and clinical data of the patient; providing, by the one or more processors, the genomic data to a genomic transformer trained to extract exomic data of the patient associated with the disease; providing, by the one or more processors, the exomic data, the medical image data, and the clinical data to a disease prediction model trained to generate the prediction of the response of the patient to the medication; and providing, by the one or more processors, the prediction to a user device.

2. The computer- implemented method of claim 1, wherein the genomic transformer includes one or more of: a short context model trained to understand a short neighborhood of exomic information; or a long context model trained to understand exomic information from at least one entire chromosome.

3. The computer-implemented method of claim 1, further comprising: obtaining, by the one or more processors, binary alignment map (BAM) data of a plurality of patients; preprocessing, by the one or more processors, the BAM data comprising: converting the BAM data to browser extensible (BED) data using BED conversion parameters, converting the BAM data to FASTQ data using FASTQ conversion parameters, generating one or more tokens representing the BED data and the FASTQ data,generating NPZ data from the tokenized BED data and the tokenized FASTQ data, wherein the NPZ data includes array data stored in a compressed format, and generating H5 data based upon analyzing the NPZ data; and training, by the one or more processors, the genomic transformer using the H5 data.

4. The computer-implemented method of claim 3, wherein the BED conversion parameters include one or more of a minimum mapping quality, a concise idiosyncratic gapped alignment report (CIGAR) score, a minimum read length, or a minimum coverage.

5. The computer-implemented method of claim 3, wherein the FASTQ conversion parameters include one or more of a reference genome, a variant calling method, or a variant filtering criteria.

6. The computer- implemented method of claim 3, wherein training the genomic transformer includes fine-tuning a genomic transformer foundational model trained using human reference genome data of a human reference genome, multispecies genomic data of multiple species, and exomic data of a plurality of humans.

7. The computer- implemented method of claim 1, wherein the exomic data indicates exomic variations of the patient respective to a human reference genome.

8. The computer- implemented method of claim 1, wherein the medical image data includes x-rays of one or more of a hand or a foot of the patient.

9. The computer-implemented method of claim 1, wherein the clinical data includes clinical notes associated with the disease.

10. The computer-implemented method of claim 1, wherein: the disease is rheumatoid arthritis; andthe prediction is a response of the patient to Methotrexate.

11. The computer- implemented method of claim 1 , wherein the genomic transformer includes a classifier head fine-tuned to make predictions associated with the disease.

12. A system for generating a prediction of a response of a patient to a medication to treat a disease, the system comprising: one or more processors; and one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the system to: obtain genomic data, medical image data, and clinical data of the patient, provide the genomic data to a genomic transformer trained to extract exomic data of the patient associated with the disease, provide the exomic data, the medical image data, and the clinical data to a disease prediction model trained to generate the prediction of the response of the patient to the medication, and provide the prediction to a user device.

13. The system of claim 12, wherein the genomic transformer includes one or more of: a short context model trained to understand a short neighborhood of exomic information; or a long context model trained to understand exomic information from at least one entire chromosome.

14. The system of claim 12, further comprising instructions that, when executed by the one or more processors, cause the system to: obtain binary alignment map (BAM) data of a plurality of patients; preprocess the BAM data comprising:converting the BAM data to browser extensible (BED) data using BED conversion parameters, converting the BAM data to FASTQ data using FASTQ conversion parameters, generating one or more tokens representing the BED data and the FASTQ data, generating NPZ data from the tokenized BED data and the tokenized FASTQ data, wherein the NPZ data includes array data stored in a compressed format, and generating H5 data based upon analyzing the NPZ data; and train the genomic transformer using the H5 data.

15. The system of claim 14, wherein the BED conversion parameters include one or more of a minimum mapping quality, a concise idiosyncratic gapped alignment report (CIGAR) score, a minimum read length, or a minimum coverage.

16. The system of claim 14, wherein the FASTQ conversion parameters include one or more of a reference genome, a variant calling method, or a variant filtering criteria.

17. The system of claim 14, wherein training the genomic transformer includes fine- tuning a genomic transformer foundational model trained using human reference genome data of a human reference genome, multispecies genomic data of multiple species, and exomic data of a plurality of humans.

18. The system of claim 12, wherein one or more of: the exomic data indicates exomic variations of the patient respective to a human reference genome; the medical image data includes x-rays of one or more of a hand or a foot of the patient; or the clinical data includes clinical notes associated with the disease.

19. The system of claim 12, wherein:the disease is rheumatoid arthritis; and the prediction is a response of the patient to Methotrexate.

20. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to at least: obtain genomic data, medical image data, and clinical data of a patient; provide the genomic data to a genomic transformer trained to extract exomic data of the patient associated with a disease; provide the exomic data, the medical image data, and the clinical data to a disease prediction model trained to generate a prediction of a response of the patient to a medication; and provide the prediction to a user device.

Citation Information

Patent Citations

  • Method and process for predicting and analyzing patient cohort response, progression, and survival

    US20200211716A1

  • Processes for Genetic and Clinical Data Evaluation and Classification of Complex Human Traits

    US20210158894A1

  • Multi-omic search engine for integrative analysis of cancer genomic and clinical data

    US20210319907A1

  • Techniques for generating predictive outcomes relating to spinal muscular atrophy using artificial intelligence

    WO2022115356A1

Cited By

  • Disease prediction and auxiliary diagnosis system construction method and system based on multi-modal large model

    CN120565123A