Generalizable transformer for disease state classification from liquid biopsies
The transformer model with a genome-wide positional encoder and inverse attention head enhances disease state classification in liquid biopsies by addressing variability and adaptability issues in traditional models, enabling detection of diverse diseases with enriched learning and operational flexibility.
Patent Information
- Application Number
- PCT/US2025/037778
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-16
- Filing Date
- 2025-07-15
- Publication Date
- 2026-01-22
AI Technical Summary
Traditional machine learning models for liquid biopsies are limited in variability of detected patterns, lack flexibility across diseases, and struggle with binary classification in clinical settings due to reliance on prevalence-based and population-based approaches, often overlooking unique molecular observations and requiring large, curated datasets.
A transformer model equipped with a genome-wide positional encoder and inverse attention head for molecular-level pattern anomaly detection, trained with general and targeted data to enhance adaptability and flexibility across various diseases.
The transformer model improves disease state classification by enriching learning experiences, detecting a multitude of diseases beyond initial training, and overcoming concentration and population-based limitations, providing operational flexibility and adaptability.
Smart Images

Figure US2025037778_22012026_PF_FP_ABST
Abstract
Description
GENERALIZABLE TRANSFORMER FOR DISEASE STATE CLASSIFICATIONFROM LIQUID BIOPSIESInventors: Soheil Damangir Hamed AminiCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] The present application claims the benefit of and priority to U.S. Provisional Application No. 63 / 672,128 filed on July 16, 2024, which is incorporated by reference in its entirety.BACKGROUND
[0002] Traditional machine learning methodologies employed in liquid biopsy applications have primarily focused on the identification of specific molecular patterns. These conventional models examine each sample, producing a score that is dependent upon the presence and prevalence of such patterns. There are two key downsides to this traditional methodology.
[0003] First, current models engaged in a uniform approach to decision-making, identifying the presence or absence of a particular pattern irrespective of the individual sample, are limited in the variability of patterns detected. They are trained to discern patterns that are abnormal with respect to the overall population, but often fail to identify unusual patterns unique to a singular sample. Relatedly, these models form deductions on the prevalence of a pattern from a set of molecules, rather than on each individually observed molecule. As a consequence, unique molecular observations that could prove significant or vital for signal detection in applications including disease diagnosis, prognostication, monitoring, treatment selection, or treatment response assessment are commonly overlooked as they lack the significant prevalence in the gathered sample.
[0004] Second, these models are trained using a limited set of samples, specifically curated for a particular disease. Procuring the data set for robust training of the models typically requires enrollment of large studies with numerous participants. With some rare diseases, such numerosity is unattainable. Moreover, each model is particularly trained on a separate training data set. The patterns identified to be correlative to one disease differ from patterns correlative to another disease. As such, the models lack flexibility and the ability toadapt to other diseases. In other words, they significantly lack the capacity to “transfer knowledge” from one disease condition to another. This inability severely restricts their functionality across a broader spectrum and curtails their application to a single specific disease.
[0005] A further limitation of enrolling a large number of participants for development of a molecular diagnostic test is that such studies often include participants with disease and healthy participants, so that the model learns to distinguish between disease and healthy in a controlled way. However, such binary classification often performs worse when used in the clinical setting, as subjects may have other diseases, comorbidities or indications. So, the task is not so simply distinguishing the disease of interest from healthy, it is rather distinguishing the disease of interest from healthy and other disease indications.SUMMARY
[0006] Recognizing the limitations inherent in the traditional models used for liquid biopsies, the present disclosure provides systems and methods that use a transformer model equipped with a genome-wide positional encoder and an inverse attention head to provide improved molecular-level pattern anomaly detection and operational flexibility, to solve the problems of the traditional models.
[0007] The transformer model trains an attention subspace on general data drawn from a broad spectrum of diseases. This enriches the system’s learning experience and expands its applicability to detect a multitude of diseases beyond those it was initially trained on, serving to redress the lack of “knowledge transfer” pointed out in traditional models. The introduction of an inverse attention head in the transformer architecture guides the transformer model in ignoring molecular data that is normal, i.e., insignificant in detecting disease signals, as a function of all other molecules present in the sample thereby overcoming the concentration prevalence-based limitation and population-based limitation of prior models. In further embodiments, the transformer model implements genome-wide positional encoding on a per molecule basis. This differs from positional encoding in other transformer models in that traditionally the positional encoding provides positional context between input tokens to model the order of the input tokens, whereas the genome-wide positional encoding provides global genome positioning context, independent of relational context between input tokens.
[0008] In summary, the generalizable transformer model presents technical improvements to disease state classification from liquid biopsy. The dual-process training ofthe attention subspace with general training data and the transformer model with targeted training data provides flexibility and adaptability. With the generalized attention subspace, the transformer model can leverage understanding across various diseases, while training the inverse attention head to inversely attend to (i.e., ignore) other disease types and / or subtypes with limited targeted training data.
[0009] The training of the machine-learned models described herein (such as the transformer model or any other model referenced herein) include the performance of one or more non-mathematical operations or implementation of non-mathematical functions at least in part by a machine or computing system, examples of which include but are not limited to data loading operations, data storage operations, data toggling or modification operations, non-transitory computer-readable storage medium modification operations, metadata removal or data cleansing operations, data compression operations, protein structure modification operations, image modification operations, noise application operations, noise removal operations, and the like. Accordingly, the training of the machine-learned models described herein may be based on or may involve mathematical concepts, but is not simply limited to the performance of a mathematical calculation, a mathematical operation, or an act of calculating a variable or number using mathematical methods.
[0010] Likewise, it should be noted that the training of these models described herein cannot be practically performed in the human mind. The models are innately complex including vast amounts of weights and parameters associated through one or more complex functions. Training and / or deployment of such models involve so great a number of operations that it is not feasibly performable by the human mind alone, nor with the assistance of pen and paper. In such embodiments, the operations may number in the hundreds, thousands, tens of thousands, hundreds of thousands, millions, billions, or trillions. Moreover, the training data may include hundreds, thousands, tens of thousands, hundreds of thousands, millions, or billions of sequence reads, each sequence read may further include anywhere from hundreds, to thousands, or to tens of thousands of molecular units. Accordingly, such models are necessarily rooted in computer-technology for their implementation and use.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] FIG. 1 is an example networking environment for an analytics system, according to one or more embodiments.
[0012] FIG. 2 is an example sequencing device for performing molecular sequencing, according to one or more embodiments.
[0013] FIG. 3 is a block diagram of the analytics system, according to one or more embodiments.
[0014] FIG. 4 is an example architecture of the transformer model for disease classification, according to one or more embodiments.
[0015] FIG. 5A is a flowchart of the process of training an inverse attention head in the transformer model, according to one or more embodiments.
[0016] FIG. 5B is a flowchart of the process of training a feed-forward network of the transformer model, according to one or more embodiments.
[0017] FIG. 6 is a flowchart of the process of deploying the transformer model to infer a disease classification for a test subject, according to one or more embodiments.DETAILED DESCRIPTIONOVERVIEW OF ADVANTAGES OF LIQUID BIOPSY FOR DISEASE PREDICTION
[0018] Traditional disease screening include blood tests, cancer screenings, infectious disease panels, imaging techniques, or some combination thereof. Ordinarily, the blood tests would identify macromolecular characteristics, present in the subject’s blood, providing some limited insight into the subject’s health state. Cancer screenings would ordinarily include blood tests, whole-body imaging, pap smears, collection and visual analysis of tumor biopsies, or some combination thereof. Infectious disease panels would include assays that target specific nucleic acid markers for the various infectious diseases, e.g., influenza, human immunodeficiency virus, hepatitis, etc.
[0019] These traditional screening paradigms had their own challenges. For one, blood tests to identify macromolecular characteristics provided a limited snapshot of the subject’s health state, being imprecise to any one particular disease. For example, abnormally high cholesterol measurements could be correlated with any number of diseases or maladies, e.g., atherosclerosis, coronary artery disease, propensity for a stroke, peripheral arterial disease, hypertension, erectile dysfunction, diabetes, obesity, chronic kidney disease, etc. For two, traditional biomarkers were highly specific to each disease or sub type. For example, pap smears are only useful for screening cervical cancer. Or, as another example, a mammography is only useful for screening breast cancer. So on and so forth with other cancer-specific or organ specific biomarkers. Moreover, these traditional screeningparadigms relied heavily on a clinician’s subjective insight in diagnosing the subject’s condition. For example, diagnosing glioma requires a neurosurgeon to visually interpret magnetic resonance imaging (MRI) images to identify and qualitatively assess the tumor. For three, infectious disease panels could hone in on specific markers, but could be limited by that specificity. For example, viruses tend to mutate, evolving to different variants with different viral strains. This ever-moving target can be hard to discover and to hit.
[0020] Moreover, following a disease diagnosis, monitoring of the disease progression can also be a challenge. Once a disease is diagnosed, there can be targeted tests for disease monitoring. These targeted tests may continue to rely on human subjective qualitative assessment. For example, to monitor recurrence of breast cancer, a treated subject may continue to undergo periodic mammograms, with the clinician making subjective calls on the recurrence. This same challenge applies to evaluating treatment efficacy. Ordinarily, treatment efficacy is tied to the traditional biomarkers and / or monitoring process. If there is recurrence of the disease, then the treatment would be deemed ineffective or at the very least less-than-wholly effective.
[0021] Liquid biopsies are a burgeoning technology that can resolve a lot of the challenges with traditional screening paradigms. However, as described more fully above, liquid biopsies still face challenges with curating a robust training data set. This issue is especially pertinent with diseases where samples are few or hard to obtain. For example, pancreatic cancer has symptoms that can often be mistaken for digestive issues (e.g., abdominal pain, weight loss, changes in stool, etc.). The location of the pancreas also makes the organ hard to access for routine biopsies or other testing. More often than not, pancreatic cancer is detected in late stages, where prognosis is bleak. Late detection creates a slim timeline for sample collection and processing for use in training diagnostic models for detection of such diseases.
[0022] The present disclosure provides systems and methods that deploy a transformer model equipped with a genome-wide positional encoder and an inverse attention head to provide improved molecular-level pattern anomaly detection and operational flexibility, to solve the problems of the traditional models. This machine-learning architecture amounts to an improvement to the field of machine learning, and more particularly to the field of disease detection from liquid biopsies.
[0023] The transformer model may be applied to detect a variety of diseases correlated with genomic or epigenomic alteration. For example, the transformer model may be applied to detect various cancers or chronic diseases. For example breast cancer, lung cancer, livercancer, colon cancer, and leukemia are often driven by genomic and epigenomic changes. Many chronic diseases such as metabolic dysfunction-associated steatotic liver disease (MASLD), metabolic dysfunction-associated steatohepatitis (MASH), liver cirrhosis, fibrosis, type 2 diabetes mellitus, obesity as also characterized by epigenetic changes. Other diseases may have undiscovered genomic or epigenomic alteration that can be, nonetheless, detected with the transformer model. The transformer model may be applied to detect inherited genetic disorders like cystic fibrosis (caused by mutations in the CFTR gene), sickle cell anemia (HBB gene), and Huntington’s disease (HTT gene). The transformer model may be applied to detect diseases correlated with epigenomic aberration, such as abnormal DNA methylation or histone modification patterns, implicated in diseases like cancer (e.g., where tumor suppressor genes may be silenced), Fragile X syndrome (due to methylation of the FMRI gene), and imprinting disorders like Prader-Willi and Angelman syndromes. Additionally, the transformer model may be applied to detect neurodegenerative diseases such as Alzheimer’s and Parkinson’s, now associated with both genetic mutations and epigenetic changes that affect gene expression and neuronal function.SYSTEM ENVIRON ENT
[0024] FIG. 1 illustrates an example networking environment 100 for an analytics system 110, according to one or more embodiments. The analytics system 110 applies a transformer model to predict a disease classification for a biological sample obtained from a subject and comprising genetic and / or epigenetic information on the subject. The system environment 100 illustrated in FIG. 1 includes the analytics system 110, a client device 120, a sequencing device 130, a third-party system 140, and a network 130. Alternative embodiments may include additional, fewer, or different components than those illustrated in FIG. 1. In other embodiments, functionality of the components may be distributed differently than as described below. Additionally, each component may perform their respective functionalities in response to a request from a human, or automatically without human intervention.
[0025] The analytics system 110 leverages the transformer model to predict disease classification for biological samples. The analytics system 110 may be a general computing device. The analytics system 110 may perform one or more computational analyses on molecular data, e.g., sequencing data derived from sequencing genetic molecules from a biological sample. In one or more embodiments, the analytics system 110 performs analyses include preprocessing of the sequencing data, training of the transformer model, deployingthe transformer model, or some combination thereof. Based on outputs of the transformer model, the analytics system 110 may generate and transmit notifications including the model outputs to the client device 120, e.g., for presentation to a user. A healthcare provider may, for example, base diagnosis, treatment determination, prognosis, or other healthcare-related steps based on the model outputs in the notification. In training the transformer model, the analytics system 110 may retrieve training data from the third-party system 140.
[0026] The client device 120 is a general computing device used by a user. The user may be a patient, a healthcare provider, a researcher, etc. For example, a patient can use a client device 120 to communicate with their healthcare provider. Alternatively, the patient can use the client device 120 to report personal information, e.g., that may be stored and used by the analytics system 110 in leveraging the transformer model. In the context of a healthcare provider, the client device 120 may be used in communication with the analytics system 110 for receiving and / or viewing model outputs or other predictions by the analytics system 110. A researcher may use the client device 120 to perform the one or more analyses by the analytics system 110 and / or guide training of the transformer model.
[0027] The sequencing device 130 performs sequencing assays on biological samples to sequence molecules present in the biological samples. In one or more embodiments, the sequencing device 130 is an instrument used to determine the precise order of molecules in a molecule polymer. In the realm of genomics, the sequencing device 130 may perform DNA or RNA sequencing. In such context, the sequencing device 130 may determine the order of nucleotide bases — adenine, guanine, cytosine, and thymine (or uracil in the case of RNA) — in a given DNA or RNA molecule. The sequencing device 130 may employ the Sanger method, next-generation sequencing (NGS) techniques, or other sequencing techniques. The sequencing device 130 may also include an electronic display for real-time data visualization and for user interaction with analysis tools. The output of the sequencing device 130 is sequencing data composed of sequence reads. Each sequence read can correspond to a molecule in the biological sample. An example sequencing device 130 is illustrated in FIG. 2 and described below.
[0028] The third-party system 140 stores reference data, e.g., that may be leveraged in training of the transformer model. The reference data may include sequencing data from a plurality of subjects, each associated with different diseases and / or disease states. For example, the sequencing data may be divided into cohorts for different disease types. One cohort of subjects (with their corresponding sequencing data) may be diagnosed with one or more cognitive disorders, whereas another cohort of subjects (with their correspondingsequencing data) may be diagnosed with one or more heart diseases. Or, in an example of subdivision by disease states, one cohort of subjects (with their corresponding sequencing data) may be diagnosed with one type of cancer (e.g., of a particular tissue or cell of origin, of a particular cancer biology, of a particular stage, or some combination thereof), whereas another cohort may be diagnosed with another type of cancer. Other examples may include other types of diseases, e.g., chronic diseases, metabolic diseases, neurodegenerative diseases, etc. The sequencing data stored by the third-party system 140 may include aggregate sequencing data or molecule-level sequencing data. Aggregate sequencing data, e.g., for one subject or for one sample from a subject, aggregates molecular sequencing data. This may take the form of averaging, combining, or otherwise identifying key signals in the molecule-level sequencing data. Molecule-level sequencing data, e.g., for one subject or for one sample from a subject, includes most or all sequencing information on molecules collected from the subject.
[0029] The network 150 facilitates communication(s) between the components of the networking environment 100. The network 130 is a collection of computing devices that communicate via wired or wireless connections. The network 130 may include one or more local area networks (LANs) or one or more wide area networks (WANs). The network 130, as referred to herein, is an inclusive term that may refer to any or all of standard layers used to describe a physical or virtual network, such as the physical layer, the data link layer, the network layer, the transport layer, the session layer, the presentation layer, and the application layer. The network 130 may include physical media for communicating data from one computing device to another computing device, such as multiprotocol label switching (MPLS) lines, fiber optic cables, cellular connections (e.g., 3G, 4G, or 5G spectra), or satellites. The network 130 also may use networking protocols, such as TCP / IP, HTTP, SSH, SMS, or FTP, to transmit data between computing devices. In some embodiments, the network 130 may include Bluetooth or near-field communication (NFC) technologies or protocols for local communications between computing devices. The network 130 may transmit encrypted or unencrypted data.EXAMPLE SEQUENCING DEVICE
[0030] FIG. 2 illustrates an example sequencing device 220 for performing molecular sequencing, according to one or more embodiments. The sequencing device 220 is an embodiment of the sequencing device 130. The sequencing device 220 performs one or more chemical assays to a sample 210 to output sequencing data 250. The sequencing data 250may include sequence reads that describe molecules in the sample 210. For example, in genomic sequencing, each sequence read may indicate nucleotides and their respective order in a nucleic acid molecule, or the state of any modifications on the nucleotides such as the methylation state of cytosines.
[0031] The sequencing device 220 may include one or more internal components for performing the chemical assays. Such components may include a sample preparation module, a sequencing chip or plate, a detection module, and a reagent delivery module. The sequencing device 220 may further include external-facing components such as a display interface 230 and a loading tray 240. Alternative embodiments may include additional, fewer, or different components than those illustrated in FIG. 1. In other embodiments, functionality of the components may be distributed differently than as described below. Additionally, each component may perform their respective functionalities in response to a request from a human, or automatically without human intervention.
[0032] The sample preparation module prepares and otherwise processes the sample 210 ahead of sequencing. For example, the sample preparation module may perform molecule extraction, purification, concentration, or some combination thereof. After extraction, a library of molecules is constructed. Library construction involves processes like fragmentation of molecules into smaller pieces, followed by the attachment of platformspecific adapter molecules to the molecules in the sample 210, which aids in sequencing. This process can also involve molecule amplification (e.g., amplification of DNA samples through polymerase chain reaction (PCR)) to increase chances that the sequencing process captures and detects each unique molecule originally in the sample 210.
[0033] The sequencing chip or plate serves as a reaction chamber where the chemical assaying processes occur. Depending on the type of device, the sequencing chip or plate can take many forms, from a high-density array of wells to a flow cell. The sequencing chip or plate may connect to the loading tray 240, where the sample 210 is deposited into the sequencing device 220. With the sample 210 loaded, the sequencing device 220 may perform the one or more chemical assaying steps on the sequencing chip or plate. For example, in genomic sequencing, the sequencing chip or plate is where nucleic acid molecules are grown on a lawn of primers and where fluorescent nucleotides are ligated onto attached DNA molecules to sequence the DNA molecules.
[0034] The detection module may include one or more detectors for sensing one or more characteristics indicative of different molecular units in the polymer. For example, the detection module may include an optical sensor for detecting wavelength of emitted light andan intensity thereof. In other examples, the detection module may include an electrical sensor for sensing changes in electrical characteristics during interaction of one or more electrical components with the molecules being sequenced. In yet other examples, the detection module may include a pH sensor for sensing changes in pH affected by ligated molecular units (e.g., nucleotides) onto a target molecule. In one particular example of genomic sequencing, each fluorescent nucleotide may be ligated onto a complementary nucleic acid molecule, with the fluorophore excited with an energy pulse (e.g., from a laser). The optical sensor may detect the emitted wavelength of light, which may indicate the particular nucleotide ligated onto the complementary nucleic acid molecule. The detection module may provide the captured data to a computer system to decode into the sequencing data.
[0035] The reagent delivery module may manage and direct reagents to the sequencing chip or array. The reagent delivery module may include one or more vessels for storing reagents, one or more channels for directing reagents from the vessels to the sequencing chip or array, one or more gates for controlling the flow of reagents, or some combination thereof. Example reagents include buffers, primers, probes, enzymes, molecular units (e.g., nucleotides), other chemicals for use in the chemical assaying processes, etc.
[0036] The display interface 230 provides a user interface for operation by a user. The display interface 230 may include an electronic display and may further include a touchscreen, e.g., for receiving user touch input. Other input devices may also be implemented. The electronic display may provide real-time progress on a sequencing assay. The user may provide inputs to guide or define the sequencing. In other embodiments, the sequencing device 220 is connected to another computing system, which may provide the instructions.
[0037] The loading tray 240 is a receptacle for receiving the sample 210. In some embodiments, the loading tray 240 serves as the space where prepared samples are positioned or loaded before the sequencing process begins. The configuration and capacity of these trays can vary significantly depending on the specific sequencing system. The loading tray is designed such that reagents can be efficiently and automatically distributed to the samples during the sequencing process. It may also be structured to facilitate the required thermal and fluidic manipulations needed throughout the process.
[0038] In one or more example implementations, peripheral blood is collected in K2- EDTA or specialized cfDNA-stabilizing blood collection tubes (e.g., Streck tubes). Following collection, a centrifuge is used to spin blood down to pellet cells. The pelleted cells are transferred to a fresh tube. Isolation and separation of pelleted cells maximizes thecfDNA signal present in the supernatant fluid, while removing genomic DNA (gDNA) from lysed blood cells. In some instances, the genomic DNA can be treated as noise to identify disease signals in the cfDNA. Spinning down the cells and plasma collection from the supernatant can be performed multiple times, to further minimize gDNA confounding signals. Cell-free DNA (cfDNA) can be extracted from the plasma specimens using silica- based or magnetic bead-based protocols optimized for low-input, fragmented DNA. Extracted cfDNA is quantified using fluorometric, gel-base, digital PCR methods, or some combination thereof to ensure sufficient yield for downstream library preparation. In some embodiments, downstream processing of the sample can be tailored to account for the quantified nucleic acid load.
[0039] Following extraction, libraries are prepared using assay protocols that empower reading the methylation state of cytosines in the CpG context. These methods include chemical processes such as bisulfite conversion of cytosines in a CpG (e.g. bisulfite conversion of unmethylated cytosine to uracil), or enzymatic conversion of cytosine in a CpG as a function of its methylation state. The protocol could involve conversion of cytosines to uracils (e.g. conversion of unmethylated Cs to preserve methylated cytosines in WGBS), endrepair and adapter ligation. Libraries are PCR-amplified, purified, and quality-controlled before being sequenced on next-generation sequencing platforms such as an Illumina high- throughput sequencing platform (e.g., NovaSeq), generating paired-end reads typically covering the entire methylome at 5X to 200X depth depending on the application.
[0040] In one or more example implementations, the sequencer writes base calls into raw Binary Base Call (BCL) files. These raw files can be converted to demultiplexed FASTQ files, which may then be securely transferred to an on-premises or cloud-compute cluster.The analytics system may perform preprocessing pipelines including: (1) Quality trimming & adapter removal to strip away low-quality bases and any residual adapter sequences. (2) Cleaned reads are mapped to the human reference genome (e.g. GRCh38) using an optimized short read aligner, producing coordinate-sorted BAM files which among others contains the genomic coordinate of each molecule. (3) Reads that are PCR and optical duplicates of the same original DNA molecule are then marked or removed. (4) For each remaining aligned read (molecule), the state of its CpG methylations are calculated. These per-molecule data points are assembled into a numerical matrix. This matrix is the direct input to the downstream analysis, e.g., input into a machine-learning model for disease prediction.
[0041] This physical-to-digital pipeline transforms a real-world biospecimen into a comprehensive, single-base-resolution dataset reflecting the methylation landscape of thecfDNA, which serves as the input for the downstream machine-learning models described herein. The upstream assay establishes a robust and scalable foundation for data-driven molecular inference.
[0042] The output and results generated by the Al model is then assembled into a report that is used by health care professionals or researchers to detect, monitor, and direct effective treatment for the disease.
[0043] Analysis of sequencing data with machine-learning is made possible by the computing power of computing systems. Without the power of computing systems, applying machine-learning models is an infeasible task, particularly for the human mind. For example, a neural network may be architected with dozens of layers, each layer sized with 24 or more nodes, and each node in one layer is weighted to each node in the next layer. In one example, a small transformer-based model can have upwards of 10 million parameters, while some larger models can be on the order of hundreds of billions of parameters. The sheer computational scale of these models would be impossible for a human to undertake. Even assuming, for the sake of argument, that a human could iteratively perform simple computations needed to process one pass of a machine-learning model, the speed at which a human could manually perform one inference task would span the order of years, if not decades. This would render such inference task not useful or obsolete, especially in the context of disease detection and / or monitoring.ANALYTICS SYSTEM ARCHITECTURE
[0044] FIG. 3 is a block diagram of the analytics system 110, according to one or more embodiments. The analytics system 110 receives sequencing data and leverages a transformer model 320 to perform predictive analyses on the sequencing data. The analytic system 110 may be a general computing system, including one or more computer processors and one or more computer-readable storage media with computer-readable instructions (or code). With the trained transformer model 320, the analytics system 110 may provide disease state classification to test samples obtained from a patient and sequenced by a sequencing device.
[0045] The analytic system 110 may include, among other components, a sequence processor 310, a sequencing database 315, a transformer model 320, a model database 325, and a machine-learning engine 330. Alternative embodiments may include additional, fewer, or different components than those illustrated in FIG. 1. In other embodiments, functionality of the components may be distributed differently than as described below. Additionally, eachcomponent may perform their respective functionalities in response to a request from a human, or automatically without human intervention.
[0046] The sequence processor 310 processes sequencing data received from one or more sequencing devices (e.g., the sequencing device 130, or the sequencing device 220). The sequencing data may be in the sequencing device raw file format (e.g., BCL or FAST5), FASTQ file format or a BAM file format. The sequence processor 310 may perform one or more preprocessing steps with the sequencing data. For example, the sequence processor 310 may assign sequence reads to various samples that were multiplex sequenced. In multiplex sequencing, the sequencing device may sequence multiple samples in one sequencing assay. To distinguish sequence reads as belonging to each sample, the sequencing device may use different cells, different indices, or some combination thereof. The sequence processor 310 may utilize unique molecular tags to assign the sequencing data to each of the samples in a multiplex sequencing assay. With a sample’s sequencing data, the sequence processor 310 may perform contamination detection to assess a contamination level. The sequence processor 310 may also bag sequence reads to identify distinct original molecules in the sample. For example, two sequence reads may pertain to amplicons of the same original molecule. The sequence processor 310 can stitch bagged sequence reads into a consensus read. The sequence processor 310 stores the sequencing data in the sequencing database 315.
[0047] In some embodiments, the sequence reads may be aligned to a reference genome (e.g., a human reference genome) using various techniques to determine alignment position information. Alignment position may generally describe a beginning position and an end position of a region in the reference genome that corresponds to a beginning nucleotide base and an end nucleotide base of a given sequence read. Corresponding to methylation sequencing, the alignment position information may be generalized to indicate a first CpG site and a last CpG site included in the sequence read according to the alignment to the reference genome. The alignment position information may further indicate methylation statuses and locations of all CpG sites in a given sequence read. A region in the reference genome may be associated with a gene or a segment of a gene; as such, the sequence processor 310 may label a sequence read with one or more genes that align to the sequence read. In one embodiment, fragment length (or size) is determined from the beginning and end positions.
[0048] In various embodiments, for example when a paired-end sequencing process is used, a sequence read is comprised of a read pair denoted as R_1 and R_2. For example, the first read R_1 may be sequenced from the first end of a double-stranded DNA (dsDNA)molecule whereas the second read R_2 may be sequenced from the second end of the doublestranded DNA (dsDNA). Therefore, nucleotide base pairs of the first read R_1 and second read R_2 may be aligned consistently (e.g., in opposite orientations) with nucleotide bases of the reference genome. Alignment position information derived from the read pair R_1 and R_2 may include a beginning position in the reference genome that corresponds to an end of a first read (e.g., R_l) and an end position in the reference genome that corresponds to an end of a second read (e.g., R_2). In other words, the beginning position and end position in the reference genome can represent the likely location within the reference genome to which the nucleic acid fragment corresponds. An output file having SAM (sequence alignment map) format or BAM (binary) format may be generated and output for further analysis.
[0049] The transformer model 320 is a machine-learning model that inputs sequencing data and outputs a prediction on disease state classification. The transformer model 320 may be, as the name suggests, a machine-learning transformer model. In general, a machinelearning transformer model handles variable-sized input using a self-attention mechanism and captures contextual relationships between tokens in the input. In one or more embodiments, the transformer model 320 implements a genome-wide positional encoder, an inverse attention head trained on an attention subspace, or some combination thereof. The attention subspace may be learned from general training data across a plurality of different geneticbased diseases, targeted training data relating to a target disease of interest, or some combination thereof. The attention subspace may be learned from aggregate level data or molecular level data. In one or more examples, each of the general training data and the targeted training data may be aggregate level data, molecular level data, or some combination thereof. Aggregate level sequencing data includes data that is an aggregate representation of a plurality of sequence reads from a sample. For example, the aggregate level sequencing data may include methylation rates across a plurality of methylation sites in a reference genome. In another example, the aggregate level sequencing data may specify frequency of particular genetic mutations identified from the sequencing data, or other aggregate characteristics derivable from the sequencing data such as epigenetic patterns. Molecular level sequencing data includes data on individual sequence reads. For example, molecular level sequencing data may include the nucleotide sequences of individual nucleic acid molecules present in the sample. The genome-wide positional encoder encodes positional information of the alignment of a sequence read into the data. The inverse attention head generates a set of weights that focus and / or defocus the transformer model 320 on particular sequence reads in a sample. The transformer model 320 also includes an embedding moduleand a feed-forward network. The transformer model 320’ s architecture is further described inFIG. 4.
[0050] The machine-learning training engine 330 performs the training of the transformer model 320 and any other models used by the analytics system 110. The learned parameters and functions of each model may be stored in the model database 325. To train the models, the machine-learning training engine 330 obtains training data. The training data may be labeled (for supervised learning) or unlabeled (for unsupervised learning). Example labels for disease states may include type of disease, subtypes of the disease, progression of disease, aggression of disease, etc. Each sample in the training data may further include one or more covariate characteristics, e.g., age, biological sex, ethnicity, smoking status, obesity, other health characteristics, other diagnosed disease or medical disorder, etc. The training data may include general training data (e.g., obtained from third-party databases, and may be aggregate level or molecular level sequencing data) and targeted training data (e.g., obtained from sequencing individuals with a particular disease type, and may be molecular-level sequencing data).
[0051] In the training process, the machine-learning training engine 330 may use the training data to determine one or more statistical distributions for different subsets of training samples. The statistical distributions may be used to infer likelihood of observing one or more sequence reads, e.g., in a test sample. The machine-learning training engine 330 may leverage the statistical distributions with optimization algorithms to determine the most probable explanation for an observed set of sequence read(s).
[0052] The machine-learning training engine 330 may also use an objective function (or conversely a loss function) that scores predictability of a model, e.g., based on labels for the training data. In such a training manner, the machine-learning training engine 330 inputs training data into a model to predict a label, which may be scored against a known label. The machine-learning training engine 330 adjusts parameters of the model to optimize the score (i.e., to maximize the objective, or conversely to minimize the loss). Training may be performed with a number of batches of samples and / or over a number of epochs. Training may leverage hard-negative mining, reinforcement learning, other refinement techniques, or other validating techniques. Example types of machine-learning models include: linear regression, logistic regression, decision trees, support vector machine, neural networks, transformers, etc.
[0053] In training the transformer model 320, the machine-learning training engine 330 may perform a dual-training with general training data and disease-specific training data (alsoreferred to as “targeted training data”). The general training data may be obtained from third- party systems and / or sequencing devices. The general training data includes data across a plurality of different genetic-based diseases. The data may be in the form of genetic sequencing data, e.g., whole genome sequencing data, small variants (e.g., insertions, deletions, short tandem repeats, single nucleotide polymorphisms), copy number variants, methylation variants, histone modifications, non-coding ribonucleic acids (RNAs), other epigenetic variants, etc. The disease-specific training data includes data from subjects diagnosed with a common disease (or disease subtype). With the general training data, the machine-learning training engine 330 trains or learns an attention subspace as a latent representation of the sequencing data. The machine-learning training engine 330 may further train or tune the attention subspace with the targeted training data. The attention subspace can thus be trained or learned with either aggregate sequencing data, molecule-level sequencing data, or some combination thereof. With the learned attention subspace, the machine-learning training engine 330 trains the transformer model 320 with the target training data. For example, the machine-learning training engine 330 trains the inverse attention head to defocus the transformer model 320 from normal sequence reads, i.e., sequence reads that are insignificant in detecting the target disease. The machine-learning training engine 330 may also train the other components of the transformer model 320 (e.g., the feed-forward network) to predict disease state for the target disease.
[0054] In some embodiments, the dual-process training may be decoupled, i.e., the machine-learning training engine 330 may independently train the attention subspace from the other components of the transformer model 320. In other embodiments, the dual-process training is performed concurrently, i.e., the machine-learning training engine 330 trains the attention subspace with the general training data (scored with one loss / objective function) and the transformer model 320 (at large) with the disease-specific training data (scored with another loss / objective function). In the concurrent dual-process training, the machinelearning training engine 330 may optimize the aggregate score of both functions.TRANSFORMER MODEL ARCHITECTURE
[0055] FIG. 4 is an example architecture of the transformer model 320 for disease classification, according to one or more embodiments. The transformer model 320 includes an embedding module 410, a positional encoder 420, an inverse attention head 430, and a feed-forward network 450. The transformer model 320 may also include an attentionsubspace 465. In other embodiments, the transformer model 320 may include additional, fewer, or different components than those listed herein.
[0056] The embedding module 410 maps sequences reads 405 into an embedding space that represents all possible values of the sequence reads. The embedding module 410 inputs a sequence read 405 and outputs a corresponding molecule embedding 415. All molecule embeddings may include coordinates in the embedding space (i.e., all have the same dimensionality). The embedding space is a latent space. In one or more embodiments, sequence reads corresponding to molecules with similar molecular units are mapped into proximate positions in the embedding space. In one or more embodiments, the embedding module 410 is independently trained prior to training of other components of the transformer model 320.
[0057] The positional encoder 420 inputs positional information for the sequence reads 405 and outputs genome-wide positional encodings 425. The genome-wide positional encodings 425 encode the genome-wide location information, such that molecules in proximity to one another in the human genome may be proximately mapped to genome-wide positional encodings.
[0058] The transformer model 320 may implement a multiplicative encoding at the positional encoder 420. In a multiplicative manner, the positional encoder 420 multiplies, for each sequence read, the corresponding molecule embedding 515 and the genome-wide positional encoding 525. In one or more embodiments, multiplicative encoding conforms to dimensionality of molecular level data and aggregate data. This equips training of the transformer model 320 with both the molecular level training data and the aggregate training data. This is an advantage over additive encoding which would exclude use of aggregate training data without precise genomic location information (additive encoding would be usable for molecular level training data having precise genomic location information).
[0059] The inverse attention head 430 inputs the multiplied product of the molecule embeddings 415 and the genome-wide positional encodings 425 into the inverse attention head 430 to output attention weights 435. The attention subspace 465 is trained on general training data 460 (optionally in combination with target training data). The inverse attention head 430 operates in the attention subspace 465 to defocus from normal molecules to better leverage the information learned by the attention subspace 465 about general biology. The inverse attention head 430 is trained to output a weight that represents what sequence reads and / or what training samples to not attend to, i.e., to defocus. This is particularly advantageous as training a model of general biology of healthy or normal from eitheraggregate or molecular level data is easier and thus it is easier to learn to defocus normal or healthy signals rather than positively attending to discriminatory signals. This concept may also be termed negative attention. The contrary approach of positive attention would require large quantities of molecule level data from the target disease to provide sufficient signal for the transformer model 320 to attend to the disease signal. This, however, may be infeasible in situations where target disease training data is limited or hard to obtain.
[0060] The transformer model 320 combines the molecule embeddings 415, the genomewide positional encodings 425, and the attention weights 435 into a sample matrix 440. The sample matrix 440 is a compact mapping of a sample’s sequencing data with positional encoding and inverse attention.
[0061] The feed-forward network 450 inputs the sample matrix 440 and outputs the prediction 455. The prediction 455 may be binary or multiclass. The binary output may indicate presence or absence of the target disease, or likelihood of presence / absence of the target disease. The multiclass output may indicate one of a plurality of disease types and / or subtypes.
[0062] To train the transformer model 320, the training engine may leverage two different score functions. The first score function computes a first score for learning the attention subspace 465. The training engine optimizes this score function when crafting the attention subspace 465 to represent learned relations of sequencing data in the general training data 460 (optionally in combination with the targeted training data). The second score function computes a second score for the disease-specific training data (not illustrated) representing the learned capabilities of the transformer model 320 in classifying the target disease (e.g., components disposed along the vertical axis such as the inverse attention head 430, the feed-forward network 450, etc.). The two scores may be disparately weighted and aggregated. During training, the training engine may adjust parameters to optimize the score.
[0063] In one example implementation, the transformation is defined as follows. For a sample including a set of F molecule fragments: §= { i, fz, fp} (here, “molecule fragment” refers to the sequence read of a molecular fragment obtained by a sequencing device), wherein each fragment f is observed at a genomic location I. The learned attention subspace is a set of H hypotheses: H= {Hi, H2, . . ., Hi, . . ., HH J . H is the learned subspace where the model uses to attend and may be learned from aggregate or molecular level data. The positional encoding is a set of L location encoders: L= {Li, L2, ..., L;, . . ., LL}. L is the positional encoding. The attention weights include W weight values: W= {Wi, W2, . . ., Wfc,. . Wn / }. W are the attention weights, that in some embodiments can be a function of the learned subspace H. The transform is defined as:wherein, index i is for hypothesis, index j is for location, and index k is for attention weights. Let H be an H F matrix representing the normalized value of each hypothesis for each fragment in the sample §. Let L be an L F matrix representing the location encoding for each fragment in the sample §. Let W be an W F matrix representing the weighting schemas for each fragment in the sample §. With a number of assumptions, the calculation can be simplified and parameterized. It can be assumed that: (1) L is a fixed function of the first non-zero element of each £§; (2) W is a function of H parametrized with fl; (3) H is a linear function of § represented as a M F matrix S parametrized with 0. For the case of negative attention, the weight schemas are a reciprocal of a linear mapping from the H hypothesis, e.g., W=l / fl.H where fl is a W H mapping. This construct provides the ability to learn the parameters 0 using data from either aggregate or molecular level from any disease, providing flexibility in training and generalizability. It only requires the parameters fl to be trained with molecular level data from the disease of interest. This crucially brings the number of training samples needed for a transformer down to a level that is feasible to be collected in a clinical study to design a liquid biopsy test.EXAMPLE METHODS
[0064] FIGs. 5 A, 5B, and 6 refer to example processes relating to the application of a transformer model. In particular, FIGs. 5 A & 5B relate to the training of the transformer model, and FIG. 6 relates to deployment or inference with the trained transformer model. The processes are described from the perspective of an analytics system leveraging the transformer model. In other embodiments, some or all of the steps may be performed by other devices or computer systems.
[0065] FIG. 5A is a flowchart of the process of attention subspace training 500 in the transformer model, according to one or more embodiments.
[0066] The analytics system obtains 505 general training data comprising molecular sequencing data and associated disease data for a plurality of diseases. The general training data may be obtained from a third-party system. The plurality of diseases may be of differing biology, affecting different organs, etc. The plurality of diseases may be genetic-based diseases or disorders. The molecular sequencing data may include aggregate sequencing dataand / or molecule-level sequencing data. The molecular sequencing data may include genetic sequencing data and / or epigenetic sequencing data. Some of the samples in the general training data may be from healthy subjects.
[0067] The analytics system trains 510 the attention subspace of the transformer model with the general training data. Training of the attention subspace generally entails transforming the general training data into embeddings in the attention subspace. The embeddings can be scored with a score function on their representation of the general training data. The analytics system may also train the attention subspace with the targeted training data.
[0068] FIG. 5B is a flowchart of the process of training a transformer model 520, according to one or more embodiments.
[0069] The analytics system obtains 525 targeted training data including molecular sequencing data from samples associated with a target disease (or set of target disease types and / or subtypes). The targeted training data may be from a curated set of subjects empaneled for a particular disease study. The molecular sequencing data may include fragment sequencing data or microarray data. The molecular sequencing data may include genetic sequencing data and / or epigenetic sequencing data. The genetic and / or epigenetic sequencing data may include sequence reads indicating the order of molecular units in a polymer molecule in the sample. Some of the samples in the targeted training data may be from healthy subjects.
[0070] The analytics system generates 530, for each sample and for each molecular sequence read, a molecule embedding based on the information in the molecule. The molecule embedding is a vector representation of the molecular information existing in an embedding space. The embedding space may be of lower dimensionality than the input molecular sequencing space. To generate the molecule embedding, the transformer model may leverage an embedding module, e.g., one that is previously trained to encode the molecular information.
[0071] The analytics system generates 535, for each sample and for each molecular sequence read, a genome-wide positional encoding based on the positional information of the molecule. In the example of genetic and / or epigenetic sequencing data, the positional information of a nucleic acid sequence read may indicate a location and / or a gene in the human genome that the nucleic acid molecule maps to.
[0072] The analytics system applies 540, for each sample and for each molecular sequence read, an inverse attention head to determine a weight vector that inversely attends toinformation in the molecule embedding. The inverse attention head operates in the learned attention subspace, e.g., trained via attention subspace training 500 described in FIG. 5A. In some embodiments, the inverse attention head inputs a multiplied product of the molecule embedding and the positional encoding to output the weight vector. The inverse attention operates to defocus the transformer model from signals that relate to healthy or normal molecules and / or samples.
[0073] The analytics system generates 545, for each sample a sample matrix as a product of the molecule embeddings, the positional encodings, and the weight vectors. The sample matrix serves as a compact representation of the test sample’s molecular sequencing data that includes positional encoding and inverse attention of molecules.
[0074] The analytics system trains 550 the transformer model to classify disease state of the target disease with the targeted training data. Training of the transformer model may entail training one or more of the components in an end-to-end process. In other embodiments, the analytics system may perform disjointed training of one or more components. For example, the analytics system may train the feed-forward network by inputting the sample matrices of the targeted training data into the feed-forward network to output predictions on disease classification. The analytic system adjusts parameters of the feed-forward network to optimize the scores (i.e., to maximize objective and / or to minimize loss). The analytics system may train the inverse attention head by inputting the molecule embeddings and the genome-wide positional encodings. The predictions can be scored with a score function on their accuracy to known disease states of the samples in the targeted training data. The analytics system adjusts parameters of the inverse attention head based to optimize the classification scores.
[0075] FIG. 6 is a flowchart of the process of deploying the transformer model 600 to infer a disease classification for a test subject, according to one or more embodiments.
[0076] The analytics system obtains 610 molecular sequencing data for a test subject. The molecular sequencing data may include fragment sequencing data or microarray data. The molecular sequencing data may include genetic sequencing data and / or epigenetic sequencing data. The genetic and / or epigenetic sequencing data may include sequence reads indicating the order of molecular units in a polymer molecule in the sample.
[0077] The analytics system generates 620 a molecule embedding for each molecule based on the information in the molecule. An embedding module may generate the molecule embedding as a mapping from the molecular sequencing space into the embedding space.
[0078] The analytics system generates 630 a positional encoding for each molecule based on the positional information of the molecule. In the example of genetic and / or epigenetic sequencing data, the positional information of a nucleic acid sequence read may indicate a location and / or a gene in the human genome that the nucleic acid molecule maps to.
[0079] The analytics system applies 640 the inverse attention head to determine a weight vector for each molecule that inversely attends to information within the molecule embedding. In some embodiments, the inverse attention head inputs a multiplied product of the molecule embedding and the positional encoding to output the weight vector.
[0080] The analytics system generates 650 a sample matrix as a product of the molecule embeddings, the positional encodings, and the weight vectors. The sample matrix serves as a compact representation of the test sample’s molecular sequencing data that includes positional encoding and inverse attention of molecules.
[0081] The analytics system applies 660 the feed-forward network to the sample matrix to classify disease state of the test subject. The output prediction may be binary, multi class, or some combination thereof. The prediction may be transmitted or reported to a healthcare provider to diagnose, recommend treatment, prognose, assess disease progression, detect minimal residual disease, etc.EXAMPLE PRACTICAL APPLICATIONS
[0082] In one or more example implementations, the transformer model is used to detect disease in a subject by analyzing the subject’s sequencing data. A nucleic acid sample is collected from the subject and sequenced by a sequencer to obtain sequence reads representing nucleotide sequences of the nucleic acid fragments in the sample. The sequencing data may be preprocessed and then input into the transformer model to output a disease prediction. In some embodiments, the disease prediction indicates a state of a disease in the subject, e.g., presence or absence. In some embodiments, the disease prediction by the transformer model may indicate gradation or progression of the disease, e.g., for cancer, the disease prediction may indicate a stage of the cancer or for metabolic dysfunction-associated steatohepatitis (MASH), the disease prediction may indicate fibrosis stage. In some embodiments, the transformer model may be configured to output a treatment recommendation. The treatment recommendation may identify one or more treatments that are predicted to be efficacious for the predicted disease. In some embodiments, the treatment recommendation may further identify targets for treatment, e.g., particular markers, cells, genes, pathway, etc., that can be the focus of available therapies.
[0083] In one or more example implementations, the transformer model (or another machine-learning model) may be configured to recommend one or more therapies for treatment of the disease. For example cancer treatment options are diverse and often tailored to the specific type and stage of the disease, as well as the patient’s overall health. The most common treatments include surgery, which involves physically removing the tumor; chemotherapy, which uses drugs to kill rapidly dividing cancer cells throughout the body; and radiation therapy, which targets and destroys cancer cells with high-energy rays. In addition, targeted therapy uses drugs designed to specifically attack cancer cells based on their unique genetic or molecular features, while immunotherapy boosts the body’s own immune system to recognize and fight cancer. Hormone therapy may be used for cancers that are sensitive to hormones, such as certain breast or prostate cancers. In some cases, stem cell or bone marrow transplants are performed to restore healthy blood-forming cells after intensive treatment. Increasingly, personalized medicine — where treatment is guided by the genetic or epigenetic makeup of the patient’s tumor — is being used to improve outcomes. Often, a combination of these therapies is employed to maximize effectiveness and minimize side effects. From the set of treatment options, the transformer model may be configured to identify one or more therapies for recommendation, tailored to the subject’s particular disease prediction and / or health state.
[0084] In one or more example implementations, the transformer model may be leveraged for disease monitoring. The model can play a crucial role in ongoing disease monitoring. With the trained model, a patient’s blood samples can be collected at regular intervals (e.g., every week, month, 3-month period, 6-month period, year, 2-year period, 5- year period, 10-year period, etc.). Each sample, e.g., with cell-free DNA, can be sequenced for disease monitoring over time. The resulting sequencing data is processed and input into the model, which analyzes the sequencing data to output the disease prediction. By comparing predictions across multiple time points, the model (and / or with a clinician) can track changes in the disease state — such as progression, remission, or recurrence — in a minimally invasive manner. This recurrent sampling and inference approach empowers realtime, personalized monitoring.
[0085] In one or more example implementations, the model is leveraged for assessment of treatment efficacy. A pre-treatment sample can be collected and analyzed by the trained model to establish a baseline disease state. After the patient undergoes treatment, a follow-up sample can be collected and similarly processed by the trained model. By evaluating changes in the model’s predictions — such as a reduction in disease-associated signals or a shift towarda non-disease state — clinicians can objectively measure the biological response to the therapy. If the therapy is deemed inefficacious, the clinician can be armed with such insight to modify the treatment plan. This can include stopping the current therapy, modifying parameters of the current therapy (e.g., dosage, duration, intensity, etc.), introducing a new therapy, or some combination thereof. Based on the determined treatment efficacy, the trained model may further be tuned (in such embodiments with functionality for treatment recommendation). This approach provides a sensitive, minimally invasive means to detect whether the treatment is effectively reducing or eliminating disease, even before clinical symptoms change or traditional imaging methods can detect a response. Ultimately, this enables more precise and timely adjustments to treatment strategies, improving patient management and outcomes.
[0086] In one or more example implementations, the model can also be used for curating high-quality training datasets from other cohorts of sequencing data. By applying the model to new, unlabeled or partially labeled datasets, researchers can identify samples that are likely to represent clear cases of disease or non-disease states based on the model’s predictions. These confidently classified samples can then be prioritized for inclusion in future training sets, helping to reduce noise and mislabeling that might otherwise hinder model performance. Additionally, the model can highlight ambiguous or borderline cases, guiding researchers to review or validate these samples more closely before adding them to the dataset. This iterative process not only streamlines the curation of large and diverse sequencing cohorts but also helps to continuously improve the accuracy and generalizability of disease prediction models as new data becomes available. Moreover, the model can be used to identify confounding signals of other disease states that can muddle the training process.
[0087] In one or more example implementations, retraining of the model involves incorporating new data to enhance its accuracy and adaptability. In this process, a fresh set of sequencing data — distinct from the original training set — is analyzed by the existing model to generate disease state predictions. These predictions are then reviewed by a clinician or healthcare provider, who corrects any misclassifications and assigns accurate labels based on clinical expertise and additional patient information. The newly labeled data, now reflecting expert-verified ground truth, is then combined with the original training data or used as an independent set to retrain the model. This iterative approach allows the model to learn from its previous errors, adapt to new patterns or rare cases, and ultimately improve its predictive performance and reliability in real-world clinical settings. The analytics system may continue to iterate or refine the transformer model. This retraining process can target a generalattention head within the model that was initially trained on non-specific sequencing data from various disease states. By introducing new, accurately labeled data and updating the attention head’s parameters, the model can refine its ability to focus on the most relevant features across diverse disease contexts. This targeted retraining helps the attention head become more sensitive to subtle, disease-specific patterns, thereby enhancing the model’s overall capacity to distinguish between different disease states and improving its generalizability to new or rare conditions. The tuned model is better equipped for training of subsequent models specific to other disease types, that may have more limited training data. This modularity in the transformer training is a technical improvement, as the generalized attention head is adaptable across varying disease states.
[0088] In one or more example implementations, a point-of-care device for molecular sequencing performs rapid analysis of liquid biopsy samples, providing key improvements and advantages for clinical diagnostics and disease monitoring. The device integrates advanced portable sequencing technology, e.g., nanopore-based platforms, with automated sample preparation and real-time data analysis capabilities. Upon receiving a biological sample, e.g., a volume of blood or another bodily fluid with nucleic acid molecules, the device extracts cell-free nucleic acids and performs high-throughput sequencing to detect disease-associated mutations, epigenetic changes or biomarkers. To accomplish this, the point-of-care device may combine both functionality of a sequencer and functionality of an analytics system. This approach eliminates the need for centralized laboratory processing, significantly reducing turnaround time and enabling clinicians to make immediate, informed decisions regarding diagnosis, treatment, or disease monitoring. The system is designed for ease of use, high sensitivity, and accuracy, supporting robust performance in diverse clinical settings while maintaining compliance with regulatory and quality standards.
[0089] Based on outputs of the analyses, the point-of-care device may perform additional analyses or other actions. For example, the point-of-care device may have a schedule for sampling and analyzing a subject’s sequencing data for disease detection. Upon identifying the disease state of the subject reflects an increased disease signal, the point-of-care device may modify the schedule to increase the frequency in sampling and analysis. Or, in another example, following a positive diagnosis and treatment, the point-of-care device may be used to evaluate recurrence of the disease. Upon detecting a decrease in disease signal, the point- of-care device may decrease the frequency in sampling and analysis.EXAMPLE ASPECTS
[0090] Clause 1. A computer-implemented method comprising: obtaining general training data from a first plurality of subjects, wherein each subject of the first plurality of subjects is associated with one of a plurality of disease types or subtypes; obtaining targeted training data from a second plurality of subjects associated with a target disease; training an attention subspace of a transformer model as a latent space of the molecular sequence reads in one or both of: the general training data and the targeted training data; applying a transformer model to molecular sequencing data for each subject in the targeted training data by: generating a molecule embedding for each molecular sequence read in the sample based at least in part on molecular information in the molecular sequence read, generating a genome-wide positional encoding for each molecular sequence read in the sample at least in part based on positional information in the molecular sequence read, applying an inverse attention head based in the attention subspace to each molecule embedding to determine a weight vector for the molecule embedding, generating a sample matrix as a product of the molecule embeddings, the genome-wide positional encodings, and the weight vectors, and applying a feed-forward network of the transformer model to the sample matrix to classify a disease state for the target disease; and training the transformer model with the targeted training data.
[0091] Clause 2. The computer-implemented method of clause 1, further comprising: obtaining molecular sequencing data sequenced from a biological sample obtained from a test subject; applying the trained transformer model to the molecular sequencing data of the test subject to determine the disease state for the target disease.
[0092] Clause 3. The computer-implemented method of clause 2, wherein the disease state further indicates one of: a prognosis of the target disease in the test subject, a progression of the target disease in the test subject, and a treatment recommendation for treating the target disease in the test subject.
[0093] Clause 4. The computer-implemented method of clause 2, wherein the biological sample is obtained post-treatment of the test subject for the target disease with a therapy, the method further comprising: determining efficacy of the therapy in treating the target disease by comparing the disease state for the target disease post-treatment to a disease state for the target disease pre-treatment; and modifying the therapy based on the determined efficacy.
[0094] Clause 5. The computer-implemented method of clause 1, wherein training the attention subspace of the transformer model as the latent space comprises training theattention subspace with aggregate-level sequencing data that represents numerous sequence reads or molecular-level sequencing data that represents individual sequence reads.
[0095] Clause 6. The computer-implemented method of clause 1, wherein training the attention subspace of the transformer model as the latent space comprises: training the attention subspace with the general training data; and tuning the attention subspace with the targeted training data.
[0096] Clause 7. The computer-implemented method of clause 1, wherein applying the inverse attention head based in the attention subspace to each molecule embedding to determine a weight vector for the molecule embedding comprises defocusing from normal sequence reads.
[0097] Clause 8. The computer-implemented method of clause 1, the attention subspace is concurrently trained with the feed-forward network of the transformer model by aggregating a score from a first function for learning the attention subspace and a score from a second function for learning the targeted training data by the transformer model.
[0098] Clause 9. The computer-implemented method of clause 1, wherein generating the molecule embedding for each molecular sequence read in the sample comprises outputting the molecular embedding in a latent space.
[0099] Clause 10. The computer-implemented method of clause 1, wherein generating the genome-wide positional encoding for each molecular sequence read in the sample comprises performing a multiplicative encoding to conform dimensionality prior to generating the sample matrix.
[0100] Clause 11. A computer-implemented method comprising: obtaining molecular sequencing data sequenced from a biological sample obtained or derived from a test subject; applying a transformer model to the molecular sequencing data of the test subject to determine a disease state for a target disease by: generating a molecule embedding for each molecular sequence read in the biological sample based on at least in part molecular information in the molecular sequence read, generating a genome-wide positional encoding for each molecular sequence read in the biological sample based at least in part on positional information in the molecular sequence read, applying an inverse attention head based in an attention subspace to each molecule embedding to determine a weight vector for the molecule embedding, the attention subspace being a latent space of molecular sequencing data trained on general training data from a first plurality of subjects associated with one of a plurality of diseases, generating a sample matrix as a product of the molecule embeddings, the genomewide positional encodings, and the weight vectors, and applying a feed-forward network ofthe transformer model to the sample matrix to classify a disease state for the target disease, wherein the transformer model is trained on targeted training data from a second plurality of subjects, wherein a subset of the second plurality of subjects is associated with the target disease.
[0101] Clause 12. The computer-implemented method of clause 11, wherein the disease state further indicates one of: a prognosis of the target disease in the test subject, a progression of the target disease in the test subject, and a treatment recommendation for treating the target disease in the test subject.
[0102] Clause 13. The computer-implemented method of clause 11, wherein the biological sample is obtained post-treatment of the test subject for the target disease with a therapy, the method further comprising: determining efficacy of the therapy in treating the target disease by comparing the disease state for the target disease post-treatment to a disease state for the target disease pre-treatment; and modifying the therapy based on the determined efficacy.
[0103] Clause 14. The computer-implemented method of clause 11, wherein the attention subspace is trained by: obtaining general training data from a plurality of subjects, wherein each subject of a subset of the first plurality of subjects is associated with one of a plurality of disease types or subtypes; and training the attention subspace of the transformer model as a latent space of the molecular sequence reads in the general training data.
[0104] Clause 15. The computer-implemented method of clause 14, wherein training the attention subspace of the transformer model as the latent space comprises training the attention subspace with aggregate-level sequencing data that represents numerous sequence reads or molecular-level sequencing data that represents individual sequence reads.
[0105] Clause 16. The computer-implemented method of clause 14, wherein training the attention subspace of the transformer model as the latent space comprises: training the attention subspace with the general training data; and tuning the attention subspace with targeted training data from a second plurality of subjects associated with the target disease.
[0106] Clause 17. The computer-implemented method of clause 11, wherein applying the inverse attention head based in the attention subspace to each molecule embedding to determine a weight vector for the molecule embedding comprises defocusing from normal sequence reads.
[0107] Clause 18. The computer-implemented method of clause 11, the attention subspace is concurrently trained with the feed-forward network of the transformer model byaggregating a score from a first function for learning the attention subspace and a score from a second function for learning the targeted training data by the transformer model.
[0108] Clause 19. The computer-implemented method of clause 11, wherein generating the molecule embedding for each molecular sequence read in the sample comprises outputting the molecular embedding in a latent space.
[0109] Clause 20. The computer-implemented method of clause 11, wherein generating the genome-wide positional encoding for each molecular sequence read in the sample comprises performing a multiplicative encoding to conform dimensionality prior to generating the sample matrix.
[0110] Clause 21. A computer-implemented method comprising: obtaining targeted training data from a plurality of subjects associated with a target disease; applying a transformer model to molecular sequencing data for each subject in the targeted training data by: generating a molecule embedding for each molecular sequence read in the sample based at least in part on molecular information in the molecular sequence read, generating a genome-wide positional encoding for each molecular sequence read in the sample at least in part based on positional information in the molecular sequence read, applying an inverse attention head based in the attention subspace to each molecule embedding to determine a weight vector for the molecule embedding, generating a sample matrix as a product of the molecule embeddings, the genome-wide positional encodings, and the weight vectors, and applying a feed-forward network of the transformer model to the sample matrix to classify a disease state for the target disease (e.g., for applications including disease diagnosis, prognostication, monitoring, treatment selection, or treatment response assessment); and training the transformer model with the targeted training data.[OHl] Clause 22. A computer-implemented method comprising: obtaining targeted training data from a plurality of subjects associated with a target disease; applying a transformer model to molecular sequencing data for each subject in the targeted training data by: generating a molecule embedding for each molecular sequence read in the sample based at least in part on molecular information in the molecular sequence read, applying an inverse attention head based in the attention subspace to each molecule embedding to determine a weight vector for the molecule embedding, generating a sample matrix as a product of the molecule embeddings and the weight vectors, and applying a feed-forward network of the transformer model to the sample matrix to classify a disease state for the target disease (e.g., for applications including disease diagnosis, prognostication, monitoring, treatment selection,or treatment response assessment); and training the transformer model with the targeted training data.
[0112] Clause 23. A computer-implemented method comprising: obtaining targeted training data from a plurality of subjects associated with a target disease; applying a transformer model to molecular sequencing data for each subject in the targeted training data by: generating a molecule embedding for each molecular sequence read in the sample based at least in part on molecular information in the molecular sequence read, generating a genome-wide positional encoding for each molecular sequence read in the sample at least in part based on positional information in the molecular sequence read, generating a sample matrix as a product of the molecule embeddings and the genome-wide positional encodings, and applying a feed-forward network of the transformer model to the sample matrix to classify a disease state for the target disease (e.g., for applications including disease diagnosis, prognostication, monitoring, treatment selection, or treatment response assessment); and training the transformer model with the targeted training data.
[0113] Clause 24. A non-transitory computer-readable storage medium storing instructions that, when executed by a computer processor, cause the computer processor to perform the method of any one of the foregoing clauses.
[0114] Clause 25. A system comprising: a computer processor; and the non-transitory computer-readable storage medium of clause 24.
[0115] Clause 26. A point-of-care device comprising: a sequencing device; and the system of clause 25.ADDITIONAL CONSIDERATIONS
[0116] The foregoing description of the embodiments has been presented for the purpose of illustration; many modifications and variations are possible while remaining within the principles and teachings of the above description.
[0117] Any of the steps, operations, or processes described herein may be performed or implemented with one or more hardware or software modules, alone or in combination with other devices. In some embodiments, a software module is implemented with a computer program product comprising one or more computer-readable media storing computer program code or instructions, which can be executed by a computer processor for performing any or all of the steps, operations, or processes described. In some embodiments, a computer-readable medium comprises one or more computer-readable media that, individually or together, comprise instructions that, when executed by one or moreprocessors, cause the one or more processors to perform, individually or together, the steps of the instructions stored on the one or more computer-readable media. Similarly, a processor may comprise one or more subprocessing units that, individually or together, perform the steps of instructions stored on a computer-readable medium.
[0118] Embodiments may also relate to a product that is produced by a computing process described herein. Such a product may store information resulting from a computing process, where the information is stored on a non-transitory, tangible computer-readable medium and may include any embodiment of a computer program product or other data combination described herein.
[0119] The description herein may describe processes and systems that use machinelearning models in the performance of their described functionalities. A “machine-learning model,” as used herein, comprises one or more machine-learning models that perform the described functionality. Machine-learning models may be stored on one or more computer- readable media with a set of weights. These weights are parameters used by the machinelearning model to transform input data received by the model into output data. The weights may be generated through a training process, whereby the machine-learning model is trained based on a set of training examples and labels associated with the training examples. The training process may include: applying the machine-learning model to a training example, comparing an output of the machine-learning model to the label associated with the training example, and updating weights associated for the machine-learning model through a back- propagation process. The weights may be stored on one or more computer-readable media, and are used by a system when applying the machine-learning model to new data.
[0120] The language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to narrow the inventive subject matter. It is therefore intended that the scope of the patent rights be limited not by this detailed description, but rather by any claims that issue on an application based hereon.
[0121] As used herein, the terms “comprises,” “comprising,” “includes,” “including,” “has,” “having,” or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, “or” refers to an inclusive “or” and not to an exclusive “or”. For example, a condition “A or B” is satisfied by any one of the following: A is true (or present) and B is false (or not present); A is false (or not present) and B is true (or present);and both A and B are true (or present). Similarly, a condition “A, B, or C” is satisfied by any combination of A, B, and C being true (or present). As a not-limiting example, the condition “A, B, or C” is satisfied when A and B are true (or present) and C is false (or not present). Similarly, as another not-limiting example, the condition “A, B, or C” is satisfied when A is true (or present) and B and C are false (or not present).
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented method comprising: obtaining general training data from a first plurality of subjects, wherein each subject of the first plurality of subjects is associated with one of a plurality of disease types or subtypes; obtaining targeted training data from a second plurality of subjects associated with a target disease; training an attention subspace of a transformer model as a latent space of the molecular sequence reads in one or both of: the general training data and the targeted training data; applying a transformer model to molecular sequencing data for each subject in the targeted training data by: generating a molecule embedding for each molecular sequence read in the sample based at least in part on molecular information in the molecular sequence read, generating a genome-wide positional encoding for each molecular sequence read in the sample at least in part based on positional information in the molecular sequence read, applying an inverse attention head based in the attention subspace to each molecule embedding to determine a weight vector for the molecule embedding, generating a sample matrix as a product of the molecule embeddings, the genome-wide positional encodings, and the weight vectors, and applying a feed-forward network of the transformer model to the sample matrix to classify a disease state for the target disease; and training the transformer model with the targeted training data.
2. The computer-implemented method of claim 1, further comprising: obtaining molecular sequencing data sequenced from a biological sample obtained from a test subject;applying the trained transformer model to the molecular sequencing data of the test subject to determine the disease state for the target disease.
3. The computer-implemented method of claim 2, wherein the disease state further indicates one of: a prognosis of the target disease in the test subject, a progression of the target disease in the test subject, and a treatment recommendation for treating the target disease in the test subject.
4. The computer-implemented method of claim 2, wherein the biological sample is obtained post-treatment of the test subject for the target disease with a therapy, the method further comprising: determining efficacy of the therapy in treating the target disease by comparing the disease state for the target disease post-treatment to a disease state for the target disease pre-treatment; and modifying the therapy based on the determined efficacy.
5. The computer-implemented method of claim 1, wherein training the attention subspace of the transformer model as the latent space comprises training the attention subspace with aggregate-level sequencing data that represents numerous sequence reads or molecular-level sequencing data that represents individual sequence reads.
6. The computer-implemented method of claim 1, wherein training the attention subspace of the transformer model as the latent space comprises: training the attention subspace with the general training data; and tuning the attention subspace with the targeted training data.
7. The computer-implemented method of claim 1, wherein applying the inverse attention head based in the attention subspace to each molecule embedding to determine a weight vector for the molecule embedding comprises defocusing from normal sequence reads.
8. The computer-implemented method of claim 1, the attention subspace is concurrently trained with the feed-forward network of the transformer model by aggregating a score from a first function for learning the attention subspace and a score from a second function for learning the targeted training data by the transformer model.
9. The computer-implemented method of claim 1, wherein generating the molecule embedding for each molecular sequence read in the sample comprises outputting the molecular embedding in a latent space.
10. The computer-implemented method of claim 1, wherein generating the genome-wide positional encoding for each molecular sequence read in the sample comprises performing a multiplicative encoding to conform dimensionality prior to generating the sample matrix.
11. A computer-implemented method comprising: obtaining molecular sequencing data sequenced from a biological sample obtained or derived from a test subject; applying a transformer model to the molecular sequencing data of the test subject to determine a disease state for a target disease by: generating a molecule embedding for each molecular sequence read in the biological sample based on at least in part molecular information in the molecular sequence read, generating a genome-wide positional encoding for each molecular sequence read in the biological sample based at least in part on positional information in the molecular sequence read, applying an inverse attention head based in an attention subspace to each molecule embedding to determine a weight vector for the molecule embedding, the attention subspace being a latent space of molecular sequencing data trained on general training data from a first plurality of subjects associated with one of a plurality of diseases, generating a sample matrix as a product of the molecule embeddings, the genome-wide positional encodings, and the weight vectors, and applying a feed-forward network of the transformer model to the sample matrix to classify a disease state for the target disease, wherein the transformer model is trained on targeted training data from a second plurality of subjects, wherein a subset of the second plurality of subjects is associated with the target disease.
12. The computer-implemented method of claim 11, wherein the disease state further indicates one of: a prognosis of the target disease in the test subject, a progression of the target disease in the test subject, and a treatment recommendation for treating the target disease in the test subject.
13. The computer-implemented method of claim 11, wherein the biological sample is obtained post-treatment of the test subject for the target disease with a therapy, the method further comprising: determining efficacy of the therapy in treating the target disease by comparing the disease state for the target disease post-treatment to a disease state for the target disease pre-treatment; and modifying the therapy based on the determined efficacy.
14. The computer-implemented method of claim 11, wherein the attention subspace is trained by: obtaining general training data from a plurality of subjects, wherein each subject of a subset of the first plurality of subjects is associated with one of a plurality of disease types or subtypes; and training the attention subspace of the transformer model as a latent space of the molecular sequence reads in the general training data.
15. The computer-implemented method of claim 14, wherein training the attention subspace of the transformer model as the latent space comprises training the attention subspace with aggregate-level sequencing data that represents numerous sequence reads or molecular-level sequencing data that represents individual sequence reads.
16. The computer-implemented method of claim 14, wherein training the attention subspace of the transformer model as the latent space comprises: training the attention subspace with the general training data; and tuning the attention subspace with targeted training data from a second plurality of subjects associated with the target disease.
17. The computer-implemented method of claim 11, wherein applying the inverse attention head based in the attention subspace to each molecule embedding to determine a weight vector for the molecule embedding comprises defocusing from normal sequence reads.
18. The computer-implemented method of claim 11, the attention subspace is concurrently trained with the feed-forward network of the transformer model by aggregating a score from a first function for learning the attention subspace and a score from a second function for learning the targeted training data by the transformer model.
19. The computer-implemented method of claim 11, wherein generating the molecule embedding for each molecular sequence read in the sample comprises outputting the molecular embedding in a latent space.
20. The computer-implemented method of claim 11, wherein generating the genome-wide positional encoding for each molecular sequence read in the sample comprises performing a multiplicative encoding to conform dimensionality prior to generating the sample matrix.
21. A computer-implemented method comprising: obtaining targeted training data from a plurality of subjects associated with a target disease; applying a transformer model to molecular sequencing data for each subject in the targeted training data by: generating a molecule embedding for each molecular sequence read in the sample based at least in part on molecular information in the molecular sequence read, generating a genome-wide positional encoding for each molecular sequence read in the sample at least in part based on positional information in the molecular sequence read, applying an inverse attention head based in the attention subspace to each molecule embedding to determine a weight vector for the molecule embedding, generating a sample matrix as a product of the molecule embeddings, the genome-wide positional encodings, and the weight vectors, and applying a feed-forward network of the transformer model to the sample matrix to classify a disease state for the target disease (e.g., for applications including disease diagnosis,prognostication, monitoring, treatment selection, or treatment response assessment); and training the transformer model with the targeted training data.
22. A computer-implemented method comprising: obtaining targeted training data from a plurality of subjects associated with a target disease; applying a transformer model to molecular sequencing data for each subject in the targeted training data by: generating a molecule embedding for each molecular sequence read in the sample based at least in part on molecular information in the molecular sequence read, applying an inverse attention head based in the attention subspace to each molecule embedding to determine a weight vector for the molecule embedding, generating a sample matrix as a product of the molecule embeddings and the weight vectors, and applying a feed-forward network of the transformer model to the sample matrix to classify a disease state for the target disease (e.g., for applications including disease diagnosis, prognostication, monitoring, treatment selection, or treatment response assessment); and training the transformer model with the targeted training data.
23. A computer-implemented method comprising: obtaining targeted training data from a plurality of subjects associated with a target disease; applying a transformer model to molecular sequencing data for each subject in the targeted training data by: generating a molecule embedding for each molecular sequence read in the sample based at least in part on molecular information in the molecular sequence read,generating a genome-wide positional encoding for each molecular sequence read in the sample at least in part based on positional information in the molecular sequence read, generating a sample matrix as a product of the molecule embeddings and the genome-wide positional encodings, and applying a feed-forward network of the transformer model to the sample matrix to classify a disease state for the target disease (e.g., for applications including disease diagnosis, prognostication, monitoring, treatment selection, or treatment response assessment); and training the transformer model with the targeted training data.
24. A non-transitory computer-readable storage medium storing instructions that, when executed by a computer processor, cause the computer processor to perform the method of any one of the foregoing claims.
25. A system comprising: a computer processor; and the non-transitory computer-readable storage medium of claim 24.
26. A point-of-care device comprising: a sequencing device; and the system of claim 25.
Citation Information
Patent Citations
Transformer-Based Neural Network including a Mask Attention Network
US20220067533A1
Disease representation and classification with machine learning
US20230222176A1
Designing biomolecule sequence variants with pre-specified attributes
US20230268026A1