Large genetic model (LGM) for genomic prediction to personalize medical treatment
Patent Information
- Application Number
- PCT/IL2026/050190
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-27
- Filing Date
- 2026-02-26
- Publication Date
- 2026-09-03
Smart Images

Figure IL2026050190_03092026_PF_FP_ABST
Abstract
Description
P-655803-PCLARGE GENETIC MODEL (LGM) FOR GENOMIC PREDICTION TO PERSONALIZE MEDICAL TREATMENTFIELD OF THE INVENTION
[0001] Embodiments of the present invention relate to the field of personalized medical treatment based on genomic modeling and bioinformatics, using a computational framework for analyzing metagenomically sequenced DNAto predict genetically encoded biological conditions. In particular, embodiments of the invention provide a Large Genetic Model (LGM) comprising an extensive library of biomarkers that, in response to genetic prompts comprising microbial and host (human) genomic DNA sequences, predicts biological conditions of the DNA’s host. Detection of those biologically-correlated biomarkers in patients prompts administering treatments for those biological conditions to patients with improved treatment efficacy.BACKGROUND OF THE INVENTION
[0002] Genomic analysis has long been hindered by the lack of internal ordering in genetic sequences, making it difficult to extract meaningful features from raw sequencing data. The loss of internal ordering is a fundamental challenge in metagenomic sequencing, where sequencing technologies fragment genetic material into short reads without preserving the original arrangement.
[0003] The Problem of Internal Ordering in Genomic Data:Unlike traditional genome sequencing, where long contiguous sequences are obtained (e.g., through long-read sequencing technologies), metagenomic sequencing (also known as "shotgun sequencing") captures millions or even billions of short “reads” of nucleic acid sequences that originate from multiple sources within a sample. These reads are generated in a randomized fashion, losing all contextual information regarding their original genomic arrangement. The implications of this disorder include:• Loss of Chromosomal Context (Human Genomes): In human genome sequencing, reads come from different chromosomes, but there is no inherent way to determine whether two reads belong to the same chromosome or are functionally related.• Microbial Complexity (Microbiome Studies): When analyzing microbiomes, sequencing produces a mixture of reads originating from different microbial species, but standard sequencing methods cannot determine whether two reads originate from the same organism.P-655803-PC• Disruption of Functional Interpretation: Many biological processes rely on the spatial arrangement of genetic elements. Regulatory elements, enhancers, and promoters influence gene expression based on their positioning, but this relationship is lost when sequencing is performed without preserving ordering.
[0004] Effects of The Problem of Internal Ordering:
[0005] The inability to determine how sequencing reads relate to each other makes conventional approaches insufficient for many tasks in both host and microbial genomic studies. Traditional gene annotation and alignment-based methods depend on the assumption that genes exist in contiguous sequences. However, metagenomic sequencing breaks down this continuity, leaving avast dataset of fragmented genetic information with no clear relationship between different segments.
[0006] Even deep learning models, such as NVIDI A's™ Evo2-40B, face limitations when applied to this type of data. These models rely on sequential data streams, meaning they perform optimally when the underlying data retains an ordered structure. Since metagenomic data lacks this structure, deep learning methods struggle to identify functional patterns that span across different reads.SUMMARY OF THE INVENTION
[0007] A Large Genetic Model (LGM) is a trained artificial intelligent model that encodes vast amount of DNA sequences into a compact and efficient embedding space of biomarkers representing cluster groups of co-occurring DNA sequences, restoring contextual co-occurrence relationships between DNA sequences conventionally lost with metagenomic sequencing. The biomarker embedding space encodes statistically significant groups or clusters of DNA sequence that co-occur in DNA segments or sub-lengths in anomalous distribution patterns that deviate from a neutral (e.g., power law) distribution and so, their expression is unlikely randomly statistical, and therefore likely associated with a phenotypic property associated with the DNA-donor ’s biological condition. A query comprising a patient’s DNA containing those biological condition-correlated cluster groups is input into the LGM, which executes the biomarker embedding space analysis to predict the patient’s biological condition correlated with those biomarker embeddings. Predicting the patient’s biological condition may prompt administering a treatment to improve that biological condition for improved personalized medical treatment.
[0008] The LGM may (i) initially be constructed from a large-scale genetic repository by clustering its DNA sequences into a biomarker embedding space, then (ii) in a training phase training its supervised learning classifier to correlate clusters of DNA sequences of multiple hosts with classification labels of those hosts’ biological conditions, then (iii) in a run-time phase executing theP-655803-PCLGM to predict a patient’s biological condition using the trained supervised classifier based on the patient’s clusters of DNA sequences. The LGM may analyze genetic prompts to map both labeled datasets in a training phase and query sequences in a run-time phase into the unified biomarker-based embedding space, enabling efficient and accurate prediction of biological conditions based on decentralized co-occurring DNA sequences previously fragmented and undiscoverable by conventional metagenomic sequencing.
[0009] In an embodiment of the invention, an LGM may be constructed by receiving a large-scale database of sequenced DNA (e.g. reads) metagenomic sequenced from genome or microbiome DNA samples of vast numbers of hosts across one or more species. Aplurality of DNA sequences (e.g., k-mers) may be extracted from the database of sequenced DNA. The plurality of DNA sequences may be grouped into clusters using unsupervised learning. Each cluster comprises multiple of the plurality of DNA sequences that co-occur with statistical significance (e.g., non-neutral and / or non-random distribution of co-occurrence of DNA sequences in the cluster) that is unlikely to be random and therefore likely as a candidate for correlation with biological conditions. The clusters of DNA sequences may be encoded as biomarkers embedded in a biomarker embedding space (e.g., a compact vector space of statistically significant biomarkers of clusters of DNA sequence). The biomarker embedding space may be iteratively updated with new biomarkers based on new sequenced DNA in the repository. The LGM may be trained for a supervised classifier to correlate clusters of DNA sequences of observed hosts embedded in the biomarker embedding space with the classification labels of the biological condition of those hosts. The trained LGM may a new predict the patient’s biological condition using the trained supervised classifier based on clusters of the patient’s DNA sequences in the biomarker embedding space. Once the patient’s biological condition is predicted, a treatment for the biological condition may be administered to the patient to improve the biological condition and / or treat the biological condition with improved efficacy e.g., compared to average populations, control groups, and / or to patients not predicted to have the biological condition or without those biomarkers (or with fewer or different combinations thereof).BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0011] The subject matter regarded as the invention is particularly pointed out and distinctly claimed in the concluding portion of the specification. The invention, however, both as to organization and method of operation, together with objects, features, and advantages thereof, may best beP-655803-PCunderstood by reference to the following detailed description when read with the accompanying drawings in which:
[0012] Fig. 1 schematically illustrates a system for executing a Large Genetic Model (LGM) comprising a compact and efficient biomarker embedding space for converting DNA sequences into predictions of, and personalize treatments for, biological conditions, according to an embodiment of the invention;
[0013] Fig. 2 schematically illustrates a portion of a DNA network 100 representing DNA sequences grouped into clusters to generate biomarkers in a biomarker embedding space, according to an embodiment of the invention; and
[0014] Fig. 3 is a flowchart of a method for executing a Large Genetic Model (LGM), according to an embodiment of the invention.
[0015] It will be appreciated that for simplicity and clarity of illustration, elements shown in the figures have not necessarily been drawn to scale. For example, the dimensions of some of the elements may be exaggerated relative to other elements for clarity. Further, where considered appropriate, reference numerals may be repeated among the figures to indicate corresponding or analogous elements.DETAILED DESCRIPTION OF THE INVENTION
[0016] Embodiments of the present invention provide a Large Genetic Model (LGM) that answers genetic prompts comprising microbiome or genomic DNA with responses that classify and predict biological conditions of the DNA host using an extensive library of biomarker-based representations extracted from diverse microbiome and genetic sources.
[0017] LGM Solution to the Problem
[0018] The Large Genetic Model (LGM) introduces a new paradigm by bypassing the need for internal ordering altogether. Instead of relying on contiguous sequences, the LGM focuses on biomarker-based representations that group DNA sequences (e.g., k-mers) into functionally relevant clusters. These clusters act as high-dimensional embeddings, allowing for a context-independent analysis of genetic information, whether from native host (human) DNA and / or the host’s microbiome DNA.
[0019] Rather than reconstructing genome structure, LGM treats sequencing data as a network of co-occurring genetic sequences, extracting patterns that are independent of their physical location in the genome. This approach enables:• Genomic prediction and classification without requiring contiguous sequences.P-655803-PC• A scalable and adaptable framework that improves as new biomarker data is introduced. • Cross-species and cross-sample analysis by comparing shared functional elements of cooccurring genetic sequences instead of raw sequences.
[0020] By leveraging biomarker embeddings rather than linear sequences, LGM provides a robust and scalable solution to the challenge of unordered metagenomic data, offering a transformative approach to genomic intelligence applicable to both host and microbial datasets.
[0021] The Large Genetic Model (LGM) may include multiple distinct phases of (1) Construction, (2) Training and (3) Prediction (run-time):
[0022] 1. Construction Phase (Unsupervised Biomarker Extraction): The LGM may initially be constructed from a large-scale genetic repository by clustering its DNA sequences into a biomarker embedding space. The LGM may be constructed by inputting a large-scale database or repository of sequenced DNA (e.g. reads) metagenomic sequenced from genome or microbiome DNA samples of many hosts (e.g., FASTQ, BAM, or VCF fdes). Example repositories include various genome projects across many samples of the same or different species. From these DNA repositories, DNA sequences (e.g., k-mers) may be extracted, for example, using a sliding window encompassing a predetermined length of (e.g., k, overlapping or non-overlapping) sequential nucleotides. Those DNA sequences may be grouped into clusters using an unsupervised learning technique, e.g., using a DNA network as described in reference to Fig. 2. Each cluster may comprise multiple of the DNA sequences that co-occur with statistical significance (e.g., non-neutral and / or non-random distribution of co-occurrence of DNA sequences in the cluster). Statistical significance may be detected when the DNA sequences in the cluster have higher (or lower) co-occurrence frequencies than a power-law distribution, for example, as disclosed in International Patent Application Publication No. WO / 2025 / 052372 published March 13, 2025 entitled Efficient Detection of Decentralized Biomarkers of Groups of DNA Sequences in Microbiome DNA, International Patent Application No. PCT / IL2026 / 050179 filed February 25, 2026 entitled Scalable Detection of Decentralized Genetic Biomarkers in Human Microbial Data to Personalize Medical Treatment, and / or International Patent Application No. PCT / IL2026 / 050180 filed February 25, 2026 entitled Detection of Cross-Chromosomal Biomarkers From Groups Of DNA and RNA Sequences in Human Genome to Personalize Medical Treatment, all of which are incorporated by reference herein in their entirety. Additionally or alternatively, statistical significance may be detected when the DNA sequences in the cluster are “dominant” such that they have highe st or higher than threshold independent or cumulative occurrence frequencies, for example, as disclosed in International Patent Application No. PCT / IL2026 / 050178 filed February 25, 2026 entitled Efficient Unsupervised Detection of DominantP-655803-PCDNA Sequences in Genomic and Microbiome Samples to Personalize Medical Treatment, which is incorporated by reference herein in its entirety. The clusters of DNA sequences may then be encoded as biomarkers embedded in a biomarker embedding space. In an embodiment, the biomarker embedding space may be a vector space of statistically significant biomarkers of clusters of DNA sequence, where each DNA sample is represented as an N-length vector of whether or not it contains each of the N biomarkers or an M-length vector of whether or not it contains each of the M DNA sequences of a cluster. Additionally or alternatively, the biomarker embedding space may be generated by executing an auto-encoder to encode the sequenced DNA as clusters of DNA sequences, to represent each new sample by a vector of size F. The auto-encoder may produce a significantly reduced number of features compared to the N or M-length vector space encoding (F « N, M), which may improve the efficiency of model training and prediction by increasing their execution speed reducing their storage usage and memory consumption, as well as reducing overfitting (e.g., especially when there is less data).
[0023] The LGM may be dynamically updated with each new genome or microbiome repository update adding new sequenced DNA metagenomic sequenced from new DNA samples of the same or different one or more hosts. This may automatically expand the biomarker embedding space with new clusters of DNA sequences to improve biomarker identification and classification accuracy based on those biomarker overtime.
[0024] In an example, the LGM may be constructed as follows:• Input: Large-scale metagenomic sequencing data (e.g., FASTQ, BAM, or VCF files) from any source (e.g., human, microbial, environmental, plant), e.g., generated by genetic sequencer 102 of Fig. 1.• Process:o Extracts DNA sequences (e.g., k-mers) from sequencing DNA.o Groups co-occurring DNA sequences into clusters using unsupervised learning (e.g., anomalous co-occurring k-mer groups). The DNA sequences may be grouped into the clusters using a DNA network, e.g., as described in reference to Fig. 2.o Stores these clusters as high-dimensional embeddings in a canonical biomarker space (e.g., an N-biomarker, M-kmer group, or F-feature space).o Iteratively updates the embedding space with new sequencing data to refine biomarker definitions.• Output: A dynamically evolving pre-trained model containing a structured library of biomarker embeddings of clusters of DNA sequences that co-occur with statistical significance.P-655803-PC
[0025] 2. Training Phase (Supervised learning using known biological conditions of hosts): The LGM may be trained using a supervised learning classifier to correlate clusters of DNA sequences of multiple hosts with classification labels of those hosts’ biological conditions. The LGM may be trained using labeled DNA sequences that are metagenomic sequenced from multiple genome or microbiome DNA samples of multiple hosts. The labels may provide a genomic classification of a biological condition of the respective hosts. The labeled DNA sequences may be mapped into the biomarker embedding space (e.g., indicating whether or not the clusters of DNA sequences are contained therein and / or a frequency or distribution thereof). A supervised classifier may be trained to correlate the clusters of DNA sequences in the biomarker embedding space with the classification labels of the biological condition (e.g., treatment positive response vs. neutral response vs. negative response or degrees of positive and / or negative response).
[0026] In an example, the LGM may be trained as follows:• Input:o Labeled metagenomic sequencing data (e.g., FASTQ, BAM, or VCF files) with known classifications (training data for a specific query task).• Process:o Maps labeled sequencing data into the biomarker embedding space.o Trains a supervised classifier using labeled embeddings.• Output: A trained LGM classifier.
[0027] 3. Prediction Phase (Queries to predict a patient’s biological condition): The trained LGM may predict a patient’s biological condition using the trained supervised classifier based on the patient’s clusters of DNA sequences. The LGM may receive an input query comprising unlabeled DNA sequences that are metagenomic sequenced from a genome or microbiome DNA sample of the patient. The patient’s unlabeled DNA sequences of the query may be mapped to the clusters of DNA sequences in the biomarker embedding space. The trained LGM may be executed to predict the patient’s biological condition using the trained supervised classifier based on the mapped clusters of DNA sequences in the biomarker embedding space, for example, to generate predictions for training / correlated and new / uncorrelated clusters of DNA sequences (e.g., with uncertainty quantification). When the clusters of DNA sequences are unknown (e.g., not grouped or uncorrelated with the classification labels of the biological condition) by the trained supervised classifier, generating a confidence score (e.g., an uncertainty estimate, which may be based on training accuracy of classification for known clusters) for the prediction of the patient’s biological condition.P-655803-PC
[0028] In an example, the LGM may predict as follows:• Input:o Unlabeled query sequence for prediction.• Process:o Maps the query sequence into the embedding space and predicts a classification or regression value.o Detects out-of-distribution sequences and assigns confidence scores to predictions.• Output: A prediction based on similarity to labeled samples in the embedding space, with an uncertainty estimate for novel DNA sequences.
[0029] Additional or alternative phases of generating the LGM may also be used.
[0030] Example LGM Features:• Separation of phases (1) Unsupervised Construction and (2) Supervised Learning: The biomarker library is generated independently and can be used for multiple prediction tasks.• Scalability: The more biomarker-based embeddings are injected into LGM, the more accurate its prediction and generalizable to various patients.• Multimodal Data Integration: Biomarkers may be extracted from a wide range of genetic sources, including for example, human, microbial, environmental, and / or plant genomes. • Cross-Domain Application: The LGM works to analyze both native host (human) genomic DNA and its microbiome DNA.
[0031] Additional or alternative features may also be provided by the LGM.
[0032] IMPROVEMENTS: Embodiments of the present invention may provide one or more of the following improvements to the state of the art:1. Discover Previously Undetectable Fragmented and Decentralized Groups of DNA Sequenceso Whereas metagenomic sequencing fragments DNA sequences thus losing information regarding their contextual co-occurrences, the LGM restores this data by detecting clusters of simultaneously occurring DNA sequences collectively responsible for group functionality of biological conditions.2. Compact Biomarker Embedding Space to Improve Computer Efficiencyo Embodiments of the invention provide an efficient and accurate LGM for detecting biomarkers in a compact biomarker embedded space based on anomalous cooccurrence patterns in donor DNA samples. The biomarker embedding space may beP-655803-PCa compact vector space to analyze and detect statistically significant clusters, efficiently flagging candidate DNA sequence groups to analyze for biological relevance, which are otherwise undetectable because there are too many DNA sequence (e.g., 4Ak k-mers) to consider by conventional techniques.o This improvement is further strengthened by the ability to incorporate relatively long DNA sequences (e.g., k-mers with large values of k, such as 30 or more).3. Generalization Across Genetic Data Types:o Unlike traditional classifiers, which require predefined gene annotations, the LGM may learns from any metagenomic sequencing data source, making it applicable to both known and novel biological datasets.4. Solving the Lack of Internal Sequence Order:o By embedding DNA sequence groups rather than relying on sequential ordering, LGM bypasses the challenges associated with unordered genetic material, providing a more robust method for detecting functional patterns.5. Scalable and Self-Improving Model:o The model may continuously integrate new biomarker knowledge, improving its predictive capabilities with increasing data volume.6. Embedding-Based Approach Overcomes Deep Learning Constraints:o Unlike deep learning models that rely on sequential inputs (such as Evo2-40B), the LGM may detect functionally significant biomarkers that are scattered across different genomic regions or across different metagenomic samples.7. Supports Both Supervised and Unsupervised Analysis:o The LGM can operate with labeled training data (supervised classification) or in an exploratory mode, identifying novel relationships between sequences (unsupervised clustering).8. Improved Personalized Medical Treatmento New biomarkers of decentralized groups of DNA sequences improves genetic detection of biological conditions (e.g., and its treatment responsive or non-responsive variants) correlated with those biomarkers. Improved prediction of biological conditions and its therapeutic response to treatments improves the efficacy of administering treatments to patients based on their treatment-correlated biomarkers, thus improving personalized medical care.
[0033] Improvements to Personalized Medicine Administered Based on Detecting Clusters of Co-Occurring DNA SequencesP-655803-PC
[0034] Once clusters of DNA sequences that co-occur with statistical significance are detected in genomic or microbiome donor(s) samples, those clusters may be correlated with the donor(s)’s biological condition(s), such as, therapeutic response to a treatment therefore, e.g., therapeutic response of an indication Y to a drug X. Those clusters of DNA sequences and the donor(s)’s biological condition may be correlated by training the LGM. The LGM may be trained to correlate any or particular threshold(s) of co-occurrence frequencies of a cluster’s DNA sequences with positive, neutral and / or negative (absolute or relative) treatment response and / or degrees of treatment response (e.g., of N-levels or a continuous scale of efficacy, which may measure either or both of positive response(s) and / or negative response(s) predicted individually or as a combined offset response). Any number of single or multiple clusters, types / levels / categories / degrees of therapeutic responses, side-effects, safety factors, treatment-administration factors (e.g., dosages, timings for administering, combinations of drugs, etc.), patient factors (e.g., demographics, diagnosis or disease(s)), iterations of analysis, etc., or combinations thereof may be used. The correspondences between any of these numbers may be one-to-one, one-to-many, many-to-one or many-to-many. In various embodiments, there may be an integer number of N levels of therapeutic response (e.g., degree ranges of therapeutic response, such as, chemical blood levels, and / or each with or without side-effects), a vector of NxM, NxP or MxP therapeutic responses or a matrix of NxMxP therapeutic responses (e.g., for multiple N levels of therapeutic response, M side-effects and / or P safety measures), M different clusters (e.g., individually or in combination) with a frequency of cooccurrence that is statistically significant (e.g., by any degree relative to the power-law distribution) in patients with therapeutic treatment response), Fm features characterizing different statistical e.g. co-occurrence patterns of cluster sequences, etc. The multiple N-levels of therapeutic response may be relative response measures (e.g., high, medium, low, no physiological benefit or Gaussian or other distribution range percentage levels), orN quantitative physiological ranges (e.g., ten absolute value ranges or percentage improvements in levels of biological sample including but not limited to blood, saliva or feces). Any combination or division of these characteristics may be used.
[0035] In an embodiment, after the correlation is trained, a new patient’s genomic or microbiome DNA sample sequences may be searched forthose treatment-correlated cluster(s) (e.g., of various trained degrees or threshold numbers of occurrences, such as, greater than or within a range of those detected in the training donor population). If detected, the patient may be predicted to correlate with the trained biological condition and / or its treatment response (e.g., preferential response of an indication to a drug) and may be administered the treatment (e.g., the drug). The new patient may thus experience improved efficacy compared to an average efficacy of random or non-P-655803-PCcorrelated patients (e.g., without the clusters(s) and / or with their occurrence numbers below the trained threshold(s)).
[0036] Embodiments of the invention train the LGM to predict the biological condition such as treatment response in patients based on the presence and statistical distribution of specific sequences of any sequenced segment of raw genomic or microbiome DNA. These genomic or microbiome DNA sequences may encode relatively short, functionally relevant nucleotide sequences, such as k-mers, that may correspond to coding regions, regulatory motifs, mobile elements, or other genomic features. By freeing analysis from the constraints of known genomic segments or microbes, the LGM can automatically detect biomarkers even in previously uncharacterized genome segments or microbial species. Additionally, unlike conventional approaches that are limited to short DNA sequences (e.g., k-mers with k<7 nucleotides), embodiments of the invention may provide a genetic analysis that can efficiently analyze long DNA sequences (e.g., k-mers with k>30 nucleotides) by a fast and compact power-law analysis that can tackle the previously computationally impossible search space (e.g., > 101830-mers). Experimental data suggest that such longer DNA sequences (e.g., in the range of 30-60 nucleotides) improve biomarker accuracy (e.g., shorter k-mers are too common, while longer k-mers are too rare, for accurate statistics). Thus, embodiments of the invention may improve the ability to detect longer uninterrupted and therefore more biologically-accurate biomarkers with improved descriptive power that were previously undiscoverable. Additionally, analyzing combinations of non-sequential, dispersed or decentralized clusters of DNA sequences, which compounds the already previously computationally impossible search space, are now also possible according to some embodiments due to the compact and efficient biomarker embedding space. Because genetic and microbial functional units are often scattered across multiple genomic segments or multiple microbes, such embodiments detect, and treatment benefits from, dispersed genetic interactions that traditional approaches would miss.
[0037] Embodiments of the invention may improve personalized medical treatment and patient safety by identifying and administering a treatment (e.g., drug x) selectively to patient(s) with decentralized clusters of DNA sequences trained or linked by the LGM to have a higher likelihood of treatment-response or a biological condition responsive to the treatment and / or a lower likelihood, fewer or reduced levels or degrees of side-effects to the treatment (e.g., improving overall benefit considering side-effects), e.g., compared to average populations and to patients without those clusters . For example, embodiments of the invention may identify and administer a treatment selectively to patient(s) in which the likelihood of toxic or detrimental effects are outweighed by therapeutically or prophylactically beneficial effects of administering the treatment. Embodiments of the invention may (a) administer a treatment to a first set of patients (e.g., a testing or training population) (e.g., basedP-655803-PCon one or more dosages, schedules and / or combinations with other drugs); (b) train the LGM to correlate DNA clusters to treatment-responsive patients by analyzing whether the patients maintain or later develop (e.g., an improved or worsened) disease or indication; and (c) administer the treatment to a second set of patients (e.g., a treatment population) predicted responsive by the LGM based on presence or absence (e.g., or threshold amounts) of those treatment-linked DNA clusters (e.g., administered at treatment-linked or optimized dosages, schedules and / or combinations with other drugs). Because the second set of patients’ DNA is correlated with preferential treatment-responsive DNA clusters, their treatment is predicted to cause an improvement to their biological conditions compared to the first set of patients. Additionally or alternatively, embodiments of the invention may identify, select, or stratify patients for patient cohorts who are predicted to be treatment-responsive (e.g., with drug x), or who need modified treatment, or with a higher therapeutic index to the treatment. Embodiments of the invention may identify, select, or stratify patients for administering the treatment in clinical trials or other studies with the treatment, thereby increasing the likelihood that the clinical trial or study will succeed and / or for optimizing or improving the treatment regimen, e.g., to select a drug’s dosages, dosing frequency, route of administration, etc. Conversely, patients without those treatment-linked DNA clusters (or with treatment-averse DNA clusters) may be predicted to not have the biological conditions and may not be administered the treatment for the biological conditions, may be administered an altered treatment (e.g., altered drug administration regimen, such as dosage, timing, drug combination, etc.), or may weigh against administering the treatment to the patient.
[0038] A treatment may be administered to those patients predicted by the LGM to be treatment-responsive (e.g., whose genomic or microbiome DNA comprise the one or more DNA clusters correlated with therapeutic response to the treatment). Additionally or alternatively, based on the LGM’s treatment-response prediction, embodiments of the invention may administer or predict therapeutic response to different treatment variants, such as response of indication Y to different drug x regimens (e.g., N dosages and / or schedules may be administered to N patient cohort groups predicted to have N respective therapeutic responses, indications and / or side-effects), administer or predict therapeutic response to a treatment such as indication Y to different combinations of drug x with other drugs (e.g., including combinations with immunomodulators, anti-inflammatory agents, corticosteroids, antibiotics, probiotics or other genomic or microbiome -modulating agents, and biologies targeting different pathways), stratify patients as responders and non-responders before administering the treatment, predict optimized timing or sequence of treatment administration, predict therapeutic response to the treatment in specific patient subgroups (such as treatment-naive vs. treatment-experienced patients, patients with early vs. late-stage disease, patients with comorbiditiesP-655803-PCor co-medications) to assess the biomarkers as a predictive or companion diagnostic in a clinical trial, predicting long-term health outcomes of administering the treatment (e.g., including relapse rates and remission duration) in biomarker-stratified cohorts (e.g., to support regulatory submissions or labeling for biomarker-enriched indications), predict and administer minimal effective dosage or maintenance dosage adjustments, predict therapeutic response to different treatment formulations, such as drug dosages and / or schedules and administer treatment formulations based thereon, e.g. to identify or reduce the likelihood of treatment-related adverse events or toxicity in biomarker-positive patients or reduce their exposure by not administering treatment to biomarker-negative patients, to guide retreatment, drug holiday, or step-down strategies based on biomarker-defmed likelihood of sustained remission, to tailor the route of administration (e.g., oral vs. parenteral) based on biomarker status (e.g., especially if gut microbiome influences drug x metabolism or absorption), predict treatment resistance or secondary non-response emerging in specific biomarker-defmed subgroups and administer the treatment or not based thereon. Accordingly, the LGM may predict a biological condition and / or its treatment response based on the DNA clusters biomarker space detected in biological sample(s) of donor(s), and administer the treatment to the donor(s) or new patients. Thus, a donor or patient predicted to have a biological condition correlated with therapeutic response to the treatment may be administered that treatment with improved efficacy compared to average efficacy of random patients.
[0039] Reference is made to Fig. 1, which schematically illustrates a system for executing a Large Genetic Model (LGM) for analyzing DNA sequences in microbiome or genomic DNA of one or more hosts to predict and personalize treatment for a biological condition, according to an embodiment of the invention.
[0040] The system of Fig. 1 may include a genetic sequencer 102 and / or a sequence analyzer 106. Units 102 and 106 may be implemented in one or more computerized devices as hardware and / or software units, for example, specifying instructions configured to be executed by one or more processors. One or more of units 102 and 106 may be implemented as separate devices or combined as an integrated device.
[0041] Genetic sequencer 102 may input DNA obtained from human or animal biological samples, such as, a sample of one or more real living host organisms and may output each organism’s genetic sequence including the host’s genetic information at one or more genetic loci, for example, a plurality of DNA sequences. In some embodiments, genetic sequencer 102 may input a biological sample, perform shotgun fragmentation (e.g., randomly breaking up the genome into small DNAP-655803-PCfragments), and sequence each fragment individually. Each single organism’s DNA sample may be sequenced for analysis as an individual set of DNA sequences or multiple organism’s DNA samples may be sequenced for analysis as a combined set of the host group’s DNA sequences.
[0042] Genetic sequencer 102 may store in one or more memorie(s) 114, and send to sequence analyzer 106 to receive and store in one or more memorie(s) 118, a plurality of DNA sequences (e.g., reads and / or k-mers), sequenced from DNA of one or more hosts, for example, each DNA sequence having a fixed or variable length of nucleotides (e.g., 100-150 nucleotides in each read and / or k nucleotides in each k-mer, such as, k=30-60).
[0043] Genetic sequencer 102 and sequence analyzer 106 may include one or more controller(s) or processor(s) 108 and 112, respectively, configured for executing operations and one or more memory unit(s) 114 and 118, respectively, configured for storing data such as genetic information or sequences and / or instructions (e.g., software) executable by a processor, for example for carrying out methods as disclosed herein. Processor(s) 108 and 112 may include, for example, a central processing unit (CPU), a digital signal processor (DSP), a microprocessor, a controller, a chip, a microchip, an integrated circuit (IC), or any other suitable multi-purpose or specific processor or controller. Processor(s) 108 and 112 may individually or collectively be configured to carry out embodiments of a method according to the present invention by for example executing software or code. Memory unit(s) 114 and 118 may include, for example, a random access memory (RAM), a dynamic RAM (DRAM), a flash memory, a volatile memory, a non-volatile memory, a cache memory, a buffer, a short term memory unit, a long term memory unit, or other suitable memory units or storage units. Genetic sequencer 102 and sequence analyzer 106 may include one or more input / output devices, such as output display 120 (e.g., such as a monitor or screen) for displaying to users results provided by sequence analyzer 106 and an input device 122 (e.g., such as a mouse, keyboard or touchscreen) for example to control the operations of the system and / or provide user input or feedback, such as, selecting one or more hosts, selecting input genetic sequences, etc.
[0044] Reference is made to Fig. 2, which schematically illustrates a portion of a DNA network 100 representing DNA sequences aggregated into clusters to generate biomarkers in a biomarker embedding space, according to an embodiment of the invention. In a training phase, the LGM is trained, e.g., using a supervised classifier, to correlate those biomarker clusters with classification labels of the DNA-donor’s biological condition. In a prediction (run-time) phase, the LGM executes the trained classifier to input a new patient’s DNA sequences (e.g., in a query prompt) and, based on its clusters, predict the patient’s biological condition. A portion of DNA network 100 is shown because it typically has too many nodes, for example, millions, to illustrate in its entirety. Data structures described herein may be stored in one or more memor(ies) (e.g., memory unit(s) 114P-655803-PCand / or 118 of Fig. 1), may be generated and controlled by one or more processor(s) (e.g., controller(s) or processor(s) 112 of a DNA sequence analyzer 106 of Fig. 1), and any visualizations, or data thereof may be displayed on one or more display(s) (e.g., output display 120 of Fig. 1 or any user display).
[0045] DNA network 100 comprises a plurality of nodes 101 representing a respective plurality of DNA sequences, each sequenced from a biological sample of one or more human or animal hosts. The plurality of nodes 101 may be pairwise connected by a plurality of edges 103. Each edge 103 may represent a co-occurrence of each pair of DNA sequences in a same sample, segment or sub-length of the microbiome DNA, such as, two reads in the same partial or whole string of DNA or two k-mers in the same read. Additionally or alternatively, different DNA sequences may co-occur in different samples, segments or sub-lengths of the DNA. Nodes 101 may be densely-connected by edges 103 (e.g., > 80% of nodes connected by edges), sparsely-connected by edges 103 (e.g., < 20% of nodes connected by edges) or moderately-connected by edges 103 (e.g., 20-80% of nodes connected by edges). Nodes 101 and / or edges may be filtered, e.g., to remove or pmne edges having extreme degree (e.g., > dmin and / or < dmax, where dmin and dmax are predefined minimum and maximum threshold degrees) or by machine learning.
[0046] Nodes 101 mayeachhaveadegreecountingthenumberofedgestowhichthenodeconnects. Nodes 101 may be bundled into distinct groups or clusters of the same or similar degree. Nodes 101 in the same cluster thus have the same or similar number of unique DNA sequences that co-occur with the DNA sequence represented by the node in the same sample, segment or sub-length of the microbiome DNA. Each bundled cluster may be evaluated to determine if it has sufficiently anomalous behavior to correlate with the biological condition of the host patient from which the DNA was sampled (e.g., being clinically validated to exhibit therapeutic response to a treatment for the biological condition). Those nodes 101 of the cluster that are inter-connected with a probability that is unlikely randomly statistical (e.g., too high or too low, indicating the nodes’ co-occurrence is likewise unlikely randomly statistical, but significant, and therefore likely correlated with the host’s biological condition). The anomaly condition may be, for example, that the internal connectivity for the cluster of nodes deviates from an expected internal connectivity in a network, such as, one that follows a power law distribution. Internal connectivity of nodes 101 by edges 103 maybe a measure of a number of nodes 101 or edges 103 in the same cluster, a density of nodes 101 or edges 103 in the network 100 or region thereof, a number of pre-defined partially or fully connected node clusters (e.g., triangles or polygons or straws or cliques of nodes) in the same cluster, a ratio of intra-cluster edges (internal to a cluster) to intercluster edges or overall edges (external to a cluster or in the entire network or region thereof), or derivative or deviation therefrom. Clusters of DNA sequences represented by nodes 101 in anomalous clusters may be used, together with patient biological conditions measured for their DNA’s hosts, toP-655803-PCgenerate a training dataset to train the LGMto correlate clusters of DNA sequences with their host(s)’ biological condition.
[0047] Due to the unmanageable number of different possible DNA sequences (e.g., k-mers) in DNA (e.g., 43030-mer combinations) and even larger number of groups of those DNA sequences (e.g., 24A3030-mer pair combinations), conventional models cannot analyze these volumes of genetic material to predict their biological effect in practical time. Embodiments of the invention may group a subset of “significant” groups or clusters of DNA sequence that exhibit anomalous co-occurrence patterns that deviate from a neutral, e.g., power law, distribution and so, their expression is unlikely randomly statistical, and therefore likely associated with a phenotypic property associated with a biological condition. The grouping of statistically significant clusters thus eliminates non-anomalous groups of genome DNA sequences, which are likely randomly statistical noise. Training the LGM only on the significant clusters of anomalous groups of genome DNA sequences makes it possible to train the machine learning model based on groups of DNA sequences in finite practical time. Eliminating insignificant k-mer groups also reduces the size of the search space from previously unmanageable volumes of all combinations of genome DNA sequences groups (e.g., 24A3030-mer pairs) to a manageable volume of only clustered groups of anomalous genome DNA sequences (e.g., an order of magnitude of k-mer groups equal to the number of distinct degrees in a network, which depends on how the groups are bundled, and is less than a linear function of the number of individual k-mers in the network). A reduction in training data size from an exponential to a linear function of the number of individual k-mers in the network reduces the storage size and increases training speed by orders of magnitude, thereby improving training efficiency of the LGM.
[0048] Reference is made to Fig. 3, which is a flowchart of a method for executing a Large Genetic Model (LGM) comprising a compact and efficient biomarker embedding space for converting DNA sequences into predictions of, and personalize treatments for, biological conditions, according to an embodiment of the invention. Operations described herein may be executed by one or more processor(s) (e.g., controller(s) or processor(s) 112 of a DNA sequence analyzer 106 of Fig. 1), data or data structures described herein may be stored in one or more memor(ies) (e.g., memory unit(s) 118 of a DNA sequence analyzer 106 of Fig. 1), and any visualizations, or data may be displayed on one or more display(s) (e.g., output display 120 of Fig. 1 or any user display).
[0049] In operation 310, a process or processor may receive and store in memorie(s) a large-scale database or repository of sequenced DNA (e.g. reads) metagenomic sequenced from genome or microbiome DNA samples of many samples from many host organisms in a single or across multiple species (e.g., FASTQ, BAM, or VCF files). The DNA sequence reads may be sequenced byP-655803-PCmetagenomic sequencing also known as "shotgun sequencing," metatranscriptomic sequencing or another sequencing technique (e.g., at genetic sequencer 102 of Fig. 1).
[0050] In operation 320, a process or processor may extract a plurality of DNA sequences (e.g., k-mers) from the database of sequenced DNA using e.g. a sliding window encompassing a predetermined length k of sequential nucleotides. Those extracted DNA sequences may be grouped into clusters using unsupervised learning. Each cluster comprises multiple of the plurality of DNA sequences that co-occur with statistical significance (e.g., non-neutral and / or non-random distribution of co-occurrence of DNA sequences in the cluster, such as higher (or lower) occurrence frequencies than a power-law distribution or groups with combinations of DNA sequences with the highest or above-threshold occurrence frequencies. In an embodiment, the plurality of DNA sequences may be grouped into the clusters by generating a DNA network (e.g., 100 of Fig. 1) comprising a plurality of nodes (e.g., 101 of Fig. 1) representing the plurality of DNA sequences and a plurality of edges (e.g., 103 of Fig. 1) representing a co-occurrence of each pair of the DNA sequences in a same segment of the DNA sequences, generating distinct clusters of the plurality of DNA sequences represented by nodes based on the node’s degree, the degree of each node quantifying a number of edges that contain the node indicating a number of unique DNA sequences that co-occur with the DNA sequence represented by the node in a same segment of the DNA sequences, and selecting the clusters of DNA sequences that each have an internal connectivity between nodes in the same cluster that satisfies a preferential condition indicating the DNA sequences of that cluster cooccur with a probability that is statistically significant. The internal connectivity condition may be satisfied e.g., if the internal connectivity for the cluster of nodes deviates from an expected internal connectivity in a network that follows a power law distribution.
[0051] In operation 330, a process or processor may encode the clusters of DNA sequences as biomarkers embedded in a biomarker embedding space. The biomarker embedding space may be a compact vector space of statistically significant biomarkers of clusters of DNA sequence. For example, each sample may be represented as an N-length vector of whether or not they contain each of the N biomarkers or an M-length vector of whether or not they contain each of the M DNA sequences of a cluster. Additionally or alternatively, the biomarker embedding space may be encoded by an auto-encoder executed on the clusters of DNA sequences to produce a significantly reduced number of features F, and then each new sample is represented by a vector of size F (F « N, M) to accelerate the model training and inference and reduce overfitting (e.g., especially when there is less data).
[0052] Operations 310-330 may be repeated by dynamically updating the LGM with each new update adding new sequenced DNA (e.g., metagenomic sequenced from new DNA samples of theP-655803-PCsame or different one or more hosts) to the database. Operations 310-330 may be repeated by receiving a new plurality of DNA sequences (operations 310), grouping them into new clusters (operations 320), and encoding the new clusters as new biomarkers (operations 330) for the new sequenced DNA to automatically expand the biomarker embedding space with new clusters of DNA sequences to improve biomarker identification and classification accuracy based on those biomarker overtime.
[0053] In operation 340, a process or processor may training the LGM by receiving a plurality of labeled DNA sequences that are metagenomic sequenced from multiple genome or microbiome DNA samples of multiple hosts, and are labeled with a genomic classification of a biological condition of the respective hosts, mapping the plurality of labeled DNA sequences into the biomarker embedding space e.g., indicating whether or not the clusters of DNA sequences are contained therein (and / or a frequency or distribution thereof), and training a supervised classifier to correlate the clusters of DNA sequences in the biomarker embedding space with the classification labels of the biological condition (e.g., treatment positive response vs. neutral response vs. negative response or degrees of positive and / or negative response).
[0054] In operation 350, a process or processor may executing the trained LGM to predict a patient’s biological condition by receiving a LGM query comprising one or more unlabeled DNA sequences that are metagenomic sequenced from a genome or microbiome DNA sample of the patient, mapping the patient’s unlabeled DNA sequences of the query to one or more of the clusters of DNA sequences in the biomarker embedding space, and executing the trained LGM to predict the patient’s biological condition using the trained supervised classifier based on the one or more mapped clusters of DNA sequences in the biomarker embedding space (e.g., to generate predictions for training / correlated and new / uncorrelated clusters of DNA sequences with uncertainty quantification). When the one or more clusters of DNA sequences are unknown by the trained supervised classifier, the process or processor may generate a confidence score (e.g., uncertainty estimate based on training accuracy of classification for known clusters) for the prediction of the patient’s biological condition.
[0055] In operation 360, a treatment for the biological condition may be administered to the human patient predicted to have the biological condition for personalized DNA-based medical treatment. The treatment may be administered by a doctor or other medical personnel, the patient, or automatically by an automated medical delivery device triggered by a treatment program, the prediction of operation 350, and / or an intermediary person or system confirming the prediction results. Because the patient is predicted to have the biological condition responsive to the treatment, administering the treatment to the patient results in an improvement to the biological condition and / orP-655803-PCtreatment efficacy e.g., compared to average populations, control groups, and / or to patients without those biomarkers (or with fewer or different combinations of their validated groups).
[0056] Other or different operations or orders of operations may be used and some operations may be omitted repeated, e.g., operations 310-360 may be repeated (e.g., periodically, upon receiving new data, etc.) to retrain the model with new DNA sequences.
[0057] Improved Cryptographic Security: Due to the private nature of genetic information, patient, host and donor DNA data may be encrypted to provide secure storage, secure transmission and allow distributed computations thereon across multiple servers without (or with reduced) security risk (e.g., to execute inter-server communication or reduced risk of a security breach at one of multiple servers).
[0058] In some embodiments of the invention, a cryptography application may automatically encrypt the private or sensitive DNA sequences of the patients, hosts and / or donors (without human observation or intervention that would vitiate privacy). The encrypted data may then be transmitted securely over a network from a first to a second computing device without the second or either devices accessing the unencrypted data. In an embodiment, the system of Fig. 1 may provide secure cryptographic communication between a first computer terminal (e.g., genetic sequencer 102) and a second computer terminal (e.g., sequence analyzer 106) over a communication channel therebetween. The first computer terminal may receive plaintext(s) of a patient’s private DNA sequences, transform the plaintext s) into a plurality of DNA sequences m (e.g., k-mers), encode each of the DNA sequences m under an encryption scheme (e.g., homomorphic encryption) to produce one or more ciphertext(s) ct. The first computer terminal may then transmit the ciphertext(s) ct to the second computer terminal over a communication channel, without exposing the DNA sequences m themselves to the second computer terminal. The second computer terminal may execute a cryptography circuit, such as, a homomorphic encryption circuit, on the ciphertext(s) ct to analyze and operate on the DNA sequences (e.g., according to embodiments of the invention) under encryption, without exposing any of the underlining unencrypted data. The private decryption key may securely stored, e.g., only at a secure user’s (e.g., patient, host or donor’s) device and not at the first and / or second computer terminals.
[0059] CONCLUSION
[0060] The Large Genetic Model (LGM) may separate unsupervised biomarker extraction from supervised prediction, enabling a scalable and reusable genomic intelligence framework. ByP-655803-PCembedding genetic information into a structured biomarker space, LGM overcomes the challenges posed by unordered genetic sequences, providing a scalable, continuously improving system for genomic classification and prediction across multiple domains.
[0061] Embodiments of the present invention bridge the gap between raw genetic sequencing data and interpretable biological insights, offering an adaptive, scalable, and biomarker-driven approach to genomic intelligence applicable to both microbial and host genomic data.
[0062] Hosts, donors or patients discussed herein are living human or animal organisms that provide biological samples sequenced into the microbiome or genome DNA sequences. One or more processor(s) (e.g., processor(s) 108 of Fig. 1) may sequence the organism’s DNA obtained from its biological sample to identify the genetic constituents of the DNA. The processor(s) used to sequence DNA may be the same processor(s) or different processor(s) (e.g., external service provider) used to analyze and model the data (e.g., processor(s) 112 of Fig. 1). In some embodiments, metagenomic or "shotgun" sequencing may be used (e.g., providing rich data, but with no semantic association to the microbes themselves).
[0063] Although embodiments of the invention describe k-mers, any other continuous or non-continuous genomic or microbiome DNA sequences may alternatively or additionally be used.
[0064] Embodiments of the invention encompass a host’s microbiome DNA alone, microbiome DNA combined with the host’s genomic DNA, or the host’s genomic DNA alone, for example, sequenced from any biological host sample or specimen or mixture thereof including, for example, bile, blood, tissue, saliva, urine, stem-cells, tumors, biopsies, etc. DNA may be sampled from a single host sample or a DNA mixture sequenced from multiple hosts or multiple samples from a single host.
[0065] As used herein, “DNA” may refer to Deoxyribonucleic acid and / or Ribonucleic acid (RNA), such as, messenger RNA (mRNA), transfer RNA (tRNA), and / or ribosomal RNA (rRNA), complementary DNA (cDNA), cell-free DNA, viral genomes, fungal genomes, bacterial genomes, archaeal genomes, host-derived nucleic acids, and any other nucleic-acid-based biological sequence obtained from a biological sample, whether via metagenomic, metatranscriptomic, or other sequencing methodologies. RNA contains adenine (A), cytosine (C), guanine (G), or uracil (U) nucleotide bases, whereas DNA contains adenine (A), cytosine (C), guanine (G) and thymine (T). DNA sequence data structures may include, for example, one or more vectors, scalar values, functions, sequences, sets, matrices, tables, lists, arrays, and / or other data structures, representing biological genetic material including one or more bases, nucleotides, genes, alleles, codons or other generic material.P-655803-PC
[0066] As used herein, “therapeutic response” or “treatment response” indicates a non-neutral and / or non-random response to a medical treatment as verified in one or more physiological test(s) that may be measured e.g., once or over time. In various embodiments, therapeutic response may indicate a completely positive therapeutic benefit or a partially positive therapeutic benefit and a partially negative therapeutic detriment, such as with side-effects. Therapeutic response may be an overall positive aggregated, e.g., where positive effects predominate over negative effects.
[0067] As used herein, “drug x” may refer for example to any of the following drugs or equivalent molecule(s) and / or molecular variants with the same or equivalent operating mechanism or biological effect that may be administered to treat biological conditions of the following respective indications Y: Vedolizumab (Entyvio®) for Crohn's Disease (CD), Vedolizumab (Entyvio®) for Ulcerative Colitis (UC), Risankizumab (Skyrizi®) for Crohn's Disease (CD), Adalimumab (Humira®) for Crohn's Disease (CD), Infliximab (Remicade®) for Crohn's Disease (CD) Ustekinumab (Stelara®) for Crohn's Disease (CD), Upadacitinib (Rinvoq®) for Ulcerative Colitis (UC), and Certolizumab Pegol (Cimzia®) for Crohn's Disease (CD). Other or different combinations of drugs and indications may also be used.
[0068] Variables used herein, e.g., k, F, M, N, etc. may be integers.
[0069] As used herein, “practical,” “possible,” or “finite” time, computations, problems or tasks may refer to those that can be executed with a standard computer system on the order of less than a few (e.g., one or two) hours or days, and “impractical” or “impossible” time, computations, problems or tasks may refer to those that can only be executed with a standard computer system on the order of greater than a few (e.g., one or two) hours or days.
[0070] In the foregoing description, various aspects of the present invention are described. For purposes of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the present invention. However, it will also be apparent to persons of ordinary skill in the art that the present invention may be practiced without the specific details presented herein. Furthermore, well known features may be omitted or simplified in order not to obscure the present invention.
[0071] Unless specifically stated otherwise, it is appreciated that throughout the specification discussions utilizing terms such as "processing," "computing," "calculating," "determining," or the like, refer to the action and / or processes of a computer or computing system, or similar electronic computing device, that manipulates and / or transforms data represented as physical, such as electronic, quantities within the computing system's registers and / or memories into other data similarly represented as physical quantities within the computing system's memories, registers or other such information storage, transmission or display devices.P-655803-PC
[0072] Embodiments of the invention may include an article such as a computer or processor readable non-transitory storage medium, such as for example a memory, a disk drive, or a USB flash memory device (e.g., memory unit(s) 114 and / or 118 of Fig. 1) encoding, including or storing instructions, e.g., computer-executable instructions, which when executed by a processor or controller (e.g., controller(s) or processor(s) 108 and / or 112 of Fig. 1), cause the processor or controller to carry out methods disclosed herein.
[0073] Different embodiments are disclosed herein. Features of certain embodiments may be combined with features of other embodiments; thus certain embodiments may be combinations of features of multiple embodiments.
[0074] The foregoing description of the embodiments of the invention has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise form disclosed. It should be appreciated by persons of ordinary skill in the art that many modifications, variations, substitutions, changes, and equivalents are possible in light of the above teaching. For example, it should be appreciated that sign conventions are equivalent and that embodiments of the invention in which values are above a lower bound or threshold are equivalent to embodiments of the invention in which values are below an upper bound or threshold, since the difference is a mere convention of sign. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of the invention.
Claims
P-655803-PCCLAIMS1. A method for constructing a large genetic model (LGM) for genetic prediction, comprising:receiving a database of sequenced DNA metagenomic sequenced from genome or microbiome DNA samples of multiple hosts;extracting a plurality of DNA sequences from the database of sequenced DNA; grouping the plurality of DNA sequences into clusters using unsupervised learning, wherein each cluster comprises multiple of the plurality of DNA sequences that co-occur with statistical significance; andencoding the clusters of DNA sequences as biomarkers embedded in a biomarker embedding space.
2. The method of Claim 1 comprising dynamically updating the LGM with each new update adding new sequenced to the database by repeating the method of Claim 1 for the new sequenced DNA.
3. The method of Claim 1 comprising training the LGM, comprising:receiving a plurality of labeled DNA sequences that are metagenomic sequenced from multiple genome or microbiome DNA samples of multiple hosts, and are labeled with a genomic classification of a biological condition of the respective hosts;mapping the plurality of labeled DNA sequences into the biomarker embedding space; andtraining a supervised classifier to correlate the clusters of DNA sequences in the biomarker embedding space with the classification labels of the biological condition.
4. The method of Claim 3 comprising executing the trained LGM to predict a patient’s biological condition, the method comprising:receiving a LGM query comprising one or more unlabeled DNA sequences that are metagenomic sequenced from a genome or microbiome DNA sample of the patient;mapping the patient’s unlabeled DNA sequences of the query to one or more of the clusters of DNA sequences in the biomarker embedding space; andexecuting the trained LGM to predict the patient’s biological condition using the trained supervised classifier based on the one or more mapped clusters of DNA sequences in the biomarker embedding space.P-655803-PC5. The method of Claim 4 comprising, when the one or more clusters of DNA sequences are unknown by the trained supervised classifier, generating a confidence score for the prediction of the patient’s biological condition.
6. The method of Claim 1, wherein the plurality of DNA sequences are grouped into the clusters by:generating a DNA network comprising a plurality of nodes representing the plurality of DNA sequences and a plurality of edges representing a co-occurrence of each pair of the DNA sequences in a same segment of the DNA sequences;generating distinct clusters of the plurality of DNA sequences represented by nodes based on the node’s degree, the degree of each node quantifying a number of edges that contain the node indicating a number of unique DNA sequences that co-occur with the DNA sequence represented by the node in a same segment of the DNA sequences; and selecting the clusters of DNA sequences that each have an internal connectivity between nodes in the same cluster that satisfies a preferential condition indicating the DNA sequences of that cluster co-occur with a probability that is statistically significant.
7. The method of Claim 6, wherein the internal connectivity condition is satisfied if the internal connectivity for the cluster of nodes deviates from an expected internal connectivity in a network that follows a power law distribution.
8. The method of Claim 4 comprising administering a treatment for the biological condition to the patient predicted to have the biological condition.
9. The method of Claim 4 comprising:training the supervised classifier to correlate the clusters of DNA sequences in the biomarker embedding space with the classification labels of multiple levels of therapeutic response to a treatment for the biological condition;predicting the patient has one of the multiple levels of therapeutic response to the treatment upon detecting the patient’s mapped clusters of DNA sequences are correlated with the therapeutic response level; andadministering the treatment to the patient in one or multiple regimens associated with the therapeutic response level.
10. A system for constructing a large genetic model (LGM) for genomic prediction, the system comprising:one or more memories configured to store the database of DNA sequenced from the genome or microbiome DNA samples; andone or more processors configured to perform the method of any of claims 1-9.P-655803-PC11. The system of Claim 10 comprising a genetic sequencer configured to process the genome or microbiome DNA samples of the multiple hosts to generate the database of sequenced DNA.
12. A non -transitory computer-readable storage medium storing instructions thereon that, when executed by a processor, cause the processor to perform the method of any of claims 1-9.
13. A non-transitory computer-readable medium storing instructions that, when executed by a processor, perform genomic classification, the instructions comprising:constructing a biomarker embedding space from metagenomic sequencing data; receiving a labeled dataset and a query sequence;mapping sequencing reads into the biomarker embedding space; andapplying supervised learning to classify the query sequence, with an adaptive mechanism for integrating newly sequenced k-mers.