Methods and systems for analyzing low-input samples
By processing low-quality and low-volume biological samples through fragmentation and hybridization with capture probes, the method achieves accurate genomic and epigenomic analysis, addressing the limitations of traditional methods in disease identification and treatment prediction.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NEOGENOMICS LAB
- Filing Date
- 2025-11-13
- Publication Date
- 2026-05-21
AI Technical Summary
Genomic analysis methods require large quantities of high-quality biological material, which is often unavailable or of poor quality, limiting accurate disease identification and treatment prediction, especially in cases like breast cancer where tissue samples are scarce.
A method involving fragmentation incubation length and washing steps to process small volumes of low-quality biological samples, followed by PCR amplification and hybridization with capture probes to enrich nucleic acid samples, enabling analysis of low-input and low-quality samples.
Enables accurate genomic and epigenomic analysis with low error rates, broad characterization of genetic data, and efficient integration of multiple biomarkers, overcoming limitations of traditional assays.
Smart Images

Figure US2025055399_21052026_PF_FP_ABST
Abstract
Description
Docket No.: 69246-750.601METHODS AND SYSTEMS FOR ANALYZING LOW-INPUT SAMPLESCRO SS-REFERENCE
[0001] This international patent application claims the benefit of U.S. Provisional Patent Application No. 63 / 721,350, filedNovember 15, 2024, which is incorporated herein by reference in its entirety.BACKGROUND
[0002] Many diseases and treatments for diseases can have a genomic basis. Genomic analysis of biological samples, such as genetic assays or epigenetic assays, can be used to determine one or more aspects of a disease, such as identifying a disease, or predicting an effective treatment for a disease. Genomic analysis can be performed on a biological sample of a subject. The genomic analysis of the biological sample can be used to identify a disease of the subject or predict a treatment for a disease of a subject, or both.SUMMARY
[0003] Genomic analysis of biological samples can include genetic assays or epigenetic assays. Genetic assays and epigenetic assays often require a minimum quantity or volume of biological sample material to successfully generate genetic or epigenetic analysis data, much less generate such analysis data with sufficient accuracy and minimal error for determining, identifying, or accurately predicting disease or effective treatments for disease of a subject. Generally, genetic sequencing assays such as next generation sequencing (NGS), or epigenetic assays such as whole exome sequencing (WES) and whole transcriptome sequencing (WTS) may require relatively large quantities of biological material for analysis. Often, hundreds of nanograms, or micrograms, of biological material from a subject may be required for obtaining sufficiently accurate and low-error data from genetic assays or epigenetic assays. This requirement for hundreds of nanograms or micrograms of biological material may not be possible to obtain from a tissue of the subject, such as a solid tumor of the subject. Often, there is little tissue available for analysis for some subjects such as breast cancer patients. In these cases, core needle biopsy may yield as little as lOng of biological material, which may not be a sufficient sample volume on which to perform genomic analysis assays. This issue often severely limits patient access to lifesaving testing, essential to clinical treatment of diseases such as cancer. The present disclosure provides methods and systems for analyzing and characterizing genetic data or epigenetic data, or both, obtained from low amounts of biological sample materials. The methodsDocket No.: 69246-750.601and systems described herein successfully analyze low amounts of biological sample materials with sufficient accuracy and low error to determine, identify, or accurately predict disease or effective treatments for disease of a subject. Applicant has recognized that fragmentation incubation length or associated washing steps, or both, in sample preparation as described in methods and systems herein may be utilized to successfully analyze these small biological sample volumes with sufficient accuracy and low error.
[0004] In addition to genetic assays and genomic assays often requiring a relatively large amount of biological material to successfully generate accurate and high-quality analysis data, these genetic assays and genomic assays often require that the biological material assayed to be high-quality biological material samples. Generally, genetic sequencing assays such as next generation sequencing (NGS) or epigenetic assays such as whole exome sequencing (WES) and whole transcriptome sequencing (WTS) may not be able to generate assay results such as analysis data with low-quality biological material samples, much less generate such analysis data with sufficient accuracy and minimal error for determining, identifying, or accurately predicting disease or effective treatments for disease of a subject. Low-quality samples can include samples that are degraded, such as samples having nucleic acid molecules with cytosine deamination, oxidative damage, nicks, gabs, abasic sites, or non-uniform ends. This degradation often results in large errors in genomic analysis, such as from chimeric reads from annealing overhangs and sequencing artifacts from introduction of erroneous or damaged bases during end repair steps, for example. The methods and systems described herein successfully analyze low-quality biological sample materials with sufficient accuracy and low error to determine, identify, or accurately predict disease or effective treatments for disease of a subject. Applicant has recognized that fragmentation incubation length or associated washing steps, or both, in sample preparation as described in methods and systems herein may be utilized to successfully analyze these low-quality biological samples with sufficient accuracy and low error.
[0005] Another issue with many genomic analysis assays and methods is that they are often limited in scope to hundreds of clinically relevant genes. This limits assessment of broad biomarkers that may affect the entire genome, such as tumor mutational burden (TMB), micro satellite instability (MSI), homologous recombination deficiency (HRD), splice variants, fusions, and disease state expression signatures. Further, identification of new biomarkers and including new and emerging targets is accordingly limited. These limitations often necessitate separate assays that are specialized and limited in scope, such as solid tumor assays and myeloid assays emphasizing a subset of genetic material. These separate assays are often expensive and difficult to coordinate for patients and healthcare providers. Data resulting from these separateDocket No.: 69246-750.601assays is often difficult to integrate to obtain an accurate and holistic analysis of the genome of a subject. Applicant has recognized that the methods and systems described herein broadly characterize genomic data such as genetic assay data and epigenetic assay data of a subject. This broad characterization by the methods and systems as described herein often results in analysis of assay data that is more efficient and accurate than analyzing assay data from multiple separate assays.
[0006] In an aspect, the present disclosure provides a method of preparing an enriched nucleic acid sample, the method comprising: obtaining a biological sample obtained or derived from a subject, wherein said biological sample is about 10 nanograms to about 100 nanograms; processing said biological sample to provide a plurality of single stranded deoxyribonucleic acid (ssDNA) fragments; amplifying said plurality of ssDNA fragments in two or more cycles of polymerase chain reaction (PCR), thereby producing a sequencing library; and contacting said sequencing library with a pool of capture probes, wherein said pool of capture probes comprises capture probes that are capable of hybridizing to at least a portion of an exome, capture probes that are capable of hybridizing to at least a portion of a biomarker, or capture probes that are capable of hybridizing to a single nucleotide polymorphism (SNP), or a combination thereof, to produce said enriched nucleic acid sample.
[0007] In some embodiments, the sequencing library comprises a double -stranded deoxyribonucleic acid (dsDNA) library.
[0008] In some embodiments, the biological sample is about 10 nanograms.
[0009] In some embodiments, the biological sample is about 25 nanograms.
[0010] In some embodiments, the biological sample is about 50 nanograms.
[0011] In some embodiments, the biological sample is about 100 nanograms.
[0012] In some embodiments, the biological sample comprises two or more sample types.
[0013] In some embodiments, the two or more sample types comprises two or more of a Formalin-Fixed Paraffin Embedded (FFPE) sample, a fine needle aspirate sample (FNA), a blood sample, a plasma sample, a serum sample, or a bone marrow sample.
[0014] In some embodiments, the biological sample comprises a Formalin -Fixed Paraffin Embedded (FFPE) sample, a fine needle aspirate sample (FNA), a blood sample, a plasma sample, a serum sample, or a bone marrow sample.
[0015] In some embodiments, the biological sample is said FFPE sample.
[0016] In some embodiments, the biological sample is said bone marrow sample.
[0017] In some embodiments, the biological sample is said blood sample, wherein said blood sample is a peripheral blood sample.Docket No.: 69246-750.601
[0018] In some embodiments, the biological sample is said peripheral blood sample.
[0019] In some embodiments, the biological sample comprises a cell line.
[0020] In some embodiments, the cell line is commercially available.
[0021] In some embodiments, the biological sample is obtained or derived from a solid tumor of the subject.
[0022] In some embodiments, the biological sample is obtained or derived from a biopsy.
[0023] In some embodiments, the biopsy is a liquid biopsy.
[0024] In some embodiments, the biopsy comprises a fine needle aspirate or a core needle biopsy.
[0025] In some embodiments, the biopsy is said fine needle aspirate.
[0026] In some embodiments, the biopsy is said core needle biopsy.
[0027] In some embodiments, the biological sample comprises ribonucleic acid (RNA), deoxyribonucleic acid (DNA), or a combination thereof.
[0028] In some embodiments, the biological sample comprises said RNA.
[0029] In some embodiments, the RNA is cell-free RNA.
[0030] In some embodiments, the RNA is exosomal RNA.
[0031] In some embodiments, the RNA is coding RNA (cRNA) or non -coding RNA (ncRNA).
[0032] In some embodiments, the biological sample comprises said DNA.
[0033] In some embodiments, the DNA is cell-free DNA or DNA from cells.
[0034] In some embodiments, the DNA is said ssDNA, double -stranded DNA (dsDNA), or a combination thereof.
[0035] In some embodiments, the DNA comprises said ssDNA.
[0036] In some embodiments, the DNA comprises said dsDNA.
[0037] In some embodiments, the DNA comprises one or more DNA fragments comprising a nick.
[0038] In some embodiments, the DNA comprises end-damaged DNA.
[0039] In some embodiments, the DNA comprises single stranded DNA overhangs on a double -stranded DNA.
[0040] In some embodiments, the amplifying comprises three or more, four or more, five or more, six or more, seven or more, eight or more, or nine or more cycles of PCR.
[0041] In some embodiments, the amplifying comprises eight cycles of PCR.
[0042] In some embodiments, the amplifying comprises five to nine cycles of PCR.
[0043] In some embodiments, the amplifying comprises six to nine cycles of PCR.
[0044] In some embodiments, the amplifying comprises six to eight cycles of PCR.Docket No.: 69246-750.601
[0045] In some embodiments, the amplifying further comprises ligating one or more adapters to said plurality of ssDNA fragments.
[0046] In some embodiments, the one or more adapters comprise directional adapters.
[0047] In some embodiments, the one or more adapters comprise non-directional adapters.
[0048] In some embodiments, the one or more adapters comprise Unique Molecular Index (UMI) adapters.
[0049] In some embodiments, the method further comprises ligating one or more UMI adapters to the at least subset of the plurality of ssDNA fragments prior to the amplification in (c).
[0050] In some embodiments, the method further comprises ligating one or more UMI adapters to the at least subset of the plurality of ssDNA fragments after the amplification in (c).
[0051] In some embodiments, the amplification further comprises providing a Unique Molecular Index (UMI) sequence to said ssDNA fragments.
[0052] In some embodiments, the UMI sequence is provided via primer extension or via PCR.
[0053] In some embodiments, the UMI sequence comprises a length of about 5 nucleotides to about 15 nucleotides.
[0054] In some embodiments, the UMI sequence comprises a length of about 9 nucleotides.
[0055] In some embodiments, the pool of capture probes comprises at least two of said capture probes that are capable of hybridizing to at least a portion of an exome, said capture probes that are capable of hybridizing to at least a portion of a biomarker, or said capture probes that are capable of hybridizing to a single nucleotide polymorphism (SNP).
[0056] In some embodiments, the pool of capture probes comprises said capture probes that are capable of hybridizing to at least a portion of an exome, said capture probes that are capable of hybridizing to at least a portion of a biomarker, and said capture probes that are capable of hybridizing to a single nucleotide polymorphism (SNP).
[0057] In some embodiments, the exome is a human exome.
[0058] In some embodiments, said pool of capture probes comprises at least two of said capture probes that are capable of hybridizing to at least a portion of a transcriptome.
[0059] In some embodiments, the transcriptome is a human transcriptome.
[0060] In some embodiments, said pool of capture probes comprises said capture probes that are capable of hybridizing to at least a portion of a transcriptome.
[0061] In some embodiments, said transcriptome is a human transcriptome.
[0062] In some embodiments, the biomarker comprises a plurality of biomarkers.
[0063] In some embodiments, the plurality of biomarkers comprises at least 100 biomarkers, at least 200 biomarkers, at least 300 biomarkers, at least 400 biomarkers, at least 500 biomarkers, atDocket No.: 69246-750.601least 600 biomarkers, at least 700 biomarkers, at least 800 biomarkers, at least 900 biomarkers, at least 1,000 biomarkers, at least 1,100 biomarkers, at least 1,200 biomarkers, at least 1,300 biomarkers, at least 1,400 biomarkers, or at least 1,500 biomarkers.
[0064] In some embodiments, the plurality of biomarkers comprises about 500 biomarkers.
[0065] In some embodiments, the plurality of biomarkers comprises about 1,100 biomarkers.
[0066] In some embodiments, the plurality of biomarkers comprises between about 20,000 biomarkers and about 25,000 biomarkers.
[0067] In some embodiments, the plurality of biomarkers comprises at least about 50,000 biomarkers.
[0068] In some embodiments, the plurality of biomarkers comprises at least about 65,000 biomarkers.
[0069] In some embodiments, the plurality of biomarkers comprises at least about 75,000 biomarkers.
[0070] In some embodiments, the plurality of biomarkers comprises at least about 100,000 biomarkers.
[0071] In some embodiments, the plurality of biomarkers comprises at least about 200,000 biomarkers.
[0072] In some embodiments, the plurality of biomarkers is associated with a disease or disorder.
[0073] In some embodiments, the disease or disorder comprises cancer.
[0074] In some embodiments, the plurality of biomarkers is associated with solid tumor cancer.
[0075] In some embodiments, the solid tumor cancer comprises carcinoma, sarcoma, or lymphoma.
[0076] In some embodiments, the plurality of biomarkers is associated with a hematologic malignancy.
[0077] In some embodiments, the method further comprises amplifying said plurality of ssDNA fragments subsequent to (d).
[0078] In some embodiments, the amplifying said plurality of ssDNA fragments subsequent to (d) comprises polymerase chain reaction (PCR).
[0079] In some embodiments, the amplifying comprises isothermal amplification, rolling circle amplification (RCA), or strand displacement amplification.
[0080] In some embodiments, the method further comprises sequencing said enriched nucleic acid sample.
[0081] In some embodiments, the sequencing comprises next generation sequencing (NGS).Docket No.: 69246-750.601
[0082] In some embodiments, the sequencing comprises DNA sequencing, RNA sequencing, whole-genome sequencing, whole-ex ome sequencing, targeted sequencing, bisulfite sequencing, or enzymatic sequencing.
[0083] In some embodiments, the sequencing comprises use of Illumina™ sequencing, Ultima™ sequencing, Ion Torrent™ sequencing, or a combination thereof.
[0084] In some embodiments, the method further comprises computer processing said enriched nucleic acid sample.
[0085] In some embodiments, the method further comprises analyzing data from said sequencing utilizing one or more computer processors.
[0086] In some embodiments, analyzing said sequencing data comprises determining genetic information, epigenetic information, or both, of said biomarkers.
[0087] In some embodiments, the genetic information or said epigenetic information comprises one or more of : tumor mutational burden (TMB), microsatellite instability (MSI), copy number variation (CNV), loss of heterozygosity (LoH), or human leukocyte antigen (HL A) identification, or any combination thereof.
[0088] In another aspect, the present disclosure provides a system comprising one or more processors and a memory operatively coupled to said one or more processors, wherein said one or more processors are individually or collectively programmed to perform the method.
[0089] In yet another aspect, the present disclosure provides a composition, comprising a pool of capture probes, wherein said pool of capture probes comprises capture probes that are capable of hybridizing to at least a portion of an ex ome, capture probes that are capable of hybridizing to at least a portion of a biomarker, or capture probes that are capable of hybridizing to a single nucleotide polymorphism (SNP), or a combination thereof.
[0090] In some embodiments, the capture probes that are capable of hybridizing to said SNP comprise customized capture probes.
[0091] In some embodiments, the pool of capture probes comprises at least two of said capture probes that are capable of hybridizing to at least a portion of an exome, said capture probes that are capable of hybridizing to at least a portion of a biomarker, or said capture probes that are capable of hybridizing to a single nucleotide polymorphism (SNP).
[0092] In some embodiments, the pool of capture probes comprises said capture probes that are capable of hybridizing to at least a portion of an exome, said capture probes that are capable of hybridizing to at least a portion of a biomarker, and said capture probes that are capable of hybridizing to a single nucleotide polymorphism (SNP).Docket No.: 69246-750.601
[0093] In yet another aspect, the present disclosure provides a kit comprising the composition; and instructions for use according to any one of the methods described herein.
[0094] Another aspect of the present disclosure provides a non -transitory computer readable medium comprising machine executable code that, upon execution by one or more computer processors, implements any of the methods above or elsewhere herein.
[0095] Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein.
[0096] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure.Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE
[0097] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicatedto be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0098] The novel features of the disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:
[0099] FIG. 1 illustrates an example of library preparation according to methods as described herein.Docket No.: 69246-750.601
[0100] FIG.2 illustrates an example of pooling, amplification and assaying biological samples according to the methods as described herein.
[0101] FIG. 3 illustrates an example of a computing device; in this case, a device with one or more processors, memory, storage, and a network interface.
[0102] FIG. 4 illustrates an example of data showing uniformity and percentage of targets of various nanograms for a variety of DNA samples.
[0103] FIG.5 illustrates an example of a box plot showing uniformity values for various groups of chemistries each grouped by a coverage value and a clasped or unclasped BAM.
[0104] FIG. 6A illustrates an example of data showing total reads of solid FFPE samples and heme (i.e. peripheral blood and bone marrow) samples.
[0105] FIG.6B illustrates an example of data showing a percentage of UMI consensus reduction for solid FFPE samples and heme (i.e. peripheral blood and bone marrow) samples.
[0106] FIG. 6C illustrates an example of data showing coverage for more than 1,000 cancer genes in solid FFPE samples and heme (i.e. peripheral blood and bone marrow) samples.
[0107] FIG.6D illustrates an example of data showing coverage for remaining variant genes in solid FFPE samples and heme (i.e. peripheral blood and bone marrow) samples.
[0108] FIG.7A illustrates an example of data showing sensitivity, specificity, and accuracy for solid FFPE samples and heme (i.e. peripheral blood and bone marrow) samples.
[0109] FIG. 7B illustrates an example of data showing concordance of CNVs for 59 genes.
[0110] FIG.7C illustrates an example of data showing concordance of CNVs for ERBB2 / HER2 genes.
[0111] FIG.7D illustrates an example of data showing concordance of CNVs for the MET gene.
[0112] FIG. 7E illustrates an example of data showing uniformity for groups of samples that failed QC, Formalin damaged control samples, and GIAB cell lines.
[0113] FIG. 7F illustrates an example of data for concordance of HLA genes in 24 FFPE samples.
[0114] FIG.8 A illustrates an example method of comparing a featured assay analysis of region 1 with a WES of a mimic region 2 gene to generate libraries for region 1.
[0115] FIG. 8B illustrates an example of data concerning tumor percentage, tumor mutational burden (TMB), and microsatellite instability (Mil).
[0116] FIG. 8C illustrates an example of data relating to SNV / indel calling sensitivity and specificity in region 1 for a variety of samples.
[0117] FIG. 8D illustrates an example of data relating to SNV / indel calling sensitivity and specificity in region 1 for various fragments of a variety of samples.Docket No.: 69246-750.601
[0118] FIG. 8E illustrates an example of data relating to SNV / indel calling sensitivity or specificity in region 2 for a variety of samples.
[0119] FIG. 8F illustrates an example of data relating to SNV / indel calling sensitivity or specificity in region 2 for various fragments of a variety of samples.
[0120] FIG. 9 illustrates an example of a WTS library preparation workflow according to methods described herein.
[0121] FIG. 10 illustrates an example of a WTS workflow according to methods described herein.
[0122] FIG. 11 A illustrates an example of data related to HPV detection sensitivity, specificity, accuracy, and positive predictive value (PPV).
[0123] FIG. 11B illustrates an example of summary data of the HPV risk in samples.
[0124] FIG. 12 illustrates an example of HPV genome organization.
[0125] FIG. 13 illustrates an example of data relating to identification of transcripts in Formalin Fixed Paraffin Embedded (FFPE) tumor samples by Whole Transcriptome Sequencing (WTS; bottom panels) and traditional exome capture (top panels) for targeted transcripts and nontargeted transcripts.
[0126] FIG. 14 illustrates data relating to identification of transcripts in hematological cancers by Whole Transcriptome Sequencing (WTS; bottom panels) and traditional exome capture (top panels) for targeted transcripts and non-targeted transcripts.DETAILED DESCRIPTION
[0127] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.Terms and Definitions
[0128] As used herein, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated.
[0129] As used herein, the phrases “at least one,” “one or more,” and “and / or” are open-ended expressions that are both conjunctive and disjunctive in operation. For example, each of the expressions “at least one of A, B and C,” “at least one of A, B, or C,” “one or more of A, B, and C”, “one or more of A, B, or C” and “A, B, and / or C” means A alone, B alone, C alone, A and BDocket No.: 69246-750.601together, A and C together, B and C together, or A, B and C together. As used herein, the phrase “at most three” can mean less than one, one, two, or three.
[0130] Reference throughout this specification to “some embodiments,” “further embodiments,” or “a particular embodiment,” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in some embodiments,” or “in further embodiments,” or “in a particular embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments
[0131] The terms "subject," "individual," and "patient" may be used interchangeably and refer to humans, as well as non-human mammals (e.g., non-human primates, canines, equines, felines, porcines, bovines, ungulates, lagomorphs, rodents, and the like). In various embodiments, the subject can be a human (e.g., adult male, adult female, adolescent male, adolescent female, male child, female child) under the care of a physician or other health worker in a hospital, as an outpatient, or other clinical context. In certain embodiments, the subject may not be under the care or prescription of a physician or other health worker. In some embodiments, the subject may be under the care of a dental professional.
[0132] As used herein, “treatment” or “treating” refers to an approach for obtaining beneficial or desired results with respect to a disease, disorder, or medical condition including, but not limited to, a therapeutic benefit and / or a prophylactic benefit. In certain embodiments, treatment or treating involves administering a therapeutic to a subject. A therapeutic benefit may include the eradication or amelioration of the underlying disorder being treated. Also, a therapeutic benefit may be achieved with the eradication or amelioration of one or more of the physiological symptoms associated with the underlying disorder, such as observing an improvement in the subject, notwithstanding that the subject may still be afflicted with the underlying disorder.
[0133] The term “about” or “approximately” means within an acceptable error range for the particular value, which may depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, “about” can mean within 1 or more than 1 standard deviation. Alternatively, “about” can mean a range of upto 20%, up to 10%, up to 5%, or up to 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within an order of magnitude, preferably within 5 -fold, and more preferably within 2-fold, of a value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value may be assumed.Docket No.: 69246-750.601
[0134] The terms “determining”, “assessing”, “assaying,” and “analyzing” are often used interchangeably herein to refer to forms of measurement. The terms can include determining if an element is present or not (for example, detection). These terms can include quantitative, qualitative, or quantitative and qualitative determinations. Assessing can be relative or absolute. “Detecting the presence of’ can include determining the amount of something present in addition to determining whether it is present or absent depending on the context.
[0135] The term “cancer” is used herein to refer to any disease characterized by uncontrolled cell division. A cancer can be a cancer of the blood (e.g., hematological cancer), e.g., leukemia, lymphoma, or multiple myeloma, or a cancer can be neoplastic, e.g., associated with an abnormal mass of tissue in which cells grow and divide more than they should or do not die when they should. Neoplastic cancers, e.g., lung, breast or liver cancer, are associated with a solid tumor.
[0136] The terms “cell-free DNA,” “cfDNA molecules,” or simply “cfDNA” refer to DNA molecules that naturally occur in a subject in extracellular form (e.g., in blood, serum, plasma, or other bodily fluids such as lymph, cerebrospinal fluid, urine, or sputum). While the cfDNA previously existed in a cell or cells in a large complex biological organism, e.g., a mammal, it has undergone release from the cell(s) into a fluid found in the organism, and may be obtained from a sample of the fluid without the need to perform an in vitro cell lysis step. cfDNA molecules may occur as DNA fragments.
[0137] The term “circulating tumor DNA” or “ctDNA” refers to DNA that originates directly from a tumor or from circulating tumor cells (CTCs), which are viable, intact tumor cells that shed from primary tumors and can enter the bloodstream or lymphatic system. The precise mechanism of how ctDNA is released is unclear, although it is postulated to involve apoptosis and necrosis from dying cells, or active release from viable tumor cells. Circulating tumor DNA (ctDNA) can be highly fragmented and in some embodiments can have a mean fragment size about 100-250 bp, e.g., 150 to 200 bp long. The amount of ctDNA in a sample of circulating cell-free DNA isolated from a cancer patient varies greatly: typical samples contain less than 10% ctDNA, although many samples from patients being assessed for MRD may have less than 0.01 % ctDNA and some samples have over 10% ctDNA. Molecules of ctDNA can be often identified because they contain tumorigenic mutations.
[0138] As used herein, the terms “neoplasm” and “tumor” are used interchangeably. They refer to abnormal growth of cells in a subject. A neoplasm or tumor can be benign, potentially malignant, or malignant. A malignant tumor is referred to as a cancer or a cancerous tumor.DocketNo.: 69246-750.601
[0139] The term “nucleic acid” and “polynucleotide” are used interchangeably herein to describe a polymer of any length, e.g., greater than about 2 bases, greater than about 10 bases, greater than about lOObases, greater than about 500 bases, greaterthan 1000 bases, greater than 10,000 bases, greaterthan 100,000 bases, greaterthan about 1,000,000, up to about 1010 or more bases composed of nucleotides, e.g., deoxyribonucleotides or ribonucleotides, and may be produced enzymatically or synthetically (e.g., PNA as described in U.S. Patent No. 5,948,902 and the references cited therein) which can hybridize with naturally occurring nucleic acids in a sequence specific manner analogous to that of two naturally occurring nucleic acids, e.g., can participate in Watson-Crick base pairing interactions. Naturally occurring nucleotides include guanine, cytosine, adenine, thymine, uracil (G, C, A, T and U respectively). DNA and RNA have a deoxyribose and ribose sugar backbone, respectively, whereas PNA's backbone is composed of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. In PNA various purine and pyrimidine bases are linked to the backbone by methylene-carbonyl bonds. A locked nucleic acid (LNA), often referred to as inaccessible RNA, is a modified RNA nucleotide. The ribose moiety of an LNA nucleotide is modified with an extra bridge connecting the 2' oxygen and 4' carbon. The bridge "locks" the ribose in the 3'-endo (North) conformation, which is often found in the A-form duplexes. LNA nucleotides can be mixed with DNA or RNA residues in the oligonucleotide whenever desired.
[0140] As used herein, a “sample” or a “biological sample” refers to anything capable of being analyzed by the methods or systems disclosed herein. A sample may be derived from a subject.
[0141] As used herein, “selective enrichment” or “selectively enriching” refers to any process where a concentration of a target sequence, or a target (gene) locus, is increased relative to its initial concentration. Such processes may amplify the target sequence or target locus to increase the concentration prior to sequencing. Such processes may capture the target sequence or target locus to increase the concentration prior to sequencing. Such processes may include hybrid capture methods.
[0142] As used herein, “sequencing” refers to any of a number of technologies used to determine the sequence (e.g., the identity and order of monomer units) of a biomolecule, e.g., a nucleic acid such as DNA or RNA. Examples of sequencing methods include, but are not limited to, targeted sequencing, single molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy termination sequencing, whole-genome sequencing, sequencing by hybridization, pyrosequencing, duplex sequencing, cycle sequencing, single-base extension sequencing, solid phase sequencing, high-Docket No.: 69246-750.601throughput sequencing, massively parallel signature sequencing, emulsion PCR, co -amplification at lower denaturation temperature-PCR (COLD-PCR), multiplex PCR, sequencing by reversible dye terminator, paired-end sequencing, near-term sequencing, exonuclease sequencing, sequencing by ligation, short-read sequencing, single molecule sequencing, sequencing-by-synthesis, real-time sequencing, reverse-terminator sequencing, long-read sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD sequencing, MS-PET sequencing, and a combination thereof. In some embodiments, sequencing can be performed by a gene analyzer such as, for example, gene analyzers commercially available from Illumina, Inc., Ultima Genomics, Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, among many others.
[0143] The term “next-generation sequencing” or “NGS”, as used herein, generally refers to sequencing technologies having increased throughput as compared to traditional Sanger- and capillary electrophoresis-based approaches, for example, with the ability to generate hundreds of thousands of relatively small sequence reads at a time. Some examples of next -generation sequencing techniques include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization. In some embodiments, next-generation sequencing includes the use of instruments capable of sequencing single molecules.
[0144] Where values are disclosed as ranges, it may be understood that such disclosure includes the disclosure of all possible sub-ranges within such ranges, as well as specific numerical values that fall within such ranges irrespective of whether a specific numerical value or specific subrange is expressly stated.Methods for Preparing Enriched Nucleic Acids from Low-Input Biological Samples
[0145] In an aspect, the present disclosure provides a method of preparing an enriched nucleic acid sample, the method comprising: obtaining a biological sample obtained or derived from a subject; denaturing the biological sample to provide a plurality of single stranded deoxyribonucleic acid (ssDNA) fragments; amplifying the plurality of ssDNA fragments in two or more cycles of polymerase chain reaction (PCR), thereby producing a sequencing library; and contacting the sequencing library with a pool of capture probes as shown in FIG. 1.
[0146] In some embodiments, producing a sequencing library as illustrated in FIG. 1 may comprise utilizing a DNA extraction protocol optimized for small numbers of cells. In some embodiments, the sequencing of the first sample can comprise SRSLY (Super -Resolution Sequencing of Low-Yield samples). SRSLY (Super-Resolution Sequencing of Low -Yield samples) is a technology designed specifically for genomic analysis of samples with limitedDocket No.: 69246-750.601material, such as those obtained through fine needle aspiration (FNA) or core needle biopsies. The SRSLY method can incorporate several components to maximize information yield from minimal input. The process may comprise a highly efficient DNA extraction protocol optimized for small numbers of cells, capable of recovering DNA from as few as 10-100 cells. In some embodiments, extracted DNA can be utilized in a library preparation process as shown in FIG. I that may include: (1) an initial repair to address damage that is common in clinical samples, particularly FFPE-preserved specimens; (2) incorporation of unique molecular identifiers (UMIs) to enable tracking of individual DNA molecules throughout the workflow; (3) a minimal -loss adapter ligation strategy that achieves high conversion efficiency of input DNA molecules to sequencing-ready fragments; and (4) a carefully optimized amplification protocol that minimizes bias while generating sufficient library material for sequencing. The resulting libraries may undergo sequencing on high-output platforms, followed by analysis with bioinformatics pipelines specifically calibrated for low-input samples. These pipelines may employ error suppression strategies based on the UMIs and statistical models that account f or the unique characteristics of data generated from minimal input samples. The SRSLY approach enables comprehensive genomic profiling from challenging clinical samples that would be inadequate for some sequencing methods, making it possible to define robust digital signatures for MRD monitoring even when only minimally invasive biopsy procedures are feasible.
[0147] In some embodiments, the nucleic acid molecules can be obtained or derived from a biological sample. The biological sample can be obtained or derived from a subject. The subject can comprise a human subject. The subject can be an adult. The subject can be a child. The subject can be an infant. The subject can be a newborn. In some embodiments, the subject canbe a fetus. In some embodiments, the subject can be less than one year old. In some embodiments, the subject can be between seconds old and months old. In some embodiments, the subject can be between about one second old and one hour old. In some embodiments, the subject can be about a day old, a week old, a month old, or more than a month old. In some embodiments, the subject can be about2 months old, 3 months old, 4 months old, 5 months old, 6 months old, 7 months old, 8 months old, 9 months old, 10 months old, 11 months old, 12 months old, or more than 12 months old. In some embodiments, the subject can be about 1 year old, 2 years old, 3 years old, 4 years old, 5 years old, 6 years old, 7 years old, 8 years old, 9 years old, 10 years old, I I years old, 12 years old, 13 years old, 14 years old, 15 years old, 16 years old, 17 years old, 18 years old, or more than 18 years old. In some embodiments, the subject canbe an adult. In some embodiments, the subject can be between about 18 and 100 years old.Docket No.: 69246-750.601
[0148] In some embodiments, the subject may be a mammal. For example, the mammal may include a mouse, a rat, a gerbil, a guinea pig, a hamster, a fox, a dog, a monkey, a sheep, a cow, a pig, or the like. The mammal may include a monkey. For example, the mammal may comprise a chimpanzee, a bonobo, an orangutan, a baboon, or the like. The subject may be a human. The subject may be an adult (e.g., at least 18 years of age). The subject may be a child (e.g., less than 18 years of age). The subject may be a male. The subject may be a female.
[0149] In some embodiments, the biological sample can comprise one or more of: a blood sample, a plasma sample, a serum sample, a urine sample, a saliva sample, a cerebrospinal fluid sample, a stool sample, a lymph node sample, or a bone marrow sample, or any combination thereof. The biological sample may comprise a blood sample. The biological sample may comprise a plasma sample. The biological sample may comprise a serum sample. The biological sample may comprise a urine sample. The biological sample may comprise a saliva sample. The biological sample may comprise a cerebrospinal fluid sample. The biological sample may comprise a stool sample. The biological sample may comprise a lymph node sample. The biological sample may comprise a bone marrow sample.
[0150] The biological samples described herein may be obtained by any suitable method. In some embodiments, the biological sample may be obtained by a blood draw. In some embodiments, the biological sample may be obtained by a biopsy. The biopsy may be a liquid biopsy. The biopsy may comprise a fine needle aspiration (FNA). The biopsy may comprise a core needle biopsy. The biopsy may comprise a stereotactic biopsy. The biopsy may comprise an excisional biopsy, a bone marrow biopsy, an endometrial biopsy, an endoscopic biopsy, an incisional biopsy, a colposcopy-directed biopsy, a needle biopsy, a skin biopsy, a lymph node biopsy, or a combination thereof. The biological sample may be obtained by a biopsy. The biological sample may be completely obtained by a biopsy.
[0151] In some embodiments, the biological sample may be a tumor sample. The biological sample may be obtained or derived from tumor tissue of a subject. The tumor tissue may be obtained using a biopsy. The biopsy may comprise an FNA. The FNA can be used to obtain a small amount of tumor tissue of a subject. The small amount of tumor tissue can be utilized as a sample in performing genomic assays as described herein.
[0152] In some embodiments, the biological sample can be a cell-free biological sample. The cell-free biological sample can comprise one or more of: a plasma sample, a serum sample, a urine sample, a saliva sample, a cerebrospinal fluid sample, a lymph sample, or any combination thereof. In some embodiments, the biological sample may comprise a cell -free biological sample. The cell-free biological sample may comprise a plasma sample. The cell -free biological sampleDocket No.: 69246-750.601may comprise a serum sample. The cell-free biological sample may comprise a urine sample. The biological sample may be at least 50% cell-free (e.g., at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 99%). The biological sample may be completely cell-free.
[0153] In some embodiments, at least a portion of the nucleic acid molecules can comprise deoxyribonucleic acid (DNA), or ribonucleic acid (RNA), or both.
[0154] The DNA can comprise single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), or a combination thereof.
[0155] The cell-free biological sample can comprise cell-free deoxyribonucleic acid (cfDNA), cell-free ribonucleic acid (cfRNA), or a combination thereof. The cell -free biological sample may comprise cfDNA. The cell-free biological sample may comprise cfRNA. The cell-free biological sample may comprise cfDNA and cfRNA.
[0156] In some embodiments, the biological sample may be about 10 nanograms (ng) to about 100 ng. The biological sample may comprise DNA extracted from the subject. The biological sample may comprise about 10 ng to about 50 ng of DNA extracted from the subject. The biological sample may be between about 1 ng and about 100 ng. The biological sample may comprise between about 1 ng and about 100 ng of DNA extracted from the subject. The biological sample may be between about 1 ng and about 50 ng. The biological sample may be about 25 ng. The biological sample may comprise between about 1 ng and about 50 ng of DNA extracted from the subject. The biological sample may be about 50 ng. The biological sample may be about 100 ng. The biological sample may be about 1 ng, 2 ng, 3 ng, 4 ng, 5 ng, 6 ng, 7 ng, 8 ng, 9 ng, 10 ng, 11 ng, 12 ng, 13 ng, 15 ng, 15 ng, 16 ng, 17 ng, 18 ng, 19 ng, 20 ng, 21 ng, 22 ng, 23 ng, 24 ng, 25 ng, 26 ng, 27 ng, 28 ng, 29 ng, 30 ng, 31 ng, 32 ng, 33 ng, 34 ng, 35 ng, 36 ng, 37 ng, 38 ng, 39 ng, 40 ng, 41 ng, 42 ng, 43 ng, 44 ng, 45 ng, 46 ng, 47 ng, 48 ng, 49 ng, 50 ng, 51 ng, 52 ng, 53 ng, 54 ng, 55 ng, 56 ng, 57 ng, 58 ng, 59 ng, 60 ng, 61 ng, 62 ng, 63 ng, 64 ng, 65 ng, 66 ng, 67 ng, 68 ng, 69 ng, 70 ng, 71 ng, 72 ng, 73 ng, 74 ng, 75 ng, 76 ng, 77 ng, 78 ng, 79 ng, 80 ng, 81 ng, 82 ng, 83 ng, 84 ng, 85 ng, 86 ng, 87 ng, 88 ng, 89 ng, 90 ng, 91 ng, 92 ng, 93 ng, 94 ng, 95 ng, 96 ng, 97 ng, 98 ng, 99 ng, or about 100 ng, or more than about 100 ng.
[0157] In some embodiments, the pool of capture probes may comprise capture probes capable of hybridizing to at least a portion of an exome. The capture probes may be capable of hybridizing to at least a portion of a biomarker. The capture probes may be capable of hybridizingto a single nucleotide polymorphism (SNP). The capture probes may be capable ofDocket No.: 69246-750.601hybridizing to at least a portion of a biomarker and an SNP. The hybridizing may produce the enriched nucleic acid sample.
[0158] In some embodiments, the sequencing library may comprise a double-stranded deoxyribonucleic acid (dsDNA) library. In some embodiments, the biological sample may be about 10 nanograms. In some embodiments, the biological sample may be about 50 nanograms. The biological sample may comprise two or more sample types. The two or more sample types may comprise two or more of a Formalin-Fixed Paraffin Embedded (FFPE) sample, a fine-needle aspirate (FNA) sample, a blood sample, a plasma sample, a serum sample, or a bone marrow sample. The biological sample may comprise an FFPE sample, a blood sample, a plasma sample, a serum sample, or a bone marrow sample. The biological sample may be the FFPE sample. The biological sample may be the bone marrow sample. The biological sample may be the blood sample. The blood sample may be a peripheral blood sample. The biological sample may be the peripheral blood sample. The biological sample may comprise a cell line. The cell line may be commercially available. The biological sample may be obtained or derived from a solid tumor of the subject. The biological sample may be obtained or derived from a biopsy. The biopsy may be a liquid biopsy. The biopsy may comprise a fine needle aspirate or a core needle biopsy. The biopsy maybe the fine needle aspirate. The biopsy may be the core needle biopsy.
[0159] In some embodiments, methods for sequencing of the first sample can comprise utilizing FNA or core needle sequencing. Fine needle aspiration (FNA) and core needle biopsies may present challenges for genomic analysis due to the limited amount of material obtained and potential heterogeneity of the samples. Various sequencing methods may address these challenges and enable comprehensive genomic profiling from these minimally invasive biopsy procedures. These methods may incorporate several key aspects including, but not limited to: (1) ultra -low input DNA extraction protocols optimized to maximize recovery from limited cellular material, sometimes yielding nanogram quantities of DNA; (2) library preparation techniques specifically designed for low-input and potentially degraded DNA (such as employing targeted approaches to focus sequencing on regions of interest); (3) unique molecular identifiers (UMIs) added during library preparation to enable identification and removal of PCR duplicates, which are particularly problematic with low-input samples; and (4) bioinformatics pipelines calibrated for the characteristics of data generated from these challenging samples, including adjustments for potential sampling bias and tumor heterogeneity. These methods may enable successful sequencing from FNA samples, which may yield very small numbers of cells collected through a thin needle, and / or core needle biopsies, which provide small tissue cores but still represent limited sampling of potentially heterogeneous tumors. By making it possible to performDocket No.: 69246-750.601comprehensive genomic analysis on these minimally invasive samples, these methods expand the accessibility of precision oncology approaches, including MRD monitoring, to subjects for whom more extensive tissue sampling may not be feasible due to tumor location, subject condition, or other clinical considerations.
[0160] In some embodiments, the biological sample may comprise ribonucleic acid (RNA), deoxyribonucleic acid (DNA), or a combination thereof. The biological sample may comprise the RNA. The RNA may be cell-free RNA (cfRNA). The RNA may be derived from exosomes, such as exosomal RNA. The RNA may be coding RNA (cRNA). The RNA may be non -coding RNA (ncRNA). The biological sample may comprise the DNA. The DNA may be cell -free DNA or DNA from cells. The DNA may be ssDNA, double-stranded DNA (dsDNA), or a combination thereof. The DNA may comprise the ssDNA. The DNA may comprise the dsDNA. The DNA may comprise one or more DNA fragments comprising a nick. The DNA may comprise end-damaged DNA. The DNA may comprise single stranded DNA overhangs on a double-stranded DNA.
[0161] In some embodiments, the method may further comprise processing the sample to produce ssDNA. In some embodiments, the processing further comprises generating the dsDNA from the extracted RNA. In some embodiments, the dsDNA comprises double-stranded circular DNA (dscDNA). In some embodiments, a reverse transcriptase is used produce cDNA amplicons using RNA as a template. In some embodiments, a DNA polymerase is used produce DNA amplicons using DNA as a template.
[0162] In some embodiments, the processing further comprises denaturing the dsDNA or the dscDNA. The temperature at which the dsDNA or the dscDNA denatures or melts is referred to as the melting temperature (Tm). The Tmis the temperature at which one-half (50%) of a DNA duplex of an oligonucleotide (such as a primer) and its perfect complement dissociates and becomes single strand DNA. In some embodiments, the method may further comprise assaying the denatured dsDNA or dscDNA.
[0163] In some embodiments, the amplifying may comprise three or more, four or more, five or more, six or more, seven or more, eight or more, or nine or more cycles of PCR. The amplifying may comprise eight cycles of PCR. The amplifying may comprise five to nine cycles of PCR. The amplifying may comprise six to nine cycles of PCR. The amplifying may comprise six to eight cycles of PCR.
[0164] In some embodiments, the amplifying may further comprise ligating one or more adapters to the plurality of ssDNA fragments. The one or more adapters may compriseDocket No.: 69246-750.601directional adapters. In some embodiments, the one or more adapters may comprise Unique Molecular Index (UMI) adapters.
[0165] In some embodiments, the method may further comprise ligating one or more UMI adapters to the at least subset of the plurality of ssDNA fragments prior to the amplification in (c). A UMI adapter may be ligated onto a portion of a single ssDNA fragment of the plurality of ssDNA fragments prior to the amplification. Two or more UMI adapters may be ligated onto separate portions of a single ssDNA fragment of the plurality of ssDNA fragments prior to the amplification. The UMI adapter may be ligated to an ssDNA fragment at the 3’ end. The UMI adapter may be ligated to an ssDNA fragment at the 5’ end. The UMI adapter ligated to the ssDNA fragments may be amplified with the ssDNA fragments during amplification. One or more UMI adapters ligated to the ssDNA fragment may comprise an UMLssDNA complex. The UMI-ssDNA complex may be amplified as a complex. Multiple UMI adapters may be ligated to an ssDNA fragment. One, two, three, four, or more than four UMI adapters may be ligated to a single ssDNA fragment.
[0166] In some embodiments, the method may further comprise ligating one or more UMI adapters to the at least subset of the plurality of ssDNA fragments after the amplification in (c). A UMI adapter may be ligated onto a portion of a single ssDNA fragment of the plurality of amplified ssDNA fragments after amplification. Two or more UMI adapters may be ligated onto separate portions of a single ssDNA fragment of the plurality of amplified ssDNA fragments after amplification. The UMI adapter may be ligated to an ssDNA fragment at the 3’ end. The UMI adapter may be ligated to an ssDNA fragment at the 5 ’ end. The UMI adapters may not be amplified. One or more UMI adapters ligated to the ssDNA fragment may comprise an UMI-ssDNA complex. Multiple UMI adapters may be ligated to an amplified ssDNA fragment. One, two, three, four, or more than four UMI adapters may be ligated to a single amplified ssDNA fragment.
[0167] In some embodiments, the amplification may further comprise providing a Unique Molecular Index (UMI) sequence to the ssDNA fragments. The UMI sequence may be provided via primer extension or via PCR. The UMI sequence may comprise a length of about 5 nucleotides to about 15 nucleotides. The UMI sequence may comprise a length of about 9 nucleotides.
[0168] In some embodiments, the pool of capture probes may comprise at least two capture probes capable of hybridizing to at least a portion of an exome. The capture probes may be capable of hybridizing to at least a portion of a biomarker. The capture probes may be capable of hybridizing to a single nucleotide polymorphism (SNP). The capture probes may be capable ofDocket No.: 69246-750.601hybridizing to both of an at least portion of a biomarker and an SNP. The pool of capture probes may comprise capture probes capable of hybridizing to at least a portion of an exome. The capture probes may be capable of hybridizing to at least a portion of a biomarker. The capture probes may be capable of hybridizing to a single nucleotide polymorphism (SNP). The exome may be a human exome.
[0169] In some embodiments, the biomarker may comprise a plurality of biomarkers. The plurality of biomarkers may comprise at least 100 biomarkers, at least 200 biomarkers, at least 300 biomarkers, at least 400 biomarkers, at least 500 biomarkers, at least 600 biomarkers, at least 700 biomarkers, at least 800 biomarkers, at least 900 biomarkers, at least 1,000 biomarkers, at least 1,100 biomarkers, at least 1,200 biomarkers, at least 1,300 biomarkers, at least 1,400 biomarkers, oratleast 1,500 biomarkers. The plurality of biomarkers may comprise about 500 biomarkers. The plurality of biomarkers may comprise about 1,100 biomarkers. The plurality of biomarkers may comprise between about 20,000 biomarkers and about 25,000 biomarkers. In some embodiments, the plurality of biomarkers may comprise about 30,000 biomarkers, 35,000 biomarkers, 40,000 biomarkers, 45,000 biomarkers, 50,000 biomarkers, 55,000 biomarkers, 60,000 biomarkers, 65,000 biomarkers, 70,000 biomarkers, 75,000 biomarkers, 80,000 biomarkers, 85,000 biomarkers, 90,000 biomarkers, 95,000 biomarkers, 100,000 biomarkers, or more than 100,000 biomarkers. The plurality of biomarkers may comprise more than 65,000 biomarkers.
[0170] In some embodiments, the plurality of biomarkers may be associated with a disease or disorder. The disease or disorder may comprise cancer. The plurality of biomarkers may be associated with solid tumor cancer. The solid tumor cancer may comprise carcinoma, sarcoma, or lymphoma. The plurality of biomarkers may be associated with a hematologic malignancy.
[0171] In some embodiments, the method may further comprise amplifying the plurality of ssDNA fragments. The amplifying the plurality of ssDNA fragments may comprise polymerase chain reaction (PCR). The PCR can comprise one or more of: multiplex PCR, blocker PCR, nested PCR, indexing PCR, emulsion PCR (ePCR), or any combination thereof. The amplifying may comprise isothermal amplification. The amplifying may comprise rolling circle amplification (RCA). The amplifying may comprise strand displacement amplification.
[0172] In some embodiments, the method may further comprise sequencing the enriched nucleic acid sample. The sequencing may comprise next generation sequencing (NGS). The sequencing may comprise one or more of: whole genome sequencing (WGS), whole transcriptome sequencing (WTS), targeted sequencing, NGS, methylomics assays, fragmentomics, or any combination thereof. The sequencing assay may comprise next generation sequencing (NGS).Docket No.: 69246-750.601The sequencing assay may comprise deoxyribonucleic acid (DNA) sequencing. The sequencing assay may comprise ribonucleic acid (RNA) sequencing. The sequencing assay may comprise pyrosequencing. The sequencing may comprise whole genome sequencing (WGS). The sequencing may comprise whole transcriptome sequencing. The sequencing may comprise whole exome sequencing (WES). The sequencing may comprise whole transcriptome sequencing (WTS). The sequencing may comprise targeted sequencing. The sequencing may comprise shotgun sequencing. The sequencing may comprise single-cell sequencing. The sequencing may comprise sanger sequencing. The sequencing may comprise SOLiD sequencing. The sequencing may comprise paired plus minus sequencing (ppmSeq). The sequencing may comprise ion semiconductor sequencing. The sequencing may comprise sequencing by synthesis. The sequencing may comprise DNA nanoball sequencing. The sequencing may comprise polony sequencing. The sequencing may comprise bridge polymerase chain reaction (PCR) sequencing. The sequencing may comprise one or more of DNA sequencing, RNA sequencing, wholegenome sequencing, whole-exome sequencing, whole-transcriptome sequencing, targeted sequencing, bisulfite sequencing, or enzymatic sequencing, or any combination thereof. The sequencing may comprise use of Illumina™ sequencing, Ultima™ sequencing, Ion Torrent™ sequencing, or a combination thereof.
[0173] In some embodiments, the method may further comprise computer processing the enriched nucleic acid sample.
[0174] In some embodiments, the method may further comprise analyzing data from the sequencing utilizing one or more computer processors. The analyzing can comprise generating sequencing data using the sequencing assay. The analyzing can further comprise comparing the generated sequencing data to reference sequencing data to determine one or more mutations of the amplified nucleic acid molecules. The generated sequence data can comprise one or more data sets. The generated sequence data can comprise strings, byte sequences, byte arrays, lists, tuples, range objects, matrices, or other data formats of data obtained or derived from sequencing the amplified nucleic acid molecules. In some embodiments, the one or more mutations may be distinguished from sequencing errors. The one or more mutations may be distinguished from sequencing errors by comparing a first set of sequencing data from a first sample of amplified nucleic acid molecules and a second set of sequencing data from a second sample of amplified nucleic acid molecules. The first sample and second sample of amplified nucleic molecules can comprise nucleic acid molecules amplified from the same primary sample of nucleic acid molecules. The mutations can comprise oneormore of: single nucleotide variants (SNVs), single nucleotide polymorphisms (SNPs), double base substitutions, insertion -deletions (indels), copyDocket No.: 69246-750.601number variations (CNVs), inversions, translocations, tandem repeats (TRs), short tandem repeats (STRs), variable number tandem repeats (VNTRs), quasi-tandem repeats (QTRs), or a combination thereof.
[0175] In some embodiments, analyzing the sequencing data may comprise determining genetic information of the biomarkers. In some embodiments, analyzing the sequencing data may comprise determining epigenetic information of the biomarkers. In some embodiments, analyzing the sequencing data may comprise determining both genetic information and epigenetic information of the biomarkers. The genetic information may comprise one or more of: tumor mutational burden (TMB), microsatellite instability (MSI), copy number variation (CNV), loss of heterozygosity (LoH), or human leukocyte antigen (HLA) identification, or any combination thereof. Genotyping can be performed for HLA identification and LoH identification. In some embodiments, the method can comprise identifying the HLAs based at least in part on HLA genotyping. The HLA genotyping can be utilized to determine or predict emerging biomarkers associated with protein markers of disease. The emerging biomarkers can comprise protein biomarkers. The protein biomarkers can be novel protein markers not previously found. The protein markers may be therapeutic targets. The biomarkers can comprise BRCA1 mutations, BRCA2 mutations, LoH, HRD status for DNA, HRD status for RNA, CNV modifications, or any combination thereof.
[0176] The epigenetic information may comprise methylation information. In some embodiments, determining the epigenetic information can comprise determining a methylation pattern corresponding to at least a portion of the subset of the nucleic acid molecules. In some embodiments, the method can further comprise determining the one or more unmethylated sites of the at least portion of the subset of the nucleic acid molecules. In some embodiments, the one or more unmethylated sites can be on a subset of the strands of the nucleic acid molecules. In some embodiments, the one or more unmethylated sites can be on one or more fragments of the nucleic acid molecules. In some embodiments, at least a portion of the one or more unmethylated sites can comprise an unmethylated cysteine nucleotide. In some embodiments, at least a portion of the one or more unmethylated sites can comprise unmethylated CpG sites.
[0177] In some embodiments, both DNA and RNA extracted from the biological sample can be analyzed. The DNA and RNA extracted from the biological sample can be analyzed in parallel. The DNA can be analyzed using WES in parallel with RNA analysis using whole transcriptome sequencing (WTS). HRD status can be determined based at least in part on WES and WTS analysis data.Docket No.: 69246-750.601Systems for Performing Assays of Low-Input Samples
[0178] In another aspect, the present disclosure provides a system comprising one or more processors and a memory operatively coupled to the one or more processors. The one or more processors may be individually or collectively programmed to perform the method.Examples of Machine Learning Techniques
[0179] As disclosed throughout, in some cases, the systems, the methods, the computer-readable media, and the techniques disclosed herein may implement one or more machine learning techniques. In some cases, ML may generally involve identifying and recognizing patterns in existing data in order to facilitate making predictions for subsequent data. ML may include a ML model (which may include, for example, a ML algorithm). Machine learning, whether analytical or statistical in nature, may provide deductive or abductive inference based on real or simulated data. The ML model may be a trained model. ML techniques may comprise one or more supervised, semi-supervised, self-supervised, or unsupervised ML techniques. For example, an ML model may be a trained model that is trained through supervised learning (e.g., various parameters are determined as weights or scaling factors). ML may comprise one or more of regression analysis, regularization, classification, dimensionality reduction, ensemble learning, meta learning, association rule learning, cluster analysis, anomaly detection, deep learning, or ultra-deep learning. ML may comprise: k-means, k-means clustering, k-nearest neighbors, learning vector quantization, linear regression, non-linear regression, least squares regression, partial least squares regression, logistic regression, stepwise regression, multivariate adaptive regression splines, ridge regression, principal component regression, least absolute shrinkage and selection operation (LASSO), least angle regression, canonical correlation analysis, factor analysis, independent component analysis, linear discriminant analysis, multidimensional scaling, non-negative matrix factorization, principal components analysis, principal coordinates analysis, projection pursuit, Sammon mapping, t-distributed stochastic neighbor embedding, AdaBoosting, boosting, gradient boosting, bootstrap aggregation, ensemble averaging, decision trees, conditional decision trees, boosted decision trees, gradient boosted decision trees, random forests, stacked generalization, Bayesian networks, Bayesian belief networks, naive Bayes, Gaussian naive Bayes, multinomial naive Bayes, hidden Markov models, hierarchical hidden Markov models, support vector machines, encoders, decoders, auto -encoders, stacked autoencoders, perceptrons, multi-layer perceptrons, artificial neural networks, feedforward neural networks, convolutional neural networks, recurrent neural networks, residual neural networks, physics-informed neural networks, long short-term memory, deep belief networks, deepDocket No.: 69246-750.601Boltzmann machines, deep convolutional neural networks, deep recurrent neural networks, large language models, transformer models, vision transformers, or generative adversarial networks.Examples of Decision Trees and Random Forests
[0180] As described above, the machine learning model may implement a decision tree. A decision tree may be a supervised ML algorithm that can be applied to both regression and classification problems. For example, a decision tree may grow from a root (base condition), and when it meets a condition (internal node / feature), it may split into multiple branches. The end of the branch that does not split anymore may be an outcome (leaf). A decision tree can be generated using a training dataset set according to the following operations: (A) starting from a root node (the entire dataset), the algorithm may split the dataset in two branches using a decision rule or branching criterion; (B) each of these two branches may generate a new child node; (C) for each new child node, the branching process may be repeated until the dataset cannot be split any further; (D) each branching criterion may be chosen to maximize information gain (e.g., a quantification of how much a branching criterion reduces a quantification of how mixed the labels are in the children nodes). The labels may be the data or the classification that is predicted by the decision tree.
[0181] A random forest regression is an extension of the decision tree model that tends to yield more robust predictions by stretching the use of the training dataset partition. Whereas a decision tree may make a single pass through the data, a random forest regression may bootstrap 50% of the data (e.g., with replacement) and build many trees. Rather than using all explanatory variables as candidates for splitting, a random subset of candidate variables may be used for splitting, which may enable trees that have different data and different variables (hence the term random). The predictions from the trees, which may be collectively referred to as the “forest,” may then be averaged to produce a final prediction. Many trees (e.g., ten trees, fifty trees, one hundred trees, one thousand trees, etc.) may be included in a random forest model, with a number (e.g., 3, 6, 10, etc.) of terms sampled per split, a minimum of number (e.g., 1, 2, 4, 10, etc.) of splits per tree, and a minimum split size (e.g., 16, 32, 64, 128, 256, etc.). Random forests may be trained in a similar way as decision trees. Specifically, training a random forest may include the following operations: (A) randomly select k features from the total number of features; (B) create a decision tree from these k features using the same operations as for generating a decision tree; and (C) repeat the previous two operations until a target number of trees is created.
[0182] As disclosed, a random forest classifier, which may comprise a plurality of decision trees where the output prediction may be the mode of the predicted classifications of the individualDocket No.: 69246-750.601trees, can be helpful in reducing overfitting to training dataset. In some cases, an ensemble of decision trees can be constructed using a random subset of features at each split or decision node. The Gini criterion may be employed, in some cases, to choose the best partition, where decision nodes having the lowest calculated Gini impurity index are selected. The Gini impurity can be used, in some cases, as a criterion to find informative features based on which the splits in each decision tree may be constructed.
[0183] In some cases, each decision tree of a random forest may comprise one or more decision nodes, where each decision node specifies a predicate condition. For example, decision node may predicate the condition that, for a given dataset, the outcome to a question is a specific outcome. At each decision node, a decision tree can be split based on whether the predicate condition attached to the decision node holds true, leading to various prediction nodes. Each prediction node can comprise output values that represent “votes” for one or more of the classifications or conditions being evaluated by the assessment model. At prediction time, a “vote” can be taken overall of the decision trees, and the majority vote (or mode of the predicted classifications) can be output as the predicted classification.
[0184] In some cases, when the dataset being queried in the assessment model reaches a “leaf’, or a final prediction node with no further downstream splits, the output values of the leaf can be output as the votes for the particular decision tree. Since a random forest model comprises a plurality of decision trees, the final votes across all trees in the forest can be summed to yield the final votes and the corresponding classification of the subject. A large number of decision trees can help reduce overfitting of the assessment model to the training dataset, by reducing the variance of each individual decision tree. For example, an assessment model can comprise, for example, at least about 3 decision trees, at least about 5 decision trees, at least about 10 decision trees, at least about 20 decision trees, at least about 50 decision trees, at least about 100 decision trees, etc.Computing Systems
[0185] Referring to FIG. 3, a block diagram is shown depicting an exemplary machine that includes a computer system 300 (e.g., a processing or computing system) within which a set of instructions can execute for causing a device to perform or execute any one or more of the aspects and / or methodologies for static code scheduling of the present disclosure. The components in FIG.3 are examples only and do not limit the scope of use or functionality of any hardware, software, embedded logic component, or a combination of two or more such components implementing particular embodiments.Docket No.: 69246-750.601
[0186] Computer system 300 may include one or more processors 301, a memory 303, and a storage 307 that communicate with each other, and with other components, via a bus 340. The bus 340 may also link a display 332, one or more input devices 333 (which may, for example, include a keypad, a keyboard, a mouse, a stylus, etc.), one or more output devices 334, one or more storage devices 335, and various tangible storage media 336. All of these elements may interface directly or via one or more interfaces or adaptors to the bus 340. For instance, the various tangible storage media 336 can interface with the bus 340 via storage medium interface 326. Computer system 300 may have any suitable physical form, including but not limited to one or more integrated circuits (ICs), printed circuit boards (PCBs), mobile handheld devices (such as mobile telephones or PDAs), laptop or notebook computers, distributed computer systems, computing grids, or servers.
[0187] Computer system 300 includes one or more processor(s) 301 (e.g., central processing units (CPUs) or general-purpose graphics processing units (GPGPUs)) that carry out functions. Processor(s) 301 optionally contains a cache memory unit 302 for temporary local storage of instructions, data, or computer addresses. Processor(s) 301 are configured to assist in execution of computer readable instructions. Computer system 300 may provide functionality for the components depicted in FIG. 3 as a result of the processor(s) 301 executing non -transitory, processor-executable instructions embodied in one or more tangible computer-readable storage media, such as memory 303, storage 308, storage devices335, and / or storage medium 336. The computer-readable media may store software that implements particular embodiments, and processor(s) 301 may execute the software. Memory 303 may read the software from one or more other computer-readable media (such as mass storage device(s) 335, 336) or from one or more other sources through a suitable interface, such as network interface 320. The software may cause processor(s) 301 to carry out one or more processes or one or more operations of one or more processes described or illustrated herein. Carrying out such processes or operations may include defining data structures stored in memory 303 and modifying the data structures as directed by the software.
[0188] The memory 303 may include various components (e.g., machine readable media) including, but not limited to, a random access memory component (e.g., RAM 304) (e.g., static RAM (SRAM), dynamic RAM (DRAM), ferroelectric random access memory (FRAM), phasechange random access memory (PRAM), etc.), a read-only memory component (e.g., ROM 305), and any combinations thereof. ROM 305 may act to communicate data and instructions unidirectionally to processor(s) 301, and RAM 304 may act to communicate data and instructions bidirectionally with processor(s) 301. ROM 305 and RAM 304 may include anyDocket No.: 69246-750.601suitable tangible computer-readable media described below. In one example, a basic input / output system 306 (BIOS), including basic routines that help to transfer information between elements within computer system 300, such as during start-up, may be stored in the memory 303.
[0189] Fixed storage 308 is connected bidirectionally to processor(s) 301, optionally through storage control unit 307. Fixed storage 307 provides additional data storage capacity and may also include any suitable tangible computer-readable media described herein. Storage 308 may be used to store operating system 309, executable(s) 310, data 311, applications 312 (application programs), and the like. Storage 308 can also include an optical disk drive, a solid-state memory device (e.g., flash-based systems), or a combination of any of the above. Information in storage 308 may, in appropriate cases, be incorporated as virtual memory in memory 303.
[0190] In one example, storage device(s) 335 may be removably interfaced with computer system 300 (e.g., via an external port connector (not shown)) via a storage device interface 325. Particularly, storage device(s) 335 and an associated machine-readable medium may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for the computer system 300. In one example, software may reside, completely or partially, within a machine-readable medium on storage device(s) 335. In another example, software may reside, completely or partially, within processor(s) 301.
[0191] Bus 340 connects a wide variety of subsystems. Herein, reference to a bus may encompass one or more digital signal lines serving a common function, where appropriate. Bus 340 may be any of several types of bus structures including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combinations thereof, using any of a variety of bus architectures. As an example and not by way of limitation, such architectures include an Industry Standard Architecture (ISA) bus, an Enhanced ISA (EISA) bus, a Micro Channel Architecture (MCA) bus, a Video Electronics Standards Association local bus (VLB), a Peripheral Component Interconnect (PCI) bus, a PCI -Express (PCI-X) bus, an Accelerated Graphics Port (AGP) bus, HyperTransport (HTX) bus, serial advanced technology attachment (SATA) bus, and any combinations thereof.
[0192] Computer system 300 may also include an input device 333. In one example, a user of computer system 300 may enter commands and / or other information into computer system 300 via input device(s) 333. Examples of an input device(s) 333 include, but are not limited to, an alpha-numeric input device (e.g., a keyboard), a pointing device (e.g., a mouse or touchpad), a touchpad, a touch screen, a multi-touch screen, a joystick, a stylus, a gamepad, an audio input device (e.g., a microphone, a voice response system, etc.), an optical scanner, a video or still image capture device (e.g., a camera), and any combinations thereof. In some embodiments, theDocketNo.: 69246-750.601input device is a Kinect, Leap Motion, or the like. Input device(s) 333 may be interfaced to bus 340 via any of a variety of input interfaces 323 (e.g., input interface 323) including, but not limited to, serial, parallel, game port, USB, FIREWIRE, THUNDERBOLT, or any combination of the above.
[0193] In particular embodiments, when computer system 300 is connected to network 330, computer system 300 may communicate with other devices, specifically mobile devices and enterprise systems, distributed computing systems, cloud storage systems, cloud computing systems, and the like, connected to network 330. Communications to and from computer system 300 may be sentthrough network interface320. For example, network interface320 may receive incoming communications (such as requests or responses from other devices) in the form of one or more packets (such as Internet Protocol (IP) packets) from network 130, and computer system 300 may store the incoming communications in memory 303 for processing. Computer system 300 may similarly store outgoing communications (such as requests or responses to other devices) in the form of one or more packets in memory 303 and communicated to network 330 from network interface 320. Processor(s) 301 may access these communication packets stored in memory 303 for processing.
[0194] Examples of the network interface 320 include, but are not limited to, a network interface card, a modem, and any combination thereof. Examples of a network 330 or network segment 330 include, but are not limited to, a distributed computing system, a cloud computing system, a wide area network (WAN) (e.g., the Internet, an enterprise network), a local area network (LAN) (e.g., a network associated with an office, a building, a campus or other relatively small geographic space), a telephone network, a direct connection between two computing devices, a peer-to-peer network, and any combinations thereof. A network, such as network 330, may employ a wired and / or a wireless mode of communication. In general, any network topology may be used.
[0195] Information anddata canbe displayed through a display 332. Examples of a display 332 include, but are not limited to, a cathode ray tube (CRT), a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT-LCD), an organic liquid crystal display (OLED) such as a passive-matrix OLED (PMOLED) or active-matrix OLED (AMOLED) display, a plasma display, and any combinations thereof. The display 332 can interface to the processor(s) 301, memory 303, and fixed storage 308, as well as other devices, such as input device(s) 333, via the bus 340. The display 332 is linked to the bus 340 via a video interface 322, and transport of data between the display 332andthebus340 canbe controlled via thegraphics control 321. In some embodiments, the display is a video projector. In some embodiments, the display is a head-Docket No.: 69246-750.601mounted display (HMD) such as a VR headset. In further embodiments, suitable VR headsets include, by way of non-limiting examples, HTC Vive, Oculus Rift, Samsung Gear VR, Microsoft HoloLens, Razer OSVR, FOVE VR, Zeiss VR One, Avegant Glyph, Freefly VR headset, and the like. In still further embodiments, the display is a combination of devices such as those disclosed herein.
[0196] In addition to a display 332, computer system 300 may include one or more other peripheral output devices 334 including, but not limited to, an audio speaker, a printer, a storage device, and any combinations thereof. Such peripheral output devices may be connected to the bus 340 via an output interface 324. Examples of an output interface 324 include, but are not limited to, a serial port, a parallel connection, a USB port, a FIREWIRE port, a THUNDERBOLT port, and any combinations thereof.
[0197] In addition or as an alternative, computer system 300 may provide functionality as a result of logic hardwired or otherwise embodied in a circuit, which may operate in place of or together with software to execute one or more processes or one or more operations of one or more processes described or illustrated herein. Reference to software in this disclosure may encompass logic, and reference to logic may encompass software. Moreover, reference to a computer-readable medium may encompass a circuit (such as an IC) storing software for execution, a circuit embodying logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware, software, or both.
[0198] The various illustrative logical blocks, modules, circuits, and algorithm operations described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and operations have been described above generally in terms of their functionality.
[0199] The various illustrative logical blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.Docket No.: 69246-750.601
[0200] The operations of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by one or more processor(s), or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium. An exemplary storage medium is coupled to the processor such the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.
[0201] In accordance with the description herein, suitable computing devices include, by way of non-limiting examples, cloud computing platforms, distributed computing platforms, server clusters, server computers, desktop computers, laptop computers, notebook computers, sub-notebook computers, netbook computers, and netpad computers.
[0202] In some embodiments, the computing device includes an operating system configured to perform executable instructions. The operating system is, for example, software, including programs and data, which manages the device’s hardware and provides services for execution of applications. Various suitable server operating systems include, by way of non -limiting examples, FreeBSD, OpenBSD, NetBSD®, Linux, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Various suitable personal computer operating systems include, by way of non -limiting examples, Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and UNIX -like operating systems such as GNU / Linux®. In some embodiments, the operating system is provided by cloud computing. Various suitable mobile smartphone operating systems include, by way of non -limiting examples, Nokia® Symbian® OS, Apple® iOS®, Research in Motion® BlackBerry OS®, Google® Android®, Microsoft® Windows Phone® OS, Microsoft® Windows Mobile® OS, Linux®, and Palm® WebOS®. Non-transitory Computer Readable Storage Medium
[0203] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more non-transitory computer readable storage media encoded with a program including instructions executable by the operating system of an optionally networked computing device. In further embodiments, a computer readable storage medium is a tangible component of a computing device. In still further embodiments, a computer readable storage medium is optionally removable from a computing device. In some embodiments, a computer readable storage medium includes, by way of non-limiting examples, CD-ROMs, DVDs, flash memoryDocket No.: 69246-750.601devices, solid state memory, magnetic disk drives, magnetic tape drives, optical disk drives, distributed computing systems including cloud computing systems and services, and the like. In some cases, the program and instructions are permanently, substantially permanently, semipermanently, or non-transitorily encoded on the media.Computer Programs
[0204] In some embodiments, the platforms, systems, media, and methods disclosed herein include at least one computer program, or use of the same. A computer program includes a sequence of instructions, executable by one or more processor(s) of the computing device’s CPU, written to perform a specified task. Computer readable instructions may be implemented as program modules, such as functions, objects, Application Programming Interfaces (APIs), computing data structures, and the like, which perform particular tasks or implement particular abstract data types. A computer program may be written in various versions of various languages.
[0205] The functionality of the computer readable instructions may be combined or distributed as desired in various environments. In some embodiments, a computer program comprises one sequence of instructions. In some embodiments, a computer program comprises a plurality of sequences of instructions. In some embodiments, a computer program is provided from one location. In other embodiments, a computer program is provided from a plurality of locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or combinations thereof.Software Modules
[0206] In some embodiments, the platforms, systems, media, and methods disclosed herein include software, server, and / or database modules, or use of the same. In view of the disclosure provided herein, software modules are created by various techniques using various machines, software, and languages. The software modules disclosed herein are implemented in a multitude of ways. In various embodiments, a software module comprises a file, a section of code, a programming object, a programming structure, a distributed computing resource, a cloud computing resource, or combinations thereof. In further various embodiments, a software module comprises a plurality of files, a plurality of sections of code, a plurality of programming objects, a plurality of programming structures, a plurality of distributed computing resources, a plurality of cloud computing resources, or combinations thereof. In various embodiments, the one or more software modules comprise, by way of non -limiting examples, a web application, aDocket No.: 69246-750.601mobile application, a standalone application, and a distributed or cloud computing application. In some embodiments, software modules are in one computer program or application. In other embodiments, software modules are in more than one computer program or application. In some embodiments, software modules are hosted on one machine. In other embodiments, software modules are hosted on more than one machine. In further embodiments, software modules are hosted on a distributed computing platform such as a cloud computing platform. In some embodiments, software modules are hosted on one or more machines in one location. In other embodiments, software modules are hosted on one or more machines in more than one location. Databases
[0207] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more databases, or use of the same. Various databases are suitable for storage and retrieval of information, for example customer incident data. In various embodiments, suitable databases include, by way of non-limiting examples, relational databases, non-relational databases, object oriented databases, object databases, entity -relationship model databases, associative databases, XML databases, document oriented databases, and graph databases. Further non-limiting examples include SQL, PostgreSQL, MySQL, Oracle, DB2, Sybase, and MongoDB. In some embodiments, a database is Internet-based. In further embodiments, a database is web-based. In still further embodiments, a database is cloud computing-based. In a particular embodiment, a database is a distributed database. In other embodiments, a database is based at least in part on one or more local computer storage devices.
[0208] In some embodiments, the system can comprise one or more processors further configured to generate information relating to genetic variations of the nucleic acid molecules. The information relating to genetic variations can comprise information relating to one or more of: single nucleotide variants (SNVs), single nucleotide polymorphisms (SNPs), double base substitutions, insertion-deletions (indels), copy number variations (CNVs), inversions, translocations, tandem repeats (TRs), short tandem repeats (STRs), variable number tandem repeats (VNTRs), quasi-tandem repeats (QTRs), or a combination thereof.
[0209] In some embodiments, the system can comprise one or more processors further configured to compare one or more data sets. In some embodiments, the system can comprise one or more models configured to compare the one or more data sets. The one or more models can be configured to perform evaluations of the one or more data sets, such as evaluations relating to genetic variations in the one or more data sets. The one or more models can be statistical models. The one or more models can be predictive models.Docket No.: 69246-750.601
[0210] In some embodiments, the system can comprise one or more processors further configured to generate sequencing data using the sequencing assay. In some embodiments, the system can comprise one or more processors further configured to compare the generated sequencing data to reference sequencing data to determine one or more mutations of the amplified nucleic acid molecules.
[0211] In some embodiments, the system can comprise one or more processors further configured to determine an error value of the sequencing data. In some cases, the error value can comprise one or more of: an error rate, an error percentage, an error proportion, or any other value representing an error of the generated sequencing data. The error value can be determined by comparing the generated sequencing data to the reference sequencing data. The error value can be determined by comparing generated sequence data of a first sample to generated sequence data of a second sample. The error value can be generated by one or more models. The error value can be generated by one or more statistical methods. The error value can be generated based at least in part on a comparison of a one or more data sets. The error value of the generated sequencing data can be reduced compared to an error value of otherwise identical generated sequencing data. In some cases, the error value can be reduced by between about 0.1% and about 1%. In some cases, the error value can be reduced by about 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.10%, 0.11%, 0.12%, 0.13%, 0.14%, 0.15%, 0.16%, 0.17%, 0.18%, 0.19%, 0.20%, 0.21%, 0.22%, 0.23%, 0.24%, 0.25%, 0.26%, 0.27%, 0.28%, 0.29%, 0.30%, 0.31%, 0.32%, 0.33%, 0.34%, 0.35%, 0.36%, 0.37%, 0.38%, 0.39%, 0.40%, 0.41%, 0.42%, 0.43%, 0.44%, 0.45%, 0.46%, 0.47%, 0.48%, 0.49%, 0.50%, 0.51%, 0.52%, 0.53%, 0.54%, 0.55%, 0.56%, 0.57%, 0.58%, 0.59%, 0.60%, 0.61%, 0.62%, 0.63%, 0.64%, 0.65%, 0.66%, 0.67%, 0.68%, 0.69%, 0.70%, 0.71%, 0.72%, 0.73%, 0.74%, 0.75%, 0.76%, 0.77%, 0.78%, 0.79%, 0.80%, 0.81%, 0.82%, 0.83%, 0.84%, 0.85%, 0.86%, 0.87%, 0.88%, 0.89%, 0.90%, 0.91%, 0.92%, 0.93%, 0.94%, 0.95%, 0.96%, 0.97%, 0.98%, 0.99%, or about 1.00%, or more than about 1.00%.
[0212] In some cases, the error value can be reduced by between about 1% and about 50%. In some cases, the error value can be reduced by about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, or about 50%, or more than about 50%.
[0213] In some cases, the error value can be reduced by between about 50% and about 99%. In some cases, the error value can be reduced by about 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%,Docket No.: 69246-750.60174%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or about 99%, or more than about 99%.
[0214] In some cases, the error value can be reduced by about 50%. In some cases, the error value can be reduced by about 60%. In some cases, the error value can be reduced by about 70%. In some cases, the error value can be reduced by about 80%. In some cases, the error value can be reduced by about 90%. In some cases, the error value can be reduced by about 95%. In some cases, the error value can be reduced by about 99%. In some cases, the error value can be reduced by about 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, or about 99.9%, or more than about 99.9%.Compositions of Capture Probes for Use in Analysis of Low-Input Samples
[0215] In yet another aspect, the present disclosure provides a composition comprising a pool of capture probes. The pool of capture probes may comprise capture probes that may be capable of hybridizing to at least a portion of an exome as shown in FIG.2. The pool of capture probes may comprise capture probes that may be capable of hybridizing to at least a portion of a biomarker. The pool of capture probes may comprise capture probes capable of hybridizing to a single nucleotide polymorphism (SNP). The capture probes may comprise customized capture probes. The pool of capture probes may comprise at least two of the capture probes. The at least two capture probes maybe capable of hybridizing to at least a portion of an exome. The at least two capture probes may be capable of hybridizing to at least a portion of a biomarker. The least two capture probes may be capable of hybridizing to an SNP. The pool of capture probes may comprise the capture probes capable of hybridizing to at least a portion of an exome. The capture probes may be capable of hybridizing to at least a portion of a biomarker. The capture probes may be capable of hybridizing to a single nucleotide polymorphism (SNP).Kits for Use in Analysis of Low-Input Samples
[0216] In yet another aspect, the present disclosure provides a kit comprising the composition and instructions for use according to any one of the methods described herein .
[0217] In some embodiments, provided herein are methods comprising providing a first sample comprising cancer DNA from the subject. In some embodiments, the method can comprise obtaining a biological sample containing cancer cells from the subject. The sample may be collected during initial diagnosis or treatment of the cancer. The sample may be a tumor biopsy obtained during surgical resection of a primary tumor from a subject diagnosed with a specific cancer.Docket No.: 69246-750.601
[0218] The cancer may comprise adrenal gland cancer, kidney cancer, aerodigestive tract cancer, biliary tract cancer, breast cancer, liver cancer, ovarian cancer, lung cancer, leukemia, lymphoma, salivary gland cancer, renal cancer, bladder cancer, brain cancer, head and neck cancer, prostate cancer, skin cancer, pancreatic cancer, cervical cancer, colorectal cancer, testicular cancer, thyroid cancer, bile duct cancer, central nervous system cancer, or esophageal cancer. The cancer may comprise a stage of a cancer, for example, stage 0 cancer, stage I cancer, stage II cancer, stage III cancer, or stage IV cancer. The cancer may comprise a hematologic cancer, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), acute promyelocytic leukemia (APL), anaplastic large cell lymphoma (ALCL), Burkitt lymphoma, chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), diffuse large B-cell lymphoma (DLBCL), eosinophilia, follicular lymphoma, hairy cell leukemia (HCL), Hodgkin lymphoma, large granular lymphocytic leukemia (LGL), MALT lymphoma, mantle cell lymphoma (MCL), marginal zone B-cell lymphoma (MZL), Mastocytosis, myelodysplastic syndrome (MDS), myeloproliferative neoplasm (MPN), non -Hodgkin lymphoma (NHL), plasma cell myeloma, PNH, t-cell lymphoma, or Waldenstrom macroglobulinemia.
[0219] In some embodiments, the method can comprise collecting a sample from a subject. In some embodiments, the sample is collected using clinical procedures appropriate for the tumor type and location, such as fine needle aspiration, core needle biopsy, or surgical excision. In some embodiments, the sample can be preserved using appropriate methods such as flash freezing in liquid nitrogen or fixation in formalin followed by embedding in paraffin (FFPE). In some embodiments, the preserved sample can be transported to a laboratory facility equipped for genomic analysis under controlled temperature conditions to maintain sample integrity.
[0220] In some embodiments, a first sample can comprise a tumor sample or a surgical sample. The tumor sample may be obtained during the initial diagnosis or treatment of the subject's cancer. For solid tumors, this may involve a surgical procedure such as a biopsy or resection. During a biopsy procedure, a small portion of the tumor is removed using techniques such as fine needle aspiration, core needle biopsy, or endoscopic biopsy, depending on the tumor location and type. During a surgical resection, the entire tumor or a substantial portion of it is removed as part of the therapeutic intervention. In either case, a portion of the removed tissue may be allocated for genomic analysis. The sample may be handled according to pathology protocols, which may include immediate flash freezing in liquid nitrogen to preserve nucleic acid integrity or fixation in formalin followed by embedding in paraffin (FFPE). For hematological malignancies, the sample may consist of bone marrow aspirate or peripheral blood containing leukemic cells. The use of tumor tissue as the first sample may provide direct access to theDocket No.: 69246-750.601cancer cells, enabling comprehensive characterization of the genomic alterations present in the subject's cancer and the definition of a robust digital signature for subsequent MRD detection.
[0221] In some embodiments, the second sample can comprise cell-free DNA (cfDNA). The cfDNA can comprise extracellular DNA fragments that circulate in bodily fluids, primarily blood. In cancer subjects, a fraction of cfDNA may be derived from tumor cells (circulating tumor DNA or ctDNA). In some embodiments, a blood sample may be collected from the subject using phlebotomy procedures, such as drawing 10-20 mL of blood into cell-free DNA collection tubes containing stabilizing agents that prevent lysis of white blood cells and consequent contamination with genomic DNA. cfDNA may be extracted from the plasma using extraction kits designed for recovery of low-abundance, fragmented DNA, e.g., yielding between 1-100 ng of cfDNA depending on the plasma volume and the patient's physiological state. The cfDNA may be prepared for sequencing with library preparation protocols optimized for low-input, fragmented DNA samples. The use of cfDNA for MRD detection offers a minimally invasive approach for longitudinal monitoring of subjects after cancer treatment, enabling repeated samplingthat may not be feasible with tissue biopsies. Less than about 50 ng of DNA may be sufficient for sample processing (e.g., library preparation, sample enrichment) and sequencing.
[0222] In some embodiments, library preparation utilizing single stranded DNA (ssDNA) produces fragment size distributions different from canonical double stranded library preparation. In some embodiments, ssDNA library preparation yields non -normally distributed fragment sizes of ssDNA fragments. The ssDNA library preparation may yield multiple populations of fragments resulting from increased fragmentation time. Library quality may be improved across sample types and result in more efficient and accurate sequencing performance.
[0223] In some embodiments, more accurate and efficient sequencing performance is achieved for poor quality samples which show failed performance on standard of care assays in the field. At least about 80% of samples were assayed that would not have been assayed otherwise due to low sample volume or low sample quality.
[0224] In some embodiments, a biological sample comprising at least 10 ng of FFPE DNA may be analyzed. In some embodiments, a sample comprising at least 10 ng of FFPE DNA may be analyzed and have a low fail rate in a cancer sample in comparison to assays requiring 50 ng or more of DNA fragments. In some embodiments, the input sample yields results sufficient for analyzing.
[0225] In some embodiments, a genomic assay such as a genetic assay or epigenetic assay may assay a sample comprising about 10 ng of DNA. The genomic assay may have from about 3 toDocket No.: 69246-750.601about 5 times lower DNA input failure rate than an otherwise similar assay requiring 50 ng of DNAforthe assay. For example, a sample comprising about 10 ng of DNA may yield more high quality sequencing reads sufficient for analysis as compared to a sample comprising about 50 ng of DNA.
[0226] The methods described herein may survey the entire exome. In some embodiments, more than 1,000 genes are analyzed. In some embodiments, between about 1,000 genes and 10,000 genes ae analyzed. In some embodiments, about 1000 genes, 1500 genes, 2000 genes, 2500 genes, 3000 genes, 3500 genes, 4000 genes, 4500 genes, 5000 genes, 5500 genes, 6000 genes, 6500 genes, 7000 genes, 7500 genes, 8000 genes, 8500 genes, 9000 genes, 9500 genes, or about 10000 genes, or more than about 10,000 genes are analyzed. In some embodiments, the analysis using the assays can comparatively generate a significantly increased amount of result data. The analysis can comprise an increase in the number of genes being surveyed. In some embodiments, hybrid capture target enrichment methods may be utilized to increase the number of genes analyzed. In some embodiments, the methods utilize exome baits, enhancement (e.g., region of interest) baits, SNP baits, or a combination thereof. In some embodiments, the term enhancement baits may comprise enhanced baits. In some embodiments, enhanced baits may achieve coverage of at least 1,000 genes. In some embodiments, between about 1,000 genes and 10,000 genes ae analyzed. In some embodiments, about 1000 genes, 1500 genes, 2000 genes, 2500 genes, 3000 genes, 3500 genes, 4000 genes, 4500 genes, 5000 genes, 5500 genes, 6000 genes, 6500 genes, 7000 genes, 7500 genes, 8000 genes, 8500 genes, 9000 genes, 9500 genes, or about 10000 genes, or more than about 10,000 genes may be covered by the assay. In some embodiments, coverage can comprise 5 times tiling across exons at respective genes. In some embodiments, a plurality of baits targeting SNPs across the genome were utilized in addition to enhanced baits and / or exome baits.
[0227] In some embodiments, the expansion in the number of genes being surveyed results in a broader assessment of current biomarkers that rely on a breadth of genomic information. The genomic information such as one or more of: tumor mutational burden, copy number variations or micro satellite instability, or any combination thereof. In some embodiments, the genes can comprise encoding genes. In some embodiments, a portion of the genes can comprise encoding genes. In some embodiments, all of the genes can comprise encoding genes. In some embodiments, the expanded number of genes being surveyed by the assays facilitates generation by the assays of additional, new, and emerging biomarkers. The emerging biomarkers comprise one or more of : tumor mutational burden (TMB), microsatellite instability (MSI), or homologous recombination deficiency (HRD), or any combination thereof.Docket No.: 69246-750.601
[0228] In some embodiments, the methods further comprise performing union analyses. The union analysis can be performed for development of expression -informed HRD or neoantigens.
[0229] The DNA molecules in cell-free DNA can be highly fragmented. In some embodiments, the cfDNA fragments may comprise a median size that is below 1 kb (e.g., in the range of 50 bp to 500 bp, 80 bp to 400 bp, or 100-1 ,000bp), although fragments having a median size outside of this range may be present. In some embodiments, cfDNA has a mean fragment size of about 100-250 bp, e.g., 150 to 200 bp long, or about 160 bp.
[0230] To provide a first sample, DNA may be extracted from a cancer sample using a nucleic acid extraction protocol appropriate for the sample type. For FFPE samples, an extraction kit designed to overcome formalin-induced DNA modifications and fragmentation may be used. In some embodiments, FFPE and heme sample types can be analyzed by the same assay. In some embodiments, FFPE and heme sample types can be assayed separately and the resulting data can be analyzed together.
[0231] In some embodiments, the extracted DNA may undergo quality control assessment (e.g., including quantification using fluorometric methods and / or fragment size analysis). In some embodiments, the DNA is prepared for sequencing using any suitable library preparation protocols that may include DNA fragmentation, end repair, adapter ligation, end blunting, amplification, or a combination thereof . In some embodiments, the prepared library undergoes WGS. In some embodiments, WGS is performed on a high-throughput sequencing platform. The sequencing may be performed at sufficient depth. In some cases, the sequencing may be performed atbetween about 1X-2O,OOOX coverage for detection of somatic mutations present at various allele frequencies within the tumor DNA. In some cases, the sequencing may be performed at about IX, 2X, 3X, 4X, 5X, 6X, 7X, 8X, 9X, 10X, or more than 10X coverage. In some embodiments, the target depth for sequencing can be between about 1 OX and 1000000X. In some embodiments, the target depth for sequencing can be between about 10X and 100X. In some embodiments, the target depth for sequencing can be about 10X, 1 IX, 12X, 13X, 14X, 15X, 16X, 17X, 18X, 19X, 20X, 21X, 22X, 23X, 24X, 25X, 26X, 27X, 28X, 29X, 30X, 31X, 32X, 33X, 34X, 35X, 36X, 37X, 38X, 39X, 40X, 41X, 42X, 43X, 44X, 45X, 46X, 47X, 48X, 49X, 50X, 51X, 52X, 53X, 54X, 55X, 56X, 57X, 58X, 59X, 60X, 61X, 62X, 63X, 64X, 65X, 66X, 67X, 68X, 69X, 70X, 71X, 72X, 73X, 74X, 75X, 76X, 77X, 78X, 79X, 80X, 81X, 82X, 83X, 84X, 85X, 86X, 87X, 88X, 89X, 90X, 9 IX, 92X, 93X, 94X, 95X, 96X, 97X, 98X, 99X, 100X, or more than 100X. In some embodiments, the target depth for sequencing can be between about 100X and 100000X. In some embodiments, the target depth for sequencing can be between about 100X and 1000X. In some embodiments, the target depth for sequencing can beDocket No.: 69246-750.601about 100X, 150X, 200X, 250X, 300X, 350X, 400X, 450X, 500X, 550X, 600X, 650X, 700X, 750X, 800X, 850X, 900 X, 950X, 1000X, or more than 1000X. In some embodiments, the target depth for sequencing can be between about 1000X and 10000X. In some embodiments, the target depth for sequencing can be about 1000X, 1500X, 2000X, 2500X, 3000X, 3500X, 4000X, 4500X, 5000X, 5500 X, 6000X, 6500X, 7000X, 7500X, 8000X, 8500X, 9000X, 9500X, 10000X, or more than 10000X. In some embodiments, the target depth for sequencing can b e between about lOOOOXand 100000 X. In some embodiments, the target depth for sequencing can be about 10000X, 15000X, 20000X, 25000X, 30000X, 35000X, 40000X, 45000X, 50000X, 55000X, 60000X, 65000X, 70000X, 75000X, 80000X, 85000X, 90000X, 95000X, 100000X, or more than 100000X.
[0232] The sequencing may generate millions of short DNA sequence reads that may be aligned to a human reference genome (e.g., using bioinformatics alignment algorithms).
[0233] In some embodiments, a cfDNA sample may comprise circulating tumor DNA (ctDNA). ctDNA is of tumor origin and originates directly from the tumor or from circulating tumor cells (CTCs), which are viable, intact tumor cells that shed from primary tumors and can enter the bloodstream or lymphatic system. The precise mechanism of how cancer DNA is released is unclear, although it is postulated to involve apoptosis and necrosis from dying cells, or active release from viable tumor cells. The amount of ctDNA in a sample of circulating cell -free DNA isolated from a cancer subject varies greatly. In some cases, at least a portion of the samples may contain less than 10% ctDNA, although many samples from subjects being assessed for MRD may have less than 0.01% ctDNA and some samples may have over 10% ctDNA. In some cases, molecules of cancer DNA can be identified based at least in part on contained tumorigenic mutations.
[0234] In some embodiments, the second sample can comprise circulating tumor cells (CTCs). The CTCs can comprise cancer cells that have detached from the primary tumor or metastatic sites and entered the bloodstream or lymphatic system. For MRD detection using CTCs, a blood sample may be collected from the subject using phlebotomy procedures. The blood sample may be processed using CTC enrichment techniques to isolate CTCs from the much more abundant normal blood cells. Enrichment methods may include immunomagnetic separation using antibodies against epithelial cell adhesion molecule (EpCAM) or other tumor-associated surface markers, size-based filtration (e.g., exploiting a larger size of CTCs compared to white blood cells), or microfluidic approaches that utilize multiple physical and biological properties of CTCs. In some embodiments, CTC analysis offers the advantage of assessing intact cancer cellsDocket No.: 69246-750.601rather than cell-free DNA fragments, providing additional information about the phenotypic characteristics of residual disease.
[0235] In some embodiments, DNA may be extracted from a low input sample (e.g., CTCs) and undergo whole genome amplification to generate sufficient material for sequencing. The amplified DNA may then be sequenced using targeted approaches focusing on the mutations included in the subject's digital signature. The detection of these mutations (e.g., cancer-specific mutations) in the low input sample may provide evidence for the presence of MRD. CTC analysis offers the advantage of assessing intact cancer cells rather than cell-free DNA fragments, potentially providing additional information about the phenotypic characteristics of residual disease.
[0236] In some embodiments, the second sample can comprise one or more of: a surgical sample, a biopsy sample, a plasma sample, a serum sample, a bone marrow sample, a dried blood spot sample, a cerebral spinal fluid sample, a fine needle aspirate sample, a pleural fluid sample, a saliva sample, a stool sample, or a urine sample, or any combination thereof. The sample type may depend on the cancer type, the expected routes of metastasis or recurrence, and practical considerations regarding sample collection. For suspected local recurrence, a surgical sample or biopsy sample may be obtained from the site of concern, providing direct access to potential residual disease. The sample may comprise a fine needle aspirate sample. The fine needle aspirate sample may originate from a lung of the subject. For systemic monitoring, liquid biopsies such as plasma or serum samples may be used, offering a minimally invasive approach that can capture ctDNA released from multiple potential disease sites. For hematological malignancies, bone marrow samples may provide a sensitive assessment of residual disease in the primary site of involvement. Dried blood spot samples may offer a convenient collection method that can be performed at home by the subject and mailed to the laboratory, thereby enabling more frequent monitoring. For example, for cancers affecting the central nervous system, cerebral spinal fluid samples may provide higher sensitivity for detecting residual disease than peripheral blood. For example, for lung cancers or malignancies with pleural involvement, pleural fluid samples may offer enriched detection of cancer-derived nucleic acids. For example, for head and neck cancers or oral lesions, saliva samples may provide localized assessment of residual disease. For example, for colorectal cancers, stool samples may contain exfoliated tumor cells or tumor DNA from residual disease in the gastrointestinal tract. For example, for urological cancers, urine samples may contain cancer cells or cell -free DNA from residual disease in the urinary tract. Regardless of the sample type, appropriate collection, preservation, and processing protocols may be employed to maintain the integrity of nucleicDocket No.: 69246-750.601acids for subsequent analysis. The sample may undergo DNA extraction using methods optimized for the specific sample type, followed by sequencing.
[0237] In some embodiments, the test sample can comprise a blood plasma sample. In some embodiments, cell-free DNA (cfDNA) can be isolated from the blood plasma sample. In some embodiments, the fraction of cancer DNA in the test sample (compared to non -cancerous DNA) may be equal or less than 0.01%, equal or less than 0.002%, equal or less than 0.005%, or equal or less than 0.001%. In some embodiments, a detectable fraction of cancer DNA in the test sample ofDNAmay befromaboutO.0001%, however the actual limit of detection may vary. In some embodiments, the test sample comprises less than 25,000 genome equivalents of DNA (e.g., cfDNA), e.g., less than 20,000, less than 10,000, less than 5,000, or less than 1 ,000 genome equivalents of DNA. In some embodiments the test sample comprises from about 100 to about 25,000 genome equivalents (i.e., enrichable or amplifiable copies) of DNA. In some embodiments, the test sample comprises from about 1 Ong to about lOOng of DNA. In some embodiments, the test sample comprises at least lOng, at least 20ng, at least 30ng, at least 40ng, at least 50ng, at least 60ng, at least 70ng, at least 80ng, at least 90ng, or at least lOOng of DNA. In some embodiments, the test sample comprises 66ng of DNA.EXAMPLES
[0238] The following examples are provided to further illustrate some embodiments of the present disclosure, but are not intended to limit the scope of the disclosure; it will be understood by their exemplary nature that other procedures, methodologies, or techniques may alternatively be used.Example 1: Uniformity for ssDNA Library Preparation for Hybridization Capture of WES Assay
[0239] As illustrated in FIG. 4, coverage and other metrics are required for a successful NGS analysis. Highly uniform DNA libraries allow clinical labs to generate more hits from sample screens, saving both time and money across the course of diagnosis thereby ultimately benefiting patient care. The commercially available single stranded DNA library prep kit tested to generate the data shown in FIG. 4 renders input molecules single-stranded before ligation therefore it recovers template molecules containing DNA nicks and labile lesions thus maximizing library complexity. 10 solid tumor FFPE derived DNA samples went through three different NGS workflows: lOOng of DNA input in an amplicon-based assay 325 gene panel (Assay A), 60ng input for a 500 gene mechanical sheared-based double strand DNA library preparation withDocket No.: 69246-750.601hybridization capture assay (-500 gene panel, Assay B), 5 Ong and 1 Ong input for an enzymatic shear-based single stranded DNA library preparation with hybridization capture assay (WES with 1000+ enhanced gene content, Assay C). As shown in FIG. 5, data from assays A and C were processed with in-house bioinformatic pipeline, while Assay B’s data was processed with a commercially available bioinformatic pipeline that was bundled with the assay. NGS QC metrics were compared across the three datasets.
[0240] As shown in FIG. 5, this single stranded DNA library preparation with hybridization capture workflow (Assay C) showed superior QC metrics for uniformity (average 99%) and percent of target with at least lOOx coverage (average 100%) compared to the other two workflows even with as low as 1 Ong input.Example 2: Analysis of Heme and Solid Tumor Samples at Low Input
[0241] Lab developed tests ‘LDT’ simultaneously evaluate many different genomic alterations. Utilizing nucleic acid extraction, library preparation, and next generation sequencing to address limited sample quantity and variable quality , which often severely limits patient access to LDTs. DNA and RNA were dual extracted from a single sample. A novel DNA library generation method with unique dual indexes and molecular identifier sequences were used to create the novel assay. The DNA libraries underwent hybrid capture and sequencing. Additionally, the RNA is used to perform RNAseq. Libraries were sequenced on NovaSeq 6000 and bioinformatically processed with custom pipelines for variant calling for SNVs, InDeis, TMB, MSI, CNVs, and HLA profiling. We compare FFPE and Heme performance to validated amplicon assays.
[0242] Over 800 libraries were generated from more than 700 unique clinical samples, including 100 heme samples from peripheral blood or bone marrow and more than 600 FFPE samples from breast, colon, lung, lymph, ovarian, and prostate (ranging from 20 to 90% tumor). The workflow illustrated in FIG. 2 facilitates unified processing of heme and FFPE samples down to lOng input. Poor-quality samples that did not yield sufficient nucleic acid or failed QC in the amplicon assays were rescued. As shown in FIGs. 6A-6D, samples met a minimum of 500x average coverage at over 1,000 cancer genes and 80x average coverage across the exome. As shown in FIG. 7A, variant calling to 5% AF reached 99% accuracy (98% sensitivity and 99% specificity). TMB concordance met 93% accuracy (95% sensitivity and 91% specificity). MSI detection reached 97% accuracy (95% sensitivity and 100% specificity). As shown in FIGs.7B-7D, CNVs were detected at 100% accuracy with high correlations for MET and ERBB2 amplifications in solid tumors (Pearson, r = 0.999 and r = 0.997, p value <= 1.9x10-13).Docket No.: 69246-750.601
[0243] As shown in FIG. 7E, Poor quality samples with bad uniformity (lOOng) on another NGS assay and formalin damaged controls had good uniformity on this assay (5 and 20ng) and matched cell line controls at 50ng input. 5 additional samples failed the other assay (low coverage) butpassed QC forthis assay. As shown in FIG. 7F, HLA genotyping of 11 loci was performed to four-digit resolution in heme and FFPE samples and reached 98% concordance.Example 3: SNV and InDei Detection in an Exome NGS Assay with Novel Features
[0244] Most known disease-causing mutations occur in exons, therefore Whole Exome Sequencing (WES) is an appropriate method to identify the vast majority of disease-causing mutations. WES target regions may range from 35-45 Mbp, which makes high coverage expensive. As shown in FIG.8A, sequencing clinically relevant targets at deeper coverages for accurate SNV / InDel calling as an enhancement (average depth > 500x, Region 1) to WES, while keeping targets outside of the Region 1 at lower coverage (average coverage > 15 Ox, Region 2) is a method that can be performed. In order to evaluate the feasibility of SNV and InDei calling in such a scenario, the biological samples were sequenced at 3500x to create a “Truth Set” variant list within the Region 1, and applied it to a hybridization capture WES with Region 1 at 500x, and again to the same samples with WES only (150x). As shown in FIG. 8B, 50ng of DNA was used for detecting somatic SNV / Indel variants in both Region 1 and 2, with as low as 30% tumor purity. Additionally, we incorporated a novel single stranded library prep method that enabled an agnostic workflow for hematological and FFPE sample types.
[0245] The Truth Set of SNPs and InDeis for eight clinical samples was created by sequencing Region 1 at > 3500x with 200ng DNA. The same samples underwent WES with Region 1 (for Region 1) and WES only (for simulating Region 2) library construction with 50ng DNA and sequenced. These samples served as ‘experimental samples’ for detection rate of the Truth Set variants. All sequencing data was processed with an in house bioinformatic pipeline for SNV / Indel calling. 500x average coverage for Region 1 and 150x for WES simulating Region 2 were achieved through down-sampling and compared to the Truth Set for sensitivity and specificity calculation. As illustrated in FIGs. 8C-8F, for the truth set, the SNV / Indel calling cutoff was AF>= 5% and DP >= 700; whereas for Region 1 it was AF>= 5%, DP >= 100, and AltRead >= 10, and for Region 2 it’s AF>= 10% and DP >= 50.
[0246] As illustrated in FIGs. 8C-8F, SNV / Indel calling with 50ng input in Region 1 was performed, the assay achieved 98-99% sensitivity and 97-99% specificity. In Region 2, the sensitivity was 94-98%, with the specificity being 97-99%. With as low as lOng input, Region 1Docket No.: 69246-750.601sensitivity was 97 - 98%, and specificity was 97 - 99%. In Region 2, lOng input sensitivity was 93 - 98%, and the specificity was 97 - 99%.Example 4: Development of a WTS Assay with Novel Features
[0247] As illustrated in FIGs. 9-10, a WTS workflow was developed and performed. Table 1 shows data relatingto MET exon 14, EGFR vIII or Epstein Barr Virus (EBV) calling for various fragments of a variety of samples.Table 1 : WTS calling for various samples
[0248] As shown in Table 1, the WTS workflow had an accuracy of 100% in determining Exon 14 skips, EGFR vIII (a mutant form of the epidermal growth factor receptor), and EBV. As shown in FIGs. 11A-11B, the WTS workflows were also highly accurate at detecting various subtypes of human papilloma virus (HPV). FIG. 12 showsan example of an HPV structure and various genetic regions.
[0249] In an additional experiment, a challenging input was tested to review the workflow’s accuracy in detecting novel fusions. Total RNA was extracted from a lymph node fine needle aspirate (FNA) sample using the Omega method. RNA yield was 85 Ong, RNA concentration was approximately 17ng / ul: RIN:2; DV20050-70%. As shown in Table 2, the method was able to detect a TNS4-CCR7 fusion.Table 2: WTS fusion calling for FNA sampleDocket No.: 69246-750.601
[0250] Table 3 shows a summary of the sample input.Table 3: FNA sample information>>
[0251] As shown in FIG. 13, WTS offers more comprehensive transcriptome analysis by detecting up to -55,000 different transcripts in solid tumor-derived FFPE samples (lower panels) as compared with typical exome capture which detected -38,000 transcripts. Both strategies detect all transcripts targeted by the exome (left panels), but WTS detects almost double the number of additional transcripts, including alternative splicing forms, short and long non -coding RNAs, etc.
[0252] Similar to solid tumors, FIG 14 illustrates that WTS offers more comprehensive transcriptome overview in hematological cancers by detecting up to -52,000 different transcripts in leukemia, lymphoma, myeloma and other heme samples (lower panels) as compared with typical exome capture which detected -39,000 transcripts. Both strategies detect all transcripts targeted by the exome (left panels), but WTS detects almost double the number of additional transcripts, including alternative splicing forms, short and long non-coding RNAs, etc.Docket No.: 69246-750.601
[0253] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are notmeantto be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
Docket No.: 69246-750.601CLAIMSWhat is claimed is:
1. A method of preparing an enriched nucleic acid sample, the method comprising:a) obtaining a biological sample obtained or derived from a subject, wherein said biological sample is about 10 nanograms to about 100 nanograms; b) processing said biological sample to provide a plurality of single stranded deoxyribonucleic acid (ssDNA) fragments;c) amplifying said plurality of ssDNA fragments in two or more cycles of polymerase chain reaction (PCR), thereby producing a sequencing library; and d) contacting said sequencing library with a pool of capture probes, wherein said pool of capture probes comprises capture probes that are capable of hybridizing to at least a portion of an exome, capture probes that are capable of hybridizing to at least a portion of a biomarker, or capture probes that are capable of hybridizing to a single nucleotide polymorphism (SNP), or a combination thereof, to produce said enriched nucleic acid sample.
2. The method of claim 1, wherein the sequencing library comprises a double -stranded deoxyribonucleic acid (dsDNA) library.
3. The method of claim 1, wherein said biological sample is about 10 nanograms.
4. The method of claim 1, wherein said biological sample is about 25 nanograms.
5. The method of claim 1, wherein said biological sample is about 50 nanograms.
6. The method of claim 1, wherein the biological sample is about 100 nanograms.
7. The method of claim 1, wherein said biological sample comprises two or more sample types.
8. The method of claim 7, wherein the two or more sample types comprises two or more of a Formalin-Fixed Paraffin Embedded (FFPE) sample, a fine-needle aspirate (FNA) sample, a blood sample, a plasma sample, a serum sample, or a bone marrow sample.
9. The method of claim 1, wherein said biological sample comprises a Formalin -Fixed Paraffin Embedded (FFPE) sample, a fine-needle aspirate (FNA) sample, a blood sample, a plasma sample, a serum sample, or a bone marrow sample.
10. The method of claim 9, wherein said biological sample is said FFPE sample.Docket No.: 69246-750.60111. The method of claim 9, wherein said biological sample is said bone marrow sample.
12. The method of claim 9, wherein said biological sample is said blood sample, wherein said blood sample is a peripheral blood sample.
13. The method of claim 12, wherein said biological sample is said peripheral blood sample.
14. The method of claim 1, wherein said biological sample comprises a cell line.
15. The method of claim 14, wherein said cell line is commercially available.
16. The method of claim 1, wherein the biological sample is obtained or derived from a solid tumor of the subject.
17. The method of claim 1, wherein said biological sample is obtained or derived from a biopsy.
18. The method of claim 17, wherein said biopsy is a liquid biopsy.
19. The method of claim 17, wherein said biopsy comprises a fine needle aspirate or a core needle biopsy.
20. The method of claim 19, wherein said biopsy is said fine needle aspirate.
21. The method of claim 19, wherein said biopsy is said core needle biopsy.
22. The method of claim 1, wherein said biological sample comprises ribonucleic acid (RNA), deoxyribonucleic acid (DNA), or a combination thereof.
23. The method of claim 22, wherein said biological sample comprises said RNA.
24. The method of claim 23, wherein the RNA is cell-free RNA.
25. The method of claim 23, wherein the RNA is exosomal RNA.
26. The method of claim 23, wherein the RNA is coding RNA (cRNA) or non-coding RNA (ncRNA).
27. The method of claim 22, wherein said biological sample comprises said DNA.
28. The method of claim 27, wherein said DNA is cell-free DNA or DNA from cells.
29. The method of claim 27, wherein said DNA is said ssDNA, double-stranded DNA (dsDNA), or a combination thereof.
30. The method of claim 29, wherein said DNA comprises said ssDNA.
31. The method of claim 29, wherein said DNA comprises said dsDNA.Docket No.: 69246-750.60132. The method of claim 27, wherein said DNA comprises one or more DNA fragments comprising a nick.
33. The method of claim 27, wherein said DNA comprises end-damaged DNA.
34. The method of claim 27, wherein said DNA comprises one or more single stranded DNA (ssDNA) overhangs on a double-stranded DNA (dsDNA).
35. The method of any one of claims 22-34, wherein the processing further comprises extracting the RNA.
36. The method of claim 35, wherein the processing further comprises generatingthe dsDNA from the extracted RNA.
37. The method of claim 36, wherein the dsDNA comprises double -stranded circular DNA (dscDNA).
38. The method of claim 36 or 37, further comprising denaturing the dsDNA or the dscDNA.
39. The method of claim 38, further comprising assaying the denatured dsDNA or dscDNA.
40. The method of claim 39, wherein the assaying comprises performing WES on the denatured dsDNA.
41. The method of claim 39, wherein the assaying comprises performing WTS on the denatured dscDNA.
42. The method of claim 1, wherein said amplifying comprises three or more, four or more, five or more, six or more, seven or more, eight or more, or nine or more cycles of PCR.
43. The method of claim 42, wherein said amplifying comprises eight cycles of PCR.
44. The method of claim 42, wherein said amplifying comprises five to nine cycles of PCR.
45. The method of claim 42, wherein said amplifying comprises six to nine cycles of PCR.
46. The method of claim 42, wherein said amplifying comprises six to eight cycles of PCR.
47. The method of claim 1, wherein said amplifying further comprises ligating one or more adapters to at least a portion of said plurality of ssDNA fragments.
48. The method of claim 47, wherein said one or more adapters comprise directional adapters.
49. The method of claim 47, wherein said one or more adapters comprise non -directional adapters.Docket No.: 69246-750.60150. The method of claim 47, wherein the one or more adapters comprise Unique Molecular Index (UMI) adapters.
51. The method of claim 1, further comprising ligating one or more UMI adapters to the at least subset of the plurality of ssDNA fragments prior to the amplification in (c).
52. The method of claim 1, further comprising ligating one or more UMI adapters to the at least subset of the plurality of ssDNA fragments after the amplification in (c).
53. The method of claim 1 , wherein said amplification further comprises providing a Unique Molecular Index (UMI) sequence to said ssDNA fragments.
54. The method of claim 53, wherein said UMI sequence is provided via primer extension or via PCR.
55. The method of claim 53, wherein said UMI sequence comprises a length of about 5 nucleotides to about 15 nucleotides.
56. The method of claim 55, wherein said UMI sequence comprises a length of about 9 nucleotides.
57. The method of claim 1, wherein said pool of capture probes comprises atleasttwo of said capture probes that are capable of hybridizing to at least a portion of an exome, said capture probes that are capable of hybridizing to at least a portion of a biomarker, or said capture probes that are capable of hybridizing to a single nucleotide polymorphism (SNP).
58. The method of claim 1, wherein said pool of capture probes comprises said capture probes that are capable of hybridizing to at least a portion of an exome, said capture probes that are capable of hybridizing to at least a portion of a biomarker, and said capture probes that are capable of hybridizing to a single nucleotide polymorphism (SNP).
59. The method of claim 1, wherein said exome is a human exome.
60. The method of claim 1, wherein said pool of capture probes comprises atleasttwo of said capture probes that are capable of hybridizing to at least a portion of a transcriptome.
61. The method of claim 60, wherein the transcriptome is a human transcriptome.
62. The method of claim 1, wherein said pool of capture probes comprises said capture probes that are capable of hybridizing to at least a portion of a transcriptome.Docket No.: 69246-750.60163. The method of claim 1, wherein said transcriptome is a human transcriptome.
64. The method of claim 1, wherein said biomarker comprises a plurality of biomarkers.
65. The method of claim 64, wherein said plurality of biomarkers comprises at least 100 biomarkers, at least 200 biomarkers, at least 300 biomarkers, at least 400 biomarkers, at least 500 biomarkers, at least 600 biomarkers, at least 700 biomarkers, at least 800 biomarkers, at least 900 biomarkers, at least 1,000 biomarkers, at least 1,100 biomarkers, at least 1,200 biomarkers, at least 1,300 biomarkers, at least 1,400 biomarkers, or at least 1,500 biomarkers.
66. The method of claim 65, wherein said plurality of biomarkers comprises about 500 biomarkers.
67. The method of claim 65, wherein said plurality of biomarkers comprises about 1,100 biomarkers.
68. The method of claim 64, wherein said plurality of biomarkers comprises between about 20,000 biomarkers and about 25,000 biomarkers.
69. The method of claim 64, wherein said plurality of biomarkers comprises at least about 50,000 biomarkers.
70. The method of claim 64, wherein said plurality of biomarkers comprises at least about 65,000 biomarkers.
71. The method of claim 64, wherein said plurality of biomarkers comprises at least about 75,000 biomarkers.
72. The method of claim 64, wherein said plurality of biomarkers comprises at least about 100,000 biomarkers.
73. The method of claim 64, wherein said plurality of biomarkers comprises at least about 200,000 biomarkers.
74. The method of claim 64, wherein said plurality of biomarkers is associated with a disease or disorder.
75. The method of claim 74, wherein said disease or disorder comprises cancer.
76. The method of claim 64, wherein said plurality of biomarkers is associated with solid tumor cancer.Docket No.: 69246-750.60177. The method of claim 76, wherein said solid tumor cancer comprises carcinoma, sarcoma, or lymphoma.
78. The method of claim 64, wherein said plurality of biomarkers is associated with a hematologic malignancy.
79. The method of claim 1, further comprising amplifying said plurality of ssDNA fragments subsequent to (d).
80. The method of claim 79, wherein said amplifying said plurality of ssDNA fragments subsequent to (d) comprises polymerase chain reaction (PCR).
81. The method of claim 1, wherein said amplifying comprises isothermal amplification, rolling circle amplification (RCA), or strand displacement amplification.
82. The method of claim 1, further comprising sequencing said enriched nucleic acid sample.
83. The method of claim 82, wherein said sequencing comprises next generation sequencing (NGS).
84. The method of claim 82, wherein said sequencing comprises DNA sequencing, RNA sequencing, whole-genome sequencing, whole-exome sequencing, whole-transcriptome sequencing, targeted sequencing, bisulfite sequencing, or enzymatic sequencing.
85. The method of claim 84, wherein said sequencing comprises use of Illumina™ sequencing, Ultima™ sequencing, Ion Torrent™ sequencing, or a combination thereof.
86. The method of claim 1, further comprising computer processing said enriched nucleic acid sample.
87. The method of claim 86, further comprising analyzing data from said sequencing utilizing one or more computer processors.
88. The method of claim 87, wherein analyzing said sequencing data comprises determining genetic information, epigenetic information, or both, of said biomarkers.
89. The method of claim 88, wherein said genetic information comprises one or more of:tumor mutational burden (TMB), micro satellite instability (MSI), copy number variation (CNV), loss of heterozygosity (LoH), human leukocyte antigen (HLA) identification, gene fusions, gene splice variants, gene expression signatures, or viral detection features, or any combination thereof.Docket No.: 69246-750.60190. The method of claim 88, wherein said epigenetic information comprises methylation information.
91. A system comprising one or more processors and a memory operatively coupled to said one or more processors, wherein said one or more processors are individually or collectively programmed to perform said method of any one of claims 1 -90.
92. A composition, comprising a pool of capture probes, wherein said pool of capture probes comprises capture probes that are capable of hybridizing to at least a portion of an exome, capture probesthat are capable of hybridizing to at least a portion of abiomarker, or capture probes that are capable of hybridizing to a single nucleotide polymorphism (SNP), or a combination thereof.
93. The composition of claim 92, wherein said capture probes that are capable of hybridizing to said SNP comprise customized capture probes.
94. The composition of claim 92, wherein said pool of capture probes comprises at least two of said capture probes that are capable of hybridizing to at least a portion of an exome, said capture probes that are capable of hybridizing to at least a portion of a biomarker, or said capture probes that are capable of hybridizing to a single nucleotide polymorphism (SNP).
95. The composition of claim 92, wherein said pool of capture probes comprises said capture probes that are capable of hybridizing to at least a portion of an exome, said capture probes that are capable of hybridizing to at least a portion of a biomarker, and said capture probes that are capable of hybridizing to a single nucleotide polymorphism (SNP).
96. A kit comprising (a) said composition of any one of claims 92-95; and (b) instructions for use according to any one of said methods of claims 1-90.