Low coverage, genome-wide identification of minority cfDNA contributors.
Patent Information
- Application Number
- JP2026507134
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-08-04
- Filing Date
- 2024-08-02
- Publication Date
- 2026-09-17
Smart Images

Figure 2026531522000001_ABST
Abstract
Description
Technical Field
[0001] Cross Reference This application claims the priority benefit of U.S. Provisional Patent Application No. 63 / 517,741, filed on August 4, 2023, which is hereby incorporated by reference in its entirety into the present specification.
[0002] In some aspects, described herein are methods for analyzing cell-free DNA (cfDNA), wherein low-coverage, genome-wide analysis is used to detect and identify minor cfDNA contributors in a biological sample. In some aspects, described herein are kits for analyzing cell-free DNA (cfDNA), which provide instructions for detecting and identifying minor cfDNA contributors in a biological sample using low-coverage, genome-wide analysis.
Background Art
[0003] Minor components in biological samples can enable early diagnosis of transplant rejection or transplant failure events, as well as cancer remission or progression.
Summary of the Invention
[0004] In some embodiments, methods for analyzing cell-free DNA (cfDNA) are described herein, each method comprising the steps of: obtaining a biological sample derived from a subject, wherein the biological sample contains cfDNA; enriching the proportion of cfDNA in the biological sample; sequencing the cfDNA-enriched biological sample using low-coverage, genome-wide nucleic acid sequencing; identifying a plurality of minority components present in the sequenced cfDNA-enriched biological sample; assigning a designation representing a low-confidence estimate of the minor variant frequency to each identified minority component present in the sequenced cfDNA-enriched biological sample; and averaging a plurality of low-confidence estimates of the minor variant frequency across a plurality of sequenced genomic loci to present an estimate of the minority component frequency in the cfDNA-enriched biological sample. In some embodiments, the assigned designation is a binary classifier for distinguishing multiple sequenced genomic loci from sequenced genomic loci that were not identified as having minor variants in the sequenced cfDNA-enriched biological sample, with respect to individual identified minority components. In some embodiments, the step of identifying multiple minority components present in the sequenced cfDNA-enriched biological sample includes aligning raw sequence data generated by low-coverage whole-genome sequencing (lcWGS) with a reference sequence, marking duplicate reads of sequenced fragments, pre-processing the BAM files generated after lcWGS by base quality score recalibration (BQSR), performing local realignment of sequences from the pre-processed BAM files to generate BAM files for analysis, and performing variant calling on the BAM files for analysis to identify minority components.In some embodiments, a reference sample containing genomic DNA derived from the subject is analyzed to distinguish the somatic cell genotype present in the subject from minority components identified in the cfDNA-enriched biological sample. In some embodiments, the somatic cell genotype is identified by high-confidence genotyping. In some embodiments, high-confidence genotyping includes sequencing with at least 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 12x, 15x, or 20x genomic coverage. In some embodiments, the sites sequenced by low-coverage, genome-wide nucleic acid sequencing are independent of predefined genomic loci. In some embodiments, the estimate of minority component frequencies in the cfDNA-enriched biological sample is the quantitative detection of minority components present in the cfDNA of the biological sample. In some embodiments, the minority components detected in the cfDNA of a biological sample represent less than 25%, less than 20%, less than 15%, less than 12.5%, less than 10%, less than 9%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, less than 1.5%, less than 1.4%, less than 1.3%, less than 1.25%, less than 1.2%, and less than 1.15% of the total cfDNA present in the biological sample. This includes percentages of 1.0%, 1.1%, 1.05%, 1.0%, 0.95%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.45%, 0.4%, 0.35%, 0.3%, 0.25%, 0.2%, 0.15%, 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, or 0.05%. In some embodiments, the number of genomic loci assayed by low-coverage, genome-wide nucleic acid sequencing to detect potential variants is at least 5000, 10000, 20000, 35000, 50000, 75000, 100000, 150000, 200000, 250000, 300000, 400000, 500000, 600000, 750000, 875000, 1000000, 2000000, 3000000, 4000000, 5000000, 7500000, or 10000000 genomic loci.In some embodiments, the biological sample is from the subject's serum, plasma, blood, saliva, urine, mucus, tears, sweat, semen, breast milk, lymph, cerebrospinal fluid, or amniotic fluid. In some embodiments, the biological sample is from the subject's plasma. In some embodiments, the volume of the biological sample taken from the subject is approximately 200 μL, 150 μL, 125 μL, 100 μL, 80 μL, 75 μL, 70 μL, 60 μL, 55 μL, 50 μL, 45 μL, 40 μL, 35 μL, 30 μL, 25 μL, 20 μL, 17.5 μL, 15 μL, 12.5 μL, 10 μL, 9 μL, 8 μL, 7 μL, 6 μL, 5 μL, 4 μL, 3 μL, 2.5 μL, 2 μL, 1.5 μL, 1 μL, or 0.5 μL. In some embodiments, the amount of blood collected from the subject as a biological sample is approximately 200 μL, 150 μL, 125 μL, 100 μL, 80 μL, 75 μL, 70 μL, 60 μL, 55 μL, 50 μL, 45 μL, 40 μL, 35 μL, 30 μL, 25 μL, 20 μL, 17.5 μL, 15 μL, 12.5 μL, 10 μL, 9 μL, 8 μL, 7 μL, 6 μL, 5 μL, 4 μL, 3 μL, 2.5 μL, 2 μL, 1.5 μL, 1 μL, or 0.5 μL. In some embodiments, the biological sample is obtained from the subject by a method using capillary-based collection. In some embodiments, the individually identified minority components in cfDNA in e) include alternative heterozygous or alternative homozygous alleles compared to the genomic DNA from the subject. In some embodiments, the minority components individually identified in cfDNA in e) include alternative homozygous alleles compared to alleles in genomic DNA from the subject. In some embodiments, the detected variants include single nucleotide polymorphisms (SNPs), minor insertions or deletions (INDELs), variable number tandem repeats (VNTRs), simple sequence repeats (SSRs), or simple tandem repeats (STRs), or any combination thereof. In some embodiments, the detected variants include SNPs.In some embodiments, low-coverage, genome-wide nucleic acid sequencing includes sequencing coverage of less than approximately 1x, less than 0.9x, less than 0.8x, less than 0.7x, less than 0.6x, less than 0.5x, less than 0.4x, less than 0.35x, less than 0.3x, less than 0.25x, less than 0.2x, less than 0.15x, less than 0.125x, less than 0.1x, less than 0.09x, less than 0.08x, less than 0.075x, less than 0.065x, less than 0.05x, less than 0.035x, less than 0.025x, less than 0.02x, less than 0.015x, less than 0.01x, or less than 0.005x of the genome in question. In some embodiments, the alternative homozygous alleles for the identified minority components include loci less than approximately 1000, less than 750, less than 500, less than 400, less than 300, less than 200, less than 150, less than 125, less than 110, less than 100, less than 90, less than 80, less than 75, less than 70, less than 65, less than 60, less than 55, less than 50, less than 45, less than 40, less than 35, less than 30, less than 25, or less than approximately 20 loci. In some embodiments, the identified minority component allele homozygotes include approximately 1000, 750, 500, 400, 300, 200, 150, 125, 110, 100, 90, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, or 20 SNP loci. In some embodiments, the surrogate heterozygote or surrogate homozygote allele is derived from cfDNA from pre-malignant or malignant cells of the subject. In some embodiments, the surrogate heterozygote or surrogate homozygote allele is derived from cfDNA from one or more infectious agents present within the subject. In some embodiments, the surrogate heterozygote or surrogate homozygote allele is derived from the donor subject. In some embodiments, the donor subject is an embryo. In some embodiments, the donor subject is a fetus. In some embodiments, the donor subject provides a tissue graft or organ graft to the host subject. In some embodiments, estimation of minority component frequencies in cfDNA-enriched biological samples is used for transplant monitoring.In some embodiments, transplant monitoring includes distinguishing non-rejection (TX) from either or both acute rejection (AR) and acute dysfunction non-rejection (ADNR). In some embodiments, transplant monitoring includes comparing the level of donor-derived cell-free DNA (dd-cfDNA) detected in a cfDNA-enriched biological sample to a predetermined threshold. In some embodiments, a host is indicated to have potential for AR or ADNR by dd-cfDNA levels that are at least 0.5%, 0.6%, 0.75%, 0.8%, 0.85%, 0.9%, 0.95%, 1.0%, 1.1%, 1.25%, 1.5%, 1.75%, 2.0%, 2.5%, 3.0%, 4.0%, 5.0%, 7.5%, or 10% higher or equal to the predetermined threshold. In some embodiments, a host is indicated to have the potential for non-rejection by dd-cfDNA levels approximately 1.5%, 1.25%, 1.0%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.45%, 0.4%, 0.35%, 0.30%, 0.25%, 0.2%, 0.15%, or 0.1% lower than a predetermined threshold. In some embodiments, an increase over time in the presence of donor-derived cell-free DNA indicates dysfunction or rejection of the transplanted organ or tissue. In some embodiments, estimates of minority component frequencies in cfDNA-enriched biological samples are used for oncological detection or monitoring. In some embodiments, at least two biological samples taken from the subject at different time points are analyzed to screen for changes in minor cfDNA components over time. In some embodiments, a significant increase in somatic variants over time is detected. In some embodiments, the method further includes identifying novel somatic cell variants detected in the biological sample taken from the subject after initial biological sample collection from the subject. In some embodiments, somatic cell SNPs are distinguished from germline SNPs in the subject. In some embodiments, the initial biological sample is obtained from a tumor biopsy.In some embodiments, the method further includes a step of analyzing the DNA methylation pattern in a cfDNA-enriched biological sample. In some embodiments, the method further includes a step of analyzing the ratio of detected single-stranded cfDNA to detected double-stranded cfDNA in a cfDNA-enriched biological sample. In some embodiments, the method includes a step of monitoring the progression of an infection. In some embodiments, the method includes a step of comparing the level of cfDNA from one or more detected infectious pathogens in a cfDNA-enriched biological sample to a predetermined threshold. In some embodiments, the step of monitoring the progression of an infection includes detecting a significant increase in cfDNA from one or more infectious pathogens over time. In some embodiments, the subject is monitored for the progression of sepsis. In some embodiments, the onset or progression of pregnancy complications is monitored. The sequencing is performed. In some embodiments, the method further includes a step of analyzing fragment patterning in cfDNA-enriched biological samples. In some embodiments, the method further includes an inputting step of supplementing missing SNP genotypes in the sequenced cfDNA-enriched biological samples. In some embodiments, the method further includes a step of calculating and evaluating regional linkage disequilibrium ratios between detected variant alleles to improve the calculated estimate of minor component frequencies in the cfDNA-enriched biological samples. In some embodiments, the method further includes a step of adjusting the level of individual subjects, including long-term sampling of biological samples obtained from subjects to screen for changes in minor cfDNA components over time. In some embodiments, low-coverage, genome-wide nucleic acid sequencing is unbiased sequencing.
[0005] In some embodiments, a kit for analyzing cell-free DNA (cfDNA) is described herein, the kit comprising a sampler configured to collect a biological sample of interest, and a set of instructions for enriching the proportion of cfDNA in the biological sample, sequencing the cfDNA-enriched biological sample using low-coverage, genome-wide nucleic acid sequencing, identifying several minority components present in the sequenced cfDNA-enriched biological sample, assigning designations representing low-confidence estimates of minor variant frequencies to each identified minority component present in the sequenced cfDNA-enriched biological sample, and presenting an estimate of minority component frequencies in the cfDNA-enriched biological sample by averaging several low-confidence estimates of minor variant frequencies across several sequenced genomic loci. In some embodiments, the sampler is configured to collect one or more biological samples, including whole blood, from the subject.
[0006] Embedding by reference All publications, patents, and patent applications referenced herein are incorporated by reference to the same extent as each individual publication, patent, or patent application is specifically and individually indicated. To the extent that any publications and patents or patent applications incorporated by reference conflict with the disclosures contained herein, this Specification is intended to supersede and / or take precedence over any such conflicting material. [Brief explanation of the drawing]
[0007] Novel features of the present invention are described in detail in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by referring to the following detailed description and appended drawings which describe exemplary embodiments in which the principles of the present invention are utilized.
[0008] [Figure 1] Schematic diagrams of this method according to several embodiments are shown below. [Figure 2A]Figures 2A and 2B show the quantification and detection limits for several embodiments. Figure 2A shows a graph representing the quantification limits for samples of different concentrations of cell-free DNA (dd-cfDNA) derived from target donors. [Figure 2B] Figures 2A and 2B show the limits of quantification and detection for several embodiments. Figure 2B shows a graph representing the limit of quantification (LoQ), limit of detection (LoD), and blank limit (LoB) for samples with different concentrations of target dd-cfDNA. [Figure 3] Several embodiments of computer systems are shown. [Modes for carrying out the invention]
[0009] Quantitative identification of minority components present in cell-free nucleic acids (e.g., cell-free DNA, cfDNA) can be utilized in various non-invasive diagnostic testing strategies. This includes transplant monitoring, oncological detection and monitoring, prenatal testing, screening and detection of monogenic diseases, and diagnosis and monitoring of infectious diseases. Depending on the application, quantitative detection of minority components at low percentages, such as 1%, 0.5%, or 0.1% of the total cell-free nucleic acid material, is desirable. In transplant monitoring, the presence of cell-free nucleic acids from transplanted tissue exceeding a certain threshold (e.g., around 1%) may indicate active rejection or organ failure of the transplanted tissue. This indication can serve as evidence for medical intervention. In oncological monitoring, cell-free nucleic acids from tumors may indicate cancer recurrence after surgery and / or chemotherapy. Furthermore, changes in the composition of cell-free nucleic acids over time may indicate the onset or progression of certain diseases. As a non-limiting example, one or more non-germline polymorphisms detected or elevated in cfDNA samples from a subject may indicate cancer progression where the DNA repair mechanisms of tumor cells are impaired. Another non-limiting example is that one or more non-germline polymorphisms detected or elevated in cfDNA samples from a subject may indicate clonal proliferation of pathological cells in a disease (e.g., autoimmune diseases, diseases with pre-malignant changes, malignant proliferation in cancers such as leukemia, or malignant proliferation and / or metastasis in cancers such as lymphoma, carcinoma, and sarcoma). Therefore, close and frequent monitoring of cell-free nucleic acids is desirable for early intervention in individuals who may be affected, for example, by transplant rejection or cancer recurrence. Non-invasive or minimally invasive methods for detecting minority components are particularly advantageous because they reduce the burden on patients and healthcare workers in collecting usable samples. This also allows patients to provide usable samples to laboratories using home testing kits.Non-invasive or minimally invasive methods for detecting minority components of cfDNA that combine sensitivity and accuracy can provide clinically valuable and diagnostically useful information at an early stage of disease or disease progression, enabling i) earlier initiation of treatment, ii) earlier modification of ongoing treatment, iii) earlier gate-like decisions for switching treatments, or iv) any combination thereof, compared to less sensitive or less accurate methods for detecting minority components of cfDNA.
[0010] While it is desirable to detect small amounts of minority components in cell-free nucleic acids derived from biological samples with high accuracy, identifying minority components can be difficult when large amounts of nucleic acids from native or healthy cell populations (e.g., host cells, patient non-cancer cells, or host cell nucleic acid contamination in cell-free nucleic acid biological samples) are present. One approach to address this challenge is to perform deep sampling at a small number of genomic loci. Deep sampling can improve statistical certainty. Sequencing can be performed using techniques such as targeted deep sequencing, digital PCR, or allele-specific PCR. The focus of such approaches is to distinguish low-frequency minority variants from the majority of background cfDNA. These methods may use tens to thousands of loci to estimate the relative frequency of minority components.
[0011] In contrast, genome-wide approaches can be used to cover hundreds of thousands to millions of potentially informative loci. Estimates of minority component frequencies can be obtained by averaging low-confidence estimates of minor variant frequencies across a large number of loci. This distinguishes it from deep-depth approaches that use high-confidence estimates of minor variant frequencies across a smaller number of loci. Because genome-wide approaches do not require deep-depth profiling, they can be used to screen minority components from small starting materials (e.g., a few microliters / a few drops of blood versus a few milliliters / a tube of blood). In other words, genome-wide approaches can be used with very small samples that would not inherently contain enough cfDNA for deep-depth profiling at a sufficient number of loci, or even provide signals from pre-selected genomic loci. Furthermore, because targeted approaches may require additional processing steps, genome-wide approaches can incorporate additional nucleic acid diagnostic features (e.g., epigenetic markers, fragment size profiles, alignment bias, first-base bias, etc.), which can further refine the sensitivity of the approach. In some embodiments, the disclosure provides low-coverage, genome-wide sequencing methods for quantitatively evaluating the abundance of minor components.
[0012] In some embodiments, this disclosure provides a method for analyzing cell-free nucleic acids. In some embodiments, the method analyzes cell-free DNA (cfDNA). Figure 1 shows a schematic diagram of the method according to some embodiments. First, a biological sample containing cfDNA can be collected from a subject. If necessary, the biological sample can be enriched so that the proportion of cfDNA in the biological sample is enriched. The sample can be sequenced using a low-coverage and / or genome-wide nucleic acid sequencing method. In some embodiments of the method described herein, the low-coverage nucleic acid sequencing method or the low-coverage, genome-wide nucleic acid sequencing method includes a sequencing depth of sampling the genome under test with coverage between 0.1x and 10x. In some embodiments of the methods described herein, a low-coverage nucleic acid sequencing method or a low-coverage, genome-wide nucleic acid sequencing method includes a sequencing depth of sampling the genome under test with coverage of approximately 0.1x, 0.2x, 0.3x, 0.4x, 0.5x, 0.6x, 0.7x, 0.8x, 0.9x, 1x, 1.1x, 1.2x, 1.3x, 1.4x, 1.5x, 1.6x, 1.7x, 1.8x, 1.9x, 2x, 2.25x, 2.5x, 2.75x, 3x, 3.5x, 4x, 4.5x, 5x, 5.5x, 6x, 6.5x, 7x, 7.5x, 8x, 8.5x, 9x, 9.5x, or 10x. In some embodiments of the methods described herein, a low-coverage nucleic acid sequencing method or a low-coverage, genome-wide nucleic acid sequencing method includes a sequencing depth of sampling the genome under test with coverage of approximately 0.1 to 2 times.In some embodiments of the methods described herein, a low-coverage nucleic acid sequencing method or a low-coverage, genome-wide nucleic acid sequencing method includes a sequencing depth of sampling the genome under test with coverage less than approximately 1x but at least approximately 0.1x. The sequencing can generate a dataset capable of identifying minority components present in the sequenced cfDNA. For each identified minority component present in the sequenced cfDNA, a designation representing a low-confidence estimate of the minor variant frequency may be assigned. By averaging multiple low-confidence estimates of minor variant frequencies, for example, across multiple sequenced genomic loci, an estimate of minority component frequencies in multiple sequenced cfDNA-enriched biological samples can be presented. Estimating the frequency of minority components in cfDNA-enriched biological samples can be a quantitative detection of minority components present in the cfDNA of the biological sample. Individuals in which minority components in cfDNA are identified may contain alternative heterozygous alleles and / or alternative homozygous alleles compared to germline alleles in the genomic DNA of the subject. In some embodiments, alternative homozygous alleles from multiple loci present in the minority components of cfDNA from a biological sample are identified and distinguished from homozygous germline alleles with different genotypes from multiple loci in the genomic DNA of the subject. In some embodiments, heterozygous alleles from multiple loci present in the minority components of cfDNA from a biological sample are identified and distinguished from homozygous germline alleles with different genotypes from multiple loci in the genomic DNA of the subject.In some embodiments, alternative heterozygous alleles from multiple loci present in minority components of cfDNA from a biological sample are identified and distinguished from heterozygous germline alleles with different genotypes from multiple loci in the genomic DNA from the subject. Each identified minority component in the cfDNA may contain a homozygous allele compared to the heterozygous or homozygous alleles of the genomic DNA from the subject. Each identified minority component in the cfDNA may contain an alternative homozygous allele compared to the alleles of the genomic DNA from the subject. Each identified minority component in the cfDNA may contain an alternative allele(s) to the majority component of the cfDNA sample. In some embodiments of the methods described herein, the identified alternative alleles may be either heterozygous or homozygous in the affected cells, or they may be minority components compared to the alleles in the majority component of the cfDNA sample.
[0013] In some embodiments, a biological sample can be enriched such that the proportion of cfDNA in the biological sample is enriched. In some embodiments, the proportion of cfDNA in the biological sample is enriched relative to the amount of nucleic acid derived from cellular nucleic acid. In some embodiments, the proportion of cfDNA in the biological sample is enriched relative to the amount of nucleic acid derived from cellular DNA. In some embodiments, the proportion of cfDNA in the biological sample is enriched relative to the amount of nucleic acid derived from cellular genomic DNA, mitochondrial DNA, or microbiome DNA, or any combination thereof. In some embodiments, the proportion of cfDNA in the biological sample is enriched by purifying the cfDNA in the biological sample. In some embodiments, the cfDNA is purified to separate the DNA from other nucleic acids in the biological sample. In some embodiments, the cfDNA is purified to increase the proportion of cfDNA to cellular DNA in the biological sample. In some embodiments, cfDNA is enriched in the biological sample by performing DNA purification to produce a sample with higher DNA purity. In some embodiments, cfDNA is enriched in the biological sample by performing DNA purification to produce a sample having an enriched subset of cfDNA in a sample with higher DNA purity. In some embodiments, enriched cfDNA samples, purified cfDNA samples, or concentrated cfDNA samples are stored prior to sequencing of the cfDNA-enriched biological sample. In some embodiments, the enriched cfDNA sample contains at least 100 ng / mL, 75 ng / mL, 50 ng / mL, 40 ng / mL, 30 ng / mL, 20 ng / mL, 15 ng / mL, 10 ng / mL, 7.5 ng / mL, 5.0 ng / mL, 3.5 ng / mL, 2.5 ng / mL, 2.0 ng / mL, 1.5 ng / mL, 1.0 ng / mL, 0.75 ng / mL, 0.5 ng / mL, 0.25 ng / mL, or 0.1 ng / mL of DNA per volume of the biological sample, as measured by liquid quantification. In some embodiments, the biological sample is combined with a drug that selectively binds to nucleic acids or subsets of nucleic acids.In some embodiments, the agent that selectively binds to nucleic acids includes magnetic beads, silica, carbide, silica carbide, chitosan, polymers, or charged materials.
[0014] In some embodiments, the biological sample to be enriched is obtained from the subject's serum, plasma, blood, saliva, urine, mucus, tears, sweat, semen, breast milk, lymph, cerebrospinal fluid, or amniotic fluid, so that the proportion of cfDNA in the biological sample is enriched. In some embodiments, when blood is obtained from the subject, peripheral venous blood is collected in a blood collection tube containing EDTA. In some embodiments, when blood is obtained from the subject, peripheral venous blood is collected using a blood collection tube that does not contain EDTA or other additives. In some embodiments, when blood is obtained from the subject, peripheral venous blood is collected using a plasma separation device. In some embodiments, the plasma separation device does not use EDTA or other chelating agents to separate plasma from the blood sample. In some embodiments, when blood is obtained from the subject, capillary blood is collected using a sample collection and plasma separation card. The enriched cfDNA sample can be prepared from any of the sample sources listed above.
[0015] For peripheral blood samples collected using EDTA-containing blood collection tubes, centrifugation may be performed at 4°C and 1600×g for 10 minutes to separate the plasma from peripheral blood cells. The plasma portion is then transferred to a new sterile tube and centrifuged at 4°C and 16,000×g for 10 minutes to precipitate any remaining cells in the biological sample. The plasma portion is then transferred again to a new sterile tube. cfDNA can be extracted from the purified plasma using the QIAamp DSP DNA Blood Mini Kit (QIAGEN®) according to the manufacturer's protocol. Next, the cfDNA-enriched biological sample can be subjected to end repair, adapter ligation, and PCR amplification using the Ion Xpress® Plus Fragment Library Kit (Thermo Fisher Scientific) to construct a library for subsequent low-coverage whole-genome sequencing. In some embodiments, cfDNA enrichment is performed during library construction, after end repair and before adapter ligation. Magnetic beads with an average particle size of 1 μm can be used for size sorting of the end-repaired DNA fragments. In some embodiments, enriching portions of cfDNA in a biological sample involves sorting the cfDNA in the biological sample based on fragment size. In some embodiments, cfDNA sorting involves isolating cfDNA fragments in the biological sample that are below a certain nucleotide length to obtain a fragment-length enriched population of cfDNA fragments. In some embodiments, the population of fragment-length enriched cfDNA fragments is less than approximately 5000 base pairs (bp), less than 2000 bp, less than 1000 bp, less than 750 bp, less than 500 bp, less than 400 bp, less than 350 bp, less than 325 bp, less than 310 bp, less than 300 bp, less than 280 bp, less than 275 bp, less than 250 bp, less than 235 bp, less than 220 bp, less than 200 bp, less than 180 bp, less than 175 bp, less than 170 bp, less than 167 bp, less than 160 bp, less than 150 bp, less than 140 bp, less than 135 bp, less than 130 bp, less than 120 bp, less than 100 bp, less than 80 bp, less than 70 bp, less than 60 bp, or less than 50 bp in length.In some embodiments, the population of fragment-length enriched cfDNA fragments has lengths greater than approximately 350 bp, greater than 325 bp, greater than 310 bp, greater than 300 bp, greater than 280 bp, greater than 275 bp, greater than 250 bp, greater than 235 bp, greater than 220 bp, greater than 200 bp, greater than 180 bp, greater than 175 bp, greater than 170 bp, greater than 167 bp, greater than 160 bp, greater than 150 bp, greater than 140 bp, greater than 135 bp, greater than 130 bp, greater than 120 bp, greater than 100 bp, greater than 80 bp, greater than 70 bp, greater than 60 bp, or greater than 50 bp. In some embodiments, the population of fragment-length enriched cfDNA fragments includes DNA fragments whose lengths are in the range of approximately 50 to 350 bp. In some embodiments, sorting cfDNA from a biological sample based on size includes fragmenting the cfDNA sample by using sonication, centrifugation, or enzymatic digestion.
[0016] Each identified minority component in cfDNA may include any combination of homozygous alleles, alternative homozygous alleles, and alternative heterozygous alleles at multiple loci, compared to alleles at multiple loci of genomic DNA from the subject. Homozygous alleles, alternative homozygous alleles, and / or alternative heterozygous alleles may originate from cfDNA from pre-malignant or malignant cells of the subject. Homozygous alleles, alternative homozygous alleles, and / or alternative heterozygous alleles may originate from circulating tumor DNA (ctDNA). In some embodiments, ctDNA is a component of cfDNA from a biological sample that is released into the bloodstream and other bodily fluids by a malignant tumor. In some embodiments, ctDNA is at least 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.10%, 0.11%, 0.12%, 0.13%, 0.14%, 0.15%, 0.16%, 0.17%, 0.18%, 0.19%, 0.20%, 0.25%, 0% of the cfDNA in the biological sample derived from the subject. This includes 0.30%, 0.35%, 0.40%, 0.45%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1.0%, 1.1%, 1.2%, 1.3%, 1.4%, 1.5%, 1.75%, 2.0%, 2.5%, 2.75%, 3.0%, 3.5%, 4.0%, 4.5%, 5%, 6%, 7%, 8%, 9%, or 10%. Homozygous alleles, surrogate homozygous alleles, and / or surrogate heterozygous alleles may originate from cfDNA from one or more infectious pathogens present in the subject. Homozygous alleles, surrogate homozygous alleles, and / or surrogate heterozygous alleles may originate from the donor subject. The donor subject may be an embryo. The donor subject may be a fetus. The donor subject may be a child or adolescent. The donor may be an adult. The donor may provide a tissue graft or organ graft to the host. The donor may be a human leukocyte antigen (HLA) match with the host. The donor may not be related to the host by blood. The donor may be related to the host by blood.For related donor subjects, the coefficient of kinship, which indicates the percentage of shared DNA with the host, may be approximately 50%, 37.5%, 25%, 12.5%, 9.38%, 6.25%, 3.13%, 0.78%, or 0.20%. In some embodiments, the estimate of the percentage of shared DNA between a related donor subject and the host, based on the coefficient of kinship, is explained by the accuracy of statistical calculations for a method of estimating the frequency of minority components in cfDNA within a biological sample. The minority components detected in the cfDNA of biological samples represent less than 25%, less than 20%, less than 15%, less than 12.5%, less than 10%, less than 9%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, less than 1.5%, less than 1.4%, less than 1.3%, less than 1.25%, less than 1.2%, less than 1.15%, and less than 1.1% of the total cfDNA present in the biological sample. This may include percentages of 1.05%, 1.0%, 0.95%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.45%, 0.4%, 0.35%, 0.3%, 0.25%, 0.2%, 0.15%, 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, or 0.05%. The minority components detected in the cfDNA of a biological sample may comprise at least approximately 25%, 20%, 15%, 12.5%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1.5%, 1.4%, 1.3%, 1.25%, 1.2%, 1.15%, 1.1%, 1.05%, 1.0%, 0.95%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.45%, 0.4%, 0.35%, 0.3%, 0.25%, 0.2%, 0.15%, 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, or 0.05% of the total cfDNA present in the biological sample.The minority components detected in the cfDNA of a biological sample may comprise approximately 25%, 20%, 15%, 12.5%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1.5%, 1.4%, 1.3%, 1.25%, 1.2%, 1.15%, 1.1%, 1.05%, 1.0%, 0.95%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.45%, 0.4%, 0.35%, 0.3%, 0.25%, 0.2%, 0.15%, 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, or 0.05% of the total cfDNA present in the biological sample.The minority components detected in cfDNA of biological samples represent approximately 0.05%~0.10%, 0.05%~0.20%, 0.05%~0.50%, 0.05%~1.0%, 0.05%~2.0%, 0.05%~5%, 0.05%~10%, 0.05%~25%, 0.10%~0.20%, 0.10%~0.50%, 0.10%~1.0%, 0.10%~2.0%, 0.10%~5%, 0.10%~10%, 0.10%~25%, and 0.15%~0.30% of the total cfDNA present in the biological sample. 0.15%~0.50%, 0.15%~0.75%, 0.15%~1.0%, 0.15%~2.0%, 0.15%~5%, 0.15%~10%, 0.15%~25%, 0.20%~0.40%, 0.20%~0.60%, 0.20%~1.0%, 0.20%~2.0%, 0.20%~5%, 0.20%~10%, 0.20%~25%, 0.25%~0.50%, 0.25%~0.75%, 0.25%~1.0%, 0.25%~2.5%, 0.25%~5%, 0.25%~10%, 0.25%~2 5%, 0.35%~0.60%, 0.35%~0.75%, 0.35%~1.0%, 0.35%~2.0%, 0.35%~5%, 0.35%~10%, 0.35%~25%, 0.50%~1.0%, 0.50%~2.0%, 0.50%~3.0%, 0.50%~5%, 0.50%~7.5%, 0.50%~10%, 0.50%~25%, 0.75%~1.5%, 0.75%~3.0%, 0.75%~5%, 0.75%~10%, 0.75%~25%, 1.0%~2.0%, 1.0%~3.5% This may include ranges such as 1.0%~5%, 1.0%~7.5%, 1.0%~10%, 1.0%~15%, 1.0%~20%, 1.0%~25%, 1.5%~3.0%, 1.5%~5%, 1.5%~7.5%, 1.5%~10%, 1.5%~15%, 1.5%~20%, 1.5%~25%, 2.0%~5.0%, 2.0%~7.5%, 2.0%~10%, 2.0%~15%, 2.0%~20%, 2.0%~25%, 3.5%~7%, 3.5%~15%, 3.5%~20%, or approximately 3.5%~25%.In some embodiments, the minority components detected in the cfDNA of a biological sample are less than 5%, less than 20%, less than 15%, less than 12.5%, less than 10%, less than 9%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, less than 1.5%, less than 1.4%, less than 1.3%, less than 1.25%, less than 1.2%, and less than 1.15% of the total cfDNA2 present in the biological sample. This includes percentages of 1.0%, 1.1%, 1.05%, 1.0%, 0.95%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.45%, 0.4%, 0.35%, 0.3%, 0.25%, 0.2%, 0.15%, 0.1%, 0.09%, 0.08%, 0.07%, 0.06%, or 0.05%.
[0017] Depending on the circumstances, this method may include the absolute quantification of minority components. The homozygous or heterozygous alleles of the identified minority components are approximately less than 2000, less than 1500, less than 1250, less than 1000, less than 900, less than 800, less than 750, less than 600, less than 500, less than 400, less than 300, less than 200, less than 150, less than 125, less than 110, less than 100, less than 90, less than 85, less than 80, less than 75, less than 70, less than 65, May include less than 60, less than 55, less than 50, less than 45, less than 40, less than 35, less than 34, less than 33, less than 32, less than 31, less than 30, less than 29, less than 28, less than 27, less than 26, less than 25, less than 24, less than 23, less than 22, less than 21, less than 20, less than 19, less than 18, less than 17, less than 16, less than 15, less than 14, less than 13, less than 12, less than 11, or approximately less than 10 loci. The number of homozygous or heterozygous alleles of the identified minority components is approximately over 2000, over 1500, over 1250, over 1000, over 900, over 800, over 750, over 600, over 500, over 400, over 300, over 200, over 150, over 125, over 110, over 100, over 90, over 85, over 80, over 75, and over 70. This may include loci greater than 65, greater than 60, greater than 55, greater than 50, greater than 45, greater than 40, greater than 35, greater than 34, greater than 33, greater than 32, greater than 31, greater than 30, greater than 29, greater than 28, greater than 27, greater than 26, greater than 25, greater than 24, greater than 23, greater than 22, greater than 21, greater than 20, greater than 19, greater than 18, greater than 17, greater than 16, greater than 15, greater than 14, greater than 13, greater than 12, greater than 11, or about 10 or more. The identified minority component homozygous or heterozygous alleles may comprise approximately 2000, 1500, 1250, 1000, 900, 800, 750, 600, 500, 400, 300, 200, 150, 125, 110, 100, 90, 85, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 34, 33, 32, 31, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, or approximately 10 loci.The identified minority components may include fewer than 1000, fewer than 750, fewer than 500, fewer than 400, fewer than 300, fewer than 200, fewer than 150, fewer than 125, fewer than 110, fewer than 100, fewer than 90, fewer than 80, fewer than 75, fewer than 70, fewer than 65, fewer than 60, fewer than 55, fewer than 50, fewer than 45, fewer than 40, fewer than 35, fewer than 30, fewer than 25, or fewer than approximately 20 loci. The identified minority component alternative homozygous alleles may include fewer than approximately 1000, fewer than 750, fewer than 500, fewer than 400, fewer than 300, fewer than 200, fewer than 150, fewer than 125, fewer than 110, fewer than 100, fewer than 90, fewer than 80, fewer than 75, fewer than 70, fewer than 65, fewer than 60, fewer than 55, fewer than 50, fewer than 45, fewer than 40, fewer than 35, fewer than 30, fewer than 25, or fewer than approximately 20 SNP loci. Alternative homozygous alleles for identified minority components may include approximately 1000, 750, 500, 400, 300, 200, 150, 125, 110, 100, 90, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, or approximately 20 loci. Alternative homozygous alleles for identified minority components may include approximately 1000, 750, 500, 400, 300, 200, 150, 125, 110, 100, 90, 80, 75, 70, 65, 60, 55, 50, 45, 40, 35, 30, 25, or 20 loci. In some embodiments, the locus includes an SNP locus. In some embodiments, the locus includes a small insertion or deletion (INDEL). In some embodiments, the locus includes a variable number of tandem repeats (VNTR). In some embodiments, the locus includes a simple sequence repeat (SSR). In some embodiments, the locus includes a simple tandem repeat (STR). In some embodiments, the locus includes an SNP, INDEL, VNTR, SSR, or STR, or any combination thereof. In some embodiments, the locus is an SNP locus. In some embodiments, the locus consists of an SNP locus.
[0018] In some embodiments, the method analyzes cell-free RNA (cfRNA). In some embodiments, the method analyzes cell-free mRNA, cell-free tRNA, cell-free rRNA, or any combination thereof. In some embodiments, the cfRNA is released from cancer cells and non-cancerous cells. In some embodiments, the method analyzes the composition of alternative splicing of cell-free RNA. In some embodiments, the cfRNA is derived from one or more non-mutant tissues. In some embodiments, the one or more non-mutant tissues include the substrate of an organ, gland, or tissue. In some embodiments, the one or more non-mutant tissues include hematopoietic tissue. In some embodiments, the cfRNA is obtained from a biological sample from a host subject. In some embodiments, the biological sample is from the serum, plasma, blood, saliva, urine, mucus, tears, sweat, semen, breast milk, lymph, cerebrospinal fluid, or amniotic fluid of a host subject. In some embodiments, the cfRNA obtained from the biological sample is enriched. In some embodiments, the cfRNA obtained from the biological sample is purified. In some embodiments, cfRNA obtained from a biological sample is reverse transcribed to cDNA for further analysis using low-coverage, genome-wide sequencing, exome sequencing, or targeted sequencing. In some embodiments, alternative heterozygous and / or alternative homozygous alleles are detected and analyzed in a minority component of cfRNA compared to a majority component of cfRNA in the sample. In some embodiments, alternative heterozygous and / or alternative homozygous alleles are detected and analyzed in a minority component of cell-free mRNA compared to a majority component of cell-free mRNA in the sample. In some embodiments, alternative heterozygous and / or alternative homozygous alleles include SNPs. In some embodiments, alternative heterozygous and / or alternative homozygous alleles include INDELs.In some embodiments, alternative alleles containing alternatively spliced mRNA transcripts are detected and analyzed in a minority component of cell-free mRNA compared to a majority component of cell-free mRNA in the sample.
[0019] In cfDNA-enriched biological samples, fragment patterns can be analyzed. In some embodiments, fragment patterns can be analyzed after PCR amplification. In some embodiments, PCR-amplified fragments can be analyzed by electrophoresis (e.g., gel or capillary electrophoresis). In some embodiments, fragment patterns can be analyzed by evaluating restriction fragment length polymorphism (RFLP). In some embodiments, RFLP analysis uses an RFLP probe containing a labeled DNA sequence that can hybridize to fragments of one or more DNA samples after digestion with one or more restriction enzymes, after separation by electrophoresis. In some embodiments, fragment patterns can be analyzed by amplified fragment length polymorphism analysis. In some embodiments, fragment patterns can be analyzed by single-strand conformational polymorphism (SSCP) analysis. In some embodiments, fragment patterns can be analyzed by next-generation sequencing.
[0020] Missing SNP genotypes in cfDNA-enriched biological samples being sequenced can be imputed. Genotype imputation is the process of inferring unobserved genotypes in one or more samples. To improve the calculated estimate of minority component frequencies in cfDNA-enriched biological samples, local linkage disequilibrium ratios between detected variant alleles may be calculated and evaluated. In some embodiments, a haplotyping program is used to imputate missing genotypes during haplotype estimation. In some embodiments, genotype imputation tools such as PLINK, TUNA, WHAP, or BEAGLE are used to imputate genotypes by concentrating the analysis on a relatively small number of proximity markers when imputing missing genotypes. In some embodiments, genotype imputation tools such as IMPUTE, MACH, or fastPHASE / BIMBAM are used to imputate genotypes by considering all observed genotypes when imputing each missing genotype.
[0021] This method can be tailored to individual subjects based on long-term sampling of biological samples obtained from subjects to screen for changes in minor cfDNA components over time. In some embodiments, long-term sampling involves tailoring the sampling to individual parameters dependent on the subject being evaluated. In some embodiments, long-term sampling is compressed for subjects exhibiting serious disease or those whose condition is rapidly deteriorating. In some embodiments, compression of long-term sampling involves reducing the total number of samples taken for long-term analysis and / or increasing the sampling frequency. In some embodiments, long-term sampling involves taking at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, or 50 biological samples from subjects at different points in time. In some embodiments, biological samples are collected once every approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 25, 28, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, or 365 days. Low-coverage, genome-wide nucleic acid sequencing can be unbiased sequencing. Low-coverage, genome-wide nucleic acid sequencing can be hypothesis-free, for example, untargeted sequencing.
[0022] Genomic DNA derived from the subject can be analyzed and used as a reference sequence. Genomic DNA derived from the subject can be used to distinguish between germline genotypes present in the subject and somatic genotypes as minority components in cfDNA-enriched biological samples. Minority components may be somatic variants resulting from a diseased state of the cell (e.g., cancer or immune-related disease). Minority components may be DNA different from the subject's germline DNA resulting from circumstances specific to the subject (e.g., pregnancy or organ transplantation). Somatic genotypes can be identified through high-confidence genotyping. High-confidence genotyping may involve sequencing with at least 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 12x, 15x, or 20x genomic coverage. In some embodiments, somatic genotypes can be identified through targeted deep sequencing, next-generation sequencing, digital PCR, or allele-specific PCR.
[0023] Methods using next-generation sequencing generate individual base call (BCL) files. After next-generation sequencing is complete, the BCL files are converted into raw sequence data by generating FASTQ files, which are then evaluated for quality control. Next, the FASTQ files are trimmed using programs such as Skewer, Cutadapt, or Trimmomatic to remove low-quality bases and adapter sequences from the sequenced reads. The trimmed raw sequence data can be aligned to a reference sequence (e.g., to the UCSC Genome Browser, Ensembl Genome Browser, or the human genome reference sequence provided by IGV (Broad Institute)). Raw data is generated by low-coverage sequencing. Raw data is generated by whole-genome sequencing. Raw data is generated by transcriptome sequencing. Raw data is generated by low-coverage whole-genome sequencing (lcWGS). Low-coverage sequencing can be performed without amplification (e.g., PCR). Duplicate reads in the sequenced fragments can be marked. Low coverage sequencing is less than approximately 5 times, less than 4.5 times, less than 4 times, less than 3.5 times, less than 3 times, less than 2.5 times, less than 2.25 times, less than 2 times, less than 1.75 times, less than 1.6 times, less than 1.5 times, less than 1.4 times, less than 1.3 times, less than 1.2 times, less than 1.1 times, less than 1 time, less than 0.9 times, less than 0.8 times, less than 0.7 times, less than 0.6 times, less than 0.5 times, less than 0.4 times of the target genome. This may include sequencing coverage of less than 0.35x, less than 0.3x, less than 0.25x, less than 0.2x, less than 0.15x, less than 0.125x, less than 0.1x, less than 0.09x, less than 0.08x, less than 0.075x, less than 0.065x, less than 0.05x, less than 0.035x, less than 0.025x, less than 0.02x, less than 0.015x, less than 0.01x, or less than 0.005x.Low-coverage sequencing may comprise sequencing coverage of more than about 5-fold, more than 4.5-fold, more than 4-fold, more than 3.5-fold, more than 3-fold, more than 2.5-fold, more than 2.25-fold, more than 2-fold, more than 1.75-fold, more than 1.6-fold, more than 1.5-fold, more than 1.4-fold, more than 1.3-fold, more than 1.2-fold, more than 1.1-fold, more than 0.9-fold, more than 0.8-fold, more than 0.7-fold, more than 0.6-fold, more than 0.5-fold, more than 0.4-fold, more than 0.35-fold, more than 0.3-fold, more than 0.25-fold, more than 0.2-fold, more than 0.15-fold, more than 0.125-fold, more than 0.1-fold, more than 0.09-fold, more than 0.08-fold, more than 0.075-fold, more than 0.065-fold, more than 0.05-fold, more than 0.035-fold, more than 0.025-fold, more than 0.02-fold, more than 0.015-fold, more than 0.01-fold, or more than 0.005-fold of the target genome. In some embodiments, sequencing comprises real-time analysis.
[0024] The generated genome dataset (e.g., BAM file) can be preprocessed. Preprocessing may include base quality score recalibration (BQSR). Subsequently, the sequences can be locally aligned to generate an analyzable BAM file. The sites sequenced by low-coverage, genome-wide nucleic acid sequencing may be agnostic to predefined genomic loci. The sites sequenced by low-coverage, genome-wide nucleic acid sequencing may be targeted to predefined genomic loci. Variant calling can be performed on the analyzable BAM file to identify locations different from the reference sequence (e.g., SNPs or INDELs). Detected variants may include single nucleotide polymorphisms (SNPs), small insertions or deletions (INDELs), variable number tandem repeats (VNTRs), simple repeat sequences (SSRs), or simple tandem repeats (STRs), or any combination thereof. Variants may include SNPs. The number of genomic loci analyzed by low-coverage, genome-wide nucleic acid sequencing to detect potential variants may be at least 5000, 10000, 20000, 35000, 50000, 75000, 100000, 150000, 200000, 250000, 300000, 400000, 500000, 600000, 750000, 875000, 1000000, 2000000, 3000000, 4000000, 5000000, 7500000, or 10000000 genomic loci. The number of genomic loci analyzed by low-coverage, genome-wide nucleic acid sequencing to detect potential variants can be up to 5000, 10000, 20000, 35000, 50000, 75000, 100000, 150000, 200000, 250000, 300000, 400000, 500000, 600000, 750000, 875000, 1000000, 2000000, 3000000, 4000000, 5000000, 7500000, or 10000000 genomic loci.Low-coverage differences between variants identified as different from a reference sequence enable the detection, identification, and quantification of low-level minor components in a sample-derived cfDNA or cfRNA pool.
[0025] The designation may be a classifier for individually identified minor components. The classifier may be a binary classifier. The classifier may be a non-binary classifier. The designation may be a probability of an individually identified minor component. The designation may be a machine learning algorithm. The machine learning algorithm may be a neural network. The machine learning algorithm may be a random forest algorithm. Individually identified minor components may be preprocessed, for example normalized, before classification or probability determination. The classifier can distinguish sequenced genomic loci derived from minor components. The classifier can distinguish sequenced genomic loci derived from non-minor components. The classifier can distinguish a plurality of sequenced genomic loci from sequenced genomic loci that have not been identified as having a minor variant in a sequenced enriched cfDNA biological sample.
[0026] Biological samples can be obtained by non-invasive, minimally invasive, or invasive procedures. Biological samples can be obtained using kits disclosed herein. Biological samples can be obtained by blood collection, bloodletting, pharyngeal swab, buccal mucosa swab, bronchial lavage, urine collection, skin or epidermal abrasion, fecal collection, menstrual blood collection, semen collection, blood collection, venous puncture, biopsy, alveolar or lung lavage, needle aspiration, fingertip puncture, dried blood spot, capillary-based collection, or any combination thereof. Biological samples may be from the serum, plasma, blood, saliva, urine, mucus, tears, sweat, semen, breast milk, lymph, cerebrospinal fluid, or amniotic fluid of the subject. Biological samples may be from the plasma or blood of the subject. The volume of biological samples obtained from the subjects is approximately less than 1000 μL, less than 900 μL, less than 800 μL, less than 700 μL, less than 600 μL, less than 500 μL, less than 400 μL, less than 300 μL, less than 200 μL, less than 150 μL, less than 125 μL, less than 100 μL, less than 80 μL, less than 75 μL, less than 70 μL, less than 60 μL, less than 55 μL, and less than 50 μL. It may be less than 45 μL, less than 40 μL, less than 35 μL, less than 30 μL, less than 25 μL, less than 20 μL, less than 17.5 μL, less than 15 μL, less than 12.5 μL, less than 10 μL, less than 9 μL, less than 8 μL, less than 7 μL, 6 μL, less than 5 μL, less than 4 μL, less than 3 μL, less than 2.5 μL, less than 2 μL, less than 1.5 μL, 1 μL, or about less than 0.5 μL.In some embodiments, when blood is obtained from a subject, the amount of blood obtained from a biological sample derived from the subject for use in a method for analyzing cfDNA is approximately less than 10 mL, less than 8 mL, less than 7.1 mL, less than 6.1 mL, less than 4.6 mL, less than 2.1 mL, less than 1 mL, less than 0.75 mL, less than 500 μL, less than 400 μL, less than 350 μL, less than 300 μL, less than 275 μL, less than 250 μL, less than 225 μL, less than 200 μL, 175 μL, less than 150 μL, and 125 μL. Less than 100 μL, less than 80 μL, less than 75 μL, less than 70 μL, less than 65 μL, less than 60 μL, less than 55 μL, less than 50 μL, less than 45 μL, less than 40 μL, 35 μL, less than 30 μL, less than 25 μL, less than 20 μL, less than 17.5 μL, less than 15 μL, less than 12.5 μL, less than 10 μL, less than 9 μL, less than 8 μL, less than 7 μL, 6 μL, less than 5 μL, less than 4 μL, less than 3 μL, less than 2.5 μL, less than 2 μL, less than 1.5 μL, less than 1 μL, or about less than 0.5 μL. The volume of the biological sample obtained from the subject may be greater than approximately 200 μL, greater than 150 μL, greater than 125 μL, greater than 100 μL, greater than 80 μL, greater than 75 μL, greater than 70 μL, greater than 60 μL, greater than 55 μL, greater than 50 μL, greater than 45 μL, greater than 40 μL, greater than 35 μL, greater than 30 μL, greater than 25 μL, greater than 20 μL, greater than 17.5 μL, greater than 15 μL, greater than 12.5 μL, greater than 10 μL, greater than 9 μL, greater than 8 μL, greater than 7 μL, greater than 6 μL, greater than 5 μL, greater than 4 μL, greater than 3 μL, greater than 2.5 μL, greater than 2 μL, greater than 1.5 μL, greater than 1 μL, or greater than approximately 0.5 μL.
[0027] This method may further include steps to complement missing genotypes, potentially improving statistical capabilities by enabling the inference of additional genotypes while employing a low-coverage, genome-wide sequencing approach. The method may include steps utilizing fragment patterning to improve accuracy and / or detection limits. The method may also include steps utilizing long-term sampling to improve the accuracy of minority component concentrations and identify changes in the target disease state or condition over long-term sampling periods. Understanding the biological basis of cfDNA generation from tissues is crucial for appreciating the usefulness of cfDNA approaches in clinical settings. While SNP profiling can clearly identify the origin of minority cfDNA components, biases in fragment size and end-position provide additional information about their possible origins, allowing for more refined estimation of minority component frequencies in cfDNA. The method may also include steps utilizing epigenetic markers to improve accuracy and / or detection limits. Epigenetic markers can provide information about the tissue origin of cfDNA for oncological applications and prenatal screening. Additional information regarding epigenetic markers (e.g., methylation) may be used to refine the likely estimates of minor component frequencies in cfDNA. This method may include learned patient-level fingerprints to improve accuracy and / or detection limits. Long-term sampling data from patients for screening changes in minor components of cfDNA over time can be used to "learn" individual-level fingerprints regarding both SNP distribution and the aforementioned properties to further refine estimates of minor component frequencies in cfDNA.
[0028] Transplant monitoring The terms “transplant” or “graft” refer to the transfer of tissue, cells, or solid organs from a donor individual to a recipient individual. The donor and recipient may be of the same species or different species. For example, in some embodiments, a human recipient may receive solid organs from a non-human animal. “Allograft” further refers to the transfer of tissue, cells, or solid organs between different individuals of the same species. In contrast, when the donor and recipient are the same individual, the graft may be called an “autograft.” Estimation of minority component frequencies in cfDNA-enriched biological samples can be used for transplant monitoring.
[0029] A graft may include an organ graft or a tissue graft. A graft may include the adrenal gland, appendix, bladder, brain, ear, esophagus, eye, gallbladder, heart, kidney, large intestine, liver, lung, mouth, muscle, nose, pancreas, parathyroid gland, pineal gland, pituitary gland, skin, small intestine, spleen, stomach, thymus, thyroid gland, trachea, uterus, appendix (vermiform), cornea, skin, heart valves, arteries, or veins. In some cases, the organ may be a glandular organ. For example, an organ may be an organ of the digestive or endocrine system, and in some cases, an organ may be both an endocrine gland and a digestive organ. An organ may originate from the endoderm, ectoderm, primitive endoderm, or mesoderm. A graft may be an angiogenic composite allograft.
[0030] Organs, tissues, or cell grafts may be complete organs, fragments of complete organs, destroyed organs, or cells derived from any of the organs disclosed herein. Donor cells may be derived from any of the donor organs disclosed herein (e.g., pancreatic cells, hepatocytes, glioma cells, etc.). Transplanted tissues may contain stem cells (e.g., pluripotent stem cells, totipotent stem cells, neural stem cells, cardiac stem cells, induced pluripotent stem cells, embryonic stem cells, umbilical cord blood-derived cells, etc.). Depending on the circumstances, the transplanted organ, tissue, or cells may include gallbladder cells, cardiomyocytes, valvular cells, glomerular cells (e.g., parietal cells, podocytes), renal proximal tubular brush boundary cells, loop of Henle fine cells, ascending limb thickening cells, renal distal tubular cells, renal collecting duct cells, or renal stromal cells, intestinal epithelial cells, goblet cells, intestinal epithelial cells, caveolaeminate tuft cells, enteroendocrine cells, ganglion neurons, parenchymal cells, non-parenchymal cells, hepatocytes, sinusoidal endothelial cells, Kupffer cells, hepatic stellate cells, tendons, cartilage, bone, blood, lymph, muscle cells, muscle fibers, pancreatic β-cells, endothelial cells, or exocrine cells. Tissues may include, but are not limited to, connective tissue, epithelial tissue, muscle tissue, nerve tissue, adipose tissue, compact fibrous tissue, skeletal muscle, cardiac muscle, or smooth muscle. Muscle tissue may include muscle fibers or muscle cells. In some cases, the tissue is bone or tendon (both are called musculoskeletal grafts).
[0031] The graft may be a cellular allogeneic graft, for example, a graft containing allogeneic cells derived from a donor. The graft may include cells taken from a donor for administration to a recipient, cells taken from a donor and genetically engineered before administration to a recipient, cells taken from a donor and cultured before administration to a recipient, cells taken from a donor and processed before administration to a recipient, and any combination thereof.
[0032] The transplant recipient may receive one or more of a variety of allogeneic cells. Allogeneic cells may include, but are not limited to, blood cells, stem cells, cardiomyocytes, neurons, lymphocytes, NK cells, NKT cells, Treg cells, macrophages, dendritic cells, and islet cells. In some embodiments, the allogeneic cells are allogeneic blood cells. Allogeneic blood cells may include hematopoietic stem cells (i.e., HSCs), T cells, B cells, CAR T cells, NK cells, NKT cells, and TILs. In some embodiments, the allogeneic cells are allogeneic T cells. In some embodiments, the allogeneic cells are administered as bone marrow, umbilical cord blood, or purified allogeneic cells. In some embodiments, the allogeneic cells are bone marrow cells. In some embodiments, the allogeneic cells are umbilical cord blood cells. In some embodiments, the graft contains HSCs. In some embodiments, the HSCs are administered as bone marrow, umbilical cord blood, or purified HSCs. In some embodiments, the HSCs are derived from a donor. In some embodiments, the HSCs are administered as a hematopoietic cell transplant.
[0033] The graft may contain stem cells. Allogeneic stem cells may be embryonic stem cells, tissue-specific stem cells, mesenchymal stem cells, induced pluripotent stem cells, hematopoietic stem cells, skeletal muscle stem cells, cardiomyocyte stem cells, neural stem cells, epidermal stem cells, or intestinal stem cells. In some embodiments, the allogeneic stem cells are hematopoietic stem cells.
[0034] Transplantation may include cell grafts of the patient's own cells, for example, grafts containing cells derived from the recipient. These include, but are not limited to, cells taken from the recipient and genetically engineered before being re-administered to the same recipient, cells taken from the recipient and cultured before being re-administered to the same receptor, cells taken from the recipient and processed before being re-administered to the same recipient, and any combination thereof. For example, immune cells such as lymphocytes, NK cells, or macrophages can be genetically engineered and used to target and kill specific cancer cells. For example, T cells can be modified to produce a special structure called a chimeric antigen receptor (CAR) on their surface, which is designed to target specific cancer antigens. When these CAR T cells are administered to a recipient patient, the CAR receptor allows the CAR T cells to bind to the targeted cancer antigen, killing the cancer cells without damaging healthy tissue.
[0035] Examples of autologous cells include blood cells, stem cells, cardiomyocytes, nerve cells, lymphocytes, NK cells, NKT cells, Treg cells, macrophages, dendritic cells, and pancreatic islet cells, which are genetically modified and / or manufactured and / or cultured before being re-administered to the same receptor. In some embodiments, the autologous cells are autologous T cells, which are genetically modified and re-administered as CAR (chimeric antigen receptor) T cells.
[0036] Donor organs, tissues, or cells may be obtained from individuals that have a certain degree of similarity or compatibility with the recipient. For example, donor organs, tissues, or cells may be obtained from donors who are age-matched, race-matched, sex-matched, blood-type compatible, or HLA-matched to the recipient. In some circumstances, due to organ availability, donor organs, tissues, or cells may be taken from donors who have one or more mismatches with the transplant recipient in terms of age, race, sex, blood type, or HLA markers. Organs may originate from living or deceased donors.
[0037] "Recipient" may refer to an individual receiving transplantation, allograft, or autotransplantation. The recipient may be human. The recipient may be monitored for at least 6 hours, 12 hours, 1 day, 2 days, 3 days, 4 days, 5 days, 10 days, 15 days, 20 days, 25 days, 1 month, 2 months, 3 months, 4 months, 5 months, 7 months, 9 months, 11 months, 1 year, 2 years, 4 years, 5 years, 10 years, 15 years, or 20 years using the methods disclosed herein. The recipient may be monitored for up to 6 hours, 12 hours, 1 day, 2 days, 3 days, 4 days, 5 days, 10 days, 15 days, 20 days, 25 days, 1 month, 2 months, 3 months, 4 months, 5 months, 7 months, 9 months, 11 months, 1 year, 2 years, 4 years, 5 years, 10 years, 15 years, or 20 years using the methods described herein. If transplant rejection is detected, the recipient may be administered immunosuppressants. If transplant rejection is detected, the recipient may receive a higher dose of immunosuppressants. If transplant rejection is detected, the recipient may receive a different immunosuppressant. If transplant rejection is detected, a surveillance biopsy may be performed on the recipient. If transplant rejection is detected, a high-coverage method may be performed to confirm rejection.
[0038] Rejection can be "acute rejection." Acute rejection refers to a condition in which the transplanted tissue is rejected by the recipient's immune system, and the transplanted tissue is damaged or destroyed unless immunosuppression is achieved. T cells, B cells, other immune cells, and antibodies from the recipient can cause lysis of graft cells or produce cytokines that recruit other inflammatory cells, ultimately leading to necrosis of the allogeneic transplanted tissue. In some cases, acute rejection can be diagnosed by biopsy of the transplanted organ. Acute rejection can occur between 3 and 12 months after transplantation. It can also occur up to 5 years after transplantation, or throughout the patient's lifetime, if immunosuppression becomes insufficient for any reason.
[0039] Rejection can be cellular rejection or antibody-mediated rejection. In some cases, there may be no evidence of rejection. In some cases, there may be no evidence of cellular rejection. In some cases, there may be no histological evidence of rejection. In some cases, there may be no symptoms of rejection.
[0040] Oncological detection and monitoring Estimation of minority component frequencies in cfDNA-enriched biological samples can be used for oncological detection or monitoring. At least two biological samples taken from a subject at different time points may be analyzed to screen for changes in minor cfDNA components over time. An increase in minority components (e.g., minimal residual disease) may indicate cancer recurrence. Significant increases in somatic variants over time may be detected, and novel somatic variants may be detected and identified in biological samples taken from a subject after initial biological sample collection. Somatic SNPs can be distinguished from germline SNPs in the subject. Initial biological samples may be obtained from tumor biopsies. DNA methylation patterns in cfDNA-enriched biological samples may be analyzed. Optionally, oncological detection and monitoring may further include analysis of DNA methylation patterns in cfDNA-enriched biological samples to enhance the accuracy, sensitivity, and / or specificity of ctDNA detection in cfDNA-enriched biological samples. In some embodiments, DNA methylation patterns in cfDNA-enriched biological samples indicate positive tumorigenesis, cancer progression, cancer metastasis, or any combination thereof. In some embodiments, detection of hypomethylation of global DNA in cfDNA-enriched biological samples indicates positive malignant transformation, cancer progression, cancer metastasis, or any combination thereof. Optionally, oncological detection and monitoring may further include analysis of DNA strand composition patterns in cfDNA-enriched biological samples. In some embodiments, an increase in the proportion of single-stranded cfDNA compared to the amount of double-stranded cfDNA in cfDNA-enriched biological samples indicates positive malignant transformation, cancer progression, cancer metastasis, or any combination thereof.
[0041] By using a low-coverage, genome-wide sequencing approach instead of a targeted deep sequencing approach, performance can be improved through the evaluation of numerous additional features available through genome-wide data acquisition. In each of these cases, the use of non-genotype data collected during whole-genome sequencing of cfDNA (e.g., DNA methylation patterns in cfDNA, detection of genome-wide DNA hypomethylation in cfDNA, and analysis of DNA stranding patterns in cfDNA) further weights the evaluation of minor component frequencies in cfDNA.
[0042] Screening and detection of single-gene disorders Estimation of minority component frequencies in cfDNA-enriched biological samples can be used for screening monogenetic disorders. Estimation of minority component frequencies in cfDNA-enriched biological samples can be used for detecting monogenetic disorders. In some embodiments, the methods described herein can be used for prenatal screening of one or more monogenetic disorders. In some embodiments, the methods described herein can be used for prenatal detection of monogenetic disorders. In some embodiments, the non-invasiveness of the sampling techniques used in the methods described herein is advantageous for prenatal screening of one or more monogenetic disorders or for prenatal detection of monogenetic disorders. In some embodiments, the methods described herein can be used in non-invasive prenatal diagnosis (NIPD) as an alternative to invasive prenatal diagnostic methods such as amniocentesis or chorionic villus sampling. In some embodiments, screening and detection of monogenetic disorders may involve assaying and identifying one or more potential maternal, paternal, parental, or de novo mutations. In some embodiments, NIPD involves collecting and assaying biological samples from pregnant subjects during approximately 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, or 41 weeks of gestation. In some embodiments, long-term sampling is utilized. In some embodiments, the methods described herein are used for cystic fibrosis, Huntington's disease, type I myotonic dystrophy, Duchenne muscular dystrophy, facioscapulohumeral muscular dystrophy, Gaucher disease, Pompe disease, Friedreich's ataxia, congenital hearing loss, familial hypercholesterolemia, hemochromatosis, sickle cell disease, Tay-Sachs disease, adrenoleukodystrophy, hemophilia, and adrenal hyperplasia due to 21-hydroxylase deficiency (21-OHDCAH), Eicardi-Goutier syndrome encephalopathy, α-1-antitrypsin (A1AT) deficiency (AATD), arrhythmic right ventricular cardiomyopathy / dysplasia (ARVC, ARVD), autosomal dominant polycystic kidney disease (ADPKD), Brugada syndrome ventricular fibrillation, catecholamine polymorphic ventricular tachycardia (CPVT), Charcot-Marie-Tooth disease / hereditary motor and sensory neuropathy, congenital adrenal hyperplasia (CAH), congenital sucrase-isomaltase deficiency (CSID), congenital bilateral vas deferens aplasia, cystinuria-lysinuria syndrome / cystinuria, cytomegalovirus-associated congenital adrenal hypoplasia (AHC) (a subtype of congenital adrenal hypoplasia), amelodysplasia (DGI), dysbetalipoproteinemia / hyperlipoproteinemia type 3, Ehlers-Danlos syndrome, familial adenomatous Polyposis (FAP), Gardner syndrome (a subtype of familial adenomatous polyposis), familial spongiform malformation, familial hypocalciuric hypercalcemia type 1 (FHH), familial isolated dilated cardiomyopathy, familial long QT syndrome (LQTS) (including Romano-Ward syndrome), fragile X syndrome / Martin-Bell syndrome, glucose-6-phosphate dehydrogenase deficiency, GM2 gangliosidosis, hemolytic anemia due to erythrocyte pyruvate kinase deficiency, hemophilia A and B, hemorrhagic telangiectasia / Osler-Weder-Rendu disease, hereditary angioedema (HAE) / angioedema, hereditary breast and ovarian cancer syndrome, hereditary fructose intolerance / fructoseemia, hereditary xanthinuria / xanthine lithiasis, anhidrotic ectodermal dysplasia (HED).Dysplasia), iminoglycinuria, Li-Fraumeni syndrome (sarcoma, breast cancer, leukemia, adrenal tumor syndrome (SBLA)), long-chain 3-hydroxyacyl-CoA dehydrogenase deficiency (LCHAD), Lynch syndrome, Marfan syndrome, maternal phenylketonuria / phenylketonuria-induced fetal disease, medium-chain acyl-CoA dehydrogenase deficiency (MCADD), type III mucolipidosis (ML3) α / β, mucopolysaccharidosis type 4A (MPS4A) / Morcio disease type A, multiple endocrine neoplasia type 2, multiple Epiphyseal dysplasia (MED), neurofibromatosis type 1 (NF1) / von Recklinghausen's disease, oculocutaneous albinism (OCA), osteogenesis imperfecta / brittle osteodystrophy, Pendred syndrome (PDS) / hearing loss with goiter, phenylketonuria (PKU) / phenylalanine hydroxylase deficiency (PAH deficiency), proximal spinal muscular atrophy (SMA), retinitis pigmentosa (RP), recessive X-linked ichthyosis (XLI), retinoblastoma, Rett syndrome, Sotos syndrome / cerebral gigantism, Stargardt disease / macular macular degeneration (Fundus flavimaculatus), Stickler syndrome / hereditary progressive articular ocular disease, supravalvular aortic stenosis (SVAS), β-thalassemia, tibial muscular dystrophy / Upp myopathy, tuberous sclerosis / Bornville syndrome, von Hippel-Lindau disease. It may be used to screen for or detect diseases or groups of diseases, including von Willebrand disease, X-linked adrenoleukodystrophy (ALD), and X-linked retinoschisis (XLRS).
[0043] kit In some embodiments, the Disclosure provides a kit for use in the quantitative identification of minority components present in cell-free nucleic acids. The kit may comprise one or more biological sample collection devices. Samples may be collected over a long period of time. Samples may be collected approximately every 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 hours. Samples may be collected approximately every 1, 2, 3, 4, 5, 6, or 7 days. Samples may be collected approximately every 1, 2, 3, 4, or 5 weeks. Samples may be collected approximately every 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 months.
[0044] In some embodiments disclosed herein, kits are provided for obtaining genetic information from biological samples. The kits can collect and examine biological samples at a user-selected location to detect the presence and / or quantity of minority components in the sample. In some cases, the kit includes a sample purifier for removing at least one component (e.g., cells, cell fragments, proteins) from the biological sample of interest. The kit may include a nucleic acid sequencer for sequencing cell-free nucleic acids in the biological sample. The kit may include a nucleic acid sequence output for communicating the sequence information to the user.
[0045] The kit can integrate multiple functions, such as purification, amplification, detection, and combination thereof of a target analyte (e.g., cfDNA). In some cases, these multiple functions are performed within a single test assembly unit or a single device. In other cases, all functions are performed outside of a single unit or device. In some examples, at least one function is performed outside of a single unit or device. In some examples, only one of the functions is performed outside of a single unit or device. In some examples, the sample purifier, oligonucleotides, and detection reagents or components are housed within a single device. The kit may include a display, means for connecting to the display, or means for communicating with the display to convey information about a biological sample to one or more people.
[0046] The kit may include additional components such as a sample transport compartment, a sample storage compartment, a receptacle for samples and / or reagents, a temperature indicator, an electronic port, a communication connection, a communication device, a sample collection device, a housing unit, or any combination thereof. In some examples, the additional components are integrated with the device. In some examples, the additional components are not integrated with the device. In some examples, the additional components are housed within a single device together with a sample purifier, detection reagent, or other components. In some examples, the additional components are not housed within a single device.
[0047] The kit may include components for sample collection, cell-free nucleic acid extraction, and cell-free nucleic acid purification. In some examples, the apparatus, systems, and kits disclosed herein include components for sample collection, cell-free nucleic acid extraction, cell-free nucleic acid purification, and preparation of cell-free nucleic acid libraries. The kit may include components for sample collection, cell-free nucleic acid extraction, cell-free nucleic acid purification, and cell-free nucleic acid sequencing. The kit may include components for sample collection, cell-free nucleic acid extraction, cell-free nucleic acid purification, preparation of cell-free nucleic acid libraries, and cell-free nucleic acid sequencing. Components for sample collection include percutaneous puncture devices and filters for obtaining plasma from blood. Components for cell-free nucleic acid extraction and purification may include buffers, beads, and magnets. The buffers, beads, and magnets may be supplied in volumes suitable for accepting typical sample volumes from fingertip punctures (e.g., 50–150 μl of blood).
[0048] The kit may include a receptacle for receiving biological samples. The receptacle may be configured to hold biological samples in volumes between 1 μl and 1 ml. The receptacle may be configured to hold biological samples in volumes between 1 μl and 500 μl. The receptacle may be configured to hold biological samples in volumes between 1 μl and 200 μl. The receptacle may have a predetermined volume that is the same as the sample volume suitable for processing and analysis by other components of the device / system. In some examples, the device, system, and kit do not have a receptacle for receiving biological samples. In some examples, a sample purifier directly receives the biological sample. The sample purifier may have a predetermined volume suitable for processing and analysis by other components of the device / system. The user can store the analyzed sample or send it to another location (e.g., a lab, clinic) for further analysis or confirmation of results obtained in clinical practice. This kit may be used to separate plasma from blood. The plasma may be analyzed in clinical practice, and the cells in the blood may be sent to another location for analysis. The kit may be equipped with a transport compartment or storage compartment for these purposes. The transport compartment or storage compartment has the capacity to accommodate a biological sample, its components, or a portion thereof. The transport compartment or storage compartment has the capacity to accommodate a biological sample, its components, or a portion thereof during transport to a location far from the immediate user. The transport compartment or storage compartment has the capacity to accommodate cells isolated from a biological sample so that the cells can be sent to a location far from the immediate user for testing. Non-limiting examples of a location far from the immediate user include a laboratory or clinic, where the immediate user is at home. In some cases, the home may not have the equipment or additional devices necessary to perform further analysis of the biological sample. The transport compartment or storage compartment has the capacity to accommodate the products of a reaction or process resulting from the addition of the biological sample to the device. In some cases, the products of the reaction or process are biological sample components bound to the binding portion described herein.In some examples, the transport compartment or storage compartment may include absorbent pads, paper, glass containers, plastic containers, polymer matrices, liquid solutions, gels, preservatives, or combinations thereof.
[0049] Transport compartments or storage compartments may contain preservatives. Preservatives are also referred to herein as stabilizers or biological stabilizers. In some examples, devices, systems, or kits contain preservatives that reduce enzyme activity during storage and / or transport. In some examples, the preservative is a whole blood preservative. Non-limiting examples of whole blood preservatives or their components include glucose, adenine, citrate, trisodium citrate, dextrose, disodium phosphate, and monosodium phosphate. In some examples, the preservative contains EDTA. EDTA can reduce enzyme activity that would normally degrade nucleic acids. In some examples, the preservative contains formaldehyde. In some examples, the preservative is a known derivative of formaldehyde. Formaldehyde, or its derivatives, can crosslink proteins, thereby stabilizing cells and preventing cell lysis.
[0050] The kit may include a sample collector. In some examples, the sample collector is provided separately from the other components of the kit. In some examples, the sample collector is integrated with a receptacle as described herein. In some examples, the sample collector is a cup, tube, capillary, or well for applying a biological fluid. In some examples, the biological fluid to be collected is whole blood. In some examples, whole blood is collected from a subject by venous puncture. In some examples, whole blood is collected by fingertip puncture to provide a capillary blood sample. In some examples, the sample collector may be a cup for providing urine. In some examples, the sample collector may include a pipette for providing urine. In some examples, the sample collector may be a capillary tube integrated with a device disclosed herein for applying blood. In some examples, the sample collector may be a plasma separation device. In some examples, the sample collector may be a dried blood spot card. In some examples, the dried blood spot card may include a plasma separation card for collecting a spot of venous or capillary blood and separating a plasma sample from the dried blood spot. In some examples, the sample collector may be a tube, well, pad, or paper integrated with the apparatus disclosed herein for applying saliva. In some examples, the sample collector may be a pad or paper for applying sweat.
[0051] The kit may include a percutaneous puncture device. Non-limiting examples of percutaneous puncture devices are needles and lancets. In some examples, a sample collector may include a percutaneous puncture device. The kit may include a microneedle, a microneedle array, or a microneedle patch. The kit may include a hollow microneedle. In a non-limiting example, the percutaneous puncture device is integrated with a well or capillary so that when a subject punctures their finger, blood is released into the well or capillary, making the system or device available for analysis of its components. In some examples, the percutaneous puncture device is a push-button device having a needle or lancet in a concave surface. In some examples, the needle is a microneedle. In some examples, the percutaneous puncture device includes an array of microneedles. By pressing an actuator, button, or a predetermined position on the non-needle side of the concave surface, the needle punctures the subject's skin in a more controlled manner than a lancet. Furthermore, this push-button device may include a vacuum source or plunger to assist in aspirating blood from the puncture site.
[0052] The kit may include a sample processor, which modifies a biological sample to remove components of the sample or separates the sample into multiple fractions (e.g., a blood cell fraction and plasma or serum). The sample processor may also include a sample purifier, which is configured to remove unwanted substances or non-target components from the biological sample, thereby modifying the sample. Depending on the source of the biological sample, unwanted substances may include, but are not limited to, proteins (e.g., antibodies, hormones, enzymes, serum albumin, lipoproteins), free amino acids and other metabolites, microvesicles, nucleic acids, lipids, electrolytes, urea, urobilin, pharmaceuticals, mucus, bacteria, other microorganisms, and combinations thereof. In some examples, the sample purifier separates components of the biological sample disclosed herein. In some examples, the sample purifier disclosed herein removes components of the sample that would interfere with, disrupt, or otherwise be detrimental to post-processing steps such as detection. In some examples, the resulting modified sample is enriched with respect to the target analyte.
[0053] In some examples, sample purifiers include separation materials for removing unwanted substances other than patient cells from biological samples. Useful separation materials may include specific binding sites that bind to or associate with such substances. The binding may be covalent or non-covalent. In some examples, sample purifiers disclosed herein include binding sites that bind to nucleic acids, proteins, cell surface markers, or microvesicle surface markers. In some examples, the binding sites include antibodies, antigen-binding fragments, ligands, receptors, peptides, small molecules, or combinations thereof.
[0054] In some examples, the sample purifiers disclosed herein include filters. In some examples, the sample purifiers disclosed herein include membranes. The filters or membranes have the ability to separate or remove from the biological samples disclosed herein blood components other than cells, cell particles, cell fragments, cell-free nucleic acids, or combinations thereof.
[0055] In some cases, sample purifiers facilitate the separation of plasma or serum from the cellular components of blood samples. In some cases, sample purifiers facilitate the separation of plasma or serum from the cellular components of blood samples before initiating a sequencing reaction. Separation of plasma or serum can be achieved by several different methods, such as centrifugation, sedimentation, or filtration. In some cases, sample purifiers include a filter matrix for receiving whole blood, the filter matrix having pore sizes that prohibit cells from passing through while allowing plasma or serum to pass through unimpeded. In some cases, the filter matrix combines large pore sizes at the top with small pore sizes at the bottom, resulting in very gentle cell treatment that prevents cells from being degraded or lysed during the filtering process. This is advantageous because cell degradation or lysing would release nucleic acids from blood cells or maternal cells, potentially contaminating the target cell-free nucleic acids.
[0056] The kit may include vertical filtration driven by capillary force to separate components or fractions from a sample (e.g., plasma from blood). The sample purifier may include a lateral filter (e.g., the sample does not move in the direction of gravity, or the sample moves perpendicular to the direction of gravity). The sample purifier may include a vertical filter (e.g., the sample moves in the direction of gravity). The sample purifier may include a vertical filter and a lateral filter. The sample purifier may be configured to receive the sample or a portion of it with a vertical filter, followed by a lateral filter. The sample purifier may be configured to receive the sample or a portion of it with a lateral filter, followed by a vertical filter. In some examples, the vertical filter includes a filter matrix. In some examples, the filter matrix of the vertical filter has pores with a pore size that prevents cells from passing through, while plasma can pass through the filter matrix without being restricted. In some examples, filter matrices may have membranes particularly suited to this application, as they combine large pore sizes at the top with small pore sizes at the bottom, resulting in very gentle cell processing that prevents cells from being degraded during the filtering process.
[0057] In some examples, the sample purifier includes appropriate separation material, e.g., a filter or membrane, to remove undesirable substances from the biological sample without removing cell-free nucleic acids. In some embodiments where the biological sample is whole blood, standard collection techniques using whole blood centrifugation are used in the preparation of cfDNA-enriched biological samples. In some examples, the separation material separates substances in the biological sample based on size, for example, the separation material has a pore size that excludes cells but is permeable to cell-free nucleic acids. Thus, if the biological sample is blood, plasma or serum can move more quickly through the separation material in the sample purifier than blood cells, and plasma or serum containing cell-free nucleic acids can permeate the pores of the separation material. In some examples, the biological sample is blood, and the cells slowed and / or captured in the separation material are red blood cells, white blood cells, or platelets. In some examples, the cells are from tissues that have come into contact with the biological sample in the body, and these include, but are not limited to, bladder or urinary tract epithelial cells (in urine) or buccal mucosa cells (in saliva), etc. In some examples, the cells are bacteria or other microorganisms.
[0058] In some cases, sample purifiers have the ability to slow and / or capture cells without damaging them, thereby avoiding the release of cellular contents, including cellular nucleic acids and other proteins or cell fragments, which could interfere with the subsequent evaluation of cell-free nucleic acids. This can be achieved, for example, by a gradual decrease in pore size along a lateral flow strip or other suitable assay form pathway, to allow for a gradual slowing of cell migration and thereby minimize the force on the cells. In some cases, at least 95%, at least 98%, at least 99%, or up to 100% of the cells in a biological sample remain intact when captured in the separation material. In addition to, or independently of, size separation, the separation material can capture or separate undesirable substances based on cellular characteristics other than size; for example, the separation material may include binding sites that bind to cell surface markers. In some cases, the binding site is an antibody or an antigen-binding antibody fragment. In some cases, the binding site is a ligand or receptor-binding protein for a receptor on blood cells or microvesicles.
[0059] The kit may include a sample purifier, a filter, and / or a separation material for moving, aspirating, extruding, or drawing in a biological sample through a membrane. In some examples, the material is a wicking material. Examples of suitable separation materials used in a sample purifier to remove cells include, but are not limited to, polyvinylidene difluoride, polytetrafluoroethylene, acetylcellulose, nitrocellulose, polycarbonate, polyethylene terephthalate, polyethylene, polypropylene, glass fiber, borosilicate, vinyl chloride, and silver. A suitable separation material may be characterized as preventing the passage of cells. In some examples, the separation material is not limited as long as it has the property of preventing the passage of red blood cells. In some cases, the separation material is a hydrophobic filter, e.g., a glass fiber filter; a composite filter, e.g., Cytosep (e.g., Ahlstrom Filtration or Pall Specialty Materials, Port Washington, NY); or a hydrophilic filter, e.g., cellulose (e.g., Pall Specialty Materials). In some cases, whole blood can be fractionated into red blood cells, white blood cells, and serum components using a commercially available kit (e.g., Arrayit Blood Card Serum Isolation Kit, Cat. ABCS, Arrayit Corporation, Sunnyvale, CA) for further processing according to the methods of this disclosure.
[0060] In some examples, the sample purifier includes at least one filter or at least one membrane characterized by at least one pore size. In some examples, the sample purifier includes multiple filters and / or membranes, where the pore size of at least one filter or membrane is different from the pore size of a second filter or membrane. In some examples, at least one pore size of at least one filter / membrane is approximately 0.05 microns to approximately 10 microns. In some examples, the pore size is approximately 0.05 microns to approximately 8 microns. In some examples, the pore size is approximately 0.05 microns to approximately 6 microns. In some examples, the pore size is approximately 0.05 microns to approximately 4 microns. In some examples, the pore size is approximately 0.05 microns to approximately 2 microns. In some examples, the pore size is approximately 0.05 microns to approximately 1 micron. In some examples, at least one pore size of at least one filter / membrane is approximately 0.1 microns to approximately 10 microns. In some examples, the pore size is approximately 0.1 microns to approximately 8 microns. In some examples, the pore size is approximately 0.1 microns to approximately 6 microns. In some cases, the pore size is approximately 0.1 microns to 4 microns. In some cases, the pore size is approximately 0.1 microns to 2 microns. In some cases, the pore size is approximately 0.1 microns to 1 micron.
[0061] In some cases, sample purifiers are characterized as gentle sample purifiers. Gentle sample purifiers, such as those containing filter matrices, vertical filters, wicking materials, or membranes with pores that do not allow cell passage, are particularly useful for analyzing cell-free nucleic acids.
[0062] In some examples, the sample processor is configured to separate blood cells from whole blood. In some examples, the sample processor is configured to separate plasma from whole blood. In some examples, the sample processor is configured to separate serum from whole blood. In some examples, the sample processor is configured to separate plasma or serum from less than 1 milliliter of whole blood. In some examples, the sample processor is configured to separate plasma or serum from less than 1 milliliter of whole blood. In some examples, the sample processor is configured to separate plasma or serum from less than 500 μL of whole blood. In some examples, the sample processor is configured to separate plasma or serum from less than 400 μL of whole blood. In some examples, the sample processor is configured to separate plasma or serum from less than 300 μL of whole blood. In some examples, the sample processor is configured to separate plasma or serum from less than 200 μL of whole blood. In some examples, the sample processor is configured to separate plasma or serum from less than 150 μL of whole blood. In some cases, the sample processing device is configured to extract plasma or serum from less than 100 μL of whole blood.
[0063] The kit may include binding sites for generating modified samples from which unwanted or irrelevant cells, cell fragments, nucleic acids, or proteins have been removed. The kit may include binding sites for reducing unwanted or irrelevant cells, cell fragments, nucleic acids, or proteins in a biological sample. The kit may include binding sites for generating modified samples enriched with target cells, target cell fragments, target nucleic acids, or target proteins.
[0064] The kit may include binding sites that have the ability to bind to nucleic acids, proteins, peptides, cell surface markers, or microvesicle surface markers. The kit may include binding sites for capturing extracellular vesicles or extracellular microparticles in a biological sample. In some examples, the extracellular vesicles contain at least one of DNA and RNA. The kit may include reagents or components for analyzing the DNA or RNA contained within the extracellular vesicles. In some examples, the binding sites include antibodies, antigen-binding fragments, ligands, receptors, proteins, peptides, small molecules, or combinations thereof.
[0065] The kit may include binding sites that have the ability to interact with or capture extracellular vesicles released from cells. In some cases, extracellular vesicles are released from organs, glands, or tissues. In non-limiting examples, organs, glands, or tissues may be diseased, aging, infected, or proliferating. Non-limiting examples of organs, glands, and tissues include the brain, liver, heart, kidney, colon, pancreas, muscle, fat, thyroid, prostate, breast tissue, and bone marrow. The kit may be configured to have the ability to capture and discard extracellular vesicles or extracellular microparticles from a maternal sample to enrich the sample with cell-free nucleic acids.
[0066] In some examples, the binding portion is bound to a solid support, and after the binding portion comes into contact with the biological sample, the solid support may be separated from the rest of the biological sample, or the biological sample may be separated from the solid support. Non-limiting examples of solid supports include beads, nanoparticles, magnetic particles, chips, microchips, fiber strips, polymer strips, membranes, matrices, columns, plates, or combinations thereof.
[0067] The kit may include cell lysis reagents. Non-limiting examples of cell lysis reagents include detergents such as NP-40 and sodium dodecyl sulfate, and salt solutions containing ammonium, chloride, or potassium. The kit may include cell lysis components. These components may be structural or mechanical and have the ability to lyse cells. Non-limiting examples include cell lysis components that can shear cells and release intracellular components such as nucleic acids. In some embodiments, the kit does not include cell lysis reagents.
[0068] Computer system This disclosure provides a computer system programmed to carry out the methods of this disclosure. For example, Figure 3 shows a computer system 301 programmed or otherwise configured to, for example, quantify the amount of minority components in a sample based on genomic data, identify multiple minority components present in a sequenced cfDNA-enriched biological sample, assign designations representing low-confidence estimates of minor variant frequencies to each identified minority component present in the sequenced cfDNA-enriched biological sample, and / or present an estimate of minority component frequencies in the cfDNA-enriched biological sample by averaging multiple low-confidence estimates of minor variant frequencies across multiple sequenced genomic loci.
[0069] The computer system 301 may modulate various aspects of the analysis, computation, and generation of this disclosure, such as quantifying the amount of minority components in a sample based on genomic data, identifying multiple minority components present in a sequenced cfDNA-enriched sample, assigning designations representing low-confidence estimates of minor variant frequencies to each identified minority component present in a sequenced cfDNA-enriched biological sample, and / or averaging multiple low-confidence estimates of minor variant frequencies across multiple sequenced genomic loci to generate an estimate of minority component frequencies in the cfDNA-enriched biological sample. The computer system 301 may be an electronic device or computer system of a user located remotely from the electronic device. The electronic device may be a mobile electronic device.
[0070] The computer system 301 includes a central processing unit (CPU, "processor" and "computer processor" as herein) 305, which can be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 301 also includes memory or memory locations 310 (e.g., random-access memory, read-only memory, flash memory), electronic storage devices 315 (e.g., hard disks), a communication interface 320 for communicating with one or more other systems (e.g., a network adapter), and peripheral devices 325 such as caches, other memories, data storage devices, and / or electronic display adapters. The memory 310, storage devices 315, interface 320, and peripheral devices 325 communicate with the CPU 305 via a communication bus (solid line), such as a motherboard. The storage unit 315 may be a data storage unit (or data repository) for storing data. The computer system 301 is operably connected to a computer network ("network") 330 with the assistance of the communication interface 320. Network 330 may be the Internet, the Internet and / or an extranet, or an intranet and / or extranet communicating with the Internet.
[0071] Network 330 may, in some cases, be a telecommunications and / or data network. Network 330 may include one or more computer servers, which may enable distributed computing such as cloud computing. For example, cloud computing on Network 330 ("Cloud") may enable one or more computer servers to perform various aspects of the analysis, computation, and generation of this disclosure, such as, for example, quantifying the amount of minority components in a sample based on genomic data, identifying multiple minority components present in a sequenced cfDNA-enriched biological sample, assigning designations representing low-confidence estimates of minor variant frequencies to each identified minority component present in a sequenced cfDNA-enriched biological sample, and / or presenting an estimate of minority component frequencies in the cfDNA-enriched biological sample by averaging multiple low-confidence estimates of minor variant frequencies across multiple sequenced genomic loci. Such cloud computing may be provided by cloud computing platforms such as, for example, Amazon Web Services (AWS), Microsoft Azure, Google Cloud Platform, and IBM Cloud. Network 330 can, in some cases, implement a peer-to-peer network with the help of computer system 301, which may allow devices connected to computer system 301 to act as clients or servers.
[0072] The CPU 305 may include one or more computer processors and / or one or more graphics processing units (GPUs). The CPU 305 can execute a set of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in a memory location such as memory 310. These instructions can be directed to the CPU 305, which can then be programmed or otherwise configured to perform the methods of this disclosure. Examples of operations performed by the CPU 305 include fetching, decoding, executing, and writing back.
[0073] The CPU 305 may be part of a circuit, such as an integrated circuit. One or more other components of the system 301 may be included in the circuit. In some cases, the circuit is an application-specific integrated circuit (ASIC).
[0074] The storage device 315 can store files such as drivers, libraries, and saved programs. The storage device 315 can also store user data, such as user preferences and user programs. The computer system 301 may include one or more additional data storage devices located outside of the computer system 301, such as being located on a remote server that is in communication with the computer system 301 via an intranet or the internet.
[0075] Computer system 301 can communicate with one or more remote computer systems via network 330. For example, computer 501 can communicate with a user's remote computer. Examples of remote computer systems include personal computers (e.g., portable PCs), slate or tablet PCs (e.g., Apple® iPad, Samsung® Galaxy Tab, etc.), telephones, smartphones (e.g., Apple® iPhone, Android-enabled devices, Blackberry®), or personal digital assistants. Users can access computer system 301 via network 330.
[0076] The methods described herein may be implemented by machine-executable code (e.g., a computer processor) stored on an electronic memory location of the computer system 301, such as memory 310 or an electronic memory unit 315. The machine-executable code or machine-readable code may be provided in the form of software. In use, the code may be executed by the processor 305. In some cases, the code may be retrieved from the storage device 315 and stored in memory 310 for immediate access by the processor 305. In some situations, the electronic storage device 315 may be excluded, and the machine-executable instructions are stored in memory 310.
[0077] The code may be pre-compiled and configured for use with a machine having a processor adapted to run the code, or it may be compiled at runtime. The code may be provided in a programming language in which it can choose to run either pre-compiled or as compiled.
[0078] Embodiments of systems and methods provided herein, such as computer system 301, can be embodied in programming. Various embodiments of this technology can typically be conceived as “products” or “manufacturing supplies” in the form of machine (or processor) executable code and / or related data that are executed or embodied on some kind of machine-readable medium. Machine-executable code can be stored in electronic storage devices such as memory (e.g., read-only memory, random-access memory, flash memory) or hard disks. “Storage” type media can include any or all of the tangible memory of a computer or processor, or its related modules, such as various semiconductor memories, tape drives, disk drives, etc., which can provide a non-temporary recording medium for programming software at any time. All or part of the software may sometimes be communicated over the Internet or various other telecommunication networks. Such communication may enable the loading of software from one computer or processor to another, for example, from a management server or host computer to an application server computer platform. Therefore, other types of media that can hold software elements include light, electricity, and electromagnetic waves, such as those used in wired and terrestrial optical communication network between local devices, and through various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., may also be considered media that hold software. Unless limited to non-temporary tangible “storage” media as used herein, terms such as computer or machine “readable media” refer to any medium involved in providing instructions to a processor to execute.
[0079] Therefore, machine-readable media such as computer executable code may take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include optical or magnetic disks, such as any storage device in any computer(s), which may be used to implement, for example, a database as shown in the drawings. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include copper wires and optical fibers, including coaxial cables and wires including buses in computer systems. Carrier transmission media may take the form of electrical or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Therefore, common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, other magnetic media, CD-ROMs, DVDs, or DVD-ROMs, other optical media, punched card paper tapes, other physical storage media with patterns of holes, RAM, ROMs, PROMs, and EPROMs, FLASH-EPROMs, other memory chips or cartridges, carriers that carry data or instructions, cables or links that carry such carriers, or other media from which a computer reads programming code and / or data. Many forms of computer-readable media may be involved in processor-transmitted one or more sequences of one or more instructions for execution.
[0080] The computer system 301 includes, or may be in communication with, an electronic display 335 that includes a user interface (UI) 340 for quantifying the amount of minority components in a sample based on genomic data, identifying multiple minority components present in a sequenced cfDNA-enriched sample, assigning designations representing low-confidence estimates of minor variant frequencies to each identified minority component present in a sequenced cfDNA-enriched biological sample, and / or averaging multiple low-confidence estimates of minor variant frequencies across multiple sequenced genomic loci to generate an estimate of minority component frequencies in the cfDNA-enriched biological sample. Examples of the UI include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0081] The methods and systems of this disclosure may be implemented by one or more algorithms. The algorithms may be implemented by software, which are executed by the central processing unit 305. The algorithms may, for example, quantify the amount of minority components in a sample based on genomic data. [Examples]
[0082] The following embodiments are provided to illustrate some of the embodiments of the disclosure and to further illustrate them, but are not intended to limit the scope of the disclosure. They are illustrative only, and it will be understood that other procedures, methods, or techniques well known to those skilled in the art may be used instead.
[0083] Example 1: Identification of low-coverage, genome-wide minority cfDNA contributors Theoretical background:
[0084] The dbSNP database contains approximately 11.4 million SNPs with minor allele frequencies exceeding 20% in the human population. Based on predictions from the Hardy-Weinberg equilibrium, when comparing two random individuals, at least 500,000 SNPs should be surrogate homozygotes. Using low-coverage whole-genome sequencing with a genome coverage of 0.1x (one base sequenced for approximately 10 bases of the genome), it is expected that in a typical sample, approximately 50,000 of these SNP loci of surrogate homozygous genotypes between two unrelated individuals will have some level of sequencing coverage. If minor components of cfDNA are present at a rate of 0.1%, it is expected that approximately 1 in 1,000 molecules will originate from a minor contributor, and this is close to 50 surrogate homozygous SNP loci. It is expected that 50 events will be sufficient to adequately estimate the presence of minor cfDNA components against background noise.
[0085] To increase the number of sampled sites and, consequently, improve the sensitivity of the assay, heterozygous and homozygous sites may be incorporated similarly. This method detects SNP alleles that are not explicitly linked to SNP alleles known to be present (or absent) in minority cfDNA sets, and which are different from sets of majority cfDNAs. Thus, highly reliable genotyping of references (majority components) can significantly improve reliability. However, the same cannot be expected for references of minority components.
[0086] Experimental estimation of small-scale cfDNA quantification:
[0087] Plasma mixtures were prepared from two individuals, one male and one female, ranging from 100% female to 0.05% female, and sequenced with approximately 0.2x genomic coverage. Furthermore, unmixed reference samples were prepared for each sample and sequenced with approximately 10x genomic coverage to identify potential SNP sites. In the mixtures, the relative abundance of the Y chromosome was used to calibrate the frequencies of experimentally observed minor components in this set.
[0088] Using the structure outlined below, we evaluated the ability to identify and screen SNP sites in a mixture and quantitatively determine the level of minor cfDNA down to a 0.1% abundance. These data support the ability to identify minor cfDNA levels of at least 0.1%, and in some cases as low as 0.05%. 1. Identify homozygous SNP loci in the recipient (major component) target. 2. On an optional basis, identify heterozygous and alternative homozygous SNP loci in the donor (minor component) target. 3. In the mixture, determine the frequency of non-recipient genotypes at the recipient SNP loci that were screened in advance. 4. Compare the obtained frequency to the level of known mixtures.
[0089] Example 2: Monitoring of transplant rejection This example provides a strategy for monitoring changes over time in the frequency of minority components of cfDNA derived from transplanted tissue. The screening test can be performed from a capillary-based collection of plasma in a volume of microliters.
[0090] An increase in the presence of graft-derived cfDNA over time may indicate organ failure and / or rejection. This method can be used to screen for different SNP loci between donors and recipients to determine an estimate of donor cfDNA levels. If this value increases over time, the patient may be notified to follow up with a physician for further care. Using the method herein, the individual background of a subject can be adapted to a test format to create individualized profiling. Long-term sampling over a period of time allows for monitoring of transplant rejection and transplant organ failure. In some embodiments, the frequency of long-term sampling is faster (e.g., several samples are taken from the subject over a period of several days or weeks) as opposed to slower sampling frequencies (e.g., monthly, bimonthly, quarterly, biennially, or annually). In some embodiments, faster-frequency long-term sampling is used when the subject presents with acute illness. In some embodiments, faster-frequency long-term sampling is used when the subject is at high risk of transplant rejection. In some embodiments, faster-frequency long-term sampling is used when the subject is at high risk of transplant organ failure. In some embodiments, longer-period sampling is used when the subject is in a stable state. In some embodiments, longer-period sampling is used to monitor changes within the stable state in the subject. Estimation of minor component frequencies of dd-cfDNA may be further evaluated by including analysis that involves data splitting to determine any changes in the detection percentage of dd-cfDNA relative to host-derived cfDNA across multiple samples. During the analysis, splitting the sequencing reads of the biological sample into a set of contaminant sequencing reads and a set of target sequencing reads may be useful for detection accuracy, sensitivity, and / or specificity in transplant rejection monitoring or transplant failure monitoring.By separating the set of contaminant sequencing reads from the set of target sequencing reads, the object can be classified and / or characterized with a significant increase in detection accuracy, sensitivity, and / or specificity in detecting and / or determining the presence of the set of target sequencing reads, and thus the classification and / or characterization of objects having at least the set of target sequencing reads can be improved.
[0091] Example 3: Oncology Screening This example provides a strategy for monitoring changes in the frequency of minor components of cfDNA derived from tumor samples over time. The screening test can be performed from a capillary-based collection of microliter volumes of plasma.
[0092] In oncology screening, the minor components of the identified cfDNA are likely to be tumor-derived. Therefore, this use case differs from some others that screen a mixture of two sets of germline SNP variants in that it aims to identify a significant increase over time of somatic variants, such as tens to hundreds of variants per megabase of tumor tissue, thus offering a large number of potential targets. Nevertheless, the above assumption must still hold true in such use cases, despite the use of a different reference approach (somatic variant set versus germline SNPs). In some versions, this use case may be informed by past patient history (e.g., tumor biopsy data), although this is not a mandatory element of the screening.
[0093] This method is used to detect increased somatic mutations in minority components of cfDNA, including ctDNA. Increased somatic mutations can be due to inefficient DNA damage repair in pre-malignant and malignant cells and represent a driving force in cancer progression. Identifying an increase in tumor cells that cannot efficiently repair DNA damage leading to increased somatic mutations in ctDNA allows for i) earlier initiation of treatment in the subject, ii) earlier modification of ongoing treatment in the subject, iii) earlier gate-like decisions to switch treatments in the subject, or iv) any combination thereof.
[0094] This method is used to detect an increase in clonal somatic mutations in minority components of cfDNA, including ctDNA. A higher percentage of clonal somatic mutations in a later sample compared to a previously taken sample from a subject may indicate cancer progression or metastasis in the subject, as the higher percentage of clonal somatic mutations may originate from a higher proportion of ctDNA in the total cfDNA sample. Identifying an increase in clonal somatic mutations in a later sample allows for i) earlier initiation of treatment in the subject, ii) earlier modification of ongoing treatment in the subject, iii) earlier gate-like decision-making regarding a change in treatment in the subject, or iv) any combination thereof.
[0095] While preferred embodiments of the Disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided only as examples. Many variations, alterations, and substitutions can be conceived by those skilled in the art without departing from the Disclosure. It should be understood that various alternatives to the embodiments of the Disclosure described herein may be used in the practice of the Disclosure. The following claims define the scope of the Disclosure, and the methods and configurations within the scope of these claims and their equivalents are intended to be encompassed thereby.
Claims
1. A method for analyzing cell-free DNA (cfDNA), a) A step of obtaining a biological sample derived from a target, wherein the biological sample contains cfDNA, b) A step of enriching the proportion of cf DNA in the biological sample, c) A step of sequencing a cfDNA-enriched biological sample using low-coverage, genome-wide nucleic acid sequencing, d) A step of identifying a plurality of minority components present in the sequenced cfDNA-enriched biological sample, e) A step of assigning a designation representing a low-confidence estimate of the minor variant frequency to each identified minority component present in the sequenced cfDNA-enriched biological sample, The process involves averaging multiple low-confidence estimates of minor variant frequencies across multiple sequenced genomic loci to present an estimate of the frequency of minor components in the cfDNA-enriched biological sample. A method that includes this.
2. The method according to claim 1, wherein the designation in e) is a binary classifier for distinguishing a plurality of sequenced genomic loci from sequenced genomic loci that were not identified as having minor variants in a sequenced cfDNA-enriched biological sample, with respect to each of the identified minority components.
3. The identification in d) i. Aligning raw sequence data generated by low-coverage whole-genome sequencing (lcWGS) in c) with a reference sequence. ii. Marking duplicate reads of sequenced fragments. iii. Preprocess the BAM files generated after lcWGS by performing Base Quality Score Rescalculation (BQSR). iv. Performing local realignment of sequences from preprocessed BAM files to generate BAM files for analysis, and v. Perform a variant call on the BAM file for analysis to identify minority components. The method according to claim 1 or 2, including the method according to claim 1 or 2.
4. The method according to any one of claims 1 to 3, wherein a reference sample containing genomic DNA derived from the subject is analyzed in order to distinguish the somatic cell genotype present in the subject from the minority components identified in the cfDNA-enriched biological sample.
5. The method according to claim 4, wherein the somatic cell genotype is identified by high-reliability genotyping.
6. The method according to claim 5, wherein the high-confidence genotyping comprises sequencing of at least 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 12x, 15x, or 20x genomic coverage.
7. The method according to any one of claims 1 to 6, wherein the sites sequenced by low-coverage, genome-wide nucleic acid sequencing are independent of predefined genomic loci.
8. The method according to any one of claims 1 to 7, wherein the estimation of the frequency of minority components in the cfDNA-enriched biological sample is the quantitative detection of minority components present in the cfDNA of the biological sample.
9. The detected minority components present in the cfDNA of the biological sample represent less than 25%, less than 20%, less than 15%, less than 12.5%, less than 10%, less than 9%, less than 8%, less than 7%, less than 6%, less than 5%, less than 4%, less than 3%, less than 2%, less than 1.5%, less than 1.4%, less than 1.3%, less than 1.25%, less than 1.2%, less than 1.15%, and less than 1.1% of the total cfDNA present in the biological sample. The method according to claim 8, wherein the amount is less than 1.05%, less than 1.0%, less than 0.95%, less than 0.9%, less than 0.8%, less than 0.7%, less than 0.6%, less than 0.5%, less than 0.45%, less than 0.4%, less than 0.35%, less than 0.3%, less than 0.25%, less than 0.2%, less than 0.15%, less than 0.1%, less than 0.09%, less than 0.08%, less than 0.07%, less than 0.06%, or less than 0.05%.
10. The method according to any one of claims 1 to 9, wherein, in order to detect potential variants, the number of genomic loci assayed by the low-coverage, genome-wide nucleic acid sequencing in c) is at least 5,000, 10,000, 20,000, 35,000, 50,000, 75,000, 100,000, 150,000, 200,000, 250,000, 300,000, 400,000, 500,000, 600,000, 750,000, 875,000, 100,000, 200,000, 300,000, 400,000, 500,000, 750,000, or 10,000,000 genomic loci.
11. The method according to any one of claims 1 to 10, wherein the biological sample is derived from the serum, plasma, blood, saliva, urine, mucus, tears, sweat, semen, breast milk, lymph, cerebrospinal fluid, or amniotic fluid of the subject.
12. The method according to claim 11, wherein the biological sample is derived from the plasma of the subject.
13. The method according to claim 11 or 12, wherein the volume of the biological sample obtained from the subject is less than approximately 200 μL, less than 150 μL, less than 125 μL, less than 100 μL, less than 80 μL, less than 75 μL, less than 70 μL, less than 60 μL, less than 55 μL, less than 50 μL, less than 45 μL, less than 40 μL, less than 35 μL, less than 30 μL, less than 25 μL, less than 20 μL, less than 17.5 μL, less than 15 μL, less than 12.5 μL, less than 10 μL, less than 9 μL, less than 8 μL, less than 7 μL, less than 6 μL, less than 5 μL, less than 4 μL, less than 3 μL, less than 2.5 μL, less than 2 μL, less than 1.5 μL, less than 1 μL, or less than 0.5 μL.
14. The method according to any one of claims 11 to 14, wherein the volume of blood obtained from the subject for the biological sample is less than approximately 200 μL, less than 150 μL, less than 125 μL, less than 100 μL, less than 80 μL, less than 75 μL, less than 70 μL, less than 60 μL, less than 55 μL, less than 50 μL, less than 45 μL, less than 40 μL, less than 35 μL, less than 30 μL, less than 25 μL, less than 20 μL, less than 17.5 μL, less than 15 μL, less than 12.5 μL, less than 10 μL, less than 9 μL, less than 8 μL, less than 7 μL, less than 6 μL, less than 5 μL, less than 4 μL, less than 3 μL, less than 2.5 μL, less than 2 μL, less than 1.5 μL, less than 1 μL, or less than 0.5 μL.
15. The method according to claim 13 or 14, wherein the biological sample is obtained from the subject by a method using capillary-based collection.
16. The method according to any one of claims 1 to 15, wherein, in e), the individual identified minority components in the cfDNA include alternative heterozygous alleles or alternative homozygous alleles compared to the alleles of the genomic DNA from the subject.
17. The method according to any one of claims 1 to 16, wherein, in e), the individual identified minority components in the cfDNA include alternative homozygous alleles compared to alleles of genomic DNA from the subject.
18. The method according to any one of claims 1 to 17, wherein the detected variant includes a single nucleotide polymorphism (SNP), a small insertion or deletion (INDEL), a variable number tandem repeat (NVTR), a simple sequence repeat (SSR), or a simple tandem repeat (STR), or any combination thereof.
19. The method according to any one of claims 1 to 18, wherein the detected variant includes an SNP.
20. The method according to any one of claims 1 to 19, wherein the low-coverage, genome-wide nucleic acid sequencing in c) includes sequencing coverage of less than approximately 1x, less than 0.9x, less than 0.8x, less than 0.7x, less than 0.6x, less than 0.5x, less than 0.4x, less than 0.35x, less than 0.3x, less than 0.25x, less than 0.2x, less than 0.15x, less than 0.125x, less than 0.1x, less than 0.09x, less than 0.08x, less than 0.075x, less than 0.065x, less than 0.05x, less than 0.035x, less than 0.025x, less than 0.02x, less than 0.015x, less than 0.01x, or less than 0.005x of the target genome.
21. The method according to any one of claims 17 to 20, wherein the alternative homozygous alleles of the identified minority components include about 1000, less than 750, less than 500, less than 400, less than 300, less than 200, less than 150, less than 125, less than 110, less than 100, less than 90, less than 80, less than 75, less than 70, less than 65, less than 60, less than 55, less than 50, less than 45, less than 40, less than 35, less than 30, less than 25, or less than 20 loci.
22. The method according to claim 21, wherein the alternative homozygous alleles of the identified minority components include about 1000, less than 750, less than 500, less than 400, less than 300, less than 200, less than 150, less than 125, less than 110, less than 100, less than 90, less than 80, less than 75, less than 70, less than 65, less than 60, less than 55, less than 50, less than 45, less than 40, less than 35, less than 30, less than 25, or less than 20 SNP loci.
23. The method according to any one of claims 16 to 21, wherein the alternative heterozygous allele or alternative homozygous allele is derived from cfDNA from the target pre-malignant or malignant cells.
24. The method according to any one of claims 16 to 21, wherein the alternative heterozygous allele or alternative homozygous allele is derived from cfDNA from one or more infectious agents present in the subject.
25. The method according to any one of claims 16 to 21, wherein the alternative heterozygous allele or the alternative homozygous allele is derived from a donor.
26. The method according to claim 25, wherein the donor is an embryo.
27. The method according to claim 25, wherein the donor is a fetus.
28. The method according to claim 25, wherein the donor subject provides tissue transplantation or organ transplantation to a host subject.
29. The method according to any one of claims 1 to 28, wherein the estimated values of the minority component frequencies in the cfDNA-enriched biological sample in f) are used for transplant monitoring.
30. The method according to claim 29, wherein transplant monitoring includes distinguishing non-rejection (TX) from one or both of acute rejection (AR) and acute dysfunctional non-rejection (ADNR).
31. The method according to claim 29 or 30, wherein transplant monitoring includes comparing the level of donor-derived cell-free DNA (dd-cfDNA) detected in the cfDNA-enriched biological sample with a predetermined threshold.
32. The method according to claim 31, wherein the host is identified as potentially having AR or ADNR by dd-cfDNA levels that are at least 0.5%, 0.6%, 0.75%, 0.8%, 0.85%, 0.9%, 0.95%, 1.0%, 1.1%, 1.25%, 1.5%, 1.75%, 2.0%, 2.5%, 3.0%, 4.0%, 5.0%, 7.5%, or 10% higher than or equal to a predetermined threshold.
33. The method according to claim 31, wherein the host is found to have the potential to not reject a dd-cfDNA level that is approximately 1.5%, 1.25%, 1.0%, 0.9%, 0.8%, 0.7%, 0.6%, 0.5%, 0.45%, 0.4%, 0.35%, 0.30%, 0.25%, 0.2%, 0.15%, or 0.1% lower than a predetermined threshold.
34. The method according to any one of claims 29 to 33, wherein the time-dependent increase in the presence of donor-derived cell-free DNA indicates dysfunction or rejection of the transplanted organ or tissue.
35. The method according to any one of claims 1 to 23, wherein the estimation of the frequency of minority components in the cfDNA-enriched biological sample in f) is used for oncological detection or oncological monitoring.
36. The method according to claim 35, comprising analyzing at least two biological samples taken from the subject at different time points in time to screen for changes in minor cfDNA components over time.
37. The method according to claim 35 or 36, wherein a significant increase in somatic variants over time is detected.
38. The method according to any one of claims 35 to 37, further comprising identifying novel somatic cell variants detected in the biological sample taken from the subject after taking an initial biological sample from the subject.
39. The method according to any one of claims 35 to 38, wherein somatic cell SNPs are distinguished from germline SNPs in the subject.
40. The method according to any one of claims 35 to 39, wherein the initial biological sample is obtained from a tumor biopsy.
41. The method according to any one of claims 35 to 40, further comprising the step of analyzing the DNA methylation pattern in the cfDNA-enriched biological sample.
42. The method according to any one of claims 35 to 41, further comprising the step of analyzing the ratio of detected single-stranded cfDNA to detected double-stranded cfDNA in the cfDNA-enriched biological sample.
43. The method according to claim 24, further comprising the step of monitoring the progression of an infectious disease.
44. The method according to claim 43, further comprising the step of comparing the level of cfDNA detected from one or more infectious pathogens in the cfDNA-enriched biological sample with a predetermined threshold.
45. The method according to claim 43 or 44, wherein the step of monitoring the progression of an infectious disease includes detecting a significant increase in cfDNA from one or more infectious pathogens over time.
46. The method according to any one of claims 43 to 45, wherein the subject is monitored for the progression of sepsis.
47. The method according to claim 25 or 26, wherein the onset or progression of pregnancy complications is monitored.
48. The method according to any one of claims 1 to 47, further comprising the step of analyzing the fragment pattern in the cfDNA-enriched biological sample.
49. The method according to any one of claims 1 to 48, further comprising the step of supplementing missing SNP genotypes in the sequenced cfDNA-enriched biological sample.
50. The method according to any one of claims 1 to 49, further comprising the step of calculating and evaluating local linkage disequilibrium ratios between detected variant alleles in order to improve the calculated estimate of minority component frequencies in the cfDNA-enriched biological sample.
51. The method according to any one of claims 1 to 50, further comprising the step of adjusting the level of individual subjects, which includes sampling over time of biological samples obtained from the subject in order to screen for changes in minor cfDNA components over time.
52. The method according to any one of claims 1 to 51, wherein the low-coverage, genome-wide nucleic acid sequencing is unbiased sequencing.
53. This is a kit for analyzing cell-free DNA (cfDNA), a) A sample collector configured to collect a biological sample of the target, b) A set of instructions, a. A step of enriching the proportion of cfDNA in the biological sample. b. A step of sequencing a cfDNA-enriched biological sample using low-coverage, genome-wide nucleic acid sequencing. c. A step of identifying multiple minority components present in the sequenced cfDNA-enriched biological sample. d. A step of assigning a designation representing a low-confidence estimate of the minor variant frequency to each identified minority component present in the sequenced cfDNA-enriched biological sample, and e. A step of averaging multiple low-confidence estimates of minor variant frequencies across multiple sequenced genomic loci to present an estimate of the frequency of minor components in the cfDNA-enriched biological sample. Instructions for and A kit that includes this.
54. The kit according to claim 53, wherein the sample collector is configured to collect one or more biological samples, including whole blood, from the subject.