Methods and systems for detection of pre-cancer
A novel genomic assay using a single-molecule sequencing platform with targeted enrichment and ultra-deep sequencing effectively detects advanced adenomas and CRC by maximizing cfDNA yield and signal amplification, achieving enhanced sensitivity and specificity.
Patent Information
- Application Number
- PCT/US2025/049724
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-07
- Filing Date
- 2025-10-06
- Publication Date
- 2026-04-16
AI Technical Summary
Current diagnostic modalities for colorectal cancer (CRC) are limited in detecting pre-cancerous lesions, such as advanced adenomas, due to challenges in target enrichment and signal-to-noise ratios in cell-free DNA analysis, leading to suboptimal detection sensitivity and specificity.
A method utilizing a single-molecule sequencing platform with specific genomic assays, bioinformatics modeling, and capture probe sets to enrich for target sequences, combined with ultra-deep sequencing and error suppression, enhances the detection of advanced adenomas and CRC by maximizing cfDNA yield and signal amplification.
The method achieves greater than 90% sensitivity for CRC and 45% sensitivity for advanced adenomas at 95% specificity, significantly improving detection accuracy compared to existing tests.
Smart Images

Figure US2025049724_16042026_PF_FP_ABST
Abstract
Description
Attorney Docket No. 45249-710601METHODS AND SYSTEMS FOR DETECTION OF PRE-CANCERCROSS-REFERENCE
[0001] This application claims the benefit of U.S. Application No. 63 / 704,483, filed October 7, 2024, which is incorporated by reference herein in its entirety.BACKGROUND
[0002] Next-generation sequencing allows countless genomes to be sequenced in a fraction of the time that it once took. Despite these technical advances, whole genome sequencing remains very expensive and as a result target enrichment is necessary.
[0003] Colorectal cancer is one of the most prevalently diagnosed cancers in the US, affecting both men, women, as well as all racial and ethnic groups, with increased occurrence with age. Despite current screening modalities, CRC detection remains high due to limitations in diagnostics for pre-cancerous lesions, or advance adenomas (AA), which normally precede invasive cancer development for at least for a number of years.SUMMARY
[0004] Provided herein are methods that may allow for targeting specific colon cancer progression pathways to screen for and detect advanced adenomas (AA) and Colorectal Cancer (CRC). The methods can utilize a single-molecule sequencing platform which can comprise the use of specific genomic assay, panel design addressing clinically relevant variants, and / or bioinformatics modeling and analysis. Multiple elements can be combined to identify advanced adenoma with high sensitivity.
[0005] In an aspect, the present disclosure provides a method for detecting a presence or an absence of cancer or pre-cancer in a subject, said method comprising: (a) receiving a biological sample obtained or derived from said subject; (b) extracting cell-free nucleic acids from said biological sample to generate extracted cell-free nucleic acids; (c) processing said extracted cell-free nucleic acids, or derivatives thereof, to generate a nucleic acid sequencing library, wherein said processing comprises contacting said extracted cell-free nucleic acids, or derivatives thereof, with a capture probe set to enrich for one or more target sequences; (d) sequencing said nucleic acid sequencing library to produce sequence data; and (e) computer processing said sequence data to detect said presence or said absence of cancer or pre-cancer in said subject. In some embodiments, the sequencing comprises a depth of at least 10,000X.Attorney Docket No. 45249-710601In some embodiments, the sequencing comprises a depth of at least 100,000X. In some embodiments, the sequencing comprises sequencing nucleic acids corresponding to no more than 50 genomic loci. In some embodiments, the sequencing comprises sequencing nucleic acids corresponding to 10 to 45 genes. In some embodiments, the extracting comprises a bead-based capture. In some embodiments, the extracting is performed in a single vessel. In some embodiments, the extracting comprises protein denaturation. In some embodiments, the extracting comprises binding said cell-free nucleic acids to a solid support. In some embodiments, the extracting comprises an efficiency of at least 80% yield of cell-free nucleic acids in said biological sample.
[0006] In some embodiments, the processing comprises amplification. In some embodiments, the amplification comprises no more than 5 amplification cycles. In some embodiments, the amplification comprises no more than 10 amplification cycles. In some embodiments, the processing comprises a selection of nucleic acid fragment size. In some embodiments, the size selection comprises contacting said cell-free nucleic acids, or derivatives thereof, to one or more beads. In some embodiments, the processing comprises ligating one or more adaptors to said extracted cell-free nucleic acids. In some embodiments, the litigating is performed for at least 4 hours. In some embodiments, the ligating is performed for at least 8 hours. In some embodiments, the ligating is performed for at least 16 hours. The method of any of claims 1 to 19, wherein said ligating is performed at a temperature no greater than 25°C. In some embodiments, the ligating is performed at a temperature no greater than 20°C. In some embodiments, the ligating is performed at a temperature no greater than 10°C. In some embodiments, the ligating is performed at a temperature no greater than 4°C. In some embodiments, the wherein said contacting is performed subsequent to said ligating. In some embodiments, the wherein said contacting is performed prior to said ligating. In some embodiments, the one or more adaptors comprise one or more identifiers. In some embodiments, the identifiers comprise unique molecule identifiers (UMIs). In some embodiments, the identifiers comprise sample barcodes. In some embodiments, the nucleic acid library comprises nucleic acids derived from at least 90% of said cell-free nucleic acids initially present in said biological sample.
[0007] In some embodiments, the capture probe set comprises probes complementary to sequences associated with one or more biomarkers. In some embodiments, the one or more biomarkers comprise colorectal cancer driver genes. In some embodiments, the one or more biomarkers comprise NRAS, PTEN, FGFR, KRAS, TP53, SMAD4, PIK3CA, FBXW7, APC, SEPT9, or BRAF. In some embodiments, the capture probe set comprises no more thanAttorney Docket No. 45249-71060140 different probes. In some embodiments, the capture probe set comprises one or more probes of at least 100 or more nucleotides. In some embodiments, the probes of said capture probe set comprise 100 or more nucleotides. In some embodiments, the capture probe set comprises one or more probes of at least 120 or more nucleotides. In some embodiments, the probes of said capture probe set comprise at least 120 or more nucleotides. In some embodiments, the capture probe set comprises one or more probes comprising sequences identical or complementary to sequences of at least 50% identity, 60% identity, 70% identity, 80% identity, 90% identity, 95% identity, or 100% identity to SEQ ID NO: 1- 39. In some embodiments, the processing comprises, subsequent to said contacting, subjecting said extracted cell-free nucleic acids, or derivatives thereof, to a low stringency wash. In some embodiments, the processing comprises, subsequent to said contacting, subjecting said extracted cell-free nucleic acids, or derivatives thereof, to a high stringency wash.
[0008] In some embodiments, the biological sample is a blood sample. In some embodiments, the biological sample is a plasma sample. In some embodiments, the cell-free nucleic acids comprise cell-free DNA (cfDNA). In some embodiments, the cell-free nucleic acids comprise cell-free RNA (cfRNA).
[0009] In some embodiments, the method further comprises detecting said presence or said absence of cancer or pre-cancer in said subject at a positive predictive value of at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, the method further comprises detecting said presence or said absence of cancer or pre-cancer in the subject at a negative predictive value sensitivity of at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
[0010] In some embodiments, the method further comprises detecting said presence or said absence of cancer or pre-cancer in said subject at a sensitivity of sensitivity of at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some embodiments, the method further comprises detecting said presence or said absence of cancer or pre-cancer in said subject at a specificity of sensitivity of at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In someAttorney Docket No. 45249-710601 embodiments, (e) comprises computer processing said sequence data to detect said presence or said absence of cancer. In some embodiments, the cancer comprises colorectal cancer. In some embodiments, (e) comprises computer processing said sequence data to detect said presence or said absence of pre-cancer. In some embodiments, (e) comprises computer processing said sequence data to detect said absence of cancer or pre-cancer.
[0011] In some embodiments, the method further comprises identifying a clinical intervention for said subject based at least in part on said presence or said absence of said cancer or precancer. In some embodiments, the method further comprises administering said clinical intervention to said subject.
[0012] In another aspect, the present disclosure provides a method for detecting a presence or an absence of cancer or pre-cancer in a subject, said method comprising: (a) receiving a biological sample obtained or derived from said subject; (b) extracting cell-free nucleic acids from said biological sample to generate extracted cell-free nucleic acids, wherein said extracted nucleic acids comprise at least 79% of cell-free nucleic acids present in said biological sample; (c) processing said extracted cell-free nucleic acids, or derivatives thereof, to generate a nucleic acid sequencing library, wherein said processing comprises contacting said extracted cell-free nucleic acids, or derivatives thereof, with a capture probe set to enrich for one or more target sequences; (d) sequencing said nucleic acid sequencing library to produce sequence data; and (e) computer processing said sequence data to detect said presence or said absence of cancer or pre-cancer in said subject.
[0013] In another aspect, the present disclosure provide a method for detecting a presence or an absence of cancer or pre-cancer in a subject, said method comprising:
[0014] (a) receiving a biological sample obtained or derived from said subject; (b) extracting cell-free nucleic acids from said biological sample to generate extracted cell-free nucleic acids; (c) processing said extracted cell-free nucleic acids, or derivatives thereof, to generate a nucleic acid sequencing library, wherein said nucleic acid sequencing library comprises nucleic acids derived from at least 79% of said extracted cell-free nucleic acids; (d) enriching said nucleic acid sequencing library at least in part by contacting said nucleic acid sequencing library with a probe set to generate an enriched sequencing library; (e) sequencing said enriched sequencing library to produce sequence data; and (f) computer processing said sequence data
[0015] Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable code that, upon execution by one or more computer processors, implements any of the methods above or elsewhere herein.Attorney Docket No. 45249-710601
[0016] Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein.
[0017] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE
[0018] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “figure” and “FIG.” herein), of which:
[0020] FIG. 1 shows an example schematic of the disclosed methods.
[0021] FIG. 2 shows a computer system that is programmed or otherwise configured to implement methods provided herein.DETAILED DESCRIPTION
[0022] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided byAttorney Docket No. 45249-710601 way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.
[0023] The present disclosure provides a novel genomics-based assay in a process that has been developed to specifically and quantitatively detect AA as well as CRC, based on plasma cfDNA. Tests based on cfDNA have historically been limited due to signal to noise levels, where within this cfDNA, only a small fraction of molecules may be of tumor origin (ctDNA). The methods and systems provided can overcome these issues thorough comprehensive and efficient sample handling, optimized to maximize yield and target ctDNA amplification of signal against a large genomic background. Combining this wet-lab process with ultra-deep sequencing and loci-specifics error suppression, biological noise can be filtered across patients, allowing for increased sensitivity and specificity in the detection of both AA and CRC. For example, using the methods disclosed, a blood test demonstrated greater than 90%sensitivity for colorectal cancer and 45% sensitivity for advanced adenomas at 95% specificity, which are values significantly improved compared to both currently approved tests, as well as prospective tests in clinical trials.
[0024] Due to these intrinsic limitation in ctDNA detections and this fleeting nature of signal to noise levels in single nucleotide variant (SNV) and polymorphism analysis against a large genomic background, more and more CRC based tests have been moving away and further diversifying their targets using methylation and fragmentomics analysis. This process presents a clear advantage with late-stage cancers where widespread changes at the cellular levels can significantly improve your signal to noise levels. In the case, of a clinically validated multi-cancer early detection (MCED) screening, tests showed as good or better clinical sensitivity using targeted methylation markers only, compared to filtered wholeblood sequenced SNV analysis. Similarly, another existing process has been able to obtain higher sensitivity and specificity in their tests by combining SNV analysis with both targeted methylation markers and fragamentomics analysis. Unfortunately, while this process might be successful for MCED based tests and detection of late-stage CRC, this advantage from methylation markers tend to not improve sensitivity earlier in cancer progression, such as AA and precancerous lesions.
[0025] In cancer, while widespread changes can build over time, these changes are normally limited at the genome level early in cancer inception. Methylation markers as such can have only limited impact in AA detection, and may yet to be clinically demonstrated for high sensitivity and specificity. This analysis for methylation markers is also critical towardsAttorney Docket No. 45249-710601 sample processing. Analyzing cfDNA from plasma for both methylation and SNV type analytes, cannot be directly and easily combined. As such, samples may need to be either split or filtered before subsequent processing, before even larger losses associated with other downstream events such as bisulfite sequencing. While these losses can be normally overcome through increased signal availability and diversity, these losses early in cancer progression are likely invaluable and unsurmountable.
[0026] Generally, there are two types of target enrichment strategies: Amplicon based and Hybridization based. Amplicon strategy relies on enrichment via Polymerase Chain Reaction (PCR) based amplification of target using short complementary nucleotide sequences called primers. However, this strategy may result in missing fragments of DNA thus missing variants and introducing errors. The hybridization strategy on the other hand can comprise binding fragments of DNA based on complementarity resulting in efficient capture of all variants for a given target. However, the hybridization strategy can suffer from problems such as strand bias, uneven coverage and inefficient binding and capture. Current preparation techniques are based on liquid biopsy -based screening tests that rely on detection of methylation signatures across broad regions of the genome. Methylation events develops during later stages of cancer progression, and therefore limits the ability of existing technologies to detect advanced adenomas. Without wishing to be bound by theory, the present disclosure provides a hybridization / capture method to overcome these challenges. As such, screening abilities that can both target and detect AA, as early as possible, are key to proficient treatment and diagnostics of CRC.
[0027] The provided methods may allow for improved or maximized extraction of cfDNA from plasma samples. Additionally, the methods may allow for preserving and amplifying target signal that can be carried through into library preparation, target enrichment and sequencing. A cfDNA extraction as described herein may allow for consistently extracting over >90% of all the cfDNA in the plasma, which may greatly improve or maximize the potential carryover into library preparations. With such a process, where the input is maximized before any signal amplification, detections rates can be improved for DNA polymorphism and tracking progression of pre-cancerous lesions.
[0028] The present disclosure provides methods that can use a small, parsimonious gene panel of specific biomarkers (e.g., CRC driver genes). Such panel can allow assaying that both maximizes the coverage of clinically relevant loci, while allowing ultra-deep sequencing >40,000X for improved signal to noise (S / N) ratio. This process also may allow for AAAttorney Docket No. 45249-710601 specific loci modeling, using background suppression, allowing clear classification of signal from AA, CRC and negative / NNF samples.The present disclosure can allow for the identification of advanced precancerous lesions, including those in various subcategories. For example, the advanced precancerous lesions may include the following subcategories: 1) Adenoma with carcinoma in situ / high grade dysplasia, any size; 2) - Adenoma with villous growth pattern (>=25%), any size; 3 ) Adenoma >= 1.0 cm in size, or 4) Sessile serrated lesion, > =1.0 cm in size. These various clinical criteria may be used for identifying any sample that is cancerous or precancerous. The provided assay / platform can identify samples from a screening population exhibiting one or more of these four characteristics. The screening population can include a random sample from the whole of the population with no bias or pre-selection of sub cohorts (i.e. sex, gender, medical history and other possible cohorts.)
[0029] This disclosure provides methods and systems that make use of cell-free DNA to detect a set of gene expression variants to identify advanced precancerous lesions from a screening population.
[0030] In various cases, the subject may be asymptomatic for cancer. For example, the cancer may not exhibit any symptoms and the subject may be unaware of the presence of pre-cancer or advanced adenoma. The methods described herein may allow for pre-cancer or cancer to be identified at an earlier stage than would otherwise be possible. The identification of the presence of the pre-cancer or cancer at an earlier stage may allow a treatment option or recommendation to be determined at an earlier stage and may allow the subject to have an improved prognosis.
[0031] In various embodiments, biological samples are obtained or derived from subjects. The biological sample may comprise nucleic acids. The biological sample be a cell-free deoxyribonucleic acid (cfDNA) sample. The nucleic acid may be a DNA (e.g. doublestranded DNA, single-stranded DNA, single- stranded DNA hairpins, cDNA, genomic DNA, germline DNA, circulating tumor DNA (ctDNA), cell-free DNA (cfDNA)), an RNA (e.g. cfRNA, mRNA, cRNA, miRNA, siRNA, miRNA, snoRNA, piRNA, tiRNA, snRNA), or a DNA / RNA hybrids. The biological sample may be a derived from or contain a biological fluid. For example, the biological sample may be a plasma sample, a serum sample, a buffy coat sample, a peripheral blood mononuclear cell (PBMC) sample, a red blood cell sample, a urine sample, a saliva sample, or other body fluid sample. The biological sample may comprise or be a pleural fluid sample, peritoneal fluid sample, amniotic fluid sample, cerebrospinal fluid sample, lymphatic fluid sample, sweat sample, tear sample, semenAttorney Docket No. 45249-710601 sample, or any combination of biological fluid. In some cases, the biological sample may comprise multiple components. For example, the biological sample may be a whole blood sample. The biological sample may be subjected to reactions such to separate or fractionate a biological sample. For example, a whole blood sample may be a fractionated and cell free nucleic acids may be obtained. The whole blood sample may be fractionated using centrifugation such that blood cells may be separated from the plasma (which may contain cell free nucleic acid). A sample may be subjected to multiple rounds of separation or fractionation.
[0032] The biological samples may be subjected to additional reactions or conditions prior to assaying or sequencing nucleic acids. For example, the biological sample may be subjected to conditions that are sufficient to extract or enrich nucleic acids (e.g., cfDNA molecules), or generate library from nucleic acids (e.g., cfDNA molecules) present in the sample .
[0033] The methods disclosed herein may comprise enriching nucleic acids. For example, one or more nucleic acid molecules in a sample may be contacted with a . The enrichment reactions may comprise contacting a sample with one or more beads or bead sets. The enrichment reactions may comprise one or more hybridization reactions. For example, the enrichment reactions may comprise contacting a sample with one or more capture probes or bait molecules that hybridize to a nucleic acid molecule of the biological sample. The enrichment reaction may comprise differential amplification of a set of nucleic acid molecules. The enrichment reaction may enrich for a plurality of genetic loci or sequences corresponding to genetic loci. The enrichment reactions may comprise the use of primers or probes that may complementarity to sequences (or sequences upstream or downstream) of a sequence that is to be enriched. For example, a capture probe may comprise sequence complementarity to a set of genomic loci and allow the enrichment of the genomic loci. For example, a capture probe may comprise sequence complementarity to a gene selected from Table 1. The enrichments reactions may comprise a set of probes or primers. A set of capture probes may comprise 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, or more different probes. A set of capture probes may comprise no more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, or less different probes.
[0034] In various embodiment, washing or rinsing may be performed on nucleic acids. Washing or rinsing may comprise incubating or passing nucleic acids into a buffer. TheAttorney Docket No. 45249-710601 washing may remove unwanted nucleic acids or contaminants. Enriching of nucleic acids may comprise the use of washes or rinses. For example, binding of nucleic acids to one or more probes may allow for the enrichment. However, contaminants or non-specific binding events may occur such that other molecules may be bound to other otherwise captured by the one or more probes. The use of washes or rinses may allow for removal of contaminants or unwanted molecules. In various embodiments, multiple washes or rinses may be performed. For example, an enrichment may comprise contacting nucleic acids with a capture probe set and performing multiple washes. Multiple washes may allow for improved binding specificity, for example, by removing additional contaminants in a second (or third, or nth wash) that were not removed in a prior wash. Washes or rinses may be performed at various stringency. For example, a wash at a higher stringency wash may comprise harsher conditions that break or otherwise remove interactions, that would otherwise remain in a lower stringency wash. By increasing the stringency of a wash, specificity of binding may be improved. For example, nucleic acids that is only partially complementary (e.g., contains mismatches) to a probe, may be removed via a high stringency wash, which would otherwise remain in a lower stringency wash. Multiple washes may be performed of various stringencies. For example, nucleic acids may be subjected to a low stringency wash, followed by a higher stringency wash. A progression of stringency in washes in the enrichment may allow for improve specificity of binding, and may allow for enrichment of sets of nucleic acids that meet a particular threshold (e.g., no mismatches, no more than 1 mismatch, no more than 2 mismatches, etc.). Based at least on the enrichment specificity, downstream sequence data may comprise more reads pertaining to sequences of interest.
[0035] The methods disclosed herein may comprise extracting nucleic acid molecules in a sample. Performing efficient or high yield extraction may be valuable or critical for the methods or systems of the disclosure. For example, an extraction of nucleic acids (e.g., cfDNA) above a yield may allow for detection of cancer or pre-cancer, that would otherwise be undetected. For example, the extraction may be performed at a yield of at least 70%, 71%, 72,%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82,%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 9 ! %, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more. The yield may be based on the amount of nucleic acids present in the sample before extraction compared to the amount after extraction. The extraction reactions may lyse cells or disrupt nucleic acid interactions with the cell such that the nucleic acids may be isolated, purified, enriched, or subjected to other reactions. The extracting may comprise contacting a sample with one or more beads or bead sets. The extracting may comprise capturing nucleic acids to a bead orAttorney Docket No. 45249-710601 solid surface. The extraction may comprise washing or rinsing the beads or solid surfaces to remove contaminants or other unwanted molecules. The extracting may comprise the use of one or more separators. The one or more separators may comprise a magnetic separator. The extracting may comprise separating bead bound nucleic acid molecules from bead free nucleic acid molecules. The extracting may comprise protein denaturation or degradation. For example, the proteins in a sample may be subjected to proteinases, peptidases, or other enzymes to degrade proteins. In various embodiments, the extracting (or multiple portions thereof) may be performed in single vessel. For example, a vessel may be used and a bead based binding and protein degradation reaction may be performed in the same vessel. The reactions may be temporally separated in the same vessel. For example, a protein degradation may be performed followed a bead based capture of nucleic acid. The use of a single vessel for multiple reactions or steps of an extraction may allow for minimizing or reducing loss of nucleic acids. For example, transfer of samples between different tubes or vessels may cause residual sample or nucleic acids to remain in a vessel and not moved on to later steps or reactions in a process. Over multiple transfers, this effect may be amplified such that the ending yield of the extraction process is significantly reduced.
[0036] In various embodiments, vessels are used to perform reactions. The vessels may be at least 1 mL, 2 mL , 3 mL 4 mL, 5 mL, 6 mL, 7 mL 8mL, 9 mL, 10 mL, 15 mL, 20 mL , 30 mL 40 mL, 50 mL, 60 mL, 70 mL 80 mL, 90 mL, 100 mL, 200 mL , 300 mL 400 mL, 500 mL, 600 mL, 700 mL 800 mL, 900 mL, 1000 mL, or more. The vessels may be no more than 1 mL, 2 mL , 3 mL 4 mL, 5 mL, 6 mL, 7 mL 8mL, 9 mL, 10 mL, 15 mL, 20 mL , 30 mL 40 mL, 50 mL, 60 mL, 70 mL 80 mL, 90 mL, 100 mL, 200 mL , 300 mL 400 mL, 500 mL, 600 mL, 700 mL 800 mL, 900 mL, 1000 mL, or less.
[0037] The methods disclosed herein may comprise amplification or extension reactions. The amplification reactions may comprise polymerase chain reaction. The amplification reaction may comprise PCR-based amplifications, non-PCR based amplifications, or a combination thereof. The one or more PCR-based amplifications may comprise PCR, qPCR, nested PCR, linear amplification, or a combination thereof. The one or more non-PCR based amplifications may comprise multiple displacement amplification (MDA), transcription- mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, circle-to- circle amplification or a combination thereof. The amplification reactions may comprise an isothermal amplification.Attorney Docket No. 45249-710601
[0038] In various embodiments, the PCR or other amplification may be performed for an amount of cycles. PCR or other amplification techniques may be valuable at generating enough nucleic acid material such that other downstream reactions (e.g., sequencing may be performed). However, PCR and amplification reactions may also introduce errors into the amplicons, and these errors may propagate to future generated amplicons, and may ultimately introduce errors into the sequencing data. Reducing error in the amplicons and sequencing data may allow for more accurate or sensitive methods. In the case of pre-cancer or cancer, mutations or variants may be associated with a condition whereas the wild type or nonmutant version may indicate the lack of a condition. As such, removal of amplification introduced error may be relevant for the methods of the disclosure. Amplification or PCR using a small number of amplification cycles may help reduce error. For example, amplification may be performed using no more than 5 cycles. Amplification may be performed using no more than 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or less, cycles. Amplification may be performed using greater than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15, or more cycles.
[0039] The method disclosed herein may comprise adding an adaptor to nucleic acids (e.g., cell free nucleic acids). A reaction may comprise the addition of an adaptor comprising a barcode or tag to the nucleic acid. The barcode may be a molecular barcode or a sample barcode. For example, a barcode may comprise a barcode sequence which may be a degenerate n-mer. The sequence may be randomly generated or generated such to synthesize a specific barcode sequence. The barcode nucleic acid may be added to a sample such to label the nucleic acid molecules in the sample. The barcodes may be specific to a sample. For example, a plurality of barcode nucleic acids may be added to a sample in which the barcode sequence is the same. Upon barcoding of the nucleic acids, those originating from a same sample may have a same barcode sequence, and may allow a nucleic acid to be identified as belonging to a particular or given sample. A molecular barcode may also be used such that each molecule (or a plurality of molecules) in a same volume have a different molecular barcode. A molecular barcode may be a unique molecule barcode. For example, the molecular barcode may also be used such that each molecule (or a plurality of molecules) in a same volume have a different molecular barcode. This barcode may be subjected to amplification such that all amplicons derived from a molecule have the same barcode. In this way, molecules originating from a same molecule may be identified. The sequences reads may be processed based on the barcode sequences. For example, the processing may reduce errors or allow a molecule to be tracked. An adaptor may comprise multiple barcodes. ForAttorney Docket No. 45249-710601 example, the adaptor may comprise a sample barcode and a molecular barcode. An adaptor may be appended to one or both sides of a nucleic acid (e.g., cell free nucleic acid).
[0040] Adaptors may be appended or otherwise added or incorporated into a sequence by various reactions, for example an amplification, extension, or ligation reaction, and may be performed enzymatically using a nucleic acid polymerase or ligase. The ligation may be an overhang or blunt end ligation and the barcodes may comprise complementarity to nucleic acids to be barcoded. This complementarity may be a sequence derived from the sample from the subject or may be constant sequence generated via a reaction performed on the nucleic acids in the sample.
[0041] Additions of adaptors may allow for nucleic acids to be sequences for example the adaptor may comprise sequences that allow for nucleic acids to be sequenced by a sequencer. For example, a sequencing adaptor sequence may allow for a nucleic acid (or a derivative thereof) to be attached to flow cell of a sequencer. As such, the addition of adaptors to nucleic acids may be important for allowing nucleic acids to be sequences. Nucleic acids in a sample in which an adaptor is not appended to may ultimately not be sequenced, and data corresponding to these nucleic acids may be lost. Improving the efficiency or yield of the adaptor addition to nucleic may allow for more sample nucleic acids to be sequenced. The methods of the disclosure may allow for at least a 70%, 71%, 72,%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82,%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 9!%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more yield from an adaptor addition reaction. For example, of the initial cell free nucleic acids subjected to an adaptor addition reaction at least 70%, 71%, 72,%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82,%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 9 !%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more may have an adaptor added successfully. Additions of adaptors may be performed via a ligation reaction. The ligation reaction may be performed at a temperature of no greater than 25°C. The ligation reaction may be performed at a temperature of no greater than 20°C. The ligation reaction may be performed at a temperature of no greater than 10°C. The ligation reaction may be performed at a temperature of no greater than 4°C. The ligation reaction may be performed for various times. For example, the ligation reaction may be performed for at least 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9, hours, 10, hours, 11, hours, 12 hours, 13 hours, 14 hours, 15 hours, 16, hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 23 hour, 24 hours, or more. The ligation reaction may be performed at temperature and a time. For example, the ligation may be performed at no greater than 20°C for at least 12 hours. The ligation at a temperature or no more than 0°C,Attorney Docket No. 45249-7106011°C, 2°C, 3°C, 4°C, 5°C, 6°C, 7°C, 8°C, 9°C, 10°C, 11°C, 12°C, 13°C, 14°C, 15°C, 16°C, 17°C, 18°C, 19°C, 20°C, 21°C, 22°C, 23°C, 24°C, 25°C, 26°C, 27°C, 28°C, 29°C, 30°C, and for at least 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 7 hours, 8 hours, 9, hours, 10, hours, 11, hours, 12 hours, 13 hours, 14 hours, 15 hours, 16, hours, 17 hours, 18 hours, 19 hours, 20 hours, 21 hours, 22 hours, 23 hour, 24 hours, or more. A ligation performed at a low temperature for a longer incubation time may allow for a higher yield of ligation products.
[0042] In various aspects described throughout the disclosure, the nucleic acids may be subjected to sequencing reactions. The sequencing the reactions may be used on DNA, RNA or other nucleic acid molecules. Example of a sequencing reaction that may be used include capillary sequencing, next generation sequencing, Sanger sequencing, sequencing by synthesis, single molecule nanopore sequencing, sequencing by ligation, sequencing by hybridization, sequencing by nanopore current restriction, or a combination thereof. Sequencing by synthesis may comprise reversible terminator sequencing, processive single molecule sequencing, sequential nucleotide flow sequencing, or a combination thereof. Sequential nucleotide flow sequencing may comprise pyrosequencing, pH-mediated sequencing, semiconductor sequencing or a combination thereof. The sequencing reactions may comprise whole genome sequencing, whole exome sequencing, targeted sequencing, methylation-aware sequencing, enzymatic methylation sequencing, bisulfite methylation sequencing.
[0043] The sequencing reactions can be performed at various sequencing depths. The sequencing depths of a sequencing reaction may be selected or modulated. The sequencing reactions may comprise sequencing at a region a depth of at least lx, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, lOx, l lx ,12x, 13x, 14x, 15x, 16x, 17x, 18x, 19x, 20x, 25x, 30x, 35x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, lOOx, 200x, 300x, 400x, 500x, 600x, 700x, 800x, 900x,1000x, 2000x, 3000x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, 10,000x, 20,000x, 30,000x, 40,000x, 50,000x, 60,000x, 70,000x, 80,000x, 90,000, 100,000x, or more. The sequencing reactions may comprise sequencing a region at a depth of no more than lx, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, lOx, l lx ,12x, 13x, 14x, 15x, 16x, 17x, 18x, 19x, 20x, 25x, 30x, 35x, 40x, 45x, 50x, 60x, 70x, 80x, 90x, lOOx, 200x, 300x, 400x, 500x, 600x, 700x, 800x, 900x,1000x, 2000x, 3000x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, 10,000x, 20,000x, 30,000x, 40,000x, 50,000x, 60,000x, 70,000x, 80,000x, 90,000, 100,000x, or less. Using high depth sequencing may be advantageous for accuracy or sensitivity of detecting a sequence. In the case of pre-cancer or cancer, sequences indicative of pre-cancer may be rare and sequencing at high depth (e.g., ultra deep sequencing) may allow for detection of the samples.Attorney Docket No. 45249-710601
[0044] In various embodiments, a sequencing reaction may be performed using a set of capture probes. The sequencing reaction using a set of capture probes may be a deep sequencing reaction or ultra-deep sequencing reaction. For example, the sequencing reaction using a set of capture probes may be performed at an sequencing depth of 50x, 60x, 70x, 80x, 90x, lOOx, 200x, 300x, 400x, 500x, 600x, 700x, 800x, 900x,1000x, 2000x, 3000x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, 10,000x, 20,000x, 30,000x, 40,000x, 50,000x, 60,000x, 70,000x, 80,000x, 90,000, 100,000x, or more.
[0045] The sequencing of nucleic acids may generate sequence data. The sequencing reads may be processed such to generate data of improved quality. The sequencing reads may be generated with a quality score. The quality score may indicate an accuracy of a sequence read or a level or signal above a nose threshold for a given base call. The quality scores may be used for filtering sequencing reads. For example, sequencing reads may be removed that do not meet a particular quality score threshold. Amplification or PCR may generate error in amplicons such that the sequences are not identical to a parent sequence. Using sample barcodes or molecular barcodes, error correction may be performed. Error correction may include identifying sequence reads that do not corroborate with other sequences from a same sample or same original parent molecules. The use of identifiers may allow the identification of a same parent or sample.
[0046] Sequence data may be processed, for example, to detect the presence of absence of a pre-cancer or cancer. For example, analysis of sequencing data can be performed to extract unique molecular identifier’s (UMI), build consensus reads, align reads to the genome, perform variant detection, analysis, and calling. Unique molecular identifiers can be used to correct errors and allow for removal of systemic errors and amplification and selection biases. For example, reads can be grouped by the start and end within a few bases and same unique molecular identifier in the sequencing read. These reads are then used to create a consensus sequence. Based in part on sequence data (e.g., consensus sequences) a determination can be made to determine the presence of a biomarker(e.g., a variant) and the presence of the biomarker may be used to detect the presence or absence of the cancer or pre-cancer.
[0047] The sequence data may be analyzed by a prediction model that can combine NGS - derived sequencing information with other database held data to detect variants that are cancerous or pre-cancerous in nature. The presence of specific variants that correspond with cancer or precancer can be used to detect the cancer or pre-cancer. Signal for specific variants can indicate reads of tumor origin.Attorney Docket No. 45249-710601
[0048] In various embodiments, the sequence data may be processed using an algorithm. For example, data may be processed with an algorithm that can identify a read as having a variant. In another example a statistical model can be used, where machine learning algorithms and statistics can be used to develop a model specific to each loci based on training data from confirmed clinical samples. The models can account for the possibility of variant mutations, insertions, and deletions at each loci and determines the likelihood that a variant is derived from precancerous or cancerous cell-free DNA.
[0049] The methods and systems may use or identify the presence or absence of one or more biomarkers. A biomarker may comprise a nucleic acid sequence of a gene or genomic loci. A biomarker may comprise a single nucleotide (e.g., a single nucleotide variant). The presence of a nucleotide at a location may be indicative of a presence or absence of a precancer or cancer. For example, the methods and systems may use a set of capture probes that is able to bind to one or more genes or genomic loci. In some cases, the one or more genes or genomic loci may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47,48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72,73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97,98, 99, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180,185, 190, 195, 200, or more. In some cases, the methods and system may use a limited or smaller set of genes or genomic loci. For example, the one or more genes or genomic loci may be no more than 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47,48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72,73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97,98, 99, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180,185, 190, 195, 200, or less. For example, the methods and systems may identify the presence of one or mor mutation or variant (e.g., single nucleotide variant, or splice variant), in the one or more genes or genomic loci. For example, the one or more genes may comprise a mutation. The presence of the mutation in the genes may indicate a presence or absence of cancer or pre-cancer.
[0050] The methods and systems may use multiple biomarkers. The presence of a biomarker may be indicative of a presence or absence of cancer or pre-cancer. In some cases, the detection of a presence of a biomarker may allow for the detection a cancer or pre-cancer. In some cases, the detection of a presence of a biomarker may be indicative of a pre-cancer orAttorney Docket No. 45249-710601 cancer, however this may have a low sensitivity, specificity, or accuracy at determining that the subject has cancer or pre-cancer. The detection of multiple biomarkers may improve the detection (e.g., increase sensitivity, specificity, or accuracy) or pre-cancer or cancer. The methods and systems may comprise sequencing multiple genes or genomic loci and using multiple nucleotides in the sequencing data to make a determination of the presence or absence of cancer. For example, the methods or systems may assay multiple genomic loci, and nucleotides in the genomic loci (e.g., each nucleotide present at in loci) may be analyzed such that nucleotides at multiple positions (e.g., all positions at a loci) are analyzed. This may allow for the methods and systems to use multiple biomarkers and may combine data for multiple biomarkers to make a determination.
[0051] Libraries
[0052] The methods provided herein allow for the capture of a nucleic acid including a target oligonucleotide sequence from a library of nucleic acids that include a nucleic acid including a target oligonucleotide sequence. For example, a library can include a plurality of different nucleic acids, where each nucleic acid includes a different nucleic acid sequence. In some embodiments, each nucleic acid within a library has the same length or has approximately the same length. In some embodiments, each nucleic acid within a library is of a different length. In some embodiments, at least two nucleic acids within a library have a different length.
[0053] A library of nucleic acids can be a library of double-stranded DNAs, a library of single stranded DNAs, a library of single-stranded RNAs, a library of double-stranded RNAs, or a library of double-stranded nucleic acids made of one strand of DNA and one strand of RNA.
[0054] In some embodiments, the nucleic acid(s) in a library can be of chromosomal, plasmid, genomic, mitochondrial, exosomal, cell-free DNA, cellular (e.g., mammalian cellular), or viral origin. In some embodiments, one or both strands of a double-stranded nucleic acid molecule (e.g., any of the double-stranded nucleic acids described herein) can be captured using any of the methods described herein.
[0055] A library of nucleic acids may include a plurality of double-stranded nucleic acids (e.g., double-stranded DNAs, double-stranded RNAs, or double-stranded nucleic acids made of one strand of DNA and one strand of RNA) having a total length of, e.g., about 20 base pairs (bp) to about 5,000 bp, about 20 bp to about 4,000 bp, about 20 bp to about 3,000 bp, about 20 bp to about 2,000 bp, about 20 bp to about 1,500 bp about 20 bp to about 1,000 bp, about 20 bp to about 500 bp, about 20 bp to about 100 bp, about 20 bp to about 60 bp, about 20 bp to about 40 bp, about 100 bp to about 5,000 bp, about 100 bp to about 4,000 bp, aboutAttorney Docket No. 45249-710601100 bp to about 2,000 bp, about 100 bp to about 1,000 bp, about 100 bp to about 500 bp, about 100 bp to about 250 bp, about 100 bp to about 200 bp, about 250 bp to about 5,000 bp, about 250 bp to about 1,000 bp, about 250 bp to about 500 bp, about 500 bp to about 5,000 bp, about 500 bp to about 2,000 bp, about 500 bp to about 1,000 bp, about 1,000 bp to about 5,000 bp, about 1,000 bp to about 2,000 bp, about 1,500 bp to about 5,000 bp, about 1,500 bp to about 2,000 bp, about 2,000 bp to about 5,000 bp, about 2,000 bp to about 4,000 bp, about 3,000 bp to about 5,000 bp, about 3,000 bp to about 4,000 bp, or about 4,500 bp to about 5,000 bp.
[0056] A library of nucleic acids may include a plurality of single-stranded nucleic acids (e.g., single-stranded DNAs or single-stranded RNAs) having a total length of, e.g., about 20 nucleotides (nt) to about 5,000 nt, about 20 bp to about 4,000 bp, about 20 bp to about 3,000 bp, about 20 bp to about 2,000 bp, about 20 bp to about 1,500 bp, about 20 bp to about 1,000 bp, about 20 bp to about 500 bp, about 20 bp to about 100 bp, about 20 bp to about 60 bp, about 20 bp to about 40 bp, about 100 bp to about 5,000 bp, about 100 bp to about 4,000 bp, about 100 bp to about 2,000 bp, about 100 bp to about 1,000 bp, about 100 bp to about 500 bp, about 100 bp to about 250 bp, about 100 bp to about 200 bp, about 250 bp to about 5,000 bp, about 250 bp to about 1,000 bp, about 250 bp to about 500 bp, about 500 bp to about 5,000 bp, about 500 bp to about 2,000 bp, about 500 bp to about 1,000 bp, about 1,000 bp to about 5,000 bp, about 1,000 bp to about 2,000 bp, about 1,500 bp to about 5,000 bp, about 1,500 bp to about 2,000 bp, about 2,000 bp to about 5,000 bp, about 2,000 bp to about 4,000 bp, about 3,000 bp to about 5,000 bp, about 3,000 bp to about 4,000 bp, or about 4,500 bp to about 5,000 bp.
[0057] A library of nucleic acids may include a plurality of at least 1 x 103different nucleic acids, at least 1 x 104different nucleic acids, at least 1 x 105different nucleic acids, at least 1 x 106different nucleic acids, at least 1 x 107different nucleic acids, at least 1 x 108different nucleic acids, at least 1 x 109different nucleic acids, at least 1 x 1010different nucleic acids, at least 1 x 1011different nucleic acids, at least 1 x 1012different nucleic acids, at least 1 x 1013different nucleic acids, at least 1 x 1014different nucleic acids, or at least 1 x 1015different nucleic acids. For example, any of the libraries described herein can include a plurality of, e.g., about 1.0 x 102different nucleic acids to about 1.0 x 109different nucleic acids, about 1.0 x 102different nucleic acids to about 0.5 x 109different nucleic acids, about 1.0 x 102different nucleic acids to about 1.0 x 108different nucleic acids, about 1.0 x 102different nucleic acids to about 0.5 x 108different nucleic acids, about 1.0 x 102different nucleic acids to about 1.0 x 107different nucleic acids, about 1.0 x 102different nucleic acidsAttorney Docket No. 45249-710601 to about 0.5 x 107different nucleic acids, about 1.0 x 102different nucleic acids to about 1.0 x 106different nucleic acids, about 1.0 x 102different nucleic acids to about 0.5 x 106different nucleic acids, about 1.0 x 102different nucleic acids to about 1.0 x 105different nucleic acids, about 1.0 x 102different nucleic acids to about 0.5 x 105different nucleic acids, about 1.0 x 102different nucleic acids to about 1.0 x 104different nucleic acids, about 1.0 x102different nucleic acids to about 0.5 x 104different nucleic acids, about 1.0 x 103different nucleic acids to about 0.5 x 109different nucleic acids, about 1.0 x 103different nucleic acids to about 1.0 x 108different nucleic acids, about 1.0 x 103different nucleic acids to about 0.5 x 108different nucleic acids, about 1.0 x 103different nucleic acids to about 1.0 x 107different nucleic acids, about 1.0 x 103different nucleic acids to about 0.5 x 107different nucleic acids, about 1.0 x 103different nucleic acids to about 1.0 x 106different nucleic acids, about 1.0 x 103different nucleic acids to about 0.5 x 106different nucleic acids, about 1.0 x103different nucleic acids to about 1.0 x 105different nucleic acids, about 1.0 x 103different nucleic acids to about 0.5 x 105different nucleic acids, about 1.0 x 103different nucleic acids to about 1.0 x 104different nucleic acids, about 1.0 x 103different nucleic acids to about 0.5 x 104different nucleic acids, about 0.5 x 104different nucleic acids to about 1.0 x 109different nucleic acids, about 0.5 x 104different nucleic acids to about 0.5 x 109different nucleic acids, about 0.5 x 104different nucleic acids to about 1.0 x 108different nucleic acids, about 0.5 x 104different nucleic acids to about 0.5 x 108different nucleic acids, about 0.5 x104different nucleic acids to about 1.0 x 107different nucleic acids, about 0.5 x 104different nucleic acids to about 0.5 x 107different nucleic acids, about 0.5 x 104different nucleic acids to about 1.0 x 106different nucleic acids, about 0.5 x 104different nucleic acids to about 0.5 x 106different nucleic acids, about 0.5 x 104different nucleic acids to about 1.0 x 105different nucleic acids, about 0.5 x 104different nucleic acids to about 0.5 x 105different nucleic acids, about 0.5 x 104different nucleic acids to about 1.0 x 104different nucleic acids, about 1.0 x 104different nucleic acids to about 1.0 x 109different nucleic acids, about 1.0 x 104different nucleic acids to about 0.5 x 109different nucleic acids, about 1.0 x 104different nucleic acids to about 1.0 x 108different nucleic acids, about 1.0 x 104different nucleic acids to about 0.5 x 108different nucleic acids, about 1.0 x 104different nucleic acids to about 1.0 x 107different nucleic acids, about 1.0 x 104different nucleic acids to about 0.5 x 107different nucleic acids, about 1.0 x 104different nucleic acids to about 1.0 x 106different nucleic acids, about 1.0 x 104different nucleic acids to about 0.5 x 106different nucleic acids, about 1.0 x 104different nucleic acids to about 1.0 x 105different nucleic acids, about 1.0 x 104different nucleic acids to about 0.5 x 105different nucleic acids, about 0.5 x 105differentAttorney Docket No. 45249-710601 nucleic acids to about 1.0 x 109different nucleic acids, about 0.5 x 105different nucleic acids to about 0.5 x 109different nucleic acids, about 0.5 x 105different nucleic acids to about 1.0 x 108different nucleic acids, about 0.5 x 105different nucleic acids to about 0.5 x 108different nucleic acids, about 0.5 x 105different nucleic acids to about 1.0 x 107different nucleic acids, about 0.5 x 105different nucleic acids to about 0.5 x 107different nucleic acids, about 0.5 x 105different nucleic acids to about 1.0 x 106different nucleic acids, about 0.5 x 105different nucleic acids to about 0.5 x 106different nucleic acids, about 0.5 x 105different nucleic acids to about 1.0 x 105different nucleic acids, about 1.0 x 105different nucleic acids to about 1.0 x 109different nucleic acids, about 1.0 x 105different nucleic acids to about 0.5 x 109different nucleic acids, about 1.0 x 105different nucleic acids to about 1.0 x 108different nucleic acids, about 1.0 x 105different nucleic acids to about 0.5 x 108different nucleic acids, about 1.0 x 105different nucleic acids to about 1.0 x 107different nucleic acids, about 1.0 x105different nucleic acids to about 0.5 x 107different nucleic acids, about 1.0 x 105different nucleic acids to about 1.0 x 106different nucleic acids, about 1.0 x 105different nucleic acids to about 0.5 x 106different nucleic acids, about 0.5 x 106different nucleic acids to about 1.0 x 109different nucleic acids, about 0.5 x 106different nucleic acids to about 0.5 x 109different nucleic acids, about 0.5 x 106different nucleic acids to about 1.0 x 108different nucleic acids, about 0.5 x 106different nucleic acids to about 0.5 x 108different nucleic acids, about 0.5 x 106different nucleic acids to about 1.0 x 107different nucleic acids, about 0.5 x106different nucleic acids to about 0.5 x 107different nucleic acids, about 0.5 x 106different nucleic acids to about 1.0 x 106different nucleic acids, about 1.0 x 106different nucleic acids to about 1.0 x 109different nucleic acids, about 1.0 x 106different nucleic acids to about 0.5 x 109different nucleic acids, about 1.0 x 106different nucleic acids to about 1.0 x 108different nucleic acids, about 1.0 x 106different nucleic acids to about 0.5 x 108different nucleic acids, about 1.0 x 106different nucleic acids to about 1.0 x 107different nucleic acids, about 1.0 x106different nucleic acids to about 0.5 x 107different nucleic acids, about 0.5 x 107different nucleic acids to about 1.0 x 109different nucleic acids, about 0.5 x 107different nucleic acids to about 0.5 x 109different nucleic acids, about 0.5 x 107different nucleic acids to about 1.0 x 108different nucleic acids, about 0.5 x 107different nucleic acids to about 0.5 x 108different nucleic acids, about 0.5 x 107different nucleic acids to about 1.0 x 107different nucleic acids, about 1.0 x 107different nucleic acids to about 1.0 x 109different nucleic acids, about 1.0 x 107different nucleic acids to about 0.5 x 109different nucleic acids, about 1.0 x107different nucleic acids to about 1.0 x 108different nucleic acids, about 1.0 x 107different nucleic acids to about 0.5 x 108different nucleic acids, about 0.5 x 108different nucleic acidsAttorney Docket No. 45249-710601 to about 1.0 x 109different nucleic acids, about 0.5 x 108different nucleic acids to about 0.5 x 109different nucleic acids, about 0.5 x 108different nucleic acids to about 1.0 x 108different nucleic acids, about 1.0 x 108different nucleic acids to about 1.0 x 109different nucleic acids, about 1.0 x 108different nucleic acids to about 0.5 x 109different nucleic acids, or about 0.5 x 109different nucleic acids to about 1.0 x 109different nucleic acids.
[0058] In some embodiments of any of the methods described herein, a nucleic acid that is present in a library (e.g., and that is captured by the methods described herein) can include or consist of a sequence that has a high GC content. In some embodiments, the GC content of a nucleic acid in a library or a portion thereof (e.g., a target oligonucleotide sequence present in a nucleic acid in a library) can have a GC percentage of about 60% and above (e.g., about 62% and above, about 64% and above, about 65% and above, about 68% and above, about70% and above, about 72% and above, about 74% and above, about 75% and above, about78% and above, about 80% and above, about 82% and above, about 84% and above, about85% and above, about 88% and above, about 90% and above, about 92% and above, about94% and above, about 95% and above, or about 98% and above, or about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or 100%).
[0059] In some embodiments of any of the methods described herein, a nucleic acid that is present in a library (e.g., and that is captured by the methods described herein) can include or consist of a sequence that has a low GC content. In some embodiments, the GC content of a nucleic acid in a library or a portion thereof (e.g., a target oligonucleotide sequence present in a nucleic acid in a library) can have a GC percentage of about 59% and below (e.g., about 58% and below, about 56% and below, about 54% and below, about 52% and below, about50% and below, about 48% and below, about 46% and below, about 44% and below, about42% and below, about 40% and below, about 38% and below, about 36% and below, about34% and below, about 32% and below, about 30% and below, about 28% and below, about26% and below, about 24% and below, about 22% and below, about 20% and below, about18% and below, about 16% and below, about 14% and below, about 12% and below, about10% and below, about 8% and below, about 6% and below, about 4% and below, about 2% and below, or about 1% and below, or about 0.5%, about 1%, about 2%, about 3%, about 4%,Attorney Docket No. 45249-710601 about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 26%, about 27%, about 28%, about 29%, about 30%, about 31%, about 32%, about 33%, about 34%, about 35%, about 36%, about 37%, about 38%, about 39%, about 40%, about 41%, about 42%, about 43%, about 44%, about 45%, about 46%, about 47%, about 48%, about 49%, about 50%, about 51%, about 52%, about 53%, about 54%, about 55%, about 56%, about 57%, about 58%, or about 59%).
[0060] Some embodiments of any of the methods described herein further include generating a library comprising the steps of: fragmenting double-stranded DNA (e.g., cell free DNA, genomic DNA or cellular DNA from a mammalian cell, e.g., mammalian cells present in a biopsy sample) or obtaining or extracting nucleic fragments from a sample (e.g. cell free DNA fragments), performing end repair and dA-tailing, ligating an adaptor, and / or performing PCR amplification, thus yielding a library. Enrichment of a library may be performed, for example via methods described throughout this disclosure (e.g. using capture probes).Probes
[0061] In various cases, one or more probes may be used. For example, a probe may bind or hybridize to a nucleic acid (e.g. a nucleic acid in a library). A probe may be a single-stranded probe that includes a sequence that is complementary to a target oligonucleotide sequence (e.g., any of the target sequences described herein). The one or more probes may be one or more capture probes. The one or more probes may be a capture probe set.
[0062] In various aspects, one or more probes may be used that correspond to one or more biomarkers. The biomarkers may correspond to one or more cancer-associated genes or loci. The cancer-associated genomic loci may correspond to a set of genes.
[0063] Various example probes are provided in Table 1 below.
[0064] Table 1Atorney Docket No. 45249-710601Attomey Docket No. 45249-710601
[0065] The one or more probes may comprise sequences comprising SEQ ID NO: 1-29. The one or more probes may comprise sequences complementary to a portion of SEQ ID NO: 1-Attorney Docket No. 45249-71060129. The one or more probes may comprise sequences comprising portions of SEQ ID NO: 1- 29. The one or more probes may comprise sequences complementary to of SEQ ID NO: 1-29. The one or more probes may comprise sequences comprising at least a 50%, 60%, 70%, 80%, 90%, 95%, or 100% identity to SEQ ID NO: 1-29. The one or more probes may comprise sequences complementary (or the reverse complement) of sequences at least a 50%, 60%, 70%, 80%, 90%, 95%, or 100% identity to SEQ ID NO: 1-29.
[0066] In some embodiments of any of the methods described herein, the probe contains a total of about 10 nucleotides (nt) to about 800 nts, about 10 nt to about 500 nt, about 10 nt to about 250 nt, about 10 nt to about 100 nt, about 10 nt to about 50 nt, about 10 nt to about 40 nt, about 10 nt to about 30 nt, about 10 nt to about 20 nt, about 10 nt to about 15 nt, about 20 nt to about 800 nt, about 20 nt to about 500 nt, about 20 nt to about 200 nt, about 20 nt to about 100 nt, about 20 nt to about 50 nt, about 20 nt to about 40 nt, about 50 nt to about 800 nt about 50 nt to about 500 nt, about 50 nt to about 250 nt, about 50 nt to about 100 nt, about 100 nt to about 800 nt, about 100 nt to about 500 nt, about 100 nt to about 250 nt, about 100 nt to about 200 nt, about 150 nt to about 800 nt, about 150 nt to about 500 nt about 150 nt to about 250 nt, about 150 nt to about 200 nt, about 200 nt to about 800 nt, about 200 nt to about 500 nt, about 200 nt to about 400 nt, about 200 nt to about 300 nt, about 200 nt to about 250 nt, about 250 nt to about 500 nt, about 500 nt to about 800 nt,.
[0067] In some embodiments, the sequence that is complementary to a target oligonucleotide sequence and / or the target oligonucleotide sequence include a total of about 8 nucleotides (nt) to about 400 nt, about 8 nt to about 200 nt, about 8 nt to about 100 nt, about 8 nt to about 50 nt, about 8 nt to about 30 nt, about 8 nt to about 20 nt, about 8 nt to about 16 nt, about 8 nt to about 10 nt, about 10 nt to about 400 nt, about 10 nt to about 200 nt, about 10 nt to about 100 nt, about 10 nt to about 50 nt, about 10 nt to about 20 nt, about 20 nt to about 400 nt, about 20 nt to about 200 nt, about 20 nt to about 100 nt, about 50 nt to about 400 nt, about 50 nt to about 100 nt, about 75 nt to about 400 nt, about 75 nt to about 200 nt, about 75 nt to about 100 nt, about 100 nt to about 400 nt, about 100 nt to about 200 nt, about 200 nt to about 400 nt, about 200 nt to about 300 nt, about 300 nt to about 400 nt.
[0068] In some embodiments of any of the methods described herein, the sequence that is complementary to a target oligonucleotide sequence and / or the target oligonucleotide sequence can include or consist of a sequence that has a high GC content. In some embodiments, the GC content of the sequence that is complementary to a target oligonucleotide sequence and / or the target oligonucleotide sequence can have a GC percentage of about 60% and above (e.g., about 62% and above, about 64% and above, aboutAttorney Docket No. 45249-71060165% and above, about 68% and above, about 70% and above, about 72% and above, about 74% and above, about 75% and above, about 78% and above, about 80% and above, about 82% and above, about 84% and above, about 85% and above, about 88% and above, about 90% and above, about 92% and above, about 94% and above, about 95% and above, or about 98% and above, or about 60%, about 61%, about 62%, about 63%, about 64%, about 65%, about 66%, about 67%, about 68%, about 69%, about 70%, about 71%, about 72%, about 73%, about 74%, about 75%, about 76%, about 77%, about 78%, about 79%, about 80%, about 81%, about 82%, about 83%, about 84%, about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or 100%).
[0069] In some embodiments of any of the methods described herein, the sequence that is complementary to a target oligonucleotide sequence and / or the target oligonucleotide sequence can include or consist of a sequence that has a low GC content. In some embodiments, the GC content of a sequence that is complementary to a target oligonucleotide sequence and / or the target oligonucleotide sequence can have a GC percentage of about 59% and below (e.g., about 58% and below, about 56% and below, about 54% and below, about 52% and below, about 50% and below, about 48% and below, about 46% and below, about 44% and below, about 42% and below, about 40% and below, about 38% and below, about 36% and below, about 34% and below, about 32% and below, about 30% and below, about 28% and below, about 26% and below, about 24% and below, about 22% and below, about 20% and below, about 18% and below, about 16% and below, about 14% and below, about 12% and below, about 10% and below, about 8% and below, about 6% and below, about 4% and below, about 2% and below, or about 1% and below, or about 0.5%, about 1%, about 2%, about 3%, about 4%, about 5%, about 6%, about 7%, about 8%, about 9%, about 10%, about 11%, about 12%, about 13%, about 14%, about 15%, about 16%, about 17%, about 18%, about 19%, about 20%, about 21%, about 22%, about 23%, about 24%, about 25%, about 26%, about 27%, about 28%, about 29%, about 30%, about 31%, about 32%, about 33%, about 34%, about 35%, about 36%, about 37%, about 38%, about 39%, about 40%, about 41%, about 42%, about 43%, about 44%, about 45%, about 46%, about 47%, about 48%, about 49%, about 50%, about 51%, about 52%, about 53%, about 54%, about 55%, about 56%, about 57%, about 58%, or about 59%).
[0070] A capture probe may comprise sequence complementarity to a set of genomic loci and allow the enrichment of the genomic loci. For example, a capture probe may comprise sequence complementary to a gene. The capture probe may comprise a sequenceAttorney Docket No. 45249-710601 complementary or identical to a sequence from a protooncogene. The capture probe may comprise a sequence complementary or identical to a sequence from an oncogene. The capture probe may comprise a sequence complementary or identical to a sequence from an oncogenic kinase fusion protein. The capture probe may comprise a sequence complementary or identical to a sequence from an oncogenic kinase fusion protein. The capture probe may comprise a sequence complementary or identical to a sequence from a tumor suppressor gene.
[0071] For example, a capture probe may comprise a sequence complementary or identical to a sequence (or portion of the sequence) of NRAS, PTEN, FGFR, KRAS, TP53, SMAD4, PIK3CA, FBXW7, APC, SEPT9,or BRAF. A set of capture probes may comprise a sequence complementary or identical to a sequence (or portion of the sequence) of NRAS, PTEN, FGFR, KRAS, TP53, SMAD4, PIK3CA, FBXW7, APC, SEPT9, BRAF, or any combination thereof. A set of capture probes may comprise a sequence complementary or identical to a sequence (or portion of the sequence) of NRAS, PTEN, FGFR, KRAS, TP53, SMAD4, PIK3CA, FBXW7, APC, SEPT9, and BRAF. . A set of capture probes may comprise at least2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, or more different probes. A set of capture probes may comprise no more than 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, or less different probes. In various embodiments, the set of probes may comprise no more than 40 different probes. A small probe set may allow for a small group of nucleic acids to be enriched. For example, enrichment may be performed using a small or parsimonious probe set specifically designed to increase fidelity for detection of cancer or pre-cancer (e.g., CRC or advance adenoma detection). For example, a selection of CRC driver genes may allow for maximizing coverage of clinically relevant loci, while maintaining small panel size, which may allow for ultra-deep sequencing for improved signal to noise ratio.
[0072] The one or more probes may be complementary or may be able to hybridize to one or more exons of one or more genes. For example, the one or more probes may be complementary to or identical to a sequence of an exon or portion of an exon of NRAS, PTEN, FGFR, KRAS, TP53, SMAD4, PIK3CA, FBXW7, APC, SEPT9, or BRAF. For example, the one or more probes may be complementary to or identical to any of a sequence, or a portion thereof, of exon 1,2, 3, 4, 5, 6, or 7, of NRAS, exon 1, 2, 3, 4, ,5, 6, 7, 8, or 9 of PTEN, exon 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16, 17, or 18of FGFR, exon 1, 2,3, 4, 5, 6, 7, or 8 of KRAS, exon 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 11 of TP53, exon 1, 2, 3, 4, 5,Attorney Docket No. 45249-7106016, 7, 8 9, 10, 11, 12, 13, or 14 of SMAD4, exon 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 21 ofPIK3CA, exon 1, 2, 3, 4, 5, 6, 7,8, 9, 10, 11, 12, or 13, of FBXW7, exon 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 of APC, exon 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 ofBRAF, or exon 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 21 of SETP9.
[0073] The one or more probes may be complementary or may be able to hybridize to one or more introns of one or more genes. For example, the one or more probes may be complementary to or identical to a sequence of an intron or portion of an intron of NRAS, PTEN, FGFR, KRAS, TP53, SMAD4, PIK3CA, FBXW7, APC, SEPT9, or BRAF. The one or more probes may be complementary or may be able to hybridize to part of an intron and part of an exon of one or more genes. For example, the one or more probes may be complementary to or identical to a sequence of portion of an intron and a portion of an exon of NRAS, PTEN, FGFR, KRAS, TP53, SMAD4, PIK3CA, FBXW7, APC, SEPT9, or BRAF.
[0074] The choice of exons targeted by the probes may be based at least in part by one or more databases or guidelines relating to genes, sequences, or variants of interest. For example, the National Comprehensive Cancer Network (NCCN) guidelines may be used to determine actionable variants. For example, exons may be selected or targeted based at least on those that harbor most (or a large number) cancer variants based on The Cancer Genome Atlas (TGCA). In another example, probes may be complementary to exons that cover maximum number of variants over large population of patients based on data from Catalog of Somatic Mutations Cancer (COSMIC).
[0075] In order to improve coverage uniformity, the probes may include extra nucleotides flanking exons, and may comprise sequences complementary or identical to exon-intron splice junction. Probes designed as such, may allow for covering the entirety of the exon uniformly without any exon dropouts. Furthermore, to have high hybridization efficiency, the probes may be designed to have the coverage to panel size ratio as high as possible. In this way, the number of probes may be small yet offer broad coverage thereby improving hybridization efficiency.
[0076] In various aspects, the systems and methods may comprise an accuracy, sensitivity, or specificity of detection of the pre-cancer or cancer. For example, the methods or systems may comprise detecting the presence or the absence of pre-cancer or cancer in the subject at an accuracy of at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at leastAttorney Docket No. 45249-710601 about 98%, or at least about 99%. The methods or systems may comprise detecting the presence or the absence of pre-cancer or cancer in the subject at a sensitivity of at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The methods or systems may comprise detecting the presence or the absence of precancer or cancer in the in the subject at a specificity of at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The the methods or systems may comprise detecting the presence or the absence of pre-cancer or cancer in the subject at a positive predictive value of at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. The methods or systems may comprise detecting the presence or the absence of pre-cancer or cancer in the subject at a negative predictive value of at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
[0077] In various aspects as described herein, a clinical intervention or a therapy may be identified at least in part based on the identification of the presence of pre-cancer or cancer, or the presence of a parameter of cancer. The clinical intervention may be a surgical resection, chemotherapy, radiotherapy, immunotherapy, adjuvant therapy, neoadjuvant therapy, or a combination thereof. In some cases, the clinical interventions may be administered to the subject.
[0078] Computer control systems
[0079] The present disclosure provides computer systems that are programmed to implement methods of the disclosure. FIG. 2 shows a computer system 201 that is programmed or otherwise configured to perform analysis or operations of the methods, for example determine a likelihood of the presence of a cancer based on a set of biomarkers of an individual or run an algorithm. The computer system 201 can regulate various aspects of methods and systems of the present disclosure, such as, for example, perform an algorithm, input training data, analyze sets of biomarker, or output a result for the user as to the presenceAttorney Docket No. 45249-710601 or absence of cancer. The computer system 201 can be an electronic device of a user or a computer system that is remotely located with respect to the electronic device. The electronic device can be a mobile electronic device.
[0080] The computer system 201 includes a central processing unit (CPU, also “processor” and “computer processor” herein) 205, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 201 also includes memory or memory location 210 (e.g., random-access memory, read-only memory, flash memory), electronic storage unit 215 (e.g., hard disk), communication interface 220 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 225, such as cache, other memory, data storage and / or electronic display adapters. The memory 210, storage unit 215, interface 220 and peripheral devices 225 are in communication with the CPU 205 through a communication bus (solid lines), such as a motherboard. The storage unit 215 can be a data storage unit (or data repository) for storing data. The computer system 201 can be operatively coupled to a computer network (“network”) 230 with the aid of the communication interface 220. The network 230 can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network 230 in some cases is a telecommunication and / or data network. The network 230 can include one or more computer servers, which can enable distributed computing, such as cloud computing. The network 230, in some cases with the aid of the computer system 201, can implement a peer-to-peer network, which may enable devices coupled to the computer system 201 to behave as a client or a server.
[0081] The CPU 205 can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 210. The instructions can be directed to the CPU 205, which can subsequently program or otherwise configure the CPU 205 to implement methods of the present disclosure. Examples of operations performed by the CPU 205 can include fetch, decode, execute, and writeback.
[0082] The CPU 205 can be part of a circuit, such as an integrated circuit. One or more other components of the system 201 can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0083] The storage unit 215 can store files, such as drivers, libraries and saved programs. The storage unit 215 can store user data, e.g., user preferences and user programs. The computer system 201 in some cases can include one or more additional data storage units that areAttorney Docket No. 45249-710601 external to the computer system 201, such as located on a remote server that is in communication with the computer system 201 through an intranet or the Internet.
[0084] The computer system 201 can communicate with one or more remote computer systems through the network 230. For instance, the computer system 201 can communicate with a remote computer system of a user (e.g., a medical professional or patient). Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC’s (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system 201 via the network 230.
[0085] Methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 201, such as, for example, on the memory 210 or electronic storage unit 215. The machine executable or machine readable code can be provided in the form of software. During use, the code can be executed by the processor 205. In some cases, the code can be retrieved from the storage unit 215 and stored on the memory 210 for ready access by the processor 205. In some situations, the electronic storage unit 215 can be precluded, and machine-executable instructions are stored on memory 210.
[0086] The code can be pre-compiled and configured for use with a machine having a processer adapted to execute the code, or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a pre-compiled or as-compiled fashion.
[0087] Aspects of the systems and methods provided herein, such as the computer system 201, can be embodied in programming. Various aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk. “Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into theAttorney Docket No. 45249-710601 computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carry such waves, such as wired or wireless links, optical links or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.
[0088] Hence, a machine readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0089] The computer system 201 can include or be in communication with an electronic display 235 that comprises a user interface (UI) 240 for providing, for example, an input of biomarkers or sequencing data, or a visual output relating to a detection, diagnosis, or prognosis. Examples of UI’s include, without limitation, a graphical user interface (GUI) and web-based user interface.
[0090] Methods and systems of the present disclosure can be implemented by way of one or more algorithms. An algorithm can be implemented by way of software upon execution byAttorney Docket No. 45249-710601 the central processing unit 205. The algorithm can, for example, determine a presence or absence of a cancer or precancer based on a set of sequence data from a sample derived from a subject.
[0091] EXAMPLES
[0092] Example 1 : Analysis of cell free DNA for detection of pre cancer or cancer
[0093] Using methods and systems of the present disclosure, precancer or cancer can be detected in a subject. FIG. 1 provides an example schematic of a method used. A biological sample was collected from a subject. The sample was subjected to extraction reactions to extract cell-free nucleic acids. The extraction reactions comprised a protein denaturation, and a capture of nucleic acids using a bead or other solid support. The extraction reaction was performed in a single tube to minimize loss of liquid from transfer to new tube. Once the cell- free nucleic were extracted, library preparation was performed. To generate a library, adaptor ligation was performed. Adaptor comprising identifiers (e.g., molecular identifiers, or sample identifier) were litigated to the nucleic acid. To drive the reaction to completion, the litigation was performed in an excess of ligase using an incubation time longer than 2 hrs (e.g., 24 hrs) and was performed at a temperature below 20°C. After ligation, amplification of the library was performed using PCR. The PCR was performed for less than 10 cycles.
[0094] Enrichment of the library was performed. Capture probes corresponding to genes associated with pre-cancer were used. To improve sensitivity, a capture probe set of no greater than 50 probes were utilized. Captured nucleic acids were then subjected to multiple washes of increasing stringency, and the capture nucleic acids were eluted.
[0095] Sequencing on the capture nucleic acids were performed at high depth (e.g., >100,000X). The resulting sequence reads were processed using a computer algorithm, such to suppress or remove errors, generate consensus sequences, and make variant calls. Based on the processing, the determination of the presence or absence of pre-cancer or cancer is made.
Claims
Attorney Docket No. 45249-710601CLAIMSWHAT IS CLAIMED IS:
1. A method for detecting a presence or an absence of cancer or pre-cancer in a subject, said method comprising:(a) receiving a biological sample obtained or derived from said subject;(b) extracting cell-free nucleic acids from said biological sample to generate extracted cell- free nucleic acids;(c) processing said extracted cell-free nucleic acids, or derivatives thereof, to generate a nucleic acid sequencing library, wherein said processing comprises contacting said extracted cell-free nucleic acids, or derivatives thereof, with a capture probe set to enrich for one or more target sequences;(d) sequencing said nucleic acid sequencing library to produce sequence data; and(e) computer processing said sequence data to detect said presence or said absence of cancer or pre-cancer in said subject.
2. The method of claim 1, wherein said sequencing comprises a depth of at least 10,000X.
3. The method of any of claim 1 or 2, wherein said sequencing comprises a depth of at least 100,000X.
4. The method of any of claim 1 to 3, wherein said sequencing comprises sequencing nucleic acids corresponding to no more than 50 genomic loci.
5. The method of any of claims 1 to 4, wherein said sequencing comprises sequencing nucleic acids corresponding to 10 to 45 genes.
6. The method of any of claims 1 to 5, wherein said extracting comprises a bead-based capture.
7. The method of any of claims 1 to 6, wherein said extracting is performed in a single vessel.
8. The method of any of claims 1 to 7, wherein said extracting comprises protein denaturation.
9. The method of any of claims 1 to 8, wherein said extracting comprises binding said cell- free nucleic acids to a solid support.
10. The method of any of claims 1 to 9, wherein said extracting comprises an efficiency of at least 80% yield of cell-free nucleic acids in said biological sample.
11. The method of any of claims 1 to 10, wherein said processing comprises amplification.
12. The method of any of claims 1 to 11, wherein said amplification comprises no more than 5 amplification cycles.Attorney Docket No. 45249-71060113. The method of any of claims 1 to 12, wherein said amplification comprises no more than 10 amplification cycles.
14. The method of any of claims 1 to 13, wherein said processing comprises a selection of nucleic acid fragment size.
15. The method of any of claims 1 to 14, wherein said size selection comprises contacting said cell-free nucleic acids, or derivatives thereof, to one or more beads.
16. The method of any of claims 1 to 15, wherein said processing comprises ligating one or more adaptors to said extracted cell-free nucleic acids.
17. The method of any of claims 1 to 16, wherein said litigating is performed for at least 4 hours.
18. The method of any of claims 1 to 17, wherein said ligating is performed for at least 8 hours.
19. The method of any of claims 1 to 18, wherein said ligating is performed for at least 16 hours.
20. The method of any of claims 1 to 19, wherein said ligating is performed at a temperature no greater than 25 °C.
21. The method of any of claims 1 to 20, wherein said ligating is performed at a temperature no greater than 20°C.
22. The method of any of claims 1 to 21, wherein said ligating is performed at a temperature no greater than 10°C.
23. The method of any of claims 1 to 22, wherein said ligating is performed at a temperature no greater than 4°C.
24. The method of any of claims 16 to 23, wherein said contacting is performed subsequent to said ligating.
25. The method of any of claims 16 to 24, wherein said contacting is performed prior to said ligating.
26. The method of any of claims 16 to 25, wherein said one or more adaptors comprise one or more identifiers.
27. The method of claim 26, wherein said identifiers comprise unique molecule identifiers (UMIs).
28. The method of any of claims 26 to 27, wherein said identifiers comprise sample barcodes.
29. The method of any of claims 1 to 28, wherein said nucleic acid library comprises nucleic acids derived from at least 79% of said cell-free nucleic acids initially present in said biological sample.Attorney Docket No. 45249-71060130. The method of any of claims 1 to 29, wherein said capture probe set comprises probes complementary to sequences associated with one or more biomarkers.
31. The method of claim 30, wherein said one or more biomarkers comprise colorectal cancer driver genes.
32. The method of any of claims 30 to 31, wherein said one or more biomarkers comprise NRAS, PTEN, FGFR, KRAS, TP53, SMAD4, PIK3CA, FBXW7, APC, SEPT9, or BRAE33. The method of any of claims 1 to 32, wherein said capture probe set comprises no more than 40 different probes.
34. The method of any of claims 1 to 33, wherein said capture probe set comprises one or more probes of at least 100 or more nucleotides.
35. The method of any of claims 1 to 34, wherein probes of said capture probe set comprise 100 or more nucleotides.
36. The method of any of claims 1 to 35, wherein said capture probe set comprises one or more probes of at least 120 or more nucleotides.
37. The method of any of claims 1 to 36, wherein probes of said capture probe set comprise at least 120 or more nucleotides.
38. The method of any of claims 1 to 37, wherein said capture probe set comprises one or more probes comprising sequences identical or complementary to sequences of at least 50% identity, 60% identity, 70% identity, 80% identity, 90% identity, 95% identity, or 100% identity to SEQ ID NO: 1- 39.
39. The method of any of claims 1 to 38, wherein said processing comprises, subsequent to said contacting, subjecting said extracted cell-free nucleic acids, or derivatives thereof, to a low stringency wash.
40. The method of any of claims 1 to 39, wherein said processing comprises, subsequent to said contacting, subjecting said extracted cell-free nucleic acids, or derivatives thereof, to a high stringency wash.
41. The method of any of claims 1 to 40, wherein said biological sample is a blood sample.
42. The method of any of claims 1 to 41, wherein said biological sample is a plasma sample.
43. The method of any of claims 1 to 42, wherein said cell-free nucleic acids comprise cell- free DNA (cfDNA).
44. The method of any of claims 1 to 43, wherein said cell-free nucleic acids comprise cell- free RNA (cfRNA).Attorney Docket No. 45249-71060145. The method of any of claims 1 to 44, further comprising detecting said presence or said absence of cancer or pre-cancer in said subject at a positive predictive value of at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
46. The method of any of claims 1 to 45, further comprising detecting said presence or said absence of cancer or pre-cancer in the subject at a negative predictive value of at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
47. The method of any of claims 1 to 46, further comprising detecting said presence or said absence of cancer or pre-cancer in said subject at a sensitivity of at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
48. The method of any of claims 1 to 47, further comprising detecting said presence or said absence of cancer or pre-cancer in said subject at a specificity of at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%.
49. The method of any of claims 1 to 48, wherein (e) comprises computer processing said sequence data to detect said presence or said absence of cancer.
50. The method of claim 49 , wherein said cancer comprises colorectal cancer.
51. The method of any of claims 1 to 50, wherein (e) comprises computer processing said sequence data to detect said presence or said absence of pre-cancer.
52. The method of any of claims 1 to 51, wherein (e) comprises computer processing said sequence data to detect said absence of cancer or pre-cancer.
53. The method of any of claims 1 to 52, further comprising identifying a clinical intervention for said subject based at least in part on said presence or said absence of said cancer or precancer.
54. The method of claim 53, further comprising administering said clinical intervention to said subject.
55. A method for detecting a presence or an absence of cancer or pre-cancer in a subject, said method comprising:Attorney Docket No. 45249-710601(a) receiving a biological sample obtained or derived from said subject;(b) extracting cell-free nucleic acids from said biological sample to generate extracted cell- free nucleic acids, wherein said extracted nucleic acids comprise at least 90% of cell-free nucleic acids present in said biological sample;(c) processing said extracted cell-free nucleic acids, or derivatives thereof, to generate a nucleic acid sequencing library, wherein said processing comprises contacting said extracted cell-free nucleic acids, or derivatives thereof, with a capture probe set to enrich for one or more target sequences;(d) sequencing said nucleic acid sequencing library to produce sequence data; and(e) computer processing said sequence data to detect said presence or said absence of cancer or pre-cancer in said subject.
56. A method for detecting a presence or an absence of cancer or pre-cancer in a subject, said method comprising:(a) receiving a biological sample obtained or derived from said subject;(b) extracting cell-free nucleic acids from said biological sample to generate extracted cell- free nucleic acids;(c) processing said extracted cell-free nucleic acids, or derivatives thereof, to generate a nucleic acid sequencing library, wherein said nucleic acid sequencing library comprises nucleic acids derived from at least 80% of said extracted cell-free nucleic acids;(d) enriching said nucleic acid sequencing library at least in part by contacting said nucleic acid sequencing library with a probe set to generate an enriched sequencing library;(e) sequencing said enriched sequencing library to produce sequence data; and(f) computer processing said sequence data.
Citation Information
Patent Citations
Methods for early detection of cancer
US20190085406A1
SYSTEMS AND METHODS FOR DETECTING CANCER VIA cfDNA SCREENING
US20210155992A1
Methods and compositions for analyses of cancer
US20230002831A1