Methods, compositions and systems for improving binding of methylated polynucleotides
The method using modified carrier nucleic acid molecules improves tumor detection in liquid biopsies by processing and sequencing divided nucleic acid sets, addressing the challenges of low and heterogeneous nucleic acid amounts in bodily fluids.
Patent Information
- Application Number
- JP2025207226
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-11-26
- Filing Date
- 2025-11-27
- Publication Date
- 2026-03-04
AI Technical Summary
Existing liquid biopsy tests for cancer detection face challenges due to low amounts of nucleic acids in bodily fluids and their heterogeneous forms, including RNA and DNA with various post-replicative modifications, making it difficult to accurately detect tumors.
A method involving the use of carrier nucleic acid molecules, modified to prevent ligation, is added to polynucleotide samples to generate divided sets, followed by processing, sequencing, and analyzing these sets to detect the presence or absence of tumors, utilizing capture agents and enrichment for specific regions of interest.
Enhances the detection of tumors by improving the sensitivity and specificity of liquid biopsy tests, allowing for accurate analysis of methylated and unmethylated nucleic acids, even in low concentrations.
Smart Images

Figure 2026035733000010 
Figure 2026035733000011 
Figure 2026035733000012
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 940,853, filed November 26, 2019, which is incorporated herein by reference for all purposes. [Background technology]
[0002] background Cancer is a leading cause of disease worldwide. Each year, tens of millions of people worldwide are diagnosed with cancer, and more than half of those ultimately die from it. In many countries, cancer ranks as the second most common cause of death after cardiovascular disease. Early detection is associated with improved outcomes for many cancers.
[0003] Cancer can be caused by the accumulation of genetic variations within an individual's normal cells, at least in part resulting in inappropriately regulated cell division. Such variations commonly include copy number variations (CNVs), single nucleotide variations (SNVs), gene fusions, insertions and / or deletions (indels), and epigenetic variations include 5-methylation of cytosine (5-methylcytosine) and the association of DNA with chromatin and transcription factors.
[0004] Cancer is often detected by tumor biopsy followed by analysis of cells, markers, or DNA extracted from the cells. However, more recently, it has been proposed that cancer can also be detected from cell-free nucleic acids in bodily fluids such as blood or urine. Such tests have the advantage of being noninvasive and can be performed without identifying suspected cancer cells in the biopsy. However, such liquid biopsy tests are complicated by the fact that the amount of nucleic acid in bodily fluids is very low and that the nucleic acid present is in heterogeneous forms (e.g., RNA and DNA, single-stranded and double-stranded, as well as various states of post-replicative modifications and association with proteins such as histones). Summary of the Invention [Means for solving the problem]
[0005] overview In certain aspects, the disclosure provides a method for detecting the presence or absence of a tumor in a subject, comprising the steps of: (i) obtaining a polynucleotide sample from the subject; (ii) adding a set of carrier nucleic acid molecules to the polynucleotide sample to generate a first sample, wherein the set of carrier nucleic acid molecules comprises: (a) at least one subset of unmethylated carrier nucleic acid molecules; and / or (b) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, wherein the unmethylated carrier nucleic acid molecules do not comprise methylated nucleotides and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides; (iii) adding a set of carrier nucleic acid molecules to the polynucleotide sample to generate a first sample, wherein the set of carrier nucleic acid molecules comprises: (a) at least one subset of unmethylated carrier nucleic acid molecules; and / or (b) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, wherein the unmethylated carrier nucleic acid molecules do not comprise methylated nucleotides and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides; The method includes: (i) using a capture agent that selectively binds to nucleotides to divide into at least two divided sets, thereby generating divided samples; (iv) processing at least a portion of the divided samples to generate processed samples, wherein this processing comprises at least one of the following steps: (a) tagging, (b) amplifying, and (c) enriching molecules for specific regions of interest; (v) sequencing at least a portion of the processed samples to generate a set of sequencing reads; and (vi) analyzing at least a portion of the set of sequencing reads to detect the presence or absence of tumors. In some embodiments, the sequencing step comprises sequencing at least a portion of the processed samples from at least two divided sets.
[0006] In another aspect, the disclosure provides a method for analyzing polynucleotides, comprising the steps of: (i) obtaining a polynucleotide sample from a subject; (ii) adding a set of carrier nucleic acid molecules to the polynucleotide sample to generate a first sample, the set of carrier nucleic acid molecules comprising: (a) at least one subset of unmethylated carrier nucleic acid molecules; and / or (b) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, and wherein the unmethylated carrier nucleic acid molecules comprise no methylated nucleotides and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides; (iii) adding the first sample to a set of methylated polynucleotides; The method includes the following steps: (i) using a capture agent that selectively binds nucleotides to divide into at least two divided sets, thereby generating divided samples; (iv) processing at least a portion of the divided samples to generate processed samples, and this processing comprises at least one of the following steps: (a) tagging, (b) amplifying, and (c) enriching molecules for specific regions of interest; (v) sequencing at least a portion of the processed samples to generate a set of sequencing reads; and (vi) analyzing at least a portion of the set of sequencing reads to detect the presence or absence of tumor.In some embodiments, the sequencing step comprises sequencing at least a portion of the processed samples from at least two divided sets.
[0007] In another aspect, the disclosure provides a method for analyzing polynucleotides, comprising the steps of: (i) adding a set of carrier nucleic acid molecules to a polynucleotide sample from a subject to generate a first sample, wherein the set of carrier nucleic acid molecules comprises: (a) at least one subset of unmethylated carrier nucleic acid molecules; and / or (b) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, wherein the unmethylated carrier nucleic acid molecules do not comprise methylated nucleotides and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides; and (iii) adding a set of carrier nucleic acid molecules to a polynucleotide sample from a subject to selectively bind to methylated polynucleotides. (iv) processing at least a portion of the distributed samples to generate processed samples, the processing comprising at least one of the following steps: (a) tagging, (b) amplifying, and (c) enriching molecules for specific regions of interest; (v) sequencing at least a portion of the processed samples to generate a set of sequencing reads; and (vi) analyzing at least a portion of the set of sequencing reads to detect the presence or absence of tumors. In some embodiments, the sequencing step comprises sequencing at least a portion of the processed samples from the at least two distributed sets.
[0008] In another aspect, the disclosure provides a method for analyzing polynucleotides, comprising: (i) obtaining a polynucleotide sample from a subject; (ii) adding a set of carrier nucleic acid molecules to the polynucleotide sample to generate a first sample, the set of carrier nucleic acid molecules comprising: (a) at least one subset of unmethylated carrier nucleic acid molecules; and / or (b) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, and wherein the unmethylated carrier nucleic acid molecules comprise no methylated nucleotides and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides; and (iii) adding a set of carrier nucleic acid molecules to the polynucleotide sample to generate a first sample, the set of carrier nucleic acid molecules comprising: (a) at least one subset of unmethylated carrier nucleic acid molecules; and / or (b) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, and wherein the unmethylated carrier nucleic acid molecules comprise no methylated nucleotides and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides; (iv) processing at least a portion of the distributed samples to generate processed samples, the processing comprising at least one of the following steps: (a) tagging; and (b) amplifying polynucleotides; (v) sequencing at least a portion of the processed samples to generate a set of sequencing reads; and (vi) analyzing at least a portion of the set of sequencing reads to detect the presence or absence of tumors. In some embodiments, the sequencing step comprises sequencing at least a portion of the processed samples from the at least two distributed sets. In some embodiments, the processing step further comprises enriching polynucleotides for specific regions of interest.
[0009] In another aspect, the disclosure provides a method for detecting the methylation state of a polynucleotide, the method comprising the steps of: (i) obtaining a polynucleotide sample from a subject; (ii) adding a set of carrier nucleic acid molecules to the polynucleotide sample to generate a first sample, the set of carrier nucleic acid molecules comprising: (a) at least one subset of unmethylated carrier nucleic acid molecules; and / or (b) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, and wherein the unmethylated carrier nucleic acid molecules comprise no methylated nucleotides and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides; (iii) adding a set of carrier nucleic acid molecules to the polynucleotide sample to generate a first sample, the set of carrier nucleic acid molecules comprising: (a) at least one subset of unmethylated carrier nucleic acid molecules; and / or (b) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, and wherein the unmethylated carrier nucleic acid molecules comprise no methylated nucleotides and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides; The method includes: (i) using a capture agent that selectively binds to nucleotides to divide into at least two divided sets, thereby generating divided samples; (iv) processing at least a portion of the divided samples to generate processed samples, wherein this processing comprises at least one of the following steps: (a) tagging, (b) amplifying, and (c) enriching molecules for specific regions of interest; (v) sequencing at least a portion of the processed samples to generate a set of sequencing reads; and (vi) analyzing at least a portion of the set of sequencing reads to detect the presence or absence of tumors. In some embodiments, the sequencing step comprises sequencing at least a portion of the processed samples from at least two divided sets.
[0010] In some embodiments, analyzing at least a portion of the set of sequencing reads includes detecting one or more somatic variations. In some embodiments, analyzing at least a portion of the set of sequencing reads includes determining the methylation status of the polynucleotides (i.e., whether the polynucleotides are methylated or not) based on the number of CpG residues (or other methylated nucleotides) and the distributed sets into which the polynucleotides are distributed. In some embodiments, prior to sequencing, the polynucleotides in at least two distributed sets may be subjected to a chemical procedure that selectively converts nucleotide bases to provide information on nucleotide base modification (e.g., methylation). In these embodiments, the chemical procedure may be bisulfite treatment, TAB-Seq, ACE-Seq, EM-Seq, hmC-Seal, TAPS, or TAPSB. In these embodiments, analyzing at least a portion of the set of sequencing reads includes determining nucleotide base modification depending on nucleotide base conversion. In some embodiments, methods and systems used for dispensing can be found in PCT Patent Application No. PCT / US2020 / 053610, which is incorporated herein by reference in its entirety.
[0011] In another aspect, the disclosure provides a set of carrier nucleic acid molecules comprising: (i) at least one subset of unmethylated carrier nucleic acid molecules; and / or (ii) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, and wherein the unmethylated carrier nucleic acid molecules do not comprise methylated nucleotides and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides.
[0012] In another aspect, the disclosure provides a population of nucleic acids comprising: (i) a set of carrier nucleic acid molecules comprising: (a) at least one subset of unmethylated carrier nucleic acid molecules; and / or (b) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, and wherein the unmethylated carrier nucleic acid molecules do not comprise a methylated nucleotide and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides; and (ii) a polynucleotide sample obtained from a subject.
[0013] In another aspect, the disclosure provides a method, when performed by at least one electronic processor, for generating a first sample, the method comprising: (i) adding a set of carrier nucleic acid molecules to a polynucleotide sample to generate a first sample, the set of carrier nucleic acid molecules comprising: (a) at least one subset of unmethylated carrier nucleic acid molecules; and / or (b) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, the unmethylated carrier nucleic acid molecules do not comprise methylated nucleotides, and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides; (ii) partitioning the first sample into at least two partitioned sets using a capture agent that selectively binds methylated polynucleotides, (iii) processing at least a portion of the distributed sample to generate a processed sample, the processing comprising at least one of the following: (a) tagging, (b) amplifying, and (c) enriching molecules for a specific region of interest; (iv) sequencing at least a portion of the processed sample to generate a set of sequencing reads; and (v) analyzing at least a portion of the set of sequencing reads to detect the presence or absence of a tumor.
[0014] In some embodiments, the carrier nucleic acid molecule is between 25 bp and 325 bp in length. In some embodiments, the first and second subsets of at least one subset of unmethylated carrier nucleic acid molecules comprise the same nucleotide sequence. In some embodiments, the first and second subsets of at least one subset of unmethylated carrier nucleic acid molecules comprise different nucleotide sequences.
[0015] In some embodiments, the first and second subsets of at least one subset of unmethylated carrier nucleic acid molecules contain one or more CpG dinucleotides in their nucleotide sequences. In some embodiments, the positions of one or more CpG dinucleotides in the first subset of at least one subset of unmethylated carrier nucleic acid molecules are different from the positions of one or more CpG dinucleotides in the second subset. In some embodiments, the number of CpG dinucleotides in the first subset of at least one subset of unmethylated carrier nucleic acid molecules is different from the number of CpG dinucleotides in the second subset. In some embodiments, the sequence of nucleotides adjacent to one or more CpG dinucleotides in the first subset of at least one subset of unmethylated carrier nucleic acid molecules is different from the sequence of nucleotides adjacent to one or more CpG dinucleotides in the second subset. In some embodiments, the sequence of nucleotides adjacent to one or more methylated nucleotides in the first subset and / or the second subset can be a sequence of 1, 2, 3, 4, or 5 nucleotides adjacent to one or more methylated nucleotides. In some embodiments, the first subset and the second subset of at least one subset of unmethylated carrier nucleic acid molecules are of different lengths.
[0016] In some embodiments, the one or more methylated nucleotides are selected from the group consisting of: (i) 5-methylcytosine, (ii) 6-methyladenine, (iii) hydroxymethylcytosine, (iv) methyluracil, and (v) any other methylated nucleotide. In some embodiments, the number of methylated nucleotides is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 15, 16, 17, 18, 19, or at least 20.
[0017] In some embodiments, the first and second subsets of at least one subset of methylated carrier nucleic acid molecules comprise the same nucleotide sequence. In some embodiments, the first and second subsets of at least one subset of methylated carrier nucleic acid molecules comprise different nucleotide sequences.
[0018] In some embodiments, the first and second subsets of at least one subset of methylated carrier nucleic acid molecules comprise one or more CpG dinucleotides in the nucleotide sequence. In some embodiments, the one or more CpG dinucleotides comprise one or more methylated cytosines.
[0019] In some embodiments, the positions of one or more methylated nucleotides in a first subset of at least one subset of methylated carrier nucleic acid molecules are different from the positions of one or more methylated nucleotides in a second subset. In some embodiments, the number of methylated nucleotides in the first subset of at least one subset of methylated carrier nucleic acid molecules is different from the number of methylated nucleotides in the second subset. In some embodiments, the sequence of nucleotides adjacent to one or more methylated nucleotides in the first subset of at least one subset of methylated carrier nucleic acid molecules is different from the sequence of nucleotides adjacent to one or more methylated nucleotides in the second subset. In some embodiments, the sequence of nucleotides adjacent to one or more methylated nucleotides in the first subset and / or the second subset can be a sequence of 1, 2, 3, 4, or 5 nucleotides adjacent to one or more methylated nucleotides.
[0020] In some embodiments, the first subset and the second subset of at least one subset of methylated carrier nucleic acid molecules are of different lengths.
[0021] In some embodiments, the amount of at least one subset of methylated carrier nucleic acid molecules relative to at least one subset of unmethylated carrier nucleic acid molecules is about 0:1, 0.1:99.9, 0.5:99.5, 0.75:99.25, 1:99, 1:95, 1:90, 1:80, 1:75, 1:70, 1:60, 1:50, 1:40, 1:30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90, 1:90, 1:90, 1:90, 1:1 ... 1:25, 1:20, 1:10, 1:5, 1:2, 1:1.15, 1:1, 1.15:1, 2:1, 5:1, 10:1, 20:1, 25:1, 30:1, 40:1, 50:1, 60:1, 70:1, 75:1, 80:1, 90:1, 95:1, 99:1, 99.25:0.75, 99.5:0.5, 99.9:0.1, or 1:0. In some embodiments, the amount of at least one subset of methylated carrier nucleic acid molecules to at least one subset of unmethylated carrier nucleic acid molecules is at a ratio of about 1:1. In some embodiments, the amount of at least one subset of methylated carrier nucleic acid molecules to at least one subset of unmethylated carrier nucleic acid molecules is at a ratio of about 2:1. In some embodiments, the amount of at least one subset of methylated carrier nucleic acid molecules relative to at least one subset of unmethylated carrier nucleic acid molecules is in a ratio of about 1:2. In some embodiments, the amount of polynucleotide sample relative to the set of carrier nucleic acid molecules is in a ratio of about 1:0.1; 1:0.2, 1:0.3, 1:4, 1:0.5, 1:6, 1:7, 1:8, 1:0.9, 1:1, 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:20, 1:30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90, 1:10 ... 1:50, 1:60, 1:70, 1:80, 1:90, 1:100, 1:200, 1:300; 1:400, 1:500, 1:600, 1:700, 1:800, 1:900, 1:1000, 1:5000, 1:10,000, 1:100,000, 1:500,000, 1:106, 1:107, 1:108, or 1:109. In some embodiments, the amount is in mass. In some embodiments, the amount is in molar concentration. In some embodiments, the polynucleotide sample is at most 1 μg. In some embodiments, the polynucleotide sample is at most 200 ng. In some embodiments, the polynucleotide sample is at most 150 ng.In some embodiments, the polynucleotide sample is up to 100 ng.
[0022] In some embodiments, the set of carrier nucleic acid molecules is added in an amount sufficient to provide a total amount of about 175 ng, 200 ng, 225 ng, 250 ng, 275 ng, 300 ng, 350 ng, 400 ng, 450 ng, 500 ng, 600 ng, 700 ng, 750 ng, 800 ng, 900 ng, 1 μg, 1.1 μg, 1.25 μg, or 1.5 μg of polynucleotide sample and set of carrier nucleic acid molecules.
[0023] In some embodiments, the sequence of the carrier nucleic acid molecule is selected from the group consisting of (i) a sequence derived from a viral genome, (ii) a sequence derived from a bacterial genome, (iii) a sequence derived from a lambda genome, and (iv) a sequence derived from a non-human genome. In some embodiments, the carrier nucleic acid molecule is synthetic DNA. In some embodiments, the carrier nucleic acid molecule comprises a uracil nucleoside. In some embodiments, the method further comprises adding uracil deglycosylase and DNA glycosylase-lyase prior to amplification (Kropachev KT et al., Biochemistry (2006); 45(39):12039-12049). In some embodiments, the carrier DNA molecule may be generated by PCR. In some embodiments, the carrier nucleic acid molecule generated by PCR may be further modified by treatment with a methyltransferase to incorporate a methyl group into one or more nucleotides in the carrier nucleic acid molecule. In some embodiments, the carrier nucleic acid molecule may be end-labeled with a polymerase to incorporate a modified nucleoside to prevent ligation of the carrier nucleic acid molecule to an adapter. In some embodiments, the carrier nucleic acid molecule comprises a non-naturally occurring nucleoside derivative. In some embodiments, the carrier nucleic acid molecule may be labeled with biotin or a fluorophore.
[0024] In some embodiments, at least one end of the carrier nucleic acid molecule comprises a C3 (propyl group) spacer. The C3 spacer aids in blocking ligation. In some embodiments, at least one end of the carrier nucleic acid molecule comprises a dideoxynucleotide. In some embodiments, at least one end of the carrier nucleic acid molecule comprises any chemical modification that prevents a hydroxyl group from acting as a nucleophile. In some embodiments, the 5' end of the carrier nucleic acid molecule comprises at least one of the following modifications: (i) an inverted (5'-5')-dideoxythymine, dideoxycytosine, dideoxyguanine, or dideoxyadenine; (ii) a propyl group; or (iii) another organic functional group, such as, but not limited to, benzyl, ethyl, or methyl. In some embodiments, the 3' end of the carrier nucleic acid molecule comprises at least one of the following modifications: (i) any dideoxy base that can be added enzymatically or during synthesis, such as dideoxythymine, dideoxycytosine, dideoxyguanine, or dideoxyadenine; (ii) a propyl group, or (iii) other organic functional groups, such as, but not limited to, benzyl, ethyl, or methyl.
[0025] In some embodiments, the polynucleotide sample is obtained from tissue, blood, plasma, serum, urine, saliva, feces, cerebrospinal fluid, buccal swab, or thoracocentesis. In some embodiments, the polynucleotide sample is obtained from tissue. In some embodiments, the polynucleotide sample obtained from tissue is fragmented by enzymatic or mechanical means. In some embodiments, the polynucleotide sample is obtained from blood. In some embodiments, the blood-derived polynucleotide sample is a cell-free DNA sample. In some embodiments, the polynucleotide sample is a cell-free DNA sample. In some embodiments, the carrier nucleic acid molecule is a double-stranded molecule. In some embodiments, the method disclosed herein does not have a denaturing step before adapter ligation. That is, the first sample is not subjected to a denaturing step (in other words, the double-stranded carrier nucleic acid molecule is subjected to ligation).
[0026] In some embodiments, the results of the system and / or method disclosed herein are used as input to generate a report. The report can be in paper or electronic format. For example, the information about the distribution of nucleic acid molecules determined by the method or system disclosed herein and / or the information derived from the distribution of nucleic acid molecules can be displayed in such a report. The method or system disclosed herein can further include a step of transmitting the report to a third party, such as the subject who obtained the sample or a medical professional.
[0027] The various steps of the methods disclosed herein or steps performed by the systems disclosed herein may be performed at the same or different times and / or in the same or different geographic locations, e.g., countries. The various steps of the methods disclosed herein may be performed by the same or different people.
[0028] In other embodiments, the invention includes kits for practicing the subject methods, comprising: (a) a set of carrier nucleic acid molecules comprising (i) at least one subset of unmethylated carrier nucleic acid molecules; and / or (ii) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, and wherein the unmethylated carrier nucleic acid molecules contain no methylated nucleotides and the methylated carrier nucleic acid molecules contain one or more methylated nucleotides; and (b) a capture agent that selectively binds to methylated polynucleotides.
[0029] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
[0030] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain embodiments and, together with the written description, serve to explain certain principles of the methods, computer-readable media, and systems disclosed herein. The description provided herein will be better understood when read in conjunction with the accompanying drawings, which are included by way of example and not by way of limitation. It will be understood that like reference numerals identify like components throughout the drawings, unless the context dictates otherwise. It will also be understood that some or all of the drawings are schematic views for illustrative purposes and do not necessarily depict the actual relative dimensions or positions of the elements shown. [Brief explanation of the drawings]
[0031] [Figure 1] FIG. 1 is a flow chart diagram of a method for detecting the presence or absence of a tumor in a subject according to one embodiment of the present disclosure.
[0032] [Figure 2] FIG. 2 is a schematic diagram of a carrier nucleic acid molecule suitable for use in some embodiments of the present disclosure.
[0033] [Figure 3] FIG. 3 is a schematic diagram of a set of carrier nucleic acid molecules suitable for use in some embodiments of the present disclosure.
[0034] [Figure 4] FIG. 4 is a schematic diagram of a set of carrier nucleic acid molecules suitable for use in some embodiments of the present disclosure.
[0035] [Figure 5] FIG. 5 is a schematic diagram of a set of carrier nucleic acid molecules suitable for use in some embodiments of the present disclosure.
[0036] [Figure 6] FIG. 6 is a schematic diagram of a set of carrier nucleic acid molecules suitable for use in some embodiments of the present disclosure.
[0037] [Figure 7] FIG. 7 is a schematic diagram of an example system suitable for use in some embodiments of the present disclosure.
[0038] [Figure 8] 8A and 8B are graphical representations of cell-free DNA molecules from a sample in the presence and absence of carrier DNA molecules in a hyper-partitioned set. DETAILED DESCRIPTION OF THE INVENTION
[0039] definition In order that this disclosure may be more readily understood, certain terms are first defined below. Additional definitions for these and other terms may be set forth throughout this specification. In the event that a definition of a term set forth below conflicts with a definition in an application or patent incorporated by reference, the definition set forth in this application should be used to understand the meaning of the term.
[0040] As used in this specification and the appended claims, the singular forms "a," "an," and "the" include the plural forms unless the context clearly dictates otherwise. Thus, for example, reference to "a method" includes one or more methods, and / or steps, etc., of the type described herein and / or that will become apparent to those skilled in the art upon reading this disclosure.
[0041] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting. Moreover, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. In describing and claiming the methods, computer-readable media, and systems, the following terminology and grammatical variants thereof will be used in accordance with the definitions set forth below.
[0042] About: As used herein, "about" or "approximately," when applied to one or more values or elements of interest, refers to a value or element that is similar to a stated reference value or element. In certain embodiments, the term "about" or "approximately" refers to a range of values or elements that are within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or a smaller percentage in either direction (greater or less) of the stated reference value or element, unless otherwise stated or apparent from the context (except where such number would exceed 100% of the possible values or elements).
[0043] Adapter: As used herein, "adapter" refers to a short nucleic acid (e.g., less than about 500 nucleotides, less than about 100 nucleotides, or less than about 50 nucleotides in length) that is typically at least partially double-stranded and attached to either or both ends of a given sample nucleic acid molecule. The adapter may include nucleic acid primer binding sites at both ends that allow amplification of the nucleic acid molecule flanked by the adapter, and / or sequencing primer binding sites, including primer binding sites for sequencing applications, such as various next-generation sequencing (NGS) applications. The adapter may also include binding sites for capture probes, such as oligonucleotides attached to flow cell supports, etc. The adapter may also include a nucleic acid tag, as described herein. The nucleic acid tag is typically positioned relative to the amplification primer and sequencing primer binding sites so that the nucleic acid tag is included in the amplicon and sequence read of a given nucleic acid molecule. Adapters of the same or different sequences can be attached to each end of a nucleic acid molecule. In some embodiments, adapters of the same sequence, except for the different nucleic acid tags, are attached to each end of a nucleic acid molecule. In some embodiments, the adapter is a Y-shaped adapter, with one end blunt or tailed as described herein for attachment to a nucleic acid molecule that is also blunt or tailed with one or more complementary nucleotides. In yet other exemplary embodiments, the adapter is a bell-shaped adapter that includes a blunt or tailed end for attachment to a nucleic acid molecule to be analyzed. Other examples of adapters include T-tailed and C-tailed adapters.
[0044] Amplify: As used herein, in the context of nucleic acids, "amplify" or "amplification" refers to the production of multiple copies of a polynucleotide, or portion of a polynucleotide, typically starting from a small amount of the polynucleotide (e.g., a single polynucleotide molecule), where the amplification product or amplicon is generally detectable. Amplification of polynucleotides encompasses a variety of chemical and enzymatic processes. Amplification includes, but is not limited to, polymerase chain reaction (PCR).
[0045] Barcode: As used herein, in the context of nucleic acids, "barcode" or "molecular barcode" refers to a nucleic acid molecule that contains a sequence that can function as a molecular identifier. A barcode or molecular barcode is a type of nucleic acid tag. For example, individual "barcode" sequences are typically added to DNA fragments during next-generation sequencing (NGS) library preparation so that sequencing reads can be identified and sorted prior to final data analysis.
[0046] Cancer type: As used herein, "cancer type" refers to the type or subtype of cancer, as defined, for example, by histopathology. Cancer type refers to the presence in a given tissue (e.g., blood cancer, central nervous system (CNS), brain cancer, lung cancer (small cell and non-small cell), skin cancer, nose cancer, throat cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, bowel cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, breast cancer, prostate cancer, ovarian cancer, lung cancer, intestinal cancer, soft tissue cancer, neuroendocrine cancer, gastroesophageal cancer, head and neck cancer, gynecological cancer, colorectal cancer, urinary tract cancer, etc.). Cancers may be defined by any conventional criteria, such as based on their type (epithelial cancer, solid cancer, heterogeneous cancer, homogeneous cancer), origin of unknown primary origin, etc., and / or based on cancers of the same cellular lineage (e.g., carcinoma, sarcoma, lymphoma, cholangiocarcinoma, leukemia, mesothelioma, melanoma, or glioblastoma) and / or exhibiting cancer markers such as, but not limited to, Her2, CA15-3, CA19-9, CA-125, CEA, AFP, PSA, HCG, hormone receptors, and NMP-22. Cancers may also be classified by stage (e.g., stage 1, 2, 3, or 4) and whether they are of primary or secondary origin.
[0047] Carrier nucleic acid molecule: As used herein, "carrier nucleic acid molecule" refers to a set of nucleic acid molecules that can be added to a polynucleotide sample obtained from a subject to improve detection of the presence or absence of a tumor in the subject. In some embodiments, the carrier nucleic acid molecule can be single-stranded or double-stranded. In some embodiments, the carrier nucleic acid molecule can be DNA or RNA. In some embodiments, the carrier nucleic acid molecule can be a synthetic oligonucleotide. In some embodiments, the carrier nucleic acid molecule can be generated by PCR via amplifying one or more specific regions of interest from a genome. In some embodiments, the carrier nucleic acid molecule generated by PCR can be further modified by treatment with a methyltransferase to incorporate a methyl group into one or more nucleotides in the carrier nucleic acid molecule. In some embodiments, the carrier nucleic acid molecule can be end-labeled with a polymerase to incorporate a modified nucleoside that prevents ligation of the carrier nucleic acid molecule to an adapter. In some embodiments, the carrier nucleic acid molecule can have a non-naturally occurring nucleic acid sequence. In some embodiments, the carrier nucleic acid molecule can have a naturally occurring nucleic acid sequence. In some embodiments, the carrier nucleic acid molecule can have a nucleic acid sequence that corresponds to a non-human genome. As non-limiting examples, these carrier nucleic acid molecules may have (i) a sequence corresponding to a region of lambda phage DNA or the human genome, (ii) a sequence corresponding to a region of a viral genome, (iii) a sequence corresponding to a region of a bacterial genome, (iv) a non-naturally occurring sequence, and / or (v) any combination of the above. In some embodiments, the carrier nucleic acid molecule comprises a non-naturally occurring nucleoside derivative. In some embodiments, the carrier DNA molecule may be generated by PCR. In some embodiments, the carrier nucleic acid molecule generated by PCR may be further modified by treatment with a methyltransferase to incorporate a methyl group into one or more nucleotides in the carrier nucleic acid molecule. In some embodiments, the carrier nucleic acid molecule may be end-labeled with a polymerase to incorporate a modified nucleoside that prevents ligation of the carrier nucleic acid molecule to an adapter. In some embodiments, the carrier nucleic acid molecule comprises a non-naturally occurring nucleoside derivative.In some embodiments, the carrier nucleic acid molecule may be labeled with biotin or a fluorophore.
[0048] Cell-free nucleic acid: As used herein, "cell-free nucleic acid" refers to nucleic acid that is not contained within or otherwise associated with cells, or, in some embodiments, nucleic acid that remains in a sample after removal of intact cells. Cell-free nucleic acid can include, for example, all unencapsulated nucleic acids provided by a subject's bodily fluids (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.). Cell-free nucleic acid includes DNA (cfDNA), RNA (cfRNA), and hybrids thereof, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acid can be double-stranded, single-stranded, or a hybrid thereof. Cell-free nucleic acid can be released into bodily fluids through secretion or cell death processes, such as cell necrosis, apoptosis, etc. Some cell-free nucleic acids, such as circulating tumor DNA (ctDNA), are released from cancer cells into body fluids. Others are released from healthy cells. ctDNA can be fragmented DNA derived from unencapsulated tumors. Cell-free nucleic acids can have one or more epigenetic modifications, for example, cell-free nucleic acids can be acetylated, 5-methylated, and / or hydroxymethylated.
[0049] Cellular nucleic acid: As used herein, "cellular nucleic acid" refers to nucleic acid that is present within the cell or cells from which the nucleic acid originates, at least at the time the sample is taken or collected from a subject, even if such nucleic acid is subsequently removed (e.g., via cell lysis) as part of a given analytical process.
[0050] Coverage: As used herein, the terms "coverage," "total molecule count," or "total allele count" are used interchangeably. They refer to the total number of DNA molecules at a particular genomic location in a given sample.
[0051] Deoxyribonucleic acid or ribonucleic acid: As used herein, "deoxyribonucleic acid" or "DNA" refers to natural or modified nucleotides that have a hydrogen group at the 2'-position of the sugar moiety. DNA typically comprises a chain of nucleotides containing four types of nucleotide bases: adenine (A), thymine (T), cytosine (C), and guanine (G). As used herein, "ribonucleic acid" or "RNA" refers to natural or modified nucleotides that have a hydroxyl group at the 2'-position of the sugar moiety. RNA typically comprises a chain of nucleotides containing four types of nucleotide bases: A, uracil (U), G, and C. As used herein, the term "nucleotide" refers to natural or modified nucleotides. Certain pairs of nucleotides specifically bind to each other in a complementary manner (called complementary base pairing). In DNA, adenine (A) pairs with thymine (T), and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U) and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand composed of nucleotides complementary to those in the first strand, the two strands combine to form a duplex. As used herein, "nucleic acid sequencing data," "nucleic acid sequencing information," "sequence information," "nucleic acid sequence," "nucleotide sequence," "genomic sequence," "gene sequence," or "fragment sequence" or "nucleic acid sequencing read" refers to any information or data that indicates the order and / or identity of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a molecule (e.g., a whole genome, whole transcriptome, exome, oligonucleotide, polynucleotide, or fragment) of a nucleic acid such as DNA or RNA.It should be understood that the present teachings contemplate sequence information obtained using all of the various available techniques, platforms, or technologies, including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion- or pH-based detection systems, and electronic signature-based systems.
[0052] Mutation: As used herein, "mutation" refers to a variation from a known reference sequence, including, for example, single nucleotide variants (SNVs) and mutations such as insertions or deletions (indels). Mutations can be germline mutations or somatic mutations. In some embodiments, the reference sequence used for comparison is the wild-type genomic sequence of the species from which the test sample is provided, typically the human genome.
[0053] Mutation caller: As used herein, "mutation caller" refers to an algorithm (typically embodied in software or otherwise computer-implemented) used to identify mutations in test sample data (e.g., sequence information obtained from a subject).
[0054] Neoplasm: As used herein, the terms "neoplasm" and "tumor" are used interchangeably. They refer to an abnormal growth of cells in a subject. A neoplasm or tumor can be benign, potentially malignant, or malignant. A malignant tumor is called a cancer or cancerous tumor.
[0055] Next-generation sequencing: As used herein, " next-generation sequencing " or " NGS " refers to a sequencing technology that has increased throughput compared to traditional Sanger-based approaches and capillary electrophoresis-based approaches, for example, has the ability to generate hundreds of thousands of relatively small sequence reads at once.Some examples of next-generation sequencing methods include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization.In some embodiments, next-generation sequencing involves the use of equipment that can sequence single molecules.
[0056] Nucleic acid tag: As used herein, "nucleic acid tag" refers to a short nucleic acid (e.g., less than about 500, 100, 50, or 10 nucleotides in length) used to distinguish nucleic acids from different samples (e.g., representing a sample index) or to distinguish different nucleic acid molecules in the same sample of different types or subjected to different treatments (e.g., representing a molecular barcode). Nucleic acid tags comprise predetermined, fixed, non-random, random, or semi-random oligonucleotide sequences. Such nucleic acid tags can be used to label different nucleic acid molecules or different nucleic acid samples or subsamples. Nucleic acid tags can be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags can have the same length or varying lengths, as appropriate. Nucleic acid tags can also comprise double-stranded molecules with one or more blunt ends, 5' or 3' single-stranded regions (e.g., overhangs), and / or one or more other single-stranded regions elsewhere within a given molecule. Nucleic acid tags can be attached to one or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Nucleic acid tags can be decoded to reveal information, such as the sample origin, form, or processing of a given nucleic acid. For example, nucleic acid tags can be used to enable pooling and / or parallel processing of multiple samples containing nucleic acids with different molecular barcodes and / or sample indices, where the nucleic acids are then deconvoluted by detecting (e.g., reading) the nucleic acid tags. Nucleic acid tags can also be referred to as identifiers (e.g., molecular identifiers, sample identifiers). Additionally or alternatively, nucleic acid tags can be used as molecular identifiers (e.g., to distinguish between different molecules or amplicons of different parent molecules in the same sample or subsample). This includes, for example, uniquely tagging different nucleic acid molecules in a given sample or non-uniquely tagging such molecules.For non-unique tagging applications, a limited number of tags (i.e., molecular barcodes) may be used to tag each nucleic acid molecule so that different molecules, in combination with at least one molecular barcode, can be distinguished based on their intrinsic sequence information (e.g., start and / or end positions, subsequences at one or both ends of the sequence, and / or sequence length when they are mapped to a selected reference genome). Typically, a sufficient number of different molecular barcodes are used so that the probability that any two molecules will have the same intrinsic sequence information (e.g., start and / or end positions, subsequences at one or both ends of the sequence, and / or length) and also have the same molecular barcode is low (e.g., less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% chance).
[0057] Partitioning: As used herein, "partitioning" and "epigenetic partitioning" are used interchangeably. This refers to separating or fractionating nucleic acid molecules based on their characteristics (e.g., the level / degree of epigenetic modification). Partitioning can be a physical partitioning of molecules. Partitioning can involve separating nucleic acid molecules into groups or sets based on the level of epigenetic modification (e.g., methylation). For example, nucleic acid molecules can be partitioned based on the level of methylation of the nucleic acid molecules. In some embodiments, methods and systems used for partitioning can be found in PCT Patent Application No. PCT / US2017 / 068329, which is incorporated herein by reference in its entirety. Methods, systems, and compositions can also be found in PCT Patent Application Nos. PCT / US2019 / 059217 and PCT / US2020 / 016120, each of which is incorporated herein by reference in its entirety.
[0058] Distributed set: As used herein, "distributed set" refers to a set of nucleic acid molecules distributed into sets / groups based on the differential binding affinity of the nucleic acid molecules to a binder. The binder preferentially binds to nucleic acid molecules containing nucleotides with epigenetic modifications. For example, if the epigenetic modification is methylation, the binder can be a methyl-binding domain (MBD) protein. In some embodiments, the distributed set can include nucleic acid molecules belonging to a specific level / degree of epigenetic modification. For example, nucleic acid molecules can be distributed into three sets: one set of highly methylated nucleic acid molecules (or hypermethylated nucleic acid molecules), which can be referred to as the hypermethylated distributed set or highly distributed set; another set of lowly methylated nucleic acid molecules (or hypomethylated nucleic acid molecules), which can be referred to as the hypomethylated distributed set or lowly distributed set; and a third set of moderately methylated nucleic acid molecules, which can be referred to as the moderately methylated distributed set or moderately distributed set. In another example, the nucleic acid molecules may be distributed based on the number of methylated nucleotides - one distributed set may have nucleic acid molecules with 9 methylated nucleotides, and another distributed set may have unmethylated nucleic acid molecules (i.e., 0 methylated nucleotides).
[0059] Polynucleotide: As used herein, "polynucleotide," "nucleic acid," "nucleic acid molecule," or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or their analogs) joined by internucleoside linkages. Typically, a polynucleotide contains at least three nucleosides. Oligonucleotides often range in size from a few monomeric units, e.g., 3-4 monomeric units, to several hundred monomeric units. When a polynucleotide is designated by a sequence of letters, such as "ATGCCTG," it is understood that the nucleotides are in 5'→3' order from left to right, and that, in the case of DNA, "A" denotes deoxyadenosine, "C" denotes deoxycytidine, "G" denotes deoxyguanosine, and "T" denotes deoxythymidine, unless otherwise specified. The letters A, C, G, and T may be used, as standard in the art, to refer to the base itself, a nucleoside, or a nucleotide that comprises the base.
[0060] Reference sequence: As used herein, "reference sequence" refers to a known sequence used for comparison with experimentally determined sequences. For example, the known sequence can be the entire genome, a chromosome, or any segment thereof. The reference typically comprises at least about 20, at least about 50, at least about 100, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 1000, or more than 1000 nucleotides. In some embodiments, the reference sequence can be the human genome. The reference sequence can be aligned with a single continuous sequence of a genome or chromosome, or can include non-contiguous segments that align with different regions of a genome or chromosome. Examples of reference sequences include, for example, the human genome, for example, hG19 and hG38.
[0061] Sample: As used herein, "sample" means anything that can be analyzed by the methods and / or systems disclosed herein.
[0062] Sequencing: As used herein, "sequencing" refers to any of several techniques used to determine the sequence (e.g., the identity and / or order of monomeric units) of a biomolecule, e.g., a nucleic acid such as DNA or RNA. Examples of sequencing methods include targeted sequencing, single-molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscope-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxytermination sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, gel electrophoresis, duplex sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, and the like. These include, but are not limited to, sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification-PCR at lower denaturation temperatures (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short-read sequencing, single-molecule sequencing, sequencing by synthesis, real-time sequencing, reverse-terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, sequencing can be performed by a genetic analyzer, such as a commercially available genetic analyzer from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, among many others.
[0063] Sequence information: As used herein, "sequence information," in the context of a nucleic acid molecule, means the order and / or identity of the monomer units (e.g., nucleotides, etc.) in that molecule, and may also include the start and end genomic coordinates of the nucleic acid molecule mapping to a reference sequence.
[0064] Somatic mutation: As used herein, the terms "somatic mutation" or "somatic variation" are used interchangeably. They refer to mutations in the genome that occur after conception. Somatic mutations can occur in any cell of the body except germ cells, and therefore are not inherited by offspring.
[0065] Subject: As used herein, "subject" refers to an animal, e.g., a mammalian species (e.g., a human) or an avian (e.g., an avian) species, or other organism, e.g., a plant. More specifically, a subject can be a vertebrate, e.g., a mammal, e.g., a mouse, a primate, a monkey, or a human. Animals include livestock (e.g., production cattle, dairy cows, poultry, horses, pigs, etc.), sport animals, and companion animals (e.g., pets or support animals). A subject can be a healthy individual, an individual having or suspected of having a disease or a predisposition to the disease, or an individual in need of treatment or suspected of needing treatment. The terms "individual" or "patient" are intended interchangeably with "subject."
[0066] For example, the subject may be an individual who has been diagnosed with cancer, an individual who is going to receive cancer treatment, and / or an individual who has received at least one cancer treatment.The subject may be in cancer remission.As another example, the subject may be an individual who has been diagnosed with an autoimmune disease.As another example, the subject may be a female individual who is pregnant or planning to become pregnant, who may be diagnosed with or suspected of having a disease, such as cancer, an autoimmune disease.
[0067] Detailed Description I. Overview Methods based on genomic / epigenetic partitioning can enable simultaneous signal detection of multiple analytes in one assay. However, the detected signals of partition-based analytes may have low resolution and are subject to variable assay conditions that alter the sensitivity and specificity of the signals. It is desirable to increase the sensitivity of liquid biopsy assays while reducing the loss of cell-free nucleic acid (original material) or data in the process. Controlling for assay variability by using one or more controls as described herein can also provide the ability to compare results across different experiments. Also desirable.
[0068] The present disclosure provides methods, compositions, and systems for analyzing polynucleotides in a partitioning assay. The present invention involves the use of a set of carrier nucleic acid molecules. In some embodiments, the use of carrier nucleic acid molecules can increase the specificity of partitioning of methylated nucleic acid molecules by preventing nonspecific binding of unmethylated nucleic acid molecules to capture agents that selectively bind to methylated nucleic acid molecules.
[0069] Nucleic acid molecules, such as cell-free polynucleotides, can differ based on epigenetic characteristics such as methylation. Nucleic acids can have different nucleotide sequences, such as specific genes or loci. Characteristics can differ in degree. For example, DNA molecules can differ in the degree of epigenetic modification. The degree of modification can refer to the number of modification events that a molecule has undergone, such as the number of methylation groups (degree of methylation) or other epigenetic changes. For example, DNA can be hypomethylated or hypermethylated.
[0070] A feature of a nucleic acid molecule can be a modification, which can include various chemical modifications (i.e., epigenetic modifications). Non-limiting examples of chemical modifications can include, but are not limited to, covalent DNA modifications, including DNA methylation. In some embodiments, DNA methylation involves the addition of a methyl group to a cytosine at a CpG site (a cytosine-phosphate-guanine site (i.e., a cytosine followed by a guanine in the 5'→3' direction of a nucleic acid sequence)). In some embodiments, DNA methylation involves the addition of a methyl group to an adenine, e.g., N 6 In some embodiments, DNA methylation is 5-methylation (modification of the carbon at the 5 position in the 6-membered ring of cytosine). In some embodiments, 5-methylation involves the addition of a methyl group to the 5C position of cytosine, generating 5-methylcytosine (m5c). In some embodiments, methylation involves derivatives of m5c. Derivatives of m5c include, but are not limited to, 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine (5-fC), and 5-caryboxylcytosine (5-caC). In some embodiments, DNA methylation is 3C methylation (modification of the carbon at the 3 position in the 6-membered ring of cytosine). In some embodiments, 3C methylation involves the addition of a methyl group to the 3C position of cytosine, generating 3-methylcytosine (3mC). Methylation can also occur at non-CpG sites, for example, methylation can occur at CpA, CpT, or CpC sites. DNA methylation can alter the activity of methylated DNA regions. For example, if DNA in a promoter region is methylated, gene transcription can be repressed. DNA methylation is important for normal development, and abnormalities in methylation can disrupt epigenetic regulation. Disruptions in epigenetic regulation, such as repression, can cause diseases such as cancer. Promoter methylation in DNA can indicate cancer.
[0071] A CpG dyad is a dinucleotide CpG (cytosine-phosphate-guanine, i.e., a cytosine followed by a guanine in the 5' to 3' direction of the nucleic acid sequence) on the sense strand of a double-stranded DNA molecule and its complementary CpG on the antisense strand. CpG dyads can be either fully methylated or hemimethylated (only one strand is methylated).
[0072] CpG dinucleotides are underrepresented in the normal human genome, and the majority of CpG dinucleotide sequences are transcriptionally inactive (e.g., DNA heterochromatin regions in pericentromeric portions of chromosomes and in repetitive elements) and methylated. However, many CpG islands, especially those around transcription start sites (TSSs), are protected from such methylation.
[0073] Cancer can be indicated by epigenetic variations, such as methylation.Examples of methylation changes in cancer include the local increase in DNA methylation in CpG islands at the TSS of genes involved in normal growth control, DNA repair, cell cycle regulation and / or cell differentiation.This hypermethylation can be associated with the abnormal loss of transcriptional capacity of the involved genes, and occurs at least as frequently as point mutations and deletions as the cause of altered gene expression.DNA methylation profiling can be used to detect regions with different degrees of methylation ("differentially methylated regions" or "DMRs") in the genome that are altered during development or perturbed by disease, such as cancer or any cancer-related disease.
[0074] Methylation profiling can involve determining methylation patterns across different regions of a genome. For example, after dividing molecules based on the degree of methylation (e.g., the relative number of methylated nucleotides per molecule) and sequencing, the sequences of molecules in different fractions can be mapped to a reference genome. This can indicate regions of the genome that are more highly methylated or less highly methylated compared to other regions. In this way, genomic regions can have different degrees of methylation, as opposed to individual molecules. In addition to methylation, other epigenetic modifications can be profiled as well.
[0075] Nucleic acid molecules in a sample can be fractionated or distributed based on one or more characteristics. Distribution of nucleic acid molecules in a sample can increase rare signals. For example, genetic variations present in hypermethylated DNA but less abundant (or absent) in hypomethylated DNA can be more easily detected by partitioning the sample into hypermethylated and hypomethylated nucleic acid molecules. By analyzing multiple fractions of a sample, multidimensional analysis of single molecules can be performed, thus achieving greater sensitivity. Distribution can include physically distributing nucleic acid molecules into subsets or groups based on the presence or absence of genomic features. Fractionation can include physically distributing nucleic acid molecules into fraction groups based on the degree to which genomic features, such as epigenetic modifications, are present. A sample can be fractionated or distributed into one or more fraction groups based on differential gene expression or features indicative of a disease state. Samples can be fractionated based on characteristics or combinations thereof that provide differences in signal between normal and diseased states during analysis of nucleic acids, e.g., cell-free DNA ("cfDNA"), non-cfDNA, tumor DNA, circulating tumor DNA ("ctDNA"), and cell-free nucleic acid ("cfNA").
[0076] The present disclosure provides methods, compositions, and systems for analyzing polynucleotides using a partitioning assay to detect the presence or absence of tumors. These methods may include adding a carrier nucleic acid molecule to a polynucleotide sample obtained from a subject. In some embodiments, the use of a carrier nucleic acid molecule helps improve partitioning of polynucleotides using a methyl-binding protein. Methyl-binding domain proteins have low affinity for unmethylated molecules. Therefore, when incubated with DNA obtained from a subject, molecules that do not contain methylated cytosines may be unintentionally captured. This reduces the ability to efficiently separate methylated molecules from unmethylated molecules. Therefore, by including a carrier nucleic acid molecule in the assay, the carrier nucleic acid molecule binds to the methyl-binding protein, preventing unmethylated molecules from the polynucleotide sample from binding to the methyl-binding protein, thereby improving the specificity of partitioning of methylated molecules. In some embodiments, the carrier nucleic acid molecule improves library preparation of molecules for next-generation sequencing.
[0077] In some embodiments, carrier nucleic acid molecules are added to a polynucleotide sample and divided into different divided sets based on the methylation level of the molecules, and then the nucleic acid molecules in each fraction are sequenced (single or together) and analyzed. In some embodiments, a fraction of nucleic acids is enriched for a specific target genomic region. In some embodiments, a fraction of nucleic acid molecules is amplified before and / or after the enrichment step. In some embodiments, enrichment can be performed after the divided sets are differentially tagged with molecular barcodes and recombined into a mixture of differentially tagged divided sets. These methods can be used in various applications, such as disease prognosis, diagnosis, and / or monitoring. In some embodiments, the disease is cancer.
[0078] Accordingly, in one aspect, the present disclosure provides a method of detecting the presence or absence of a tumor in a subject, comprising the steps of: (i) obtaining a polynucleotide sample from the subject; (ii) adding a set of carrier nucleic acid molecules to the polynucleotide sample to generate a first sample, wherein the set of carrier nucleic acid molecules comprises: (a) at least one subset of unmethylated carrier nucleic acid molecules; and / or (b) at least one subset of methylated carrier nucleic acid molecules, wherein the unmethylated carrier nucleic acid molecules do not comprise methylated nucleotides and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides; and (iii) adding the first sample to a capture nucleic acid molecule that selectively binds methylated polynucleotides. (iv) processing at least a portion of the distributed samples to generate processed samples, the processing comprising at least one of the following: (a) tagging, (b) amplifying, and (c) enriching DNA molecules for a specific region of interest; (v) sequencing at least a portion of the processed samples to generate a set of sequencing reads; and (vi) analyzing at least a portion of the set of sequencing reads to detect the presence or absence of a tumor.
[0079] In another aspect, the disclosure provides a method for detecting the presence or absence of a tumor in a subject, comprising the steps of: (i) obtaining a polynucleotide sample from the subject; (ii) adding a set of carrier nucleic acid molecules to the polynucleotide sample to generate a first sample, wherein the set of carrier nucleic acid molecules comprises: (a) at least one subset of unmethylated carrier nucleic acid molecules; and / or (b) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, wherein the unmethylated carrier nucleic acid molecules comprise no methylated nucleotides and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides; (iii) adding a set of carrier nucleic acid molecules to the polynucleotide sample to generate a first sample, wherein the set of carrier nucleic acid molecules comprises: (a) at least one subset of unmethylated carrier nucleic acid molecules; and / or (b) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, wherein the unmethylated carrier nucleic acid molecules comprise no methylated nucleotides and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides; (iv) processing at least a portion of the distributed samples to generate processed samples, the processing comprising at least one of the following steps: (a) tagging, (b) amplifying, and (c) enriching DNA molecules for specific regions of interest; (v) sequencing at least a portion of the processed samples to generate a set of sequencing reads; and (vi) analyzing at least a portion of the set of sequencing reads to detect the presence or absence of a tumor.
[0080] FIG. 1 shows an exemplary embodiment of a method 100 for detecting the presence or absence of a tumor in a subject. At 102, a polynucleotide sample from a subject is obtained. In some embodiments, the polynucleotide sample is obtained from the subject's tissue, blood, plasma, serum, urine, saliva, feces, cerebrospinal fluid, buccal swab, or thoracocentesis. In some embodiments, the polynucleotide sample is obtained from tissue. In some embodiments, the polynucleotide sample obtained from tissue is fragmented by either enzymatic or mechanical means / methods. In some embodiments, a fragmentase enzyme may be used to fragment DNA obtained from tissue. In some embodiments, the polynucleotide sample is obtained from blood. In some embodiments, the polynucleotide sample obtained from blood is a cell-free DNA sample. In some embodiments, the polynucleotide sample is a cell-free DNA sample.
[0081] At 104, a set of carrier nucleic acid molecules is added to the polynucleotide sample to generate a first sample. In some embodiments, the set of carrier nucleic acid molecules includes at least one subset of unmethylated carrier nucleic acid molecules and / or at least one subset of methylated carrier nucleic acid molecules, where the unmethylated carrier nucleic acid molecules do not include a methylated nucleotide and the methylated carrier nucleic acid molecules include one or more methylated nucleotides. In some embodiments, at least one end of the carrier nucleic acid molecules is modified to prevent ligation. In some embodiments, at least one end of the carrier nucleic acid molecules includes a C3 (propyl group) spacer. In some embodiments, at least one end of the carrier nucleic acid molecules includes a dideoxynucleotide. In some embodiments, at least one end of the carrier nucleic acid molecules includes any chemical modification that prevents a hydroxyl group from acting as a nucleophile. In some embodiments, the 5' end of the carrier nucleic acid molecule comprises at least one of the following modifications: (i) an inverted (5'-5')-dideoxythymine, dideoxycytosine, dideoxyguanine, or dideoxyadenine; (ii) a propyl group; or (iii) other organic functional groups, such as, but not limited to, benzyl, ethyl, or methyl. In some embodiments, the 3' end of the carrier nucleic acid molecule comprises at least one of the following modifications: (i) any dideoxy base that can be added enzymatically or during synthesis, such as, for example, dideoxythymine, dideoxycytosine, dideoxyguanine, or dideoxyadenine; (ii) a propyl group; or (iii) other organic functional groups, such as, but not limited to, benzyl, ethyl, or methyl.
[0082] In some embodiments, the one or more methylated nucleotides are at least one of the following: (i) 5-methylcytosine, (ii) 6-methyladenine, (iii) hydroxymethylcytosine, (iv) methyluracil, or (v) any other methylated nucleotide. In some embodiments, the methylated carrier nucleic acid molecule may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 15, 16, 17, 18, 19, or at least 20 methylated nucleotides. In some embodiments, the carrier nucleic acid molecule may be between 25 bp and 325 bp in length. In some embodiments, the unmethylated carrier nucleic acid molecules of one subset may have the same sequence as the unmethylated carrier nucleic acid molecules of another subset. In some embodiments, the methylated carrier nucleic acid molecules of one subset may have the same sequence as the methylated carrier nucleic acid molecules of another subset. In some embodiments, the sequence of the unmethylated carrier nucleic acid molecules in one subset is different from the sequence of the unmethylated carrier nucleic acid molecules in the other subset(s). In some embodiments, the sequence of the methylated carrier nucleic acid molecules in one subset differs from the sequence of the methylated carrier nucleic acid molecules in the other subset(s).
[0083] In some embodiments, one or more subsets of the unmethylated carrier nucleic acid molecules and / or the methylated carrier nucleic acid molecules may have one or more CpG dinucleotides in the nucleotide sequence. In some embodiments, one or more CpG dinucleotides in the methylated carrier nucleic acid molecules may have one or more methylated cytosines.
[0084] In some embodiments, the position of one or more CpG dinucleotides in one subset of unmethylated carrier nucleic acid molecules is different from the position of one or more CpG dinucleotides in the other subset(s). In some embodiments, the number of CpG dinucleotides in one subset of unmethylated carrier nucleic acid molecules is different from the number of CpG dinucleotides in the other subset(s). In some embodiments, the sequence of nucleotides adjacent to one or more CpG dinucleotides in one subset of unmethylated carrier nucleic acid molecules is different from the sequence of nucleotides adjacent to one or more CpG dinucleotides in the other subset(s). In some embodiments, the length of the unmethylated carrier nucleic acid molecules in one subset of unmethylated carrier nucleic acid molecules is different from the length of the unmethylated carrier nucleic acid molecules in the other subset(s).
[0085] In some embodiments, the position of one or more methylated nucleotides in one subset of methylated carrier nucleic acid molecules (i.e., relative to the end of the molecule) is different from the position of one or more methylated nucleotides in the other subset(s). In some embodiments, the number of methylated nucleotides in one subset of methylated carrier nucleic acid molecules is different from the number of methylated nucleotides in the other subset(s). In some embodiments, the sequence of nucleotides adjacent to one or more methylated nucleotides in one subset of unmethylated carrier nucleic acid molecules is different from the sequence of nucleotides adjacent to one or more methylated nucleotides in the other subset(s). In some embodiments, the length of the methylated carrier nucleic acid molecules in one subset of methylated carrier nucleic acid molecules is different from the length of the methylated carrier nucleic acid molecules in the other subset(s). In some embodiments, the carrier nucleic acid molecules comprise uracil nucleosides. In these embodiments, the method further comprises adding uracil deglycosylase and a DNA glycosylase-lyase (e.g., endonuclease VIII) prior to the amplification step.
[0086] In some embodiments, the amount of methylated carrier nucleic acid molecule to unmethylated carrier nucleic acid molecule is about 0:1, 0.1:99.9, 0.5:99.5, 0.75:99.25, 1:99, 1:95, 1:90, 1:80, 1:75, 1:70, 1:60, 1:50, 1:40, 1:30, 1:25, 1:20, 1:10, Ratios of 1:5, 1:2, 1:1.15, 1:1, 1.15:1, 2:1, 5:1, 10:1, 20:1, 25:1, 30:1, 40:1, 50:1, 60:1, 70:1, 75:1, 80:1, 90:1, 95:1, 99:1, 99.25:0.75, 99.5:0.5, 99.9:0.1 or 1:0. In some embodiments, the amount of polynucleotide sample to the set of carrier nucleic acid molecules is in a ratio of about 1:0.1; 1:0.2, 1:0.3, 1:4, 1:0.5, 1:6, 1:7, 1:8, 1:0.9, 1:1, 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:20 , 1:30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90, 1:100, 1:200, 1:300; 1:400, 1:500, 1:600, 1:700, 1:800, 1:900, 1:1000, 1:5000, 1:10,000, 1:100,000, 1:500,000, 1:10 6 , 1:10 7 , 1:10 8 or 1:10 9 In some embodiments, the amount is in units of mass. In some embodiments, the amount is in units of molar concentration. In some embodiments, the polynucleotide sample is at most 1 μg. In some embodiments, the polynucleotide sample is at most 200 ng. In some embodiments, the polynucleotide sample is at most 150 ng. In some embodiments, the polynucleotide sample is at most 100 ng. In some embodiments, the set of carrier nucleic acid molecules is added in an amount sufficient to make the total amount of the polynucleotide sample and the set of carrier nucleic acid molecules about 175 ng, 200 ng, 225 ng, 250 ng, 275 ng, 300 ng, 350 ng, 400 ng, 450 ng, 500 ng, 600 ng, 700 ng, 750 ng, 800 ng, 900 ng, 1 μg, 1.1 μg, 1.25 μg, or 1.5 μg.
[0087] In some embodiments, the sequence of the carrier nucleic acid molecule is selected from one of the following: (i) a sequence derived from a viral genome, (ii) a sequence derived from a bacterial genome, (iii) a sequence derived from a lambda genome, or (iv) a sequence derived from a non-human genome. In some embodiments, the carrier nucleic acid molecule is synthetic DNA. In some embodiments, the carrier nucleic acid molecule comprises a uracil nucleoside. In some embodiments, the carrier DNA molecule may be generated by PCR. In some embodiments, the carrier nucleic acid molecule generated by PCR may be further modified by treatment with a methyltransferase to incorporate a methyl group into one or more nucleotides in the carrier nucleic acid molecule. In some embodiments, the carrier nucleic acid molecule may be end-labeled with a polymerase to incorporate a modified nucleoside that prevents ligation of the carrier nucleic acid molecule to an adapter. In some embodiments, the carrier nucleic acid molecule comprises a non-naturally occurring nucleoside derivative. In some embodiments, the carrier nucleic acid molecule may be labeled with biotin or a fluorophore.
[0088] At 106, at least a portion of the first sample is partitioned or fractionated into at least two partitioned sets using capture agents that selectively bind to methylated polynucleotides, thereby generating partitioned samples. In some embodiments, the partitioning is based on differential binding affinities of nucleic acid molecules to the binders. Examples of binders include, but are not limited to, methyl-binding domains (MBDs) and methyl-binding proteins (MBPs). Examples of MBPs contemplated herein include, but are not limited to: (a) MeCP2 is a protein that preferentially binds 5-methyl-cytosine over unmodified cytosine; (b) RPL26, PRP8, and the DNA mismatch repair protein MHS6 preferentially bind 5-hydroxymethyl-cytosine over unmodified cytosine; (c) FOXK1, FOXK2, FOXP1, FOXP4, and FOXI3 bind 5-formyl-cytosine more preferentially than unmodified cytosine (Iurlaro et al., Genome Biol. 14, R119 (2013)); and (d) an antibody specific for one or more methylated nucleotide bases;
[0089] For some affinity agents and modifications, binding to the agent can occur essentially in an all-or-none manner depending on whether the nucleic acid has a modification (e.g., methylation), but separation can be a matter of degree. In such embodiments, nucleic acids that are overrepresented with the modification will bind to the agent to a greater extent than nucleic acids that are underrepresented with the modification. Alternatively, nucleic acids with the modification may bind in an all-or-none manner. However, various levels of modification can be sequentially eluted from the binding agent.
[0090] For example, in some embodiments, the partitioning can be binary or based on the degree / level of methylation. For example, all methylated molecules can be partitioned from unmethylated molecules using a methyl-binding domain protein (e.g., MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific)). Further partitioning can then involve eluting fragments with different levels of methylation by adjusting the salt concentration in a solution containing the methyl-binding domain and bound fragments. As the salt concentration increases, molecules with greater methylation levels are eluted.
[0091] In some embodiments, partitioning comprises partitioning the nucleic acid molecules based on their differential binding affinity to binding agents that preferentially bind nucleic acid molecules that include methylated nucleotides.
[0092] In some embodiments, the distributed set represents nucleic acids with different degrees of modification (overrepresentation or underrepresentation of the modification). Overrepresentation and underrepresentation can be defined by the number of methylated nucleotides that a nucleic acid has compared to the median number of methylated nucleotides per molecule in the population. For example, if the median number of 5-methylcytosine nucleotides in nucleic acid molecules in a sample is 2, nucleic acid molecules containing more than two 5-methylcytosine residues will be overrepresented in this modification, and nucleic acids with one or zero 5-methylcytosine residues will be underrepresented. The effect of affinity separation is to separate nucleic acids that are overrepresented in the modification in the binding phase and nucleic acids that are underrepresented in the modification in the non-binding phase (i.e., in solution). The nucleic acids in the binding phase can be eluted before subsequent processing.
[0093] When using the MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific), various levels of methylation can be separated using sequential elution. For example, the low-methylated fraction (no methylation) can be separated from the methylated fraction by contacting the nucleic acid population with MBD from the kit bound to magnetic beads. The beads are used to separate methylated nucleic acids from unmethylated nucleic acids. One or more elution steps are then performed sequentially to elute nucleic acids with different levels of methylation. For example, the first set of methylated nucleic acids can be eluted at a salt concentration of about 150 mM or about 160 mM or higher, for example, at least 150 mM, 200 mM, 300 mM, 400 mM, 500 mM, 600 mM, 700 mM, 800 mM, 900 mM, 1000 mM, or 2000 mM. After elution of such methylated nucleic acids, magnetic separation is again used to separate highly methylated nucleic acids from nucleic acids with low levels of methylation. The elution and magnetic separation steps can be repeated to generate various fractions, such as a low methylated fraction (representative of no methylation), a moderately methylated fraction (representative of low levels of methylation), and a high methylated fraction (representative of high levels of methylation).
[0094] In some methods, nucleic acids bound to the agent used for affinity separation are subjected to a washing step. The washing step washes away nucleic acids that are weakly bound to the affinity agent. Such nucleic acids can be enriched for nucleic acids with modifications to a degree close to the average or median (i.e., intermediate between nucleic acids that remain bound to the solid phase and nucleic acids that do not bind to the solid phase when the agent is first contacted with the sample). Affinity separation results in at least two, and sometimes three or more, fractions of nucleic acids with different degrees of modification. The distribution of nucleic acid molecules can be analyzed by sequencing the distributed nucleic acid molecules, by digital droplet PCR (ddPCR), or by quantitative PCR (qPCR). In some embodiments, instead of adding carrier nucleic acid molecules before the distribution step, carrier nucleic acid molecules can be added either during the washing step (i.e., while collecting the moderately distributed set) or during the elution step (i.e., while collecting the highly distributed set).
[0095] At 108, at least a portion of the distributed sample is processed to generate a processed sample. The processing step includes tagging; amplifying and / or enriching molecules for specific regions of interest. In some embodiments, prior to amplification, each of at least two distributed sets is differentially tagged. The tagged distributed sets are then pooled together prior to amplification. Differential tagging of distributed sets aids in tracking nucleic acid molecules belonging to a particular distributed set. Tags are typically provided as components of adapters. Nucleic acid molecules in different distributed sets receive different tags that can distinguish members of one distributed set from members of another distributed set. Tags linked to nucleic acid molecules of the same fraction set can be the same or different from each other. However, if different from each other, the tags can share a portion of their sequence to identify the molecules to which they are attached as belonging to a particular distributed set. For example, if the molecules of a first sample are distributed into two distributed sets—P1 and P2—the molecules in P1 can be tagged with A1, A2, A3, etc., and the molecules in P2 can be tagged with B1, B2, B3, etc. Such a tagging system allows for differentiation of distributed sets and between molecules within a distributed set. In some embodiments, the tag (i.e., molecular barcode) is part of an adapter, and the adapter contains a universal primer binding site. The adapter containing the tag is attached via ligation. In some embodiments, the end of the carrier nucleic acid molecule is modified to prevent ligation. In these embodiments, the adapter (which contains the tag, i.e., molecular barcode) is not attached to the carrier nucleic acid molecule. Therefore, the carrier nucleic acid molecule cannot be amplified using a universal primer. After tagging, these molecules are amplified using primers that bind to the primer binding region present in the adapter (ligated to the molecule). This amplification amplifies only the polynucleotides obtained from the subject, not the carrier nucleic acid molecule. After amplification, the molecules are enriched for the specific region of interest.
[0096] At 110, at least a portion of the processed sample is sequenced to generate a set of sequencing reads. In some embodiments, at least a portion of the processed sample from at least two distributed sets is sequenced to generate a set of sequencing reads. The obtained sequence information includes sequences of tags (i.e., molecular barcodes) attached to the nucleic acid molecules and polynucleotides. From the sequences of the tags (i.e., molecular barcodes) attached to the polynucleotides, the tags (i.e., molecular barcodes) can be correlated with distributed sets of polynucleotides. The sequence information is used to identify the polynucleotides (obtained from the subject) and their corresponding distributed sets. At 112, at least a portion of the set of sequencing reads is analyzed to detect the presence or absence of a tumor. In some embodiments, the analyzing step includes determining the methylation status of molecules. For example, specific regions of interest have previously been determined to be unmethylated in healthy individuals and methylated in individuals with malignant tumors. The analyzing step includes determining whether molecules are methylated in these regions of interest. This is determined based on the number of CpG residues in the molecule and the distributed set to which the molecule is distributed, which is then used to detect the presence or absence of a tumor.
[0097] II. Carrier Nucleic Acid Molecules Carrier nucleic acid molecules are used in analyzing polynucleotides in partitioning assays. In some embodiments, the use of carrier nucleic acid molecules helps increase the specificity of partitioning methylated polynucleotides using methyl-binding proteins. Methyl-binding domain proteins have low affinity for unmethylated molecules. Therefore, when incubated with DNA obtained from a subject, molecules that do not contain methylated cytosines may be unintentionally captured. This incomplete specificity reduces the ability to efficiently separate methylated molecules from unmethylated molecules. Therefore, by including carrier nucleic acid molecules in the assay, the carrier nucleic acid molecules can bind to the methyl-binding proteins, preventing unmethylated molecules from the polynucleotide sample from binding to the methyl-binding proteins, thereby improving the specificity of partitioning methylated molecules. Similarly, carrier nucleic acid molecules can be used to improve the partitioning specificity of other methylated nucleic acid-specific binding reagents, such as antibodies, antibody derivative molecules, etc. In some embodiments, carrier nucleic acid molecules improve the preparation of libraries of molecules for next-generation sequencing.
[0098] In some embodiments, the carrier nucleic acid molecule may have a naturally occurring nucleic acid sequence. In some embodiments, the carrier nucleic acid molecule may have a non-naturally occurring nucleic acid sequence. In some embodiments, the carrier nucleic acid molecule may be a synthetic oligonucleotide. In some embodiments, the carrier nucleic acid molecule may have a nucleic acid sequence corresponding to a non-human genome. For example, these molecules may have (i) a sequence corresponding to a region of lambda phage DNA or the human genome, (ii) a sequence corresponding to a region of a viral or bacterial genome, (iii) a non-naturally occurring sequence, and / or (iv) any combination of the above. In some embodiments, the carrier DNA molecule may be generated by PCR. In some embodiments, the carrier nucleic acid molecule generated by PCR may be further modified by treatment with a methyltransferase to incorporate a methyl group into one or more nucleotides in the carrier nucleic acid molecule. In some embodiments, the carrier nucleic acid molecule may be end-labeled with a polymerase to incorporate a modified nucleoside that prevents ligation of the carrier nucleic acid molecule to an adapter. In some embodiments, the carrier nucleic acid molecule comprises a non-naturally occurring nucleoside derivative. In some embodiments, the carrier nucleic acid molecule may be labeled with biotin or a fluorophore.
[0099] In another aspect, the present disclosure provides a set of carrier nucleic acid molecules comprising: (i) at least one subset of unmethylated carrier nucleic acid molecules; and / or (ii) at least one subset of methylated carrier nucleic acid molecules, wherein the unmethylated carrier nucleic acid molecules do not comprise a methylated nucleotide and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides. In some embodiments, at least one end of the carrier nucleic acid molecules is modified to prevent ligation.
[0100] In another aspect, the disclosure provides a set of carrier nucleic acid molecules comprising: (i) at least one subset of unmethylated carrier nucleic acid molecules; and / or (ii) at least one subset of methylated carrier nucleic acid molecules, wherein at least one subset of the unmethylated carrier nucleic acid molecules does not comprise methylated nucleotides, and at least one subset of the methylated carrier nucleic acid molecules comprises methylated nucleotides, and at least one end of the carrier nucleic acid molecules is modified to prevent ligation.
[0101] In some embodiments, at least one end of the carrier nucleic acid molecule comprises a C3 (propyl group) spacer. In some embodiments, at least one end of the carrier nucleic acid molecule comprises a dideoxynucleotide. In some embodiments, at least one end of the carrier nucleic acid molecule comprises any chemical modification that prevents a hydroxyl group from acting as a nucleophile. In some embodiments, the 5' end of the carrier nucleic acid molecule comprises at least one of the following modifications: (i) an inverted (5'-5')-dideoxythymine, dideoxycytosine, dideoxyguanine, or dideoxyadenine; (ii) a propyl group; or (iii) another organic functional group, such as, but not limited to, benzyl, ethyl, or methyl. In some embodiments, the 3' end of the carrier nucleic acid molecule comprises at least one of the following modifications: (i) any dideoxy base that can be added enzymatically or during synthesis, such as dideoxythymine, dideoxycytosine, dideoxyguanine, or dideoxyadenine; (ii) a propyl group, or (iii) other organic functional groups, such as, but not limited to, benzyl, ethyl, or methyl.
[0102] In some embodiments, the one or more methylated nucleotides are at least one of the following: (i) 5-methylcytosine, (ii) 6-methyladenine, (iii) hydroxymethylcytosine, (iv) methyluracil, or (v) any other methylated nucleotide. In some embodiments, the methylated carrier nucleic acid molecule may have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 15, 16, 17, 18, 19, or at least 20 methylated nucleotides. In some embodiments, the carrier nucleic acid molecule may be between 25 bp and 325 bp in length. In some embodiments, the unmethylated carrier nucleic acid molecules of one subset may have the same sequence as the unmethylated carrier nucleic acid molecules of another subset. In some embodiments, the methylated carrier nucleic acid molecules of one subset may have the same sequence as the methylated carrier nucleic acid molecules of another subset. In some embodiments, the sequence of the unmethylated carrier nucleic acid molecules in one subset is different from the sequence of the unmethylated carrier nucleic acid molecules in the other subset(s). In some embodiments, the sequence of the methylated carrier nucleic acid molecules in one subset differs from the sequence of the methylated carrier nucleic acid molecules in the other subset(s).
[0103] In some embodiments, one or more subsets of the unmethylated carrier nucleic acid molecules and / or the methylated carrier nucleic acid molecules may have one or more CpG dinucleotides in the nucleotide sequence. In some embodiments, one or more CpG dinucleotides in the methylated carrier nucleic acid molecules may have one or more methylated cytosines.
[0104] In some embodiments, the position of one or more CpG dinucleotides in one subset of unmethylated carrier nucleic acid molecules is different from the position of one or more CpG dinucleotides in the other subset(s). In some embodiments, the number of CpG dinucleotides in one subset of unmethylated carrier nucleic acid molecules is different from the number of CpG dinucleotides in the other subset(s). In some embodiments, the sequence of nucleotides adjacent to one or more CpG dinucleotides in one subset of unmethylated carrier nucleic acid molecules is different from the sequence of nucleotides adjacent to one or more CpG dinucleotides in the other subset(s). In some embodiments, the length of the unmethylated carrier nucleic acid molecules in one subset of unmethylated carrier nucleic acid molecules is different from the length of the unmethylated carrier nucleic acid molecules in the other subset(s).
[0105] In some embodiments, the position of one or more methylated nucleotides in one subset of methylated carrier nucleic acid molecules is different from the position of one or more methylated nucleotides in the other subset(s). In some embodiments, the number of methylated nucleotides in one subset of methylated carrier nucleic acid molecules is different from the number of methylated nucleotides in the other subset(s). In some embodiments, the sequence of nucleotides adjacent to one or more methylated nucleotides in one subset of unmethylated carrier nucleic acid molecules is different from the sequence of nucleotides adjacent to one or more methylated nucleotides in the other subset(s). In some embodiments, the length of the methylated carrier nucleic acid molecules in one subset of methylated carrier nucleic acid molecules is different from the length of the methylated carrier nucleic acid molecules in the other subset(s).
[0106] In some embodiments, the amount of methylated carrier nucleic acid molecule to unmethylated carrier nucleic acid molecule is about 0:1, 0.1:99.9, 0.5:99.5, 0.75:99.25, 1:99, 1:95, 1:90, 1:80, 1:75, 1:70, 1:60, 1:50, 1:40, 1:30, 1:25, 1:20, 1:10, In some embodiments, the ratio is 1:5, 1:2, 1:1.15, 1:1, 1.15:1, 2:1, 5:1, 10:1, 20:1, 25:1, 30:1, 40:1, 50:1, 60:1, 70:1, 75:1, 80:1, 90:1, 95:1, 99:1, 99.25:0.75, 99.5:0.5, 99.9:0.1, or 1:0. In some embodiments, the amount is in units of mass. In some embodiments, the amount is in units of molar concentration. In some embodiments, the carrier nucleic acid molecule comprises a uracil nucleoside.
[0107] Figure 2 is a schematic diagram of a carrier nucleic acid molecule according to one embodiment of the present disclosure. The carrier nucleic acid molecule in Figure 2 is a double-stranded DNA molecule. The "---" region in the double-stranded DNA sequence indicates an optional DNA sequence. R1 and R3 indicate modifications at the 5' end of the carrier nucleic acid molecule that prevent ligation of the carrier nucleic acid molecule. R2 and R4 indicate modifications at the 3' end of the carrier nucleic acid molecule that prevent ligation of the carrier nucleic acid molecule. In some embodiments, the end of the carrier nucleic acid molecule comprises a C3 (propyl group) spacer. In some embodiments, the end of the carrier nucleic acid molecule comprises a dideoxynucleotide or an inverted dideoxynucleotide. In some embodiments, the 5' end of the carrier nucleic acid molecule comprises an inverted dideoxynucleotide. In some embodiments, the 3' end of the carrier nucleic acid molecule comprises a dideoxynucleotide. In some embodiments, the end of the carrier nucleic acid molecule comprises any chemical modification that prevents the hydroxyl group from acting as a nucleophile. In some embodiments, the 5'-end of the carrier nucleic acid molecule comprises at least one of the following modifications: (i) an inverted (5'-5')-dideoxythymine, dideoxycytosine, dideoxyguanine, or dideoxyadenine; (ii) a propyl group; or (iii) other organic functional groups, such as, but not limited to, benzyl, ethyl, or methyl. In some embodiments, the 3'-end of the carrier nucleic acid molecule comprises at least one of the following modifications: (i) any dideoxy base that can be added enzymatically or during synthesis, such as, for example, dideoxythymine, dideoxycytosine, dideoxyguanine, or dideoxyadenine; (ii) a propyl group; or (iii) other organic functional groups, such as, but not limited to, benzyl, ethyl, or methyl. In some embodiments, the DNA sequence of the carrier nucleic acid molecule comprises one or more CpG dinucleotides. In some embodiments, the carrier nucleic acid molecule comprises one or more methylated nucleotides.
[0108] FIG. 3A is a schematic diagram of a set of carrier nucleic acid molecules according to one embodiment of the present disclosure. In this embodiment, the carrier nucleic acid molecules are double-stranded DNA molecules and are referred to as carrier DNA molecules. Also, in this embodiment, the ends of the carrier DNA molecules are modified to prevent ligation. In other embodiments, the ends of the carrier DNA molecules need not be modified to prevent ligation. For illustrative purposes, only one strand of a single carrier DNA molecule is shown in the figure for all subsets. In FIG. 3A, the "---" region in the carrier DNA molecule indicates any other sequence other than CpG, where M indicates 5-methylcytosine, C indicates cytosine, and G indicates guanine. In this embodiment, the carrier DNA molecules have one subset of unmethylated carrier DNA molecules (Subset 1) and one subset of methylated carrier DNA molecules (Subset A). In this embodiment, the sequences of the carrier DNA molecules in both Subset 1 and Subset A are the same. Both the carrier DNA molecules in Subset 1 and Subset A have three CpG dinucleotides. In this embodiment, subset A, the cytosines in all three CpG dinucleotides are methylated (hence designated MG in Figure 3A). In some embodiments, the carrier DNA molecule can be of any length between 25 bp and 325 bp. In some embodiments, the carrier DNA molecule can be at least 50 bp, at least 60 bp, at least 80 bp, at least 100 bp, at least 120 bp, at least 150 bp, at least 200 bp, at least 250 bp, or at least 300 bp.
[0109] Figure 3B is a schematic diagram of a set of carrier nucleic acid molecules according to one embodiment of the present disclosure. In this embodiment, the carrier nucleic acid molecules are double-stranded DNA molecules and are referred to as carrier DNA molecules. Also, in this embodiment, the ends of the carrier DNA molecules are modified to prevent ligation. In other embodiments, the ends of the carrier DNA molecules need not be modified to prevent ligation. For illustrative purposes, only one strand of a single carrier DNA molecule is shown in the figure for all subsets. In Figure 3B, the "---" region in the carrier DNA molecule indicates any other sequence other than CpG, where M represents 5-methylcytosine, C represents cytosine, and G represents guanine. In this embodiment, the carrier DNA molecules have two subsets of unmethylated carrier DNA molecules (subsets 1 and 2) and two subsets of methylated carrier DNA molecules (subsets A and B). In this embodiment, each subset (subsets 1, 2, A, and B) has a different nucleotide sequence, and each subset has a different number of CpG dinucleotides. That is, subset 1 has two CpG dinucleotides, subset 2 has four CpG dinucleotides, subset A has three CpG dinucleotides, and subset B has five CpG dinucleotides. In this embodiment, all cytosines of the CpG dinucleotides in the methylated carrier DNA molecule are methylated (hence, designated MG in FIG. 3B). In other embodiments, not all cytosines of the CpG dinucleotides in the methylated carrier DNA molecule need to be methylated (see, e.g., FIG. 4). In some embodiments, the carrier DNA molecule can be of any length between 25 bp and 325 bp. In some embodiments, the carrier DNA molecule can be at least 50 bp, at least 60 bp, at least 80 bp, at least 100 bp, at least 120 bp, at least 150 bp, at least 200 bp, at least 250 bp, or at least 300 bp.
[0110] FIG. 4 is a schematic diagram of a set of carrier nucleic acid molecules according to one embodiment of the present disclosure. In this embodiment, the carrier nucleic acid molecules are double-stranded DNA molecules and are referred to as carrier DNA molecules. Also, in this embodiment, the ends of the carrier DNA molecules are modified to prevent ligation. In other embodiments, the ends of the carrier DNA molecules do not need to be modified to prevent ligation. In this embodiment, the carrier DNA molecules described herein may take into account the methylation effect of methylated cytosines in CpG dinucleotides during partitioning of the nucleic acid molecules. For illustrative purposes, only one strand of a single carrier DNA molecule is shown in the figure for all subsets. In FIG. 4, the "---" region in the carrier DNA molecule indicates any other sequence other than CpG, M indicates 5-methylcytosine, C indicates cytosine, and G indicates guanine. In this embodiment, the carrier DNA molecules have one subset of unmethylated carrier DNA molecules (subset 1) and three subsets of methylated carrier DNA molecules (subsets A, B, and C). In this embodiment, the sequences of the carrier DNA molecules in all subsets are the same. That is, subsets 1, A, B, and C have the same nucleotide sequence, and all of them have four CpG dinucleotides. However, the number of methylated nucleotides (in this embodiment, methylated cytosines in CpG dinucleotides) is different in each subset of methylated carrier DNA molecules. In this embodiment, subset A has two methylated cytosines in CpG nucleotides, subset B has three methylated cytosines in CpG nucleotides, and subset C has four methylated cytosines in CpG nucleotides (shown as MG in Figure 4). In some embodiments, the carrier DNA molecules can be any length between 25 bp and 325 bp. In some embodiments, the carrier DNA molecules can be at least 50 bp, at least 60 bp, at least 80 bp, at least 100 bp, at least 120 bp, at least 150 bp, at least 200 bp, at least 250 bp, or at least 300 bp.
[0111] FIG. 5 is a schematic diagram of a set of carrier nucleic acid molecules according to one embodiment of the present disclosure. In this embodiment, the carrier nucleic acid molecules are double-stranded DNA molecules and are referred to as carrier DNA molecules. Also, in this embodiment, the ends of the carrier DNA molecules are modified to prevent ligation. In other embodiments, the ends of the carrier DNA molecules do not need to be modified to prevent ligation. In this embodiment, the carrier DNA molecules described herein may take into account the position-specific effects and / or methylation effects of CpG dinucleotides during distribution of the nucleic acid molecules. For illustrative purposes, only one strand of a single carrier DNA molecule is shown in the figure for all subsets. In FIG. 5, the "---" region in the carrier DNA molecule indicates any other sequence other than CpG, M indicates 5-methylcytosine, C indicates cytosine, and G indicates guanine. In this embodiment, the set of carrier DNA molecules has two subsets of unmethylated carrier DNA molecules (subsets 1 and 2) and two subsets of methylated carrier DNA molecules (subsets A and B). In this embodiment, the sequence of the carrier DNA molecules in each subset is different from the other subsets. That is, subsets 1, 2, A, and B have different nucleotide sequences. In this embodiment, the number of CpG dinucleotides in all subsets is the same. That is, all subsets (subsets 1, 2, A, and B) have four CpG dinucleotides. In this embodiment, the number of methylated nucleotides (in this embodiment, methylated cytosines in CpG dinucleotides) in each subset of methylated carrier DNA molecules is the same, and all four cytosines in the four CpG dinucleotides are methylated (shown as MG in subsets A and B in FIG. 5). In this embodiment, the positions of CpG dinucleotides differ between subset 1 and subset 2 and between subset A and subset B. However, the positions of CpG dinucleotides in subset 1 and subset A are the same, and the positions of CpG dinucleotides in subset 2 and subset B are the same. In some embodiments, the carrier DNA molecules can be of any length between 25 bp and 325 bp.In some embodiments, the carrier DNA molecule may be at least 50 bp, at least 60 bp, at least 80 bp, at least 100 bp, at least 120 bp, at least 150 bp, at least 200 bp, at least 250 bp, or at least 300 bp.
[0112] Figure 6 is a schematic diagram of a set of carrier nucleic acid molecules according to one embodiment of the present disclosure. In this embodiment, the carrier nucleic acid molecules are double-stranded DNA molecules and are referred to as carrier DNA molecules. Also, in this embodiment, the ends of the carrier DNA molecules are modified to prevent ligation. In other embodiments, the ends of the carrier DNA molecules do not need to be modified to prevent ligation. In this embodiment, the carrier DNA molecules described herein may take into account the sequence-specific effects of nucleotides adjacent to CpG dinucleotides during partitioning of the nucleic acid molecules. For illustrative purposes, only one strand of a single carrier DNA molecule is shown in the figure for all subsets. In Figure 6, the "---" region in the carrier DNA molecule indicates any other sequence apart from the CpG dinucleotide; M indicates 5-methylcytosine, C indicates cytosine, G indicates guanine; X indicates 5-methylcytosine, C indicates cytosine, G indicates guanine; and X indicates 5-methylcytosine, C indicates 5-methylcytosine, G indicates guanine, ... ni and Y ni can be any two different nucleotide sequences of the same length n i , where n i can be n l , n 2 , n 3 , n 4 , n 5 or n 6 . That is, for example, X n1 and Y n1 are two different nucleotide sequences of length n1, where n1, n2, n3, n4, n5, and n6 can be any integer between 0 and 30. In this embodiment, the set of carrier DNA molecules has two subsets of unmethylated carrier DNA molecules (subsets 1 and 2) and two subsets of methylated carrier DNA molecules (subsets A and B), and the number of CpG dinucleotides in all subsets is the same. That is, subsets 1, 2, A, and B have four CpG dinucleotides. However, the sequences of nucleotides adjacent to the CpG dinucleotides differ between the two subsets of unmethylated carrier DNA molecules. That is, the sequences of nucleotides adjacent to the CpG dinucleotides (X niand Y ni ) differs between subset 1 and subset 2. Similarly, the sequence of nucleotides adjacent to the CpG dinucleotide differs between the two subsets of methylated carrier DNA molecules. That is, the sequence of nucleotides adjacent to the CpG dinucleotide (X ni and Y ni ) is different in subset A and subset B. However, the sequence of nucleotides adjacent to the CpG dinucleotide (X ni ) is the same in subset 1 and subset A, and is also the same in subset 1, and the sequence of nucleotides adjacent to the CpG dinucleotide (Y ni ) are the same in subset 2 and subset B. In some embodiments, the carrier DNA molecule can be of any length between 25 bp and 325 bp. In some embodiments, the carrier DNA molecule can be at least 50 bp, at least 60 bp, at least 80 bp, at least 100 bp, at least 120 bp, at least 150 bp, at least 200 bp, at least 250 bp, or at least 300 bp. In some embodiments, n i can be a nucleotide sequence up to 5 bp in length.
[0113] In another aspect, the disclosure provides a population of nucleic acids comprising: (i) a set of carrier nucleic acid molecules comprising i) at least one subset of unmethylated carrier nucleic acid molecules; and / or (ii) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, and wherein the unmethylated carrier nucleic acid molecules do not comprise methylated nucleotides and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides; and (ii) a polynucleotide sample obtained from a subject.
[0114] In some embodiments, the carrier nucleic acid molecule may have a sequence corresponding to a region of (i) a viral genome, (ii) a bacterial genome, (iii) a lambda phage genome, (iv) a human genome, (iv) any naturally occurring sequence, (v) a non-naturally occurring sequence, (vi) a non-human genome, and / or (vii) any combination of the above. In some embodiments, the carrier nucleic acid molecule may comprise a non-naturally occurring sequence. In some embodiments, the carrier nucleic acid molecule may comprise a naturally occurring sequence. In some embodiments, the carrier DNA molecule may be generated by PCR. In some embodiments, the carrier nucleic acid molecule generated by PCR may be further modified by treatment with a methyltransferase to incorporate a methyl group into one or more nucleotides in the carrier nucleic acid molecule. In some embodiments, the carrier nucleic acid molecule may be end-labeled with a polymerase to incorporate a modified nucleoside that prevents ligation of the carrier nucleic acid molecule to an adapter. In some embodiments, the carrier nucleic acid molecule comprises a non-naturally occurring nucleoside derivative. In some embodiments, the carrier nucleic acid molecule may be labeled with biotin or a fluorophore.
[0115] In some embodiments, the polynucleotide sample is obtained from tissue, blood, plasma, serum, urine, saliva, feces, cerebrospinal fluid, buccal swab or thoracentesis. In some embodiments, the polynucleotide sample is obtained from tissue. In some embodiments, the polynucleotide sample obtained from tissue is fragmented by enzymatic or mechanical means. In some embodiments, the polynucleotide sample is obtained from blood. In some embodiments, the polynucleotide sample obtained from blood is a cell-free DNA sample. In some embodiments, the polynucleotide sample is a cell-free DNA sample.
[0116] In some embodiments, the polynucleotide sample is a DNA sample, an RNA sample, a cell-free polynucleotide sample, a cell-free DNA sample, or a cell-free RNA sample. In some embodiments, the polynucleotide sample is a cell-free DNA sample.
[0117] In some embodiments, the amount of at least one subset of methylated carrier nucleic acid molecules relative to at least one subset of unmethylated carrier nucleic acid molecules is about 0:1, 0.1:99.9, 0.5:99.5, 0.75:99.25, 1:99, 1:95, 1:90, 1:80, 1:75, 1:70, 1:60, 1:50, 1:40, 1:30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90, 1:90, 1:90, 1:90, 1:1 ... :25, 1:20, 1:10, 1:5, 1:2, 1:1.15, 1:1, 1.15:1, 2:1, 5:1, 10:1, 20:1, 25:1, 30:1, 40:1, 50:1, 60:1, 70:1, 75:1, 80:1, 90:1, 95:1, 99:1, 99.25:0.75, 99.5:0.5, 99.9:0.1 or 1:0 ratios. In some embodiments, the amount of polynucleotide sample to the set of carrier nucleic acid molecules is in a ratio of about 1:0.1; 1:0.2, 1:0.3, 1:4, 1:0.5, 1:6, 1:7, 1:8, 1:0.9, 1:1, 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:20 , 1:30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90, 1:100, 1:200, 1:300; 1:400, 1:500, 1:600, 1:700, 1:800, 1:900, 1:1000, 1:5000, 1:10,000, 1:100,000, 1:500,000, 1:10 6 , 1:10 7 , 1:10 8 or 1:10 9 In some embodiments, the amount is in units of mass. In some embodiments, the amount is in units of molar concentration.
[0118] In some embodiments, the polynucleotide sample is at least 1 ng, at least 5 ng, at least 10 ng, at least 15 ng, at least 20 ng, at least 30 ng, at least 50 ng, at least 75 ng, at least 100 ng, at least 150 ng, at least 200 ng, at least 250 ng, at least 300 ng, at least 350 ng, at least 400 ng, at least 450 ng, at least 500 ng, at least 750 ng, or at least 1 μg. In some embodiments, the polynucleotide sample is at most 1 μg. In some embodiments, the polynucleotide sample is at most 200 ng. In some embodiments, the polynucleotide sample is at most 150 ng. In some embodiments, the polynucleotide sample is at most 100 ng. In some embodiments, the set of carrier nucleic acid molecules is added in an amount sufficient to provide a total amount of about 175 ng, 200 ng, 225 ng, 250 ng, 275 ng, 300 ng, 350 ng, 400 ng, 450 ng, 500 ng, 600 ng, 700 ng, 750 ng, 800 ng, 900 ng, 1 μg, 1.1 μg, 1.25 μg, or 1.5 μg of polynucleotide sample and set of carrier nucleic acid molecules.
[0119] III. General Features of the Method A. Sample The sample may be any biological sample isolated from a subject, including bodily tissue, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsies (e.g., biopsies from known or suspected solid tumors), cerebrospinal fluid, synovial fluid, lymphatic fluid, peritoneal fluid, interstitial or extracellular fluid (e.g., fluid from the spaces between cells), gingival crevicular fluid, gingival crevicular fluid, gingival ductal fluid, endothelial cells ... Crevicular fluid, bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat and urine The sample may include. Samples may be bodily fluids, such as blood and its fractions, and urine. Such samples may contain nucleic acids shed from tumors. Nucleic acids may include DNA and RNA, and may be in double-stranded and single-stranded forms. Samples may be in the form originally isolated from a subject, or may be further processed to remove or add components such as cells, enrich one component for another, or convert one type of nucleic acid into another type of nucleic acid, for example, convert RNA into DNA or single-stranded nucleic acid into double-stranded nucleic acid. Thus, for example, the bodily fluid for analysis may be plasma or serum containing cell-free nucleic acid, for example, cell-free DNA (cfDNA).
[0120] In some embodiments, the sample volume of bodily fluid collected from a subject depends on the desired read depth of the region being sequenced. Example volumes are about 0.4 to 40 milliliters (mL), about 5 to 20 mL, or about 10 to 20 mL. For example, the volume can be about 0.5 mL, about 1 mL, about 5 mL, about 10 mL, about 20 mL, about 30 mL, about 40 mL, or more. The volume of sampled plasma is typically between about 5 mL and about 20 mL.
[0121] Samples can contain varying amounts of nucleic acid. Typically, the amount of nucleic acid in a given sample is considered to be equivalent to multiple genome equivalents. For example, a sample of about 30 nanograms (ng) of DNA will contain about 10,000 (10 4 ) haploid human genome equivalents, and in the case of cfDNA, approximately 200 billion (2 × 10 11 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA may contain about 30,000 haploid human genome equivalents, or about 600 billion individual molecules in the case of cfDNA.
[0122] In some embodiments, the sample contains nucleic acids from different sources, for example, from cells and from acellular sources (e.g., blood samples, etc.). Typically, the sample contains nucleic acids having mutations. For example, the sample optionally contains DNA having germline mutations and / or somatic mutations. Typically, the sample contains DNA having cancer-associated mutations (e.g., cancer-associated somatic mutations).
[0123] Exemplary amounts of cell-free nucleic acid in a sample prior to amplification typically range from about 1 femtogram (fg) to about 1 microgram (μg), e.g., from about 1 picogram (pg) to about 200 nanograms (ng), from about 1 ng to about 100 ng, or from about 10 ng to about 1000 ng. In some embodiments, the sample contains up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. Optionally, the amount is at least about 1 fg, at least about 10 fg, at least about 100 fg, at least about 1 pg, at least about 10 pg, at least about 100 pg, at least about 1 ng, at least about 10 ng, at least about 100 ng, at least about 150 ng, or at least about 200 ng of cell-free nucleic acid molecules. In some embodiments, the amount is up to about 1 fg, 10 fg, 100 fg, 1 pg, 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, 200 ng, 300 ng, 400 ng, 500 ng, 600 ng, 700 ng, 800 ng, 900 ng, or 1 μg of cell-free nucleic acid molecules. In some embodiments, the method includes obtaining between about 1 fg and about 200 ng of cell-free nucleic acid molecules from the sample.
[0124] Cell-free nucleic acids typically have a size distribution between about 100 and about 500 nucleotides in length, with molecules between about 110 and about 230 nucleotides in length representing about 90% of the molecules in the sample, with the mode being about 168 nucleotides in length (in samples from human subjects), and a second, smaller peak ranging between about 240 and about 440 nucleotides in length. In some embodiments, the cell-free nucleic acids are between about 160 and about 180 nucleotides in length, or between about 320 and about 360 nucleotides in length, or between about 440 and about 480 nucleotides in length.
[0125] In some embodiments, cell-free nucleic acids are isolated from bodily fluids via a partitioning step, in which cell-free nucleic acids found in solution are separated from intact cells and other insoluble components of the bodily fluid. In some embodiments, partitioning includes techniques such as centrifugation or filtration. Alternatively, cells in the bodily fluid can be lysed, and both cell-free and cellular nucleic acids can be processed. Generally, after the addition of buffer and washing steps, the cell-free nucleic acids can be precipitated, for example, with alcohol. In some embodiments, additional cleaning steps, such as silica-based columns to remove contaminants or salts, are used. Exemplary procedural aspects, such as optimizing yield, include the optional addition of nonspecific bulk carrier nucleic acids throughout the reaction. After such processing, the sample typically contains various forms of nucleic acids, including double-stranded DNA, single-stranded DNA, and / or single-stranded RNA. If necessary, the single-stranded DNA and / or single-stranded RNA are converted to double-stranded forms so that they can be included in subsequent processing and analysis steps.
[0126] B. Distribution and Tagging In some embodiments, nucleic acid molecules (derived from a sample of polynucleotides) can be tagged with a sample index and / or a molecular barcode (commonly referred to as a "tag"). The tag can be incorporated into or otherwise attached to an adapter by chemical synthesis, ligation (e.g., blunt-end ligation or sticky-end ligation), or overlap-extension polymerase chain reaction (PCR), among other methods. Such an adapter can ultimately be attached to a target nucleic acid molecule. In other embodiments, one or more rounds of amplification cycles (e.g., PCR amplification) are commonly applied to introduce a sample index into a nucleic acid molecule using conventional nucleic acid amplification methods. Amplification can be performed in one or more reaction mixtures (e.g., multiple microwells in an array). The molecular barcode and / or sample index can be introduced simultaneously or in any sequential order. In some embodiments, the molecular barcode and / or sample index are introduced before and / or after performing a sequence capture step. In some embodiments, only the molecular barcode is introduced before probe capture, and the sample index is introduced after performing the sequence capture step. In some embodiments, both the molecular barcode and the sample index are introduced before performing the probe-based capture step. In some embodiments, the sample index is introduced after performing the sequence capture step. In some embodiments, the molecular barcode is incorporated into the nucleic acid molecule (e.g., cfDNA molecule) in the sample via an adapter via ligation (e.g., blunt-end ligation or cohesive-end ligation). In some embodiments, the sample index is incorporated into the nucleic acid molecule (e.g., cfDNA molecule) in the sample via overlap-extension polymerase chain reaction (PCR). Typically, the sequence capture protocol involves introducing a single-stranded nucleic acid molecule complementary to a targeted nucleic acid sequence, e.g., a coding sequence of a genomic region, where mutations in such region are associated with cancer type.
[0127] In some embodiments, tag can be located at one end or both ends of sample nucleic acid molecule.In some embodiments, tag is a predetermined or random or semi-random sequence oligonucleotide.In some embodiments, tag can be less than about 500, 200, 100, 50, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2 or 1 nucleotide in length.Tag can be linked to sample nucleic acid randomly or non-randomly.
[0128] In some embodiments, each sample is uniquely tagged with a sample index or a combination of sample indexes. In some embodiments, each nucleic acid molecule of a sample or sub-sample is uniquely tagged with a molecular barcode or a combination of molecular barcodes. In other embodiments, multiple molecular barcodes may be used such that the molecular barcodes are not necessarily unique to one another among the multiple molecular barcodes (e.g., non-unique molecular barcodes). In these embodiments, generally, molecular barcodes are attached to individual molecules (e.g., by ligation) such that the combination of the molecular barcode and the sequence to which it can be attached creates a unique sequence that can be individually tracked. Typically, detection of the non-unique molecular barcode in combination with endogenous sequence information (e.g., the start (start) and / or end (end) genomic location / position corresponding to the sequence of the original nucleic acid molecule in the sample, the start and end genomic positions corresponding to the sequence of the original nucleic acid molecule in the sample, the start (start) and / or end (end) genomic location / position of the sequence read mapped to the reference sequence, the start and end genomic positions of the sequence read mapped to the reference sequence, the subsequence of the sequence read at one or both ends, the length of the sequence read, and / or the length of the original nucleic acid molecule in the sample) allows assignment of a unique identity to a particular molecule. In some embodiments, the start region comprises the first 1, first 2, first 5, first 10, first 15, first 20, first 25, first 30, or at least the first 30 base positions at the 5' end of the sequencing read that aligns with the reference sequence. In some embodiments, the end region comprises the last 1, last 2, last 5, last 10, last 15, last 20, last 25, last 30, or at least last 30 base positions at the 3' end of the sequencing read that aligns with the reference sequence. The length or number of base pairs of each sequence read is also used, as needed, to assign a unique identity to a given molecule. As described herein, fragments from a single strand of nucleic acid that have been assigned a unique identity can subsequently be identified from the parental strand and / or complementary strand.
[0129] In certain embodiments, the number of different tags used to uniquely identify a number z of molecules in a class is 2 * z, 3 * z, 4 * z, 5 * z, 6 * z, 7 * z, 8 * z, 9 * z, 10 * z, 11 * z, 12 * z, 13 * z, 14 * z, 15 * z, 16 * z, 17 * z, 18*z, 19 * z, 20 * z or 100 * Any of z (for example, the lower limit) and 100,000*z, 10,000 * z, 1000 * z or 100 * z can be anywhere between 0 and 1 (e.g., an upper limit). In some embodiments, molecular barcodes are introduced into a sample at an expected ratio of identifier sets (e.g., combinations of unique or non-unique molecular barcodes) to molecules. One exemplary format uses about 2 to about 1,000,000 different molecular barcode sequences, or about 5 to about 150 different molecular barcode sequences, or about 20 to about 50 different molecular barcode sequences ligated to both ends of a target molecule. Alternatively, about 25 to about 1,000,000 different molecular barcode sequences can be used. For example, 20-50 x 20-50 molecular barcode sequences (i.e., one of 20-50 different molecular barcode sequences can be attached to each end of a target molecule) can be used. Such a number of identifiers is typically sufficient so that different molecules with the same start and end points have a high probability (e.g., at least 94%, 99.5%, 99.99%, or 99.999%) of receiving different combinations of identifiers. In some embodiments, about 80%, about 90%, about 95%, or about 99% of the molecules have the same combination of molecular barcodes.
[0130] In some embodiments, the assignment of unique or non-unique molecular barcodes in the reactions is performed using methods and systems described, for example, in U.S. Patent Application Nos. 20010053519, 20030152490, and 20110160078, and U.S. Patent Nos. 6,582,908, 7,537,898, 9,598,731, and 9,902,992, each of which is hereby incorporated by reference in its entirety. Alternatively, in some embodiments, different nucleic acid molecules of a sample can be identified using only intrinsic sequence information (e.g., start and / or stop positions, subsequences at one or both ends of the sequence, and / or length).
[0131] In certain embodiments described herein, populations of different forms of nucleic acids (e.g., hypermethylated and hypomethylated DNA in a sample) can be physically partitioned prior to analysis, e.g., sequencing, or tagging and sequencing. This approach can be used, for example, to determine whether hypermethylated variable epigenetic target regions exhibit hypermethylation characteristic of tumor cells, or whether hypomethylated variable epigenetic target regions exhibit hypomethylation characteristic of tumor cells. In addition, partitioning heterogeneous nucleic acid populations can increase rare signals, for example, by enriching for rare nucleic acid molecules that are more abundant in one fraction (or fractions) of the population. For example, genetic variations that are present in hypermethylated DNA but less abundant (or absent) in hypomethylated DNA can be more easily detected by partitioning the sample into hypermethylated and hypomethylated nucleic acid molecules. By analyzing multiple fractions of a sample, multidimensional analysis of a single genomic locus or nucleic acid species can be performed, thus achieving greater sensitivity.
[0132] In some examples, a heterogeneous nucleic acid sample is divided into two or more fractions (e.g., at least 3, 4, 5, 6, or 7 fractions). In some embodiments, each fraction is differentially tagged. That is, each fraction can have a different set of molecular barcodes. The tagged fractions can then be pooled together for collective sample preparation and / or sequencing. The dividing-tagging-pooling steps can be performed more than once, with each round of dividing being based on a different feature (examples provided herein) and tagged using a differential tag that is distinct from the other fractions and dividing means.
[0133] Examples of characteristics that can be used for partitioning include sequence length, methylation level, nucleosome binding, sequence mismatch, immunoprecipitation, and / or proteins that bind to DNA. The resulting fractions can contain one or more of the following nucleic acid forms: single-stranded DNA (ssDNA), double-stranded DNA (dsDNA), short DNA fragments, and long DNA fragments. In some embodiments, a heterogeneous population of nucleic acids is partitioned into nucleic acids with one or more epigenetic modifications and nucleic acids without one or more epigenetic modifications. Examples of epigenetic modifications include the presence or absence of methylation; the level of methylation; the type of methylation (e.g., 5-methylcytosine versus other types of methylation, such as adenine methylation and / or cytosine hydroxymethylation); and the level of association with one or more proteins, such as histones. Alternatively or additionally, a heterogeneous population of nucleic acids can be partitioned into nucleic acid molecules associated with nucleosomes and nucleic acid molecules lacking nucleosomes. Alternatively or additionally, the heterogeneous population of nucleic acids may be distributed between single-stranded DNA (ssDNA) and double-stranded DNA (dsDNA). Alternatively or additionally, the heterogeneous population of nucleic acids may be distributed based on nucleic acid length (e.g., molecules up to 160 bp and molecules having a length greater than 160 bp).
[0134] In some examples, each fraction (representing a different nucleic acid form) is differentially tagged with a molecular barcode, and the fractions are pooled together before sequencing. In other examples, the different forms are sequenced separately. In some embodiments, a single tag may be used to label a specific fraction. In some embodiments, multiple different tags may be used to label a specific fraction. In embodiments using multiple different tags to label specific fractions, the set of tags used to label one fraction can be easily distinguished from the set of tags used to label other fractions. In some embodiments, the tag may be multifunctional. That is, the tag can simultaneously act as a molecular identifier (i.e., molecular barcode), a fraction identifier (i.e., fraction tag), and a sample identifier (i.e., sample index). For example, if there are four DNA samples and each DNA sample is divided into three fractions, the DNA molecules in each of the 12 fractions (i.e., 12 fractions for the four DNA samples in total) can be tagged with a separate set of tags, such that the tag sequence attached to the DNA molecule reveals the identity of the DNA molecule, the fraction to which it belongs, and the sample from which it originated. In some embodiments, tags can be used both as molecular barcodes and as fraction tags. For example, if a DNA sample is divided into three fractions, the DNA molecules in each fraction are tagged with a separate set of tags, such that the tag sequence attached to the DNA molecule reveals the identity of the DNA molecule and the fraction to which it belongs. In some embodiments, tags can be used both as molecular barcodes and as sample indexes. For example, if there are four DNA samples, the DNA molecules in each sample are tagged with a separate set of tags that can be distinguished from each sample, such that the tag sequence attached to the DNA molecule serves as both a molecular identifier and a sample identifier.
[0135] In one embodiment, fraction tagging comprises tagging molecules in each fraction with a fraction tag. After the fractions are recombined and the molecules are sequenced, the fraction tag identifies the source fraction. In another embodiment, different fractions are tagged with different sets of molecular tags, such as barcode pairs. In this way, each molecular barcode is useful for distinguishing source fractions and molecules within the fraction. For example, a first set of 35 barcodes can be used to tag molecules in a first fraction, while a second set of 35 barcodes can be used to tag molecules in a second fraction.
[0136] In some embodiments, after partitioning and tagging with fraction tags, the molecules may be pooled for sequencing in a single run. In some embodiments, sample tags are added to the molecules, for example, in a step after adding fraction tags and pooling. Sample tags can facilitate pooling materials generated from multiple samples for sequencing in a single sequencing run.
[0137] Alternatively, in some embodiments, fraction tag can be associated with sample and fraction.As a simple example, the first tag can represent the first fraction of the first sample; the second tag can represent the second fraction of the first sample; the third tag can represent the first fraction of the second sample; and the fourth tag can represent the second fraction of the second sample.
[0138] Tags may be attached to molecules that have already been sorted based on one or more epigenetic features, but the final tagged molecules in the library may no longer possess those epigenetic features. For example, single-stranded DNA molecules may be sorted and tagged, but the final tagged molecules in the library will likely be double-stranded. Similarly, DNA may be fractionated based on different levels of methylation, but the tagged molecules derived from these molecules in the final library will likely be unmethylated. Thus, tags attached to molecules in the library typically represent the characteristics of the "parent molecules" from which the final tagged molecules are derived, and not necessarily the characteristics of the tagged molecules themselves.
[0139] For example, use barcode 1, 2, 3, 4 etc. to tag and label the molecules in the first fraction; use barcode A, B, C, D etc. to tag and label the molecules in the second fraction; and use barcode a, b, c, d etc. to tag and label the molecules in the third fraction.Differentially tagged fractions can be pooled before sequencing.Differentially tagged fractions can be sequenced separately, or can be sequenced together simultaneously, for example, in the same flow cell of Illumina sequencer.
[0140] After sequencing, analysis of reads to detect genetic variants can be performed at the level of each fraction and at the level of the entire nucleic acid population. Tags are used to sort reads from different fractions. Analysis can include in silico analysis to determine genetic and epigenetic variations (one or more of methylation, chromatin structure, etc.) using sequence information, length of genome coordinates, coverage and / or copy number. In some embodiments, higher coverage can be correlated with higher nucleosome occupancy in genome regions, while lower coverage can be correlated with lower nucleosome occupancy or nucleosome-depleted regions (NDRs).
[0141] C. Amplification The sample nucleic acid may be flanked by adapters and can be amplified by PCR and other amplification methods using nucleic acid primers that bind to primer binding sites in the adapters adjacent to the DNA molecules to be amplified. In some embodiments, the amplification method involves cycles of extension, denaturation, and annealing resulting from thermal cycling, or may be isothermal, as in transcription-mediated amplification, for example. Other examples of amplification methods that can be used as needed include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and sequence-based self-sustained replication.
[0142] Typically, the amplification reaction generates multiple nucleic acid amplicons non-uniquely or uniquely tagged with molecular barcodes and sample indices ranging in size from about 150 nucleotides (nt) to about 700 nt, 250 nt to about 350 nt, or about 320 nt to about 550 nt. In some embodiments, the amplicons have a size of about 180 nt. In some embodiments, the amplicons have a size of about 200 nt.
[0143] D. Enrichment / Capture In some embodiments, sequences are enriched before sequencing nucleic acids. Enrichment is performed for specific target regions or non-specifically ("target sequences") as needed. In some embodiments, the targeted regions of interest can be enriched / captured using nucleic acid capture probes ("baits") selected for one or more bait set panels using differential tiling and capture schemes. Differential tiling and capture schemes generally use bait sets with different relative concentrations to differentially tile (e.g., at different "resolutions") across the genomic regions associated with the baits, subject to a set of constraints (e.g., sequencer constraints, e.g., sequencing load, availability of each bait, etc.), to capture the targeted nucleic acids at a desired level for downstream sequencing. These targeted genomic regions of interest optionally include natural or synthetic nucleotide sequences of nucleic acid constructs. In some embodiments, biotin-labeled beads bearing probes for one or more regions of interest can be used to capture target sequences and, optionally, subsequently used to amplify those regions to enrich for the regions of interest.
[0144] Sequence capture typically involves the use of oligonucleotide probes that hybridize to target nucleic acid sequences. In some embodiments, the probe set strategy involves tiling probes across a region of interest. Such probes can be, for example, about 60 to about 120 nucleotides in length. The set can have a depth (e.g., depth of coverage) of about 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, 10x, 15x, 20x, 50x, or greater than 50x. The effectiveness of sequence capture generally depends in part on the length of the sequence in the target molecule that is complementary (or nearly complementary) to the sequence of the probe.
[0145] In some embodiments, the enriched DNA molecules (or captured set) may contain DNA corresponding to a set of sequence variable target regions and a set of epigenetic target regions. In some embodiments, the amount of captured sequence variable target region DNA is greater than the amount of captured epigenetic target region DNA when normalized for differences in the size (footprint size) of the targeted regions. In some embodiments, the compositions, methods, and systems described in PCT Patent Application No. PCT / US2020 / 016120, the entire contents of which are hereby incorporated by reference herein.
[0146] Alternatively, first and second captured sets may be provided, each containing DNA corresponding to the set of sequence variable target regions and DNA corresponding to the set of epigenetic target regions, and the first and second captured sets may be combined to provide a combined captured set.
[0147] In captured sets that include DNA corresponding to a sequence variable target region set and an epigenetic target region set, including the combined captured sets discussed above, the DNA corresponding to the sequence variable target region set is present at a higher concentration than the DNA corresponding to the epigenetic target region set, e.g., 1.1-1.2 fold higher, 1.2-1.4 fold higher, 1.4-1.6 fold higher, 1.6-1.8 fold higher, 1.8-2.0 fold higher, 2.0-2.2 fold higher, 2.2-2.4 fold higher, 2.4-2.6 fold higher, 2.6-2.8 fold higher, 2.8-3.0 fold higher, 3.0-3.5 fold higher, 3.1-3.2 fold higher, 3.2-3.4 fold higher, 3.3-3.5 fold higher, 3.4-3.6 fold higher, 3.5-3.6 fold higher, 3.6-3.8 fold higher, 3.7-3.8 fold higher, 3.8-3.9 fold higher, 3.9 ... The compound may be present at a concentration of 0.5 to 4.0, 4.0 to 4.5 times higher, 4.5 to 5.0 times higher, 5.0 to 5.5 times higher, 5.5 to 6.0 times higher, 6.0 to 6.5 times higher, 6.5 to 7.0 times higher, 7.0 to 7.5 times higher, 7.5 to 8.0 times higher, 8.0 to 8.5 times higher, 8.5 to 9.0 times higher, 9.0 to 9.5 times higher, 9.5 to 10.0 times higher, 10 to 11 times higher, 11 to 12 times higher, 12 to 13 times higher, 13 to 14 times higher, 14 to 15 times higher, 15 to 16 times higher, 16 to 17 times higher, 17 to 18 times higher, 18 to 19 times higher, or 19 to 20 times higher. The degree of density difference accounts for normalization with respect to the footprint size of the target region, as discussed in the definitions section.
[0148] i. Epigenetic target region set The epigenetic target region set may include one or more types of target regions that may distinguish DNA from neoplastic (e.g., tumor or cancer) cells from DNA from healthy cells, such as non-neoplastic circulating cells. Exemplary types of such regions are discussed in detail herein. In some embodiments, the method according to the present disclosure includes determining whether cfDNA molecules corresponding to the epigenetic target region set contain or exhibit cancer-related epigenetic modifications (e.g., hypermethylation in one or more hypermethylated variable target regions; one or more perturbations of CTCF binding; and / or one or more perturbations of transcription start sites) and / or copy number variations (e.g., local amplification). The epigenetic target region set may also include one or more control regions, for example, as described herein.
[0149] In some embodiments, the set of epigenetic target regions has a footprint of at least 100 kb, e.g., at least 200 kb, at least 300 kb, or at least 400 kb. In some embodiments, the set of epigenetic target regions has a footprint in the range of 100 to 1,000 kb, e.g., 100 to 200 kb, 200 to 300 kb, 300 to 400 kb, 400 to 500 kb, 500 to 600 kb, 600 to 700 kb, 700 to 800 kb, 800 to 900 kb, and 900 to 1,000 kb.
[0150] 1. Hypermethylated variable target regions In some embodiments, the epigenetic target region set comprises one or more hypermethylated variable target regions.Generally, hypermethylated variable target regions refer to regions where the observed increase in methylation level indicates an increased likelihood that the sample (e.g., cfDNA sample) contains DNA produced by neoplastic cells, such as tumor or cancer cells.For example, hypermethylation of promoters of tumor suppressor genes has been repeatedly observed.For example, Kang et al. et al., Genome Biol. 18:53 (2017) and the references cited therein.
[0151] An extensive discussion of methylation variable target regions in colorectal cancer is provided in Lam et al., Biochim Biophys Acta. 1866:106-20 (2016). These include VIM, SEPT9, ITGA4, OSM4, GATA4, and NDRG4. An exemplary set of hypermethylated variable target regions containing genes or portions thereof based on studies of colorectal cancer (CRC) is provided in Table 1. Many of these genes likely have relevance to cancers other than colorectal cancer; for example, TP53 is widely recognized as a crucial tumor suppressor, and hypermethylation-based inactivation of this gene may be a common mechanism of tumorigenesis. [Table 1-1] [Table 1-2]
[0152] In some embodiments, the hypermethylated variable target region includes multiple genes or portions thereof set forth in Table 1, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the genes or portions thereof set forth in Table 1. For example, for each locus included as a target region, there may be one or more probes having hybridization sites that bind between the transcription start site of the gene and the stop codon (or the final stop codon for alternatively spliced genes). In some embodiments, the one or more probes bind within 300 bp, e.g., within 200 or 100 bp upstream and / or downstream of a gene or portion thereof set forth in Table 1.
[0153] Methylation variable target regions in various types of lung cancer are described, for example, in Ooki et al., Clin. Cancer Res. 23:7141-52 (2017); Belinksy, Annu. Rev. Physiol. 77:453-74 (2015); Hulbert et al., Clin. Cancer Res. 23:1998-2005 (2017);Shi et al., BMC Genomics 18:901 (2017);Schneider et al., BMC Cancer. 11:102 (2011);Lissa et al., Transl Lung Cancer Res 5(5):492-504 (2016);Skvortsova et al., Br. J. Cancer. 94(10):1492-1495 (2006);Kim et al., Cancer Res. 61:3419-3424 (2001);Furonaka et al., Pathology International 55:303-309 (2005);Gomes et al., Rev. Port. Pneumol. 20:20-30 (2014);Kim et al. al., Oncogene. 20:1765-70 (2001);Hopkins-Donaldson et al., Cell Death Differ. 10:356-64 (2003);Kikuchi et al., Clin. Cancer Res. 11:2954-61 (2005);Heller et al., Oncogene 25:959-968 (2006);Licchesi et al., Carcinogenesis. 29:895-904 (2008); Guo et al., Clin. Cancer Res. 10:7917-24 (2004); Palmisano et al., Cancer Res. 63:4620-4625 (2003); and Toyooka et al., Cancer Res. 61:4556-4560, (2001).
[0154] An exemplary set of hypermethylated variable target regions containing genes or portions thereof based on lung cancer studies is provided in Table 2. Many of these genes likely have relevance to cancers other than lung cancer; for example, Casp8 (caspase 8) is a key enzyme in programmed cell death, and hypermethylation-based inactivation of this gene may be a common tumorigenesis mechanism not limited to lung cancer. In addition, several genes appear in both Tables 1 and 2, indicating generality. [Table 2-1] [Table 2-2]
[0155] Any of the foregoing embodiments relating to target regions identified in Table 2 may be combined with any of the above embodiments relating to target regions identified in Table 1. In some embodiments, the hypermethylated variable target regions include multiple genes or portions thereof set forth in Table 1 or Table 2, e.g., at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the genes or portions thereof set forth in Table 1 or Table 2.
[0156] Additional hypermethylated target regions may be obtained, for example, from the Cancer Genome Atlas. Kang et al., Genome Biology 18:53 (2017) describe the construction of a probabilistic method called Cancer Locator using hypermethylated target regions from breast, colon, kidney, liver, and lung. In some embodiments, the hypermethylated target regions may be specific to one or more types of cancer. Thus, in some embodiments, the hypermethylated target regions include one, two, three, four, or five subsets of hypermethylated target regions that collectively exhibit hypermethylation in one, two, three, four, or five of breast cancer, colon cancer, kidney cancer, liver cancer, and lung cancer.
[0157] ii. Hypomethylated variable target regions Global hypomethylation is a phenomenon commonly observed in various cancers. For example, see Hon et al., Genome Res. 22:246-258 (2012) (breast cancer); Ehrlich, Epigenomics 1:239-259 (2009) (review article describing findings on hypomethylation in colon cancer, ovarian cancer, prostate cancer, leukemia, hepatocellular carcinoma, and cervical cancer). For example, regions such as repetitive elements, such as LINE1 elements, Alu elements, centromeric tandem repeats, paracentromeric tandem repeats, and satellite DNA, and intergenic regions that are normally methylated in healthy cells, may show reduced methylation in tumor cells. Thus, in some embodiments, the epigenetic target region set includes hypomethylated variable target regions, and the observed decrease in methylation level indicates an increased likelihood that the sample (e.g., cfDNA sample) contains DNA produced by neoplastic cells, such as tumor cells or cancer cells.
[0158] In some embodiments, the hypomethylated variable target region comprises a repetitive element and / or an intergenic region, hi some embodiments, the repetitive element comprises one, two, three, four, or five of a LINE1 element, an Alu element, a centromeric tandem repeat, a paracentromeric tandem repeat, and / or satellite DNA.
[0159] Exemplary specific genomic regions that exhibit cancer-associated hypomethylation include, for example, nucleotides 8403565-8953708 and 151104701-151106035 of human chromosome 1, according to the hg19 or hg38 human genome constructs. In some embodiments, the hypomethylated variable target region overlaps or includes one or both of these regions.
[0160] iii.CTCF binding region CTCF is a DNA-binding protein that contributes to chromatin organization and often co-localizes with cohesin. Disturbance of CTCF binding site has been reported in a variety of different cancers. For example, see Katainen et al., Nature Genetics, doi:10.1038 / ng.3335; Guo et al., Nat. Commun. 9:1520 (2018), published online on June 8, 2015. CTCF binding results in a recognizable pattern of cfDNA, which can be detected by sequencing, for example, through fragment length analysis. For example, details about sequencing-based fragment length analysis are provided in Snyder et al., Cell 164:57-68 (2016); WO2018 / 009723; and US Patent Application Publication No. 20170211143A1, each of which is incorporated herein by reference.
[0161] Thus, perturbation of CTCF binding leads to variations in the fragmentation pattern of cfDNA, and therefore, CTCF binding sites represent one type of fragmentation variable target region.
[0162] There are many known CTCF binding sites, see, for example, CTCFBSDB (CTCF Binding Site Database), available online at insulatordb.uthsc.edu / ; Cuddapah et al., Genome Res. 19:24-32 (2009); Martin et al., See Nat. Struct. Mol. Biol. 18:708-14 (2011); Rhee et al., Cell. 147:1408-19 (2011), each of which is incorporated herein by reference. Exemplary CTCF binding sites are nucleotides 56014955-56016161 on chromosome 8 and nucleotides 95359169-95360473 on chromosome 13, according to the hg19 or hg38 human genome construct.
[0163] Thus, in some embodiments, the set of epigenetic target regions comprises CTCF binding regions, hi some embodiments, the CTCF binding regions comprise at least 10, 20, 50, 100, 200, or 500 CTCF binding regions, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 CTCF binding regions, such as those listed above or in the CTCFBSDB or one or more of the above-cited articles by Cuddapah et al., Martin et al., or Rhee et al.
[0164] In some embodiments, at least a portion of the CTCF sites can be methylated or unmethylated, and the methylation status correlates with whether the cell is a cancer cell. In some embodiments, the set of epigenetic target regions includes regions at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and / or downstream of the CTCF binding site.
[0165] iv. Transcription start site Transcription start site can also show perturbation in neoplastic cell.For example, the nucleosome organization at various transcription start sites in healthy cells of hematopoietic lineage contributes substantially to cfDNA in healthy individuals, but can be different from the nucleosome organization at those transcription start sites in neoplastic cell.This results in different cfDNA patterns, which can be detected by sequencing, for example, as generally discussed in Snyder et al., Cell 164:57-68 (2016); WO2018 / 009723; and US Patent Application Publication No. 20170211143A1.
[0166] Thus, perturbations in transcription start sites also result in variations in cfDNA fragmentation patterns, and therefore represent a type of variable fragmentation target region.
[0167] Human transcription start sites are available from DBTSS (DataBase of Human Transcription Start Sites), available on the internet at dbtss.hgc.jp, and are described in Yamashita et al., Nucleic Acids Res. 34(Database issue): D86-D89 (2006), which is incorporated herein by reference.
[0168] Thus, in some embodiments, the set of epigenetic target regions includes a transcription start site. In some embodiments, the transcription start sites include at least 10, 20, 50, 100, 200, or 500 transcription start sites, or 10-20, 20-50, 50-100, 100-200, 200-500, or 500-1000 transcription start sites, e.g., transcription start sites described in the DBTSS. In some embodiments, at least a portion of the transcription start sites can be methylated or unmethylated, and the methylation status correlates with whether the cell is a cancer cell. In some embodiments, the set of epigenetic target regions includes at least 100 bp, at least 200 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 750 bp, or at least 1000 bp upstream and / or downstream of the transcription start site.
[0169] v. Copy number variation; local amplification Copy number variations such as local amplification are somatic mutations, and they can be detected by sequencing based on read frequency in a similar manner to the approach used to detect certain epigenetic changes such as methylation changes.Therefore, the region that may show copy number variations such as local amplification in cancer can be included in the epigenetic target region set, and these regions can include one or more of AR, BRAF, CCND1, CCND2, CCNE1, CDK4, CDK6, EGFR, ERBB2, FGFR1, FGFR2, KIT, KRAS, MET, MYC, PDGFRA, PIK3CA, and RAF1.For example, in some embodiments, the epigenetic target region set includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, or 18 of the above-mentioned targets.
[0170] iv. Methylation control region It may be useful to include a control region to facilitate data validation. In some embodiments, the set of epigenetic target regions includes a control region that is expected to be methylated or unmethylated in essentially all samples, regardless of whether the DNA is derived from cancer cells or normal cells. In some embodiments, the set of epigenetic target regions includes a control hypomethylated region that is expected to be hypomethylated in essentially all samples. In some embodiments, the set of epigenetic target regions includes a control hypermethylated region that is expected to be hypermethylated in essentially all samples.
[0171] b. Set of sequence-variable target regions In some embodiments, the set of sequence variable target regions comprises a plurality of regions known to undergo somatic mutations in cancer (herein referred to as cancer-associated mutations). Thus, the method can comprise determining whether the cfDNA molecules corresponding to the set of sequence variable target regions comprise cancer-associated mutations.
[0172] In some embodiments, the sequence variable target region set targets a plurality of different genes or genomic regions ("panel") that are selected so that a predetermined proportion of subjects with cancer exhibits genetic variants or tumor markers in one or more different genes or genomic regions in the panel. The panel can be selected to limit the sequencing region to a fixed number of base pairs. The panel can be selected to sequence a desired amount of DNA, for example, by adjusting the affinity and / or amount of probes as described elsewhere herein. The panel can also be selected to achieve a desired depth of sequence reads. The panel can be selected to achieve a desired depth or coverage of sequence reads in terms of the amount of sequenced base pairs. The panel can be selected to achieve a theoretical sensitivity, theoretical specificity, and / or theoretical accuracy for detecting one or more genetic variants in a sample.
[0173] The probes for detecting the panel of regions can include probes for detecting genomic regions of interest (hotspot regions) and nucleosome recognition probes (e.g., KRAS codons 12 and 13), and can be designed to optimize capture based on the analysis of cfDNA coverage and fragment size variation, which are affected by nucleosome binding patterns and GC sequence composition.Regions as used herein can also include non-hotspot regions that are optimized based on nucleosome position and GC model.
[0174] Examples of lists of genomic locations of interest can be found in Tables 3 and 4. In some embodiments, the set of sequence variable target regions used in the methods of the disclosure includes at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the genes in Table 3. In some embodiments, the set of sequence variable target regions used in the methods of the disclosure includes at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, or 70 of the SNVs in Table 3. In some embodiments, the set of sequence variable target regions used in the methods of the disclosure includes at least 1, at least 2, at least 3, at least 4, at least 5, or 6 of the fusions in Table 3. In some embodiments, the set of sequence variable target regions used in the methods of the disclosure includes at least a portion of at least 1, at least 2, or 3 of the indels in Table 3. In some embodiments, the set of sequence variable target regions used in the methods of the disclosure includes at least a portion of at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the genes in Table 4. In some embodiments, the set of sequence variable target regions used in the methods of the disclosure includes at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, or 73 of the SNVs in Table 4.In some embodiments, the set of sequence variable target regions used in the methods of the disclosure includes at least one, at least two, at least three, at least four, at least five, or six of the fusions in Table 4. In some embodiments, the set of sequence variable target regions used in the methods of the disclosure includes at least a portion of at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, at least ten, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, or 18 of the indels in Table 4. Each of these genomic locations of interest can be identified as a scaffold region or a hotspot region for a given panel. An example list of hotspot genomic locations of interest can be found in Table 5. The coordinates in Table 5 are based on the hg19 assembly of the human genome, although one of skill in the art is familiar with other assemblies and can identify coordinate sets corresponding to the exons, introns, codons, etc. represented in the selected assembly. In some embodiments, the set of sequence variable target regions used in the methods of the present disclosure includes at least a portion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, or at least 20 of the genes in Table 5. Each hotspot genomic region is described along with several characteristics, including the associated gene, the chromosome on which it resides, the genomic start and end positions representing the locus, the base pair length of the locus, the exons covered by the gene, and important features (e.g., types of mutations) that a given genomic region of interest may seek to capture. [Table 3] [Table 4] [Table 5-1] [Table 5-2] [Table 5-3]
[0175] Additionally or alternatively, suitable target region set can be obtained from literature.For example, Gale et al., PLoS One 13: e0194630 (2018), which is incorporated herein by reference, describes a panel of 35 cancer-related gene targets that can be used as part or all of sequence variable target region set.These 35 targets are AKT1, ALK, BRAF, CCND1, CDK2A, CTNNB1, EGFR, ERBB2, ESR1, FGFR1, FGFR2, FGFR3, FOXL2, GATA3, GNA11, GNAQ, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MED12, MET, MYC, NFE2L2, NRAS, PDGFRA, PIK3CA, PPP2R1A, PTEN, RET, STK11, TP53 and U2AF1.
[0176] In some embodiments, the set of sequence variable target regions includes target regions from at least 10, 20, 30, or 35 genes associated with cancer, such as those genes associated with cancers listed above.
[0177] E. Sequencing The sample nucleic acid, optionally flanked by adapter, is generally subjected to sequencing, with or without prior amplification.Sequencing method or optionally used commercially available format includes, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing by synthesis, single molecule sequencing, nanopore-based sequencing, semiconductor sequencing, sequencing by ligation, sequencing by hybridization, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next-generation sequencing (NGS), single molecule sequencing by synthesis (SMSS) (Helicos), massively parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxam-Gilbert sequencing, primer walking, PacBio, SOLiD, Ion Torrent or sequencing by Nanopore platform. Sequencing reactions can be performed in a variety of sample processing units, which may include multiple lanes, multiple channels, multiple wells, or other means for processing multiple sets of samples substantially simultaneously. Sample processing units may also include multiple sample chambers capable of processing multiple runs simultaneously.
[0178] Sequencing reactions can be performed on one or more nucleic acid fragment types or regions known to contain cancer or other disease markers.Sequencing reactions can also be performed on any nucleic acid fragments present in a sample.Sequencing reactions can be performed on at least about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of the genome.In other examples, sequencing reactions can be performed on less than about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of the genome.
[0179] Simultaneous sequencing reaction can be carried out using multiplex sequencing method.In some embodiments, cell-free polynucleotide is sequenced by at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000 or 100,000 sequencing reactions.In other embodiments, cell-free polynucleotide is sequenced by less than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000 or 100,000 sequencing reactions.Sequencing reaction is typically carried out sequentially or simultaneously.Subsequent data analysis is generally carried out for all or part of sequencing reaction. In some embodiments, data analysis is performed on at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other embodiments, data analysis may be performed on less than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. Exemplary read depths are about 1000 to about 50,000 reads per locus (e.g., base position).
[0180] F. Analysis Sequencing can generate multiple sequencing reads or reads.Sequencing reads or reads can include sequences of nucleotide data that are less than about 150 bases in length or less than about 90 bases in length.In some embodiments, the reads are between about 80 bases and about 90 bases in length, for example, about 85 bases in length.In some embodiments, the methods of the present disclosure are applied to very short reads, for example, less than about 50 bases or about 30 bases in length.Sequencing read data can include sequence data and meta information.Sequence read data can be stored in any suitable file format, including, for example, VCF file, FASTA file, or FASTQ file.
[0181] FASTA can refer to a computer program for searching sequence databases, and the name FASTA can also refer to a standard file format. For example, FASTA is described in, for example, Pearson & Lipman, 1988, Improved tools for biological sequence comparison, PNAS 85:2444-2448, which is hereby incorporated by reference in its entirety. A sequence in FASTA format begins with a one-line description: Multiple lines of sequence data follow. Description lines are separated from the sequence data by a greater-than sign (>) in the first column. The word following the ">" sign is the sequence identifier, and the rest of the line is the description (both are optional). There should be no space between the ">" and the first character of the identifier. It is recommended that all lines of text be less than 80 characters long. A sequence ends when another line begins with ">", indicating the beginning of another sequence.
[0182] The FASTQ format is a text-based format for storing both biological sequences (usually nucleotide sequences) and their corresponding quality scores. It is similar to the FASTA format, but with the quality scores following the sequence data. For simplicity, both the sequence character and the quality score are coded by a single ASCII character. The FASTQ format is described, for example, in Cock et al. (“The Sanger FASTQ file format for sequences with quality scores,” which is hereby incorporated by reference in its entirety. and the Solexa / Illumina FASTQ variants,” Nucleic Acids Res 38(6):1767-1771, 2009).
[0183] For FASTA and FASTQ files, the meta-information includes a description line and does not include a line of sequence data. In some embodiments, for FASTQ files, the meta-information includes a quality score. For FASTA and FASTQ files, the sequence data begins after the description line and is typically present using a subset of IUPAC ambiguity codes, with "-" as needed. In one embodiment, the sequence data may use A, T, C, G, and N characters, optionally with "-" or U (e.g., representing a gap or uracil) as needed.
[0184] In some embodiments, at least one master sequence read file and output file are stored as plain text files (e.g., using coding such as ASCII, ISO / IEC 646, EBCDIC, UTF-8, or UTF-16). The computer system provided by the present disclosure may include a text editor program capable of opening plain text files. A text editor program may refer to a computer program that can present the contents of a text file (e.g., a plain text file) on a computer screen and allow a person to edit the text (e.g., using a monitor, keyboard, and mouse). Examples of text editors include, without limitation, Microsoft Word, emacs, pico, vi, BBEdit, and TextWrangler. A text editor program can display a plain text file on a computer screen and present meta information and sequence reads in a human-readable format (e.g., not using binary code, but instead using alphanumeric characters as used in printing or handwriting).
[0185] Although the methods are discussed with reference to FASTA or FASTQ files, the methods and systems of the present disclosure may be used to compress any suitable sequence file format, including, for example, files in Variant Call Format (VCF) format. A typical VCF file may include a header section and a data section. The header includes any number of meta-information lines, each beginning with the characters "##", and TAB-separated field definition lines, each beginning with a single "#" character. The field definition lines name eight required columns, and the body section contains lines of data that fill the columns defined by the field definition lines. The VCF format is described, for example, in Danecek et al. ("The variant call format and VCF tools," Bioinformatics 27(15):2156-2158, 2011), which is hereby incorporated by reference in its entirety. The header section The entries can be treated as meta information to be written to the compressed file, and the data sections can be treated as lines that are stored in the master file only if each is unique.
[0186] Some embodiments provide for the assembly of sequencing reads.For example, in the assembly by alignment, sequencing reads are aligned with each other or aligned with reference sequence.By aligning each read with reference genome in turn, all of the reads are positioned relative to each other to generate assembly.In addition, aligning or mapping sequencing reads with reference sequence can also be used to identify variant sequences in sequencing reads.Identifying variant sequences in combination with the methods and systems described herein can be used to further help in the diagnosis or prognosis of disease or condition, or to guide treatment decisions.
[0187] In some embodiments, any or all of the steps are automated. Alternatively, the disclosed methods may be embodied, in whole or in part, in one or more dedicated programs, for example, each written in a compiled language such as C++ as needed, and then compiled and distributed as a binary. The disclosed methods may also be implemented, in whole or in part, as modules within an existing sequence analysis platform or by invoking functions therein. In some embodiments, the disclosed methods include several steps that are all invoked automatically in response to a single initiating cue (e.g., one or a combination of triggering events caused by human activity, another computer program, or a machine). Thus, the present disclosure provides methods in which any or any combination of steps can occur automatically in response to a cue. "Automatically" generally means without intervening human input, influence, or interaction (e.g., responding only to human activity prior to the original or cue).
[0188] The disclosed methods may also include various forms of output, including accurate and sensitive interpretations of the subject's nucleic acid sample. The search output may be provided in a computer file format. In some embodiments, the output is a FASTA file, a FASTQ file, or a VCF file. The output may be processed to generate a text file or an XML file containing sequence data, for example, aligning the nucleic acid sequence to the sequence of a reference genome. In other embodiments, the processing results in an output containing coordinates or strings describing one or more mutations of the subject nucleic acid relative to the reference genome. Alignment strings may include Simple Ungapped Alignment Report (SUGAR), Verbose Useful Labeled Gapped Alignment Report (VULGAR), and Compact Idiosyncratic Gapped Alignment Report (CIGAR) (e.g., as described in Ning et al., Genome Research 11(10):1725-9, 2001, which is hereby incorporated by reference in its entirety). These strings can be implemented, for example, in the Exonerate sequence alignment software of the European Bioinformatics Institute (Hinxton, UK).
[0189] In some embodiments, a sequence alignment is generated, such as a sequence alignment map (SAM) or binary alignment map (BAM) file containing a CIGAR string (the SAM format is described, for example, in Li et al., "The Sequence Alignment / Map format and SAMtools," Bioinformatics, 25(16):2078-9, 2009, which is hereby incorporated by reference in its entirety). In some embodiments, the CIGAR exhibits or contains one gap alignment per line. The CIGAR is a condensed pairwise alignment format reported as a CIGAR string. The CIGAR string may be useful for representing long (e.g., genomic) pairwise alignments. The CIGAR string may be used in the SAM format to represent the alignment of a read to a reference genome sequence.
[0190] CIGAR strings may follow established motifs. Each letter is preceded by a number giving the base count of the event. Letters used may include M, I, D, N, and S (M=match, I=insertion, D=deletion, N=gap, S=substitution). CIGAR strings define a sequence of matches / mismatches and deletions (or gaps). For example, the CIGAR string 2MD3M2D2M may indicate that the alignment contains two matches, one deletion (the number 1 is omitted to save some space), three matches, two deletions, and two matches.
[0191] In some embodiments, a population of nucleic acids is prepared for sequencing by enzymatically creating blunt ends in double-stranded nucleic acids having single-stranded overhangs at one or both ends. In these embodiments, the population is typically treated with an enzyme having 5'-3' DNA polymerase activity and 3'-5' exonuclease activity in the presence of nucleotides (e.g., A, C, G, and T or U). Examples of enzymes or catalytic fragments thereof that can be used as needed include Klenow large fragment and T4 polymerase. For 5' overhangs, the enzyme typically extends the recessed 3' end on the opposing strand until it overlaps with the 5' end, creating a blunt end. For 3' overhangs, the enzyme generally digests from the 3' end to, and sometimes beyond, the 5' end of the opposing strand. If this digestion proceeds beyond the 5' end of the opposing strand, the gap can be filled in with the same enzyme with polymerase activity used for the 5' overhang. Creating blunt ends in double-stranded nucleic acids facilitates, for example, adapter binding and subsequent amplification.
[0192] In some embodiments, the population of nucleic acids is subjected to further processing, such as converting single-stranded nucleic acids to double-stranded nucleic acids and / or converting RNA to DNA (e.g., complementary DNA, or cDNA). These forms of nucleic acids are also optionally ligated to adapters and amplified.
[0193] With or without prior amplification, the nucleic acid that is subjected to the above-mentioned blunt-end forming process, and other nucleic acids in sample as needed, can be sequenced to produce sequenced nucleic acid.Sequenced nucleic acid can be referred to as the sequence (for example, sequence information) of nucleic acid, or the nucleic acid whose sequence is determined.Sequencing can be carried out so that the sequence data of each nucleic acid molecule in sample is provided directly or indirectly from the consensus sequence of the amplification product of each nucleic acid molecule in sample.
[0194] In some embodiments, double-stranded nucleic acids with single-stranded overhangs in the sample after blunt-end formation are ligated at both ends to adapters containing barcodes, and sequencing determines the nucleic acid sequence and the in-line barcode introduced by the adapter. Blunt-ended DNA molecules are optionally ligated to the blunt ends of at least partially double-stranded adapters (e.g., Y-shaped or bell-shaped adapters). Alternatively, the blunt ends of the sample nucleic acid and adapter can be tailed with complementary nucleotides to facilitate ligation (e.g., for cohesive end ligation).
[0195] A nucleic acid sample is typically contacted with a sufficient number of adapters so that the probability that any two copies of the same nucleic acid will receive the same adapter barcode combination from adapters ligated to both ends is low (e.g., less than about 1 or 0.1%). Using adapters in this manner allows for the identification of families of nucleic acid sequences that have the same start and stop points on the reference nucleic acid and are ligated to the same barcode combination. Such families can represent the sequences of amplification products of nucleic acids in a sample before amplification. The sequences of family members can be compiled to derive a consensus nucleotide or complete consensus sequence of the nucleic acid molecules in the original sample modified by blunt-end formation and adapter ligation. In other words, a nucleotide occupying a specified position in a nucleic acid in a sample can be determined to be the consensus of the nucleotide occupying the corresponding position in the family member sequences. A family can include sequences of one or both strands of a double-stranded nucleic acid. If family members include sequences of both strands from a double-stranded nucleic acid, the sequence of one strand can be converted to its complement for the purpose of compiling the sequences to derive a consensus nucleotide or sequence. Some families contain only a single member sequence, in which case this sequence can be considered the sequence of the nucleic acid in the sample before amplification, or alternatively, families with only a single member sequence can be excluded from further analysis.
[0196] Nucleotide variations (e.g., SNVs or indels) in sequenced nucleic acids can be determined by comparing the sequenced nucleic acids with a reference sequence. The reference sequence is often a known sequence, such as a known full or partial genome sequence from a subject (e.g., the entire genome sequence of a human subject). The reference sequence can be, for example, hG19 or hG38. As described above, the sequenced nucleic acid can represent a sequence determined directly for a nucleic acid in a sample or a consensus sequence of an amplification product of such a nucleic acid. Comparison can be performed at one or more designated positions of the reference sequence. A subset of sequenced nucleic acids can be identified that contains positions corresponding to designated positions of the reference sequence when the respective sequences are maximally aligned. Within such a subset, it can be determined whether the sequenced nucleic acids, if any, contain nucleotide variations at designated positions, or, if necessary, contain reference nucleotides (e.g., identical to the reference sequence). If the number of sequenced nucleic acids in the subset containing a nucleotide variant exceeds a selected threshold, the variant nucleotide can be called at the designated position. The threshold may be, among other possibilities, a simple number of sequenced nucleic acids in the subset containing the nucleotide variant, such as at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, or a ratio of at least 0.5, 1, 2, 3, 4, 5, 10, 15, or 20, etc., of sequenced nucleic acids in the subset containing the nucleotide variant. Comparisons can be repeated for any designated position of interest in the reference sequence. Sometimes, comparisons can be performed for designated positions occupying at least about 20, 100, 200, or 300 contiguous positions of the reference sequence, such as about 20-500, or about 50-300 contiguous positions.
[0197] Further details regarding nucleic acid sequencing, including the formats and applications described herein, can be found in, for example, Levy et al., Annual Review of Genomics and Human Genetics, 17: 95-115 (2016); Liu et al., J. of Biomedicine and Biotechnology, Volume 2012, Article ID 251364:1-11 (2012); Voelkerding et al., Clinical Chem., 55: 641-658 (2009); MacLean et al., Nature Rev. Microbiol., 7: 287-296 (2009), Astier et al., J Am Chem Soc., 128(5):1705-10 (2006), U.S. Patent No. 6,210,891, U.S. Patent No. 6,258,568, U.S. Patent No. 6,833,246, U.S. Patent No. 7,115,400, U.S. Patent No. 6,969,488, U.S. Patent No. 5,912,148, U.S. Patent No. 6,130,073, U.S. Patent No. 7,169,560, U.S. Patent No. 7,282,337, U.S. Patent No. 7,482,120, U.S. Patent No. 7,501,245, U.S. Patent No. 6,818,395, U.S. Patent No. 6,911,345, U.S. Patent No. 7,501,245, U.S. Patent No. 7,329,492, U.S. Patent No. 7,170,050, U.S. Patent No. 7,302,146, U.S. Patent No. 7,313,308, and U.S. Patent No. 7,476,503.
[0198] IV. Computer Systems The disclosed method can be implemented using or with the aid of a computer system. For example, such a method can include the following steps: (i) adding a set of carrier nucleic acid molecules to a polynucleotide sample to generate a first sample, wherein the set of carrier nucleic acid molecules includes (a) at least one subset of unmethylated carrier nucleic acid molecules; and / or (b) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, and the unmethylated carrier nucleic acid molecules do not include a methylated nucleotide, and the methylated carrier nucleic acid molecules include one or more methylated nucleotides; (ii) capturing the first sample using a capture agent that selectively binds to methylated polynucleotides, (iii) process at least a portion of the distributed samples to generate processed samples, wherein the process comprises at least one of the following steps: (a) tagging, (b) amplifying, and (c) enriching molecules for specific regions of interest; (iv) sequence at least a portion of the processed samples to generate a set of sequencing reads; and (v) analyze at least a portion of the set of sequencing reads to detect the presence or absence of tumor. In this embodiment, the system comprises the components for adding carrier nucleic acid molecules, distributing, tagging, amplifying, enriching, and sequencing.
[0199] 7 illustrates a computer system 701 programmed or otherwise configured to implement the methods of the present disclosure. The computer system 701 can control various aspects of sample preparation, sequencing, and / or analysis. In some examples, the computer system 701 is configured to perform sample preparation and sample analysis, including nucleic acid sequencing.
[0200] The computer system 701 includes a central processing unit (CPU, also referred to herein as a "processor" and a "computer processor") 705, which may be a single-core or multi-core processor, or may include multiple processors for parallel processing. The computer system 701 also includes memory or memory locations 710 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 715 (e.g., a hard disk), a communication interface 720 (e.g., a network adapter) for communicating with one or more other systems, and peripheral devices 725, such as cache, other memory, data storage, and / or an electronic display adapter. The memory 710, the storage unit 715, the interface 720, and the peripheral devices 725 communicate with the CPU 705 through a communication network or bus (solid lines), such as a motherboard. The storage unit 715 may be a data storage unit (or data repository) for storing data. The computer system 701 may be operably coupled to a computer network 730 with the aid of the communication interface 720. The computer network 730 may be the Internet, an internet and / or an extranet, or an intranet and / or an extranet in communication with the Internet. The computer network 730 may, in some cases, be a telecommunications and / or data network. The computer network 730 may include one or more computer servers, which may enable distributed computing, such as cloud computing. The computer network 730 may, in some cases with the aid of the computer system 701, implement a peer-to-peer network, which may enable devices coupled to the computer system 701 to act as clients or servers.
[0201] CPU 705 may execute sequences of machine-readable instructions, which may be embodied in a program or software. The instructions may be stored in memory locations, such as memory 710. Examples of operations performed by CPU 705 may include fetch, decode, execute, and writeback.
[0202] The storage unit 715 can store files, such as drivers, libraries, and saved programs. The storage unit 715 can store user-generated programs and recorded sessions, as well as program-related output. The storage unit 715 can store user data, such as user preferences and user programs. The computer system 701 in some cases can include one or more additional data storage units, such as those located external to the computer system 701, for example, on a remote server that communicates with the computer system 701 through an intranet or the Internet. Data may be transferred from one location to another, for example, using a communications network or physical data transfer (e.g., using a hard drive, thumb drive, or other data storage mechanism).
[0203] Computer system 701 can communicate with one or more remote computer systems through network 730. For an embodiment, computer system 701 can communicate with a remote computer system of a user (e.g., an operator). Examples of remote computer systems include a personal computer (e.g., a mobile PC), a slate or tablet PC (e.g., an Apple® iPad®, a Samsung® Galaxy Tab), a telephone, a smartphone (e.g., an Apple® iPhone®, an Android-enabled device, a Blackberry®), or a personal digital assistant. A user can access computer system 701 through network 730.
[0204] The methods described herein can be implemented by machine (e.g., a computer processor) executable code stored in an electronic storage location of the computer system 701, such as memory 710 or electronic storage unit 715. The machine-executable or machine-readable code can be provided in the form of software. During use, the code can be executed by the processor 705. In some cases, the code is retrieved from the storage unit 715 and stored in memory 710 for easy access by the processor 705. In some situations, the electronic storage unit 715 can be omitted, and the machine-executable instructions are stored in memory 710.
[0205] In certain aspects, the disclosure provides a method, when performed by at least one electronic processor, for generating a first sample, comprising the steps of: (i) adding a set of carrier nucleic acid molecules to a polynucleotide sample to generate a first sample, the set of carrier nucleic acid molecules comprising: (a) at least one subset of unmethylated carrier nucleic acid molecules; and / or (b) at least one subset of methylated carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, wherein the unmethylated carrier nucleic acid molecules comprise no methylated nucleotides and the methylated carrier nucleic acid molecules comprise one or more methylated nucleotides; (ii) capturing the first sample using a capture agent that selectively binds to methylated polynucleotides to generate a first sample; Provided is a non-transitory computer-readable medium that includes computer-executable instructions for carrying out a method that includes: (i) dividing the nucleic acid molecules into two divided sets, thereby generating divided samples; (ii) processing at least a portion of the divided samples to generate processed samples, and this processing includes at least one of the following steps: (a) tagging, (b) amplifying, and (c) enriching molecules for specific regions of interest; (iv) sequencing at least a portion of the processed samples to generate a set of sequencing reads; and (v) analyzing at least a portion of the set of sequencing reads to detect the presence or absence of tumor.In this embodiment, the computer-readable medium includes the computer-executable instructions required for adding carrier nucleic acid molecules, dividing, tagging, amplifying, enriching, and sequencing.
[0206] The code may be precompiled and configured for use on a machine having a processor adapted to execute the code, or may be compiled at run time. The code may be supplied precompiled or written in a programming language that may be selected to enable the code to be executed as it is compiled.
[0207] Aspects of the systems and methods provided herein, e.g., computer system 701, may be embodied in programming. Various aspects of the present technology can be thought of as "products" or "articles of manufacture," typically in the form of machine (or processor) executable code and / or associated data contained or embodied in a type of machine-readable medium. The machine-executable code can be stored in an electronic storage unit, e.g., memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. "Storage" type media includes any or all of a computer's tangible memory, processor, or the like, or its associated modules, e.g., various semiconductor memories, tape drives, disk drives, and the like, which may provide non-transitory storage at any time for software programming.
[0208] All or a portion of the software may sometimes be communicated over the Internet or various other telecommunications networks. Such communication may, for example, enable loading of the software from one computer or processor to another, e.g., from a management server or host computer to an application server computer platform. Accordingly, other types of media that may bear software elements include optical, electrical, and electromagnetic waves, such as those used across physical interfaces between local devices, through wired and optical terrestrial communications networks, and over various air links. Physical elements that carry such waves, e.g., wired or wireless links, optical links, or the like, may also be considered media bearing the software. As used herein, without limitation to non-transitory, tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that contributes to providing instructions to a processor for execution.
[0209] Thus, machine-readable media, e.g., computer-executable code, may take many forms, including, but not limited to, tangible storage media, carrier wave media, or physical transmission media. Non-volatile storage media include optical or magnetic disks, such as any of the storage devices of any computer, such as those used to implement the databases, etc., shown in the figures. Volatile storage media include dynamic memory, such as the main memory of such a computer platform. Tangible transmission media include coaxial cables, copper wire, and optical fibers, including the wires that comprise a bus in a computer system. Carrier-wave transmission media can take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer readable media thus include, for example, floppy disks, flexible disks, hard disks, magnetic tape, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punch cards, paper tape, any other physical storage media with a pattern of holes, RAM, ROM, PROMs and EPROMs, FLASH-EPROMs, any other memory chip or cartridge, a carrier wave transporting data or instructions, a cable or link transporting such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0210] The computer system 701 may include or be in communication with an electronic display that includes, for example, a user interface (UI) for providing one or more results of a sample analysis. Examples of UIs include, without limitation, graphical user interfaces (GUIs) and web-based user interfaces.
[0211] Further details regarding computer systems and networks, databases, and computer program products can be found, for example, in Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011), Kurose, Computer Networking: A Top-Down Approach, Pearson, 7 th Ed. (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010), Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11 th Ed. (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006), and Rhoton, Cloud It is also provided in Computing Architected: Solution Design Handbook, Recursive Press (2011).
[0212] V. Application A. Cancer and other diseases In some embodiments, the disease under consideration is a type of cancer. Non-limiting examples of such cancers include cholangiocarcinoma, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary nonpolyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms' tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLCL), chronic lymphocytic leukemia (CLCL), chronic lymphocytic leukemia (CRL), chronic lymphocytic leukemia (CL ... Leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin's lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin's lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal carcinoma, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm neoplasm), acinic cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, gastric cancer, gastrointestinal stromal tumor (GIST), uterine cancer or uterine sarcoma.
[0213] Non-limiting examples of other genetic diseases, disorders, or conditions that may be evaluated, if desired, using the methods and systems disclosed herein include achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), cricket cricket, Crohn's disease, cystic fibrosis, Dercum's disease, Down's syndrome, Duanne's syndrome, Duchenne muscular dystrophy, factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, and Goniohashi disease. Examples of conditions that may be present include: rhesus disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter's syndrome, Marfan's syndrome, myotonic dystrophy, neurofibroma, Noonan's syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland variant, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (SCID), sickle cell disease, spinal muscular atrophy, Tay-Sachs disease, thalassemia, trimethylaminuria, Turner syndrome, palatocardiofacial syndrome, WAGR syndrome, Wilson's disease, or the like.
[0214] B. Methods for determining the risk of cancer recurrence in a test subject and / or classifying test subjects as candidates for subsequent cancer treatment In some embodiments, the methods provided herein are methods for determining the risk of cancer recurrence in a test subject. In some embodiments, the methods provided herein are methods for classifying a test subject as a candidate for subsequent cancer treatment.
[0215] Any of these methods may include collecting DNA (e.g., originating from or derived from tumor cells) from a test subject diagnosed with cancer at one or more preselected time points after one or more previous cancer treatments for the test subject. The subject may be any of the subjects described herein. The DNA may be cfDNA. The DNA may be obtained from a tissue sample.
[0216] Any such method may include capturing a set of multiple target regions from DNA from a subject, the set of multiple target regions including a set of sequence variable target regions and a set of epigenetic target regions, to produce a set of captured DNA molecules. The capturing step may be performed according to any of the embodiments described elsewhere herein.
[0217] In any of such methods, the prior cancer treatment may include surgery, administration of a therapeutic composition, and / or chemotherapy.
[0218] Any of these methods includes sequencing the captured DNA molecules, thereby generating a set of sequence information. The captured DNA molecules of the set of sequence-variable target regions can be sequenced to a higher sequencing depth than the captured DNA molecules of the set of epigenetic target regions.
[0219] Any such method may include using the set of sequence information to detect the presence or absence of DNA originating from or derived from the tumor cell at a preselected time point. Detecting the presence or absence of DNA originating from or derived from the tumor cell may be performed according to any of the embodiments described elsewhere herein.
[0220] The method for determining the risk of cancer recurrence in a test subject can include determining a cancer recurrence score for the test subject, which indicates the presence or absence or amount of DNA originating from or derived from tumor cells.The cancer recurrence score can be further used to determine a cancer recurrence status.The cancer recurrence status can be, for example, a risk of cancer recurrence when the cancer recurrence score is above a predetermined threshold.The cancer recurrence status can be, for example, a low or lower risk of cancer recurrence when the cancer recurrence score is above a predetermined threshold.In certain embodiments, a cancer recurrence score equal to a predetermined threshold can result in a cancer recurrence status of a risk of cancer recurrence, or a low or lower risk of cancer recurrence.
[0221] A method for classifying a test subject as a candidate for subsequent cancer treatment includes comparing the test subject's cancer recurrence score with a predetermined cancer recurrence threshold, and classifying the test subject as a candidate for subsequent cancer treatment if the cancer recurrence score is above the cancer recurrence threshold, or as not a candidate for treatment if the cancer recurrence score is below the cancer recurrence threshold. In certain embodiments, a cancer recurrence score equal to the cancer recurrence threshold may result in classification as a candidate for subsequent cancer treatment or as not a candidate for treatment. In some embodiments, the subsequent cancer treatment includes administration of chemotherapy or a therapeutic composition.
[0222] Any such method may include determining a disease-free survival (DFS) period for the test subject based on the cancer recurrence score, for example, the DFS period may be 1 year, 2 years, 3 years, 4 years, 5 years, or 10 years.
[0223] In some embodiments, the set of sequence information comprises a sequence variable target region sequence, and determining the cancer recurrence score may comprise determining at least a first subscore indicative of the amount of SNV, insertion / deletion, CNV, and / or fusion present in the sequence variable target region sequence.
[0224] In some embodiments, the number of mutations in the sequence variable target regions selected from 1, 2, 3, 4, or 5 is sufficient to result in a cancer recurrence score in which the first subscore is classified as positive for cancer recurrence. In some embodiments, the number of mutations is selected from 1, 2, or 3.
[0225] In some embodiments, the set of sequence information includes epigenetic target region sequences, and determining the cancer recurrence score includes determining a second subscore indicative of an alteration in an epigenetic feature in the epigenetic target region sequences, e.g., methylation of a hypermethylated variable target region and / or perturbed fragmentation of a fragmented variable target region, where "perturbed" means different from the DNA found in a corresponding sample from a healthy subject.
[0226] In any embodiment in which the Cancer Recurrence Score is classified as positive for cancer recurrence, the subject's cancer recurrence status may be classified as being at risk for cancer recurrence and / or the subject may be classified as a candidate for subsequent cancer treatment.
[0227] In some embodiments, the cancer is any one of the types of cancer described elsewhere herein, for example, colorectal cancer.
[0228] C. Treatment and Related Administration In certain embodiments, the methods disclosed herein relate to identifying and administering customized therapies to patients given the status of nucleic acid variants of somatic or germline origin. In some embodiments, essentially any cancer therapy (e.g., surgery, radiation, chemotherapy, and / or the like) can be included as part of these methods. Typically, customized therapies include at least one immunotherapy (or immunotherapeutic agent). Immunotherapy generally refers to methods that enhance immune responses against a given type of cancer. In certain embodiments, immunotherapy refers to methods that enhance T-cell responses against tumors or cancers.
[0229] In certain embodiments, the nucleic acid variant status of the sample from the subject that is the origin of somatic or germline cell lineage is compared with the database of comparator results from a reference population, and identify the customized or targeted therapy for the subject.Typically, the reference population comprises patients with the same type of cancer or disease as the test subject, and / or patients who are undergoing or have undergone the same therapy as the test subject.If the nucleic acid variant and the comparator results meet certain classification criteria (for example, are substantially or approximately the same), customized or targeted therapy (one or more therapies) can be identified.
[0230] In certain embodiments, the customized therapies described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutic agents are typically administered intravenously. Certain therapeutic agents are administered orally. However, the customized therapies (e.g., immunotherapeutic agents, etc.) can also be administered by any method known in the art, including, for example, buccal, sublingual, rectal, vaginal, urethral, topical, ocular, nasal, and / or auricular, and administration can include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, salves, ointments, or the like. [Example]
[0231] Example 1: Distribution of cell-free DNA samples Two DNA samples (sample 1 and sample 2) obtained from the GM12878 cell line by extracting DNA released into the culture medium (i.e., DNA was extracted from the supernatant of the culture medium after the cells were separated by centrifugation) were analyzed here. In this example, each cell-free DNA sample was divided into two aliquots, each containing 10 ng of cell-free DNA. 140 ng of carrier DNA molecules was added to one aliquot of the sample. In this embodiment, the carrier DNA molecules used were double-stranded, with approximately 50% of the carrier DNA molecules being methylated and 50% of the carrier DNA molecules being unmethylated. Both the methylated and unmethylated carrier DNA molecules had the same sequence, and these carrier DNA molecules contained three CpG dinucleotides. All cytosines in the three CpG dinucleotides of the methylated carrier DNA molecules were methylated. No carrier DNA molecules were added to the other aliquot of the sample.
[0232] Two aliquots from Sample 1 and Sample 2 (with and without carrier DNA molecules) are combined with methyl-binding domain (MBD) buffer and magnetic beads conjugated to MBD protein (MethylMiner Methylated DNA Enrichment Kit (ThermoFisher Scientific)) and incubated overnight. During this incubation, the MBD protein binds to methylated DNA molecules (in the cell-free DNA sample and in the carrier DNA molecules, if present). Unmethylated or less methylated DNA is washed from the beads with buffers containing increasing salt concentrations. Finally, a high salt buffer is used to wash highly methylated DNA from the MBD protein. These washes result in three fractions of DNA with increasing methylation (three distributed sets). The three partitioned sets are divided into three groups: low, medium, and high. The partitioned DNA present in the partitioned sets includes DNA from the cell-free DNA sample and carrier DNA molecules. The partitioned DNA in the three partitioned sets is cleaned, desalted, and concentrated in preparation for the enzymatic steps of library preparation.
[0233] After enriching the DNA in the distributed sets, the terminal overhangs of the distributed DNA are extended, and adenosine residues are added to the 3' ends of the fragments. The 5' ends of each fragment are phosphorylated. These modifications make the distributed DNA ligable. DNA ligase and adapters are added to ligate each distributed DNA molecule to an adapter at each end. These adapters contain non-unique barcodes, and each distributed set is ligated to an adapter with a non-unique barcode that is distinguishable from the barcodes in the adapters used in the other distributed sets. After ligation, the three distributed sets are pooled together and amplified by PCR. Because the ends of the carrier DNA molecules are modified to prevent ligation, the adapters are not ligated to the carrier DNA molecules. Therefore, these molecules are not ligated to the adapters, and because the adapters contain primer binding sites, the carrier DNA molecules are not amplified.
[0234] After PCR, the amplified DNA is cleaned again and concentrated before enrichment.After enrichment, the amplified DNA is combined with salt buffer and the biotinylated RNA probe that targets the specific region of interest, and this mixture is incubated overnight.The biotinylated RNA probe is captured by streptavidin magnetic beads and separated from the amplified DNA that is not eluted by a series of salt washes, thereby enriching the sample.In this step, if residual carrier DNA molecules still exist in the sample, the carrier DNA molecules do not have sequence similarity to bind to the probe (i.e., the carrier DNA molecules have a sequence derived from non-human genome in this example), and the carrier DNA molecules are not captured, so the probe does not bind to the carrier DNA molecules, thereby ensuring that the carrier DNA molecules are not sequenced.
[0235] After enrichment, a sample index is incorporated into the enriched molecules via PCR amplification. After PCR amplification, the amplified molecules from different samples (within a batch) are pooled together and sequenced using an Illumina NovaSeq sequencer. The sequence reads generated by the sequencer are then analyzed using bioinformatics tools / algorithms. The analysis step includes determining the methylation status of the molecules. For example, specific regions of interest have previously been determined to be unmethylated in healthy individuals and methylated in individuals with malignant tumors. In this example, the analysis step includes determining whether the molecules are methylated in these regions of interest. This is determined based on the number of CpG residues in the molecules and the distributed set into which the molecules are distributed, which is then used to detect the presence or absence of tumors.
[0236] Figures 8A and 8B show graphical representations of the distribution of cell-free DNA molecules from a sample in the presence and absence of carrier DNA molecules. It has been found that the methylation status of certain regions in the human genome often fluctuates / does not change, always remaining the same or consistent with different subjects and / or different types / stages of disease. For example, a particular human genomic region can be either predominantly methylated or predominantly unmethylated, regardless of whether the subject has cancer or not. Figure 8A shows the percentage of cell-free DNA molecules in a particular human genomic region known to be predominantly unmethylated in a highly distributed set. Figure 8A clearly shows that unmethylated cell-free DNA molecules, which are unlikely to be in the highly distributed set, were distributed to the highly distributed set. Addition of carrier DNA molecules to the sample significantly reduced the amount of unmethylated cell-free DNA molecules distributed to the highly distributed set. For example, in Sample 1 without carrier DNA molecules, the percentage of cell-free DNA molecules for a particular human genomic region known to be predominantly unmethylated in the highly partitioned set was 0.37% (Sample 1 without carrier in Figure 8A), but this was reduced to 0.07% upon addition of carrier DNA molecules (Sample 1 with carrier in Figure 8A). Similarly, in Sample 2, the percentage of cell-free DNA molecules for a particular human genomic region known to be predominantly unmethylated in the highly partitioned set (in the absence of carrier DNA molecules) was 0.46% (Sample 2 without carrier in Figure 8A), but this was reduced to 0.06% upon addition of carrier DNA molecules (Sample 2 with carrier in Figure 8A). Figure 8B shows the distribution of cell-free DNA molecules with zero CpG dinucleotides in the highly partitioned set. Ideally, these molecules should not be partitioned into the highly partitioned set. Figure 8B clearly shows that approximately 0.15% and 0.22% of the cell-free DNA molecules with zero CpG dinucleotides in Sample 1 (Sample 1 without carrier in Figure 8B) and Sample 2 (Sample 2 without carrier in Figure 8B), respectively, are partitioned into highly partitioned sets.However, with the addition of carrier DNA molecules, the percentage of cell-free DNA molecules with zero CpG dinucleotides in the highly partitioned set was reduced to 0.014% for Sample 1 (Sample 1 with carrier in Figure 8B) and 0.01% for Sample 2 (Sample 2 with carrier in Figure 8B), respectively. Figure 8 clearly shows that the use of carrier DNA molecules increases the reliability of the partitioning assay by reducing mispartitioning of unmethylated DNA molecules, assay noise, and therefore improving the molecular specificity of the partitioning assay, which leads to improved clinical performance.
[0237] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present invention be limited by the specific examples provided herein. While the present invention has been described with reference to the above specification, the descriptions and explanations of the embodiments herein are not meant to be construed in a limiting sense. Numerous modifications, changes, and substitutions will occur to those skilled in the art without departing from the invention. Furthermore, it is to be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions set forth herein, which depend upon a variety of conditions and variables. It is to be understood that various alternatives to the disclosed embodiments described herein may be employed in practicing the invention. It is therefore intended that the present disclosure should embrace any such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of the invention, and that methods and structures within the scope of these claims and their equivalents are covered thereby.
[0238] Although the foregoing disclosure has been described in some detail by way of illustration and example for purposes of clarity and understanding, it will be apparent to those skilled in the art upon reading this disclosure that various changes in form and detail may be made therein without departing from the true scope of the disclosure and may be practiced within the purview of the appended claims. For example, all features, steps, elements, or other aspects of the methods, systems, computer-readable media, and / or components may be used in various combinations.
[0239] All patents, patent applications, websites, other publications and documents, accession numbers, and the like cited herein are incorporated by reference in their entirety for all purposes to the same extent as if each separate item were specifically and separately indicated to be incorporated by reference. Where different versions of a sequence are associated with accession numbers at different times, the version associated with the accession number as of the effective filing date of this application is meant. The effective filing date means the earlier of the actual filing date or, if applicable, the filing date of the priority application that references the accession number. Similarly, where different versions of a publication, website, or the like were published at different times, the version published closest to the effective filing date of this application is meant, unless otherwise indicated. The present invention provides, for example, the following items. (Item 1) 1. A method for detecting the presence or absence of a tumor in a subject, comprising: (i) obtaining a polynucleotide sample from said subject; (ii) adding a set of carrier nucleic acid molecules to the polynucleotide sample to generate a first sample, wherein the set of carrier nucleic acid molecules comprises: (a) at least a subset of unmethylated carrier nucleic acid molecules; and / or (b) at least a subset of methylated carrier nucleic acid molecules wherein at least one end of the carrier nucleic acid molecule is modified to prevent ligation, the unmethylated carrier nucleic acid molecule does not contain any methylated nucleotides, and the methylated carrier nucleic acid molecule contains one or more methylated nucleotides; (iii) partitioning the first sample into at least two partitioned sets using a capture agent that selectively binds to methylated polynucleotides, thereby generating partitioned samples; (iv) processing at least a portion of the distributed sample to produce a processed sample, the processing comprising at least one of the following: (a) tagging, (b) amplifying, and (c) enriching molecules for specific regions of interest; (v) sequencing at least a portion of the processed sample to obtain a sequencing readout. generating a set of codes; and (vi) analyzing at least a portion of the set of sequencing reads to detect the presence or absence of a tumor. (Item 2) 2. The method of claim 1, wherein the carrier nucleic acid molecule is between 25 bp and 325 bp in length. (Item 3) 2. The method of claim 1, wherein the first subset and the second subset of at least one subset of unmethylated carrier nucleic acid molecules comprise the same nucleotide sequence. (Item 4) 2. The method of claim 1, wherein the first subset and the second subset of at least one subset of unmethylated carrier nucleic acid molecules comprise different nucleotide sequences. (Item 5) 5. The method of claim 3 or 4, wherein the first subset and the second subset of at least one subset of the unmethylated carrier nucleic acid molecules comprise one or more CpG dinucleotides in their nucleotide sequences. (Item 6) 6. The method of claim 5, wherein the position of the one or more CpG dinucleotides in the first subset of at least one subset of unmethylated carrier nucleic acid molecules is different from the position of the one or more CpG dinucleotides in the second subset. (Item 7) 6. The method of claim 5, wherein the number of CpG dinucleotides in the first subset of at least one subset of unmethylated carrier nucleic acid molecules is different from the number of CpG dinucleotides in the second subset. (Item 8) 6. The method of claim 5, wherein the sequence of nucleotides adjacent to the one or more CpG dinucleotides in the first subset of at least one subset of unmethylated carrier nucleic acid molecules is different from the sequence of nucleotides adjacent to the one or more CpG dinucleotides in the second subset. (Item 9) 3. The method of claim 2, wherein the first and second subsets of at least one subset of unmethylated carrier nucleic acid molecules are of different lengths. (Item 10) 2. The method of claim 1, wherein the one or more methylated nucleotides are selected from the group consisting of: (i) 5-methylcytosine, (ii) 6-methyladenine, (iii) hydroxymethylcytosine, (iv) methyluracil, and (v) any other methylated nucleotide. (Item 11) 11. The method of claim 10, wherein the number of methylated nucleotides is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 15, 16, 17, 18, 19 or at least 20. (Item 12) 11. The method of claim 10, wherein the first subset and the second subset of at least one subset of methylated carrier nucleic acid molecules comprise the same nucleotide sequence. (Item 13) 11. The method of claim 10, wherein the first subset and the second subset of at least one subset of methylated carrier nucleic acid molecules comprise different nucleotide sequences. (Item 14) 14. The method of claim 12 or 13, wherein the first subset and the second subset of at least one subset of methylated carrier nucleic acid molecules comprise one or more CpG dinucleotides in their nucleotide sequences. (Item 15) 15. The method of claim 14, wherein the one or more CpG dinucleotides comprise one or more methylated cytosines. (Item 16) 14. The method of claim 12, wherein the positions of the one or more methylated nucleotides in the first subset of at least one subset of methylated carrier nucleic acid molecules are different from the positions of the one or more methylated nucleotides in the second subset. (Item 17) 14. The method of claim 12 or 13, wherein the number of methylated nucleotides in the first subset of at least one subset of methylated carrier nucleic acid molecules is different from the number of methylated nucleotides in the second subset. (Item 18) 14. The method of claim 13, wherein the sequence of nucleotides adjacent to the one or more methylated nucleotides in the first subset of at least one subset of methylated carrier nucleic acid molecules is different from the sequence of nucleotides adjacent to the one or more methylated nucleotides in the second subset. (Item 19) 11. The method of claim 10, wherein the first subset and the second subset of at least one subset of methylated carrier nucleic acid molecules are of different lengths. (Item 20) The amount of at least one subset of methylated carrier nucleic acid molecules relative to at least one subset of unmethylated carrier nucleic acid molecules is about 0:1, 0.1:99.9, 0.5:99.5, 0.75:99.25, 1:99, 1:95, 1:90, 1:80, 1:75, 1:70, 1:60, 1:50, 1:40, 1:30, 1:25, 1:20 , 1:10, 1:5, 1:2, 1:1.15, 1:1, 1.15:1, 2:1, 5:1, 10:1, 20:1, 25:1, 30:1, 40:1, 50:1, 60:1, 70:1, 75:1, 80:1, 90:1, 95:1, 99:1, 99.25:0.75, 99.5:0.5, 99.9:0.1 or 1:0 ratio. (Item 21) The amount of the polynucleotide sample relative to the set of carrier nucleic acid molecules is in a ratio of about 1:0.1; 1:0.2, 1:0.3, 1:4, 1:0.5, 1:6, 1:7, 1:8, 1:0.9, 1:1, 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:20, 1: 30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90, 1:100, 1:200, 1:300; 1:400, 1:500, 1:600, 1:700, 1:800, 1:900, 1:1000, 1:5000, 1:10,000, 1:100,000, 1:500,000, 1:10 6 , 1:10 7 , 1:10 8 or 1:10 9 Item 1. The method according to item 1, (Item 22) 2. The method of claim 1, wherein the polynucleotide sample is at most 1 μg. (Item 23) 2. The method of claim 1, wherein the polynucleotide sample is at most 200 ng. (Item 24) 2. The method of claim 1, wherein the polynucleotide sample is at most 150 ng. (Item 25) 2. The method of claim 1, wherein the polynucleotide sample is at most 100 ng. (Item 26) The set of carrier nucleic acid molecules is present in an amount of about 175 ng, 200 ng, 225 ng, 250 ng, 275 ng, 300 ng, 350 ng, 400 ng, 450 ng, 500 ng, 600 ng, 700 ng, 750 ng, 800 ng, 900 ng, 1000 ng, 1100 ng, 1200 ng, 1300 ng, 1400 ng, 1500 ng, 1600 ng, 1750 ng, 1800 ng, 1900 ng, 2150 ng, 2250 ng, 2300 ng, 2450 ng, 2500 ng, 2650 ng, 2750 ng, 26. The method of any one of items 22 to 25, wherein the antibody is added in an amount sufficient to provide 0 ng, 800 ng, 900 ng, 1 μg, 1.1 μg, 1.25 μg or 1.5 μg. (Item 27) 2. The method of claim 1, wherein the sequence of the carrier nucleic acid molecule is selected from the group consisting of: (i) a sequence derived from a viral genome, (ii) a sequence derived from a bacterial genome, (iii) a sequence derived from a lambda genome, and (iv) a sequence derived from a non-human genome. (Item 28) 2. The method of claim 1, wherein the carrier nucleic acid molecule is synthetic DNA. (Item 29) 2. The method of claim 1, wherein at least one end of the carrier nucleic acid molecule comprises a C3 (propyl group) spacer. (Item 30) 2. The method of claim 1, wherein at least one end of the carrier nucleic acid molecule comprises a dideoxynucleotide. (Item 31) 2. The method of claim 1, wherein at least one end of the carrier nucleic acid molecule comprises any chemical modification that prevents a hydroxyl group from acting as a nucleophile. (Item 32) 2. The method of claim 1, wherein the carrier nucleic acid molecule comprises a uracil nucleoside. (Item 33) 33. The method of claim 32, further comprising the step of adding uracil deglycosylase and DNA glycosylase-lyase before the amplifying step. (Item 34) 2. The method of claim 1, wherein the 5' end of the carrier nucleic acid molecule comprises at least one of the following modifications: (i) an inverted (5'-5'), such as dideoxythymine, dideoxycytosine, dideoxyguanine, or dideoxyadenine; (ii) a propyl group; or (iii) another organic functional group, such as, but not limited to, benzyl, ethyl, or methyl. (Item 35) 2. The method of claim 1, wherein the 3' end of the carrier nucleic acid molecule comprises at least one of the following modifications: (i) any dideoxy base that can be added enzymatically or during synthesis, such as dideoxythymine, dideoxycytosine, dideoxyguanine, or dideoxyadenine; (ii) a propyl group; or (iii) other organic functional groups, such as, but not limited to, benzyl, ethyl, or methyl. (Item 36) 2. The method of claim 1, wherein the polynucleotide sample is obtained from tissue, blood, plasma, serum, urine, saliva, feces, cerebrospinal fluid, buccal swab, or thoracentesis. (Item 37) 2. The method of claim 1, wherein the polynucleotide sample is obtained from a tissue. (Item 38) 38. The method of claim 37, wherein the polynucleotide sample obtained from the tissue is fragmented by enzymatic or mechanical means. (Item 39) 2. The method of claim 1, wherein the polynucleotide sample is obtained from blood. (Item 40) 40. The method of claim 39, wherein the blood-derived polynucleotide sample is a cell-free DNA sample. (Item 41) (i) at least one subset of unmethylated carrier nucleic acid molecules; and / or (ii) at least a subset of methylated carrier nucleic acid molecules a set of carrier nucleic acid molecules comprising: A set of carrier nucleic acid molecules, wherein at least one end of the carrier nucleic acid molecules is modified to prevent ligation, the unmethylated carrier nucleic acid molecules do not contain methylated nucleotides, and the methylated carrier nucleic acid molecules contain one or more methylated nucleotides. (Item 42) 42. The set of carrier nucleic acid molecules according to item 41, wherein the length of the carrier nucleic acid molecules is between 25 bp and 325 bp. (Item 43) 42. The set of carrier nucleic acid molecules of Item 41, wherein the first subset and the second subset of at least one subset of unmethylated carrier nucleic acid molecules comprise the same nucleotide sequence. (Item 44) 42. The set of carrier nucleic acid molecules of Item 41, wherein the first subset and the second subset of at least one subset of unmethylated carrier nucleic acid molecules comprise different nucleotide sequences. (Item 45) 45. The set of carrier nucleic acid molecules of item 43 or 44, wherein the first subset and the second subset of at least one subset of the unmethylated carrier nucleic acid molecules comprise one or more CpG dinucleotides in their nucleotide sequences. (Item 46) 46. The set of carrier nucleic acid molecules of item 45, wherein the position of the one or more CpG dinucleotides in the first subset of at least one subset of unmethylated carrier nucleic acid molecules is different from the position of the one or more CpG dinucleotides in the second subset. (Item 47) 46. The set of carrier nucleic acid molecules of item 45, wherein the number of CpG dinucleotides in the first subset of at least one subset of the unmethylated carrier nucleic acid molecules is different from the number of CpG dinucleotides in the second subset. (Item 48) 46. The set of carrier nucleic acid molecules of Item 45, wherein the sequence of nucleotides adjacent to the one or more CpG dinucleotides in the first subset of at least one subset of unmethylated carrier nucleic acid molecules is different from the sequence of nucleotides adjacent to the one or more CpG dinucleotides in the second subset. (Item 49) 42. The set of carrier nucleic acid molecules of Item 41, wherein the first subset and the second subset of at least one subset of unmethylated carrier nucleic acid molecules are of different lengths. (Item 50) 42. The set of carrier nucleic acid molecules of Item 41, wherein the one or more methylated nucleotides are selected from the group consisting of: (i) 5-methylcytosine, (ii) 6-methyladenine, (iii) hydroxymethylcytosine, (iv) methyluracil, and (v) any other methylated nucleotide. (Item 51) 51. The set of carrier nucleic acid molecules of item 50, wherein the number of methylated nucleotides is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 15, 16, 17, 18, 19, or at least 20. (Item 52) 51. The set of carrier nucleic acid molecules of Item 50, wherein the first subset and the second subset of at least one subset of methylated carrier nucleic acid molecules comprise the same nucleotide sequence. (Item 53) 51. The set of carrier nucleic acid molecules of Item 50, wherein the first subset and the second subset of at least one subset of methylated carrier nucleic acid molecules comprise different nucleotide sequences. (Item 54) 54. The set of carrier nucleic acid molecules of item 52 or 53, wherein the first subset and the second subset of at least one subset of methylated carrier nucleic acid molecules comprise one or more CpG dinucleotides in their nucleotide sequences. (Item 55) 55. The set of carrier nucleic acid molecules according to Item 54, wherein the one or more CpG dinucleotides comprise one or more methylated cytosines. (Item 56) 54. The set of carrier nucleic acid molecules of item 52 or 53, wherein the position of the one or more methylated nucleotides in the first subset of at least one subset of methylated carrier nucleic acid molecules is different from the position of the one or more methylated nucleotides in the second subset. (Item 57) 54. The set of carrier nucleic acid molecules of item 52 or 53, wherein the number of methylated nucleotides in the first subset of at least one subset of methylated carrier nucleic acid molecules is different from the number of methylated nucleotides in the second subset. (Item 58) 54. The set of carrier nucleic acid molecules of Item 53, wherein the sequence of nucleotides adjacent to the one or more methylated nucleotides in the first subset of at least one subset of methylated carrier nucleic acid molecules is different from the sequence of nucleotides adjacent to the one or more methylated nucleotides in the second subset. (Item 59) 51. The set of carrier nucleic acid molecules of Item 50, wherein the first subset and the second subset of at least one subset of methylated carrier nucleic acid molecules are of different lengths. (Item 60) The amount of at least one subset of methylated carrier nucleic acid molecules relative to at least one subset of unmethylated carrier nucleic acid molecules is about 0:1, 0.1:99.9, 0.5:99.5, 0.75:99.25, 1:99, 1:95, 1:90, 1:80, 1:75, 1:70, 1:60, 1:50, 1:40, 1:30, 1:25, 1:20, 1:1 51. The set of carrier nucleic acid molecules of item 50, wherein the ratio is 0, 1:5, 1:2, 1:1.15, 1:1, 1.15:1, 2:1, 5:1, 10:1, 20:1, 25:1, 30:1, 40:1, 50:1, 60:1, 70:1, 75:1, 80:1, 90:1, 95:1, 99:1, 99.25:0.75, 99.5:0.5, 99.9:0.1, or 1:0. (Item 61) 42. The set of carrier nucleic acid molecules according to Item 41, wherein the sequences of the carrier nucleic acid molecules are selected from the group consisting of: (i) a sequence derived from a viral genome, (ii) a sequence derived from a bacterial genome, (iii) a sequence derived from a lambda genome, and (iv) a sequence derived from a non-human genome. (Item 62) 42. The set of carrier nucleic acid molecules according to Item 41, wherein the carrier nucleic acid molecules are synthetic DNA. (Item 63) 42. The set of carrier nucleic acid molecules according to Item 41, wherein at least one end of the carrier nucleic acid molecule comprises a C3 (propyl group) spacer. (Item 64) 42. The set of carrier nucleic acid molecules according to Item 41, wherein at least one end of the carrier nucleic acid molecule comprises a dideoxynucleotide. (Item 65) 42. The set of carrier nucleic acid molecules according to Item 41, wherein at least one end of the carrier nucleic acid molecule comprises any chemical modification that prevents a hydroxyl group from acting as a nucleophile. (Item 66) 42. The set of carrier nucleic acid molecules according to Item 41, wherein the carrier nucleic acid molecules comprise uracil nucleosides. (Item 67) 67. The set of carrier nucleic acid molecules according to Item 66, further comprising the step of adding uracil deglycosylase and DNA glycosylase-lyase before the amplifying step. (Item 68) 42. The set of carrier nucleic acid molecules according to Item 41, wherein the 5'-end of the carrier nucleic acid molecule comprises at least one of the following modifications: (i) an inverted (5'-5')-dideoxythymine, dideoxycytosine, dideoxyguanine, or dideoxyadenine; (ii) a propyl group; or (iii) another organic functional group, such as, but not limited to, benzyl, ethyl, or methyl. (Item 69) 42. The set of carrier nucleic acid molecules according to Item 41, wherein the 3' end of the carrier nucleic acid molecule comprises at least one of the following modifications: (i) any dideoxy base that can be added enzymatically or during synthesis, such as dideoxythymine, dideoxycytosine, dideoxyguanine, or dideoxyadenine; (ii) a propyl group; or (iii) other organic functional groups, such as, but not limited to, benzyl, ethyl, or methyl. (Item 70) (i) (a) at least a subset of unmethylated carrier nucleic acid molecules; and / or (b) at least a subset of methylated carrier nucleic acid molecules a set of carrier nucleic acid molecules comprising: at least one end of said set of carrier nucleic acid molecules has been modified to prevent ligation, said unmethylated carrier nucleic acid molecules comprising no methylated nucleotides, and said methylated carrier nucleic acid molecules comprising one or more methylated nucleotides; (ii) a polynucleotide sample obtained from the subject; A population of nucleic acids comprising: (Item 71) 71. The population of nucleic acids of item 70, wherein the carrier nucleic acid molecules are between 25 bp and 325 bp in length. (Item 72) 71. The population of nucleic acids of Item 70, wherein the first subset and the second subset of at least one subset of unmethylated carrier nucleic acid molecules comprise the same nucleotide sequence. (Item 73) 71. The population of nucleic acids of Item 70, wherein the first subset and the second subset of at least one subset of unmethylated carrier nucleic acid molecules comprise different nucleotide sequences. (Item 74) 74. The population of nucleic acids of item 72 or 73, wherein the first subset and the second subset of at least one subset of the unmethylated carrier nucleic acid molecules comprise one or more CpG dinucleotides in their nucleotide sequences. (Item 75) 74. The population of nucleic acids of item 72 or 73, wherein the position of the one or more CpG dinucleotides in the first subset of at least one subset of the unmethylated carrier nucleic acid molecules is different from the position of the one or more CpG dinucleotides in the second subset. (Item 76) 74. The population of nucleic acids of item 72 or 73, wherein the number of CpG dinucleotides in the first subset of at least one subset of the unmethylated carrier nucleic acid molecules is different from the number of CpG dinucleotides in the second subset. (Item 77) 75. The population of nucleic acids of item 74, wherein the sequence of nucleotides adjacent to the one or more CpG dinucleotides in the first subset of at least one subset of unmethylated carrier nucleic acid molecules is different from the sequence of nucleotides adjacent to the one or more CpG dinucleotides in the second subset. (Item 78) 71. The population of nucleic acids of item 70, wherein the first subset and the second subset of at least one subset of unmethylated carrier nucleic acid molecules are of different lengths. (Item 79) 71. The population of nucleic acids of item 70, wherein the one or more methylated nucleotides are selected from the group consisting of: (i) 5-methylcytosine, (ii) 6-methyladenine, (iii) hydroxymethylcytosine, (iv) methyluracil, and (v) any other methylated nucleotide. (Item 80) 80. The population of nucleic acids of item 79, wherein the number of methylated nucleotides is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 12, 14, 15, 16, 17, 18, 19, or at least 20. (Item 81) 80. The population of nucleic acids of Item 79, wherein the first subset and the second subset of at least one subset of methylated carrier nucleic acid molecules comprise the same nucleotide sequence. (Item 82) 80. The population of nucleic acids of Item 79, wherein the first subset and the second subset of at least one subset of methylated carrier nucleic acid molecules comprise different nucleotide sequences. (Item 83) 83. The population of nucleic acids of item 81 or 82, wherein the first subset and the second subset of at least one subset of methylated carrier nucleic acid molecules comprise one or more CpG dinucleotides in their nucleotide sequences. (Item 84) 84. The population of nucleic acids of item 83, wherein the one or more CpG dinucleotides comprise one or more methylated cytosines. (Item 85) 83. The population of nucleic acids of item 81 or 82, wherein the position of the one or more methylated nucleotides in the first subset of at least one subset of the methylated carrier nucleic acid molecules is different from the position of the one or more methylated nucleotides in the second subset. (Item 86) 83. The population of nucleic acids of item 81 or 82, wherein the number of methylated nucleotides in the first subset of at least one subset of methylated carrier nucleic acid molecules is different from the number of methylated nucleotides in the second subset. (Item 87) 83. The population of nucleic acids of Item 82, wherein the sequence of nucleotides adjacent to the one or more methylated nucleotides in the first subset of at least one subset of methylated carrier nucleic acid molecules is different from the sequence of nucleotides adjacent to the one or more methylated nucleotides in the second subset. (Item 88) 80. The population of nucleic acids of Item 79, wherein the first subset and the second subset of at least one subset of methylated carrier nucleic acid molecules are of different lengths. (Item 89) The amount of at least one subset of methylated carrier nucleic acid molecules relative to at least one subset of unmethylated carrier nucleic acid molecules is about 0:1, 0.1:99.9, 0.5:99.5, 0.75:99.25, 1:99, 1:95, 1:90, 1:80, 1:75, 1:70, 1:60, 1:50, 1:40, 1:30, 1:25, 1:20, 1:30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90, 1:90, 1:1 ... 80:1, 90:1, 95:1, 99:1, 99.25:0.75, 99.5:0.5, 99.9:0.1, or 1:0 ratio. (Item 90) The amount of the polynucleotide sample relative to the set of carrier nucleic acid molecules is in a ratio of about 1:0.1; 1:0.2, 1:0.3, 1:4, 1:0.5, 1:6, 1:7, 1:8, 1:0.9, 1:1, 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:20, 1: 30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90, 1:100, 1:200, 1:300; 1:400, 1:500, 1:600, 1:700, 1:800, 1:900, 1:1000, 1:5000, 1:10,000, 1:100,000, 1:500,000, 1:10 6 , 1:10 7 , 1:10 8 or 1:10 9 71. The population of nucleic acids according to Item 70, wherein (Item 91) 71. The population of nucleic acids of item 70, wherein the polynucleotide sample is at most 1 μg. (Item 92) 71. The population of nucleic acids of item 70, wherein the polynucleotide sample is up to 200 ng. (Item 93) 71. The population of nucleic acids of item 70, wherein the polynucleotide sample is up to 150 ng. (Item 94) 71. The population of nucleic acids of item 70, wherein the polynucleotide sample is up to 100 ng. (Item 95) 95. The population of nucleic acids of any one of paragraphs 91 to 94, wherein the set of carrier nucleic acid molecules is added in an amount sufficient to provide a total amount of the polynucleotide sample and the set of carrier nucleic acid molecules of about 175 ng, 200 ng, 225 ng, 250 ng, 275 ng, 300 ng, 350 ng, 400 ng, 450 ng, 500 ng, 600 ng, 700 ng, 750 ng, 800 ng, 900 ng, 1 μg, 1.1 μg, 1.25 μg, or 1.5 μg. (Item 96) 71. The population of nucleic acids of claim 70, wherein the sequence of the carrier nucleic acid molecule is selected from the group consisting of: (i) a sequence derived from a viral genome, (ii) a sequence derived from a bacterial genome, (iii) a sequence derived from a lambda genome, and (iv) a sequence derived from a non-human genome. (Item 97) 71. The population of nucleic acids of item 70, wherein at least one end of the carrier nucleic acid molecule comprises a dideoxynucleotide. (Item 98) 71. The population of nucleic acids of item 70, wherein the at least one end of the carrier nucleic acid molecule comprises any chemical modification that prevents a hydroxyl group from acting as a nucleophile. (Item 99) 71. The population of nucleic acids of item 70, wherein the carrier nucleic acid molecules comprise uracil nucleosides. (Item 100) 100. The population of nucleic acids of item 99, further comprising the step of adding uracil deglycosylase and DNA glycosylase-lyase prior to the amplifying step. (Item 101) The 5' end of the carrier nucleic acid molecule may be modified with one of the following: (i) an inverted (5'-5')-dideoxy 71. The population of nucleic acids of item 70, comprising at least one of: thymine, dideoxycytosine, dideoxyguanine, or dideoxyadenine; (ii) a propyl group, or (iii) other organic functional groups, for example, but not limited to, benzyl, ethyl, or methyl. (Item 102) 71. The population of nucleic acids of item 70, wherein the 3' end of the carrier nucleic acid molecule comprises at least one of the following modifications: (i) any dideoxy base that can be added enzymatically or during synthesis, such as dideoxythymine, dideoxycytosine, dideoxyguanine, or dideoxyadenine; (ii) a propyl group, or (iii) other organic functional groups, such as, but not limited to, benzyl, ethyl, or methyl. (Item 103) 71. The population of nucleic acids of item 70, wherein the polynucleotide sample is obtained from tissue, blood, plasma, serum, urine, saliva, feces, cerebrospinal fluid, buccal swab, or thoracentesis. (Item 104) 71. The population of nucleic acids of item 70, wherein the polynucleotide sample is obtained from a tissue. (Item 105) 105. The population of nucleic acids of item 104, wherein the polynucleotide sample obtained from the tissue is fragmented by enzymatic or mechanical means. (Item 106) 71. The population of nucleic acids according to item 70, wherein the polynucleotide sample is obtained from blood. (Item 107) 107. The population of nucleic acids of item 106, wherein the polynucleotide sample obtained from the blood is a cell-free DNA sample. (Item 108) The method of any one of the preceding items, wherein the carrier nucleic acid molecule is a double-stranded molecule. (Item 109) Item 10. The method of any one of the preceding items, wherein the first sample is not subjected to a denaturing step. (Item 110) Item 10. The method or set of carrier nucleic acid molecules of any one of the preceding items, wherein the sequence of nucleotides adjacent to the one or more CpG nucleotides or methylated nucleotides in the first subset and / or the second subset can be a sequence of 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 nucleotides, or 5 nucleotides adjacent to the one or more CpG nucleotides or methylated nucleotides. (Item 111) 2. The method of claim 1, wherein the sequencing step comprises sequencing at least a portion of the processed samples from at least two distributed sets. (Item 112) (i) (a) at least a subset of unmethylated carrier nucleic acid molecules, and / or (b) at least a subset of methylated carrier nucleic acid molecules a set of carrier nucleic acid molecules comprising: at least one end of the set of carrier nucleic acid molecules is modified to prevent ligation, the unmethylated carrier nucleic acid molecules do not contain methylated nucleotides, and the methylated carrier nucleic acid molecules contain one or more methylated nucleotides; and (ii) a capture agent that selectively binds to methylated polynucleotides Includes a kit.
Claims
[Claim 1] The invention described in the specification.