Methods, compositions, and systems for improving recovery of nucleic acid molecules
By incorporating fragment size control molecules to correct for size bias and contamination, the method improves the analysis of cell-free nucleic acids, offering a more comprehensive assessment of tumor conditions beyond somatic mutations.
Patent Information
- Application Number
- JP2025066850
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-12-20
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Current methods for analyzing cell-free nucleic acids in cancer diagnostics focus primarily on somatic mutations, neglecting the valuable information provided by the size distribution and fragmentation pattern of cell-free DNA, which can offer a more comprehensive assessment of tumor conditions.
The method involves adding fragment size control molecules to a sample of cell-free polynucleotides, extracting and processing these molecules, enriching specific subsets, and sequencing to generate fragment size scores, thereby correcting for size bias and detecting contamination.
This approach enhances the accuracy of nucleic acid analysis by addressing size bias and contamination, providing a more comprehensive understanding of tumor conditions through improved fragment size distribution analysis.
Smart Images

Figure 2025106543000001 
Figure 2025106543000002 
Figure 2025106543000003
Abstract
Description
Technical Field
[0001] Cross-reference This application claims the benefit and priority of U.S. Provisional Patent Application No. 62 / 783,046, filed Dec. 20, 2018, which is hereby incorporated by reference in its entirety.
Background Art
[0002] Background Current methods for cell-free nucleic acid (e.g., cell-free DNA or cell-free RNA) cancer diagnostic assays focus on the detection of tumor-associated somatic variants, including single nucleotide variants (SNVs), copy number variants (CNVs), fusions, and indels (i.e., insertions or deletions), all of which are the main targets of liquid biopsies. There is increasing evidence that the size distribution and fragmentation pattern of cell-free DNA can provide information regarding the origin and disease level of cell-free DNA. Combining the size distribution and fragmentation pattern of cell-free DNA with the calling of somatic mutations can provide a more comprehensive assessment of the tumor situation than can be obtained from either approach alone.
Summary of the Invention
Means for Solving the Problems
[0003] Abstract The present disclosure provides methods, compositions, and systems for analyzing nucleic acids using fragment size molecules.
[0004] In one aspect, the present disclosure provides a method for analyzing nucleic acid molecules in a sample of polynucleotides, the method comprising: a) adding a subset of fragment size control molecules to nucleic acid molecules in a sample of cell-free polynucleotides, thereby producing a first spike-in sample; b) extracting nucleic acids from the first spike-in sample; c) processing at least a subset of the extracted nucleic acids, thereby producing a processed sample, wherein processing comprises fractionating, tagging, and / or amplifying at least a subset of the first spike-in sample; d) enriching at least a subset of the processed sample, thereby producing an enriched sample; e) sequencing at least a subset of the enriched sample to generate a plurality of sequence reads; and f) analyzing the plurality of sequence reads to generate a plurality of fragment size scores for the subset of fragment size control molecules. In some embodiments, the method further comprises, prior to c), adding a second subset of fragment size control molecules, thereby producing a second spike-in sample. In some embodiments, the method further comprises, prior to d), adding a third subset of fragment size control molecules, thereby producing a third spike-in sample. In some embodiments, the method comprises, prior to e), adding a fourth subset of fragment size control molecules, thereby producing a fourth spike-in sample.
[0005] In some embodiments, the sample of polynucleotides is a sample of cell-free polynucleotides. In some embodiments, the sample of polynucleotides is selected from the group consisting of a sample of cell-free DNA and a sample of cell-free RNA. In some embodiments, the sample of cell-free polynucleotides is cell-free DNA. In some embodiments, the cell-free DNA is between 1 ng and 500 ng. In some embodiments, the concentration of the fragment size control molecules is between 1 attomole and 10 picomoles.
[0006] In another aspect, the present disclosure provides a set of fragment size control molecules comprising at least one subset of a predefined set of fragment size control molecules, wherein at least one subset of the predefined set of fragment size control molecules comprises a plurality of fragment size control molecules that include a fragment size region. In some embodiments, the fragment size control molecules further include an identifier region.
[0007] In some embodiments, the at least one subset includes at least one group of fragment size control molecules.
[0008] In some embodiments, a subset of the fragment size control molecules comprises a plurality of fragment size control molecules that include a fragment size region. In some embodiments, the fragment size control molecules further include an identifier region. In some embodiments, the subset includes at least one group of fragment size control molecules.
[0009] In some embodiments, the fragment size regions of the fragment size control molecules in a group are regions of the same length. In some embodiments, the length of the fragment size region in a first group of fragment size control molecules is different from the length of the fragment size region in a second group of fragment size control molecules.
[0010] In some embodiments, the identifier region is present on one or both sides of the fragment size region. In some embodiments, the identifier region includes a molecular barcode.
[0011] In some embodiments, the plurality of fragment control molecules include one or more primer binding sites. In some embodiments, the one or more primer binding sites are present in the identifier region.
[0012] In some embodiments, the fragment size regions of the fragment size control molecules in a group contain the same oligonucleotide sequence. In some embodiments, the fragment size regions of the fragment size control molecules in a group contain at least two distinguishable oligonucleotide sequences. In some embodiments, the fragment size regions of the fragment size control molecules in a first subset of a given fragment size control molecule contain an oligonucleotide sequence distinguishable from the oligonucleotide sequence of the fragment size regions of the fragment size control molecules in a second subset of the given fragment size control molecule.
[0013] In some embodiments, the length of the fragment size region is at least 10 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 120 bp, at least 150 bp, at least 200 bp, at least 250 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, or at least 1000 bp. In some embodiments, the length of the fragment size region is between 10 bp and 1000 bp.
[0014] In some embodiments, each subset of at least one subset of a given fragment size control molecule is present at an equimolar concentration. In some embodiments, each subset of at least one subset of a given fragment size control molecule is present at a non-equimolar concentration. In some embodiments, each group of at least one group of the fragment size control molecules in at least one subset is present at an equimolar concentration. In some embodiments, each group of at least one group of the fragment size control molecules in at least one subset is present at a non-equimolar concentration.
[0015] In some embodiments, each of the subsets of fragment size control molecules is present at equimolar concentrations. In some embodiments, each of the subsets of fragment size control molecules is present at non-equimolar concentrations. In some embodiments, each of the groups of fragment size control molecules in the subset is present at equimolar concentrations. In some embodiments, each of the groups of fragment size control molecules in the subset is present at non-equimolar concentrations.
[0016] In another aspect, the present disclosure provides a method for evaluating fragment size bias in the analysis of nucleic acid molecules in a sample of cell-free polynucleotides, the method comprising: a) adding a first subset of fragment size control molecules to nucleic acid molecules in a sample of polynucleotides, thereby producing a first spike-in sample; b) extracting nucleic acids from the first spike-in sample; c) adding a second subset of fragment size control molecules to the extracted nucleic acids, thereby producing a second spike-in sample; d) processing at least a subset of the second spike-in sample, thereby producing a processed sample, wherein processing comprises fractionating, tagging, and / or amplifying at least a subset of the second spike-in sample; e) adding a third subset of fragment size control molecules to the processed sample, thereby producing a third spike-in sample; f) enriching at least a subset of the third spike-in sample, thereby producing an enriched sample; g) adding a fourth subset of fragment size control molecules to at least a subset of the enriched sample, thereby producing a fourth spike-in sample; h) sequencing the fourth spike-in sample to generate a plurality of sequence reads; and i) analyzing the plurality of sequence reads to generate a plurality of fragment size scores for the first subset of fragment size control molecules, the second subset of fragment size control molecules, the third subset of fragment size control molecules, and / or the fourth subset of fragment size control molecules. In some embodiments, the method further comprises comparing the plurality of fragment size scores to a plurality of fragment size thresholds. In some embodiments, the method further comprises optimizing the analysis of nucleic acid molecules in the sample of polynucleotides based on the plurality of fragment size scores. In some embodiments, the method further comprises correcting for fragment size bias in the analysis of nucleic acid molecules in a sample of cell-free polynucleotides using the plurality of fragment size scores.In some embodiments, the method further includes classifying the method as successful if at least one of the plurality of fragment size scores is within the corresponding fragment size threshold of the plurality of fragment size thresholds; or as a failure if at least one of the plurality of fragment size scores is not within the corresponding fragment size threshold of the plurality of fragment size thresholds.
[0017] In another aspect, the present disclosure provides a method for detecting contamination of a first sample by a second sample, comprising, for each of the first sample and the second sample: (a) adding a subset of fragment size control molecules to generate a first spike-in sample, wherein the subset of fragment size control molecules added to the first sample is distinguishable from the subset of fragment size control molecules added to the second sample; (b) extracting nucleic acid from the first spike-in; (c) processing at least a subset of the extracted nucleic acid to produce a processed sample, wherein processing includes fractionating, tagging, and / or amplifying at least a subset of the first spike-in sample; (d) enriching at least a subset of the processed sample; (e) sequencing at least a subset of the enriched sample to generate a plurality of sequence reads; and (f) analyzing the plurality of sequence reads to generate one or more contamination scores for the subset of fragment size control molecules. In some embodiments, the method further comprises, prior to (c), adding a second subset of fragment size control molecules to generate a second spike-in sample, wherein the subset of fragment size control molecules added to the first sample is distinguishable from the subset of fragment size control molecules added to the second sample. In some embodiments, the method further comprises, prior to (d), adding a third subset of fragment size control molecules to generate a third spike-in sample, wherein the subset of fragment size control molecules added to the first sample is distinguishable from the subset of fragment size control molecules added to the second sample. In some embodiments, the method further comprises, prior to (e), adding a fourth subset of fragment size control molecules to generate a fourth spike-in sample, wherein the subset of fragment size control molecules added to the first sample is distinguishable from the subset of fragment size control molecules added to the second sample.
[0018] In another aspect, the present disclosure provides a method for detecting contamination of a first sample by a second sample, comprising, for each of the first sample and the second sample: (a) adding a first subset of fragment size control molecules to generate a first spike-in sample, wherein the first subset of fragment size control molecules added to the first sample is distinguishable from the first subset of fragment size control molecules added to the second sample; (b) extracting nucleic acid from the first spike-in; (c) adding a second subset of fragment size control molecules to the extracted nucleic acid to thereby produce a second spike-in sample, wherein the second subset of fragment size control molecules added to the first sample is distinguishable from the second subset of fragment size control molecules added to the second sample; (d) processing at least a subset of the extracted nucleic acid to thereby produce a processed sample, wherein processing comprises fractionating, tagging, and / or amplifying at least a subset of the first spike-in sample; (e) adding a third subset of fragment size control molecules to the extracted nucleic acid to thereby produce a third spike-in sample, wherein the third subset of fragment size control molecules added to the first sample is distinguishable from the third subset of fragment size control molecules added to the second sample; (f) enriching at least a subset of the processed sample; (g) adding a fourth subset of fragment size control molecules to the extracted nucleic acid molecules to thereby produce a fourth spike-in sample, wherein the fourth subset of fragment size control molecules added to the first sample is distinguishable from the fourth subset of fragment size control molecules added to the second sample; (h) sequencing at least a subset of the enriched sample to generate a plurality of sequence reads; and (i) analyzing the plurality of sequence reads to generate one or more contamination scores for a subset of the fragment size control molecules.
[0019] In some embodiments, the method further comprises comparing at least one or more contamination scores to at least one or more contamination thresholds. In some embodiments, the method comprises classifying a first sample as contaminated by a second sample if (i) at least one or more of the contamination scores are not within the corresponding contamination threshold of one or more contamination thresholds; or (ii) not contaminated by the second sample if at least one or more of the contamination scores are within the corresponding contamination threshold of one or more contamination thresholds.
[0020] In some embodiments, fractionating comprises fractionating nucleic acid molecules of at least a subset of the second spike-in sample into a plurality of fractionation sets. In some embodiments, the plurality of fractionation sets comprise nucleic acid molecules of the second spike-in sample fractionated based on the level of epigenetic modification of the nucleic acid molecules of the second spike-in sample.
[0021] In some embodiments, tagging comprises attaching a set of tags to a nucleic acid to produce a population of tagged nucleic acids, wherein the tagged nucleic acids comprise one or more tags. In some embodiments, the set of tags used in the first fractionation set of the plurality of fractionation sets resulting from fractionation is different from the set of tags used in the second fractionation set of the plurality of fractionation sets. In some embodiments, the set of tags is attached to the nucleic acid by ligation of an adapter to the nucleic acid, and the adapter comprises one or more tags.
[0022] In some embodiments, the sample of polynucleotides is a sample of cell-free polynucleotides. In some embodiments, the sample of polynucleotides is selected from the group consisting of a sample of cell-free DNA and a sample of cell-free RNA. In some embodiments, the sample of cell-free polynucleotides is cell-free DNA. In some embodiments, the cell-free DNA is between 1 ng and 500 ng. In some embodiments, the concentration of the fragment size control molecule is between 1 attomolar and 10 picomoles.
[0023] In another aspect, the present disclosure provides a set of fragment size control molecules comprising at least one subset of a set of predetermined fragment size control molecules, wherein at least one subset of the predetermined fragment size control molecules comprises a plurality of fragment size control molecules that include a fragment size region. In some embodiments, the fragment size control molecules further include an identifier region.
[0024] In some embodiments, the at least one subset includes at least one group of fragment size control molecules.
[0025] In some embodiments, a subset of the fragment size control molecules includes a plurality of fragment size control molecules that include a fragment size region. In some embodiments, the fragment size control molecules further include an identifier region. In some embodiments, the subset includes at least one group of fragment size control molecules.
[0026] In some embodiments, the fragment size regions of the fragment size control molecules in a group are regions of the same length. In some embodiments, the length of the fragment size region in a first group of fragment size control molecules is different from the length of the fragment size region in a second group of fragment size control molecules.
[0027] In some embodiments, the identifier region is present on one or both sides of the fragment size region. In some embodiments, the identifier region includes a molecular barcode.
[0028] In some embodiments, the plurality of fragment control molecules include one or more primer binding sites. In some embodiments, one or more primer binding sites are present in the identifier region.
[0029] In some embodiments, the fragment size regions of the fragment size control molecules in the group contain the same oligonucleotide sequence. In some embodiments, the fragment size regions of the fragment size control molecules in the group contain at least two distinguishable oligonucleotide sequences. In some embodiments, the fragment size regions of the fragment size control molecules in the first subset of a given fragment size control molecule contain an oligonucleotide sequence distinguishable from the oligonucleotide sequence of the fragment size regions of the fragment size control molecules in the second subset of the given fragment size control molecule.
[0030] In some embodiments, the length of the fragment size region is at least 10 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 120 bp, at least 150 bp, at least 200 bp, at least 250 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, or at least 1000 bp. In some embodiments, the length of the fragment size region is between 10 bp and 1000 bp.
[0031] In some embodiments, each subset of at least one subset of a given fragment size control molecule is present at an equimolar concentration. In some embodiments, each subset of at least one subset of a given fragment size control molecule is present at a non-equimolar concentration. In some embodiments, each group of at least one group of the fragment size control molecules in at least one subset is present at an equimolar concentration. In some embodiments, each group of at least one group of the fragment size control molecules in at least one subset is present at a non-equimolar concentration.
[0032] In some embodiments, each of the subsets of fragment size control molecules is present at equimolar concentrations. In some embodiments, each of the subsets of fragment size control molecules is present at non-equimolar concentrations. In some embodiments, each of the groups of fragment size control molecules in a subset is present at equimolar concentrations. In some embodiments, each of the groups of fragment size control molecules in a subset is present at non-equimolar concentrations.
[0033] In another aspect, the present disclosure provides a method for generating a sequencing library of a sample of cell-free polynucleotides, the method comprising: a) adding a subset of fragment size control molecules to the sample, thereby producing a first spike-in sample; b) extracting nucleic acids from the first spike-in sample; c) processing at least a subset of the extracted nucleic acids, thereby producing a processed sample, wherein processing comprises fractionating, tagging, and / or amplifying at least a subset of the first spike-in sample; and d) enriching at least a subset of the processed sample. In some embodiments, the method further comprises, prior to c), adding a second subset of fragment size control molecules, thereby producing a second spike-in sample. In some embodiments, the method further comprises, prior to d), adding a third subset of fragment size control molecules, thereby producing a third spike-in sample. In some embodiments, the method further comprises e) adding a fourth subset of fragment size control molecules, thereby producing a fourth spike-in sample.
[0034] In some embodiments, at least one subset comprises at least one group of fragment size control molecules.
[0035] In some embodiments, a subset of fragment size control molecules comprises a plurality of fragment size control molecules that include a fragment size region. In some embodiments, a fragment size control molecule further comprises an identifier region. In some embodiments, a subset comprises at least one group of fragment size control molecules.
[0036] In some embodiments, the fragment size region of the fragment size control molecules in a group is a region of the same length. In some embodiments, the length of the fragment size region of the fragment size control molecules in the first group is different from the length of the fragment size region of the fragment size control molecules in the second group.
[0037] In some embodiments, the identifier region is present on one or both sides of the fragment size region. In some embodiments, the identifier region contains a molecular barcode.
[0038] In some embodiments, the plurality of fragment control molecules include one or more primer binding sites. In some embodiments, one or more primer binding sites are present in the identifier region.
[0039] In some embodiments, the fragment size regions of the fragment size control molecules in a group contain the same oligonucleotide sequence. In some embodiments, the fragment size regions of the fragment size control molecules in a group contain at least two distinguishable oligonucleotide sequences. In some embodiments, the fragment size regions of the fragment size control molecules in the first subset of a given fragment size control molecule contain an oligonucleotide sequence distinguishable from the oligonucleotide sequence of the fragment size regions of the fragment size control molecules in the second subset of the given fragment size control molecule.
[0040] In some embodiments, the length of the fragment size region is at least 10bp, at least 50bp, at least 60bp, at least 70bp, at least 80bp, at least 90bp, at least 100bp, at least 120bp, at least 150bp, at least 200bp, at least 250bp, at least 300bp, at least 400bp, or at least 500bp. In some embodiments, the length of the fragment size region is between 10bp and 1000bp.
[0041] In some embodiments, each subset of at least one subset of the predefined fragment size control molecules is present at equimolar concentrations. In some embodiments, each subset of at least one subset of the predefined fragment size control molecules is present at non-equimolar concentrations. In some embodiments, each group of at least one group of the fragment size control molecules in at least one subset is present at equimolar concentrations. In some embodiments, each group of at least one group of the fragment size control molecules in at least one subset is present at non-equimolar concentrations.
[0042] In yet another aspect, the present disclosure provides a population of nucleic acids comprising: (a) a set of fragment size control molecules comprising at least one subset of a set of predefined fragment size control molecules, wherein at least one subset of the predefined fragment size control molecules comprises a plurality of fragment size control molecules comprising a fragment size region; and (b) a set of nucleic acid molecules in a sample of polynucleotides from a subject. In some embodiments, the at least one subset comprises at least one group of fragment size control molecules. In some embodiments, the fragment size regions of the fragment size control molecules in the group are regions of the same length. In some embodiments, the length of the fragment size region of the fragment size control molecules in the first group is different from the length of the fragment size region of the fragment size control molecules in the second group. In some embodiments, the fragment size control molecules further comprise an identifier region. In some embodiments, the identifier region is present on one or both sides of the fragment size region. In some embodiments, the identifier region comprises a molecular barcode. In some embodiments, the plurality of fragment control molecules comprises one or more primer binding sites. In some embodiments, the one or more primer binding sites are present in the identifier region. In some embodiments, the fragment size regions of the fragment size control molecules in the group comprise the same oligonucleotide sequence. In some embodiments, the fragment size regions of the fragment size control molecules in the group comprise at least two distinguishable oligonucleotide sequences. In some embodiments, the oligonucleotide sequence of the fragment size region of the fragment size control molecules in the first subset of the predefined fragment size control molecules is distinguishable from the oligonucleotide sequence of the fragment size region of the fragment size control molecules in the second subset of the predefined fragment size control molecules. In some embodiments, the length of the fragment size region is at least 10 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 120 bp, at least 150 bp, at least 200 bp, at least 250 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, or at least 1000 bp.In some embodiments, the length of the fragment size region is between 10 bp and 1000 bp. In some embodiments, each subset of at least one subset of the predefined fragment size control molecules is present at an equimolar concentration. In some embodiments, each group of at least one group of the fragment size control molecules in at least one subset is present at an equimolar concentration. In some embodiments, each group of at least one group of the fragment size control molecules in at least one subset is present at a non-equimolar concentration.
[0043] In another aspect, the disclosure includes a computer-readable medium containing or accessible to a controller including non-transitory computer-executable instructions that, when executed by at least one electronic processor, implement a method, the method comprising: a) adding a subset of fragment size control molecules to nucleic acid molecules in a sample of polynucleotides, thereby producing a first spike-in sample; b) extracting nucleic acid from the first spike-in sample; c) processing at least a subset of the extracted nucleic acid, thereby producing a processed sample, where processing includes fractionating, tagging, and / or amplifying at least a subset of the first spike-in sample; d) enriching at least a subset of the processed sample, thereby producing an enriched sample; e) sequencing at least a subset of the enriched sample to generate a plurality of sequence reads; and f) analyzing the plurality of sequence reads to generate a plurality of fragment size scores for the subset of fragment size control molecules. In some embodiments, the method further includes, prior to c), adding a second subset of fragment size control molecules, thereby producing a second spike-in sample. In some embodiments, the method further includes, prior to d), adding a third subset of fragment size control molecules, thereby producing a third spike-in sample. In some embodiments, the method further includes, prior to e), adding a fourth subset of fragment size control molecules, thereby producing a fourth spike-in sample.
[0044] In some embodiments, the sample of polynucleotide is a sample of cell-free polynucleotide. In some embodiments, the sample of polynucleotide is selected from the group consisting of a sample of cell-free DNA and a sample of cell-free RNA. In some embodiments, the sample of cell-free polynucleotide is cell-free DNA. In some embodiments, the cell-free DNA is between 1 ng and 500 ng. In some embodiments, the concentration of the fragment size control molecule is between 1 attomole and 10 picomoles.
[0045] In another aspect, the present disclosure provides a set of fragment size control molecules comprising at least one subset of a given fragment size control molecule, wherein at least one subset of the given fragment size control molecule comprises a plurality of fragment size control molecules including a fragment size region. In some embodiments, the fragment size control molecule further comprises an identifier region.
[0046] In some embodiments, the at least one subset comprises at least one group of fragment size control molecules.
[0047] In some embodiments, the subset of fragment size control molecules comprises a plurality of fragment size control molecules including a fragment size region. In some embodiments, the fragment size control molecule further comprises an identifier region. In some embodiments, the subset comprises at least one group of fragment size control molecules.
[0048] In some embodiments, the fragment size regions of the fragment size control molecules in the group are regions of the same length. In some embodiments, the length of the fragment size region in the first group of fragment size control molecules is different from the length of the fragment size region in the second group of fragment size control molecules.
[0049] In some embodiments, the identifier region is present on one or both sides of the fragment size region. In some embodiments, the identifier region comprises a molecular barcode.
[0050] In some embodiments, the plurality of fragment control molecules include one or more primer binding sites. In some embodiments, the one or more primer binding sites are present in the identifier region.
[0051] In some embodiments, the fragment size regions of the fragment size control molecules in a group include the same oligonucleotide sequence. In some embodiments, the fragment size regions of the fragment size control molecules in a group include at least two distinguishable oligonucleotide sequences. In some embodiments, the fragment size regions of the fragment size control molecules in a first subset of a given fragment size control molecule include an oligonucleotide sequence distinguishable from the oligonucleotide sequence of the fragment size regions of the fragment size control molecules in a second subset of the given fragment size control molecule.
[0052] In some embodiments, the length of the fragment size region is at least 10 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 120 bp, at least 150 bp, at least 200 bp, at least 250 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, or at least 1000 bp. In some embodiments, the length of the fragment size region is between 10 bp and 1000 bp.
[0053] In some embodiments, each subset of at least one subset of a given fragment size control molecule is present in equimolar concentration. In some embodiments, each subset of at least one subset of a given fragment size control molecule is present in non-equimolar concentration. In some embodiments, each group of at least one group of fragment size control molecules in at least one subset is present in equimolar concentration. In some embodiments, each group of at least one group of fragment size control molecules in at least one subset is present in non-equimolar concentration.
[0054] In some embodiments, each of the subsets of fragment size control molecules is present at equimolar concentrations. In some embodiments, each of the subsets of fragment size control molecules is present at non-equimolar concentrations. In some embodiments, each of the groups of fragment size control molecules in the subset is present at equimolar concentrations. In some embodiments, each of the groups of fragment size control molecules in the subset is present at non-equimolar concentrations.
[0055] In another aspect, the present disclosure includes a computer-readable medium containing, or accessible to, non-transitory computer-executable instructions that, when executed by at least one electronic processor, implement a method, or a system including a controller, the method comprising: a) adding a first subset of fragment size control molecules to nucleic acid molecules in a sample of polynucleotides, thereby producing a first spike-in sample; b) extracting nucleic acids from the first spike-in sample; c) adding a second subset of fragment size control molecules to the extracted nucleic acids, thereby producing a second spike-in sample; d) processing at least a subset of the second spike-in sample, thereby producing a processed sample, wherein processing includes fractionating, tagging, and / or amplifying at least a subset of the second spike-in sample; e) adding a third subset of fragment size control molecules to the processed sample, thereby producing a third spike-in sample; f) enriching at least a subset of the third spike-in sample, thereby producing an enriched sample; g) adding a fourth subset of fragment size control molecules to at least a subset of the enriched sample, thereby producing a fourth spike-in sample; h) sequencing the fourth spike-in sample to generate a plurality of sequence reads; and i) analyzing the plurality of sequence reads to generate a plurality of fragment size scores for the first subset of fragment size control molecules, the second subset of fragment size control molecules, the third subset of fragment size control molecules, and / or the fourth subset of fragment size control molecules. In some embodiments, the method further includes comparing the plurality of fragment size scores to a plurality of fragment size thresholds. In some embodiments, the method further includes optimizing the analysis of nucleic acid molecules in the sample of polynucleotides based on the plurality of fragment size scores. In some embodiments, the method further includes correcting for fragment size bias in the analysis of nucleic acid molecules in the sample of polynucleotides using the plurality of fragment size scores.In some embodiments, the method further includes classifying the method as successful if each of a plurality of fragment size scores is within a corresponding fragment size threshold of a plurality of fragment size thresholds; or as failed if at least one of the plurality of fragment size scores is not within a corresponding fragment size threshold of the plurality of fragment size thresholds. In some embodiments, the method further includes comparing at least one of the plurality of fragment size scores to at least one of a plurality of contamination thresholds. In some embodiments, the method further includes classifying the sample as contaminated by another sample if at least one of the plurality of fragment size scores is not within a corresponding contamination threshold of the plurality of contamination thresholds; or as not contaminated by another sample if at least one of the plurality of fragment size scores is within a corresponding contamination threshold of the plurality of contamination thresholds.
[0056] In some embodiments, the sample of polynucleotide is a sample of cell-free polynucleotide. In some embodiments, the sample of polynucleotide is selected from the group consisting of a sample of cell-free DNA and a sample of cell-free RNA. In some embodiments, the sample of cell-free polynucleotide is cell-free DNA. In some embodiments, the cell-free DNA is between 1 ng and 500 ng. In some embodiments, the concentration of the fragment size control molecule is between 1 attomole and 10 picomoles.
[0057] In another aspect, the disclosure provides a set of fragment size control molecules comprising at least one subset of a set of predetermined fragment size control molecules, wherein at least one subset of the predetermined fragment size control molecules comprises a plurality of fragment size control molecules that include a fragment size region. In some embodiments, the fragment size control molecule further includes an identifier region.
[0058] In some embodiments, the at least one subset includes at least one group of fragment size control molecules.
[0059] In some embodiments, a subset of fragment size control molecules comprises a plurality of fragment size control molecules that include a fragment size region. In some embodiments, the fragment size control molecules further include an identifier region. In some embodiments, the subset includes at least one group of fragment size control molecules.
[0060] In some embodiments, the fragment size regions of the fragment size control molecules in a group are regions of the same length. In some embodiments, the length of the fragment size region in a first group of fragment size control molecules is different from the length of the fragment size region in a second group of fragment size control molecules.
[0061] In some embodiments, the identifier region is present on one or both sides of the fragment size region. In some embodiments, the identifier region includes a molecular barcode.
[0062] In some embodiments, the plurality of fragment control molecules includes one or more primer binding sites. In some embodiments, one or more primer binding sites are present in the identifier region.
[0063] In some embodiments, the fragment size regions of the fragment size control molecules in a group include the same oligonucleotide sequence. In some embodiments, the fragment size regions of the fragment size control molecules in a group include at least two distinguishable oligonucleotide sequences. In some embodiments, the fragment size regions of the fragment size control molecules in a first subset of a given fragment size control molecule include an oligonucleotide sequence distinguishable from the oligonucleotide sequence of the fragment size regions of the fragment size control molecules in a second subset of the given fragment size control molecule.
[0064] In some embodiments, the length of the fragment size region is at least 10 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 120 bp, at least 150 bp, at least 200 bp, at least 250 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, or at least 1000 bp. In some embodiments, the length of the fragment size region is between 10 bp and 1000 bp.
[0065] In some embodiments, each subset of at least one subset of the predefined fragment size control molecules is present at an equimolar concentration. In some embodiments, each subset of at least one subset of the predefined fragment size control molecules is present at a non - equimolar concentration. In some embodiments, each group of at least one group of the fragment size control molecules in at least one subset is present at an equimolar concentration. In some embodiments, each group of at least one group of the fragment size control molecules in at least one subset is present at a non - equimolar concentration.
[0066] In some embodiments, each of the subsets of the fragment size control molecules is present at an equimolar concentration. In some embodiments, each of the subsets of the fragment size control molecules is present at a non - equimolar concentration. In some embodiments, each of the groups of the fragment size control molecules in the subset is present at an equimolar concentration. In some embodiments, each of the groups of the fragment size control molecules in the subset is present at a non - equimolar concentration.
[0067] In another aspect, the present disclosure includes a computer-readable medium including, or accessible to, computer-executable instructions that, when executed by at least one electronic processor, implement a method, the method comprising: a) adding a subset of fragment size control molecules to a sample, thereby producing a first spike-in sample; b) extracting nucleic acids from the first spike-in sample; c) processing at least a subset of the extracted nucleic acids, thereby producing a processed sample, wherein processing includes fractionating, tagging, and / or amplifying at least a subset of the first spike-in sample; and d) enriching at least a subset of the processed sample. In some embodiments, the method further includes, prior to c), adding a second subset of fragment size control molecules, thereby producing a second spike-in sample. In some embodiments, the method further includes, prior to d), adding a third subset of fragment size control molecules, thereby producing a third spike-in sample. In some embodiments, the method further includes e) adding a fourth subset of fragment size control molecules, thereby producing a fourth spike-in sample.
[0068] In some embodiments, the sample of polynucleotides is a sample of cell-free polynucleotides. In some embodiments, the sample of polynucleotides is selected from the group consisting of a sample of cell-free DNA and a sample of cell-free RNA. In some embodiments, the sample of cell-free polynucleotides is cell-free DNA. In some embodiments, the cell-free DNA is between 1 ng and 500 ng. In some embodiments, the concentration of the fragment size control molecules is between 1 attomolar and 10 picomolar.
[0069] In another aspect, the present disclosure provides a set of fragment size control molecules comprising at least one subset of a predefined set of fragment size control molecules, wherein at least one subset of the predefined set of fragment size control molecules comprises a plurality of fragment size control molecules that include a fragment size region. In some embodiments, the fragment size control molecules further include an identifier region.
[0070] In some embodiments, the at least one subset comprises at least one group of fragment size control molecules.
[0071] In some embodiments, a subset of the fragment size control molecules comprises a plurality of fragment size control molecules that include a fragment size region. In some embodiments, the fragment size control molecules further include an identifier region. In some embodiments, the subset comprises at least one group of fragment size control molecules.
[0072] In some embodiments, the fragment size regions of the fragment size control molecules in a group are regions of the same length. In some embodiments, the length of the fragment size region in a first group of fragment size control molecules is different from the length of the fragment size region in a second group of fragment size control molecules.
[0073] In some embodiments, the identifier region is present on one or both sides of the fragment size region. In some embodiments, the identifier region includes a molecular barcode.
[0074] In some embodiments, the plurality of fragment control molecules includes one or more primer binding sites. In some embodiments, one or more primer binding sites are present in the identifier region.
[0075] In some embodiments, the fragment size regions of the fragment size control molecules in a group contain the same oligonucleotide sequence. In some embodiments, the fragment size regions of the fragment size control molecules in a group contain at least two distinguishable oligonucleotide sequences. In some embodiments, the fragment size regions of the fragment size control molecules in a first subset of a given fragment size control molecule contain an oligonucleotide sequence distinguishable from the oligonucleotide sequence of the fragment size regions of the fragment size control molecules in a second subset of the given fragment size control molecule.
[0076] In some embodiments, the length of the fragment size region is at least 10 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 120 bp, at least 150 bp, at least 200 bp, at least 250 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, or at least 1000 bp. In some embodiments, the length of the fragment size region is between 10 bp and 1000 bp.
[0077] In some embodiments, the fragment size control molecule is a synthetic molecule. In some embodiments, the fragment size control molecule is generated as an amplicon by PCR amplification.
[0078] In some embodiments, each subset of at least one subset of a given fragment size control molecule is present at an equimolar concentration. In some embodiments, each subset of at least one subset of a given fragment size control molecule is present at a non-equimolar concentration. In some embodiments, each group of at least one group of the fragment size control molecules in at least one subset is present at an equimolar concentration. In some embodiments, each group of at least one group of the fragment size control molecules in at least one subset is present at a non-equimolar concentration.
[0079] In some embodiments, each of the subsets of fragment size control molecules is present at equimolar concentrations. In some embodiments, each of the subsets of fragment size control molecules is present at non-equimolar concentrations. In some embodiments, each of the groups of fragment size control molecules in a subset is present at equimolar concentrations. In some embodiments, each of the groups of fragment size control molecules in a subset is present at non-equimolar concentrations.
[0080] The present disclosure also provides a kit for performing any of the above methods. An exemplary kit includes: (a) a set of fragment size control molecules comprising at least one subset of a predefined fragment size control molecule, wherein at least one subset of the predefined fragment size control molecule comprises a plurality of fragment size control molecules that include a fragment size region.
[0081] In some embodiments, the method or system further comprises creating a report that includes information regarding the analysis of nucleic acid molecules, and / or information derived therefrom, as needed.
[0082] In some embodiments, the results of the systems and / or methods disclosed herein are used as input to create a report. The report can be in written or electronic format. For example, information regarding the analysis of nucleic acid molecules determined by the methods or systems disclosed herein and / or information derived from the analysis can be displayed in such a report. The methods or systems disclosed herein can further comprise transmitting the report to a third party, such as a subject from whom the sample is derived or a healthcare provider.
[0083] The various steps of the methods disclosed herein or the steps performed by the systems disclosed herein can be performed at the same time or different times, and / or at the same geographical location or different geographical locations, e.g., countries. The various steps of the methods disclosed herein can be performed by the same person or different people.
[0084] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, which illustrates and describes merely exemplary embodiments of the present disclosure. As will be recognized, the present disclosure is capable of other and different embodiments, and some of the details thereof are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not restrictive.
[0085] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate certain embodiments and together with the written description serve to explain certain principles of the methods, computer-readable media, and systems disclosed herein. The description provided herein is to be read in conjunction with the accompanying drawings, which are included by way of example and not limitation, for a better understanding. Similar reference symbols are understood to identify similar components throughout the drawings unless the context clearly indicates otherwise. Similarly, some or all of the figures may be schematic for illustrative purposes and not necessarily depict actual relative sizes or positions of the elements shown.
Brief Description of the Drawings
[0086]
Figure 1-1
Figure 1-2
[0087]
Figure 2-1
Figure 2-2
[0088]
Figure 3
[0089]
Figure 4
[0090]
Figure 5
Embodiments for Carrying Out the Invention
[0091] Definitions To more readily understand the present disclosure, certain terms are first defined below. Additional definitions of the following terms and other terms may be described throughout this specification. If the definitions of the terms described below are inconsistent with the definitions in the applications or patents incorporated by reference, the definitions described in this application should be used to understand the meaning of the terms.
[0092] As used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to "a method" includes one or more methods and / or steps of the type described herein and / or made apparent to one of ordinary skill in the art reading the present disclosure.
[0093] Likewise, it is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. Furthermore, unless otherwise defined, all scientific and technical terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. In the description and claims of methods, computer-readable media, and systems, the following terminology and its grammatical variations are used in accordance with the definitions set forth below.
[0094] About: As used herein, "about" or "approximately" when applied to one or more values or elements of interest refers to a value or element that is similar to the recited reference value or element. In certain embodiments, the term "about" or "approximately" refers to a range of values or elements that fall within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less in either direction (greater than or less than) of the recited reference value or element (except where such numbers would exceed 100% of the possible value or element).
[0095] Adapter: As used herein, "adapter" refers to a short nucleic acid that is typically at least partially double-stranded (e.g., less than about 500 nucleotides in length, less than about 100 nucleotides in length, or less than about 50 nucleotides in length), and the adapter can be attached to one or both ends of a given sample nucleic acid molecule. The adapter can include a nucleic acid primer binding site for enabling amplification of a nucleic acid molecule with adapters adjacent to both ends, and / or a sequencing primer binding site for sequencing applications, such as various next-generation sequencing (NGS) applications. The adapter can also include a binding site for a capture probe, such as an oligonucleotide attached to a flow cell support. The adapter can also include a nucleic acid tag as described herein. The nucleic acid tag is typically arranged relative to the amplification primer and the sequencing primer binding site such that the nucleic acid tag is included in the amplicon and sequence reads of a given nucleic acid molecule. Adapters of the same or different sequences can be ligated to each end of the nucleic acid molecule. In some embodiments, the same sequence adapters are ligated to each end of the nucleic acid molecule, except that the nucleic acid tags are different. In some embodiments, the adapter is a Y-shaped adapter with one end having a blunt end or a tail with one or more complementary nucleotides for joining to a nucleic acid molecule having a blunt end or a tail at one end as described herein, and the other end of the Y-shaped adapter includes a non-complementary sequence that does not hybridize to form a double strand. In still other exemplary embodiments, the adapter is a bell-shaped adapter including a blunt or tailed end for joining to the nucleic acid molecule being analyzed. Other examples of adapters include T-tail and C-tail adapters.
[0096] Amplify: As used herein, "amplify" or "amplification" in the context of a nucleic acid typically refers to the production of multiple copies of a polynucleotide or a portion of a polynucleotide starting from a small amount of polynucleotide (e.g., a single polynucleotide molecule) where the amplification product or amplicon is generally detectable. Amplification of polynucleotides encompasses a variety of chemical and enzymatic processes. Amplification includes, but is not limited to, polymerase chain reaction (PCR).
[0097] Cancer type: As used herein, "cancer type" refers to the type or subtype of cancer defined, for example, by histopathology. Cancer types can be defined by any conventional criteria, such as the site of origin in a given tissue (e.g., blood cancers, central nervous system (CNS), brain cancer, lung cancer (small cell and non-small cell), skin cancer, nasal cancer, laryngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, bowel cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, breast cancer, prostate cancer, ovarian cancer, lung cancer, intestinal cancer, soft tissue cancer, neuroendocrine cancer, gastrointestinal cancer, head and neck cancer, gynecological cancer, colorectal cancer, urothelial cancer, solid state cancer, heterogeneous cancer, homogeneous cancer), unknown primary, etc., and / or the same cell lineage (e.g., carcinoma, sarcoma, lymphoma, cholangiocarcinoma, leukemia, mesothelioma, melanoma, or glioblastoma), and / or based on cancers that exhibit cancer markers such as, but not limited to, Her2, CA15-3, CA19-9, CA-125, CEA, AFP, PSA, HCG, hormone receptors, and NMP-22. Cancers can also be classified by stage (e.g., stage 1, 2, 3, or 4), and whether they are primary or secondary.
[0098] Cell-free nucleic acid: As used herein, "cell-free nucleic acid" refers to nucleic acids that are not contained within cells or otherwise bound to cells, or in some embodiments, nucleic acids that remain in a sample after removal of intact cells. Cell-free nucleic acids can include, for example, all unencapsulated nucleic acids sourced from a bodily fluid (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.) from a subject. Cell-free nucleic acids include DNA (cfDNA), RNA (cfRNA), and hybrids thereof, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or hybrids thereof. Cell-free nucleic acids can be released into bodily fluids through secretion or cell death processes, such as necrosis, apoptosis, etc. Some cell-free nucleic acids, such as circulating tumor DNA (ctDNA), are released from cancer cells into bodily fluids. Others are released from healthy cells. CtDNA can be unencapsulated tumor-derived fragmented DNA. Cell-free nucleic acids can have one or more epigenetic modifications, for example, cell-free nucleic acids can be acetylated, 5-methylated, and / or hydroxymethylated.
[0099] Cell nucleic acid: As used herein, "cell nucleic acid" means nucleic acids that are located within one or more cells from which the nucleic acids are derived, at least at the time the sample is taken or collected from a subject, even if the nucleic acids are subsequently removed as part of a given analysis process (e.g., via cell lysis).
[0100] Contamination: As used herein, the term "contamination" or "sample contamination" refers to any chemical or digital contamination of one sample by another sample. Contamination can result from physical carryover of liquid between samples (e.g., pipetting, sample preparation or automated liquid handling via a sequencer system, manipulation of amplified material); reverse multiplexing artifacts (e.g., basecall error cross-linked sample indices with limited pairwise Hamming distance; insertion / deletion cross-linked sample indices with limited pairwise edit distance), and reagent impurities (e.g., sample index oligos contaminated by oligos containing another sample index (through any carryover of synthesis errors)) and can be due to a variety of origins including but not limited to these.
[0101] Contamination Score: As used herein, "Contamination Score" refers to a score in a first sample representative of the presence of fragment size control molecules added to a second sample. In some embodiments, the subset identifier barcode used in one subset of one sample may be different from the subset identifiers used in other subsets of the same sample and subsets of other samples. In these embodiments, the presence of fragment size control molecules belonging to different second samples can be identified from the sequence of the subset identifier barcode present in the fragment size control molecules. In some embodiments, the contamination score can be specific for each type / group of fragment size control molecules used in a subset (i.e., an individual contamination score for different lengths / groups of fragment size control molecules used in a subset), or can be a total score representing different lengths / groups of fragment size control molecules. In some embodiments, the contamination score can be estimated based on the number of fragment size control molecules belonging to different second samples. In some embodiments, the contamination score can be estimated based on the number of sequencing reads of fragment size control molecules belonging to different second samples. In some embodiments, the contamination score of a sample can be estimated based on the fraction or percentage of the number of sequencing reads of fragment size control molecules belonging to other samples relative to the number of sequencing reads of fragment size control molecules added to that sample. In some embodiments, the contamination score of a sample can be estimated based on the fraction or percentage of the number of molecules of fragment size control molecules belonging to other samples relative to the number of molecules of fragment size control molecules added to that sample. In these embodiments, the fragment size control molecules added to that molecule can be identified from the subset identifier barcode of the fragment size control molecules.
[0102] Contamination Threshold: As used herein, “contamination threshold” refers to a value or range of values of a predefined threshold used to assess contamination of a sample in the analysis of nucleic acid molecules in the sample. These thresholds can also be used to optimize the assay at any of the steps such as extraction, library preparation, enrichment, washing / purification, and sequencing. In some embodiments, the contamination threshold can be specific for each type / group of fragment size control molecules used in a subset (i.e., an individual contamination threshold for different lengths / groups of fragment size control molecules used in a subset). In some embodiments, the contamination threshold can be a total threshold for different lengths / groups of fragment size control molecules in a subset, or a total threshold for fragment size control molecules used in one or more subsets added in the assay. In some embodiments, each group in a subset has a specific contamination threshold. For example, a set of fragment size control molecules includes two subsets, subset 1 and subset 2. If each subset includes two groups, G11 and G12 (for subset 1) and G21 and G22 (for subset 2), then each of the four groups can have a contamination score (S), S11 for G11, S12 for G12, S21 for G21, and S22 for G22, based on the recovery rate of the fragment size control molecules. Each group in a subset can have an individual predefined contamination threshold. In this example, T11, T12, T21, and T22 are the contamination thresholds for G11, G12, G21, and G22, respectively. The threshold can be a threshold for a percentage or fraction, and the threshold can be a range of threshold values instead of a specific threshold value. For the method to be considered successful, at least one of these contamination scores must be within its corresponding contamination threshold. In some embodiments, if each of the multiple contamination scores is within its corresponding contamination threshold, the method is classified as having been successfully implemented.
[0103] Coverage: As used herein, the terms “coverage,” “total number of molecules,” or “total number of alleles” are used interchangeably. They refer to the total number of DNA molecules at a specific genomic location in a given sample.
[0104] Deoxyribonucleic acid or ribonucleic acid: As used herein, "deoxyribonucleic acid" or "DNA" refers to natural or modified nucleotides having a hydrogen group at the 2'-position of the sugar moiety. DNA typically comprises a chain of nucleotides containing four types of nucleotide bases: adenine (A), thymine (T), cytosine (C), and guanine (G). As used herein, "ribonucleic acid" or "RNA" refers to natural or modified nucleotides having a hydroxyl group at the 2'-position of the sugar moiety. RNA typically comprises a chain of nucleotides containing four types of nucleotide bases; A, uracil (U), G, and C. As used herein, the term "nucleotide" refers to natural nucleotides or modified nucleotides. Certain pairs of nucleotides specifically bind to each other in a complementary manner (referred to as complementary base pairing). In DNA, adenine (A) pairs with thymine (T), and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U), and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand composed of nucleotides complementary to the nucleotides of the first strand, the two strands bind to form a double strand. As used herein, "nucleic acid sequencing data", "nucleic acid sequencing information", "sequence information", "nucleic acid sequence", "nucleotide sequence", "genomic sequence", "gene sequence", or "fragment sequence, or "nucleic acid sequencing read" refers to any information or data indicating the order and identity of nucleotide bases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a nucleic acid molecule such as DNA or RNA (e.g., whole genome, whole transcriptome, exome, oligonucleotide, polynucleotide, or fragment).This disclosure contemplates sequence information obtained using all available diverse techniques, platforms, or technologies including, but not limited to, capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide discrimination systems, pyrosequencing, ion or pH-based detection systems, and electronic signature-based systems.
[0105] DNA Sequence: As used herein, "DNA sequence" or "sequence" refers to "raw sequence reads" and / or "consensus sequences". Raw sequence reads are the output of a DNA sequencer and typically include, for example, overlapping sequences of the same parental molecule after amplification. A "consensus sequence" is a sequence derived from overlapping sequences of a parental molecule that is intended to represent the sequence of the original parental molecule. A consensus sequence can be produced by voting (in a sequence, the most commonly observed nucleotide at each given base position, e.g., is the consensus nucleotide), or by other approaches such as comparison to a reference genome. A consensus sequence can be produced by tagging the original parental molecule with unique or non-unique molecular tags, thereby enabling tracking of progeny sequences (e.g., after amplification) by tracking the tags and / or using internal information within the sequence reads. Examples of tagging or barcoding, and the use of tags or barcodes are provided, for example, in U.S. Patent Application Publication Nos. 2015 / 0368708, 2015 / 0299812, 2016 / 0040229, and 2016 / 0046986, each of which is incorporated herein by reference in its entirety.
[0106] Epigenetic modification: As used herein, "epigenetic modification" refers to a modification of the base of a nucleotide in a nucleic acid molecule that affects the regulation of that particular nucleic acid sequence and / or gene expression. The modification can be a chemical modification of the base of the nucleotide. In some examples, the modification can be methylation of the base of the nucleotide. For example, the modification can be methylation of cytosine, thereby resulting in 5-methylcytosine.
[0107] Epigenetic state: As used herein, "epigenetic state" refers to the level / degree of epigenetic modification of a nucleic acid molecule. For example, when the epigenetic modification is DNA methylation (or hydroxymethylation), the epigenetic state can refer to the presence or absence of methylation in a DNA base (e.g., cytosine), or the degree of methylation in a nucleic acid sequence (e.g., high methylation, low methylation, medium methylation, or unmethylated nucleic acid molecule). The epigenetic state can also refer to the number of nucleotides having the epigenetic modification. For example, when the epigenetic modification is DNA methylation, the epigenetic state can also refer to the number of methylated nucleotides in the nucleic acid molecule.
[0108] Enriched sample: As used herein, an "enriched sample" refers to a sample in which a specific region of interest is enriched. The sample can be enriched by selectively amplifying the region of interest or by using a double-stranded DNA / RNA probe (e.g., a probe from Twist Biosciences) or a single-stranded DNA / RNA probe that can hybridize to the nucleic acid molecule of interest (e.g., a SureSelect® probe, Agilent Technol). In some embodiments, the enriched sample is a subset of the processed sample to be enriched, where the subset of the processed sample to be enriched contains nucleic acid molecules from a sample of cell-free polynucleotides and from a first and / or second subset of fragment size control molecules. In other embodiments, the enriched sample is a subset of a third spike-in sample to be enriched, where the subset of the third spike-in sample to be enriched contains nucleic acid molecules from a sample of cell-free polynucleotides and from a first, second, and / or third subset of fragment size control molecules.
[0109] First spike-in sample: As used herein, a "first spike-in sample" is a sample in which a first subset of fragment size control molecules is added to a sample of cell-free polynucleotides from a subject.
[0110] Fourth spike-in sample: As used herein, a "fourth spike-in sample" refers to a sample in which a fourth subset of fragment size control molecules is added to a subset of the enriched sample.
[0111] Fragment size bias: As used herein, "fragment size bias" refers to any artificial bias in the length or size of the fragments of nucleic acid molecules being analyzed in an assay. Artifact bias can be due to operator handling or the reagents and / or steps used in the assay. This bias does not include the true biological bias in fragment size observed between healthy and diseased subjects. The assay can include one or more steps such as, but not limited to, handling of liquids, extraction, library preparation, enrichment, washing / purification, and sequencing. In some embodiments, there can be different fragment size biases for each sample. Based on the reaction conditions and the procedures for each of these steps, the recovery rate of nucleic acid molecules belonging to a particular fragment size can be biased. In fragmentome analysis, having accurate quantitative measurements of cfDNA molecules of a particular size is useful in the estimation of the fragment length distribution and fragmentation pattern of cell-free DNA. By estimating the fragment size bias in an assay, it is possible to correct for the fragment size bias (artificial bias) introduced by each step in the assay workflow and have a better estimate of the original fragment size distribution of cfDNA molecules (reflecting the true biological bias).
[0112] Fragment size control molecule: As used herein, "fragment size control molecule" refers to a set of nucleic acid molecules that are added to a sample of polynucleotides to evaluate and / or optimize the analysis of nucleic acid molecules in the sample. A fragment size control molecule can have two regions, a fragment size region and an identifier region. In some embodiments, the fragment size control molecule consists of only the fragment size region. In some embodiments, the fragment size control molecule can have both a fragment size region and an identifier region. A set of fragment size control molecules can be classified into subsets of fragment size control molecules, and each subset of fragment size control molecules can be added at one or more different steps in an assay to evaluate the fragment length distribution, QC metrics, and / or sample contamination in any of the steps of extraction, library preparation, enrichment, and / or sequencing of the polynucleotide sample. In some embodiments, fragment size control molecules can be used to estimate how well the original ends of the polynucleotides in a sample are preserved through an assay workflow. The lengths of these fragment size control molecules are predefined and can mimic the natural size distribution of cell-free DNA. Each subset of fragment size control molecules can be further classified into groups of fragment size control molecules based on the length of the fragment size region, and each group can have a different length of the fragment size region. For example, a subset of fragment size control molecules can be classified into three groups based on the length of the fragment size region, where the first group can have a fragment size region of length 120 bp, the second group can have a fragment size region of length 160 bp, and the third group can have a fragment size region of length 320 bp. In some embodiments, the fragment size control molecule can be a synthetic molecule, i.e., it can be synthesized in vitro with or without the use of enzymes. In some embodiments, the fragment size control molecule can have a nucleic acid sequence that does not occur naturally. In some embodiments, the fragment size control molecule can have a nucleic acid sequence that occurs naturally. In some embodiments, an amplicon generated through the amplification of a specific region (i.e., of a specific length) in any genome, plasmid, or vector, or a part thereof, can be used as a fragment size control molecule.In some embodiments, a plasmid, vector, or any genome can be digested using restriction enzymes, and the digestion products can be used as fragment size control molecules. In some embodiments, the fragment size control molecules can have nucleic acid sequences corresponding to non-human genomes. For example, these molecules can have any of (i) sequences corresponding to regions of lambda phage DNA, (ii) non-naturally occurring sequences, and / or (iii) combinations of (i) and (ii). In some embodiments, the fragment size control molecules can contain non-naturally occurring nucleotide analogs.
[0113] Fragment size region: As used herein, the term "fragment size region" refers to the region of a fragment size control molecule that represents the length of the fragment size control molecule. Each group of fragment size control molecules has a different length of the fragment size region. For example, a subset of fragment size control molecules can be classified into three groups based on the length of the fragment size region. The first group can have a fragment size region of 120 bp in length, the second group can have a fragment size region of 160 bp in length, and the third group can have a fragment size region of 320 bp in length. The length of the fragment size region can be between 10 bp and 1000 bp.
[0114] Fragment Size Score: As used herein, "Fragment Size Score" refers to a score that represents the recovery rate of fragment size control molecules belonging to a particular group and a particular subset. The identity of the fragment size control molecules and the subset to which the fragment size control molecules belong is maintained through the use of barcodes. In some embodiments, the fragment size score can be estimated based on the number of fragment size control molecules belonging to a particular group and a particular subset. In some embodiments, the fragment size score can be estimated based on the number of sequencing reads of fragment size control molecules belonging to a particular group and a particular subset. In some embodiments, the fragment size score can be estimated for a particular group with respect to a particular subset. In some embodiments, the fragment size score can be estimated for each of the groups in every subset. In some embodiments, the fragment size score can be the total score of the fragment size control molecules used in a particular subset (i.e., with respect to all single scores representing the fragment size control molecules within the subset), or the total score of the fragment size control molecules added in one or more steps of the assay (i.e., single scores representing the fragment size control molecules in one or more subsets added in the assay). In some embodiments, the fragment size score can be measured as the difference between the amount of a subset of fragment size control molecules added in a particular step and the amount of another subset of fragment size control molecules added in a different step, and in some embodiments, the amount can be an amount related to either mass (e.g., fg, pg, ng, μg), the number of sequencing reads, or the number of fragment size molecules (belonging to the subset or belonging to a particular group). In some embodiments, the fragment size score can be measured as the recovery rate (fraction or percentage) of the fragment size control molecules (belonging to the subset or belonging to a particular group), and the recovery rate can be calculated using an orthogonal measurement of the relative abundance of the fragment size control molecules within a particular subset added to the sample. In some embodiments, the orthogonal measurement can be obtained from gel electrophoresis, qPCR, ddPCR, or PCR-free sequencing.In some embodiments, the fragment size score can be measured as the difference between the amount of fragment size control molecules of a particular length and the amount of fragment size control molecules of a different length.
[0115] Fragment size threshold: As used herein, "fragment size threshold" refers to the value or range of a predetermined threshold used to evaluate fragment size bias in the analysis of nucleic acid molecules in a sample. These thresholds can also be used to correct any fragment length distribution in any of the steps such as extraction, library preparation, enrichment, washing / purification, and sequencing. In some embodiments, the fragment size threshold can be estimated for a particular group in a particular subset. In some embodiments, each group in the subset has a particular fragment size threshold. In some embodiments, the fragment size threshold can be estimated for each of the groups in any subset. In some embodiments, the fragment size threshold can be the total threshold of the fragment size control molecules used in a particular subset (i.e., with respect to all the individual thresholds of the fragment size control molecules within the subset), or the total threshold of the fragment size control molecules added in one or more steps of the assay (i.e., the individual thresholds of the fragment size control molecules in one or more subsets added in the assay). For example, a set of fragment size control molecules includes two subsets, subset 1 and subset 2. If each subset includes two groups, G11 and G12 (with respect to subset 1) and G21 and G22 (with respect to subset 2), then each of the four groups can have a fragment size score (S), S11 with respect to G11, S12 with respect to G12, S21 with respect to G21, and S22 with respect to G22, based on the recovery rate of the fragment size control molecules. Each group in the subset can have an individual predetermined fragment size threshold. In this example, T11, T12, T21, and T22 are the fragment size thresholds for G11, G12, G21, and G22, respectively. The threshold can be a threshold with respect to a percentage or fraction, and the threshold can be a range of threshold values instead of a particular threshold value. For the method to be considered successful, at least one of these fragment size scores must be within its corresponding fragment size threshold. In some embodiments, if each of a plurality of fragment size scores is within its corresponding fragment size threshold, the method is classified as having been successfully implemented.
[0116] Genomic region: As used herein, "genomic region" refers to any region of a genome, such as the entire genome, a chromosome, a gene, or an exon (e.g., a range of base pair positions). A genomic region can be a continuous or discontinuous region. A "locus" (or "site") can be part or all of a genomic region (e.g., a gene, a part of a gene, or a single nucleotide of a gene).
[0117] Identifier Region: As used herein, an "identifier region" refers to a region of a fragment size control molecule that is used to distinguish the fragment size control molecule from other fragment size control molecules. The identifier region is also used to distinguish fragment size control molecules from one subset from those of another subset. The identifier region may have a molecular barcode. The identifier region may be present on one or both sides of the fragment size region. The identifier region may be either continuous or discontinuous. The molecular barcode serves as an identifier for the fragment size control molecule. The identifier region may have an additional region (primer binding site) that facilitates the binding of one or more primers. In some embodiments, the identifier region may also have an additional flow cell binding site, such as P5 and P7, that enables the fragment size control molecule to attach to the flow cell surface of a next-generation sequencer (e.g., an Illumina sequencer). In some embodiments, the identifier region comprises (i) a molecular identifier barcode, e.g., a barcode used to identify each fragment size control molecule and to distinguish one fragment size control molecule from another, and (ii) a subset identifier barcode, e.g., a barcode that functions as a barcode used to identify the subset to which the fragment size control molecule belongs (i.e., whether the fragment size control molecule belongs to subset 1 or subset 2). The subset identifier barcode may be the same for all fragment size control molecules in the subset, or the subset identifier barcode for one subset may be different from that of another subset. In some embodiments, the subset identifier barcode for one subset in one sample may be different from the subset identifier barcode for the corresponding subset in another sample. For example, when evaluating sample contamination using fragment size control molecules, the subset identifier barcode for one subset is different from the subset identifier barcode for the corresponding subset used in all other samples within the same flow cell / batch.In some embodiments, the identifier region can include a sample index for distinguishing fragment size control molecules of one sample from those of other samples, and a sequence for attaching to the flow cell of the sequencer. For example, the identifier region of a subset of fragment size control molecules added after the enrichment step can include a sample index sequence. In some embodiments, the identifier region can be attached to the fragment size region via ligation. In some embodiments, the identifier region can be added to the fragment size region via PCR.
[0118] Mutation: As used herein, "mutation" refers to a variation from a known reference sequence and includes mutations such as single nucleotide variants (SNVs), and insertions or deletions (indels). Mutations can be germline or somatic mutations. In some embodiments, the reference sequence for comparison purposes is the wild-type genomic sequence of the species of the subject providing the test sample, typically the human genome.
[0119] Neoplasm: As used herein, the terms "neoplasm" and "tumor" are used interchangeably. They refer to the abnormal growth of cells in a subject. A neoplasm or tumor can be benign, potentially malignant, or malignant. Malignant tumors are referred to as cancers or cancerous tumors.
[0120] Next-generation sequencing: As used herein, "next-generation sequencing" or "NGS" refers to sequencing technologies that have increased throughput compared to conventional Sanger and capillary electrophoresis-based approaches and have the ability to generate, for example, hundreds of thousands of relatively small sequence reads at once. Some examples of next-generation sequencing techniques include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization. In some embodiments, next-generation sequencing includes the use of an instrument capable of sequencing single molecules.
[0121] Nucleic acid tag: As used herein, a "nucleic acid tag" is a short nucleic acid (e.g., less than about 500 nucleotides in length, less than about 100 nucleotides in length, less than about 50 nucleotides in length, or less than about 10 nucleotides in length) used to distinguish nucleic acids from nucleic acids from different samples (e.g., representing sample indices) or different nucleic acid molecules in the same sample (e.g., representing molecular barcodes), different types of nucleic acids, or nucleic acids that have undergone different treatments. Nucleic acid tags include predefined, fixed, non-random, random, or semi-random oligonucleotide sequences. Such nucleic acid tags can be used to label different nucleic acid molecules or different nucleic acid samples or subsamples. Nucleic acid tags can be single-stranded, double-stranded, or at least partially double-stranded. Nucleic acid tags can have the same length or diverse lengths as needed. Nucleic acid tags include double-stranded molecules with one or more blunt ends, 5' or 3' single-stranded regions (e.g., overhangs), and / or one or more other single-stranded regions at other positions within a given molecule. Nucleic acid tags can be attached to one or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). Decoding the nucleic acid tag can reveal information such as the origin sample, form, or treatment of a given nucleic acid. For example, nucleic acid tags can also be used to enable pooling and / or parallel processing of multiple samples containing nucleic acids with different molecular barcodes and / or sample indices, in which case the nucleic acids are then deconvoluted by detecting (e.g., reading) the nucleic acid tags. Nucleic acid tags (molecular barcodes) can also be referred to as identifiers (e.g., molecular identifiers, sample identifiers). Additionally or alternatively, nucleic acid tags can be used as molecular identifiers (e.g., to distinguish amplicons of different molecules or different parental molecules in the same sample or subsample). This includes, for example, uniquely tagging different nucleic acid molecules in a given sample or non-uniquely tagging such molecules.In the case of an application that tags non-uniquely, a finite number of tags (i.e., molecular barcodes) can be used to tag each nucleic acid molecule, such that different molecules can be distinguished based on their endogenous sequence information (e.g., the start and / or end positions to which they map in a selected reference genome, subsequences at one or both ends of the sequence, and / or the length of the sequence) in combination with at least one molecular barcode. Typically, a sufficient number of different molecular barcodes are used such that the probability that any two molecules have the same endogenous sequence information (e.g., start and / or end positions, subsequences at one or both ends of the sequence, and / or length) and also have the same molecular barcode is low (e.g., less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% likelihood).
[0122] Fractionation: As used herein, "fractionation" and "epigenetic fractionation" are used interchangeably. This term refers to separating or sorting nucleic acid molecules based on a characteristic of the nucleic acid molecule (e.g., level / degree of epigenetic modification). Fractionation can be a physical fractionation of the molecules. Fractionation can involve separating nucleic acid molecules into groups or sets based on the level of epigenetic modification (i.e., epigenetic state). For example, nucleic acid molecules can be fractionated based on the level of methylation of the nucleic acid molecule. In some embodiments, the methods and systems used for fractionation can be found in PCT Patent Application No. PCT / US2017 / 068329, which is hereby incorporated by reference in its entirety.
[0123] Fractionation set: As used herein, a "fractionation set" refers to a set of nucleic acid molecules fractionated into sets / groups based on their differential binding affinity for a binder of the nucleic acid molecule. The binder preferentially binds to nucleic acid molecules containing nucleotides with epigenetic modifications. For example, when the epigenetic modification is methylation, the binder can be a methyl-binding domain (MBD) protein. In some embodiments, the fractionation set can include nucleic acid molecules belonging to a particular level / degree of epigenetic modification (i.e., epigenetic state). For example, the nucleic acid molecules can be fractionated into three sets: one set of highly methylated nucleic acid molecules (or hypermethylated nucleic acid molecules) that can be called the hypermethylated fractionation set or high fractionation set, another set of hypomethylated nucleic acid molecules (or hypomethylated nucleic acid molecules) that can be called the hypomethylated fractionation set or low fractionation set, and a third set of moderately methylated nucleic acid molecules that can be called the moderately methylated fractionation set or mid-fractionation set. In another example, the nucleic acid molecules can be fractionated based on the number of nucleotides with epigenetic modifications, and one fractionation set can have nucleic acid molecules with 9 methylated nucleotides, and another fractionation set can have non-methylated nucleic acid molecules (zero methylated nucleotides).
[0124] Polynucleotide: As used herein, "polynucleotide", "nucleic acid", "nucleic acid molecule", or "oligonucleotide" refers to a linear polymer of nucleosides joined by internucleoside linkages (including deoxyribonucleosides, ribonucleosides, or analogs thereof). Typically, a polynucleotide contains at least three nucleosides. The size of an oligonucleotide often ranges from a few monomer units, e.g., 3 - 4, to several hundred monomer units. When a polynucleotide is represented by a character string such as "ATGCCTG", unless otherwise specified, it is understood that the nucleotides are in the 5'→3' order from left to right, and in the case of DNA, "A" represents deoxyadenosine, "C" represents deoxycytidine, "G" represents deoxyguanosine, and "T" represents deoxythymidine. The letters A, C, G, and T can be used to refer to the base itself, the nucleoside, or the nucleotide containing the base.
[0125] Processing: As used herein, "processing" refers to a set of steps used to generate a library of nucleic acids suitable for sequencing. The set of steps can include, but is not limited to, nucleic acid fractionation, end repair, addition of sequencing adapters, tagging, and / or PCR amplification.
[0126] Processed sample: As used herein, "processed sample" refers to a sample that has been processed as described elsewhere herein. In some embodiments, a processed sample can refer to a subset of nucleic acids extracted from a processed first spike-in sample, and the subset of nucleic acids extracted from the first spike-in sample contains nucleic acid molecules from a sample of cell-free polynucleotides and from a first subset of fragment size control molecules. In other embodiments, a processed sample refers to a subset of a processed second spike-in sample, and the subset of the processed second spike-in sample contains nucleic acid molecules from a sample of cell-free polynucleotides and from a first subset and a second subset of fragment size control molecules.
[0127] Quantitative measurement: As used herein, "quantitative measurement" refers to an absolute or relative measurement. A quantitative measurement can be, but is not limited to, a number, a statistical measurement (e.g., frequency, mean, median, standard deviation, or quartiles), or a degree or relative amount (e.g., high, medium, or low). A quantitative measurement can be a ratio of two quantitative measurements. A quantitative measurement can be a linear combination of quantitative measurements. A quantitative measurement can be a normalized measurement.
[0128] Reference sequence: As used herein, "reference sequence" refers to a known sequence used for the purpose of comparison with an experimentally determined sequence. For example, the known sequence can be a whole genome, a chromosome, or any segment thereof. In some embodiments, the reference sequence can be at least about 20, at least about 50, at least about 100, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 1000, or more than 1000 nucleotides. The reference sequence can align with a single continuous sequence of a genome or chromosome, or can include discontinuous segments that align with different regions of a genome or chromosome. In some embodiments, the reference sequence can be a whole genome. Examples of reference sequences include, for example, the human genomes such as hG19 and hG38.
[0129] Sample: As used herein, "sample" means anything that can be analyzed by the methods and / or systems disclosed herein.
[0130] Second spike-in sample: As used herein, "second spike-in sample" is a sample in which a second subset of fragment size control molecules has been added to nucleic acids extracted from a first spike-in sample, and the nucleic acids extracted from the first spike-in sample contain nucleic acid molecules from a sample of cell-free polynucleotides and from a first subset of fragment size control molecules.
[0131] Sequencing: As used herein, "sequencing" refers to any of several techniques used to determine the sequence (e.g., the identity and order of monomer units) of a biomolecule, such as a nucleic acid like DNA or RNA. Examples of sequencing methods include targeted sequencing, single molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxynucleotide termination sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, duplex sequencing, cycle sequencing, single nucleotide extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification-PCR at low denaturation temperature (COLD-PCR), multiplex PCR, sequencing by reversible terminators, paired-end sequencing, short-read sequencing, single molecule sequencing, sequencing by synthesis, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, SOLiD™ sequencing, MS-PET sequencing, and combinations thereof, but not limited thereto. In some embodiments, sequencing can be performed by a genetic analysis instrument, such as a genetic analysis instrument commercially available from, for example, Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific, among many others.
[0132] Array information: As used herein, "array information" in the context of a nucleic acid polymer means the order and identity of monomer units (e.g., nucleotides, etc.) within that polymer.
[0133] Somatic mutation: As used herein, the terms "somatic mutation" or "somatic variation" are used interchangeably. These terms refer to mutations within the genome that occur after conception. Somatic mutations can occur in any cell of the body other than germ cells and are therefore not passed on to offspring.
[0134] Subject: As used herein, "subject" refers to an animal such as a mammalian species (e.g., human) or an avian species (e.g., bird), or another organism such as a plant. More specifically, a subject can be a vertebrate, such as a mammal like a mouse, primate, monkey, or human. Animals include farm animals (e.g., beef cattle, dairy cows, poultry, horses, pigs, etc.), sport animals, and companion animals (e.g., pets or service animals). A subject can be a healthy individual, an individual having or suspected of having a disease or predisposition to a disease, or an individual in need of or suspected of being in need of treatment. The terms "individual" or "patient" are intended to be interchangeable with "subject".
[0135] For example, a subject can be an individual diagnosed with having cancer, scheduled to receive cancer treatment, and / or having received at least one cancer treatment. A subject can be in a state of cancer remission. As another example, a subject can be an individual diagnosed with having an autoimmune disease. As another example, a subject can be a female individual during or planning to be pregnant who is diagnosed with or suspected of having a disease, e.g., cancer, autoimmune disease.
[0136] Third spike-in sample: As used herein, a "third spike-in sample" is a sample that has been added to a sample in which a third subset of fragment size control molecules has been processed. I. Overview
[0137] Fragment size distributions and fragmentation patterns in circulating cell-free DNA can also provide information regarding the origin of cell-free DNA and information regarding disease levels. However, biases can be introduced into the fragment size distribution and fragmentation patterns through the various extraction processes and methods used to read or analyze the fragment lengths and patterns (e.g., library preparation for next-generation sequencing). Thus, it is essential to have controls that can be used to understand the origin of the bias and at which step in the process the bias occurs, and also to use these controls to correct for the bias or to normalize for sample-by-sample and batch-level fragment size biases.
[0138] The present disclosure provides methods, compositions, and systems for normalizing the recovery rate of fragment lengths, optimizing assays to recover all / most of the fragments of a desired size, and as QC metrics. Such methods can include using fragment size control molecules as controls that can be added to samples of cell-free polynucleotides at different steps from extraction through sequencing. These controls can have predefined fragment lengths to mimic the natural size distribution of cell-free DNA. The identity of the fragment size control molecules added at different steps can be maintained by using different barcodes so that the fragment size bias at a particular step can be ascertained. These control molecules can be used in many applications including, but not limited to: (i) monitoring fragment size bias, (ii) optimizing processes to reduce fragment size bias, (iii) correcting or normalizing for fragment size bias introduced during the process by analyzing the recovery rate of the fragment size control molecules, (iv) evaluating the performance of an assay based on the recovery rate of the fragment size control molecules as a QC metric, (v) as a QC metric for analyzing any contamination of a sample, and (v) optimizing an assay to minimize any contamination.
[0139] Cancer formation and progression can result from both genetic and epigenetic modifications of deoxyribonucleic acid (DNA). The present disclosure provides methods for analyzing epigenetic modifications of DNA, such as cell-free DNA (cfDNA). Such epigenetic analysis includes methylation and DNA fragment patterns that can be identified by measuring changes in fragment length distribution or the frequency of fragment endpoints mapped to genomic positions. Such “fragmentome” analysis can be used alone or in combination with existing techniques to determine the presence or absence of a disease or condition, the prognosis of a diagnosed disease or condition, the treatment of a diagnosed disease or condition, or the outcome of a predicted treatment of a disease or condition. An example of such a disease or condition is cancer.
[0140] Circulating cell-free DNA (cfDNA) can be mainly short DNA fragments (e.g., having a length of about 100-400 base pairs and a mode of about 165 bp) shed from dying tissue cells into body fluids such as peripheral blood (plasma or serum). Analysis of cfDNA can reveal, in addition to cancer-related gene variants, the epigenetic footprint of dying cells and the signature of removal by phagocytes, thereby providing the aggregated nucleosome occupancy profile of this malignancy (e.g., tumor) and its microenvironmental components.
[0141] Components or factors that can contribute to plasma fragmentome signals (e.g., signals obtained from the analysis of cfDNA fragments) include malignant solid tumors that contain tumor - associated normal cells, epithelial, and stromal cells, immune cells, as well as vascular cells, and all of these can contribute to and be represented in a cfDNA sample (e.g., obtainable from a subject's body fluid). Thus, (i) the type of cell death during DNA destruction and associated chromatin condensation events, (ii) clearance mechanisms that may involve various types of phagocytosis regulated by the subject's immune system, (iii) non - malignant variations in blood composition that can be affected by the underlying combination of cell types in circulation, (iv) multiple origins or causes of non - malignant cell death in a given type of organ or tissue, and (v) the heterogeneity of cell types within cancer are included.
[0142] Cell - free DNA in the form of histone - protected complexes can be released by various host cells including neutrophils, macrophages, eosinophils, as well as tumor cells. Circulating DNA typically has a short half - life (e.g., about 10 - 15 minutes), and the liver is typically the major organ that removes circulating DNA fragments from the bloodstream. The accumulation of cfDNA in circulation can be due to an increase in cell death and / or activation, impairment of cfDNA clearance, and / or a decrease in endogenous DNase enzyme levels. Cell - free DNA circulating in a subject's bloodstream can typically be encapsulated in membrane - covered structures (e.g., apoptotic bodies) or form complexes with biomacromolecules (e.g., histones or DNA - binding plasma proteins). The processes of DNA fragmentation and subsequent transport can be analyzed with respect to their effects on the characteristics of cell - free DNA signals detected by fragmentome analysis.
[0143] In the cell nucleus (e.g., of a human), DNA is typically present in nucleosomes, which are structures that organize into a structure containing approximately 145 base pairs (bp) of DNA that wraps around a core histone octamer. The electrostatic and hydrogen bond interactions of DNA and nucleosomes can lead to energetically unfavorable bending of the DNA on the protein surface. Such bending can sometimes cause steric hindrance to other DNA-binding proteins and thus can help regulate access to DNA in the cell nucleus. The positioning of nucleosomes in a cell can vary dynamically (e.g., over time and across various cell states and conditions), for example, by DNA unwrapping and rewrapping. Since fragmentome signals can reflect nucleosome-protected DNA fragments derived from conformations affected by nucleosome units, the stability and dynamics of nucleosomes can influence such fragmentome signals. These nucleosome dynamics can result from a variety of factors, such as: (i) ATP-dependent remodeling complexes that can use the energy of ATP hydrolysis to slide nucleosomes, exchange or remove histones from chromatin fibers; (ii) histone variants that have properties different from those of canonical histones and can create local specific domains within chromatin fibers; (iii) histone chaperones that control the supply of free histones and can cooperate with chromatin remodelers in histone deposition and removal; (iv) post-translational modifications (PTMs) of histones (e.g., acetylation, methylation, phosphorylation, and ubiquitination) that can directly or indirectly affect the structure of chromatin; and (v) active transcription by transcription factors and RNA polymerase.
[0144] Therefore, the fragmentation signals or patterns in cfDNA can exhibit aggregated cfDNA signals resulting from multiple events related to the heterogeneity of chromatin organization across the genome. Such chromatin organization can vary depending on factors such as the overall cellular identity, metabolic state, local regulatory state, local gene activity, and DNA clearance mechanisms in dying cells. Moreover, the cell-free DNA fragmentome signal can only partially originate from the underlying chromatin structure of the contributing cells. Such cfDNA fragmentome signals can show more complex footprints of chromatin compaction and DNA protection from enzymatic digestion during cell death. Therefore, a chromatin map specific to a given cell type or cell lineage type can only partially contribute to the inherent heterogeneity of DNA accessibility due to changes in nucleosome stability, conformation, and composition at various stages of cell death or fragment transport. As a result, some nucleosomes may preferentially be present or absent in cell-free DNA (e.g., there may be a filtering mechanism that affects cfDNA clearance and releases it into the bloodstream), and these can depend on factors such as the mode and mechanism of death and the clearance of cell corpses.
[0145] The fragmentome signal can be generated in cells and released into the bloodstream as cfDNA as a result of nuclear DNA fragmentation during cellular processes such as apoptosis and necrosis. Such fragmentation can be produced as a result of different nuclease enzymes acting on DNA at different stages of the cell and can result in sequence-specific DNA cleavage patterns that can be analyzed in the cfDNA fragmentome signal. Classifying such cleavage patterns can be clinically relevant markers of the cellular environment (e.g., tumor microenvironment, inflammation, disease state, tumorigenesis, etc.).
[0146] The present disclosure provides methods, compositions, and systems for assessing and correcting fragment size bias in the analysis of nucleic acid molecules in a sample of polynucleotides (in some embodiments, the polynucleotides can be cell-free polynucleotides). These methods can be used in a variety of applications, for example, for prognosis determination, diagnosis, and / or monitoring of diseases.
[0147] Analysis of nucleic acid molecules in a sample of polynucleotides can be optimized and corrected for fragment size bias by measuring the recovery rate of fragment size control molecules. The fragment size control molecules can be synthetic nucleic acid molecules having a predefined fragment length. A set of fragment size control molecules can include nucleic acid molecules that include at least one subset of fragment size control molecules added to the sample of polynucleotides at a particular step. In some embodiments, the set of fragment size control molecules can include two or more subsets of fragment size control molecules, each subset being added at a different step of the assay. Each subset can include one or more groups of fragment size control molecules, each group having a different fragment length or size. For example, if there are three steps in an assay for which fragment size bias is to be analyzed, the set of fragment size control molecules can include three subsets (S1, S2, and S3) of fragment size control molecules, each subset being added to the sample prior to each step (i.e., S1 is added prior to step 1, S2 is added prior to step 2, and S3 is added prior to step 3).
[0148] In some embodiments, the fragment size control molecules can be synthetic nucleic acid molecules. In some embodiments, amplicons generated by amplifying a specific region (i.e., of a specific length) in any genome, plasmid, or vector, or a portion thereof, can be used as fragment size control molecules. The sequence and length of the fragment size control molecules can be known prior to analysis. In some embodiments, the fragment size control molecules are designed such that these molecules do not form any secondary structure. In some embodiments, the sequence of the fragment size control molecules is designed not to overlap with any human genomic region. Thus, by adding fragment size control molecules to a sample of polynucleotides and tracking the fragment size control molecules in a subset, the recovery rate of the fragment size control molecules can be analyzed, and thereby, in some embodiments, the fragment size bias can be estimated.
[0149] Accordingly, in one aspect, the present disclosure provides a method for analyzing nucleic acid molecules in a sample of polynucleotides, the method comprising: (a) adding a subset of fragment size control molecules to nucleic acid molecules in a sample of polynucleotides, thereby producing a first spike-in sample; (b) extracting nucleic acids from the first spike-in sample; (c) processing at least a subset of the extracted nucleic acids, thereby producing a processed sample, wherein processing comprises fractionating, tagging, and / or amplifying at least a subset of the first spike-in sample; (d) enriching at least a subset of the processed sample, thereby producing an enriched sample; (e) sequencing at least a subset of the enriched sample to generate a plurality of sequence reads; and (f) analyzing the plurality of sequence reads to generate a plurality of fragment size scores for the subset of fragment size control molecules. In some embodiments, the sample of polynucleotides can be a sample of cell-free polynucleotides (e.g., cell-free DNA).
[0150] In some embodiments, the method further includes comparing a plurality of fragment size scores to a plurality of fragment size thresholds. In some embodiments, the method can be used to optimize the analysis of nucleic acid molecules in a sample of polynucleotides based on the plurality of fragment size scores. In some embodiments, the method can be used to correct for fragment size bias in the analysis of nucleic acid molecules in a sample of polynucleotides using the plurality of fragment size scores. In some embodiments, the method can use quality control (QC) metrics, and the method is classified as successful if (i) at least one of the plurality of fragment size scores is within the corresponding fragment size threshold of the plurality of fragment size thresholds; or (ii) as a failure if at least one of the plurality of fragment size scores is not within the corresponding fragment size threshold of the plurality of fragment size thresholds.
[0151] Figure 1A is a schematic diagram of a method for analyzing nucleic acid molecules in a sample of cell-free polynucleotides according to an embodiment of the present disclosure. In some embodiments, the set of fragment size control molecules can include at least one subset of a set of predefined fragment size control molecules. In some embodiments, the set of fragment size control molecules can include one subset of fragment size control molecules. In 101A, a subset of fragment size control molecules (subset 1) can be added to a sample of polynucleotides whose fragment size bias can be analyzed to generate a first spike-in sample prior to extraction of the polynucleotides. In some embodiments, the fragment size control molecules are tagged with one or more tags or molecular barcodes that can be useful for identifying each individual fragment size control molecule (molecular identifier) within the subset and likewise the subset to which it belongs, i.e., subset 1 (subset identifier).
[0152] In some embodiments, a subset of fragment size control molecules can include one or more groups of fragment size control molecules. In some embodiments, one or more groups of fragment size control molecules can include nucleic acid molecules of different lengths and / or different sequences. In some embodiments, a group of fragment size control molecules can include fragment size control molecules of the same length and the same sequence. In some embodiments, each group can include fragment size control molecules of the same length but different sequences.
[0153] In 102A, nucleic acid molecules of the first spike-in sample are extracted. In 103A, the extracted nucleic acid molecules are processed to generate a library of nucleic acid molecules suitable for sequencing, including steps such as, but not limited to, end repair, addition of sequencing adapters, tagging, and / or PCR amplification. In some embodiments, the processing of the extracted nucleic acid molecules to generate a library of nucleic acid molecules includes steps such as, but not limited to, fractionation, end repair, addition of sequencing adapters, tagging, washing / purification, and / or PCR amplification. In 104A, the nucleic acids in the processed sample are enriched with respect to (i) the nucleic acids in the sample of cell-free polynucleotides belonging to a specific region of interest, and (ii) the fragment size control molecules of subset 1. In 105A, the enriched sample is sequenced so that the recovery rate of the fragment size control molecules in subset 1 can be analyzed. In 106A, the sequence reads generated from the sequencer are analyzed to measure the recovery rate of the subset 1 fragment size control molecules. In some embodiments, the recovery rate of the fragment size control molecules can be calculated using an orthogonal measurement of the relative abundance of the fragment size control molecules within a specific subset added to the sample. In some embodiments, the orthogonal measurement can be obtained from gel electrophoresis, qPCR, ddPCR, or PCR-free sequencing.
[0154] Figure 1B illustrates an embodiment as an example of method 100B for analyzing nucleic acid molecules in a sample of polynucleotides (e.g., the sample can be a cell-free DNA sample). In 101B, a subset of fragment size control molecules (subset 1) is added to the sample of polynucleotides prior to the extraction step to generate a first spike-in sample. The fragment size control molecules are added to the sample to monitor a fragment size bias that can be introduced at any of the steps. In some embodiments, the calculated fragment length distribution of the polynucleotides in the sample can be optimized using the fragment size bias estimated using the fragment size control molecules. In some embodiments, the set of fragment size control molecules can include at least one subset of predefined fragment size control molecules. In some embodiments, the set of fragment size control molecules can include one subset of fragment size control molecules (e.g., subset 1). In some embodiments, each subset includes at least one group of fragment size control molecules. In some embodiments, each group includes fragment size control molecules of the same length and the same sequence. In some embodiments, the length of the fragment size control molecules of one group is different from the length of the fragment size control molecules of other groups. In some embodiments, each group includes fragment size control molecules of the same length but different sequences. In some embodiments, the fragment size control molecules are tagged with one or more tags or molecular barcodes that can be useful for identifying each individual fragment size control molecule (molecular identifier) within the subset, as well as the subset to which it belongs, i.e., subset 1 (subset identifier). In some embodiments, each of the subsets of fragment size control molecules is present at an equimolar concentration. In some embodiments, each of the subsets of fragment size control molecules is present at a non-equimolar concentration.
[0155] In 102B, the nucleic acid molecules of the first spike-in sample are extracted. In 103B, the extracted nucleic acids are processed to generate a processed sample, and the processed sample includes (i) nucleic acid molecules in the sample of cell-free polynucleotides, and (ii) the fragment size control molecules of subset 1.
[0156] The nucleic acid molecules are processed in 103B to generate a library of nucleic acids suitable for sequencing (e.g., by a next-generation sequencer). The processing can include steps such as, but not limited to, end repair of the nucleic acids, addition of sequencing adapters, tagging, washing / purification, and / or amplification. In some embodiments, the processing can include steps such as, but not limited to, fractionation of the nucleic acids, end repair, addition of sequencing adapters, tagging, washing / purification, and / or amplification.
[0157] In some embodiments, fractionation includes fractionating the nucleic acid molecules based on their differential binding affinity for a binder that preferentially binds to nucleic acid molecules that include nucleotides having a chemical modification (e.g., methylation). Examples of binders include, but are not limited to, methyl-binding domain (MBD) and methyl-binding protein (MBP). Examples of MBP contemplated herein include: (a) MeCP2, a protein that preferentially binds 5-methyl-cytosine over unmodified cytosine; (b) RPL26, PRP8, and DNA mismatch repair protein MHS6, which preferentially bind 5-hydroxymethyl-cytosine over unmodified cytosine; (c) FOXK1, FOXK2, FOXP1, FOXP4, and FOXI3, which preferably bind 5-formyl-cytosine over unmodified cytosine (e.g., as described in Iurlaro et al., Genome Biol. 14, R119 (2013), which is hereby incorporated by reference in its entirety); and (d) antibodies specific for one or more methylated nucleotide bases are included, but not limited to these.
[0158] Fractionation can refer to the separation or sorting of nucleic acid molecules based on the characteristics of the nucleic acid molecules. Fractionation can be a physical fractionation of the molecules. Fractionation can involve separating nucleic acid molecules into groups or sets based on the level of epigenetic modification (e.g., epigenetic state). For example, nucleic acid molecules can be fractionated based on the level of methylation of the nucleic acid molecules. In some embodiments, the methods and systems used for fractionation can be found in PCT Patent Application No. PCT / US2017 / 068329, which is hereby incorporated by reference in its entirety. In those embodiments, nucleic acids are fractionated based on different levels of methylation (different numbers of methylated nucleotides). In some embodiments, nucleic acids can be fractionated into two or more fractionation sets (e.g., at least 3, 4, 5, 6, or 7 fractionation sets). In some embodiments, the fractionation sets represent nucleic acids having different degrees of modification (overrepresentation or underrepresentation of modification). Overrepresentation and underrepresentation can be defined by the number of modifications a nucleic acid has compared to the median number of modifications per strand in the population. For example, if the median number of 5-methylcytosine nucleotides in the nucleic acid molecules in a sample is 2, nucleic acid molecules containing more than two 5-methylcytosine residues are overrepresented in this modification, and nucleic acids having one or zero 5-methylcytosine residues are underrepresented. The effect of affinity separation is to enrich nucleic acids that are overrepresented in modification in the binding phase and nucleic acids that are underrepresented in modification in the non-binding phase (i.e., in solution). The nucleic acids in the binding phase can be eluted prior to subsequent processing. In some embodiments, each of the plurality of fractionation sets is differentially tagged. The tagged fractionation sets are then pooled together for collective sample preparation, enrichment, and / or sequencing. Differential tagging of the fractionation sets helps to identify nucleic acid molecules belonging to a particular fractionation set. The tags can be provided as components of adapters. Nucleic acid molecules in different fractionation sets are given different tags that can distinguish a member of one fractionation set from another. The tags linked to nucleic acid molecules in the same fractionation set can be the same or different from each other.However, even if they are different from each other, tags can have a common part of their sequences so that the molecules to which they are attached are identified as molecules of a specific fractionation set. For example, if the molecules of a spike-in sample are fractionated into two fractionation sets, P1 and P2, the molecules within P1 can be tagged with A1, A2, A3, etc., and the molecules within P2 can be tagged with B1, B2, B3, etc. Such a tagging system makes it possible to distinguish between fractionation sets and between molecules within a fractionation set.
[0159] In 104B, the processed sample is enriched to generate an enriched sample. The enriched sample contains (i) nucleic acid molecules in a sample of cell-free polynucleotides belonging to a specific region of interest, and (ii) fragment size control molecules of subset 1. In 105B, in order to calculate the fragment size bias, the enriched sample can be sequenced to generate a plurality of sequence reads, and the recovery rate of the fragment size control molecules in subset 1 can be analyzed. The obtained sequence information includes the sequences of the nucleic acid molecules and the tags attached to the nucleic acid molecules. From the sequences of the tags attached to the fragment size control molecules, the tags can be correlated with the individual fragment size control molecules and the subsets to which the fragment size control molecules belong. This information is used to analyze the recovery rate of the fragment size control molecules in the subset.
[0160] In 106B, array reads are analyzed to generate fragment size scores for the fragment size control molecules. The fragment size scores represent the recovery rate of the fragment size control molecules belonging to a particular group and a particular subset. The identity of the fragment size control molecules and the subset to which the fragment size control molecules belong is maintained through the use of tags or barcodes. In some embodiments, the fragment size scores can be estimated based on the number of fragment size control molecules belonging to a particular group and a particular subset. In some embodiments, the fragment size scores can be estimated based on the number of sequencing reads of the fragment size control molecules belonging to a particular group and a particular subset. In some embodiments, the fragment size scores can be estimated for a particular group with respect to a particular subset. In some embodiments, the fragment size scores can be estimated for each of the groups in every subset. In some embodiments, the fragment size score can be the total score of the fragment size control molecules used in a particular subset (i.e., with respect to all single scores representing the fragment size control molecules within the subset), or the total score of the fragment size control molecules added in one or more steps of the assay (i.e., the single scores representing the fragment size control molecules in one or more subsets added in the assay). In some embodiments, the fragment size score can be measured as the difference between the amount of a subset of fragment size control molecules added in a particular step and the amount of another subset of fragment size control molecules added in a different step, and in some embodiments, the amount can be an amount related to either mass (e.g., fg, pg, ng, μg), the number of sequencing reads, or the number of fragment size molecules (belonging to the subset or belonging to a particular group). In some embodiments, the fragment size score can be measured as the recovery rate (fraction or percentage) of the fragment size control molecules (belonging to the subset or belonging to a particular group), and the recovery rate can be calculated using an orthogonal measurement of the relative abundance of the fragment size control molecules within a particular subset added to the sample. In some embodiments, the orthogonal measurement can be obtained from gel electrophoresis, qPCR, ddPCR, or PCR-free sequencing.In some embodiments, the fragment size score can be measured as the difference between the amount of fragment size control molecules of a particular length and the amount of fragment size control molecules of a different length.
[0161] In 107B, the fragment size score is compared to a corresponding fragment size threshold to determine the fragment size bias in the assay. The fragment size threshold is a value or range of values of a predefined threshold used to evaluate or optimize the fragment size bias in the analysis of nucleic acid molecules in a sample. These thresholds can also be used to correct any fragment length distribution in any of the steps such as extraction, library preparation, enrichment, washing / purification, and sequencing. In some embodiments, the fragment size threshold can be estimated for a particular group in a particular subset. In some embodiments, each group in a subset has a particular fragment size threshold. In some embodiments, the fragment size threshold can be estimated for each of the groups in any subset. In some embodiments, the fragment size threshold can be the total threshold of the fragment size control molecules used in a particular subset (i.e., with respect to all of the individual thresholds of the fragment size control molecules within the subset), or the total threshold of the fragment size control molecules added in one or more steps of the assay (i.e., the individual thresholds of the fragment size control molecules in one or more subsets added in the assay). For example, a set of fragment size control molecules includes two subsets, subset 1 and subset 2. If each subset includes two groups, G11 and G12 (with respect to subset 1) and G21 and G22 (with respect to subset 2), then each of the four groups can have a fragment size score (S), S11 with respect to G11, S12 with respect to G12, S21 with respect to G21, and S22 with respect to G22, based on the recovery rate of the fragment size control molecules. Each group in a subset can have an individual predefined fragment size threshold. In this example, T11, T12, T21, and T22 are the fragment size thresholds for G11, G12, G21, and G22, respectively. The threshold can be a threshold with respect to a percentage or fraction, and the threshold can be a range of threshold values instead of a particular threshold value. For the method to be considered successful, at least one of these fragment size scores must be within its corresponding fragment size threshold. In some embodiments, if each of the multiple fragment size scores is within its corresponding fragment size threshold, the method is classified as having been successfully implemented.In some embodiments, the fragment size threshold can be a threshold related to a percentage or fraction. In some embodiments, the fragment size threshold can be a range of thresholds instead of a specific threshold value.
[0162] In some embodiments, prior to 103B, a second subset of fragment size control molecules (subset 2) may be added to the extracted nucleic acid to generate a second spike-in sample. In these embodiments, the second spike-in sample is processed to generate a processed sample. The processed sample includes (i) nucleic acid molecules in the sample of cell-free polynucleotides, and (ii) fragment size control molecules of subset 1 and subset 2. In some embodiments, prior to 104B, a third subset of fragment size control molecules (subset 3) may be added to the processed sample to generate a third spike-in sample. In these embodiments, the third spike-in sample is enriched to generate an enriched sample. The enriched sample includes (i) nucleic acid molecules in the sample of cell-free polynucleotides belonging to a specific region of interest, and (ii) fragment size control molecules of subset 1, subset 2, and subset 3. In some embodiments, prior to 105B, a fourth subset of fragment size control molecules (subset 4) may be added to the enriched sample to generate a fourth spike-in sample. The fourth spike-in sample includes (i) nucleic acid molecules in the sample of cell-free polynucleotides belonging to a specific region of interest, and (ii) fragment size control molecules of subset 1, subset 2, subset 3, and subset 4. In these embodiments, the fourth spike-in sample is sequenced to generate a plurality of sequence reads.
[0163] In another aspect, the present disclosure provides a method for analyzing nucleic acid molecules in a sample of polynucleotides, the method comprising: (a) adding a first subset of fragment size control molecules to nucleic acid molecules in a sample of polynucleotides, thereby producing a first spike-in sample; (b) extracting nucleic acids from the first spike-in sample; (c) adding a second subset of fragment size control molecules to the extracted nucleic acids, thereby producing a second spike-in sample; (d) processing at least a subset of the second spike-in sample, thereby producing a processed sample, wherein processing comprises fractionating, tagging, and / or amplifying at least a subset of the second spike-in sample; (e) adding a third subset of fragment size control molecules to the processed sample, thereby producing a third spike-in sample; (f) enriching at least a subset of the third spike-in sample, thereby producing an enriched sample; (g) adding a fourth subset of fragment size control molecules to at least a subset of the enriched sample, thereby producing a fourth spike-in sample; (h) sequencing the fourth spike-in sample to generate a plurality of sequence reads; (i) analyzing the plurality of sequence reads to generate a plurality of fragment size scores for the first subset of fragment size control molecules, the second subset of fragment size control molecules, the third subset of fragment size control molecules, and / or the fourth subset of fragment size control molecules; and (j) comparing the plurality of fragment size scores to a plurality of fragment size thresholds.
[0164] In some embodiments, the method further includes comparing a plurality of fragment size scores to a plurality of fragment size thresholds. In some embodiments, the method can be used to optimize the analysis of nucleic acid molecules in a sample of polynucleotides based on the plurality of fragment size scores. In some embodiments, the method can be used to correct for fragment size bias in the analysis of nucleic acid molecules in a sample of polynucleotides using the plurality of fragment size scores. In some embodiments, the method can use quality control (QC) metrics, and the method is considered successful if (i) at least one of the plurality of fragment size scores is within the corresponding fragment size threshold of the plurality of fragment size thresholds; or (ii) is considered a failure if at least one of the plurality of fragment size scores is not within the corresponding fragment size threshold of the plurality of fragment size thresholds. In some embodiments, the sample of polynucleotides can be a sample of cell-free polynucleotides (e.g., cell-free DNA).
[0165] Figure 2A is a schematic diagram of a method for analyzing nucleic acid molecules in a sample of polynucleotides (e.g., cell-free polynucleotides) according to an embodiment of the present disclosure. In some embodiments, the set of fragment size control molecules can include at least one subset of a predefined set of fragment size control molecules. In 201A, a first subset of fragment size control molecules (subset 1) is added to a sample of polynucleotides whose fragment size bias is to be analyzed to generate a first spike-in sample prior to extraction of the polynucleotides. In some embodiments, the fragment size control molecules can be tagged with one or more tags or barcodes that can help identify each individual fragment size control molecule (molecular identifier) within the subset and likewise the subset to which it belongs, i.e., subset 1 (subset identifier).
[0166] In some embodiments, a subset of fragment size control molecules can include one or more groups of fragment size control molecules. In some embodiments, one or more groups of fragment size control molecules can include nucleic acid molecules of different lengths and / or different sequences. In some embodiments, a group of fragment size control molecules can include fragment size control molecules of the same length and the same sequence. In some embodiments, each group can include fragment size control molecules of the same length but different sequences.
[0167] In 202A, cell-free nucleic acid molecules of the first spike-in sample are extracted. In 203A, a second subset (subset 2) of fragment size control molecules is added to the extracted nucleic acid molecules to generate a second spike-in sample. In 204A, the second spike-in sample is processed to generate a library of nucleic acid molecules suitable for sequencing, including steps such as, but not limited to, end repair, addition of sequencing adapters, tagging, washing / purification, and / or PCR amplification. In some embodiments, the processing of the first spike-in sample to generate a library of nucleic acid molecules includes steps such as, but not limited to, fractionation, end repair, addition of sequencing adapters, tagging, washing / purification, and / or PCR amplification. In 205A, a third subset (subset 3) of fragment size control molecules is added to the processed sample to generate a third spike-in sample. In 206A, the nucleic acids in the third spike-in sample are enriched with respect to (i) the nucleic acid molecules in the sample of cell-free polynucleotides belonging to a specific region of interest, and (ii) the fragment size control molecules of subset 1, subset 2, and subset 3. In 207A, a fourth subset (subset 4) of fragment size control molecules is added to the enriched sample to generate a fourth spike-in sample. In 208A, to determine the fragment size bias, the fourth spike-in sample is sequenced to analyze the recovery rates of the fragment size control molecules in subset 1, subset 2, subset 3, and subset 4. In 209A, the sequence reads generated from the sequencer are analyzed to measure the recovery rates of the subset 1, subset 2, subset 3, and subset 4 fragment size control molecules.
[0168] FIG. 2B illustrates an embodiment as an example of method 200B for analyzing nucleic acid molecules in a sample of polynucleotides (e.g., cell-free polynucleotides). In 201B, a subset (subset 1) of fragment size control molecules is added to a sample of cell-free polynucleotides prior to the extraction step to generate a first spike-in sample. The fragment size control molecules are added to the sample to monitor and / or correct for fragment size bias introduced at any of the steps. In some embodiments, the set of fragment size control molecules may include at least one subset of predefined fragment size control molecules. In some embodiments, the set of fragment size control molecules may include one subset (e.g., subset 1) of fragment size control molecules. In some embodiments, each subset includes at least one group of fragment size control molecules. In some embodiments, each group includes fragment size control molecules of the same length and the same sequence. In some embodiments, the length of the fragment size control molecules in one group is different from the length of the fragment size control molecules in other groups. In some embodiments, each group includes fragment size control molecules of the same length but different sequences. In some embodiments, the fragment size control molecules are tagged by one or more tags or barcodes that can help identify each individual fragment size control molecule (molecular identifier) within the subset and also the subset to which it belongs, i.e., subset 1 (subset identifier). In some embodiments, each of the subsets of fragment size control molecules is present at an equimolar concentration. In some embodiments, each of the subsets of fragment size control molecules is present at a non-equimolar concentration.
[0169] In 202B, nucleic acid molecules of the first spike-in sample are extracted. In 203B, a second subset (subset 2) of fragment size control molecules is added to the extracted nucleic acid molecules to generate a second spike-in sample. In some embodiments, the fragment size control molecules in the first subset are identified from the fragment size control molecules in the second subset by using different tags, i.e., the tags (subset identifiers) of the fragment size control molecules in subset 1 are different from the tags of the fragment size control molecules in subset 2. For example, all of the fragment size control molecules in subset 1 may have a subset identifier tag (S1), and all of the fragment size control molecules in subset 2 may have a subset identifier tag (S2). Apart from the subset identifier tag, each fragment size control molecule includes a molecular identifier tag that is used to identify each individual fragment size control molecule.
[0170] In 204B, the second spike-in sample is processed to generate a processed sample, and the processed sample includes (i) nucleic acid molecules in a sample of cell-free polynucleotides, and (ii) fragment size control molecules of subset 1 and subset 2. The second spike-in sample is processed in 204B to generate a library of nucleic acids suitable for sequencing (e.g., by a next-generation sequencer). The processing may include steps such as, but not limited to, end repair of nucleic acids, addition of sequencing adapters, tagging, washing / purification, and / or amplification.
[0171] In some embodiments, the process can include steps such as, but not limited to, fractionation of nucleic acids, end repair, addition of sequencing adapters, tagging, washing / purification, and / or amplification. In some embodiments, fractionation includes fractionating nucleic acid molecules based on differential binding affinities of nucleic acid molecules for binders that preferentially bind to nucleic acids containing nucleotides with chemical modifications (e.g., methylation). In some embodiments, nucleic acids can be fractionated into two or more fractionation sets (e.g., at least 3, 4, 5, 6, or 7 fractionation sets). In some embodiments, each of the plurality of fractionation sets is differentially tagged such that the set of tags used in the first fractionation set of the plurality of fractionation sets is different from the set of tags used in the second fractionation set of the plurality of fractionation sets. The tagged fractionation sets are then pooled together for collective sample preparation, enrichment, and / or sequencing. Differential tagging of fractionation sets helps to identify nucleic acid molecules belonging to a particular fractionation set. Tags can be provided as a component of an adapter in some embodiments. Nucleic acid molecules in different fractionation sets are given different tags that can distinguish members of one fractionation set from another. Tags linked to nucleic acid molecules in the same fractionation set can be the same or different from each other. However, even if different from each other, the tags can have a common portion of their sequences such that the molecules to which they are attached are identified as molecules of a particular fractionation set.
[0172] In some embodiments, tagging comprises attaching a set of tags to a nucleic acid to produce a population of tagged nucleic acids, wherein the tagged nucleic acids contain one or more tags. In some embodiments, the set of tags is attached to the nucleic acid by ligation of the adapter to the nucleic acid, and the adapter contains one or more tags.
[0173] In 205B, a third subset of fragment size control molecules (subset 3) is added to the processed sample to generate a third spike-in sample. In 206B, the third spike-in sample is enriched to generate an enriched sample. The enriched sample includes (i) nucleic acid molecules in the sample of cell-free polynucleotides belonging to a specific region of interest, and (ii) fragment size control molecules of subset 1, subset 2, and subset 3. In 207B, a fourth subset of fragment size control molecules (subset 4) is added to the enriched sample to generate a fourth spike-in sample. In 208B, to calculate the fragment size bias, the fourth spike-in sample is sequenced to generate a plurality of sequence reads, and the recovery rates of the fragment size control molecules in subset 1, subset 2, subset 3, and subset 4 can be analyzed. The obtained sequence information includes the sequences of the nucleic acid molecules and the tags attached to the nucleic acid molecules. From the sequences of the tags attached to the fragment size control molecules, the tags can be correlated with the individual fragment size control molecules and the subsets to which the molecules belong. This information is used to analyze the recovery rates of the fragment size control molecules in the subsets.
[0174] In 209B, array reads are analyzed to generate fragment size scores for fragment size control molecules. The fragment size score represents the recovery rate of fragment size control molecules belonging to a particular group and a particular subset. In some embodiments, the recovery rate of the fragment size control molecules can be calculated using an orthogonal measurement of the relative abundance of the fragment size control molecules within a particular subset added to the sample. In some embodiments, the orthogonal measurement can be obtained from gel electrophoresis, qPCR, ddPCR, or PCR-free sequencing. The identity of the fragment size control molecules, as well as the group and subset to which the fragment size control molecules belong, is maintained through the use of tags or barcodes. In some embodiments, the fragment size score can be estimated based on the number of fragment size control molecules belonging to a particular group and a particular subset. In some embodiments, the fragment size score can be estimated based on the number of sequencing reads of the fragment size control molecules belonging to a particular group and a particular subset. In some embodiments, the fragment size score can be estimated for a particular group with respect to a particular subset. In some embodiments, the fragment size score can be estimated for each of the groups in every subset. In some embodiments, the fragment size score can be the total score of the fragment size control molecules used in a particular subset (i.e., a single score representing the fragment size control molecules within the subset), or the total score of the fragment size control molecules added to one or more steps of the assay (i.e., a single score representing the fragment size control molecules in one or more subsets added in the assay). In some embodiments, the fragment size score can be measured as the difference between the amount of a subset of fragment size control molecules added at a particular step and the amount of another subset of fragment size control molecules added at a different step, and in some embodiments, the amount can be an amount related to either mass (e.g., fg, pg, ng, μg), the number of sequencing reads, or the number of fragment size molecules (belonging to the subset or belonging to a particular group).In some embodiments, the fragment size score can be measured as the recovery rate (fraction or percentage) of fragment size control molecules (belonging to a subset or belonging to a particular group), and the recovery rate can be calculated using an orthogonal measurement of the relative abundance of fragment size control molecules within a particular subset added to the sample. In some embodiments, the orthogonal measurement can be obtained from gel electrophoresis, qPCR, ddPCR, or PCR-free sequencing. In some embodiments, the fragment size score can be measured as the difference between the amount of fragment size control molecules of a particular length and the amount of fragment size control molecules of a different length.
[0175] In 210B, the fragment size score is compared to the corresponding fragment size threshold to determine the fragment size bias in the assay. The fragment size threshold is a value or range of default threshold values used to evaluate or optimize the fragment size bias in the analysis of nucleic acid molecules in a sample. These thresholds can also be used to correct any fragment length distribution in any of the steps such as extraction, library preparation, enrichment, washing / purification, and sequencing. In some embodiments, the fragment size threshold can be estimated for a particular group in a particular subset. In some embodiments, each group in the subset has a particular fragment size threshold. In some embodiments, the fragment size threshold can be estimated for each of the groups in any subset. In some embodiments, the fragment size threshold can be the total threshold of the fragment size control molecules used in a particular subset (i.e., the single threshold of the fragment size control molecules within the subset), or the total threshold of the fragment size control molecules added in one or more steps of the assay (i.e., the single threshold of the fragment size control molecules in one or more subsets added in the assay). For example, the set of fragment size control molecules includes two subsets, subset 1 and subset 2. If each subset contains two groups, G11 and G12 (for subset 1) and G21 and G22 (for subset 2), then each of the four groups can have a fragment size score (S), S11 for G11, S12 for G12, S21 for G21, and S22 for G22, based on the recovery rate of the fragment size control molecules. Each group in the subset can have an individual default fragment size threshold. In this example, T11, T12, T21, and T22 are the fragment size thresholds for G11, G12, G21, and G22, respectively. The threshold can be a threshold for a percentage or fraction, and the threshold can be a range of threshold values instead of a particular threshold value. For the method to be considered successful, at least one of these fragment size scores must be within its corresponding fragment size threshold. In some embodiments, if each of the multiple fragment size scores is within the corresponding fragment size threshold, the method is classified as having been successfully implemented.In some embodiments, the fragment size threshold can be a threshold related to a percentage or fraction. In some embodiments, the fragment size threshold can be a range of thresholds instead of a specific threshold value.
[0176] In some embodiments, the calculated fragment length distribution of the polynucleotides in a sample can be optimized using an analysis of the nucleic acid molecules with respect to the fragment size bias estimated using a fragment size control molecule. In some embodiments, a plurality of fragment size scores are used to correct for fragment size bias in the analysis of nucleic acid molecules. In some embodiments, the step of correcting for fragment size bias includes using a quantitative measurement derived from the ratio of the fragment size scores to the corresponding fragment size threshold of a group within a subset, and / or using a quantitative measurement derived from the ratio of the fragment size scores of two groups within a subset and / or the ratio of the fragment size scores of two subsets. In some embodiments, the method is classified as successful if at least one of the plurality of fragment size scores is within the corresponding fragment size threshold of the plurality of fragment size thresholds; or as failed if at least one of the plurality of fragment size scores is not within the corresponding fragment size threshold of the plurality of fragment size thresholds.
[0177] In some embodiments, a method of analyzing a nucleic acid molecule can be classified as successful if the fragment size scores of the fragment size control molecules for all groups in all subsets are within the corresponding fragment size thresholds for all groups. Otherwise, if any one of the fragment size scores is outside its corresponding fragment size threshold, the method of analyzing the nucleic acid molecule can be classified as a failure. For example, a set of fragment size control molecules includes two subsets of fragment size control molecules, subset 1 and subset 2. If each subset includes two groups, G11 and G12 (for subset 1) and G21 and G22 (for subset 2), then each of the four groups can have a fragment size score (S), S11 for G11, S12 for G12, S21 for G21, and S22 for G22, based on the recovery rate of the fragment size control molecules. Each group in a subset has an individual fragment size threshold. In this example, T11, T12, T21, and T22 are the fragment size thresholds for G11, G12, G21, and G22, respectively. If S11 < T11, S12 < T12, S21 < T21, and S22 < T22, the method can be classified as successful. If any one of the fragment size scores is outside its corresponding fragment size threshold (i.e., not inside), the method can be classified as a failure.
[0178] In some embodiments, fragment size control molecules can be used within individual samples to correct for fragment size biases in the observed fragment length distribution that are due to technical (artificial) rather than true biological causes. For example, cfDNA molecules derived from tumor cells typically exhibit a shorter average length than fragments derived from hematopoietic cells, and summary statistics including, but not limited to, the mean, median, mode, and IQR of the fragment length distribution may be used as features for the detection of malignancies, either alone or in combination with other features. However, these summary statistics can be affected by technical factors, transportation, sample and fluid handling, and artificial biases introduced during any of the steps (e.g., extraction, fractionation, tagging, amplification, washing / purification, enrichment, and sequencing), and these effects can confound the detection of malignancies. To avoid this potential source of confounding, the observed fragment length distribution may be adjusted based on the relative recovery rates of fragment size control molecules of different lengths. For example, if the expected recovery rate of a fragment size control molecule of a particular length is 75%, but the observed recovery rate of that particular length of fragment size control molecule in a given sample is only 25%, the density of the component of the fragment length distribution of the sample polynucleotide corresponding to the length of the fragment size control molecule may be increased three-fold to correct for the low recovery rate. Conversely, if the expected recovery rate of a fragment size control molecule of a particular length is 60%, but the observed recovery rate of that particular length of fragment size control molecule in a given sample is 80%, the density of the fragment length distribution of the sample polynucleotide may be decreased by 25% corresponding to the length of the fragment size control molecule.
[0179] Adjustments to the observed fragment length distribution can occur within a window around the length of the recovered fragment size control molecule. For example, if the fragment size control molecule has a length of 200 base pairs, the observed length distribution density may be adjusted for sample polynucleotides with lengths between 180 bp and 220 bp, or between 170 bp and 230 bp. The window need not be symmetric, and the adjustment applied need not be the same for every length of polynucleotide that falls within the window. As an example, the adjustment applied may use a Gaussian kernel centered on the length of the fragment size control molecule rather than a uniform kernel.
[0180] The expected recovery rates of fragment size control molecules of different lengths can be determined by performing one or more control experiments. A fragment size threshold may be set with respect to the recovery rate of any fragment size control molecule of a particular length. If the observed recovery rate of a fragment size control molecule is below the value of the fragment size threshold or does not fall within the range of the fragment size threshold, the sample may be considered to have poor quality control. These fragment size thresholds may vary for fragment size control molecules of different lengths.
[0181] In another aspect, the present disclosure provides a method for generating a sequencing library of a sample of cell-free polynucleotides, comprising: (a) adding a subset of fragment size control molecules to the sample, thereby producing a first spike-in sample; (b) extracting nucleic acids from the first spike-in sample; (c) processing at least a subset of the extracted nucleic acids, thereby producing a processed sample, wherein processing comprises fractionating, tagging, and / or amplifying at least a subset of the first spike-in sample; and (d) enriching at least a subset of the processed sample. In some embodiments, the method further comprises, prior to processing, adding a second subset of fragment size control molecules, thereby producing a second spike-in sample. In some embodiments, the method further comprises, prior to enriching, adding a third subset of fragment size control molecules, thereby producing a third spike-in sample. In some embodiments, the method further comprises (e) adding a fourth subset of fragment size control molecules, thereby producing a fourth spike-in sample.
[0182] In another aspect, the present disclosure provides a method for detecting contamination of a first sample by a second sample, for each first sample and second sample: (a) adding a subset of fragment size control molecules to generate a first spike-in sample, wherein the subset of fragment size control molecules added to the first sample is distinguishable from the subset of fragment size control molecules added to the second sample; (b) extracting nucleic acids from the first spike-in; (c) processing at least a subset of the extracted nucleic acids to produce a processed sample, wherein processing includes fractionating, tagging, and / or amplifying at least a subset of the first spike-in sample; (d) enriching at least a subset of the processed sample; (e) sequencing at least a subset of the enriched sample to generate a plurality of sequence reads; and (f) analyzing the plurality of sequence reads to generate one or more contamination scores for a subset of the fragment size control molecules. In some embodiments, the method further includes, prior to processing, adding a second subset of fragment size control molecules to generate a second spike-in sample, wherein the subset of fragment size control molecules added to the first sample is distinguishable from the subset of fragment size control molecules added to the second sample. In some embodiments, the method further includes, prior to enriching, adding a third subset of fragment size control molecules to generate a third spike-in sample, wherein the subset of fragment size control molecules added to the first sample is distinguishable from the subset of fragment size control molecules added to the second sample. In some embodiments, the method further includes, prior to sequencing, adding a fourth subset of fragment size control molecules to generate a fourth spike-in sample, wherein the subset of fragment size control molecules added to the first sample is distinguishable from the subset of fragment size control molecules added to the second sample.In some embodiments, a subset of the fragment size control molecules added to the first sample can be distinguished from a subset of the fragment size control molecules added to the second sample by using a subset identifier barcode in the first sample that is different from the subset identifier barcode used in other samples.
[0183] The contamination score refers to the score in the first sample that represents the presence of the fragment size control molecules added to the second sample. In some embodiments, the subset identifier barcode used in one subset in one sample can be different from the subset identifier used in other subsets in other samples. In these embodiments, the presence of fragment size control molecules belonging to different second samples can be identified from the sequence of the subset identifier barcode present in the fragment size control molecules. In some embodiments, the contamination score can be specific for each type / group of fragment size control molecules used in the subset (i.e., an individual contamination score for different lengths / groups of fragment size control molecules used in the subset), or it can be a total score representing different lengths / groups of fragment size control molecules. In some embodiments, the contamination score can be estimated based on the number of fragment size control molecules belonging to different second samples. In some embodiments, the contamination score can be estimated based on the number of sequencing reads of the fragment size control molecules belonging to different second samples. In some embodiments, the contamination score can be estimated based on the number of sequencing reads of the fragment size control molecules belonging to different second samples. In some embodiments, the contamination score of a sample can be estimated based on the fraction or percentage of the number of sequencing reads of the fragment size control molecules belonging to other samples to the number of sequencing reads of the fragment size control molecules added to that sample. In some embodiments, the contamination score of a sample can be estimated based on the fraction or percentage of the number of molecules of the fragment size control molecules belonging to other samples to the number of molecules of the fragment size control molecules added to that sample. In these embodiments, the fragment size control molecules added to the molecule can be identified from the subset identifier barcode of the fragment size control molecules.
[0184] In some embodiments, the method further includes comparing at least one or more contamination scores to at least one or more contamination thresholds. A contamination threshold refers to a value or range of values of a predefined threshold used to evaluate the contamination of a sample in the analysis of nucleic acid molecules in the sample. These thresholds can also be used to optimize the assay at any of the steps such as extraction, library preparation, enrichment, washing / purification, and sequencing. In some embodiments, the contamination threshold can be specific to each type / group of fragment size control molecules used in a subset (i.e., individual contamination thresholds for different lengths / groups of fragment size control molecules used in a subset). In some embodiments, the contamination threshold can be a total threshold for different lengths / groups of fragment size control molecules in a subset, or a total threshold for fragment size control molecules used in one or more subsets added in the assay. In some embodiments, each group in a subset has a specific contamination threshold. For example, a set of fragment size control molecules includes two subsets, subset 1 and subset 2. If each subset includes two groups, G11 and G12 (for subset 1) and G21 and G22 (for subset 2), then each of the four groups can have a contamination score (S), S11 for G11, S12 for G12, S21 for G21, and S22 for G22, based on the recovery rate of the fragment size control molecules. Each group in a subset can have an individual predefined contamination threshold. In this example, T11, T12, T21, and T22 are the contamination thresholds for G11, G12, G21, and G22, respectively. The threshold can be a threshold regarding a percentage or fraction, and the threshold can be a range of threshold values instead of a specific threshold value. For the method to be considered successful, at least one of these contamination scores must be within its corresponding contamination threshold. In some embodiments, if each of the plurality of contamination scores is within its corresponding contamination threshold, the method is classified as having been successfully implemented.In some embodiments, the method further includes classifying a first sample as being contaminated by a second sample if (i) at least one or more of the contamination scores are not within the corresponding contamination thresholds of one or more contamination thresholds; or (ii) not contaminated by the second sample if at least one or more of the contamination scores are within the corresponding contamination thresholds of one or more contamination thresholds. II. Fragment Size Control Molecules
[0185] Fragment size control molecules are nucleic acid molecules that are added to a sample of polynucleotides to evaluate and / or optimize the analysis of nucleic acid molecules in the sample. The fragment size control molecules can have two regions, a fragment size region and an identifier region. A set of fragment size control molecules can be classified into subsets (plural possible) of fragment size control molecules, and each subset of fragment size control molecules can be added at one or more different steps in an assay to evaluate the fragment length distribution of a sample of cell-free polynucleotides, QC metrics, and / or contamination of the sample at any of the steps of extraction, library preparation, enrichment, and / or sequencing. The lengths of these fragment size control molecules are predefined and, in some embodiments, can mimic the natural size distribution of cell-free DNA. Each subset of fragment size control molecules can be further classified into groups of fragment size control molecules based on the length of the fragment size region, and each group can have a different length of fragment size region. For example, a subset of fragment size control molecules can be classified into three groups based on the length of the fragment size region, where the first group can have a fragment size region of length 120 bp, the second group can have a fragment size region of length 160 bp, and the third group can have a fragment size region of length 320 bp. In some embodiments, the fragment size control molecules can be synthetic oligonucleotides. In some embodiments, the fragment size control molecules can have nucleic acid sequences that do not occur in nature. In some embodiments, the fragment size control molecules can have nucleic acid sequences that occur in nature. In some embodiments, amplicons generated by amplifying a specific region (i.e., of a specific length) in any genome, plasmid, or vector, or a portion thereof, can be used as fragment size control molecules. In some embodiments, the fragment size control molecules can have nucleic acid sequences corresponding to non-human genomes. For example, these molecules can have any of (i) sequences corresponding to regions of lambda phage DNA, (ii) sequences that do not occur in nature, and / or (iii) combinations of (i) and (ii). In some embodiments, the fragment size control molecules can include nucleotide analogs that do not occur in nature.
[0186] In another aspect, the present disclosure provides a set of fragment size control molecules comprising at least one subset of a predefined fragment size control molecule, wherein at least one subset of the predefined fragment size control molecule comprises a plurality of fragment size control molecules including a fragment size region. In some embodiments, the fragment size control molecule further comprises an identifier region. The fragment size region is a region of the fragment size control molecule that represents the length of the fragment size control molecule. In some embodiments, the at least one subset comprises at least one group of fragment size control molecules. In some embodiments, the fragment size regions of the fragment size control molecules in the group are regions of the same length. The length of the fragment size region of each group of fragment size control molecules can be different. For example, the subset of fragment size control molecules can be classified into three groups based on the length of the fragment size region. The first group can have a fragment size region with a length of 120 bp, the second group can have a fragment size region with a length of 160 bp, and the third group can have a fragment size region with a length of 320 bp. The length of the fragment size region can be between 10 bp and 1000 bp. In some embodiments, the length of the fragment size region in a group is different from the length of the fragment size region in other groups.
[0187] The identifier region is the region of the fragment size control molecule that is used to distinguish the fragment size control molecule from other fragment size control molecules. The identifier region is also used to distinguish fragment size control molecules from one subset from those of another subset. In some embodiments, the identifier region is present on one or both sides of the fragment size region. In some embodiments, the identifier region includes a molecular barcode. The molecular barcode serves as an identifier for the fragment size control molecule. The identifier region can be present on one or both sides of the fragment size region. The molecular barcode serves as an identifier for the fragment size control molecule. In some embodiments, the identifier region includes (i) a molecular identifier barcode, e.g., a barcode used to identify each fragment size control molecule and to distinguish one fragment size control molecule from another, and (ii) a subset identifier barcode, e.g., a barcode that acts as an identifier for the subset to which the fragment size control molecule belongs (i.e., whether the fragment size control molecule belongs to subset 1 or subset 2). The subset identifier barcode may be the same for all fragment size control molecules in the subset, and the subset identifier barcode for one subset may be different from that of another subset. For example, all fragment size control molecules in subset 1 can have the subset identifier tag S1, and all fragment size control molecules in subset 2 can have the subset identifier tag S2. In some embodiments, the subset identifier barcode for one subset in one sample can be different from the subset identifier barcode for the corresponding subset in another sample. For example, when evaluating sample contamination using fragment size control molecules, the subset identifier barcode for one subset is different from the subset identifier barcode for the corresponding subset used in all other samples within the same flow cell / batch. In some embodiments, the identifier region can include a sample index for distinguishing the fragment size control molecules of one sample from those of another sample.For example, the identifier region of a subset of fragment size control molecules added after the enrichment step may include a sample index sequence. In some embodiments, the identifier region can be attached to the fragment size region via ligation.
[0188] In some embodiments, the fragment size control molecules include one or more primer binding sites. In some embodiments, the primer binding site is present in the identifier region. In some embodiments, the identifier region may also have additional flow cell binding sites such as P5 and P7 that attach the fragment size control molecule to the surface of a next-generation sequencer (e.g., an Illumina sequencer).
[0189] In some embodiments, the fragment size regions of the fragment size control molecules in a group include the same oligonucleotide sequence. In some embodiments, the fragment size regions of the fragment size control molecules in a group include at least two distinguishable oligonucleotide sequences. In some embodiments, the fragment size region of the fragment size control molecules in a first subset includes an oligonucleotide sequence distinguishable from the oligonucleotide sequence of the fragment size region of the fragment size control molecules in a second subset.
[0190] In some embodiments, the fragment size region can be at least 10 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 120 bp, at least 150 bp, at least 200 bp, at least 250 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, or at least 1000 bp in length. In some embodiments, the length of the fragment size region can be between 10 bp and 1000 bp. In some embodiments, each of the subsets of fragment size control molecules is present at equimolar concentrations. In some embodiments, each of the subsets of fragment size control molecules is present at non-equimolar concentrations. In some embodiments, each of the groups of fragment size control molecules in the subset is present at equimolar concentrations. In some embodiments, each of the groups of fragment size control molecules in the subset is present at non-equimolar concentrations.
[0191] Figure 3 is a schematic diagram of fragment size control molecules suitable for use with some embodiments of the present disclosure. The set of fragment size control molecules described herein has lengths similar to those of the natural size distribution of cell-free DNA. In Figure 3, as an example, the set of fragment size control molecules is broadly divided into two subsets, subset 1 and subset 2. The fragment size control molecules in Figure 3 are double-stranded DNA molecules. For illustrative purposes, only one strand of the double-stranded fragment size control molecule is shown. In this embodiment, each subset is further divided into three groups, group 1, group 2, and group 3. In Figure 3, the fragment size regions of all the fragment size control molecules within each group have the same sequence, and the length of the fragment size region in one group is different from the length of the fragment size region in other groups. In Figure 3, the region of "------" represents the fragment size region of the fragment size control molecule. In this embodiment, the identifier regions are present on both sides of the fragment size region. The identifier regions have primer binding sites at both ends of the fragment size region. Similarly, herein, the sequence of the fragment size region of group 1 in subset 1 is the same as the sequence of the fragment size region of group 1 in subset 2. Similarly, the sequences of the fragment size regions of groups 2 and 3 in subset 1 are the same as the sequences of the fragment size regions of groups 2 and 3 in subset 2, respectively.
[0192] The identifier regions on both sides of the fragment size region have molecular identifier barcodes (MBs), while the subset identifier barcode (S) is present on only one side. The molecular barcode is used as an identifier for an individual fragment size control molecule, and each fragment size control molecule has a unique molecular barcode (i.e., molecule 1 has MB1 and MB2, molecule 2 has MB3 and MB4, molecule 3 has MB5 and MB6, etc.). The subset identifier barcode can be used as an identifier for the subset to which the fragment size control molecule belongs. Here, all the fragment size control molecules of subset 1 and subset 2 have subset identifier barcodes of S1 and S2, respectively. In this example, the subset identifier barcode is present on one side of the fragment size region.
[0193] In some embodiments, the molecular barcode can be present on one or both sides of the fragment size region. In some embodiments, the subset identifier barcode can be present on one or both sides of the fragment size region.
[0194] In this embodiment, the identifier region has primer binding sites on both sides of the fragment size region. Here, for subset 1, the fragment size control molecules of group 1, group 2, and group 3 have Pr1 and Pr2, Pr3 and Pr4, and Pr5 and Pr6 primer binding sites on both sides of the fragment size region, respectively. In some embodiments, the identifier region can have additional regions (primer binding sites) that facilitate the binding of one or more primers. In some embodiments, the primer binding sites of the identifier region in one subset are different from those in other subsets. In some embodiments, these primer binding sites are used to analyze the recovery rate of the fragment size control molecules. In some embodiments, instead of analyzing the recovery rate of the fragment size control molecules by sequencing, the recovery rate of the fragment size control molecules can be analyzed by digital droplet PCR (ddPCR), quantitative qPCR, or gel electrophoresis using primers that bind to these primer binding sites.
[0195] Figure 4 is a schematic diagram of fragment size control molecules suitable for use with some embodiments of the present disclosure. In Figure 4, as an example, the set of fragment size control molecules is broadly divided into two subsets, subset 1 and subset 2. The fragment size control molecules in Figure 4 are double-stranded DNA molecules. For illustrative purposes, only one strand of the double-stranded fragment size control molecule is shown. In this embodiment, each subset is further divided into three groups, group 1, group 2, and group 3. In Figure 4, the fragment size regions of all the fragment size control molecules within each group are regions of the same length, but the sequences of the fragment size regions within a group can be different. In Figure 4, the "------" and "~~~~" regions represent the fragment size regions of the fragment size control molecules. The "----" and "~~~~" regions indicate that each group contains two different sequences, but they are sequences of the same length. The length of the fragment size region in one group is different from the length of the fragment size region in other groups. In this embodiment, the identifier regions are present on both sides of the fragment size region. Similarly, here, the two sequences of the fragment size region of group 1 in subset 1 are the same as the two sequences of the fragment size region of group 1 in subset 2. Similarly, the two sequences of the fragment size regions of group 2 and group 3 in subset 1 are respectively the same as the two sequences of the fragment size regions of group 2 and group 3 in subset 2.
[0196] The identifier regions on both sides have molecular identifier barcodes (MB), while the subset identifier barcode (S) is present on only one side. The molecular barcode is used as an identifier for an individual fragment size control molecule, and each fragment size control molecule has a unique molecular barcode (i.e., molecule 1 has MB1 and MB2, molecule 2 has MB3 and MB4, molecule 3 has MB5 and MB6, etc.). The subset identifier barcode can be used as an identifier for the subset to which the fragment size control molecule belongs. Here, all the fragment size control molecules of subset 1 and subset 2 have subset identifier barcodes S1 and S2, respectively. In this example, the subset identifier barcode is present on one side of the fragment size region.
[0197] In some embodiments, the molecular barcode can be present on one or both sides of the fragment size region. In some embodiments, the subset identifier barcode can be present on one or both sides of the fragment size region.
[0198] In some embodiments, the fragment size control molecule can have a nucleic acid sequence that does not occur naturally. In some embodiments, the fragment size control molecule can have a nucleic acid sequence that occurs naturally. In some embodiments, the fragment size control molecule can have a nucleic acid sequence corresponding to a non-human genome. For example, these molecules can have any of (i) a sequence corresponding to a region of lambda phage DNA, (ii) a sequence that does not occur naturally, and / or (iii) a combination of (i) and (ii). In some embodiments, the fragment size control molecule can include non-naturally occurring nucleotide analogs.
[0199] In some embodiments, the sample of polynucleotide is a sample of DNA, a sample of RNA, a sample of cell-free polynucleotide, a sample of cell-free DNA, or a sample of cell-free RNA. In some embodiments, the sample of polynucleotide is a sample of cell-free DNA.
[0200] In some embodiments, the cell-free DNA is at least 1 ng, at least 5 ng, at least 10 ng, at least 15 ng, at least 20 ng, at least 30 ng, at least 50 ng, at least 75 ng, at least 100 ng, at least 150 ng, at least 200 ng, at least 250 ng, at least 300 ng, at least 350 ng, at least 400 ng, at least 450 ng, or at least 500 ng.
[0201] In some embodiments, the amount of the fragment size control molecule is at least 1 attomole, at least 2 attomoles, at least 5 attomoles, at least 10 attomoles, at least 15 attomoles, at least 20 attomoles, at least 50 attomoles, at least 75 attomoles, at least 100 attomoles, at least 1 femtomole, at least 2 femtomoles, at least 5 femtomoles, at least 10 femtomoles, at least 15 femtomoles, at least 20 femtomoles, at least 50 femtomoles, at least 75 femtomoles, at least 100 femtomoles, at least 125 femtomoles, at least 150 femtomoles, or at least 200 femtomoles, at least 300 femtomoles, at least 400 femtomoles, at least 500 femtomoles, at least 600 femtomoles, at least 700 femtomoles, at least 800 femtomoles, at least 900 femtomoles, at least 1 picomole, at least 2 picomoles, at least 5 picomoles, or at least 10 picomoles. In some embodiments, the amount of the fragment size control molecule can be between 1 attomole and 10 picomoles.
[0202] Additional embodiments of the present disclosure include compositions to which size fragment control molecules are added. For example, a cell-free DNA sample to which a size fragment control molecule is added is one embodiment of the present disclosure. Similarly, a number of compositions containing one or more different subsets of size fragment control molecules produced during the practice of the method are contemplated to be embodiments of the present disclosure. III. General Features of the Method A. Sample
[0203] The sample can be any biological sample isolated from a subject. The sample can be body tissue, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells (white blood cell or leucocyte), endothelial cells, tissue biopsy (e.g., biopsy known or suspected to be a solid tumor), cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial fluid or extracellular fluid (e.g., exudate from the intercellular space), gingival exudate, gingival sulcus exudate, bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. The sample can be blood and its fractions, and body fluids such as urine. Such samples can contain nucleic acids released from tumors. The nucleic acids can include DNA and RNA and can be in double-stranded and single-stranded forms. The sample can be in the form initially isolated from the subject or can be subjected to further processing such as removing or adding a constituent component, enriching one constituent component relative to another, or converting one form of nucleic acid to another, e.g., converting RNA to DNA or single-stranded nucleic acid to double-stranded. Thus, for example, a body fluid sample for analysis can be plasma or serum containing cell-free nucleic acids, e.g., cell-free circulating DNA (cfDNA).
[0204] In some embodiments, the sample volume of the body fluid collected from the subject depends on the desired depth of reads for the region to be sequenced. Examples of volumes are from about 0.4 to 40 milliliters (mL), from about 5 to 20 mL, from about 10 to 20 mL. For example, the volume can be about 0.5 mL, about 1 mL, about 5 mL, about 10 mL, about 20 mL, about 30 mL, about 40 mL, or more. The volume of plasma from which the sample is taken is typically between about 5 mL and about 20 mL.
[0205] The sample can contain various amounts of nucleic acids. Typically, the amount of nucleic acid in a given sample is considered equivalent to a number of genome equivalents. For example, a sample of about 30 nanograms (ng) of DNA is about 10,000 (10 4 ) genome equivalents for a diploid human genome and, in the case of cfDNA, about 200 billion (2×10 11) may contain. Similarly, a sample of about 100 ng of DNA may contain about 30,000 haploid human genome equivalents, and in the case of cfDNA, about 600 billion individual molecules.
[0206] In some embodiments, the sample comprises nucleic acids from different sources, e.g., from cells and cell-free sources (e.g., blood samples). Typically, the sample comprises nucleic acids having mutations. For example, the sample may contain DNA having germline mutations and / or somatic mutations as needed. Typically, the sample contains DNA having cancer-related mutations (e.g., cancer-related somatic mutations).
[0207] Examples of amounts of cell-free nucleic acids in the sample prior to amplification are typically in the range of about 1 femtogram (fg) to about 1 microgram (μg), e.g., about 1 picogram (pg) to about 200 nanograms (ng), about 1 ng to about 100 ng, about 10 ng to about 1000 ng. In some embodiments, the sample contains up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. Optionally, the amount is at least about 1 fg, at least about 10 fg, at least about 100 fg, at least about 1 pg, at least about 10 pg, at least about 100 pg, at least about 1 ng, at least about 10 ng, at least about 100 ng, at least about 150 ng, or at least about 200 ng of cell-free nucleic acid molecules. In some embodiments, the amount is up to about 1 fg, about 10 fg, about 100 fg, about 1 pg, about 10 pg, about 100 pg, about 1 ng, about 10 ng, about 100 ng, about 150 ng, or about 200 ng of cell-free nucleic acid molecules. In some embodiments, the method comprises obtaining from the sample between about 1 fg and about 200 ng of cell-free nucleic acid.
[0208] Cell-free nucleic acids typically have a size distribution between about 100 nucleotides in length and about 500 nucleotides in length, with molecules between about 110 nucleotides in length and about 230 nucleotides in length representing about 90% of the molecules in the sample, a mode of about 168 nucleotides in length (in samples from human subjects), and a second minor peak in the range of about 240 nucleotides to about 440 nucleotides in length. In some embodiments, the cell-free nucleic acids are about 160 nucleotides to about 180 nucleotides in length, or about 320 nucleotides to about 360 nucleotides in length, or about 440 nucleotides to about 480 nucleotides in length.
[0209] In some embodiments, cell-free nucleic acids are isolated from a body fluid through a fractionation step in which the cell-free nucleic acids found in solution are separated from intact cells and other insoluble components of the body fluid. In some embodiments, fractionation includes techniques such as centrifugation or filtration. Alternatively, the cells in the body fluid can be lysed and the cell-free and cellular nucleic acids can be processed together. Generally, after addition of a buffer and washing steps, the cell-free nucleic acids can be precipitated, for example, by alcohol. In some embodiments, additional purification steps, such as silica-based columns, are used to remove contaminants or salts. For example, a large amount of non-specific carrier nucleic acid can be added to the entire reaction as needed to optimize aspects of the procedure, such as yield. After such treatment, the sample typically contains nucleic acids in various forms, including double-stranded DNA, single-stranded DNA, and / or single-stranded RNA. If desired, single-stranded DNA and / or single-stranded RNA are converted to the double-stranded form for inclusion in subsequent processing and analysis steps. B. Tagging
[0210] In some embodiments, nucleic acid molecules (from a sample of polynucleotides and fragment size control molecules) may be tagged with a sample index and / or a molecular barcode (commonly referred to as a “tag”). The tag can be incorporated into an adapter or otherwise ligated thereto, among other methods, by chemical synthesis, ligation (e.g., blunt-end ligation or sticky-end ligation), or overlap extension polymerase chain reaction (PCR). Such an adapter can ultimately be ligated to the target nucleic acid molecule. In other embodiments, one or more rounds of an amplification cycle (e.g., PCR amplification) are generally applied to introduce a sample index into the nucleic acid molecule using conventional nucleic acid amplification methods. Amplification can be performed in one or more reaction mixtures (e.g., multiple microwells in an array). The molecular barcode and / or sample index may be introduced simultaneously or in any order. In some embodiments, the molecular barcode and / or sample index are introduced before and / or after performing a sequence capture step. In some embodiments, only the molecular barcode is introduced prior to probe capture and the sample index is introduced after performing the sequence capture step. In some embodiments, both the molecular barcode and the sample index are introduced prior to performing a probe-based capture step. In some embodiments, the sample index is introduced after performing the sequence capture step. In some embodiments, the molecular barcode is incorporated into nucleic acid molecules (e.g., cfDNA molecules) in the sample through an adapter via ligation (e.g., blunt-end ligation or sticky-end ligation). In some embodiments, the sample index is incorporated into nucleic acid molecules (e.g., cfDNA molecules) in the sample through overlap extension polymerase chain reaction (PCR). Typically, a sequence capture protocol involves introducing a single-stranded nucleic acid molecule complementary to a targeted nucleic acid sequence, e.g., a coding sequence of a genomic region, and mutations in such regions are associated with cancer types.
[0211] In some embodiments, the tag can be located at one or both ends of the sample nucleic acid molecule. In some embodiments, the tag is a predefined, or random or semi-random sequence oligonucleotide. In some embodiments, the tag can be less than about 500, 200, 100, 50, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide in length. The tag can be ligated to the sample nucleic acid either randomly or non-randomly.
[0212] In some embodiments, each sample is uniquely tagged by a sample index or a combination of sample indexes. In some embodiments, each nucleic acid molecule of a sample or subsample is uniquely tagged by a molecular barcode or a combination of molecular barcodes. In other embodiments, multiple molecular barcodes may be used such that the molecular barcodes are not necessarily unique relative to each other (e.g., non-unique molecular barcodes). In these embodiments, the molecular barcodes are generally attached to individual molecules (e.g., by ligation), resulting in the creation of unique sequences that can be individually tracked by the combination of the molecular barcode and the sequence to which it can be attached. Detection of the non-uniquely tagged molecular barcodes in combination with endogenous sequence information (e.g., the starting (beginning) and / or ending (terminating) genomic location / position corresponding to the sequence of the original nucleic acid molecule in the sample, partial sequences of the sequence reads at one or both ends, the length of the sequence reads, and / or the length of the original nucleic acid molecule in the sample) typically allows for the assignment of a unique identity to a particular molecule. In some embodiments, detection of the non-uniquely tagged molecular barcodes in combination with endogenous sequence information (e.g., the starting (beginning) and / or ending (terminating) regions of the alignment of the sequence reads to a reference sequence, partial sequences of the sequence reads at one or both ends, the length of the sequence reads, and / or the length of the original nucleic acid molecule in the sample) typically allows for the assignment of a unique identity to a particular molecule. In some embodiments, the starting region includes the genomic start position of the sequencing read where it is determined that the 5' end of the sequencing read begins an alignment to the reference sequence, and the ending region includes the genomic end position of the sequencing read where it is determined that the 3' end of the sequencing read terminates an alignment to the reference sequence. In some embodiments, the starting region includes the first 1, first 2, first 5, first 10, first 15, first 20, first 25, first 30, or at least the first 30 base positions of the 5' end of the sequencing read that aligns to the reference sequence.In some embodiments, the end region includes the last 1, last 2, last 5, last 10, last 15, last 20, last 25, last 30, or at least the last 30 base positions of the 3' end of a sequencing read that aligns to a reference array.
[0213] The length or number of base pairs of individual array reads may also be used as needed to assign unique identity to a given molecule. As described herein, fragments from a single strand of a nucleic acid to which a unique identity has been assigned may enable identification of subsequent fragments from the parental and / or complementary strands.
[0214] In some embodiments, molecular barcodes are introduced at an expected ratio for the molecules in a sample of a set of identifiers (e.g., a combination of unique or non-unique molecular barcodes). One example format uses from about 2 to about 1,000,000 different molecular barcode sequences, or from about 5 to about 150 different molecular barcode sequences, or from about 20 to about 50 different molecular barcode sequences ligated to both ends of a target molecule. Alternatively, from about 25 to about 1,000,000 different molecular barcode sequences may be used. For example, 20 - 50 × 20 - 50 molecular barcode sequences (i.e., one of 20 - 50 different molecular barcode sequences can be attached to each end of a target molecule) can be used. The number of such identifiers is typically sufficient such that different molecules having the same start and end points are likely (e.g., at least 94%, 99.5%, 99.99%, or 99.999%) to receive different combinations of identifiers. In some embodiments, about 80%, about 90%, about 95%, or about 99% of the molecules have the same combination of molecular barcodes.
[0215] In some embodiments, the assignment of unique or non-unique molecular barcodes in the reaction is performed using, for example, the methods and systems described in U.S. Patent Application Nos. 20010053519, 20030152490, and 20110160078, and U.S. Patents Nos. 6,582,908, 7,537,898, 9,598,731, and 9,902,992, each of which is hereby incorporated by reference in its entirety. Alternatively, in some embodiments, only endogenous sequence information (e.g., start and / or stop positions, partial sequences at one or both ends of the sequence, and / or length) may be used to identify different nucleic acid molecules of a sample.
[0216] The subset identifier barcode can be part of the identifier region of the fragment size control molecule. The subset identifier barcode is used to identify the subset to which the fragment size control molecule belongs (i.e., whether the fragment size control molecule belongs to subset 1 or subset 2). The subset identifier barcode can be the same for all fragment size control molecules in the subset, and similarly, the subset identifier barcode of one subset can be different from the subset identifier barcodes of other subsets in the same sample and / or other samples used in the same batch. In some embodiments, the subset identifier barcode of one subset in one sample can be different from the subset identifier barcode of the corresponding subset in other samples. For example, when using fragment size control molecules to assess sample contamination, the subset identifier barcode of one subset is different from the subset identifier barcode of the corresponding subset used in all other samples within the same flow cell / batch. In some embodiments, the identifier region can include a sample index to distinguish the fragment size control molecules of one sample from those of other samples. For example, the identifier region of a subset of fragment size control molecules added after the enrichment step can include a sample index sequence.
[0217] In some embodiments, the assignment of unique or non-unique molecular barcodes in the reaction is performed using, for example, the methods and systems described in U.S. Patent Application Nos. 20010053519, 20030152490, and 20110160078, and U.S. Patents Nos. 6,582,908, 7,537,898, 9,598,731, and 9,902,992, each of which is hereby incorporated by reference in its entirety. C. Amplification
[0218] The sample nucleic acid and the fragment size control molecules can be amplified by PCR and other amplification methods using nucleic acid primers that anneal to primer binding sites in the adapters adjacent to the DNA molecules to be amplified. In some embodiments, the amplification method involves cycles of extension, denaturation, and annealing due to thermocycling or can be isothermal, for example, in transcription-mediated amplification. Other examples of amplification methods that can be utilized as needed include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustained sequence-based replication.
[0219] Typically, the amplification reaction generates a plurality of non-uniquely or uniquely tagged nucleic acid molecules having molecular barcodes and sample indices, sized in the range of about 150 nucleotides (nt) to about 700 nt, 250 nt to about 350 nt, or about 320 nt to about 550 nt. In some embodiments, the amplicons have a size of about 180 nt. In some embodiments, the amplicons have a size of about 200 nt. D. Enrichment
[0220] In some embodiments, the array is enriched prior to sequencing the nucleic acid. The enrichment is performed either specifically for a target region of interest or non-specifically (a "target sequence"), as needed. In some embodiments, the target region of interest can be enriched by nucleic acid capture probes ("baits") selected for one or more bait set panels using differential tiling and capture schemes. Differential tiling and capture schemes generally use bait sets of various relative concentrations to differentially tile (e.g., at different "resolutions") across genomic regions associated with the baits while undergoing a series of constraints (e.g., sequencer constraints such as sequencing load, usefulness of each bait), and capture the targeted nucleic acids at the desired level for downstream sequencing. These targeted genomic regions of interest optionally contain natural or synthetic nucleotide sequences of nucleic acid constructs. In some embodiments, biotinylated beads having probes for one or more regions of interest can be used to capture the target sequences and, optionally, then amplify those regions to enrich the regions of interest.
[0221] Array capture typically involves the use of oligonucleotide probes that hybridize to target nucleic acid sequences. In some embodiments, the probe set strategy involves tiling probes across the region of interest. Such probes can be, for example, about 60 to about 120 nucleotides in length. The set can have a depth of about 2×, 3×, 4×, 5×, 6×, 7×, 8×, 9×, 10×, 15×, 20×, 50×, or greater than 50× (e.g., depth of coverage). The effectiveness of array capture generally depends in part on the length of the sequence within the target molecule that is complementary (or nearly complementary) to the sequence of the probe. E. Sequencing
[0222] Optionally, sample nucleic acids and fragment size control molecules, which may or may not have been pre-amplified, are generally subjected to sequencing adjacent to the adapter. Sequencing methods or commercially available formats that may be used, for example, include Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing by synthesis, single molecule sequencing, nanopore-based sequencing, semiconductor sequencing, sequencing by ligation, sequencing by hybridization, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next-generation sequencing (NGS), Molecule Sequencing by Synthesis (SMSS) (Helicos), massively parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, PacBio, SOLiD, Ion Torrent, or sequencing using a Nanopore platform. The sequencing reaction can be carried out in a variety of sample processing apparatuses that may include multiple lanes, multiple channels, multiple wells, or other means for substantially simultaneously processing multiple sets of samples. The sample processing apparatus may also include multiple sample chambers to enable the simultaneous processing of multiple operations.
[0223] The sequencing reaction can be carried out on one or more nucleic acid fragment types or regions containing markers for cancer or other diseases. The sequencing reaction can also be carried out on any nucleic acid fragment present in the sample. The sequencing reaction can be carried out on at least about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100% of the genome. In other cases, the sequencing reaction can be carried out on less than about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100% of the genome.
[0224] The simultaneous sequencing reaction can be carried out using multiplex sequencing techniques. In some embodiments, the cell-free polynucleotides are sequenced in at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other embodiments, the cell-free polynucleotides are sequenced in less than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. The sequencing reactions are typically carried out sequentially or simultaneously. Subsequent data analysis is generally carried out on all or part of the sequencing reactions. In some embodiments, the data analysis is carried out on at least about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. In other embodiments, the data analysis may be carried out on less than about 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000, or 100,000 sequencing reactions. An example of read depth is about 1000 to about 50000 reads per locus (e.g., base position). F. Analysis
[0225] By sequencing, a plurality of sequence reads or reads can be generated. The sequencing reads or reads can include nucleotide sequence data of a length less than about 150 bases, or less than about 90 bases. In some embodiments, the reads are between about 80 bases and about 90 bases, for example, about 85 bases in length. In some embodiments, the methods of the present disclosure are applied to very short reads, for example, less than about 50 bases or about 30 bases in length. The sequence read data can include sequence data as well as meta-information. The sequence read data can be stored in any suitable file format, including, for example, a VCF file, a FASTA file, or a FASTQ file.
[0226] FASTA can refer to a computer program for searching a sequence database, and the name FASTA can also refer to a standard file format. FASTA is described, for example, by Pearson & Lipman, 1988, Improved tools for biological sequence comparison, PNAS 85:2444-2448, which is hereby incorporated by reference in its entirety. Sequences in the FASTA format begin with a single-line description line, followed by lines of sequence data. The description line is distinguished from the sequence data by a greater-than (">") symbol in the first column. The word following the ">" symbol is the identifier of the sequence, and the rest of the line is an explanation (both are optional). There may be no space between the ">" and the first character of the identifier. It is recommended that all lines of the string be less than 80 characters. The appearance of another line starting with ">" indicates the end of the sequence, which indicates the start of another sequence.
[0227] The FASTQ format is a text-based format for storing both biological sequences (usually nucleotide sequences) and their corresponding quality scores. The FASTQ format is similar to the FASTA format, but the quality scores follow the sequence data. For brevity, both the sequence characters and the quality scores are encoded with a single ASCII character. The FASTQ format is, for example, the de facto standard for storing the output of high-throughput sequencing instruments such as the Illumina Genome Analyzer, as described by Cock et al. (“The Sanger FASTQ file format for sequences with quality scores, and the Solexa / Illumina FASTQ variants,” Nucleic Acids Res 38 (6): 1767-1771, 2009), the entirety of which is incorporated herein by reference.
[0228] For FASTA and FASTQ files, the meta-information includes description lines and does not include the lines of sequence data. In some embodiments, for FASTQ files, the meta-information includes the quality scores. For FASTA and FASTQ files, the sequence data begins after the description lines and generally uses some subset of the IUPAC ambiguity codes and may be present with a “-” as necessary. In one embodiment, the sequence data can use the A, T, C, G, and N characters and may include a “-” or, if necessary, a U (e.g., to represent a gap or uracil) as necessary.
[0229] In some embodiments, at least one master array read file and output file are stored as plain text files (e.g., using an encoding such as ASCII; ISO / IEC 646; EBCDIC; UTF-8; or UTF-16). The computer system provided by the present disclosure may include a text editor program capable of opening plain text files. A text editor program may refer to a computer program that can present the contents of a text file (e.g., a plain text file) on a computer screen and allow a person to edit the text (e.g., using a monitor, keyboard, and mouse). Examples of text editors include, but are not limited to, Microsoft Word, emacs, pico, vi, BBEdit, and TextWrangler. The text editor program may be capable of displaying the plain text file on a computer screen and presenting the meta information and array reads in a format that can be read by a person (e.g., using alphanumeric characters that are not binary coded and can be used when printed or written by a person).
[0230] The methods have been considered in relation to FASTA or FASTQ files, but using the methods and systems of the present disclosure, any suitable array file format, including for example a file in Variant Call Format (VCF) format, can be compressed. A typical VCF file can include a header section and a data section. The header can contain any number of meta-information lines, each starting with the "##" character and field definition lines separated by TAB start with a single "#" character. The field definition lines name eight mandatory columns, and the body section contains lines of data that group the columns defined by the field definition lines. The VCF format is described, for example, by Danecek et al. ("The variant call format and VCF tools," Bioinformatics 27(15):2156-2158, 2011), which is hereby incorporated by reference in its entirety. The header section can be treated as meta-information for writing to the compressed file, and the data section can be treated as lines, each of which can be stored in the master file only if unique.
[0231] Some embodiments provide an assembly of sequence reads. In assembly by alignment, for example, sequence reads are aligned to each other or to a reference sequence. By aligning each read in turn to the reference genome, all of the reads are placed in relation to each other to create an assembly. Additionally, aligning or mapping sequence reads to a reference sequence can also be used to identify variant sequences within the sequence reads. Identifying variant sequences can be used in combination with the methods and systems described herein to further assist in the diagnosis or prognosis of a disease or condition or to guide treatment decisions.
[0232] In some embodiments, any or all of the steps are automated. Alternatively, the methods of the present disclosure can be fully or partially embodied in one or more dedicated programs, each written in a compiled language such as C++ as needed, then compiled and distributed as binary. The methods of the present disclosure can be implemented fully or partially as modules within an existing array analysis platform or by calling functions within an existing array analysis platform. In some embodiments, the methods of the present disclosure include some steps that are automatically invoked in response to a single starting queue (e.g., one or a combination of induced events supplied by human work, another computer program, or a machine). Thus, the present disclosure provides a method in which any or any combination of steps can occur automatically in response to a queue. "Automatically" generally means without human input, influence, or interaction (e.g., only in response to human work before the original or queue).
[0233] The methods of the present disclosure can also include various forms of output, including accurate and sensitive interpretation of a nucleic acid sample of interest. The output of the search can be provided in the format of a computer file. In some embodiments, the output is a FASTA file, a FASTQ file, or a VCF file. It is possible to process the output to create a text file or an XML file containing sequence data such as the sequence of nucleic acids aligned against the sequence of a reference genome. In other embodiments, the processing results in an output containing coordinates or strings that describe one or more mutations in the nucleic acid of interest compared to the reference genome. The alignment string is a Simple UnGapped Alignment Report (SUGAR), Verbose Useful Labeled Gapped Alignment Report (VULGAR), and Compact It may include an Idiosyncratic Gapped Alignment Report (CIGAR) (e.g., as described by Ning et al., Genome Research 11 (10): 1725-9, 2001, which is hereby incorporated by reference in its entirety). These strings can be implemented, for example, in the Exonerate sequence alignment software of the European Bioinformatics Institute (Hinxton, UK).
[0234] In some embodiments, for example, an alignment map such as a Sequence Alignment / Map (SAM) or Binary Alignment / Map (BAM) file that includes a CIGAR string is created (the SAM format is described, for example, by Li et al., “The Sequence Alignment / Map format and SAMtools,” Bioinformatics, 25 (16): 2078-9, 2009, which is hereby incorporated by reference in its entirety). In some embodiments, the CIGAR displays or includes an alignment having one gap per row. The CIGAR is a compressed pairwise alignment format reported as a CIGAR string. The CIGAR string may be useful for representing long (e.g., genomic) pairwise alignments. The CIGAR string may be used in the SAM format to represent the alignment of reads to a reference genome sequence.
[0235] A CIGAR string can follow an established motif. Before each character, a number giving the number of bases of the event is shown. The characters used can include M, I, D, N, and S (M = match; I = insertion; D = deletion; N = gap; S = substitution). A CIGAR string defines an array of matches and / or mismatches and deletions (or gaps). For example, the CIGAR string 2MD3M2D2M can indicate that the alignment contains two matches, one deletion (the number 1 is omitted to save space), three matches, two deletions, and two matches.
[0236] In some embodiments, a nucleic acid population is prepared for sequencing by enzymatically forming blunt ends on a double-stranded nucleic acid having single-stranded overhangs at one or both ends. In these embodiments, the population is typically treated with an enzyme having 5'-3' DNA polymerase activity and 3'-5' exonuclease activity in the presence of nucleotides (e.g., A, C, G, and T or U). Examples of enzymes or catalytic fragments thereof that can be used as needed include the Klenow large fragment and T4 polymerase. For a 5' overhang, the enzyme typically extends the recessed 3' end on the opposite strand until it is flush with the 5' end to produce a blunt end. For a 3' overhang, the enzyme generally digests from the 3' end to the 5' end of the opposite strand and sometimes beyond. If this digestion proceeds beyond the 5' end of the opposite strand, the gap can be filled by an enzyme having the same polymerase activity used for a 5' overhang. The formation of blunt ends on a double-stranded nucleic acid facilitates, for example, the attachment of adapters and subsequent amplification.
[0237] In some embodiments, the nucleic acid population is subjected to additional processing such as conversion from single-stranded nucleic acid to double-stranded nucleic acid and / or conversion from RNA to DNA (e.g., complementary DNA or cDNA). Nucleic acids in these forms are also ligated to adapters and amplified as needed.
[0238] Regardless of the presence or absence of pre-amplification, nucleic acids subjected to the process of forming the above blunt ends and, optionally, other nucleic acids in the sample can be sequenced to produce sequenced nucleic acids. The sequenced nucleic acids can refer to either the sequence of the nucleic acid (e.g., sequence information) or the nucleic acid whose sequence has been determined. Sequencing can be performed to directly provide sequence data of individual nucleic acid molecules in the sample or indirectly from the consensus sequence of the amplification products of individual nucleic acid molecules in the sample.
[0239] In some embodiments, double-stranded nucleic acids with single-stranded overhangs in the sample after blunt-end formation are ligated to adapters containing barcodes at both ends, and sequencing is used to determine the nucleic acid sequence and the inline barcodes introduced by the adapters. Blunt-end DNA molecules are ligated to the blunt ends of at least partially double-stranded adapters (e.g., Y-shaped or bell-shaped adapters) as needed. Alternatively, to facilitate ligation (e.g., for sticky-end ligation), tails of complementary nucleotides can be added to the blunt ends of the sample nucleic acids and the adapters.
[0240] A nucleic acid sample is typically contacted with a sufficient number of adapters such that the probability that any two copies of the same nucleic acid receive the same combination of adapter barcodes from adapters ligated to both ends is low (e.g., less than about 1% or 0.1%). Such use of adapters can enable the identification of families of nucleic acid sequences that have the same starting and ending points in a reference nucleic acid and are ligated to the same combination of barcodes. Such families can represent the sequences of the amplification products of the nucleic acids in the sample prior to amplification. The sequences of the family members can be compiled to derive a consensus nucleotide or a complete consensus sequence of the nucleic acid molecules in the original sample that have been modified by blunt-end formation and adapter attachment. In other words, the nucleotide occupying a particular position in the nucleic acid in the sample can be determined to be the consensus of the nucleotides occupying that corresponding position in the family member sequences. A family can include the sequences of one or both strands of a double-stranded nucleic acid. If the members of a family include the sequences of both strands of a double-stranded nucleic acid, the sequence of one strand can be converted to its complement for the purpose of compiling the sequences to derive a consensus nucleotide(s) or sequence. Some families include only a single member sequence. In this case, this sequence can be considered to be the sequence of the nucleic acid in the sample prior to amplification. Alternatively, families having only a single member sequence can be excluded from subsequent analysis.
[0241] Nucleotide variations (e.g., SNVs or indels) in a sequenced nucleic acid can be determined by comparing the sequenced nucleic acid to a reference sequence. The reference sequence is often a known sequence, e.g., a known full or partial genomic sequence from the subject (e.g., the full genomic sequence from a human subject). The reference sequence can be, for example, hG19 or hG38. The sequenced nucleic acid can represent, as described above, the sequence directly determined for the nucleic acid in the sample or the consensus of the sequences of amplification products of such nucleic acids. The comparison can be performed at one or more designated positions of the reference sequence. When the respective sequences are maximally aligned, a subset of the sequenced nucleic acid containing the positions corresponding to the designated positions of the reference sequence can be identified. Within such a subset, if any, it can be determined which of the sequenced nucleic acids contain a nucleotide variation at the designated position and, if necessary, if any, which contain the reference nucleotide (e.g., the same as in the reference sequence). If the number of sequenced nucleic acids in the subset containing the nucleotide variant exceeds a selected threshold, the variant nucleotide can be called at the designated position. The threshold can be a simple number of sequenced nucleic acids within the subset containing the nucleotide variant, e.g., at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, or, among other possibilities, a ratio of the sequenced nucleic acids within the subset containing the nucleotide variant, e.g., at least 0.5, 1, 2, 3, 4, 5, 10, 15, or 20. The comparison can be repeated for any designated position of interest within the reference sequence. Sometimes, the comparison can be performed for designated positions occupying at least about 20, 100, 200, or 300 consecutive positions of the reference sequence, e.g., about 20 - 500, or about 50 - 300 consecutive positions.
[0242] Additional details regarding nucleic acid sequencing, including the formats and applications described herein, are also available, e.g., Levy, each of which is hereby incorporated by reference in its entirety. et al., Annual Review of Genomics and Human Genetics, 17: 95-115 (2016); Liu et al., J. of Biomedicine and Biotechnology, Volume 2012, Article ID 251364:1-11 (2012); Voelkerding et al., Clinical Chem., 55: 641-658 (2009); MacLean et al., Nature Rev. Microbiol., 7: 287-296 (2009); Astier et al., J Am Chem Soc., 128(5):1705-10 (2006); U.S. Patent No. 6,210,891; U.S. Patent No. 6,258,568; U.S. Patent No. 6,833,246; U.S. Patent No. 7,115,400; U.S. Patent No. 6,969,488; U.S. Patent No. 5,912,148; U.S. Patent No. 6,130,073; U.S. Patent No. 7,169,560; U.S. Patent No. 7,282,337; U.S. Patent No. 7,482,120; U.S. Patent No. 7,501,245; U.S. Patent No. 6,818,395; U.S. Patent No. 6,911,345; U.S. Patent No. 7,501,245; U.S. Patent No. 7,329,492; U.S. Patent No. 7,170,050; U.S. Patent No. 7,302,146; U.S. Patent No. 7,313,308; and U.S. Patent No. 7,476,503 are also provided. IV. COMPUTER SYSTEM
[0243] The methods of the present disclosure can be implemented using or with the aid of a computer system. For example, (a) adding a first subset of fragment size control molecules to nucleic acid molecules in a sample of cell-free polynucleotides, thereby producing a first spike-in sample; (b) extracting nucleic acids from the first spike-in sample; (c) adding a second subset of fragment size control molecules to the extracted nucleic acids, thereby producing a second spike-in sample; (d) processing at least a subset of the second spike-in sample, thereby producing a processed sample, where processing includes fractionating, tagging, and / or amplifying at least a subset of the second spike-in sample; (e) adding a third subset of fragment size control molecules to the processed sample, thereby producing a third spike-in sample; (f) enriching at least a subset of the third spike-in sample, thereby producing an enriched sample; (g) adding a fourth subset of fragment size control molecules to at least a subset of the enriched sample, thereby producing a fourth spike-in sample; (h) sequencing the fourth spike-in sample to generate a plurality of sequence reads; (i) analyzing the plurality of sequence reads to generate a plurality of fragment size scores for the fragment size control molecules; and (j) comparing the plurality of fragment size scores to a plurality of fragment size thresholds. Such a method, which may include the above steps, can be implemented by a computer processor. In this embodiment, the system includes components for adding, fractionating, amplifying, enriching, and sequencing fragment size control molecules.
[0244] FIG. 5 shows a computer system 501 programmed or otherwise configured to implement the methods of the present disclosure. The computer system 501 can regulate various aspects of sample preparation, sequencing, and / or analysis. In some examples, the computer system 501 is configured to perform sample preparation and sample analysis, including nucleic acid sequencing.
[0245] Computer system 501 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 505, which can be a single-core or multi-core processor, or multiple processors for parallel processing. The computer system 501 also includes a memory or memory location 510 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 515 (e.g., hard disk), a communication interface 520 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 525 such as cache, other memory, data storage, and / or an electronic display adapter. The memory 510, storage unit 515, interface 520, and peripheral devices 525 communicate with the CPU 505 through a bus (solid line) such as a communication network or motherboard. The storage unit 515 can be a data storage unit (or data repository) for storing data. The computer system 501 can be operably coupled to a computer network 530 with the aid of the communication interface 520. The computer network 530 can be the Internet, the Internet and / or an extranet, or an intranet and / or an extranet that communicates with the Internet. The computer network 530 can be, in some cases, a telecommunications and / or data network. The computer network 530 can include one or more computer servers, thereby enabling distributed computing such as cloud computing. The computer network 530 can, in some cases, implement a peer-to-peer network with the aid of the computer system 501, thereby enabling devices coupled to the computer system 501 to operate as clients or servers.
[0246] The CPU 505 can execute a series of machine-readable instructions that can be embodied in a program or software. The instructions can be stored in a memory location such as the memory 510. Examples of operations performed by the CPU 505 can include fetch, decode, execute, and write-back.
[0247] The storage unit 515 can store files such as drivers, libraries, and saved programs. The storage unit 515 can store programs created by the user and recorded sessions, as well as outputs related to the program(s). The storage unit 515 can store user data, such as user preferences and user programs. The computer system 501 can, in some cases, include one or more additional external data storage units for the computer system 501, such as those located on a remote server that communicates with the computer system 501, for example, through an intranet or the Internet. Data can be transferred from one location to another using, for example, a communication network or physical data transfer (such as using a hard drive, thumb drive, or other data storage mechanism).
[0248] The computer system 501 can communicate with one or more remote computer systems through the network 530. For implementation, the computer system 501 can communicate with a remote computer system of a user (such as an operator). Examples of remote computer systems include personal computers (such as laptop PCs), slate or tablet PCs (such as Apple® iPad®, Samsung® Galaxy Tab), telephones, smartphones (such as Apple® iPhone®, Android-compatible devices, Blackberry®), or personal digital assistants. The user can access the computer system 501 via the network 530.
[0249] The methods described herein can be implemented by code executable by a machine (e.g., a computer processor) stored in an electronic storage location of a computer system 501, such as in memory 510 or an electronic storage unit 515. The machine-executable or machine-readable code can be provided in the form of software. In use, the code can be executed by a processor 505. In some cases, the code can be retrieved from the storage unit 515 and stored in the memory 510 for easy access by the processor 505. In some situations, the electronic storage unit 515 can be excluded and the machine-executable instructions are stored in the memory 510.
[0250] In one aspect, the present disclosure provides a non-transitory computer-readable medium comprising computer-executable instructions that, when executed by at least one electronic processor, implement at least a portion of a method, the method comprising: (a) adding a first subset of fragment size control molecules to nucleic acid molecules in a sample of cell-free polynucleotides, thereby producing a first spike-in sample; (b) extracting nucleic acids from the first spike-in sample; (c) adding a second subset of fragment size control molecules to the nucleic acids, thereby producing a second spike-in sample; (d) processing at least a subset of the second spike-in sample, thereby producing a processed sample, wherein processing comprises fractionating, tagging, and / or amplifying at least a subset of the second spike-in sample; (e) adding a third subset of fragment size control molecules to the processed sample, thereby producing a third spike-in sample; (f) enriching at least a subset of the third spike-in sample, thereby producing an enriched sample; (g) adding a fourth subset of fragment size control molecules to at least a subset of the enriched sample, thereby producing a fourth spike-in sample; (h) sequencing the fourth spike-in sample to generate a plurality of sequence reads; (i) analyzing the plurality of sequence reads to generate a plurality of fragment size scores for the fragment size control molecules; and (j) comparing the plurality of fragment size scores to a plurality of fragment size thresholds.
[0251] The code can be configured for use on a machine having a processor adapted to precompile and execute the code, or can be compiled during runtime. The code can be supplied in a programming language selected to enable the code to be executed in a precompiled manner or in an as-compiled manner each time.
[0252] Aspects of the systems and methods provided herein, such as computer system 501, can be embodied in programming. Various aspects of the technology can typically be thought of as a "product" or "manufacture" in the form of code and / or associated data that is executable by a machine (or processor) that is performed on or embodied in a type of machine-readable medium. Machine-executable code can be stored on an electronic storage unit, such a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. A "storage" type medium can include any or all of the tangible memories of a computer, processor, etc. that can provide non-transitory storage for software programming at any time, or modules associated therewith such as various semiconductor memories, tape drives, disk drives, etc.
[0253] All or part of the software can sometimes be communicated through the Internet or various other telecommunications networks. Such communication can make it possible, for example, to load software from one computer or processor to another computer or processor, such as from a management server or host computer to an application server computer platform. Thus, another type of medium that can have software elements includes light waves, radio waves, and electromagnetic waves such as those used across physical interfaces between local devices through wired and optical landline networks as well as various air links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be considered media having software. As used herein, unless restricted to non-transitory tangible "storage" media, terms such as computer or machine "readable media" refer to any medium involved in providing instructions for execution to a processor.
[0254] Thus, machine-readable media such as computer-executable code can take many forms including, but not limited to, tangible storage media, carrier wave media or physical transmission media. Non-volatile memory media includes, for example, any storage device in any computer such as optical or magnetic disks, such as those that can be used to implement a database as shown in the drawings. Volatile memory media includes dynamic memory such as the main memory of such a computer platform. Tangible transmission media includes coaxial cables; copper wires and fiber optics including wires that form part of a bus within a computer system. Carrier wave transmission media can take the form of electrical or electromagnetic signals, or acoustic or optical waves, such as those generated during high frequency (RF) and infrared (IR) data communications. Thus, common forms of computer-readable media include, for example, floppy (registered trademark) disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROM, DVD or DVD-ROM, any other optical media, punch cards, paper tapes, any other physical storage media with patterns of holes, RAM, ROM, PROM and EPROM, FLASH-EPROM, any other memory chip or cartridge, carrier wave transmitted data or instructions, cables or links that transport such carrier waves, or any other media that a computer can read programming code and / or data from. Many of these forms of computer-readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0255] Computer system 501 can include, or communicate with, for example, an electronic display that includes a user interface (UI) for providing one or more results of sample analysis. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0256] Additional details regarding computer systems and networks, databases, and computer program products can be found, for example, in Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Ed. (2011), Kurose, Computer Networking: A Top-Down Approach, Pearson, 7 th Ed. (2016), Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Ed. (2010), Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11 th Ed. (2014), Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Ed. (2006), and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011) as well. V. Applications A. Cancer and Other Diseases
[0257] In some embodiments, the methods and systems disclosed herein are used to identify customized or targeted therapies to treat a given disease or condition in a patient based on a classification of whether a nucleic acid variant is of somatic or germline origin. Typically, the disease considered is a type of cancer. Non-limiting examples of such cancers include biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary non-polyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, intraocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, hepatocarcinoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, pseudopapillary neoplasm, acinar cell carcinoma, prostate cancer, prostate adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, stomach cancer, gastric cancer, gastrointestinal stromal tumor (GIST), uterine cancer, or uterine sarcoma.
[0258] Non-limiting examples of other gene-based diseases, disorders, or conditions that may be evaluated as needed using the methods and systems disclosed herein include achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth disease (CMT), cat cry syndrome, Crohn's disease, cystic fibrosis, Dercum's disease, Down syndrome, Duane syndrome, Duchenne muscular dystrophy, factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher's disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (scid), sickle cell disease, spinal muscular atrophy, Tay-Sachs disease, thalassemia, trimethylaminuria, Turner syndrome, velocardiofacial syndrome, WAGR syndrome, Wilson's disease, and the like. B. Treatment and Related Administration
[0259] In certain embodiments, the methods disclosed herein relate to identifying and administering customized treatment to a patient, taking into account the context of nucleic acid variants of somatic or germline origin. In some embodiments, essentially any cancer treatment (e.g., surgery, radiation therapy, chemotherapy, and / or the like) may be included as part of these methods. Typically, the customized treatment includes at least one immunotherapy (or immunotherapy agent). Immunotherapy generally refers to methods of enhancing the immune response against a given type of cancer. In certain embodiments, immunotherapy refers to methods of enhancing the T cell response against a tumor or cancer.
[0260] In certain embodiments, the status of nucleic acid variants from a sample of a subject of somatic or germline origin can be compared to a database of control agent results from a reference population to identify customized or targeted therapy for that subject. Typically, the reference population includes patients having the same cancer or disease type as the test subject, and / or patients who are or have been treated with the same therapy as the test subject. If the nuclear variants and control agent results meet certain classification criteria (e.g., substantially or nearly match), customized or targeted therapy (or therapies) can be identified.
[0261] In certain embodiments, the customized therapies described herein are typically administered parenterally (e.g., intravenously or subcutaneously). Pharmaceutical compositions containing immunotherapeutic agents are typically administered intravenously. Certain therapeutic agents are administered orally. However, customized therapies (e.g., immunotherapeutic agents, etc.) may also be administered by methods such as buccal, sublingual, rectal, vaginal, intraurethral, topical, intraocular, intranasal, and / or intratympanic, and the administration may include tablets, capsules, granules, aqueous suspensions, gels, sprays, suppositories, plasters, ointments, etc.
[0262] Preferred embodiments of the present invention are shown and described herein, but it will be apparent to those skilled in the art that such embodiments are provided merely as examples. It is not intended that the present invention be limited by the specific examples provided herein. Although the present invention has been described in connection with the foregoing specification, the description of the embodiments and the figures herein are not meant to be construed in a limiting sense. Those skilled in the art will readily conceive of numerous variations, changes, and substitutions without departing from the present invention. Furthermore, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions described herein, which depend on various conditions and variables. It should be understood that various alternatives to the embodiments of the present disclosure described herein can be used in the practice of the present invention. Accordingly, it is intended that the present disclosure also cover any such alternatives, modifications, variations, or equivalents. The following claims define the scope of the present invention, and it is intended that methods and structures within the scope of these claims and their equivalents be covered thereby.
[0263] The foregoing disclosure has been described in some detail by way of example and as an example for purposes of clarity and understanding, but it will be apparent to those skilled in the art that various changes in form and detail can be made without departing from the true scope of the present disclosure and can be practiced within the scope of the appended claims. For example, the features, steps, elements, or other aspects of the methods, systems, computer-readable media, and / or components can all be used in various combinations.
[0264] Any patents, patent applications, websites, other publications or documents, accession numbers, etc. cited in this specification are hereby incorporated by reference in their entirety for all purposes to the same extent as if each individual item were specifically and individually indicated to be incorporated by reference. If different versions of sequences are associated with an accession number at different times, the version associated with the accession number as of the effective filing date of this application is intended. The effective filing date means the earlier of either the actual filing date or, if applicable, the filing date of the priority application that references the accession number. Similarly, if different versions of publications, websites, etc. are published at different times, unless otherwise specified, the version published on the date closest to the effective filing date of this application is intended. The present invention provides, for example, the following items. (Item 1) A set of fragment size control molecules comprising at least one subset of predetermined fragment size control molecules, wherein said at least one subset of predetermined fragment size control molecules comprises a plurality of fragment size control molecules including a fragment size region. (Item 2) The set of fragment size control molecules according to Item 1, wherein said at least one subset comprises at least one group of fragment size control molecules. (Item 3) The set of fragment size control molecules according to Item 2, wherein the fragment size region of the fragment size control molecules in the group is a region of the same length. (Item 4) The set of fragment size control molecules according to Item 2 or 3, wherein the length of the fragment size region in the first group of fragment size control molecules is different from the length of the fragment size region in the second group of fragment size control molecules. (Item 5) The set of fragment size control molecules according to Item 1, wherein said fragment size control molecules further comprise an identifier region. (Item 6) The set of fragment size control molecules according to Item 5, wherein said identifier region is present on one or both sides of said fragment size region. (Item 7) The set of fragment size control molecules according to item 5, wherein the identifier region contains a molecular barcode. (Item 8) The set of fragment size control molecules according to any of the above items, wherein the plurality of fragment control molecules contain one or more primer binding sites. (Item 9) The set of fragment size control molecules according to item 8, wherein the one or more primer binding sites are present in the identifier region. (Item 10) The set of fragment size control molecules according to any of items 2 to 4, wherein the fragment size region of the fragment size control molecules in the group contains the same oligonucleotide sequence. (Item 11) The set of fragment size control molecules according to any of items 2 to 9, wherein the fragment size region of the fragment size control molecules in the group contains at least two distinguishable oligonucleotide sequences. (Item 12) The set of fragment size control molecules according to any of the above items, wherein the fragment size region of the fragment size control molecules in the first subset of the predetermined fragment size control molecules contains an oligonucleotide sequence distinguishable from the oligonucleotide sequence of the fragment size region of the fragment size control molecules in the second subset of the predetermined fragment size control molecules. (Item 13) The set of fragment size control molecules according to any of the above items, wherein the length of the fragment size region is at least 10 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 120 bp, at least 150 bp, at least 200 bp, at least 250 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, or at least 1000 bp. (Item 14) A set of fragment size control molecules according to any of the above items, wherein the length of the fragment size region is between 10 bp and 1000 bp. (Item 15) A set of fragment size control molecules according to any of the above items, wherein each subset of the at least one subset of the predefined fragment size control molecules is present at an equimolar concentration. (Item 16) A set of fragment size control molecules according to any of the above items, wherein each subset of the at least one subset of the predefined fragment size control molecules is present at a non-equimolar concentration. (Item 17) A set of fragment size control molecules according to any of the above items, wherein each group of the at least one group of fragment size control molecules in the at least one subset is present at an equimolar concentration. set. (Item 18) A set of fragment size control molecules according to any of items 2 to 16, wherein each group of the at least one group of fragment size control molecules in the at least one subset is present at a non-equimolar concentration. (Item 19) a. A set of fragment size control molecules comprising at least one subset of predefined fragment size control molecules, wherein the at least one subset of predefined fragment size control molecules comprises a plurality of fragment size control molecules comprising a fragment size region; and b. A set of nucleic acid molecules in a sample of polynucleotides from a subject A population of nucleic acids. (Item 20) The population of nucleic acids according to item 19, wherein the at least one subset comprises at least one group of fragment size control molecules. (Item 21) The population of nucleic acids according to item 20, wherein the fragment size region of the fragment size control molecules in the group is a region of the same length. (Item 22) The population of nucleic acids according to item 20 or 21, wherein the length of the fragment size region in the first group of fragment size control molecules is different from the length of the fragment size region in the second group of fragment size control molecules. (Item 23) The population of nucleic acids according to item 19, wherein the fragment size control molecule further comprises an identifier region. (Item 24) The population of nucleic acids according to item 23, wherein the identifier region is present on one or both sides of the fragment size region. (Item 25) The population of nucleic acids according to item 23, wherein the identifier region contains a molecular barcode. (Item 26) The population of nucleic acids according to any one of items 19 to 25, wherein the plurality of fragment control molecules comprises one or more primer binding sites. (Item 27) The population of nucleic acids according to item 26, wherein the one or more primer binding sites are present in the identifier region. (Item 28) The population of nucleic acids according to any one of items 20 to 22, wherein the fragment size region of the fragment size control molecule in the group contains the same oligonucleotide sequence. (Item 29) The set of fragment size control molecules according to any one of items 20 to 27, wherein the fragment size region of the fragment size control molecule in the group contains at least two distinguishable oligonucleotide sequences. (Item 30) The population of nucleic acids according to any one of items 19 to 29, wherein the fragment size region of the fragment size control molecule in the first subset of the predetermined fragment size control molecules contains an oligonucleotide sequence distinguishable from the oligonucleotide sequence of the fragment size region of the fragment size control molecule in the second subset of the predetermined fragment size control molecules. (Item 31) The length of the fragment size region is at least 10 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least A population of nucleic acids according to any of items 19 to 30, which is 100 bp, at least 120 bp, at least 150 bp, at least 200 bp, at least 250 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, or at least 1000 bp. (Item 32) A population of nucleic acids according to any of items 19 to 31, wherein the length of the fragment size region is between 10 bp and 1000 bp. (Item 33) A population of nucleic acids according to any of items 19 to 32, wherein each subset of the at least one subset of the predetermined fragment size control molecules is present at an equimolar concentration. (Item 34) A population of nucleic acids according to any of items 19 to 33, wherein each subset of the at least one subset of the predetermined fragment size control molecules is present at a non - equimolar concentration. (Item 35) A population of nucleic acids according to any of items 19 to 34, wherein each group of the at least one group of the fragment size control molecules in the at least one subset is present at an equimolar concentration. (Item 36) A population of nucleic acids according to any of items 20 to 34, wherein each group of the at least one group of the fragment size control molecules in the at least one subset is present at a non - equimolar concentration. (Item 37) A method for analyzing nucleic acid molecules in a sample of polynucleotides, comprising: a) adding a subset of fragment size control molecules to the nucleic acid molecules in the sample of polynucleotides, thereby producing a first spike - in sample; b) extracting nucleic acids from the first spike - in sample; c) Processing at least a subset of the extracted nucleic acids to thereby produce a processed sample, wherein said processing comprises fractionating, tagging, and / or amplifying at least a subset of the first spike-in sample; d) Enriching at least a subset of the processed sample to thereby produce an enriched sample; e) Sequencing at least a subset of the enriched sample to generate a plurality of sequence reads; and f) Analyzing the plurality of sequence reads to generate a plurality of fragment size scores for the subset of fragment size control molecules A method comprising. (Item 38) The method according to item 37, further comprising, prior to c), adding a second subset of fragment size control molecules to thereby produce a second spike-in sample. (Item 39) The method according to item 38, further comprising, prior to d), adding a third subset of fragment size control molecules to thereby produce a third spike-in sample. (Item 40) The method according to item 39, further comprising, prior to e), adding a fourth subset of fragment size control molecules to thereby produce a fourth spike-in sample. (Item 41) A method for analyzing nucleic acid molecules in a sample of polynucleotides, comprising: a. Adding a first subset of fragment size control molecules to the nucleic acid molecules in the sample of polynucleotides to thereby produce a first spike-in sample; b. Extracting nucleic acids from the first spike-in sample; c. Adding a second subset of fragment size control molecules to the extracted nucleic acids to thereby produce a second spike-in sample; d. Processing at least a subset of the second spike-in sample, thereby producing a processed sample, wherein said processing comprises fractionating, tagging, and / or amplifying said at least a subset of the second spike-in sample; e. Adding a third subset of fragment size control molecules to the processed sample, thereby producing a third spike-in sample; f. Enriching at least a subset of the third spike-in sample, thereby producing an enriched sample; g. Adding a fourth subset of fragment size control molecules to at least a subset of the enriched sample, thereby producing a fourth spike-in sample; h. Sequencing the fourth spike-in sample to generate a plurality of sequence reads; and i. Analyzing the plurality of sequence reads to generate a plurality of fragment size scores for the first subset of fragment size control molecules, the second subset of fragment size control molecules, the third subset of fragment size control molecules, and / or the fourth subset of fragment size control molecules A method comprising. (Item 42) The method according to item 37 or 41, further comprising comparing the plurality of fragment size scores with a plurality of fragment size thresholds. (Item 43) The method according to item 42, further comprising optimizing the analysis of the nucleic acid molecules in the sample of polynucleotides based on the plurality of fragment size scores. (Item 44) The method according to item 42, further comprising correcting fragment size bias in the analysis of the nucleic acid molecules in the sample of polynucleotides using the plurality of fragment size scores. (Item 45) The method of claim 42, further comprising classifying the method as (i) successful if at least one of the plurality of fragment size scores is within the corresponding fragment size threshold of the plurality of fragment size thresholds; or (ii) failed if at least one of the plurality of fragment size scores is not within the corresponding fragment size threshold of the plurality of fragment size thresholds. (Item 46) A method for detecting contamination of a first sample by a second sample, comprising, for each of the first sample and the second sample: a. adding a subset of fragment size control molecules to generate a first spike-in sample, wherein the subset of fragment size control molecules added to the first sample is distinguishable from the subset of fragment size control molecules added to the second sample; b. extracting nucleic acid from the first spike-in; c. processing at least a subset of the extracted nucleic acid to produce a processed sample, wherein the processing includes fractionating, tagging, and / or amplifying at least a subset of the first spike-in sample; d. enriching at least a subset of the processed sample to produce an enriched sample; e. sequencing at least a subset of the enriched sample to generate a plurality of sequence reads; and f. analyzing the plurality of sequence reads to generate one or more contamination scores for the subset of fragment size control molecules A method comprising. (Item 47) The method of claim 46, further comprising, prior to c), adding a second subset of fragment size control molecules to produce a second spike-in sample, wherein the subset of fragment size control molecules added to the first sample is distinguishable from the subset of fragment size control molecules added to the second sample. (Item 48) Before (d), adding a third subset of fragment size control molecules to produce a third spike-in sample, wherein the subset of fragment size control molecules added to the first sample can be distinguished from the subset of fragment size control molecules added to the second sample, the method according to item 47, further comprising this step. (Item 49) Before (e), adding a fourth subset of fragment size control molecules to produce a fourth spike-in sample, wherein the subset of fragment size control molecules added to the first sample can be distinguished from the subset of fragment size control molecules added to the second sample, the method according to item 48, further comprising this step. (Item 50) A method for detecting contamination of a first sample by a second sample, for each of the first sample and the second sample: a. Adding a first subset of fragment size control molecules to generate a first spike-in sample, wherein the first subset of fragment size control molecules added to the first sample can be distinguished from the first subset of fragment size control molecules added to the second sample; b. Extracting nucleic acid from the first spike-in; c. Adding a second subset of fragment size control molecules to the extracted nucleic acid to produce a second spike-in sample, wherein the second subset of fragment size control molecules added to the first sample can be distinguished from the second subset of fragment size control molecules added to the second sample; d. Processing at least a subset of the extracted nucleic acid to produce a processed sample, wherein the processing includes fractionating, tagging, and / or amplifying at least a subset of the first spike-in sample; e. Adding a third subset of fragment size control molecules to the extracted nucleic acid, thereby producing a third spike-in sample, wherein the third subset of fragment size control molecules added to the first sample can be distinguished from the third subset of fragment size control molecules added to the second sample; f. Enriching at least a subset of the processed sample; g. Adding a fourth subset of fragment size control molecules to the extracted nucleic acid molecules, thereby producing a fourth spike-in sample, wherein the fourth subset of fragment size control molecules added to the first sample can be distinguished from the fourth subset of fragment size control molecules added to the second sample; h. Sequencing at least a subset of the enriched sample to generate a plurality of sequence reads; and i. Analyzing the plurality of sequence reads to generate one or more contamination scores for the subset of fragment size control molecules A method comprising. (Item 51) The method according to items 46 to 50, further comprising comparing at least one or more contamination scores with at least one or more contamination thresholds. (Item 52) The method according to item 51, further comprising classifying the first sample as (i) contaminated by the second sample if at least one or more of the contamination scores are not within the corresponding contamination thresholds of the one or more contamination thresholds; or (ii) not contaminated by the second sample if at least one or more of the contamination scores are within the corresponding contamination thresholds of the one or more contamination thresholds. (Item 53) The method according to any of items 37 to 52, wherein the subset of fragment size control molecules comprises a plurality of fragment size control molecules including a fragment size region. (Item 54) The method according to item 53, wherein the fragment size control molecule further comprises an identifier region. (Item 55) The method according to item 53 or 54, wherein the subset includes at least one group of fragment size control molecules. (Item 56) The method according to item 55, wherein the fragment size regions of the plurality of fragment size control molecules in the group are regions of the same length. (Item 57) The method according to item 55, wherein the length of the fragment size region in the first group of fragment size control molecules is different from the length of the fragment size region in the second group of fragment size control molecules. (Item 58) The method according to item 54, wherein the identifier region is present on one or both sides of the fragment size region. (Item 59) The method according to item 54, wherein the identifier region includes a molecular barcode. (Item 60) The method according to any of the above items, wherein the fragment size control molecule includes one or more primer binding sites. (Item 61) The method according to item 60, wherein the primer binding site is present in the identifier region. (Item 62) The method according to any of items 37 to 61, wherein the fragment size regions of the fragment size control molecules in the group include the same oligonucleotide sequence. (Item 63) The method according to item 55, wherein the fragment size regions of the fragment size control molecules in the group include at least two distinguishable oligonucleotide sequences. (Item 64) The method according to any of items 37 to 63, wherein the fragment size regions of the fragment size control molecules in the first subset of fragment size control molecules include an oligonucleotide sequence distinguishable from the oligonucleotide sequence of the fragment size regions of the fragment size control molecules in the second subset of fragment size control molecules. (Item 65) The method according to any one of items 37 to 64, wherein the length of the fragment size region is at least 10 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 120 bp, at least 150 bp, at least 200 bp, at least 250 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, or at least 1000 bp. (Item 66) The method according to any one of items 37 to 64, wherein the length of the fragment size region is between 10 bp and 1000 bp. (Item 67) The method according to any one of items 38 to 66, wherein each of the subsets of the fragment size control molecules is present at an equimolar concentration. (Item 68) The method according to any one of items 38 to 66, wherein each of the subsets of the fragment size control molecules is present at a non-equimolar concentration. (Item 69) The method according to any one of items 37 to 68, wherein each of the groups of the fragment size control molecules in the subset is present at an equimolar concentration. (Item 70) The method according to any one of items 37 to 68, wherein each of the groups of the fragment size control molecules in the subset is present at a non-equimolar concentration. (Item 71) The method according to item 37 or 41, wherein the fractionating comprises fractionating the nucleic acid molecules of at least the subset of the second spike-in sample into a plurality of fractionation sets. (Item 72) The method according to item 71, wherein the plurality of fractionation sets comprises nucleic acid molecules of the second spike-in sample fractionated based on the epigenetic modification level of the nucleic acid molecules of the second spike-in sample. (Item 73) The method according to item 37 or 41, wherein said tagging comprises attaching a set of tags to said nucleic acid to produce a population of tagged nucleic acids, said tagged nucleic acids comprising one or more tags. (Item 74) The method according to item 73, wherein the set of tags used in the first set of partitioning sets resulting from said partitioning is different from the set of tags used in the second set of partitioning sets of said plurality of partitioning sets. (Item 75) The method according to item 74, wherein the set of tags is attached to the nucleic acid by ligation of an adapter to the nucleic acid, said adapter comprising one or more tags. (Item 76) A method for producing a sequencing library of a sample of polynucleotides, comprising: a) adding a subset of fragment size control molecules to said sample to thereby produce a first spike-in sample; b) extracting nucleic acids from said first spike-in sample; c) processing at least a subset of said extracted nucleic acids to thereby produce a processed sample, said processing comprising fractionating, tagging, and / or amplifying at least a subset of said first spike-in sample; and d) enriching at least a subset of said processed sample A method comprising. (Item 77) The method according to item 76, further comprising, prior to c), adding a second subset of fragment size control molecules to thereby produce a second spike-in sample. (Item 78) The method according to item 77, further comprising, prior to d), adding a third subset of fragment size control molecules to thereby produce a third spike-in sample. (Item 79) The method according to item 78, further comprising the step of adding a fourth subset of fragment size control molecules, thereby producing a fourth spike-in sample. (Item 80) The method according to any one of items 37 to 79, wherein the concentration of the fragment size control molecule is between 1 attomole and 10 picomoles. (Item 81) The method according to any one of items 37 to 80, wherein the sample of polynucleotide is a sample of cell-free polynucleotide. (Item 82) The method according to item 81, wherein the sample of cell-free polynucleotide is selected from the group consisting of a sample of cell-free DNA and a sample of cell-free RNA. (Item 83) The method according to item 81, wherein the sample of cell-free polynucleotide is a sample of cell-free DNA. (Item 84) The method according to item 81, wherein the cell-free DNA is between 1 ng and 500 ng. (Item 85) A system comprising a computer-readable medium including, or accessible to, a controller including non-transitory computer-executable instructions that, when executed by at least one electronic processor, implement a method, the method comprising: a. adding a subset of fragment size control molecules to nucleic acid molecules in a sample of polynucleotide, thereby producing a first spike-in sample; b. extracting nucleic acid from the first spike-in sample; c. processing at least a subset of the extracted nucleic acid, thereby producing a processed sample, the processing including fractionating, tagging, and / or amplifying at least a subset of the first spike-in sample; d. enriching at least a subset of the processed sample, thereby producing an enriched sample; e. Sequencing at least a subset of the enriched sample to generate a plurality of sequence reads; and f. Analyzing the plurality of sequence reads to generate a plurality of fragment size scores for the subset of fragment size control molecules, a system comprising. (Item 86) The system according to item 85, further comprising the step of adding a second subset of fragment size control molecules before c) to thereby produce a second spike-in sample. (Item 87) The system according to item 86, further comprising the step of adding a third subset of fragment size control molecules before d) to thereby produce a third spike-in sample. (Item 88) The system according to item 87, wherein the method further comprises the step of adding a fourth subset of fragment size control molecules before e) to thereby produce a fourth spike-in sample. (Item 89) A system comprising a computer-readable medium including, or accessible to, a controller including non-transitory computer-executable instructions that, when executed by at least one electronic processor, implement a method, the method comprising: a. Adding a first subset of fragment size control molecules to nucleic acid molecules in a sample of polynucleotides to thereby produce a first spike-in sample; b. Extracting nucleic acid from the first spike-in sample; c. Adding a second subset of fragment size control molecules to the extracted nucleic acid to thereby produce a second spike-in sample; d. Processing at least a subset of the second spike-in sample to thereby produce a processed sample, the processing comprising fractionating, tagging, and / or amplifying the at least a subset of the second spike-in sample; e. Adding a third subset of fragment size control molecules to the processed sample, thereby producing a third spike-in sample; f. Enriching at least a subset of the third spike-in sample, thereby producing an enriched sample; g. Adding a fourth subset of fragment size control molecules to at least a subset of the enriched sample, thereby producing a fourth spike-in sample; h. Sequencing the fourth spike-in sample to generate a plurality of sequence reads; and i. Analyzing the plurality of sequence reads to generate a plurality of fragment size scores for the first subset of fragment size control molecules, the second subset of fragment size control molecules, the third subset of fragment size control molecules, and / or the fourth subset of fragment size control molecules A system comprising: (Item 90) The system according to item 85 or 89, wherein the method further comprises comparing the plurality of fragment size scores with a plurality of fragment size thresholds. (Item 91) The system according to item 90, wherein the method further comprises optimizing the analysis of the nucleic acid molecules in the sample of polynucleotides based on the plurality of fragment size scores. (Item 92) The system according to item 90, wherein the method further comprises correcting fragment size bias in the analysis of the nucleic acid molecules in the sample of polynucleotides using the plurality of fragment size scores. (Item 93) The system according to item 90, wherein the method further comprises classifying the method as (i) successful if each of the plurality of fragment size scores is within a corresponding fragment size threshold of the plurality of fragment size thresholds; or (ii) failed if at least one of the plurality of fragment size scores is not within the corresponding fragment size threshold of the plurality of fragment size thresholds. (Item 94) The system according to item 85 or 89, wherein the method further comprises comparing at least one of the plurality of fragment size scores with at least one of a plurality of contamination thresholds. (Item 95) (Item 96) The system according to item 94, further comprising classifying the sample as: (i) contaminated by another sample if at least one of the plurality of fragment size scores is outside the corresponding contamination threshold of the plurality of contamination thresholds; or (ii) not contaminated by another sample if at least one of the plurality of fragment size scores is within the corresponding contamination threshold of the plurality of contamination thresholds. (Item 96) The system according to any one of items 85 to 95, wherein the subset of fragment size control molecules comprises a plurality of fragment size control molecules including a fragment size region. (Item 97) The system according to item 96, wherein the fragment size control molecule further comprises an identifier region. (Item 98) The system according to item 96, wherein the subset of fragment size control molecules comprises at least one group of fragment size control molecules. (Item 99) The system according to item 98, wherein the fragment size regions of the plurality of fragment size control molecules in the group are regions of the same length. (Item 100) The system according to item 98, wherein the length of the fragment size region in the first group of fragment size control molecules is different from the length of the fragment size region in the second group of fragment size control molecules. (Item 101) The system according to item 97, wherein the identifier region is present on one or both sides of the fragment size region. (Item 102) The system according to item 97, wherein the identifier region comprises a molecular barcode. (Item 103) The system according to any one of items 85 to 102, wherein the fragment size control molecule comprises one or more primer binding sites. (Item 104) The system according to item 103, wherein the primer binding site is present in the identifier region. (Item 105) The system according to any one of items 85 to 104, wherein the fragment size region of the fragment size control molecule in the group contains the same oligonucleotide sequence. (Item 106) The system according to item 96, wherein the fragment size region of the fragment size control molecule in the group contains at least two distinguishable oligonucleotide sequences. (Item 107) The system according to item 96, wherein the fragment size region of the fragment size control molecule in the first subset of the fragment size control molecules contains an oligonucleotide sequence distinguishable from the oligonucleotide sequence of the fragment size region of the fragment size control molecule in the second subset of the fragment size control molecules. (Item 108) The system according to any one of items 85 to 107, wherein the length of the fragment size region is at least 10 bp, at least 50 bp, at least 60 bp, at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, at least 120 bp, at least 150 bp, at least 200 bp, at least 250 bp, at least 300 bp, at least 400 bp, at least 500 bp, at least 600 bp, at least 700 bp, at least 800 bp, at least 900 bp, or at least 1000 bp. (Item 109) The system according to any one of items 85 to 107, wherein the length of the fragment size region is between 10 bp and 1000 bp. (Item 110) The system according to any one of items 85 to 109, wherein each of the subsets of the fragment size control molecules is present at an equimolar concentration. (Item 111) The system according to any one of items 85 to 109, wherein each of the subsets of the fragment size control molecules is present at a non-equimolar concentration. (Item 112) The system according to any one of items 85 to 111, wherein each of the groups of the fragment size control molecules in the subset is present at a non-equal molar concentration. (Item 113) The system according to any one of items 85 to 111, wherein each of the groups of the fragment size control molecules in the subset is present at an equimolar concentration. (Item 114) The system according to item 85 or 89, wherein fractionating comprises fractionating the nucleic acid molecules of at least the subset of the second spike-in sample into a plurality of fractionation sets. (Item 115) The system according to item 114, wherein the plurality of fractionation sets comprises nucleic acid molecules of the second spike-in sample fractionated based on the epigenetic modification level of the nucleic acid molecules of the second spike-in sample. (Item 116) The system according to item 85 or 89, wherein tagging comprises attaching a set of tags to the nucleic acid to produce a population of tagged nucleic acids, the tagged nucleic acids comprising one or more tags. (Item 117) The system according to item 116, wherein the set of tags used in a first fractionation set of the plurality of fractionation sets resulting from fractionation is different from the set of tags used in a second fractionation set of the plurality of fractionation sets. (Item 118) The system according to item 116, wherein the set of tags is attached to the nucleic acid by ligation of an adapter to the nucleic acid, the adapter comprising one or more tags. (Item 119) A system comprising a computer-readable medium comprising or accessible to a controller that includes non-transitory computer-executable instructions that, when executed by at least one electronic processor, implement a method, the method comprising: a. Adding a subset of fragment size control molecules to a sample of polynucleotides, thereby producing a first spike-in sample; b. Extracting nucleic acids from the first spike-in sample; c. Processing at least a subset of the extracted nucleic acids, thereby producing a processed sample, wherein the processing includes fractionating, tagging, and / or amplifying at least a subset of the first spike-in sample; and d. Enriching at least a subset of the processed sample A system comprising. (Item 120) The system according to item 119, wherein the method further includes, prior to c), adding a second subset of fragment size control molecules, thereby producing a second spike-in sample. (Item 121) The system according to item 120, wherein the method further includes, prior to d), adding a third subset of fragment size control molecules, thereby producing a third spike-in sample. (Item 122) The system according to item 121, wherein the method further includes e) adding a fourth subset of fragment size control molecules, thereby producing a fourth spike-in sample. (Item 123) The system according to any one of items 85 to 122, wherein the concentration of the fragment size control molecules is between 1 attomole and 10 picomoles. (Item 124) The system according to any one of items 85 to 123, wherein the sample of polynucleotides is a sample of cell-free polynucleotides. (Item 125) The system according to item 123, wherein the sample of polynucleotides is selected from the group consisting of a sample of cell-free DNA and a sample of cell-free RNA. (Item 126) The system according to item 123, wherein the sample of cell-free polynucleotides is a sample of cell-free DNA. (Item 127) The system according to item 126, wherein the cell-free DNA is between 1 ng and 500 ng. (Item 128) Further comprising the step of creating a report containing information regarding the analysis of the nucleic acid molecule and / or information derived from the analysis, or the system or method according to any one of items 37 to 127. (Item 129) The method or system according to item 128, further comprising the step of transmitting the report to a third party such as the subject from whom the sample is derived or a medical practitioner. (Item 130) The method or system according to any of the above items, wherein the fragment size control molecule is a synthetic molecule. (Item 131) The method or system according to any of the above items, wherein the fragment size control molecule is generated as an amplicon by PCR amplification.
Claims
【Claim 1】 The invention described in the specification.