Methods and systems for cell-free nucleic acid processing
Patent Information
- Application Number
- HK62026126687
- Authority / Receiving Office
- HK · HK
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-28
- Filing Date
- 2026-07-27
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2044-04-11
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
(19) State Intellectual Property Office (12) Invention Patent Application (10) Publication Number (43) Publication Date (21) Application Number 202480038630.X (22) Application Date 2024.04.12 (30) Priority Data 63 / 496,347 2023.04.14 US 63 / 501,359 2023.05.10 US 63 / 511,441 2023.06.30 US 63 / 517,327 2023.08.02 US 63 / 588,120 2023.10.05 US 63 / 591,732 2023.10.19 US 63 / 594,365 2023.10.30 US 63 / 602,156 2023.11.22 US 63 / 549,294 2024.02.02 US 63 / 571,139 2024.03.28 US (85) PCT International Application Enters National Phase Date 2025.12.09 (86) PCT International Application Application Data PCT / US2024 / 024491 2024.04.12 (87) PCT International Application Publication Data WO2024 / 216205 EN 2024.10.17 (71) Applicant Adela Company Address California, USA (72) Inventors Daniel Diniz de Calvajo Abel Liken Julia Newton Min Jun Julia Kerlan Shen Shuyi Felicia Vencheri Zhang Junjun (74) Patent Agency Beijing Anxin Fangda Intellectual Property Agency Co., Ltd. 11262 Patent Attorney Xu Aiwen Wu Jingjing (51) Int.Cl. C12Q 1 / 6804 (2006.01) C12N 15 / 00 (2006.01) C12Q 1 / 6869 (2006.01) C12Q 1 / 6886 (2006.01) (54) Invention Title: Method and System for Cell-Free Nucleic Acid Processing (57) Abstract: This document discloses a method and system for targeted detection of circulating tumor DNA (ctDNA) molecules. In some cases, molecular sequencing libraries depleted of methylated DNA can be generated and used to reliably detect ctDNA in cell-free DNA samples at lower sequencing depth and lower cost than existing methods. Claims (10 pages), Description (55 pages), Drawings (32 pages), CN 121464223 A 2026.02.03 CN 1 21 46 42 23 A 1. A method comprising: (a) obtaining a first set of nucleic acid molecules from a cell-free sample of an object; (b) generating a second set of nucleic acid molecules from the first set of nucleic acid molecules or derivatives thereof, wherein the second set of nucleic acid molecules...(ii) The method is enriched at the methylation level relative to the first group of nucleic acid molecules; (c) enriching one or more targets from the second group of nucleic acid molecules or their derivatives to obtain a third group of nucleic acid molecules; and (d) sequencing the third group of nucleic acid molecules or their derivatives. 2. The method of claim 1, wherein the enrichment comprises contacting the second group of nucleic acid molecules or their derivatives with one or more nucleic acid capture probes. 3. The method of any one of claims 1 to 2, wherein the generation comprises contacting the first group of nucleic acid molecules or their derivatives with a methylated nucleic acid capture reagent. 4. The method of claim 3, wherein the methylated nucleic acid capture reagent is formed by incubating a methylation-binding molecule with (ii) a solid matrix. 5. The method of claim 4, wherein the solid matrix is a bead. 6. The method of any one of claims 4 to 5, wherein the solid matrix is a magnetic solid matrix. 7. The method of any one of claims 4 to 6, wherein the solid matrix comprises protein A. 8. The method of any one of claims 4 to 6, wherein the solid matrix comprises streptavidin. 9. The method of any one of claims 4 to 8, wherein the methylation-binding molecule is an antibody. 10. The method of any one of claims 4 to 9, wherein the methylation-binding molecule comprises biotin. 11. The method of any one of claims 4 to 10, wherein the methylation-binding molecule binds to methylated cytosine. 12. The method of any one of claims 4 to 11, further comprising amplifying the second group of molecules. 13. The method of claim 12, wherein the amplification is performed simultaneously with the binding of a subgroup of the first group of nucleic acids to the methylated nucleic acid capture reagent. 14. The method of any one of claims 1 to 13, wherein the sequencing is performed via a synthesis reaction. 15. The method of any one of claims 1 to 14, wherein the sequencing does not include bisulfite sequencing. 16. The method of any one of claims 1 to 15, wherein, prior to the sequencing, the third group of nucleic acid molecules undergoes one or more library preparation reactions. 17. The method of claim 16, further comprising, after the one or more library preparation reactions and prior to the sequencing, incubating the third group of nucleic acid molecules with a plurality of magnetic beads that interact with nucleic acids. 18. The method of claim 17, further comprising, after incubating the nucleic acid sample with a plurality of magnetic beads interacting with the nucleic acid, magnetically trapping the sample to remove the plurality of magnetic beads interacting with the nucleic acid. 19. The method of claim 18, further comprising performing additional magnetic trapping.20. The method of any one of claims 1 to 19, further comprising, prior to (b), adding a quantity of filler DNA to the first nucleic acid molecule. Claims 1 / 10 page 2 CN 121464223 A 21. A method comprising: (a) providing a plurality of nucleic acid molecules generated from a cell-free deoxynucleic acid (cfDNA) sample of a subject; (b) sequencing the plurality of nucleic acid molecules or derivatives thereof to generate a plurality of sequencing reads; (c) computer processing the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules; and (d) computer processing the methylation profiles to determine that the subject has cancer when the area under the receiver operating characteristic (AUROC) is at least about 91%, wherein the cancer is a low-shedding cancer. 22. The method of claim 21, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from a low-shedding tumor. 23. The method of claim 21, wherein the cancer is bladder cancer, breast cancer, endometrial cancer, prostate cancer, or kidney cancer. 24. The method of claim 23, wherein the cancer is endometrial cancer or prostate cancer. 25. A method comprising: (a) providing a plurality of nucleic acid molecules generated from a cell-free deoxyribonucleic acid (cfDNA) sample of a subject; (b) sequencing the plurality of nucleic acid molecules or derivatives thereof to generate a plurality of sequencing reads; (c) computer processing the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules; and (d) computer processing the methylation profiles to determine that the subject has cancer when the area under the receiver operating characteristic (AUROC) is at least about 94%, and wherein the cancer is an early-stage cancer. 26. The method of claim 25, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from an early-stage tumor. 27. The method of claim 26, wherein the early-stage tumor is a stage I tumor. 28. The method of claim 26, wherein the early-stage tumor is a stage II tumor. 29. The method of any one of claims 25 to 28, wherein the cancer is bladder cancer, breast cancer, colorectal cancer, endometrial cancer, esophageal cancer, head and neck cancer, hepatobiliary cancer, lung cancer, ovarian cancer, prostate cancer, or kidney cancer. 30. The method of claim 29, wherein the cancer is esophageal cancer, hepatobiliary cancer, or ovarian cancer. 31. A method comprising: (a) providing a plurality of nucleic acid molecules generated from a cell-free deoxyribonucleic acid (cfDNA) sample of a subject; (b) sequencing the plurality of nucleic acid molecules or derivatives thereof without bisulfite conversion to generate a plurality of sequencing reads; (c) computer processing the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules; and(d) The methylation spectrum is processed by a computer to determine that the subject has cancer, wherein the cancer is endometrial cancer, esophageal cancer, hepatobiliary cancer, ovarian cancer, prostate cancer, bladder cancer, breast cancer, colorectal cancer, head and neck cancer, lung cancer, pancreatic cancer, or kidney cancer. 32. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from endometrial cancer. 33. The method of claim 32, wherein the methylation spectrum is processed by a computer to determine that the subject has endometrial cancer when the area under the subject's operating characteristic curve (AUROC) is at least about 90%. 34. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from esophageal cancer. 35. The method of claim 34, wherein the methylation spectrum is processed by a computer to determine that the subject has esophageal cancer when the area under the subject's operating characteristic curve (AUROC) is at least about 99%. 36. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from hepatobiliary cancer. 37. The method of claim 36, wherein a computer processes the methylation spectrum to determine if the subject has hepatobiliary cancer when the area under the receiver operating characteristic (AUROC) is at least about 99%. 38. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from ovarian cancer. 39. The method of claim 38, wherein a computer processes the methylation spectrum to determine if the subject has ovarian cancer when the area under the receiver operating characteristic (AUROC) is at least about 97%. 40. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from prostate cancer. 41. The method of claim 40, wherein a computer processes the methylation spectrum to determine if the subject has prostate cancer when the area under the receiver operating characteristic (AUROC) is at least about 89%. 42. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from bladder cancer. 43. The method of claim 42, wherein the methylation spectrum is processed by a computer to determine that the subject has bladder cancer when the area under the receiver operating characteristic (AUROC) is at least about 95%. 44. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from breast cancer. 45. The method of claim 44, wherein the methylation spectrum is processed by a computer to determine that the subject has breast cancer when the area under the receiver operating characteristic (AUROC) is at least about 92%.46. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from colorectal cancer. 47. The method of claim 46, wherein a computer processes the methylation spectrum to determine if the subject has colorectal cancer when the area under the receiver operating characteristic (AUROC) is at least about 98%. 48. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from head and neck cancer. 49. The method of claim 48, wherein a computer processes the methylation spectrum to determine if the subject has head and neck cancer when the area under the receiver operating characteristic (AUROC) is at least about 96%. 50. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from lung cancer. 51. The method of claim 50, wherein a computer processes the methylation spectrum to determine if the subject has lung cancer when the area under the receiver operating characteristic (AUROC) is at least about 96%. 52. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from pancreatic cancer. 53. The method of claim 52, wherein the methylation profile is processed by a computer to determine that the subject has pancreatic cancer when the area under the receiver operating characteristic curve (AUROC) is at least about 99%. 54. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from renal cell carcinoma. Claims 3 / 10 Page 4 CN 121464223 A 55. The method of claim 54, wherein the methylation profile is processed by a computer to determine that the subject has renal cell carcinoma when the area under the receiver operating characteristic curve (AUROC) is at least about 91%. 56. The method of any one of claims 21 to 55, further comprising, prior to (b), adding a set of nucleic acid molecules not derived from the subject to the plurality of nucleic acid molecules. 57. The method of any one of claims 21 to 56, wherein the methylation profile is genome-wide. 58. The method of claim 57, wherein the methylation profile comprises a complete methylome. 59. The method of any one of claims 21 to 58, wherein (d) comprises a supervised machine learning method, wherein the supervised machine learning method is regression, support vector machine, tree-based method, neural network, or nearest neighbor method. 60. The method of any one of claims 21 to 58, wherein (d) comprises an unsupervised machine learning method, wherein the unsupervised machine learning method is clustering, neural network, principal component analysis, or matrix factorization. 61. The method of any one of claims 21 to 60, wherein the subject has previously received cancer treatment andAnd substantially no cancer, wherein (d) includes determining that the object has a recurrence of the cancer. 62. The method of any one of claims 21 to 61, the method further comprising, prior to (b), adding an amount of filler DNA to the plurality of nucleic acid molecules or derivatives thereof. 63. The method of claim 62, wherein the filler DNA comprises double-stranded DNA. 64. The method of any one of claims 62 to 63, wherein the amount of filler DNA is from about 20 nanograms (ng) to about 100 ng. 65. The method of any one of claims 62 to 64, wherein at least a portion of the filler DNA is methylated. 66. The method of claim 65, wherein between 10% and 40% of the filler DNA is methylated, and the remainder is unmethylated filler DNA. 67. The method of any one of claims 21 to 66, the method further comprising, prior to (b), contacting the cfDNA sample with a methylated nucleic acid capture reagent to generate the plurality of nucleic acid molecules, wherein the plurality of nucleic acids comprise one or more methylated regions. 68. The method of claim 67, wherein the methylated nucleic acid capture reagent comprises a conjugate and a solid matrix. 69. The method of claim 68, wherein the methylated nucleic acid capture reagent is generated by coupling the conjugate to the solid matrix by incubating the conjugate together with the solid matrix. 70. The method of claim 69, wherein the coupling of the conjugate to the solid matrix is performed prior to the contact of the cfDNA sample with the methylated nucleic acid capture reagent. 71. The method of any one of claims 68 to 70, wherein the solid matrix is beads. 72. The method of any one of claims 68 to 70, wherein the solid matrix is protein A beads. 73. The method of any one of claims 68 to 72, wherein the solid matrix is a magnetic solid matrix. 74. The method of any one of claims 68 to 73, wherein the conjugate comprises an antibody. 75. The method of any one of claims 68 to 74, wherein the conjugate is selected from anti-5-methylcytosine antibody or a derivative thereof, anti-5-carboxycytosine antibody or a derivative thereof, anti-5-formylcytosine antibody or a derivative thereof, anti-5-hydroxymethylcytosine antibody or a derivative thereof, anti-3-methylcytosine antibody or a derivative thereof, and any combination thereof. 76. The method according to any one of claims 67 to 75, wherein one or more methylated regions of the plurality of nucleic acids are enriched with at least about 99% specificity. Claims 4 / 10 pages 5 CN 121464223 A77. The method of any one of claims 21 to 76, further comprising, prior to (b), amplifying the plurality of nucleic acid molecules to generate an amplicon, wherein (c) comprises sequencing the amplicon. 78. The method of claim 77, wherein the amplification of the plurality of nucleic acids is performed simultaneously with binding to a solid support. 79. The method of any one of claims 77 to 78, wherein the amplification comprises PCR amplification. 80. The method of claim 79, wherein the PCR amplification comprises at least 13 or at least 14 cycles. 81. The method of any one of claims 21 to 80, further comprising, prior to (b), contacting the plurality of nucleic acid molecules with one or more nucleic acid capture probes to enrich one or more target sequences. 82. The method of claim 81, wherein the one or more target sequences comprise one or more genes. 83. A method of processing a nucleic acid sample from an object, the method comprising: (a) generating a nucleic acid sample mixture comprising a plurality of methylated nucleic acids and a plurality of filler nucleic acids from the object, wherein the plurality of filler nucleic acids comprise at least one methylated nucleic acid molecule; (b) incubating (i) a methylation-binding molecule with (ii) a solid matrix to form a methylated nucleic acid capture reagent; and (c) capturing the methylated nucleic acids by adding the methylated nucleic acid capture reagent to the nucleic acid sample mixture, thereby enriching the plurality of methylated nucleic acids in the nucleic acid sample mixture. 84. The method of claim 83, further comprising, after the capture, amplifying the captured methylated nucleic acids to generate amplicons of the plurality of methylated nucleic acids. 85. The method of claim 84, wherein the amplification is performed simultaneously with the binding of the plurality of methylated nucleic acids to the methylated nucleic acid capture reagent. 86. The method of claim 84 or 85, wherein the captured methylated nucleic acids are not eluted prior to amplification. 87. The method of any one of claims 83 to 86, further comprising subjecting the plurality of methylated nucleic acids or derivatives thereof to a sequencing reaction. 88. The method of claim 87, wherein the sequencing reaction is performed via a synthesis reaction. 89. The method of claim 87 or 88, wherein the sequencing reaction does not include bisulfite sequencing. 90. The method of any one of claims 83 to 89, wherein the plurality of methylated nucleic acids comprise cell-free nucleic acids. 91. The method of any one of claims 83 to 90, wherein the solid matrix is beads. 92. The method of any one of claims 83 to 91, wherein the solid matrix is a magnetic solid matrix. 93. The method of any one of claims 83 to 92, wherein the solid matrix comprises protein A.94. The method of any one of claims 83 to 93, wherein the solid matrix comprises streptavidin. 95. The method of any one of claims 83 to 94, wherein the methylation-binding molecule is an antibody. 96. The method of any one of claims 83 to 95, wherein the methylation-binding molecule comprises biotin. 97. The method of any one of claims 83 to 96, wherein the methylation-binding molecule binds methylated cytosine. 98. The method of any one of claims 83 to 97, further comprising, prior to (a), obtaining the nucleic acid sample from the object and performing one or more library preparation reactions on the nucleic acid sample. 99. The method of claim 98, further comprising, after performing the one or more library preparation reactions and prior to (a), incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acid. 100. The method of claim 99, further comprising, after incubating the nucleic acid sample with a plurality of magnetic beads interacting with nucleic acids, magnetically trapping the sample to remove the plurality of magnetic beads interacting with nucleic acids. 101. The method of claim 100, further comprising performing additional magnetic trapping. 102. The method of any one of claims 83 to 101, further comprising, after (a) and before (c), denaturing the nucleic acids in the nucleic acid sample mixture. 103. The method of any one of claims 83 to 102, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched at least 2-fold. 104. The method of any one of claims 83 to 103, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched at least 100-fold. 105. The method of any one of claims 83 to 104, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched with at least 99% specificity. 106. The method of any one of claims 83 to 105, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched with at least 99.5% specificity. 107. The method of any one of claims 83 to 106, further comprising contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences. 108. The method of claim 107, wherein the one or more target sequences comprise one or more genes. 109. A method of processing a nucleic acid sample selected from objects, the method comprising: (a) providing a nucleic acid sample comprising a plurality of methylated nucleic acids;(b) Incubating (i) the methylation-binding molecule with (ii) a solid matrix to form a methylated nucleic acid capture reagent; (c) Enriching the plurality of methylated nucleic acids in the nucleic acid sample mixture by adding the methylated nucleic acid capture reagent to the nucleic acid sample mixture, wherein the plurality of methylated nucleic acids are enriched with a specificity greater than 99%. 110. The method of claim 109, further comprising, after the capture, amplifying the captured methylated nucleic acids to generate amplicons of the plurality of methylated nucleic acids. 111. The method of claim 110, wherein the amplification is performed simultaneously with the binding of the plurality of methylated nucleic acids to the methylated nucleic acid capture reagent. 112. The method of claim 109 or 110, wherein the captured methylated nucleic acids are not eluted prior to amplification. 113. The method of any one of claim 109 or 112, further comprising subjecting the plurality of methylated nucleic acids or derivatives thereof to a sequencing reaction. 114. The method of claim 113, wherein the sequencing reaction is performed by a synthetic reaction. 115. The method of claim 113 or 114, wherein the sequencing reaction does not include bisulfite sequencing. 116. The method of any one of claims 109 to 115, wherein the plurality of methylated nucleic acids comprise cell-free nucleic acids. 117. The method of any one of claims 109 to 116, wherein the solid matrix is beads. 118. The method of any one of claims 109 to 117, wherein the solid matrix is a magnetic solid matrix. 119. The method of any one of claims 109 to 118, wherein the solid matrix comprises protein A. 120. The method of any one of claims 109 to 119, wherein the solid matrix comprises streptavidin. 121. The method of any one of claims 109 to 120, wherein the methylation-binding molecule is an antibody. 122. The method of any one of claims 109 to 121, wherein the methylation-binding molecule comprises biotin. 123. The method of any one of claims 109 to 122, wherein the methylation-binding molecule binds methylated cytosine. 124. The method of any one of claims 109 to 123, further comprising, prior to (c), performing one or more library preparation reactions on the methylated nucleic acid. 125. The method of claim 124, further comprising, after performing the one or more library preparation reactions and prior to (c), incubating the nucleic acid sample together with a plurality of magnetic beads that interact with the nucleic acid.126. The method of claim 125, further comprising, after incubating the nucleic acid sample with a plurality of magnetic beads interacting with nucleic acids, magnetically trapping the sample to remove the plurality of magnetic beads interacting with nucleic acids. 127. The method of claim 126, further comprising performing additional magnetic trapping. 128. The method of any one of claims 109 to 127, further comprising, after (a) and before (c), denaturing the nucleic acids in the nucleic acid sample. 129. The method of any one of claims 109 to 128, wherein the plurality of methylated nucleic acids in the nucleic acid sample are enriched at least 2-fold. 130. The method of any one of claims 109 to 129, wherein the plurality of methylated nucleic acids in the nucleic acid sample are enriched at least 100-fold. 131. The method of any one of claims 109 to 130, wherein the plurality of methylated nucleic acids in the nucleic acid sample are enriched with at least 99.5% specificity. 132. The method of any one of claims 109 to 131, further comprising contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences. 133. The method of claim 132, wherein the one or more target sequences comprise one or more genes. 134. A method of processing a nucleic acid sample from an object, the method comprising: (a) providing a nucleic acid sample comprising a plurality of methylated nucleic acids; (b) incubating (i) a methylation-binding molecule with (ii) a solid matrix to form a methylated nucleic acid capture reagent; (c) capturing the methylated nucleic acids by adding the methylated nucleic acid capture reagent to the nucleic acid sample, thereby generating methylated nucleic acids bound to the solid matrix; (d) amplifying the methylated nucleic acids bound to the solid matrix to generate amplicons of the plurality of methylated nucleic acids. 135. The method of claim 134, wherein the amplification is performed simultaneously with the binding of the plurality of methylated nucleic acids to the methylated nucleic acid capture reagent. 136. The method of claim 134 or 135, wherein the captured methylated nucleic acids are not eluted prior to amplification. 137. The method of any one of claims 134 to 136, further comprising performing a sequencing reaction on the amplicon of the methylated nucleic acid. (Claims 7 / 10, Page 8, CN 121464223 A) 138. The method of claim 137, wherein the sequencing reaction is performed via a synthesis reaction. 139. The method of claim 137 or 138, wherein the sequencing reaction does not include bisulfite sequencing. 140. The method of any one of claims 134 to 139, wherein the plurality of methylated nucleic acids comprise cell-free nucleic acids.141. The method of any one of claims 134 to 140, wherein the solid matrix is a bead. 142. The method of any one of claims 134 to 141, wherein the solid matrix is a magnetic solid matrix. 143. The method of any one of claims 134 to 142, wherein the solid matrix comprises protein A. 144. The method of any one of claims 134 to 143, wherein the solid matrix comprises streptavidin. 145. The method of any one of claims 134 to 144, wherein the methylation-binding molecule is an antibody. 146. The method of any one of claims 134 to 145, wherein the methylation-binding molecule comprises biotin. 147. The method of any one of claims 134 to 146, wherein the methylation-binding molecule binds to methylated cytosine. 148. The method of any one of claims 134 to 147, wherein, prior to (c), the methylated nucleic acid is subjected to one or more library preparation reactions. 149. The method of claim 148, further comprising, after performing the one or more library preparation reactions and before (c), incubating the nucleic acid sample with a plurality of magnetic beads that interact with nucleic acids. 150. The method of claim 149, further comprising, after incubating the nucleic acid sample with the plurality of magnetic beads that interact with nucleic acids, subjecting the sample to magnetic trapping to remove the plurality of magnetic beads that interact with nucleic acids. 151. The method of claim 150, further comprising performing additional magnetic trapping. 152. The method of any one of claims 134 to 151, further comprising, after (a) and before (c), denaturing the nucleic acids in the nucleic acid sample. 153. The method of any one of claims 134 to 152, wherein the plurality of methylated nucleic acids in the nucleic acid sample are enriched at least 2-fold. 154. The method of any one of claims 134 to 153, wherein the plurality of methylated nucleic acids in the nucleic acid sample are enriched at least 100-fold. 155. The method of any one of claims 134 to 154, wherein the plurality of methylated nucleic acids in the nucleic acid sample are enriched with at least 99% specificity. 156. The method of any one of claims 134 to 155, wherein the plurality of methylated nucleic acids in the nucleic acid sample are enriched with at least 99.5% specificity. 157. The method of any one of claims 134 to 156, further comprising contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences. 158. The method of claim 157, wherein the one or more target sequences comprise one or more genes.159. A method for processing a nucleic acid sample from an object, the method comprising: (a) generating a nucleic acid sample mixture comprising a plurality of methylated nucleic acids and a plurality of filler nucleic acids from the object, wherein the plurality of filler nucleic acids comprise at least one methylated nucleic acid molecule; (b) capturing the methylated nucleic acids by adding a capture reagent comprising a solid matrix to the nucleic acid sample mixture, thereby generating methylated nucleic acids bound to the solid matrix; and (c) amplifying the methylated nucleic acids bound to the solid matrix to generate an amplicon of the methylated nucleic acid. 160. The method of claim 159, wherein the captured methylated nucleic acids are not eluted prior to amplification. 161. The method of claim 159 or 160, further comprising performing a sequencing reaction on the amplicon of the methylated nucleic acid. 162. The method of claim 161, wherein the sequencing reaction is performed by a synthesis reaction. 163. The method of claim 161 or 162, wherein the sequencing reaction does not include bisulfite sequencing. 164. The method of any one of claims 159 to 163, wherein the plurality of methylated nucleic acids comprise cell-free nucleic acids. 165. The method of any one of claims 159 to 164, wherein the solid matrix is a bead. 166. The method of any one of claims 159 to 165, wherein the solid matrix is a magnetic solid matrix. 167. The method of any one of claims 159 to 166, wherein the solid matrix comprises protein A. 168. The method of any one of claims 159 to 167, wherein the solid matrix comprises streptavidin. 169. The method of any one of claims 159 to 168, wherein the capture agent is a methylated nucleic acid capture agent. 170. The method of claim 169, wherein the methylated nucleic acid capture agent comprises a methylation-binding molecule attached to the solid matrix. 171. The method of claim 170, wherein the methylation-binding molecule is an antibody. 172. The method of claim 170, wherein the methylation-binding molecule comprises biotin. 173. The method of claim 172, wherein the methylation-binding molecule binds methylated cytosine. 174. The method of any one of claims 159 to 173, further comprising, prior to (a), obtaining the nucleic acid sample from the object and performing one or more library preparation reactions on the nucleic acid. 175. The method of claim 174, further comprising, after performing the one or more library preparation reactions and prior to (a), incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acid.176. The method of claim 175, further comprising, after incubating the nucleic acid sample with a plurality of magnetic beads interacting with nucleic acids, magnetically trapping the sample to remove the plurality of magnetic beads interacting with nucleic acids. 177. The method of claim 176, further comprising performing additional magnetic trapping. 178. The method of any one of claims 159 to 177, further comprising, after (a) and before (b), denaturing the nucleic acids in the nucleic acid sample mixture. 179. The method of any one of claims 159 to 178, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched at least 2-fold. 180. The method of any one of claims 159 to 179, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched at least 100-fold. 181. The method of any one of claims 159 to 180, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched with at least 99% specificity. 182. The method of any one of claims 159 to 181, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched with at least 99.5% specificity. 183. The method of any one of claims 159 to 182, further comprising contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences. 184. The method of claim 183, wherein the one or more target sequences comprise one or more genes. Claims 10 / 10 pages 11 CN 121464223 A Method and system for cell-free nucleic acid processing
[0001] Cross-reference This application claims U.S. Provisional Application No. 63 / 496,347 filed April 14, 2023; U.S. Provisional Application No. 63 / 501,359 filed May 10, 2023; U.S. Provisional Application No. 63 / 511,441 filed June 30, 2023; U.S. Provisional Application No. 63 / 517,327 filed August 2, 2023; U.S. Provisional Application No. 63 / 517,327 filed October 5, 2023. U.S. Provisional Application No. 588,120, filed October 19, 2023; U.S. Provisional Application No. 63 / 591,732, filed October 30, 2023; U.S. Provisional Application No. 63 / 594,365, filed November 22, 2023; U.S. Provisional Application No. 63 / 602,156, filed February 2, 2024; and U.S. Provisional Application No. 63 / 549,294, filed March 28, 2024; and U.S. Provisional Application No. 63 / 571,139, filed March 28, 2024.The rights of these documents are hereby infringed, and each of these documents is incorporated herein by reference in its entirety. Background Art
[0002] Circulating tumor DNA (ctDNA) is increasingly showing potential as a non-invasive tumor-specific biomarker for routine clinical applications. ctDNA originates from tumor cells that have undergone primary cell death and is released into circulation in various bodily fluids, including blood. In most cancer patients, most blood-derived cell-free DNA originates from healthy (e.g., non-cancerous) tissue. Furthermore, depending on several factors including the primary site of tumor and disease burden, the fraction of ctDNA observed at diagnosis may range from <0.1% to 90% of total cell-free DNA. ctDNA provides a non-invasive pathway to the tumor molecular landscape and disease burden.
[0003] All publications, patents, and patent applications mentioned in this specification are incorporated by reference to the extent that each individual publication, patent, or patent application is specifically and individually indicated to be incorporated by reference.
[0004] In some aspects, this invention provides a method comprising: (a) providing a plurality of nucleic acid molecules generated from a cell-free deoxynucleic acid (cfDNA) sample of a subject; (b) sequencing the plurality of nucleic acid molecules or derivatives thereof to generate a plurality of sequencing reads; (c) computer processing the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules; and (d) computer processing the methylation profiles to determine that the subject has cancer when the area under the receiver operating characteristic curve (AUROC) is at least about 91%, wherein the cancer is a low-shedding cancer. In some embodiments, the cfDNA sample comprises circulating tumor nucleic acid molecules derived from a low-shedding tumor. In some embodiments, the cancer is bladder cancer, breast cancer, endometrial cancer, prostate cancer, or kidney cancer. In some embodiments, the cancer is endometrial cancer or prostate cancer.
[0005] In some aspects, this document provides a method comprising: (a) providing a plurality of nucleic acid molecules generated from a cell-free deoxynucleic acid (cfDNA) sample of an object; (b) sequencing the plurality of nucleic acid molecules or derivatives thereof to generate a plurality of sequencing reads; (c) computer processing the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules; and (d) computer processing the methylation profiles to determine that the object has cancer when the area under the receiver operating characteristic curve (AUROC) is at least about 94%, and wherein said cancer is early-stage cancer. In some embodiments, the cfDNA sample comprises circulating tumor nucleic acid molecules derived from an early-stage tumor. In some embodiments, the early-stage tumor isStage I tumor. In some embodiments, the early-stage tumor is a Stage II tumor. In some embodiments, the cancer is bladder cancer, breast cancer, colorectal cancer, endometrial cancer, esophageal cancer, head and neck cancer, hepatobiliary cancer, lung cancer, ovarian cancer, prostate cancer, or kidney cancer. In some embodiments, the cancer is esophageal cancer, hepatobiliary cancer, or ovarian cancer.
[0006] In some aspects, this document provides a method comprising: (a) providing a plurality of nucleic acid molecules generated from a cell-free deoxyribonucleic acid (cfDNA) sample of a subject; (b) sequencing the plurality of nucleic acid molecules or derivatives thereof without bisulfite conversion to generate a plurality of sequencing reads; (c) computer processing the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules; and (d) computer processing the methylation profiles to determine that the subject has cancer, wherein the cancer is endometrial cancer, esophageal cancer, hepatobiliary cancer, ovarian cancer, prostate cancer, bladder cancer, breast cancer, colorectal cancer, head and neck cancer, lung cancer, pancreatic cancer, or kidney cancer. In some embodiments, the cfDNA sample comprises circulating tumor nucleic acid molecules derived from endometrial cancer. In some embodiments, the methylation profiles are computer-processed to determine that the subject has endometrial cancer when the area under the receiver operating characteristic (AUROC) is at least about 90%. In some embodiments, the cfDNA sample comprises circulating tumor nucleic acid molecules derived from esophageal cancer. In some embodiments, the methylation spectrum is processed by a computer to determine if the subject has esophageal cancer when the area under the receiver operating characteristic (AUROC) is at least about 99%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from hepatobiliary cancer. In some embodiments, the methylation spectrum is processed by a computer to determine if the subject has hepatobiliary cancer when the area under the receiver operating characteristic (AUROC) is at least about 99%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from ovarian cancer. In some embodiments, the methylation spectrum is processed by a computer to determine if the subject has ovarian cancer when the area under the receiver operating characteristic (AUROC) is at least about 97%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from prostate cancer. In some embodiments, the methylation spectrum is processed by a computer to determine if the subject has prostate cancer when the area under the receiver operating characteristic (AUROC) is at least about 89%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from bladder cancer. In some embodiments, the methylation spectrum is processed by a computer to determine the area under the receiver operating characteristic (AUROC) of the subject.In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from breast cancer. In some embodiments, the methylation spectrum is processed by a computer to determine that the subject has breast cancer when the area under the receiver operating characteristic (AUROC) is at least about 92%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from colorectal cancer. In some embodiments, the methylation spectrum is processed by a computer to determine that the subject has colorectal cancer when the area under the receiver operating characteristic (AUROC) is at least about 98%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from head and neck cancer. In some embodiments, the methylation spectrum is processed by a computer to determine that the subject has head and neck cancer when the area under the receiver operating characteristic (AUROC) is at least about 96%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from lung cancer. In some embodiments, the methylation spectrum is processed by a computer to determine that the subject has lung cancer when the area under the receiver operating characteristic (AUROC) is at least about 96%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from pancreatic cancer. In some embodiments, the methylation profile is processed by a computer to determine if the subject has pancreatic cancer when the area under the receiver operating characteristic curve (AUROC) is at least about 99%. In some embodiments, the cfDNA sample contains circulating tumor nucleic acid molecules derived from renal cancer. In some embodiments, the methylation profile is processed by a computer to determine if the subject has renal cancer when the area under the receiver operating characteristic curve (AUROC) is at least about 91%.
[0007] In some embodiments, the method further includes, prior to (b), adding a set of nucleic acid molecules not derived from the subject to the plurality of nucleic acid molecules. In some embodiments, the methylation profile is genome-wide. In some embodiments, the methylation profile contains a complete methylome. In some embodiments, a method (e.g., step description 2 / 55 page 13 CN 121464223 A (d)) includes a supervised machine learning method, wherein the supervised machine learning method is regression, support vector machine, tree-based method, neural network, or nearest neighbor method. In some embodiments, a method (e.g., step (d)) includes an unsupervised machine learning method, wherein the unsupervised machine learning method is clustering, neural networks, principal component analysis, or matrix factorization. In some embodiments, the subject has previously received cancer treatment and is substantially free of said cancer, wherein (d) includes determining that the subject has a recurrence of said cancer. In some embodiments, a method further includes, prior to (b), feeding the plurality of kernels...A certain amount of filler DNA is added to an acid molecule or a derivative thereof. In some embodiments, the amount of filler DNA comprises double-stranded DNA. In some embodiments, the amount of filler DNA is about 20 nanograms (ng) to about 100 ng. In some embodiments, at least a portion of the filler DNA is methylated. In some embodiments, between 10% and 40% of the filler DNA is methylated, with the remainder being unmethylated filler DNA. In some embodiments, the method further includes, prior to (b), contacting the cfDNA sample with a methylated nucleic acid capture reagent to generate the plurality of nucleic acid molecules, wherein the plurality of nucleic acids contain one or more methylated regions. In some embodiments, the methylated nucleic acid capture reagent comprises a conjugate and a solid matrix. In some embodiments, the methylated nucleic acid capture reagent is generated by conjugating the conjugate to the solid matrix by incubating the conjugate together with the solid matrix. In some embodiments, the conjugate is conjugated to the solid matrix prior to the contact of the cfDNA sample with the methylated nucleic acid capture reagent. In some embodiments, the solid matrix is beads. In some embodiments, the solid matrix is protein A beads. In some embodiments, the solid matrix is a magnetic solid matrix. In some embodiments, the conjugate comprises an antibody. In some embodiments, the conjugate is selected from anti-5-methylcytosine antibody or a derivative thereof, anti-5-carboxycytosine antibody or a derivative thereof, anti-5-formylcytosine antibody or a derivative thereof, anti-5-hydroxymethylcytosine antibody or a derivative thereof, anti-3-methylcytosine antibody or a derivative thereof, and any combination thereof. In some embodiments, one or more methylated regions of the plurality of nucleic acids are enriched with at least about 99% specificity. In some embodiments, the method further includes, prior to (b), amplifying the plurality of nucleic acid molecules to generate amplicons, wherein (c) includes sequencing the amplicons. In some embodiments, the amplification of the plurality of nucleic acids is performed while the plurality of nucleic acids are bound to a solid support. In some embodiments, the amplification comprises PCR amplification. In some embodiments, the PCR amplification comprises at least 13 or at least 14 cycles. In some embodiments, the method further includes, prior to (b), contacting the plurality of nucleic acid molecules with one or more nucleic acid capture probes to enrich one or more target sequences. In some embodiments, the one or more target sequences comprise one or more genes.
[0008] In some aspects, this document provides a method for processing a nucleic acid sample from an object, the method comprising: (a) generating a mixture of nucleic acid samples comprising a plurality of methylated nucleic acids and a plurality of filler nucleic acids from said object, whereinThe plurality of filler nucleic acids comprise at least one methylated nucleic acid molecule; (b) incubating (i) the methylation-binding molecule with (ii) a solid matrix to form a methylated nucleic acid capture reagent; (c) enriching the plurality of methylated nucleic acids in the nucleic acid sample mixture by capturing the methylated nucleic acids by adding the methylated nucleic acid capture reagent to the nucleic acid sample mixture. In some embodiments, the method further includes, after the capture, amplifying the captured methylated nucleic acids to generate amplicons of the plurality of methylated nucleic acids. In some embodiments, the amplification is performed simultaneously with the binding of the plurality of methylated nucleic acids to the methylated nucleic acid capture reagent. In some embodiments, the captured methylated nucleic acids are not eluted prior to amplification. In some embodiments, the method further includes performing a sequencing reaction on the plurality of methylated nucleic acids or derivatives thereof. In some embodiments, the sequencing reaction is performed by a synthetic reaction. In some embodiments, the sequencing reaction does not include bisulfite sequencing. In some embodiments, the plurality of methylated nucleic acids comprise cell-free nucleic acids. In some embodiments, the solid matrix is beads. In some embodiments, the solid matrix is a magnetic solid matrix. In some embodiments, the solid matrix comprises protein specification page 3 / 55 14 CN 121464223 AA. In some embodiments, the solid matrix is streptavidin. In some embodiments, the methylation-binding molecule is an antibody. In some embodiments, the methylation-binding molecule comprises biotin. In some embodiments, the methylation-binding molecule binds methylated cytosine. In some embodiments, the method further includes, before (a), obtaining the nucleic acid sample from the object and performing one or more library preparation reactions on the nucleic acid sample. In some embodiments, the method further includes, after performing the one or more library preparation reactions and before (a), incubating the nucleic acid sample with a plurality of magnetic beads that interact with nucleic acids. In some embodiments, the method further includes, after incubating the nucleic acid sample with a plurality of magnetic beads that interact with nucleic acids, magnetically trapping the sample to remove the plurality of magnetic beads that interact with nucleic acids. In some embodiments, the method further includes performing additional magnetic trapping. In some embodiments, the method further includes, after (a) and before (c), denaturing the nucleic acids in the nucleic acid sample mixture. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched at least 2-fold. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched at least 100-fold. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched to at least 100-fold.At least 99% specific enrichment. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched with at least 99.5% specificity. In some embodiments, the method further includes contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences. In some embodiments, the one or more target sequences comprise one or more genes.
[0009] In some aspects, this document provides a method for processing a nucleic acid sample from an object, the method comprising: (a) providing a nucleic acid sample containing a plurality of methylated nucleic acids; (b) incubating (i) a methylation-binding molecule with (ii) a solid matrix to form a methylated nucleic acid capture reagent; (c) capturing the plurality of methylated nucleic acids by adding the methylated nucleic acid capture reagent to the nucleic acid sample mixture, thereby enriching the plurality of methylated nucleic acids in the nucleic acid sample mixture, wherein the plurality of methylated nucleic acids are enriched with greater than 99% specificity. In some embodiments, the method further includes, after the capture, amplifying the captured methylated nucleic acids to generate amplicons of the plurality of methylated nucleic acids. In some embodiments, the amplification is performed simultaneously with the binding of the plurality of methylated nucleic acids to the methylated nucleic acid capture reagent. In some embodiments, the captured methylated nucleic acids are not eluted prior to amplification. In some embodiments, the method further includes sequencing the plurality of methylated nucleic acids or their derivatives. In some embodiments, the sequencing reaction is performed via a synthesis reaction. In some embodiments, the sequencing reaction does not include bisulfite sequencing. In some embodiments, the plurality of methylated nucleic acids comprise cell-free nucleic acids. In some embodiments, the solid matrix is beads. In some embodiments, the solid matrix is a magnetic solid matrix. In some embodiments, the solid matrix contains protein A. In some embodiments, the solid matrix contains streptavidin. In some embodiments, the methylation-binding molecule is an antibody. In some embodiments, the methylation-binding molecule comprises biotin. In some embodiments, the methylation-binding molecule binds methylated cytosine. In some embodiments, the method further includes performing one or more library preparation reactions on the methylated nucleic acids prior to (c). In some embodiments, the method further includes incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acids after performing the one or more library preparation reactions and prior to (c). In some embodiments, the method further includes magnetically trapping the nucleic acid sample after incubating it with a plurality of magnetic beads that interact with the nucleic acid to remove the plurality of magnetic beads interacting with the nucleic acid. In some embodiments, the method...The method also includes performing additional magnetic capture. In some embodiments, the method further includes denaturing the nucleic acids in the nucleic acid sample after (a) and before (c). In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample are enriched at least 2-fold. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample are enriched at least 100-fold. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample are enriched with at least 99.5% specificity. In some embodiments, the method further includes contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences. In some embodiments, the one or more target sequences comprise one or more genes.
[0010] In some aspects, this document provides a method for processing a nucleic acid sample from an object, the method comprising: (a) providing a nucleic acid sample containing a plurality of methylated nucleic acids; (b) incubating (i) a methylation-binding molecule with (ii) a solid matrix to form a methylated nucleic acid capture reagent; (c) capturing the methylated nucleic acids by adding the methylated nucleic acid capture reagent to the nucleic acid sample, thereby generating methylated nucleic acids bound to the solid matrix; and (d) amplifying the methylated nucleic acids bound to the solid matrix to generate amplicons of the plurality of methylated nucleic acids. In some embodiments, the amplification is performed simultaneously with the binding of the plurality of methylated nucleic acids to the methylated nucleic acid capture reagent. In some embodiments, the captured methylated nucleic acids are not eluted prior to amplification. In some embodiments, the method further includes a sequencing reaction of the amplicons of the methylated nucleic acids. In some embodiments, the sequencing reaction is performed via a synthesis reaction. In some embodiments, the sequencing reaction does not include bisulfite sequencing. In some embodiments, the plurality of methylated nucleic acids comprise cell-free nucleic acids. In some embodiments, the solid matrix is beads. In some embodiments, the solid matrix is a magnetic solid matrix. In some embodiments, the solid matrix comprises protein A. In some embodiments, the solid matrix comprises streptavidin. In some embodiments, the methylation-binding molecule is an antibody. In some embodiments, the methylation-binding molecule comprises biotin. In some embodiments, the methylation-binding molecule binds methylated cytosine. In some embodiments, the method further includes performing one or more library preparation reactions on the methylated nucleic acid prior to (c). In some embodiments, the method further includes incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acid after performing the one or more library preparation reactions and prior to (c). In some embodiments, the method further includes incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acid prior to (c).After incubating the sample with a plurality of magnetic beads that interact with nucleic acids, the sample is subjected to magnetic trapping to remove the magnetic beads that interact with the nucleic acids. In some embodiments, the method further includes performing additional magnetic trapping. In some embodiments, the method further includes denaturing the nucleic acids in the nucleic acid sample after (a) and before (c). In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample are enriched at least 2-fold. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample are enriched at least 100-fold. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample are enriched with at least 99% specificity. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample are enriched with at least 99.5% specificity. In some embodiments, the method further includes contacting the plurality of methylated nucleic acids with one or more nucleic acid trapping probes to enrich one or more target sequences. In some embodiments, the one or more target sequences comprise one or more genes.
[0011] In some aspects, this document provides a method for processing a nucleic acid sample from an object, the method comprising: (a) generating a nucleic acid sample mixture comprising a plurality of methylated nucleic acids and a plurality of filler nucleic acids from the object, wherein the plurality of filler nucleic acids comprise at least one methylated nucleic acid molecule; (b) capturing the methylated nucleic acids by adding a capture reagent comprising a solid matrix to the nucleic acid sample mixture, thereby generating methylated nucleic acids bound to the solid matrix; and (c) amplifying the methylated nucleic acids bound to the solid matrix to generate amplicons of the methylated nucleic acids. In some embodiments, the captured methylated nucleic acids are not eluted prior to amplification. In some embodiments, the method further comprises a sequencing reaction of the amplicons of the methylated nucleic acids. In some embodiments, the sequencing reaction is performed by a synthesis reaction. In some embodiments, the sequencing reaction does not include bisulfite sequencing. In some embodiments, the plurality of methylated nucleic acids comprise cell-free nucleic acids. In some embodiments, the solid matrix is a bead. In some embodiments, the solid matrix is a magnetic solid matrix. In some embodiments, the solid matrix comprises protein A. In some embodiments, the solid matrix comprises streptavidin. In some embodiments, the capture agent is a methylated nucleic acid capture agent. In some embodiments, the methylated nucleic acid capture agent comprises a methylated binding molecule attached to the solid matrix. In some embodiments, the methylated binding molecule is an antibody. In some embodiments, the methylated binding molecule comprises biotin. In some embodiments, the methylated binding...The molecule binds to methylated cytosine. In some embodiments, the method further includes, prior to (a), obtaining the nucleic acid sample from the object and performing one or more library preparation reactions on the nucleic acid. In some embodiments, the method further includes, after performing the one or more library preparation reactions and prior to (a), incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acid. In some embodiments, the method further includes, after incubating the nucleic acid sample with the plurality of magnetic beads that interact with the nucleic acid, magnetically trapping the sample to remove the plurality of magnetic beads that interact with the nucleic acid. In some embodiments, the method further includes performing additional magnetic trapping. In some embodiments, the method further includes, after (a) and prior to (b), denaturing the nucleic acids in the nucleic acid sample mixture. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched at least 2-fold. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched at least 100-fold. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched with at least 99% specificity. In some embodiments, the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched with at least 99.5% specificity. In some embodiments, the method further includes contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences. In some embodiments, the one or more target sequences comprise one or more genes.
[0012] In some aspects, this document provides a method comprising: (a) obtaining a first set of nucleic acid molecules from a cell-free sample of an object; (b) generating a second set of nucleic acid molecules from the first set of nucleic acid molecules or derivatives thereof, wherein the second set of nucleic acid molecules is enriched at a methylation level relative to the first set of nucleic acid molecules; (c) enriching one or more targets from the second set of nucleic acid molecules or derivatives thereof to obtain a third set of nucleic acid molecules; and (d) sequencing the third set of nucleic acid molecules or derivatives thereof. In some embodiments, the enrichment comprises contacting the second set of nucleic acid molecules or derivatives thereof with one or more nucleic acid capture probes. In some embodiments, the generation comprises contacting the first set of nucleic acid molecules or said derivatives thereof with a methylated nucleic acid capture reagent. In some embodiments, the methylated nucleic acid capture reagent is formed by incubating a methylation-binding molecule together with (ii) a solid matrix. In some embodiments, the solid matrix is a bead. In some embodiments, the solid matrix is a magnetic solid matrix. In some embodiments, the solid matrix contains protein A. In some embodiments, the solid matrix contains streptavidin. In some embodiments, the methylation-binding molecule is an antibody. In some embodiments, the methylation-binding molecule...The molecules include biotin. In some embodiments, the methylation-binding molecule binds to methylated cytosine. In some embodiments, the method further includes amplifying the second group of molecules. In some embodiments, the amplification is performed simultaneously with the binding of a subgroup of the first group of nucleic acids to the methylated nucleic acid capture reagent. In some embodiments, the sequencing is performed via a synthesis reaction. In some embodiments, the sequencing does not include bisulfite sequencing. In some embodiments, the third group of nucleic acid molecules undergoes one or more library preparation reactions prior to the sequencing. In some embodiments, after the one or more library preparation reactions and prior to the sequencing, the third group of nucleic acid molecules is incubated with a plurality of magnetic beads that interact with nucleic acids. In some embodiments, the method further includes magnetically capturing the sample after incubating the nucleic acid sample with the plurality of magnetic beads that interact with nucleic acids to remove the plurality of magnetic beads that interact with nucleic acids. In some embodiments, the method further includes performing additional magnetic capture. In some embodiments, the method further includes adding a certain amount of filler DNA to the first nucleic acid molecules prior to (b).
[0013] These and other features of the preferred embodiments of the invention will become more apparent in the following detailed description with reference to the accompanying drawings, in which: Figure 1 shows a graph illustrating the process of collecting unmethylated / hypomethylated DNA fragments.
[0014] Figure 2 shows a schematic diagram of a computer system according to an embodiment of the present disclosure.
[0015] Figure 3 illustrates receiver operating characteristic (ROC) curves for an overall cohort of 12 cancers.
[0016] Figures 4A-4L illustrate early detection of various cancers according to cancer type. Figure 4A shows the ROC curve (95% confidence interval) for bladder cancer, and the area under the ROC curve (AUC) for all stages, stages I, II, III, and IV. Figure 4B shows the ROC curve (95% confidence interval) for breast cancer, and the AUC for all stages, stages I, II, III, and IV. Figure 4C shows the ROC curve (95% confidence interval) for colorectal cancer, and the AUC for all stages, stages I, II, III, and IV. Figure 4D shows the ROC curve (95% confidence interval) for endometrial cancer, and the AUC for all stages, stages I, II, III, and IV. Figure 4E shows the ROC curve (95% confidence interval) for esophageal cancer, and the AUC for all stages, stages I, II, III, and IV. Figure 4F shows the ROC curve (95% confidence interval) for head and neck cancer, and the AUC for all stages, stages I, II, III, IV, and IV.Figure 4G shows the ROC curves (95% confidence intervals) for hepatobiliary cancer, and the AUCs for all stages, stages I, II, III, and IV. Figure 4H shows the ROC curves (95% confidence intervals) for lung cancer, and the AUCs for all stages, stages I, II, III, and IV. Figure 4I shows the ROC curves (95% confidence intervals) for ovarian cancer, and the AUCs for all stages, stages I, II, III, and IV. Figure 4J shows the ROC curves (95% confidence intervals) for pancreatic cancer, and the AUCs for all stages, stages I, II, III, and IV. Figure 4K shows the ROC curves (95% confidence intervals) for prostate cancer, and the AUCs for all stages, stages I, II, III, and IV. Figure 4L shows the ROC curves (95% confidence intervals) for renal cancer, and the AUCs for all stages, stages I, II, III, and IV.
[0017] Figures 5A-5B illustrate ctDNA quantification and prognostic prediction in renal cell carcinoma. Figure 5A shows ctDNA quantification scores for renal cell carcinoma generated based on the average normalized count of reduced-size fragments to identify cancer-related methylation in 2027 regions, with these scores adjusted for methylation specificity. A threshold for ctDNA quantification scores was set such that 95% of recurrence-free or progression-free samples fell below the threshold. Figure 5B shows a Kaplan-Meier plot of recurrence-free / progression-free survival probability over time in renal cell carcinoma.
[0018] Figures 6A-6B illustrate ctDNA quantification and prognostic prediction in head and neck cancer. Figure 6A shows ctDNA quantification scores for head and neck cancer generated based on the average normalized count of reduced-size fragments to identify cancer-related methylation in 2027 regions, with these scores adjusted for methylation specificity. A threshold for ctDNA quantification scores was set such that 95% of recurrence-free or progression-free samples fell below the threshold. Figure 6B shows a Kaplan-Meier plot of recurrence-free / progression-free survival probability over time in head and neck cancer.
[0019] Figure 7 shows a Kaplan-Meier plot depicting the event-free survival over time for individuals with head and neck cancer (stratified by ctDNA quantification).
[0020] Figure 8 shows the age of individuals with renal cell carcinoma (RCC) at the time of sample collection.
[0021] Figure 9 shows a Kaplan-Meier plot depicting the event-free survival over time for individuals with all stages of RCC.
[0022] Figure 10 shows a Kaplan-Meier plot depicting the event-free survival over time for individuals with stages I-III RCC.
[0023] Figure 11 shows a Kaplan-Meier plot depicting the recurrence-free survival over time for individuals with early-stage non-small cell lung cancer (NSCLC).
[0024] Figure 12 shows a schematic diagram of patient treatment and blood collection, sample processing, and a genome-wide methylation enrichment platform.
[0025] Figures 13A-13B illustrate Kaplan-Meier plots depicting the probability of recurrence-free survival over time for individuals predicted to have positive or negative head and neck cancer based on ctDNA. Figure 13A shows a Kaplan-Meier plot depicting the probability of recurrence-free survival over time for individuals predicted to have positive or negative head and neck cancer based on ctDNA at a landmark time point. Figure 13B shows a Kaplan-Meier plot depicting the probability of recurrence-free survival over time for individuals predicted to have positive or negative head and neck cancer based on ctDNA.
[0026] Figure 14 shows estimated ctDNA quantification from pre-treatment to post-treatment in patients with recurrence-free or recurrent head and neck cancer.
[0027] Figure 15 shows a representative case study of ctDNA dynamics in individual patients before and after therapeutically intended treatment.
[0028] Figure 16 shows an example of a schematic overview of the whole-genome methylation enrichment platform.
[0029] Figure 17 shows the experimental design used to evaluate the detection limit of the whole-genome methylation enrichment platform.
[0030] Figure 18 shows the methylation-specific and unique molecular distributions of all non-cancer samples and designed cancer samples treated by the whole-genome methylation enrichment platform disclosed herein.
[0031] Figure 19 shows the ctDNA methylation score of different cfDNA sources titrated to pooled non-cancer donor-derived cfDNA in a titration series targeting less than 1% ctDNA levels.
[0032] Figure 20 shows the experimental design used to evaluate the accuracy of the whole-genome methylation enrichment platform disclosed herein.
[0033] Figure 21 shows the consistency and variability of different levels of ctDNA obtained using the ctDNA methylation score obtained from the whole-genome methylation enrichment platform disclosed herein.
[0034] Figure 22 shows the percentage of variance components (CV%) calculated for different operators, sequencing runs, and antibody reagent batches for different levels of ctDNA subjected to the whole-genome methylation enrichment platform disclosed herein.
[0035] Figure 23 shows genomic contamination of cell-free DNA.
[0036] Figure 24 shows the 0 CpG counts for samples with low and high binding specificity.
[0037] Figure 25 shows four different workflows for processing plasma-derived cell-free DNA using the whole-genome methylation enrichment platform.
[0038] Figure 26 shows the methylation specificity of the different workflows outlined in Figure 25.
[0039] Figure 27 shows the improved detection limit methylation group in the cancer methylation group method compared to the full methylation group method, expressed as a fold change.
[0040] Figure 28 shows a workflow of processing plasma-derived cell-free DNA using a full methylation group enrichment platform and then capturing target regions with probes. Detailed Description
[0041] This disclosure provides methods and systems for processing and analyzing nucleic acids present in biological samples, which can be used to determine the risk or likelihood of a subject having cancer or tumor with high sensitivity and / or high specificity. The methods and systems provided herein may include generating highly methylated cell-free nucleic acid molecules that can be processed to distinguish between cancerous and non-cancerous tissues, for example, circulating cell-free DNA (cfDNA). Specification 8 / 55 pages 19 CN 121464223 A
[0042] For example, the use and analysis of highly methylated nucleic acids can detect and / or characterize circulating tumor DNA (ctDNA) in fluid samples (such as blood samples) obtained from a subject with high sensitivity and high specificity. In some cases, the use and analysis of hypermethylated nucleic acids can improve sensitivity, specificity, and / or efficiency in determining whether a subject is at risk of having or developing a tumor or cancer. Methods for detecting ctDNA with improved sensitivity are needed, particularly in subjects with low ctDNA abundance.
[0043] The term “subject” as used herein is generally defined as any member of the animal kingdom (e.g., human, non-human primate, mouse, rat, cow, sheep, equine, canine, feline, rabbit, horse, or goat). Therefore, the methods described herein can be applied to human and veterinary diseases as well as animal models. A preferred subject is a “patient,” i.e., a living person being investigated to determine whether a disease or condition requires treatment or medical care; or receiving medical care for a disease or condition (e.g., cancer).
[0044] The term “genome” as used herein generally refers to genomic information derived from a subject, which may be at least a portion or all of the subject’s genetic information. A genome may be encoded in DNA or RNA. A genome may contain coding regions (e.g., regions that encode proteins) as well as non-coding regions. A genome may include sequences containing all chromosomes in an organism. For example, the human genome typically has a total of 46 chromosomes. The sequences of all these chromosomes together constitute the human genome.
[0045] As used herein, the term "nucleic acid" generally refers to a polynucleotide comprising two or more nucleotides, that is, a polymer of nucleotides of any length (deoxyribonucleotides (dNTPs) or ribonucleotides (rNTPs) or similar molecules). Non-limiting examples of nucleic acids include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), coding or non-coding regions of genes or gene segments, and loci defined by linkage analysis.locus), exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, short interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), ribozymes, cDNA, recombinant nucleic acids, branched-chain nucleic acids, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. Nucleic acids may contain one or more modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, the nucleotide structure may be modified before or after nucleic acid assembly. The nucleotide sequence of a nucleic acid may be interrupted by non-nucleotide components. Nucleic acids may be further modified after polymerization, such as by conjugation or binding to a reporter. A “variant” nucleic acid is a polynucleotide whose nucleotide sequence is identical to that of its original nucleic acid, except for having at least one modified nucleotide, such as being deleted, inserted, or substituted, respectively. Variant nucleic acids may have at least about 80%, at least about 90%, at least about 95%, or at least about 99% identity with the nucleotide sequence of the original nucleic acid.
[0046] Cell-free methylated DNA is DNA that can be one or more nucleic acid molecules that circulate freely in the bloodstream. In some cases, cell-free methylated DNA can be methylated at various regions of DNA. For example, plasma samples can be taken to analyze cell-free methylated DNA. Studies have shown that a large amount of circulating nucleic acids in the blood come from necrotic or apoptotic cells, and a significant increase in nucleic acid levels caused by apoptosis has been observed in diseases such as cancer. In particular, for cancer, circulating DNA has been a hallmark of the disease, including mutations in oncogenes, microsatellite alterations, and for some cancers, viral genome sequences, DNA, or RNA in plasma have been increasingly studied as potential biomarkers of the disease. For example, the quantification of low levels of circulating tumor DNA in total circulating DNA can serve as a better marker for detecting colorectal cancer recurrence compared to carcinoembryonic antigen, a standard biomarker used clinically. Cell-free DNA (e.g., circulating cfDNA) may contain circulating tumor DNA (ctDNA).
[0047] As used herein, “library preparation” typically includes one or more of a series of end repairs, A-tailing, adaptor ligation, or any other preparation of cell-free DNA to allow for subsequent DNA sequencing.
[0048] As used herein, “replenished DNA” (e.g., “filler DNA”) may be non-coding DNA, or may consist of amplicons as described on page 9 / 55 of the specification, 20 CN 121464223 A.
[0049] In some embodiments, the fragment length metric is fragment length. In some preferred embodiments, the target cell-free methylated DNA is limited to lengths of <170 bp, <165 bp, <160 bp, <155 bp, <150 bp, <145 bp.Fragments of length <100 bp, <140 bp, <135 bp, <130 bp, <125 bp, <120 bp, <115 bp, <110 bp, <105 bp, or <100 bp. In other preferred embodiments, the target cell-free methylated DNA is limited to fragments with lengths between about 100 and about 150 bp, 110 and 140 bp, or 120 and 130 bp.
[0050] In some embodiments, fragment length is measured as the fragment length distribution of the target cell-free methylated DNA. In some preferred embodiments, the target cell-free methylated DNA is limited to fragments based on the 50th, 45th, 40th, 35th, 30th, 25th, 20th, 15th, or 10th percentile of the length.
[0051] When the term “at least,” “greater than,” or “greater than or equal to” precedes the first value in a series of two or more values, the term “at least,” “greater than,” or “greater than or equal to” applies to each value in that series. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.
[0052] When the term “not greater than,” “less than,” or “less than or equal to” precedes the first value in a series of two or more values, the term “not greater than,” “less than,” or “less than or equal to” applies to each value in that series. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.
[0053] DNA methylation can be performed using any of the methods disclosed herein to process and enrich cell-free DNA (cfDNA) to obtain methylated cfDNA. As shown in Figure 1, cfDNA can undergo end repair, A-tailing, and adaptor ligation, or other preparations thereof, to produce a DNA library for downstream sequencing. The library can be combined with supplementally processed DNA (e.g., filler DNA) and / or spiked-in DNA to produce a sample mixture, which is then subjected to thermal denaturation and immunoprecipitation to enrich methylated cfDNA. Immunoprecipitation may include combining the sample mixture with any of the conjugates disclosed herein and a solid matrix (e.g., multiple magnetic beads). In some cases, immunoprecipitation may include combining the sample mixture with a methylated nucleic acid capture reagent, wherein the methylated nucleic acid capture reagent comprises an incubation mixture of the conjugate and a solid matrix. After enrichment of methylated cfDNA, the methylated cfDNA may be amplified and then sequenced (e.g., Illumina sequencing reaction).
[0054] Cell-free DNA (cfDNA) may be present in non-invasively collectable biological samples (e.g., blood, urine, saliva, cerebrospinal fluid (CSF), etc.) and may contain cfDNA derived from healthy tissue and cfDNA derived from tumors or cancer cells.Heterogeneous populations of cfDNA (e.g., circulating tumor DNA (ctDNA)). Cancer development may be associated with focal increases in 5'-methylcytosine (5mC), for example, increases at cytosine-phosphate-guanine (CpG) islands and CpG island shores. Cancer development may also be associated with global (e.g., genome-wide) cytosine demethylation (e.g., global loss of 5mC). In some cases, ctDNA can be distinguished from cfDNA molecules derived from healthy tissues (e.g., non-tumor and / or non-cancerous tissues) by the methylation level of the nucleic acid molecule (e.g., the percentage of methylated nucleotide residues). In some cases, nucleic acid molecules from tumor tissue and / or cancerous tissue may be hypomethylated (e.g., may contain lower levels of methylation, such as fewer methylated nucleotide residues and / or a lower percentage of methylated nucleotide residues) compared to nucleic acid molecules from healthy tissue or derived from healthy tissue (e.g., nucleic acid molecules from healthy tissue or derived from healthy tissue that consist of or contain nucleotide sequences corresponding to the same regions of the genome of the subject). For example, tumor-derived nucleic acid molecules (e.g., ctDNA molecules) may contain one or more regions that have fewer methylated nucleotide residues than nucleic acid molecules from healthy tissue (e.g., non-tumor and / or non-cancer tissue) derived from the same biological sample. In some cases, nucleic acid molecules from tumor tissue and / or cancerous tissue, or nucleic acid molecules derived from healthy tissue (e.g., nucleic acid molecules composed of or containing nucleotide sequences corresponding to the same regions of the genome of the subject, as specified in the manual on page 10 / 55, 21 CN 121464223 A), may be highly methylated (e.g., may contain a higher level of methylation, such as a larger number of methylated nucleotide residues and / or a higher percentage of methylated nucleotide residues). For example, a tumor-derived nucleic acid molecule (e.g., a ctDNA molecule) may contain one or more regions having a greater number of methylated nucleotide residues than a nucleic acid molecule derived from healthy tissue (e.g., non-tumor and / or non-cancer tissue) from the same biological sample. In some cases, tumor-derived portions of multiple cell-free DNA molecules (e.g., ctDNA) can be distinguished from cfDNA molecules derived from healthy tissue by one or more biophysical properties (e.g., the length of the cfDNA molecule or the presence of morphological 5' and 3' end sequence motifs) and / or one or more fragment omics patterns. For example,The nucleic acid length of ctDNA molecules may be shorter than that of cfDNA molecules derived from healthy tissue. In some cases, ctDNA molecules may contain defined 5' and 3' end motifs. In some cases, one or more of these distinguishing features can be used to deplete cfDNA derived from healthy tissue in a population of nucleic acid molecules and / or enrich ctDNA in a population of nucleic acid molecules. ctDNA typically has a shorter fragment length compared to cfDNA derived from healthy tissue.
[0055] Nucleic acid molecules (e.g., ctDNA) derived from tumor or cancer cells or tissues may be present in biological samples (and / or populations of nucleic acids derived from biological samples) in much lower amounts than nucleic acid molecules (e.g., cfDNA) derived from healthy tissue. It may be difficult to detect or sequence ctDNA present in biological samples or in multiple nucleic acid molecules (e.g., cfDNA) derived from biological samples, for example, because they are present in samples in lower amounts relative to cfDNA derived from healthy tissue (e.g., this may require a larger amount of potentially scarce biological samples and / or may require much higher sequencing depths).
[0056] Depletion (e.g., removal) of all or part of a population of methylated DNA molecules (e.g., molecules having increased nucleotide methylation levels in the whole or a subset of a genomic region represented by a population of nucleic acid molecules in a biological sample) from a plurality of nucleic acid molecules (e.g., a plurality of cell-free nucleic acid molecules or their amplicones constituting a biological sample) can produce a residual population of a plurality of nucleic acids in a biological sample, which can be used to determine the presence and / or sequence identity of ctDNA molecules in a biological sample. Typically, depletion / removal can be performed by pulling down methylated DNA molecules using a binding agent that is specific to them. Typically, the pull-down is collected and the flow containing unmethylated / hypomethylated DNA molecules is discarded. This disclosure provides for the first time methods and systems for collecting such flow containing unmethylated / hypomethylated DNA molecules and using methylated / hypomethylated DNA molecules or their derivatives to generate sequencing libraries.
[0057] In some cases, the depleted sequencing libraries of the methods, systems, compositions and kits disclosed herein may consist of or may contain such a residual population of nucleic acid molecules. In some cases, nucleic acid molecules depleted of methylation in one or more specific regions of the genomic sequence of a nucleic acid molecule (e.g., CpG islands, CpG island shores, or repetitive sequences in the genome, such as long dispersed nuclear elements (LINEs), short dispersed nuclear elements (SINEs), or LTRs (long terminal repeats)) can be sufficient to increase the effectiveness of assays used to determine the presence or absence of ctDNA molecules or sequence identity in those nucleic acids.Sensitivity and / or increased specificity. In some cases, multiple nucleic acids (e.g., cfDNA molecules or their amplicones derived from a biological sample) may undergo genome-wide depletion of nucleic acid molecules methylated in one or more specific regions of the genomic sequence of the nucleic acid molecule (e.g., CpG islands, CpG island shores, or repetitive sequences of the genome, such as long dispersed nuclear elements (LINE), short dispersed nuclear elements (SINE), or LTRs (long terminal repeats)) to achieve increased sensitivity and / or increased specificity in assays used to determine the presence or absence of ctDNA molecules or sequence identity in these multiple nucleic acids. In some cases, the remaining population (e.g., multiple nucleic acid fragments that can be used to create a depleted library) may lack CpG islands. In some cases, the remaining population (e.g., multiple nucleic acid fragments that can be used to create a depleted library) may contain one or more of the following: long dispersed nuclear elements (LINE), short dispersed nuclear elements (SINE), or long terminal repeat (LTR) elements.
[0058] Enriching all or a portion of a population of methylated DNA molecules (e.g., molecules having increased nucleotide methylation levels in the whole or a subset of a genomic region represented by multiple nucleic acid molecules in a biological sample) from multiple nucleic acid molecules (e.g., multiple cell-free nucleic acid molecules or their amplicones constituting a biological sample) can generate multiple nucleic acid populations in a biological sample that can be used to determine the presence and / or sequence identity of ctDNA molecules in the biological sample. Enrichment can be performed by pulling down the methylated DNA molecules using a conjugate that is specific to them. The pull-down can be collected, and the flow containing unmethylated / hypomethylated DNA molecules can be discarded or alternatively collected (e.g., for generating a depleted library as described in this disclosure). The enriched portion can then be sequenced.
[0059] Depletion or enrichment of all or a portion of methylated nucleic acid molecules in multiple nucleic acid molecules of a biological sample can include contacting the methylated nucleic acid molecules with a conjugate (e.g., an affinity molecule, such as an antibody or protein, that is specific to methylated nucleotide residues). For example, the creation of sequencing libraries may include contacting multiple nucleic acid molecules (e.g., cfDNA molecules) or their amplicon with conjugates that are selective for methylated regions of nucleic acid molecules (e.g., methylcytosine conjugates (MBDs), such as MBD-Fc fusion proteins). In some cases, the conjugate may be specific to one or more methylated nucleotide species (e.g., 5-methylcytosine (5mC)). Cell-free methylated DNA immunoprecipitation sequencing (cfMeDIP-seq), a whole-genome molecular profiling technique, can be performed by using conjugates (such as anti-5-methylcytosine...)Iridium (anti-5mC) antibodies or methyl-CpG-binding domain (MBD) proteins (e.g., MBD-Fc fusion proteins) can be used to enrich methylated cfDNA fragments. As described herein, cfMeDIP-seq can include as part of methods and systems for depleting methylated DNA fragments in cfDNA samples, leaving hypomethylated or unmethylated cfDNA fragments (such as ctDNA). Thus, the identification of hypomethylated or unmethylated cell-free DNA in clinical samples can be used to determine the presence of tumors or cancer in the subject.
[0060] In some cases, the depletion of multiple nucleic acid molecules (e.g., in creating depleted sequencing libraries and / or determining the presence or sequence identity of nucleic acid molecules) can include the removal of one or more nucleic acid molecules with methylation levels above a threshold methylation level (e.g., where the removed one or more nucleic acid molecules are hypermethylated, e.g., relative to one or more nucleic acid molecules not removed during depletion). In some cases, the enrichment of multiple nucleic acid molecules (e.g., in creating enriched sequencing libraries and / or determining the presence or sequence identity of nucleic acid molecules) may involve the removal of one or more nucleic acid molecules with methylation levels below a threshold methylation level (e.g., where the removed one or more nucleic acid molecules are hypomethylated, e.g., relative to one or more unenriched or unmethylated nucleic acid molecules). In some cases, the methylation level of a specific nucleic acid fragment (e.g., a DNA fragment) may be considered to have reached a threshold methylation level when a binder with sufficient specificity for methylated cytosine is able to bind to the specific nucleic acid fragment with or without the filler DNA described herein. In some cases, the methylation level of a specific nucleic acid fragment (e.g., a DNA fragment) may be considered to be below a threshold methylation level when a binder with sufficient specificity for methylated cytosine cannot bind to the specific nucleic acid fragment with or without the filler DNA described herein. In some cases, the depletion of multiple nucleic acid molecules (e.g., in creating a depleted sequencing library and / or determining the presence or sequence of nucleic acid molecules) produces (e.g., provides) a population of remaining multiple nucleic acid molecules, wherein the remaining portion of the multiple nucleic acid molecules includes (or, in some cases, consists of) nucleic acid molecules with methylation levels below a threshold methylation level (e.g., wherein the remaining population is hypomethylated / unmethylated relative to one or more nucleic acid molecules removed from the multiple nucleic acid molecules during depletion). The methylation level can be calculated as a percentage of highly methylated nucleic acid fragments compared to all nucleic acid fragments contained in the sample. In some cases, the threshold methylation level can be 0.1% to 1%, 1% to 5%, 5% to 10%, 10% to 15%, 15% to 20%, 20% to 25%, 25% to 30%, 30% to 35%, 35%. [Specification 12 / 55 pages 23 CN 121464223 A]Up to 40%, 40% to 45%, 45% to 50%, 50% to 55%, 55% to 60%, 65% to 70%, 70% to 75%, 75% to 80%, 80% to 85%, 85% to 90%, 95% to 100%, at least 1%, at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%. At least 35%, at least 40%, at least 45%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at most 1%, at most 5%, at most 10%, at most 15%, at most 20%, at most 25%, at most 30%, at most 35%, at most 40%, at most 45%, at most 50%, at most 55%, at most 60%, at most 65%, at most 70%, at most 75%, at most 80%, at most 85%, at most 90%, at most 95%, or at most 100%.
[0061] In some cases, a first plurality of nucleic acid molecules (e.g., including nucleic acid molecules from a biological sample of an object, such as cfDNA) may be combined (e.g., mixed) with a second plurality of nucleic acid molecules (e.g., wherein the second plurality of nucleic acid molecules do not come from an object from which a biological sample is obtained), for example, as shown in FIG1. In some cases, the second plurality of nucleic acid molecules include supplementally processed DNA (e.g., including lDNA). In some cases, each of the second plurality of nucleic acid molecules is not aligned with the human genome.
[0062] In some cases, the methods or systems disclosed herein may include, for example, using a sequencer to determine or identify all or a portion of the sequence of a depleted population of nucleic acid molecules (e.g., the remaining population of multiple nucleic acid fragments in a biological sample after pulling down a hypermethylated nucleic acid fragment). In some cases, the remaining population of nucleic acid molecules may be purified (e.g., after library creation) to produce a plurality of purified nucleic acid molecules, for example, before or as part of a process for determining or identifying all or a portion of the sequence of the depleted population of nucleic acid molecules. In some cases, all or a portion of the purified nucleic acid molecules may be amplified (e.g., via polymerase chain reaction), for example, before or as part of a process for determining or identifying all or a portion of the sequence of the depleted population of nucleic acid molecules. In some cases, the amplified population of nucleic acid molecules or derivatives thereof (e.g., amplicons comprising all or a portion of the purified nucleic acid molecules) may be sequenced (e.g., for determining and / or identifying the sequence of nucleic acid molecules). In some cases, sequencing may be performed using a sequencer, as described herein. In some cases, arrays or polymerase chain reactions can be used to identify or determine the sequences of multiple nucleic acid molecules (or their derivatives) in a biological sample. In some cases, the presence of tumor-derived nucleic acid molecules...The presence of tumor-derived nucleic acid molecules can be determined by calculating the sum of reads per million kilobases (RPKM) in regions of the genome (e.g., all or part of the genome, such as CpG islands only or CpG island shores only). In some cases, the presence of tumor-derived nucleic acid molecules can be indicated when a depleted sequencing library (e.g., a population of the remaining nucleic acids) is observed to have a low sum of RPKM in one or more regions of interest (e.g., CpG islands or CpG island shores), such as less than 70,000, less than 60,000, less than 50,000, less than 40,000, or less than 30,000.
[0063] Supplemental processed DNA (filler DNA) In some cases, supplemental processed DNA (e.g., filler DNA, filler nucleic acid) can be added to a first plurality of nucleic acids (e.g., a plurality of nucleic acids from a biological sample, which may include cfDNA from healthy tissue and / or cfDNA from tumor tissue, such as ctDNA). In some cases, adding supplementally treated DNA (e.g., a second or more nucleic acid molecules) to a first or more nucleic acid molecules can increase the specificity and / or sensitivity of the methods, systems, or kits described herein, for example, regarding the detection and / or identification of the nucleic acid sequences of the first or more nucleic acid molecules. In some cases, adding supplementally treated DNA (e.g., a second or more nucleic acid molecules) to a first or more nucleic acid molecules can increase the depletion rate of methylated regions of the nucleic acid sequence, for example, in practice of some embodiments of the methods and systems described herein. In some cases, adding supplementally treated DNA (e.g., a second or more nucleic acid molecules) to a first or more nucleic acid molecules (e.g., cfDNA comprising a biological sample) can increase the selectivity of the conjugate for one or more (e.g., multiple) methylated regions of the first or more nucleic acid molecules. In some cases, supplementally treated DNA (e.g., a second or more nucleic acid molecules) can be added to a first or more nucleic acid molecules in an amount sufficient to achieve the desired total mass of the combined mixture of nucleic acid molecules. In some cases, the total mass required for the method or system described on page 13 / 55 of this specification (CN 121464223 A) may be 20 ng to 30 ng, 30 ng to 40 ng, 40 ng to 50 ng, 50 ng to 60 ng, 60 ng to 70 ng, 70 ng to 80 ng, 80 ng to 90 ng, 90 ng to 100 ng, 100 ng to 110 ng, 110 ng to 120 ng, 120 ng to 130 ng, 130 ng to 140 ng, 140 ng to 150 ng, 150 ng to 160 ng, 160 ng to 170 ng, 170 ng to 180 ng, 180 ng to 190 ng, 190 ng to 200 ng, or greater than 200 ng.ng or less than 20 ng. In some cases, 1 ng to 5 ng, 5 ng to 10 ng, 10 ng to 20 ng, 20 ng to 30 ng, 30 ng to 40 ng, 40 ng to 50 ng, 50 ng to 60 ng, 60 ng to 70 ng, 70 ng to 80 ng, 80 ng to 90 ng, 90 ng to 100 ng, 100 ng to 110 ng, 110 ng to 120 ng, 120 ng to 130 ng, 130 ng to 140 ng, 140 ng to 150 ng, 150 ng to 160 ng, 160 ng to 170 ng, 170 ng to 180 ng, 180 ng to 190 ng, 190 ng to 200 ng, greater than 200 ng, less than 20 A supplemental processed DNA amount of ng, less than 10 ng, or less than 5 ng may be added to the first plurality of nucleic acid molecules (e.g., to bring the total mixture of the supplemental processed DNA and the first plurality of nucleotide molecules to the desired total mass). In some embodiments, this disclosure includes methods and systems for filling a sample with a certain amount of supplemental processed DNA (e.g., filler DNA) to generate a mixture sample, wherein the mixture sample comprises a total amount of nucleic acid mixture of at least about 50 ng, 55 ng, 60 ng, 65 ng, 70 ng, 75 ng, 80 ng, 85 ng, 90 ng, 95 ng, 100 ng, 120 ng, 140 ng, 160 ng, 180 ng, 200 ng, or any amount between these numbers. In some embodiments, the supplemental DNA comprises at least about 5%, 10%, 15%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% methylated supplemental DNA, wherein the remainder is unmethylated supplemental DNA, and in some cases between 5% and 50%, between 10% and 40%, or between 15% and 30% methylated supplemental DNA. In some embodiments, the mixture sample comprises 20 ng to 100 ng, in some cases 30 ng to 100 ng, and in some cases 50 ng to 100 ng of supplemental DNA. In some embodiments, cell-free DNA from the sample and the first amount of supplemental DNA together comprise at least 50 ng of total DNA, and in some cases at least 100 ng of total DNA.
[0064] In some cases, the supplemental DNA may be generated by fragmentation (e.g., by sonication). In some implementations, the supplemental DNA can be 50 bp to 800 bp in length, in some cases 100 bp to 600 bp in length, and in some cases 200 bp in length.The length ranges from bp to 600 bp. In some embodiments, the supplemental DNA is double-stranded. The supplemental DNA can be double-stranded DNA. For example, the supplemental DNA can be useless DNA. The supplemental DNA can also be endogenous or exogenous DNA. For example, the supplemental DNA can be non-human DNA, and in some cases, λ DNA. As used herein, "λ DNA" generally refers to Enterobacterial phage λ DNA. In some embodiments, the supplemental DNA is substantially not aligned with human DNA.
[0065] In some cases, the supplemental DNA (e.g., filler DNA) increases the enrichment rate of one or more methylated regions by at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, at least about 90, at least about 95. At least approximately 100 times, at least approximately 150 times, at least approximately 200 times, at least approximately 300 times, at least approximately 400 times, at least approximately 500 times, at least approximately 600 times, at least approximately 700 times, at least approximately 800 times, at least approximately 900 times, or at least 1000 times. In some cases, supplemental treatment of DNA (e.g., filler DNA) will increase the enrichment ratio by up to approximately 1x, up to approximately 2x, up to approximately 3x, up to approximately 4x, up to approximately 5x, up to approximately 6x, up to approximately 7x, up to approximately 8x, up to approximately 9x, up to approximately 10x, up to approximately 15x, up to approximately 20x, up to approximately 25x, up to approximately 30x, up to approximately 35x, up to approximately 40x, up to approximately 45x, up to approximately 50x, up to approximately 55x, up to approximately 60x, up to approximately 65x, up to approximately 70x, up to approximately 75x, up to approximately 80x, up to approximately 85x, up to approximately 90x, up to approximately 95x, up to approximately 100x, up to approximately 150x, up to approximately 200x, up to approximately 300x, up to approximately [the specified value is missing from the original text]. Page 25 CN 121464223 A 400 times, up to about 500 times, up to about 600 times, up to about 700 times, up to about 800 times, up to about 900 times, or up to about 1000 times. In some cases, supplemental treatment of DNA (e.g., filler DNA) will increase the enrichment ratio by about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 15, about 20, about 25, about 30, about35 times, about 40 times, about 45 times, about 50 times, about 55 times, about 60 times, about 65 times, about 70 times, about 75 times, about 80 times, about 85 times, about 90 times, about 95 times, about 100 times, about 150 times, about 200 times, about 300 times, about 400 times, about 500 times, about 600 times, about 700 times, about 800 times, about 900 times, or 1000 times.
[0066] The sample can be any biological sample separated from the object. For example, samples may include, but are not limited to, bodily fluids, whole blood, platelets, serum, plasma, feces, leukocytes or white blood cells, endothelial cells, tissue biopsy, synovial fluid, lymph, ascites, interstitial or extracellular fluid, fluids in the intercellular space including gingival crevicular fluid, bone marrow, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, urine, fluid from a nasal brush, fluid from a Pap smear, or any other bodily fluid. Bodily fluids may include saliva, blood, or serum. Samples may also be tumor samples, which may be obtained from the subject by various methods (including but not limited to venipuncture, excretion, ejaculation, massage, biopsy, needle aspiration, irrigation, curettage, surgical incision or intervention or other methods). Samples may be cell-free samples (e.g., substantially cell-free). DNA samples may be denatured, for example, using sufficient heat.
[0067] Samples may be taken from subjects suffering from a disease or condition. Samples may be taken from subjects suspected of having a disease or condition. Samples may be taken from subjects suffering from more than one disease or condition. Samples may be taken from subjects suspected of having more than one disease or condition. In some embodiments, samples may be obtained before and / or after treatment of a subject with a disease or condition. Samples may be obtained from the subject during treatment or a treatment regimen. Multiple samples may be obtained from the subject to monitor the effectiveness of treatment over time. Diseases or conditions may be cancer. Specific examples of cancer types include those suitable for detection using the methods according to this disclosure, including acute lymphoblastic leukemia, acute myeloid leukemia, adrenocortical carcinoma, AIDS-related cancer, AIDS-related lymphoma, anal cancer, appendiceal cancer, astrocytoma, basal cell carcinoma, bile duct cancer, bladder cancer, bone cancer, brain tumors (such as cerebellar astrocytoma, cerebral astrocytoma / malignant glioma, ependymoma, medulloblastoma, supratentorial primitive neuroectodermal tumor, visual pathway and hypothalamic glioma). Tumors, breast cancer, colorectal cancer, hepatobiliary cancer, bronchial adenoma, Burkitt lymphoma, cancers of unknown primary origin, central nervous system lymphoma, cerebellar astrocytoma, cervical cancer, childhood cancer, chronic lymphocytic leukemia, chronic myeloid leukemia, chronic myelodysplastic syndrome, colon cancer, cutaneous T-cell lymphoma, desmoplastic small round cell tumor, endometrial cancer, ependymoma, esophageal cancer, Ewing sarcoma, germ cell tumors, gallbladder cancer, gastric cancer, gastrointestinal carcinoid tumors, gastrointestinal stromal tumors, gliomasTumors, hairy cell leukemia, head and neck cancer, heart cancer, hepatocellular carcinoma, Hodgkin's lymphoma, hypopharyngeal cancer, intraocular melanoma, islet cell carcinoma, Kaposi's sarcoma, kidney cancer, laryngeal cancer, lip and oral cancer, liposarcoma, liver cancer, lung cancer (such as non-small cell and small cell lung cancer), lymphoma, leukemia, macroglobulinemia, bone / osteosarcoma, malignant fibrous histiocytoma, medulloblastoma, melanoma, mesothelioma, occult primary and metastatic squamous cell carcinoma of the neck, oral cancer, multiple endocrine neoplasia syndrome, myelodysplastic syndrome, myeloid leukemia, nasal cavity and sinus cancer, nasopharyngeal carcinoma, neuroblastoma, non-Hodgkin's lymphoma, non-small cell lung cancer, oral cancer Cancer), oropharyngeal cancer, osteosarcoma / malignant fibrous histiocytoma of bone, ovarian cancer, ovarian epithelial cancer, ovarian germ cell tumor, pancreatic cancer, islet cell pancreatic cancer, sinus cancer and nasal cavity cancer, parathyroid cancer, penile cancer, pharyngeal cancer, pheochromocytoma, pineal astrocytoma, pineal germ cell tumor, pituitary adenoma, pleural pulmonary blastoma, plasmacytoma, primary central nervous system lymphoma, prostate cancer, rectal cancer, renal cell carcinoma, transitional cell carcinoma of the renal pelvis and ureter, retinoblastoma, rhabdomyosarcoma, salivary gland cancer, sarcoma, skin cancer, Merkel cell skin cancer, small intestine cancer, soft tissue sarcoma, squamous cell carcinoma, gastric cancer, T-cell lymphoma, laryngeal cancer, thymus. (Instructions for use 15 / 55, page 26, CN 121464223 A) Tumors, thymic carcinomas, thyroid cancers, trophoblastic tumors (during pregnancy), cancers of unknown primary site, urethral cancers, uterine sarcomas, vaginal cancers, vulvar cancers, Waldenström macroglobulinemia, and Wilms' tumors. In some embodiments, the cancer is a squamous cell carcinoma of the head and neck.
[0068] In some embodiments, the cancer may be an early-stage cancer. In some embodiments, an early-stage cancer may be a cancer confined to a small area and / or that has not yet spread to distant areas or nearby tissues. In some embodiments, an early-stage cancer may be a stage I cancer. In some embodiments, the cancer may be a stage I, II, III, or IV cancer.
[0069] In some embodiments, the cancer may be a low-shedding tumor (e.g., bladder cancer, breast cancer, endometrial cancer, prostate cancer, or kidney cancer). Low-shedding tumors may shed low levels of ctDNA. In some cases, low-shedding tumors may have a low ctDNA burden. In some cases, low levels of ctDNA may be found in early-stage cancers.
[0070] Samples may be taken from healthy individuals. In some cases, samples may be obtained longitudinally from the same individual. In some cases, longitudinally acquired samples can be analyzed for the purpose of monitoring individual health and early detection of health problems (e.g., early-stage cancer). In some implementations, samples can be collected in a home or point-of-care setting and subsequently analyzed.Previously transported via mail, courier, or other means of transport. For example, a home user may collect a blood spot sample via finger puncture, which can be dried and then delivered by mail prior to analysis. In some cases, longitudinally acquired samples may be used to monitor responses to stimuli expected to affect health, athletic performance, or cognitive performance. Non-limiting examples include responses to pharmacological treatments, diets, or exercise programs.
[0071] In some embodiments, this disclosure provides a system, method, or kit that includes or uses one or more biological samples. One or more samples as used herein may include any substance containing or presumed to contain nucleic acids. Samples may include biological samples obtained from an object. In some embodiments, the biological sample is a liquid sample.
[0072] In some embodiments, the sample contains less than about 100 ng, 90 ng, 80 ng, 75 ng, 70 ng, 60 ng, 50 ng, 40 ng, 30 ng, 20 ng, 10 ng, 5 ng, 1 ng, or any amount between these numbers of cell-free nucleic acid molecules. Further, in some embodiments, the sample contains less than about 1 pg, less than about 5 pg, less than about 10 pg, less than about 20 pg, less than about 30 pg, less than about 40 pg, less than about 50 pg, less than about 100 pg, less than about 200 pg, less than about 500 pg, less than about 1 ng, less than about 5 ng, less than about 10 ng, less than about 20 ng, less than about 30 ng, less than about 40 ng, less than about 50 ng, less than about 100 ng, less than about 200 ng, less than about 500 ng, less than about 1000 ng, or any amount between these numbers.
[0073] In some cases, creating or providing multiple nucleic acid molecules from a biological sample may include end repair, A-tailing, and adaptor ligation of multiple nucleic acid molecules (e.g., after purification from the biological sample).
[0074] In some embodiments, samples can be collected and sequenced at a first time point, and then another sample can be collected and sequenced at a subsequent time point. This approach can be used, for example, for longitudinal monitoring purposes, to track the development or progression of a disease. In some embodiments, disease progression can be tracked before, after, or during treatment to determine the effectiveness of the treatment. For example, the methods described herein can be performed on subjects before and after drug treatment to measure the progression or regression of the disease response to drug treatment.
[0075] After obtaining samples from subjects, the samples can be processed to generate a dataset indicating the subject's disease or condition. For example, cell-free nucleic acid analysis of samples at a set of cancer-associated genomic loci or microbiome-associated loci.The presence, absence, or quantitative assessment of a molecule (e.g., ctDNA) can indicate cancer in the subject. Processing a sample obtained from the subject may include (i) subjecting the sample to conditions sufficient to isolate, enrich, or extract multiple cell-free nucleic acid molecules, and (ii) determining multiple cell-free nucleic acid molecules to generate a dataset (e.g., nucleic acid sequences). In some embodiments, multiple cell-free nucleic acid molecules are extracted from the sample and sequenced to generate multiple sequencing reads. Specification 16 / 55 pages 27 CN 121464223 A
[0076] In some embodiments, cell-free nucleic acid molecules may include cell-free ribonucleic acid (cfRNA) or cell-free deoxyribonucleic acid (cfDNA). Cell-free nucleic acid molecules (e.g., cfRNA or cfDNA) can be extracted from the sample by multiple methods. Cell-free nucleic acid molecules can be enriched by multiple probes configured to enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to a set of cancer-associated genomic loci. The probes may have sequence complementarity with nucleic acid sequences from one or more of the set of cancer-associated genomic loci. This set of cancer-associated genomic loci may contain at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, at least about 90, at least about 95, at least about 100 or more different cancer-associated genomic loci. The probe may be a nucleic acid molecule (e.g., RNA or DNA) that is sequence complementary to the nucleic acid sequence (e.g., RNA or DNA) of one or more genomic loci (e.g., cancer-associated genomic loci). These nucleic acid molecules can be primers or enriched sequences. Assaying a sample using probes selective for one or more genomic loci (e.g., cancer-associated genomic loci or microbiome-associated loci) can include using array hybridization, polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., RNA sequencing or DNA sequencing).
[0077] Certain methods for capturing cell-free methylated DNA are described in WO 2017 / 190215 and WO 2019 / 010564, both of which are incorporated herein by reference in their entirety for all purposes.
[0078] Various sequencing library assays can be used in the methods of this disclosure, such as library preparation (which may include polymerase chain reaction).(PCR), followed by sequencing (e.g., next-generation sequencing, Sanger sequencing, etc.). Next-generation sequencing (NGS) technology, also known as high-throughput sequencing, can include a variety of sequencing technologies, including: Illumina (Solexa) sequencing, Roche 454 sequencing, ion flooding: proton / PGM sequencing, SOLiD sequencing, or long-read sequencing. NGS allows for sequencing of DNA and RNA much faster and cheaper than the previously used Sanger sequencing. In some embodiments, the sequencing is optimized for short-read sequencing.
[0079] Hypermethylated sequencing libraries can improve the specificity, sensitivity, and / or efficiency of methods and systems used to process nucleic acids. For example, hypermethylated sequencing libraries can improve the specificity, sensitivity, and / or efficiency of assays used to determine the presence and / or sequence identity of nucleic acid sequences. Hypermethylated sequencing libraries can contain multiple nucleic acids and / or fragments thereof. In some cases, hypermethylated sequencing libraries can contain multiple nucleic acid molecules (e.g., nucleic acid populations and / or fragments thereof). Multiple nucleic acid molecules may comprise all or a portion of a first plurality of nucleic acid molecules, for example, wherein the first plurality of nucleic acid molecules comprises one or more nucleic acid molecules containing methylated nucleic acid residues and one or more nucleic acid molecules not containing methylated nucleotide residues. In some cases, methylated nucleic acids may comprise one or more methylated nucleic acid residues. For example, methylated nucleic acids may comprise one or more methylated cytosines (e.g., one or more 5-methylcytosine (5mC) and / or one or more 5-hydroxymethylcytosine (5hmC)). Multiple nucleic acid molecules (e.g., multiple nucleic acid molecules derived from a biological sample) can be hypermethylated and enriched using conjugates, for example, as described herein, to form a hypermethylated sequencing library that can be used as a novel background instead of a whole-genome background for cfDNA analysis. In some cases, DNA can be hypermethylated and then a sequencing library with a novel background can be created using conjugates. The novel background sequencing library may comprise a set of background genomic regions enriched by the conjugates.
[0080] In some cases, multiple nucleic acids can be targeted-specifically enriched. For example, a given sequence in a sequencing library can be enriched by pulling down nucleic acids using a capture probe. In another example, nucleic acids can be amplified using primers to enrich a given sequence in a sequencing library. Capture probes or primers can be specific to a particular gene, non-coding region, or other sequence. (Instructions 17 / 55, page 28, CN 121464223 A) For example, one or more genes in nucleic acids can be enriched. These genes may be cancer-related genes. For example, these genes may be genes previously identified as being associated with a certain type of cancer. For example, genes or sequences in nucleic acids having previously identified methylation states or previously identified numbers of methylated nucleotides can be enriched. TargetingSpecific enrichment can be performed during the preparation of the sequencing library. For example, methylated nucleic acids in nucleic acids can be enriched or depleted (as described in this disclosure), and then the nucleic acids can be targeted for specific enrichment.
[0081] Nucleic Acid Molecule Sequencing This disclosure provides methods and techniques for determining the nucleotide base sequence in one or more polynucleotides. The polynucleotide can be, for example, a nucleic acid molecule, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), including its variants or derivatives (e.g., single-stranded DNA). Sequencing can be performed by next-generation sequencing.
[0082] Further, any sequencing method that provides fragment length can be utilized, such as paired-end sequencing. Alternatively or additionally, sequencing can be performed using nucleic acid amplification, polymerase chain reaction (PCR) (e.g., digital PCR, quantitative PCR, or real-time PCR), or isothermal amplification. Such a system can provide multiple raw genetic data corresponding to the genetic information of an object (e.g., a human), such as those generated by the system from a sample provided by the object. In some examples, such systems provide sequencing reads (also referred to herein as “reads”). A read can include a string of nucleic acid bases corresponding to the sequence of a nucleic acid molecule that has been sequenced. In some cases, the systems and methods provided herein can be used in conjunction with proteomics information.
[0083] In some embodiments, sequencing reads are obtained via next-generation sequencing methods or next-next-generation sequencing methods. In some embodiments, the sequencing methods include cfMeDIP sequencing, for example, including processes or systems as described by Shen et al. (“Sensitive tumor detection and classification using plasma cell-free DNA methylomes,” (2018) Nature), which is incorporated herein in its entirety. In some embodiments, sequencing may be whole-genome sequencing. In some embodiments, sequencing may be whole-exome sequencing. In some embodiments, sequencing may be partial-genome sequencing. In some embodiments, sequencing may be targeted-genome sequencing. In some embodiments, sequencing may be whole-methylome sequencing. In some embodiments, sequencing may be whole-methylome sequencing without whole-genome sequencing. In some embodiments, whole-methylome sequencing may be captured without the need for predefined target panels. In some embodiments, sequencing may be partial-methylome sequencing. In some embodiments, sequencing may be targeted-methylome sequencing. In some implementations, sequencing can be performed using methyl CpG-binding domain sequencing (MBD-seq). In some cases, MBD-seq may include capture...(For example, via a conjugate, such as an antibody specific to a class of methylated nucleotides) double-stranded methylated DNA fragments are used to sequence a library of methylated DNA fragments. In some embodiments, the sequencing method includes cancer personalized analysis via deep sequencing (CAPP-Seq), a next-generation sequencing-based method for quantifying circulating DNA (ctDNA) in cancer. This method can be generalized to any cancer type shown to have recurrent mutations and can detect one mutant DNA molecule in 10,000 healthy DNA molecules. In some embodiments, sequencing may include chemical transformation. In some embodiments, sequencing includes bisulfite sequencing. In some embodiments, sequencing does not include bisulfite sequencing.
[0084] Sequencing may include targeted sequencing. For example, the sequencing reaction may include a capture probe specific to the region of interest. The use of targeted sequencing can increase the amount of reads specific to informative regions (e.g., related to DMR, or that can be used to distinguish healthy subjects from those with disease). The capture probe may comprise one or more probes that are complementary or homologous to a region containing one or more sites susceptible to enzymatic methylation. The capture probe may contain one or more probes that are complementary to or homologous to one or more sites contained in healthy or non-disease controls that are substantially unmethylated. The capture probe may contain one or more probes that are complementary to or homologous to regions having a known methylation state. Targeted sequencing may target one or more regions known to be present in multiple cancer types.
[0085] In some embodiments, sequencing maintains fragment length and / or terminal motifs. Non-limiting examples of terminal motifs include 5' terminal motifs, 3' terminal motifs, 6-mer terminal motifs, 5-mer terminal motifs, 4-mer terminal motifs (e.g., CCCA, CCTG, CCAG, CCAA, CCAT, CCTG, CCAA, CCCT, CCTC, TGTG, TGTT, CCTA, TATT, CCAC, TCTT, CCCC, TATA, TAAA, AAAA, TTTT or other variants thereof). Fragment length and end motif retention can be assessed using a single assay for multi-omics evaluation, which can improve efficiency and performance in cases of low ctDNA levels. In some embodiments, sequencing may include enzyme transformation. In some cases, enzyme transformation may include transformation with TET2 and an oxidation enhancer. In some cases, enzyme transformation may also include APOBEC. In some embodiments, sequencing does not include enzyme transformation. In some embodiments, the absence of transformation (e.g., chemical or enzymatic) can preserve DNA and the four DNA bases (e.g., adenine, cytosine, etc.).The quality of guanine, thymine, and other compounds (such as guanine and thymine) is important. Maintaining DNA quality allows for more reads to pass through quality control filters that are uniquely aligned with the human genome.
[0086] In some cases, the sample or a portion thereof (e.g., multiple nucleic acids of the sample) can be library prepared prior to sequencing. Briefly, after end repair and A-tailing, the sample is ligated to a nucleic acid adaptor and digested with an enzyme.
[0087] In some embodiments, sequencing includes modification of nucleic acid molecules or fragments thereof, for example, by attaching a barcode, a unique molecular identifier (UMI), or another tag to the nucleic acid molecule or fragment thereof. Attaching a barcode, UMI, or tag to one end of a nucleic acid molecule or fragment thereof facilitates post-sequencing analysis of the nucleic acid molecule or fragment thereof. In some embodiments, the barcode is a unique barcode (e.g., a UMI). In some implementations, the barcode is not unique, and the barcode sequence can be used in conjunction with endogenous sequence information, such as the start and stop sequences of the target nucleic acid (e.g., barcodes and barcode sequences flanking the target nucleic acid, which bind to the sequences at the start and end of the target nucleic acid to form a uniquely tagged molecule). The barcode, UMI, or tag can be a known sequence used to associate a polynucleotide or fragment thereof with an input or target nucleic acid molecule or fragment thereof. The barcode, UMI, or tag can include natural nucleotides or non-natural (e.g., modified) nucleotides (e.g., as described herein). The barcode sequence can be included in an adaptor sequence such that the barcode sequence can be included in sequencing reads. The length of the barcode sequence can include at least 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, or more nucleotides. In some cases, barcode sequences can be long enough and sufficiently different from another barcode sequence to allow for sample identification based on the barcode sequence associated with them. Barcode sequences or combinations of barcode sequences can be used to tag and subsequently identify “original” nucleic acid molecules or fragments thereof (e.g., nucleic acid molecules or fragments thereof present in a sample from a subject). In some cases, barcode sequences or combinations of barcode sequences are used in conjunction with endogenous sequence information to identify original nucleic acid molecules or fragments thereof. For example, barcode sequences or combinations of barcode sequences can be used with endogenous sequences adjacent to barcodes, UMIs, or tags (e.g., the beginning and end of the endogenous sequence).
[0088] As described herein, the prepared library can be combined with filler nucleic acids (e.g., filler λ DNA) to minimize the impact of low abundance ctDNA in the prepared library and generate a mixed sample. In some embodiments, when the disease / condition is localized (non-metastatic) cancer, the amount of ctDNA can be low and can be difficult and not easily and accurately measured and quantified. In this case, the mixed sample can reach at least about 50 ng, 80 ng, 100 ng.ng, 120 ng, 150 ng, or 200 ng, and further enrichment.
[0089] Processing nucleic acid molecules or fragments thereof may include performing nucleic acid amplification. Amplification prior to sequencing of nucleic acids (e.g., enriched methylated nucleic acids) can generate more sequence reads. Amplification can allow the generation of more stable nucleic acids (e.g., double-stranded nucleic acids compared to partially single-stranded or single-stranded nucleic acids), which in turn can allow nucleic acids to be stored for longer periods without degradation. For example, any type of nucleic acid amplification reaction can be used to amplify target nucleic acid molecules or fragments thereof and generate amplification products. Non-limiting examples of nucleic acid amplification methods include reverse transcription, primer extension, polymerase chain reaction (PCR), ligase chain reaction, asymmetric amplification, rolling circle amplification, and multiplex substitution amplification (MDA). Examples of PCR include, but are not limited to, quantitative PCR, real-time PCR, digital PCR, emulsion PCR, hot-start PCR, multiplex PCR, asymmetric PCR, nested PCR, and assembly PCR. Nucleic acid amplification may involve one or more reagents, such as one or more primers, probes, polymerases, buffers, enzymes, and deoxyribonucleotides. Nucleic acid amplification may be isothermal or may include thermal cycling. And / or is related to the length of the endogenous sequence. In some cases, PCR amplification includes amplification of at least 10 cycles, at least 11 cycles, at least 12 cycles, at least 13 cycles, or at least 14 cycles.
[0090] Amplification to generate nucleic acids suitable for sequencing can be performed on nucleic acids bound to or attached to a solid matrix (e.g., beads). For example, as described elsewhere in this disclosure, methylation-binding molecules may be attached to a solid matrix, and methylation-binding molecules may bind nucleic acids. The bound nucleic acids may be amplified, and amplicons that do not bind to the binding molecules or the solid matrix may be generated. Amplification of nucleic acids bound to a solid matrix can increase throughput, for example, by reducing the need for washing or buffer exchanges that may occur during elution or removal of nucleic acids from the solid matrix.
[0091] Processing nucleic acid molecules or fragments thereof may include adding beads that can bind to nucleic acids. Nucleic acid binding allows nucleic acids to be specifically bound and allows for the removal of enzymes or contaminants from nucleic acid samples. Beads (e.g., SPRI beads) can bind to nucleic acids in a library and can be washed or separated to remove contaminants. Nucleic acids can be eluted from the beads and then collected. This can be done by removing the beads from the nucleic acid sample. For example, the beads can be magnetic beads and can be subjected to a magnetic field to remove the beads. The beads can undergo multiple iterations of the removal process to reduce or minimize bead residue in subsequent processing reactions.
[0092] Conjugates can be used to deplete or enrich populations of nucleic acid molecules (e.g., multiple nucleic acid molecules derived from biological samples).In some cases, the conjugate can be used to deplete or enrich one or more nucleic acid molecules with methylation levels equal to or higher than a threshold methylation level (e.g., by binding to one or more methylated nucleotides of one or more nucleic acid molecules). The conjugate can be used to enrich a population of nucleic acid molecules (e.g., multiple nucleic acids derived from a biological sample). The conjugate can be a molecule that specifically binds to methylated nucleic acids or methylated nucleotides. In some cases, the conjugate can be specific to one or more types of methylated nucleotides (e.g., 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), 4-methylcytosine (4mC), or 6-methyladenine (6mA)). In some cases, the conjugate can be selected from anti-5-methylcytosine antibodies or derivatives thereof, anti-5-carboxycytosine antibodies or derivatives thereof, anti-5-formylcytosine antibodies or derivatives thereof, anti-5-hydroxymethylcytosine antibodies or derivatives thereof, anti-3-methylcytosine antibodies or derivatives thereof, and any combination thereof. In some cases, the conjugate can be an anti-5-methylcytosine antibody or a derivative thereof. In some embodiments, the conjugate is a protein containing a methyl-CpG-binding domain. One such protein is the MBD2 protein. As used herein, "methyl-CpG-binding domain (MBD)" generally refers to certain domains of proteins and enzymes that are approximately 70 residues long and bind to DNA containing one or more symmetrically methylated CpGs. The MBDs, MBD1, MBD2, MBD4, and BAZ2 of MeCP2 mediate binding to DNA, and in the case of MeCP2, MBD1, and MBD2, binding to methylated CpGs is preferentially mediated. Human proteins MECP2, MBD1, MBD2, MBD3, and MBD4 contain a family of nucleoproteins associated with the presence of each methyl-CpG-binding domain (MBD). Each of these proteins, except MBD3, is capable of specifically binding to methylated DNA. The conjugate may include biotin, and may allow the conjugate to couple or bind to streptavidin.
[0093] In other embodiments, the conjugate is an antibody, and capturing cell-free methylated DNA includes immunoprecipitation of cell-free methylated DNA using the antibody. As used herein, “immunoprecipitation” generally refers to a technique of precipitating antigens (such as peptides and nucleotides) from solution using an antibody that specifically binds to a particular antigen. This process can be used to isolate and concentrate specific proteins or DNA from a sample and may require the antibody to be coupled to a solid matrix at some point in the process. Solid matrices include, for example, beads, such as magnetic beads. Other types of beads and solid matrices may be used. Various protein or chemical moieties may be present on a solid matrix that allows coupling or binding with the conjugate. For example, the solid matrix may be able to bind to...A methylated nucleic acid is specifically bound or coupled to a molecule. For example, a solid matrix may contain streptavidin and may be able to bind biotinylated antibodies. In another example, a solid matrix may contain protein A and may be able to bind the Fc domain of an antibody.
[0094] In various aspects, a conjugate (e.g., a methylated binding molecule) is added to a sample to bind nucleic acids. The conjugate may be coupled to a solid matrix (e.g., magnetic beads), and such a complex (e.g., a methylated nucleic acid capture reagent) may be added to a sample. The complex may initially be generated by pre-incubation without a sample. For example, an anti-5-mC antibody may be incubated with protein A beads, allowing the antibody to bind to or couple the protein A beads via the Fc domain. This may saturate the beads with the antibody, and such a complex may be able to bind methylated nucleic acids. After the complex is generated, it may be added to a sample to bind nucleic acids. The conjugate and the solid matrix may also be added to the sample simultaneously (or substantially simultaneously), or they may be added sequentially. For example, magnetic protein A beads and antibodies can be added to a sample simultaneously, allowing the antibody to bind to nucleic acids while simultaneously allowing the antibody to bind to the magnetic protein A beads. In another example, the antibody can be initially added to the sample to allow nucleic acids to bind to the antibody, and then beads are added to bind to the antibody that binds to the nucleic acids. Each method has advantages in generating the complex by pre-incubation, simultaneous addition, or sequential addition. For example, pre-incubation can allow the formation of a stable complex of the conjugate and a solid matrix without potentially sterically interfering with the nucleic acids.
[0095] For example, a 5-mC antibody (e.g., wherein the 5-mC antibody specifically binds to 5-methylcytosine) can be used as the conjugate. For the immunoprecipitation process, in some embodiments, at least 0.05 µg of antibody is added to the sample, while in some embodiments, at least 0.16 µg of antibody is added to the sample. In some cases, antibodies of 0.05 µg to 0.80 µg, 0.16 µg to 0.80 µg, 0.40 µg to 0.80 µg, 0.16 µg to 0.40 µg, 0.10 µg to 0.80 µg, 0.20 µg to 0.60 µg, 0.30 µg to 0.50 µg, or 0.40 µg to 0.50 µg can be used. To confirm the immunoprecipitation reaction, in some embodiments, the methods described herein further include the addition of a second amount of control DNA to the sample.
[0096] In some embodiments, the immunoprecipitation process is optimized, wherein optimizing the immunoprecipitation may include changing the conjugate used (e.g., antibody), adjusting the concentration of the conjugate, and / or adjusting the length of time the conjugate captures cell-free methylated DNA.
[0097] Methods for processing methylated nucleic acids This disclosure provides methods for processing cell-free nucleic acid samples from subjects to detect methylation events andSystem.
[0098] In some embodiments, the method shown in FIG25 (Workflow 1) includes: (a) providing a plurality of nucleic acid molecules (e.g., cfDNA, cfDNA with spike-in DNA) derived from a nucleic acid sample; (b) preparing a library using one or more custom adaptors (e.g., end repair, A-tailing, adaptor ligation) to generate a library; (c) adding a plurality of filler nucleic acid molecules to generate a sample mixture; (d) thermally denaturing and rapidly cooling the sample mixture; (e) performing immunoprecipitation on the sample mixture to obtain an enriched sample containing a plurality of methylated nucleic acid molecules; and (f) preparing the enriched sample for sequencing. In some instances of the method, (e) performing immunoprecipitation on the sample mixture further includes: i) adding a conjugate disclosed herein to the sample mixture; ii) adding a solid matrix (e.g., a magnetic solid matrix) to the sample mixture; iii) incubating the conjugate and solid matrix with the sample mixture for a sufficient time (e.g., overnight) to capture an enriched sample containing a plurality of methylated nucleic acid molecules; and iii) separating the conjugate and solid matrix from the enriched sample. In some specifications of this method, page 21 / 55, 32 CN 121464223 A, (f) preparing the enriched sample for sequencing further includes: i) cleaning the enriched sample to generate a cleaned enriched sample; ii) performing PCR amplification on the cleaned enriched sample to obtain multiple PCR amplicons; and iii) cleaning the PCR amplicons and then sequencing. In some cases, PCR amplification can be performed for up to 14 cycles.
[0099] In some embodiments, the method shown in FIG25 (Workflow 2) includes: (a) providing a plurality of nucleic acid molecules (e.g., cfDNA, cfDNA with spike-in DNA) derived from a nucleic acid sample; (b) preparing a library using one or more custom adaptors (e.g., end repair, A-tailing, adaptor ligation) to generate a library; (c) cleaning the library after preparation with a solid matrix (e.g., a magnetic solid matrix); (d) adding a plurality of filler nucleic acid molecules to generate a sample mixture; (e) thermally denaturing and rapidly cooling the sample mixture; (f) performing immunoprecipitation on the sample mixture to obtain an enriched sample containing a plurality of methylated nucleic acid molecules; and (g) preparing the enriched sample for sequencing. In some instances of these methods, (c) cleaning the library after preparation includes capturing any remaining solid matrix by placing the prepared library on a device (e.g., a magnetic rack). In some instances of these methods, (f) immunoprecipitation on the sample mixture further includes: i) adding a conjugate disclosed herein to the sample mixture; ii)Add a solid matrix (e.g., a magnetic solid matrix) to the sample mixture; iii) incubate the conjugate and solid matrix with the sample mixture for a sufficient time (e.g., overnight) to capture an enriched sample containing multiple methylated nucleic acid molecules; and iii) separate the conjugate and solid matrix from the enriched sample. In some of these methods, (g) preparing the enriched sample for sequencing further includes: i) cleaning the enriched sample to generate a cleaned enriched sample; ii) performing PCR amplification on the cleaned enriched sample to obtain multiple PCR amplicons; and iii) cleaning the PCR amplicons and then sequencing. In some cases, PCR amplification can be performed for up to 14 cycles.
[0100] In some embodiments, the method shown in FIG25 (Workflow 3) includes: (a) providing multiple nucleic acid molecules (e.g., cfDNA, cfDNA with spike-in DNA) derived from a nucleic acid sample; (b) performing library preparation (e.g., end repair, A-tailing, adaptor ligation) with one or more custom adaptors to generate a library; (c) performing post-library preparation cleaning with a solid matrix (e.g., magnetic solid matrix); (d) adding multiple filler nucleic acid molecules to generate a sample mixture; (e) thermally denaturing and rapidly cooling the sample mixture; (f) performing immunoprecipitation on the sample mixture to obtain an enriched sample containing multiple methylated nucleic acid molecules; and (g) preparing the enriched sample for sequencing. In some cases of this method, (c) performing post-library preparation cleaning includes additionally capturing the solid matrix (e.g., magnetic solid matrix) by placing the prepared library on a device (e.g., a magnetic rack) to capture any remaining solid matrix. In some of these methods, (f) immunoprecipitation of the sample mixture further includes: i) adding the conjugate disclosed herein to the sample mixture; ii) adding a solid matrix (e.g., a magnetic solid matrix) to the sample mixture; iii) incubating the conjugate and solid matrix with the sample mixture for a sufficient time (e.g., overnight) to capture an enriched sample containing multiple methylated nucleic acid molecules; and iii) separating the conjugate and solid matrix from the enriched sample. In some of these methods, (g) preparing the enriched sample for sequencing further includes: i) cleaning the enriched sample to generate a cleaned enriched sample; ii) performing PCR amplification on the cleaned enriched sample to obtain multiple PCR amplicons; and iii) cleaning the PCR amplicons and then sequencing. In some cases, PCR amplification can be performed for up to 13 cycles.
[0101] In some embodiments, the method shown in FIG25 (workflow 4) includes: (a) providing multiple nucleic acid molecules (e.g., cfDNA, cfDNA with spike-in DNA) derived from a nucleic acid sample; (b) using one or more custom-designed linkers.The methods involve: (a) performing library preparation (e.g., end repair, A-tailing, adaptor ligation) to generate a library; (b) cleaning the library after preparation with a solid matrix (e.g., a magnetic solid matrix); (c) adding multiple filler nucleic acid molecules to generate a sample mixture; (d) thermally denaturing and rapidly cooling the sample mixture; (e) performing immunoprecipitation on the sample mixture to obtain an enriched sample containing multiple methylated nucleic acid molecules; and (g) preparing the enriched sample for sequencing. In some cases described in this specification (page 22 / 55, 33 CN 121464223 A), (c) cleaning the library after preparation includes additionally capturing the solid matrix (e.g., a magnetic solid matrix) by placing the prepared library on a device (e.g., a magnetic rack) to capture any remaining solid matrix. In some instances of these methods, (f) immunoprecipitation of the sample mixture further includes: i) incubating the conjugate disclosed herein with a solid matrix (e.g., a magnetic solid matrix) to generate a methylated nucleic acid capture reagent; ii) adding the methylated nucleic acid capture reagent to the sample mixture; and ii) separating the methylated nucleic acid capture reagent after incubating the methylated nucleic acid capture reagent and the sample mixture together for a sufficient time (e.g., overnight) to obtain an enriched sample containing multiple methylated nucleic acid molecules. In some instances of these methods, (g) preparing the enriched sample for sequencing further includes: i) cleaning the enriched sample to generate a cleaned enriched sample; ii) performing PCR amplification on the cleaned enriched sample to obtain multiple PCR amplicons; and iii) cleaning the PCR amplicons and then sequencing. In some cases, PCR amplification can be performed for up to 13 cycles.
[0102] In some embodiments, the method shown in FIG28 includes: (a) providing a plurality of nucleic acid molecules (e.g., cfDNA, cfDNA with spike-in DNA) derived from a nucleic acid sample; (b) preparing a library using one or more custom adaptors (e.g., end repair, A-tailing, adaptor ligation) to generate a library; (c) cleaning the library after preparation with a solid matrix (e.g., a magnetic solid matrix); (d) adding a plurality of filler nucleic acid molecules to generate a sample mixture; (e) thermally denaturing and rapidly cooling the sample mixture; (f) performing immunoprecipitation on the sample mixture to obtain an enriched sample containing a plurality of methylated nucleic acid molecules; and (g) preparing the enriched sample for sequencing. In some instances of these methods, (c) cleaning the library after preparation includes capturing any remaining solid matrix by placing the prepared library on a device (e.g., a magnetic rack). In some instances of these methods, (f) immunoprecipitation on the sample mixture further includes: i) combining the conjugate disclosed herein with the solid matrix (e.g., a magnetic solid matrix)The process involves: ii) incubating the sample mixture to generate a methylated nucleic acid capture reagent; ii) adding the methylated nucleic acid capture reagent to the sample mixture; and ii) separating the methylated nucleic acid capture reagent after incubating the sample mixture with the methylated nucleic acid capture reagent for a sufficient time (e.g., overnight) to obtain an enriched sample containing multiple methylated nucleic acid molecules. In some cases of these methods, (g) preparing the enriched sample for sequencing further includes: i) cleaning the enriched sample to generate a cleaned enriched sample; ii) performing PCR amplification on the cleaned enriched sample to obtain multiple PCR amplicons; iii) cleaning the PCR amplicons; and iv) contacting the multiple methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences, followed by sequencing. In some cases, PCR amplification can be performed for up to 13 cycles. In some cases, one or more target sequences contain one or more genes.
[0103] In one aspect, this document discloses a method or system for processing a nucleic acid sample from an object, the method or system comprising: (a) generating a nucleic acid sample mixture containing a plurality of methylated nucleic acids from the object and a certain amount of supplementally treated DNA (e.g., filler DNA); and (b) incubating the nucleic acid sample mixture with (i) a methylation-binding molecule and (ii) a solid matrix; and (c) capturing the methylated nucleic acids to enrich the plurality of methylated nucleic acids (e.g., methylated single-stranded DNA) in the nucleic acid sample mixture. In some cases, a certain amount of supplementally treated DNA (e.g., filler DNA) is not required in (a). In some cases, the certain amount of supplementally treated DNA (e.g., filler DNA) contains at least one methylated DNA molecule. In some cases, the methylation-binding molecule is an antibody (e.g., an anti-5-methylcytosine (anti-5mC) antibody, a methyl-CpG-binding domain (MBD) protein). In some cases, the methylation-binding molecule includes biotin. In some cases, the methylation-binding molecule binds methylated cytosine. In some cases, the solid matrix is beads. In some cases, the solid matrix is a magnetic solid matrix. In some cases, the solid matrix contains protein A. In some cases, the solid matrix contains streptavidin. In some cases, the method further includes denaturing the nucleic acids in the nucleic acid sample mixture after (a) and before (b). In some cases, the method further includes obtaining a nucleic acid sample from the sample and performing one or more library preparation reactions on the nucleic acids before (a). In some cases, exogenous DNA (e.g., spike-in DNA) is mixed with the nucleic acid sample before performing one or more library preparation reactions.Following the library preparation reaction, prior to (a), the prepared library is incubated with multiple DNA capture beads (e.g., SPRI beads) and then removed. In some cases, the method further includes amplifying the methylated nucleic acid after capture to generate a plurality of amplicones of the methylated nucleic acid. In some cases, the captured methylated nucleic acid is subjected to a buffer exchange or washing reaction prior to amplification. In some cases, the captured methylated nucleic acid is eluted prior to amplification. In some cases, the captured methylated nucleic acid is not eluted prior to amplification. In some cases, amplification is performed by PCR. In some cases, PCR amplification includes at least 10 cycles, at least 11 cycles, at least 12 cycles, at least 13 cycles, or at least 14 cycles of amplification. In some cases, PCR amplification is 14 cycles of amplification. In some cases, PCR amplification is 13 cycles of amplification. In some cases, the generated amplicones of the methylated nucleic acid are subjected to a sequencing reaction. In some cases, the amplicones are cleaned prior to sequencing. In some cases, the amplicones are not cleaned prior to sequencing. In some cases, the sequencing reaction is performed via a synthesis reaction. In some cases, the sequencing reaction does not include bisulfite sequencing. In some cases, the method or system further includes contacting multiple methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences. In some cases, the one or more target sequences comprise one or more genes.
[0104] In one aspect, a method or system for processing a nucleic acid sample from an object is disclosed herein, the method or system comprising: (a) generating a nucleic acid sample mixture comprising multiple methylated nucleic acids from the object and a certain amount of supplementally treated DNA (e.g., filler DNA); and (b) incubating (i) a methylation-binding molecule with (ii) a solid matrix to form a methylated nucleic acid capture reagent; and (c) enriching the multiple methylated nucleic acids in the nucleic acid mixture by capturing the methylated nucleic acids by adding the methylated nucleic acid capture reagent to the nucleic acid sample mixture. In some cases, a certain amount of supplementally treated DNA (e.g., filler DNA) is not required in (a). In some cases, the certain amount of supplementally treated DNA (e.g., filler DNA) comprises at least one methylated DNA molecule. In some cases, the methylation-binding molecule is an antibody (e.g., anti-5-methylcytosine (anti-5mC) antibody, methyl-CpG-binding domain (MBD) protein). In some cases, the methylation-binding molecule includes biotin. In some cases, the methylation-binding molecule binds methylated cytosine. In some cases, the solid matrix is a bead. In some cases, the solid matrix is a magnetic solid matrix. In some cases, the solid matrix contains protein A. In some cases, the solid matrix contains streptavidin. In some cases, the formula...The method also includes denaturing the nucleic acids in the nucleic acid sample mixture prior to (b). In some cases, the method further includes obtaining a nucleic acid sample from the sample and performing one or more library preparation reactions on the nucleic acid prior to (a). In some cases, exogenous DNA (e.g., spike-in DNA) is mixed with the nucleic acid sample prior to performing one or more library preparation reactions. In some cases, after performing one or more library preparation reactions, prior to (a), the prepared library is incubated with a plurality of magnetic beads (e.g., SPRI beads) that interact with the nucleic acid and eluted from the magnetic beads that interact with the nucleic acid. In some cases, the sample is magnetically captured to remove the magnetic beads that interact with the nucleic acid. In some cases, the sample is further magnetically captured to remove any remaining magnetic beads that interact with the nucleic acid. In some cases, the method further includes amplifying the methylated nucleic acid after capture to generate amplicones of a plurality of methylated nucleic acids. In some cases, amplification is performed simultaneously with the binding of the methylated nucleic acid to a methylated nucleic acid capture reagent. In some cases, the captured methylated nucleic acid is subjected to a buffer exchange or washing reaction prior to amplification. In some cases, the captured methylated nucleic acids are eluted before amplification. In some cases, the captured methylated nucleic acids are not eluted before amplification (e.g., when amplification occurs simultaneously with the binding of the methylated nucleic acid to the methylated nucleic acid capture reagent). In some cases, amplification is performed by PCR. In some cases, PCR amplification includes at least 10 cycles, at least 11 cycles, at least 12 cycles, at least 13 cycles, or at least 14 cycles of amplification. In some cases, PCR amplification is performed for 14 cycles. In some cases, PCR amplification is performed for 13 cycles. In some cases, the amplicon of the generated methylated nucleic acid is sequenced. In some specifications (page 24 / 55, CN 121464223 A), the amplicon is cleaned before sequencing. In some cases, the amplicon is not cleaned before sequencing. In some cases, the sequencing reaction is performed by a synthesis reaction. In some cases, the sequencing reaction does not include bisulfite sequencing. In some cases, the method or system further includes contacting a plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences. In some cases, the one or more target sequences comprise one or more genes.
[0105] In one aspect, a method or system for processing a nucleic acid sample from an object is disclosed herein, the method or system comprising: (a) providing a nucleic acid sample comprising a plurality of methylated nucleic acids; (b) incubating (i) a methylation-binding molecule with (ii) a solid matrix to form a methylated nucleic acid capture reagent; (c) adding the methylated nucleic acid capture reagent to the sample.(d) Capturing the methylated nucleic acid in the nucleic acid sample to generate a methylated nucleic acid bound to a solid matrix; and (b) amplifying the methylated nucleic acid bound to the solid matrix to generate an amplicon of the methylated nucleic acid. In some cases, amplification is performed simultaneously with the binding of the methylated nucleic acid to the methylated nucleic acid capture reagent. In some cases, the captured methylated nucleic acid is not eluted prior to amplification. In some cases, the methylation-binding molecule is an antibody (e.g., anti-5-methylcytosine (anti-5mC) antibody, methyl-CpG-binding domain (MBD) protein). In some cases, the methylation-binding molecule includes biotin. In some cases, the methylation-binding molecule binds to methylated cytosine. In some cases, the solid matrix is a bead. In some cases, the solid matrix is a magnetic solid matrix. In some cases, the solid matrix contains protein A. In some cases, the solid matrix contains streptavidin. In some cases, the method further includes denaturing the nucleic acids in the nucleic acid sample mixture after (a) and before (b). In some cases, the method further includes obtaining a nucleic acid sample from the sample prior to (a) and performing one or more library preparation reactions on the nucleic acid. In some cases, exogenous DNA (e.g., spike-in DNA) is mixed with the nucleic acid sample prior to performing one or more library preparation reactions. In some cases, after performing one or more library preparation reactions, prior to (a), the prepared library is incubated with a plurality of magnetic beads (e.g., SPRI beads) that interact with the nucleic acid and eluted from the magnetic beads that interact with the nucleic acid. In some cases, the sample is magnetically captured to remove the magnetic beads that interact with the nucleic acid. In some cases, the sample is further magnetically captured to remove any remaining magnetic beads that interact with the nucleic acid. In some cases, the captured methylated nucleic acid is not eluted prior to amplification. In some cases, amplification is performed by PCR. In some cases, PCR amplification includes at least 10 cycles, at least 11 cycles, at least 12 cycles, at least 13 cycles, or at least 14 cycles of amplification. In some cases, PCR amplification is 14 cycles of amplification. In some cases, PCR amplification is a 13-cycle amplification. In some cases, sequencing reactions are performed on the amplicones of the generated methylated nucleic acids. In some cases, the amplicones are cleaned before sequencing. In some cases, the amplicones are not cleaned before sequencing. In some cases, sequencing is performed via a synthesis reaction. In some cases, the sequencing reaction does not include bisulfite sequencing. In some cases, the method or system further includes contacting multiple methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences. In some cases, one or more target sequences contain one or more genes.
[0106] In one aspect, this document discloses a method or system for processing nucleic acid samples from an object, the method or system comprising: (a) generating a nucleic acid sample mixture comprising a plurality of methylated nucleic acids from said object and a certain amount of supplementally treated DNA (e.g., filler DNA), wherein said filler DNA comprises at least one methylated DNA molecule; (b) capturing said methylated nucleic acids by adding a capture reagent comprising a solid matrix to said nucleic acid sample mixture, thereby generating methylated nucleic acids bound to the solid matrix; and (c) amplifying the methylated nucleic acids bound to the solid matrix to generate an amplicon of the methylated nucleic acid. In some cases, the certain amount of supplementally treated DNA (e.g., filler DNA) comprises at least one methylated DNA molecule. In some cases, amplification is performed simultaneously with the binding of the methylated nucleic acids to the methylated nucleic acid capture reagent. In some cases, the captured methylated nucleic acids are not eluted prior to amplification. In some cases, the methylated binding molecule is an antibody (e.g., an anti-5-methylcytosine (anti-5mC) antibody, a methyl-CpG-binding domain (MBD) protein). In some cases, the methylation-binding molecule includes biotin. In some cases, the methylation-binding molecule binds to methylated cytosine. In some cases, the solid matrix is beads. In some cases, the solid matrix is a magnetic solid matrix. In some cases, the solid matrix contains protein A. In some cases, the solid matrix contains streptavidin. In some cases, the method further includes denaturing the nucleic acids in the nucleic acid sample mixture after (a) and before (b). In some cases, the method further includes obtaining a nucleic acid sample from the sample before (a) and performing one or more library preparation reactions on the nucleic acids. In some cases, exogenous DNA (e.g., spike-in DNA) is mixed with the nucleic acid sample before performing one or more library preparation reactions. In some cases, after performing one or more library preparation reactions, before (a), the prepared library is incubated with a plurality of magnetic beads (e.g., SPRI beads) that interact with the nucleic acids and eluted from the magnetic beads that interact with the nucleic acids. In some cases, the sample is magnetically trapped to remove the magnetic beads that interact with the nucleic acids. In some cases, the sample undergoes additional magnetic trapping to remove residual magnetic beads that have interacted with nucleic acids. In some cases, the trapped methylated nucleic acids are not eluted before amplification. In some cases, amplification is performed by PCR. In some cases, PCR amplification includes at least 10 cycles, at least 11 cycles, at least 12 cycles, at least 13 cycles, or at least 14 cycles of amplification. In some cases, PCR amplification is 14 cycles of amplification. In some cases, PCR amplification is...Amplification in 13 cycles. In some cases, sequencing reactions are performed on the amplicones of the generated methylated nucleic acids. In some cases, the amplicones are cleaned before sequencing. In some cases, the amplicones are not cleaned before sequencing. In some cases, the sequencing reaction is performed via a synthesis reaction. In some cases, the sequencing reaction does not include bisulfite sequencing. In some cases, the method or system further includes contacting a plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences. In some cases, the one or more target sequences contain one or more genes.
[0107] In one aspect, a method or system is disclosed herein, comprising: (a) obtaining a first nucleic acid molecule from a cell-free sample of an object; (b) generating a second set of nucleic acid molecules from the first set of nucleic acid molecules or derivatives thereof, wherein the second set of nucleic acid molecules is enriched at the methylation level relative to the first set of nucleic acid molecules; (c) enriching one or more targets in the second set of nucleic acid molecules or derivatives thereof to obtain a third set of nucleic acid molecules; and (d) sequencing the third set of nucleic acid molecules or derivatives thereof. In some cases, enrichment includes contacting the second set of nucleic acid molecules or derivatives thereof with one or more nucleic acid capture probes. In some cases, generation involves contacting a first group of nucleic acid molecules or derivatives thereof with a methylated nucleic acid capture reagent. In some cases, the methylated nucleic acid capture reagent is formed by incubating methylation-binding molecules together with (ii) a solid matrix. In some cases, the solid matrix is a bead. In some cases, the solid matrix is a magnetic solid matrix. In some cases, the solid matrix contains protein A. In some cases, the solid matrix contains streptavidin. In some cases, the methylation-binding molecule is an antibody (e.g., anti-5-methylcytosine (anti-5mC) antibody, methyl-CpG-binding domain (MBD) protein). In some cases, the methylation-binding molecule includes biotin. In some cases, the methylation-binding molecule binds to methylated cytosine. In some cases, the method or system further includes amplifying a second group of molecules. In some cases, amplification is performed simultaneously with the binding of a subgroup of the first group of nucleic acids to the methylated nucleic acid capture reagent. In some cases, amplification is performed by PCR amplification. In some cases, PCR amplification includes at least 10, 11, 12, 13, or 14 cycles of amplification. In some cases, PCR amplification is 14 cycles. In some cases, PCR amplification is 13 cycles. In some cases, the amplicon generated from the amplification can be cleaned up. In some cases, sequencing is performed via a synthesis reaction. In some cases, sequencing does not include bisulfite sequencing. In some cases, a third set of nucleic acid molecules is subjected to a pre-sequencing reaction.Or multiple library preparation reactions. In some cases, the method or system further includes incubating a third set of nucleic acid molecules with a plurality of magnetic beads that interact with the nucleic acid after performing one or more library preparation reactions and before sequencing. In some cases, the method or system further includes magnetically trapping the nucleic acid sample to remove the plurality of beads that interact with the nucleic acid after incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acid. In some cases, the method or system further includes performing additional magnetic trapping. In some cases, the method or system further includes adding a certain amount of filler DNA to the first nucleic acid molecule before (b).
[0108] In any method or system for processing nucleic acids described herein, quality control analysis can be performed on the captured methylated nucleic acids. For example, quality control analysis may include measuring methylation binding specificity. In some cases, methylation binding specificity is measured by calculating the recovery rate of exogenous methylated fragments. In some cases, methylation binding specificity is measured by calculating the events of detected methylated fragments. In some cases, the methods and systems described herein for processing nucleic acid samples from subjects result in methylation-specific enrichment of methylated single-stranded DNA with at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, or at least about 99.9%. In some cases, the methods and systems described herein for processing nucleic acid samples from subjects result in enrichment of methylated nucleic acids by at least about 1, at least about 2, at least about 3, at least about 4, at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 55, at least about 60, at least about 65, at least about 70, at least about 75, at least about 80, at least about 85, at least about 90, at least about 95, at least about 100, at least about 150, or at least about 200.
[0109] Methylation Spectroscopy This disclosure provides methods and systems for generating methylation spectra of subjects who have or are suspected of having a disease / condition, wherein the methylation spectra can be used to determine whether the subject has the disease / condition or is in a state of being affected by it.The risk of this disease / condition is present. In some cases, methylation profiling can be used to determine whether a subject has the disease / condition (e.g., cancer) or whether the disease / condition (e.g., cancer) has recurred. In some cases, methylation profiling can be used to determine whether a subject has the disease / condition (e.g., cancer) or whether the disease / condition (e.g., cancer) will not recur. In some cases, methylation profiling may include the analysis of multiple nucleic acids (e.g., multiple nucleic acid molecules of a depleted sequencing library as described herein) (e.g., including sequencing). In some cases, methylation profiling may include the detection of methylated nucleotides and / or the quantification of methylated nucleotide counts. In some cases, methylation profiling may include the quantification of circulating tumor DNA (ctDNA). In some cases, ctDNA may be quantified over time (e.g., ctDNA kinetics). In some cases, methylation profiling may include the determination of methylation signals, for example, in a population of nucleic acids in a depleted sequencing library as described herein. In some cases, methylation profiling is compared to a genome-wide background profile. In some cases, methylation profiling is compared to a novel background profile created using highly methylated cfDNA.
[0110] In some cases, methylation profiles can be analyzed to distinguish cancer cases from controls (e.g., non-cancer controls), wherein the area under the receiver operating characteristic curve (AUROC) is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, or at least about 99.9%. In some cases, methylation profiles can be analyzed to distinguish cancer cases from controls (e.g., non-cancer controls), where AUROC is up to about 90%, up to about 91%, up to about 92%, up to about 93%, up to about 94%, up to about 95%, up to about 96%, up to about 97%, up to about 98%, up to about 99%, up to about 99.1%, up to about 99.2%, up to about 99.3%, up to about 99.4%, up to about 99.5%, up to about 99.6%, up to about 99.7%, up to about 99.8%, or up to about 99.9% (see page 27 / 55 of the specification, CN 121464223 A). In some cases, methylation profiles can be analyzed to distinguish cancer cases from controls (e.g., non-cancer controls), where AUROC values are approximately 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, and 99.1%.Approximately 99.2%, approximately 99.3%, approximately 99.4%, approximately 99.5%, approximately 99.6%, approximately 99.7%, approximately 99.8%, or approximately 99.9%.
[0111] The Genome Mutation Profiles disclosure provides methods, systems, and kits for generating mutation profiles of subjects who have or are suspected of having a disease / condition, wherein the mutation profile can be used to determine whether a subject has or is at risk of having the disease / condition. The samples disclosed in this paper can be used for library preparation and next-generation deep sequencing, for example, to depths of 1 million (M) to 60 M single reads, 10 M to 60 M single reads, 10 M to 100 M single reads, 40 M to 60 M single reads, 40 M to 100 M single reads, 60 M to 100 M single reads, 60 M to 200 M single reads, 1 M to 10 M single reads, 1 M to 40 M single reads, 1 M single read to 100 M single reads, 1 M single read to 200 M single reads, at least 1 M single read, at least 10 M single read, at least 40 M single read, at least 60 M single read, at least 100 M single read, or at least 200 M single read. In some cases, sequencing can be performed at lower sequencing depths (e.g., 10 M single reads, 20 M single reads, 30 M single reads, 40 M single reads, 1 M single read to 10 M single reads, 10 M single read to 20 M single reads, 20 M single read to 30 M single reads, 30 M single read to 40 M single reads, up to 10 M single reads, up to 20 M single reads, up to 30 M single reads, or up to 40 M single reads). In some cases, the samples disclosed herein may be in the following sizes: 0.1X to 100X, 0.1X to 60X, 0.1X to 40X, 0.1X to 30X, 0.1X to 20X, 0.1X to 10X, 0.1X to 5.0X, 0.5X to 100X, 0.5X to 60X, 0.5X to 40X, 0.5X to 30X, 0.5X to 20X, 0.5X to 10X, 0.5X to 5.0X, 1.0X to 100X, 1.0X to 60X, 1.0X to 40X, 1.0X to 30X, 1.0X. Up to 20X, 1.0X to 10X, 1.0X to 5.0X, at least 0.1X, at least 0.5X, at least 1.0X, at least 2.0X, at least 3.0X, at least 4.0X, at least 5.0X, at least 10.0X, at least 20.0X, at least 30.0X, at least 40.0X, at least 50.0X, at least 60.0X, at least 100X, at least 200X, at most 0.1X, at most 0.5X, at most 1.0X, at most 2.0X, at most 3.0X, at most 4.0X, at mostSequencing is performed once at depths of 5.0X, up to 10.0X, up to 20.0X, up to 30.0X, up to 40.0X, up to 50.0X, up to 60.0X, up to 100X, or up to 200X. Multiple sequencing reads are generated and analyzed. In some embodiments, deep sequencing can be configured to maximize the identification of genomic mutations associated with the disease / condition.
[0112] In some embodiments, the relative measurement of ctDNA abundance is calculated by the mean mutant allele fraction (MAF). In some embodiments, the mean MAF of mutations identified in the subject and included in his / her mutation profile is in the range of at least about 0.01% to at least about 10%. In some cases, the MAF of the ctDNA fraction of a sample can be at least about 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.15%, 0.2%, 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, 10%, or any percentage therebetween.
[0113] In some embodiments, a generated mutation profile of the object can be generated from the sequencing results. In some embodiments, the mutation profile includes genetic polymorphisms such as missense variants, nonsense variants, deletion variants, insertion variants, duplication variants, inversion variants, frameshift variants, or amplified repetitive sequences. In some embodiments, the mutation profile can include mutations derived from a fraction of cell-free nucleic acid molecules of a specific size range. This disclosure provides methods, systems, and kits for generating mutation profiles of subjects who have or are suspected of having a disease / condition, wherein methylation profiles can be used to determine whether the subject has the disease / condition or is at risk of having it. Generating a genomic mutation profile may include library preparation of multiple nucleic acid molecules and next-generation deep sequencing (e.g., MeDIP-seq). Multiple sequencing reads can be generated and analyzed, and in some cases, deep sequencing can be configured to maximize the identification of genomic mutations associated with the disease / condition. For example, a set of classic cancer driver genes may be included in selectors used for sequencing results analysis. In some embodiments, including genes that have not been shown to have a driver effect in a particular cancer type in the analysis of sequencing data can increase the sensitivity of ctDNA detection.
[0114] In some embodiments, the relative measurement of ctDNA abundance is calculated by the mean mutant allele fraction (MAF). In some embodiments, the mean MAF of mutations that identify the subject and are included in his / her mutation profile is at least approximatelyThe ctDNA fraction of the samples disclosed herein is in the range of at least about 0.01%, 0.02%, 0.03%, 0.04%, 0.05%, 0.06%, 0.07%, 0.08%, 0.09%, 0.1%, 0.15%, 0.2%, 0.5%, 1%, 1.5%, 2%, 2.5%, 3%, 3.5%, 4%, 4.5%, 5%, 5.5%, 6%, 6.5%, 7%, 7.5%, 8%, 8.5%, 9%, 9.5%, 10%, or any percentage therebetween.
[0115] In some embodiments, the mutation profile generated for the object does not include mutant variants derived from cell-free nucleic acid molecules derived from biological samples. In some embodiments, the mutation profile includes genetic polymorphisms such as missense variants, nonsense variants, deletion variants, insertion variants, duplication variants, inversion variants, frameshift variants, or amplified repetitive sequences. In some embodiments, the mutation profile may include mutant variants derived from fractions of cell-free nucleic acid molecules within a specific size range.
[0116] In some embodiments, the length of the ctDNA fragment is shorter than that of a cell-free nucleic acid molecule derived from a healthy subject. In some embodiments, the length of the ctDNA containing at least one mutation is shorter than that of a cell-free nucleic acid molecule containing the corresponding reference allele.
[0117] In some embodiments, sequencing does not utilize bisulfite sequences because it causes degradation of the ctDNA fragments and prevents the preservation of the ctDNA length distribution. In some embodiments, the fragment lengths of the multiple nucleic acids of this disclosure (e.g., cfDNA molecules containing a mixture of tumor or cancer tissue and healthy tissue, cfDNA molecules containing only healthy tissue, and / or ctDNA only) can be 1 to about 800 base pairs, about 50 bp to about 800 bp, about 100 bp to about 200 bp, about 120 bp to about 150 bp, about 60 to about 500 bp, about 80 to about 300 bp, 90 to about 250 bp, 80 to 170 bp, or about 100 to about 150 bp. In some embodiments, the fragment lengths of the plurality of nucleic acids disclosed herein (e.g., comprising cfDNA molecules derived from a mixture of tumor or cancer tissue and healthy tissue, comprising only cfDNA molecules derived from healthy tissue, and / or comprising only ctDNA) may be at least 800 base pairs (bp), at least 700 base pairs, at least 600 base pairs, at least 500 base pairs, at least 400 base pairs, at least 300 base pairs, at least 200 base pairs, at least 150 base pairs, at least 100 base pairs, or at least 50 base pairs. In some embodiments, the fragment lengths of the plurality of nucleic acids disclosed herein (e.g., comprising cfDNA molecules derived from a mixture of tumor or cancer tissue and healthy tissue, comprising only cfDNA molecules derived from healthy tissue, and / or comprising only ctDNA) may be at least 800 base pairs (bp), at least 700 base pairs, at least 600 base pairs, at least 500 base pairs, at least 400 base pairs, at least 300 base pairs, at least 200 base pairs, at least 150 base pairs, at least 100 base pairs, or at least 50 base pairs) may be at least 800 base pairs (bp), at least 700 base pairs, at least 600 base pairs, at least 500 base pairs, at least 100 base pairs, or at least 50 base pairs.The fragment length (containing only ctDNA) can be up to 800 base pairs (bp), up to 700 base pairs, up to 600 base pairs, up to 500 base pairs, up to 400 base pairs, up to 300 base pairs, up to 200 base pairs, up to 150 base pairs, up to 100 base pairs, or up to 50 base pairs. In some embodiments, this disclosure provides enrichment of cell-free nucleic acid samples based on the selection of cell-free molecules of a certain size. In some embodiments, multimodal analysis includes utilizing the mutation profiles and fragment length profiles described herein by selectively including multiple nucleic acid molecules in a mutation profile based on their fragment length. In some embodiments, multimodal analysis includes utilizing the methylation profiles and fragment length profiles described herein by selectively including multiple nucleic acid molecules in a methylation profile based on their fragment length. In some embodiments, multimodal analysis includes utilizing mutation profiles, methylation profiles, and fragment length profiles together by selectively including multiple nucleic acid molecules in a mutation profile based on their fragment length and by selectively including multiple nucleic acid molecules in a methylation profile based on their fragment length, respectively.
[0118] Tumor Detection and Prognosis This disclosure provides methods and systems for determining whether a subject has a disease or is at risk of having a disease, wherein these methods and systems include sequencing multiple nucleic acid molecules derived from a cell-free nucleic acid sample obtained from a subject to generate at least one of (i) a methylation profile, (ii) a mutation profile, and (iii) a fragment length profile; and processing said at least one profile to determine whether the subject has said disease or is at risk of saying disease with at least 80% sensitivity or at least about 90% specificity, wherein said cell-free nucleic acid sample contains less than 30 ng / ml of said multiple nucleic acid molecules. In some embodiments, the sensitivity is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these figures. In some embodiments, the specificity is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%.At least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these figures.
[0119] In some embodiments, these methods and systems may include sequencing multiple nucleic acid molecules derived from cell-free nucleic acid samples obtained from said object to generate at least two of (i) methylation profiles, (ii) mutation profiles, and (iii) fragment length profiles. These methods provide a sensitivity of at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these figures. In some embodiments, the sensitivity when using two spectra is increased by at least about 0.5%, at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, or any percentage between these figures, compared to the sensitivity when using one spectrum. In some implementations, the sensitivity when using three spectra is increased by at least about 0.5%, at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, or any percentage between these numbers, compared to the sensitivity when using two spectra.
[0120] Furthermore, these methods can provide specificity of at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these numbers. In some embodiments, the specificity when using two spectra is increased by at least about 0.5%, at least about 1%, at least about 2%, at least about 3%, at least about 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, or any percentage between these figures, compared to the specificity when using one spectrum. In some embodiments, the specificity when using three spectra is increased by at least about 0.5%, at least about 1%, at least about 2%, at least about 3%, or any percentage between these figures, compared to the specificity when using two spectra.Percentages of at least 4%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, at least about 10%, or any of these figures.
[0121] This disclosure provides methods and systems for processing cell-free nucleic acid samples of subjects to determine whether the subjects have a disease or are at risk of having a disease, the methods and systems comprising providing the cell-free nucleic acid sample comprising a plurality of nucleic acid molecules; sequencing the plurality of nucleic acid molecules or derivatives thereof to generate a plurality of sequencing reads; computer processing the plurality of sequencing reads to identify (i) methylation profiles, (ii) mutation profiles, and (iii) fragment length profiles of the plurality of nucleic acid molecules; and using at least the methylation profiles, the mutation profiles, and the fragment length profiles to determine whether the subject has the disease or is at risk of having the disease. In some implementations, these methods provide a sensitivity of at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, to page 30 / 55 of the specification, 41 CN 121464223 A, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these figures. These methods provide specificity of at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these figures.
[0122] This disclosure provides methods and systems for determining the tissue origin of a tumor, including identifying a nucleotide sequence specific to a particular cancer (e.g., breast cancer, colon cancer, prostate cancer, HSNCC, or lung cancer), wherein a fraction of cell-free nucleic acid molecules is derived from that nucleotide sequence. In some embodiments, the fraction of cell-free nucleic acid molecules is derived from ctDNA. In some implementations, these methods provide at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, and at least about 90%.Sensitivity of at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these numbers. These methods provide specificity of at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these numbers.
[0123] This disclosure provides methods and systems for determining whether a subject has multiple diseases (e.g., multiple cancers) or is at risk of having multiple diseases (e.g., multiple cancers), wherein these methods and systems include sequencing a plurality of nucleic acid molecules derived from a cell-free nucleic acid sample obtained from the subject to generate at least one of (i) a methylation profile, (ii) a mutation profile, and (iii) a fragment length profile; and processing the at least one profile to determine whether the subject has the disease or is at risk of the disease with at least 80% sensitivity or at least about 90% specificity, wherein the cell-free nucleic acid sample contains less than 30 ng / ml of the plurality of nucleic acid molecules. In some implementations, the sensitivity is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these figures. In some implementations, the specificity is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, at least about 99.9%, or any percentage between these figures. In some cases, it is determined whether the subject has multiple diseases (e.g., multiple cancers) or is in a state of...Methods for determining the risk of having multiple diseases (e.g., multiple cancers) may include identifying two, three, four, five, or more diseases. In some cases, when determining whether a subject has multiple diseases (e.g., multiple cancers) or is at risk of having multiple diseases (e.g., multiple cancers), it may be possible to identify no disease, one disease (e.g., cancer), or more than one disease (e.g., multiple cancers).
[0124] This disclosure provides methods and systems for determining whether a subject has cancer (e.g., low-shedding cancer) or is at risk of cancer (e.g., low-shedding cancer), wherein these methods and systems include (a) providing a plurality of nucleic acid molecules generated from a cfDNA sample of the subject, and (b) sequencing the plurality of nucleic acid molecules or derivatives thereof to generate a plurality of sequencing reads. (Specification 31 / 55 pages 42 CN 121464223 A) (c) computer processing of the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules, and (d) computer processing of the methylation profiles to determine that the subject has cancer when the area under the receiver operating characteristic curve (AUROC) is at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 99.1%, at least about 99.2%, at least about 99.3%, at least about 99.4%, at least about 99.5%, at least about 99.6%, at least about 99.7%, at least about 99.8%, or at least about 99.9%. In some cases, AUROC can be up to approximately 90%, up to approximately 91%, up to approximately 92%, up to approximately 93%, up to approximately 94%, up to approximately 95%, up to approximately 96%, up to approximately 97%, up to approximately 98%, up to approximately 99%, up to approximately 99.1%, up to approximately 99.2%, up to approximately 99.3%, up to approximately 99.4%, up to approximately 99.5%, up to approximately 99.6%, up to approximately 99.7%, up to approximately 99.8%, or up to approximately 99.9%. In some cases, AUROC can be about 90%, about 91%, about 92%, about 93%, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, about 99.1%, about 99.2%, about 99.3%, about 99.4%, about 99.5%, about 99.6%, about 99.7%, about 99.8%, or about 99.9%.
[0125] This disclosure provides methods and systems for determining whether a subject has cancer (e.g., endometrial cancer, esophageal cancer, hepatobiliary cancer, prostate cancer, bladder cancer, breast cancer, colorectal cancer, head and neck cancer, lung cancer, pancreatic cancer, or kidney cancer), wherein these methods and systems include (a) providing a plurality of nucleic acid molecules generated from a cell-free deoxyribonucleic acid (cfDNA) sample of the subject, (b)Without bisulfite conversion, the plurality of nucleic acid molecules or their derivatives are sequenced to generate a plurality of sequencing reads; (c) the plurality of sequencing reads are processed by computer to generate methylation profiles of the plurality of nucleic acid molecules; and (d) the methylation profiles are processed by computer to determine that the subject has cancer.
[0126] This disclosure provides methods and systems for determining whether a subject will experience cancer recurrence, wherein these methods and systems include (a) providing a plurality of nucleic acid molecules generated from a cell-free deoxynucleic acid (cfDNA) sample of the subject; (b) sequencing the plurality of nucleic acid molecules or their derivatives without bisulfite conversion to generate a plurality of sequencing reads; (c) processing the plurality of sequencing reads by computer to generate methylation profiles of the plurality of nucleic acid molecules; and (d) processing the methylation profiles by computer to determine that the subject has cancer.
[0127] This disclosure provides methods and systems for determining whether a subject has one or more diseases or is at risk of having one or more diseases, wherein these methods and systems include sequencing a plurality of nucleic acid molecules derived from a cell-free nucleic acid sample obtained from said subject to generate a methylation profile, and comparing the sequence of captured cell-free methylated DNA with control cell-free methylated DNA sequences from healthy individuals and cancer patients; detecting the subject's disease by determining statistically significant similarity between one or more sequences of captured cell-free methylated DNA and cell-free methylated DNA sequences from cancer individuals.
[0128] In some embodiments, control cell-free methylated DNA sequences from healthy individuals and cancer individuals are included in a database of differentially methylated regions (DMRs) between healthy individuals and cancer individuals.
[0129] This disclosure describes methods and systems for providing prognoses to subjects after receiving treatment for a disease / condition. For example, treatments include surgical resection of a tumor, chemotherapy designed for a specific type of cancer, radiotherapy, or immunotherapy (e.g., TCR, CAR, etc.). In some embodiments, these methods or systems include sequencing multiple nucleic acid molecules derived from a cell-free nucleic acid sample obtained from said subject to generate at least one of (i) a methylation profile, (ii) a mutation profile, and (iii) a fragment length profile; and monitoring or detecting minimal residual disease (MRD) based on at least one profile.
[0130] Once the subject is accurately diagnosed and receives treatment for cancer, such as surgical resection, chemotherapy, radiotherapy, etc., monitoring the effectiveness of the treatment and predicting patient survival may be important. Further, detecting minimal residual disease in cancer cells may be important.
[0131] In some embodiments, the method further includes adding a second amount of control DNA to the sample to confirm the immunoassay. (Page 32 / 55 of the instruction manual)43 CN 121464223 A Procedure for precipitation reaction.
[0132] As used herein, “control” may include both positive and negative controls, or at least a positive control.
[0133] In some embodiments, the method further includes adding a second amount of control DNA to the sample to confirm the capture of cell-free methylated DNA.
[0134] In some embodiments, identifying the presence of DNA from cancer cells also includes identifying tissue of cancer cell origin.
[0135] In some cases, tumor tissue sampling may be challenging or risky, in which case it may be desirable to diagnose and / or subtype cancer without requiring tumor tissue sampling. For example, lung tumor tissue sampling may require invasive procedures such as mediastinoscopy, open-chest surgery, or percutaneous biopsy; these procedures may result in the need for hospitalization, chest tubes, mechanical ventilation, antibiotics, or other medical interventions. Some individuals may not undergo the invasive procedures required for tumor tissue sampling due to medical comorbidities or preferences. In some cases, the actual procedure for obtaining tumor tissue may depend on the suspected cancer subtype. In other cases, cancer subtypes can evolve within the same body over time; a series of assessments using invasive tumor tissue sampling procedures are often impractical and poorly tolerated by patients. Therefore, non-invasive cancer typing via blood tests has many advantageous applications in clinical oncology practice.
[0136] Thus, in some embodiments, identifying the tissue of origin of cancer cells also includes identifying the cancer subtype. In some cases, cancer subtypes can be distinguished based on staging (e.g., early-stage cancer treated surgically versus late-stage cancer treated with chemotherapy), histology (e.g., small cell carcinoma versus adenocarcinoma versus squamous cell carcinoma in lung cancer), gene expression patterns or transcription factor activity (e.g., ER status in breast cancer), copy number abnormalities (e.g., HER2 status in breast cancer), specific rearrangements (e.g., FLT3 in AML), specific gene point mutation status (e.g., IDH gene point mutation), and DNA methylation patterns (e.g., MGMT gene promoter methylation in brain cancer).
[0137] In some embodiments, comparisons can be based on fitting using a statistical classifier. In some cases, statistical classifiers using DNA methylation data can be used to assign samples to multiple cancers for multi-cancer detection. In some cases, multi-cancer detection can be multi-cancer early detection (MCED). In some cases, statistical classifiers using DNA methylation data can be used to assign samples to minimal residual disease (MRD). In some cases, statistical classifiers using DNA methylation data can be used to assign samples to specific disease states, such as cancer type or subtype (e.g., early-stage cancer, low-shedding tumors).(tumor). In some cases, when a classifier can distinguish multiple cancer types (or subtypes) from one another, the classifier may have differentially methylated regions from pairwise comparisons of each cancer type (or subtype) of interest. In some cases, a statistical classifier may have one or more DNA methylation variables in a statistical model, and / or the output of the statistical model may have one or more thresholds to distinguish different disease states. In some cases, a statistical classifier may have features and / or thresholds that may be derived from prior knowledge of cancer types or subtypes, prior knowledge of the most informative features, machine learning, or a combination of two of these methods. In some embodiments, the classifier may be machine learning-derived. In some cases, the classifier may be an elastic net classifier, a lasso, a support vector machine, a random forest, or a neural network.
[0138] In some embodiments, comparisons may be made across the entire genome. In other embodiments, comparisons may be restricted from the whole genome to specific regulatory regions, such as, but not limited to, long scattered nuclear elements (LINEs), short scattered nuclear elements (SINEs), long terminal repeats (LTRs), FANTOM5 enhancers, CpG islands, CpG shores, CpG skeletons, or any combination thereof.
[0139] In some embodiments, the methods herein are used for cancer detection. In some embodiments, the methods herein are used for early cancer detection (e.g., early-stage cancer). In some embodiments, the methods herein are used for multi-cancer early detection (e.g., early-stage cancer). In some embodiments, the methods herein are used for low-shedding tumor detection. In some embodiments, the methods herein are used for MRD detection. In some embodiments, the methods herein are used for low ctDNA detection.
[0140] In some embodiments, the methods herein are used for monitoring cancer treatment.
[0141] Data Analysis Systems and Methods The methods and systems disclosed herein may include algorithms or their uses. One or more algorithms may be used to classify one or more samples from one or more objects. One or more algorithms can be used to quantify ctDNA. One or more algorithms can be used to predict treatment response. One or more algorithms can be used to predict disease / condition (e.g., cancer). One or more algorithms can be applied to data from one or more samples. This data may include biomarker expression data. In some embodiments, these methods or systems include sequencing multiple nucleic acid molecules derived from a cell-free nucleic acid sample obtained from said object to generate at least one of (i) methylation profile, (ii) mutation profile, and (iii) fragment length profile.Spectra; and monitoring or detecting minimal residual disease (MRD) based on at least one spectrum. The methods disclosed herein may include assigning classifications to one or more samples from one or more objects. Assigning classifications to samples may include applying algorithms to methylation spectra, mutation spectra, and fragment length spectra. In some cases, at least one spectrum is input to a data analysis system that includes a trained algorithm for classifying samples obtained from objects with disease or minor injury.
[0142] The data analysis system may be a trained algorithm. The algorithm may include a linear classifier. In some cases, the linear classifier includes one or more of linear discriminant analysis, Fisher's linear discriminant analysis, Naive Bayes classifier, logistic regression, perceptron, support vector machine, or a combination thereof. The linear classifier may be a support vector machine (SVM) algorithm. The algorithm may include a bidirectional classifier. The bidirectional classifier may include one or more decision trees, random forests, Bayesian networks, support vector machines, neural networks, or logistic regression algorithms.
[0143] The algorithm may include one or more linear discriminant analysis (LDA), basic perceptron, elastic net, logistic regression, (kernel) support vector machine (SVM), diagonal linear discriminant analysis (DLDA), Golub classifier, Parzen-based, (kernel) Fisher discriminant classifier, k-nearest neighbor, iterative RELIEF, classification tree, maximum likelihood classifier, random forest, nearest centroid, microarray predictive analysis (PAM), k-median clustering, fuzzy C-means clustering, Gaussian mixture model, quantitative response (GR), gradient boosting method (GBM), elastic net logistic regression, logistic regression, or combinations thereof. The algorithm may include the diagonal linear discriminant analysis (DLDA) algorithm. The algorithm may include the nearest centroid algorithm. The algorithm may include the random forest algorithm. In some implementations, logistic regression, random forest, and gradient boosting method (GBM) outperform linear discriminant analysis (LDA), neural networks, and support vector machines (SVM) in distinguishing between preeclampsia and non-preeclampsia.
[0144] This disclosure provides methods and systems for determining whether a subject has a disease or is at risk of having a disease, wherein these methods and systems include sequencing a plurality of nucleic acid molecules derived from a cell-free nucleic acid sample obtained from said subject to generate at least one of (i) a methylation profile, (ii) a mutation profile, and (iii) a fragment length profile; and processing said at least one profile to determine whether the subject has said disease or is at risk of said disease with at least 80% sensitivity or at least about 90% specificity, wherein said cell-free nucleic acid sample contains less than 30 ng / ml of said plurality of nucleic acid molecules. In some embodiments, the sensitivity is at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, or 87%.88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or any percentage between these numbers. In some embodiments, the specificity is at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or any percentage between these numbers.
[0145] In some embodiments, these methods and systems may include sequencing multiple nucleic acid molecules derived from a cell-free nucleic acid sample obtained from said object to generate at least two of (i) a methylation profile, (ii) a mutation profile, and (iii) a fragment length profile. These methods provide a sensitivity of at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or any percentage between these numbers. In some embodiments, the sensitivity when using two spectra is increased by at least about 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or any percentage between these numbers compared to the sensitivity when using one spectrum. In some embodiments, the sensitivity when using three spectra is increased by at least about 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or any percentage between these numbers compared to the sensitivity when using two spectra.
[0146] Further, these methods can provide specificity of at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or any percentage between these numbers. In some embodiments, the specificity when using two spectra is increased by at least about 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or any percentage between these numbers compared to the specificity when using one spectrum. In some embodiments, the specificity when using three spectra is increased by at least about 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or any percentage between these numbers compared to the specificity when using two spectra.
[0147] This disclosure provides methods and systems for processing cell-free nucleic acid samples of a subject to determine whether the subject has a disease or is at risk of having a disease. These methods and systems include providing the cell-free nucleic acid sample comprising a plurality of nucleic acid molecules; sequencing the plurality of nucleic acid molecules or derivatives thereof to generate a plurality of sequencing reads; computer processing the plurality of sequencing reads to identify (i) methylation profiles, (ii) mutation profiles, and (iii) fragment length profiles of the plurality of nucleic acid molecules; and using at least the methylation profiles, the mutation profiles, and the fragment length profiles to determine whether the subject has the disease or is at risk of having the disease. In some embodiments, these methods provide a sensitivity of at least about 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or any percentage between these numbers. These methods can provide specificity of at least about 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or any percentage between these numbers.
[0148] This disclosure describes methods and systems for providing prognoses to subjects after treatment for a disease / condition. For example, treatments include surgical resection of a tumor, chemotherapy designed for a specific type of cancer, radiotherapy, or immunotherapy (e.g., TCR, CAR, etc.). In some embodiments, these methods or systems include sequencing a plurality of nucleic acid molecules derived from a cell-free nucleic acid sample obtained from said subject to generate at least one of (i) a methylation profile, (ii) a mutation profile, and (iii) a fragment length profile; and monitoring or detecting minimal residual disease (MRD) based on at least one profile.
[0149] Computer System This disclosure provides computer systems programmed to implement the methods of this disclosure. Figure 2 shows a computer system 201 programmed or otherwise configured to generate sequencing libraries of nucleic acid molecules depleted of hypermethylated regions (e.g., ctDNA). The computer system 201 can be adapted to various aspects of this disclosure. The computer system 201 can be a user's electronic device or a computer system located remotely relative to an electronic device. The electronic device can be a mobile electronic device.
[0150] The computer system 201 includes a central processing unit (CPU, also referred to herein as a "processor" and "computer processor") 205, which can be a single-core or multi-core processor or multiple processors for parallel processing. The computer system 201 also...The system includes a memory or memory location 210 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 215 (e.g., hard disk), a communication interface 220 for communicating with one or more other systems (e.g., a network adapter), and peripheral devices 225 (such as cache, other memory, data storage, and / or electronic display adapters). The memory specification 35 / 55 pages 46 CN 121464223 A 210, storage unit 215, interface 220, and peripheral devices 225 communicate with the CPU 205 via a communication bus (solid line) such as a motherboard. Storage unit 215 may be a data storage unit (or data repository) for storing data. The computer system 201 may be operatively coupled to a computer network (“network”) 230 by means of communication interface 220. Network 230 may be the Internet, the Internet, and / or an extranet, or an intranet and / or extranet communicating with the Internet. In some cases, network 230 is a telecommunications and / or data network. Network 230 may include one or more computer servers that can support distributed computing such as cloud computing. In some cases, network 230 may be implemented as a peer-to-peer network by means of computer system 201, which allows devices coupled to computer system 201 to act as clients or servers.
[0151] CPU 205 can execute machine-readable instruction sequences that may be embodied in a program or software. Instructions may be stored in a memory location such as memory 210. Instructions may be directed to CPU 205, which may then program or otherwise configure CPU 205 to implement the methods of this disclosure. Examples of operations performed by CPU 205 may include fetching, decoding, executing, and writing back.
[0152] CPU 205 may be part of circuitry such as an integrated circuit. One or more other components of system 201 may be included in this circuitry. In some cases, the circuitry is an application-specific integrated circuit (ASIC).
[0153] Storage unit 215 may store files such as drivers, libraries, and saved programs. Storage unit 215 may store user data, such as user preferences and user programs. Computer system 201 may, in some cases, include one or more additional data storage units that are external to computer system 201, such as those located on a remote server communicating with computer system 201 via an intranet or the Internet.
[0154] Computer system 201 may communicate with one or more remote computer systems via network 230. For example, computer system 201 may communicate with a user's remote computer system. Examples of remote computer systems include personal computers (e.g., personal computers).Such as portable PCs, tablets or tablet PCs (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, smartphones (e.g., Apple® iPhone, Android-enabled devices, Blackberry®), or personal digital assistants. Users can access computer system 201 via network 230.
[0155] The methods described herein can be implemented by machine (e.g., computer processor) executable code stored on an electronic storage location of computer system 201 (such as, for example, memory 210 or electronic storage unit 215). The machine executable code or machine readable code can be provided in the form of software. During use, the code can be executed by processor 205. In some cases, the code can be retrieved from storage unit 215 and stored on memory 210 for access by processor 205 at any time. In some cases, electronic storage unit 215 may not be included, and machine executable instructions are stored on memory 210.
[0156] The code can be pre-compiled and configured for use by a machine having a processor suitable for executing the code, or it can be compiled during runtime. The code may be provided in a programming language, which may be selected to enable the code to be executed in a pre-compiled or just-in-time (JIT) manner.
[0157] In some embodiments, the computer system may include computer processing performed by supervised or unsupervised machine learning methods. In some cases, supervised machine learning methods may be regression, support vector machines, tree-based methods, neural networks, or nearest neighbor methods. In some cases, unsupervised machine learning methods may be clustering, neural networks, principal component analysis, or matrix factorization.
[0158] Various aspects of the systems and methods provided herein (such as computer system 201) may be embodied in programming. Various aspects of the technology may be considered as “products” or “articles of manufacture”, which are generally in the form of machine (or processor) executable code and / or associated data carried or embodied on a machine-readable medium of a certain type. The machine-executable code may be stored in an electronic storage unit, such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. "Storage" media can include any or all of a computer's tangible memory, processor, etc., or associated modules such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time. All or part of the software can sometimes communicate via the Internet or various other telecommunications networks. For example, such communication can enable software to be loaded from one computer or processor to another, such as from a pipe...The server or host is loaded into the computer platform of the application server. Therefore, another type of medium that can carry software elements includes light waves, radio waves, and electromagnetic waves, such as physical interfaces between local devices, used through wired and optical ground networks, and through various air links. Physical elements carrying such waves, such as wired or wireless links, optical links, etc., can also be considered as media carrying software. As used herein, unless limited to non-transitory tangible "storage" media, terms such as computer or machine "readable medium" refer to any medium that participates in providing instructions to a processor for execution.
[0159] Therefore, machine-readable media such as computer executable code can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media include, for example, optical discs or disks, such as any storage device in any computer, such as those used to implement databases as shown in the figures. Volatile storage media include dynamic memory, such as the main memory of a computer platform. Tangible transmission media include coaxial cables, copper wires, and optical fibers, including lines that constitute a bus within a computer system. Carrier transmission media can take the form of electrical or electromagnetic signals, or sound or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Therefore, common forms of computer-readable media include, for example: floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched card tapes, any other physical storage media with a perforated pattern, RAM, ROM, PROM and EPROM, FLASH-EPROM, any other memory chips or cartridges, carrier waves for transmitting data or instructions, cables or links for transmitting such carrier waves, or any other media from which a computer can read programming code and / or data. Many of these forms of computer-readable media can participate in carrying one or more sequences of one or more instructions to a processor for execution.
[0160] Computer system 201 may include or communicate with an electronic display 1135, which includes a user interface (UI) 240. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.
[0161] The methods and systems of this disclosure may be implemented by one or more algorithms. The algorithm can be implemented in software when executed by the central processing unit 205.
[0162] Although preferred embodiments of the invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. The invention is not intended to be limited to the specific examples provided herein. Although the invention has been described with reference to the foregoing description, the description of embodiments herein...The descriptions and illustrations are not intended to be interpreted in a limiting sense. Many variations, alterations, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it should be understood that all aspects of the invention are not limited to the specific descriptions, configurations, or relative proportions set forth herein, and depend on a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of the invention. Therefore, the invention is also intended to cover any such alternatives, modifications, variations, or equivalents. The following claims are intended to define the scope of the invention and thereby cover the methods and structures and their equivalents within the scope of these claims.
[0163] Kit This disclosure provides a kit for identifying or monitoring a disease or condition (e.g., cancer) in a subject. Kit Specification 37 / 55 pages 48 CN 121464223 A May contain probes for quantitative measurements (e.g., indicating presence, absence, or relative amount) of the sequence at each locus in a group of cancer-related genomic loci in a subject sample. Quantitative measurements (e.g., indicating presence, absence, or relative amount) of the sequence at each locus in a group of cancer-related genomic loci in a sample may indicate a disease or condition (e.g., cancer) in the subject. The probes may be selective for sequences of a group of cancer-related genomic loci in a sample. The kit may include instructions for processing the sample using the probes to generate datasets indicating quantitative measurements (e.g., indicating presence, absence, or relative amounts) of sequences at each locus in the group of cancer-related genomic loci in the sample.
[0164] The probes in the kit may be selective for sequences of a group of cancer-related genomic loci in a sample. The probes in the kit may be configured to selectively enrich nucleic acid (e.g., RNA or DNA) molecules corresponding to the group of cancer-related genomic loci. The probes in the kit may be nucleic acid primers. The probes in the kit may have sequence complementarity with one or more nucleic acid sequences from a group of cancer-related genomic loci or genomic regions. The group of cancer-associated genomic loci or microbial-associated genomic loci or genomic regions may include at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20 or more different groups of cancer-associated genomic loci or genomic regions.
[0165] The instructions in the kit may include instructions for using probes selective for sequences of the group of cancer-associated genomic loci in cell-free biological samples to determine the sample. These probes may be selective for sequences of cancer-associated genomic loci from multiple cancer-associated loci.Because one or more nucleic acid sequences (e.g., RNA or DNA) in a group of cancer-related genomic loci have sequence complementarity. These nucleic acid molecules may be primers or enriched sequences. Instructions for assaying cell-free biological samples may include performing array hybridization, polymerase chain reaction (PCR), or nucleic acid sequencing (e.g., DNA sequencing or RNA sequencing) to process the sample, thereby generating a dataset indicating quantitative measurements (e.g., indicating presence, absence, or relative amount) of the sequence at each locus in the group of cancer-related genomic loci in the sample. Quantitative measurements (e.g., indicating presence, absence, or relative amount) of the sequence at each locus in the group of cancer-related genomic loci in the sample may indicate a disease or condition (e.g., cancer).
[0166] Instructions in the kit may include instructions for measuring and interpreting assay reads that can be quantified at one or more sites in the group of cancer-related genomic loci to generate a dataset indicating quantitative measurements (e.g., indicating presence, absence, or relative amount) of the sequence at each locus in the group of cancer-related genomic loci in the sample. For example, array hybridization or polymerase chain reaction (PCR) quantification corresponding to groups of cancer-associated genomic loci can generate datasets that indicate quantitative measurements of the sequence at each locus in the groups of cancer-associated genomic loci in a sample (e.g., indicating presence, absence, or relative amount). Assay readouts may include quantitative PCR (qPCR) values, digital PCR (dPCR) values, digital droplet PCR (ddPCR) values, fluorescence values, etc., or their normalized values.
[0167] Examples Example 1: Processing plasma-derived cell-free DNA using a full methylome enrichment platform Cell-free methylated DNA immunoprecipitation and high-throughput sequencing (cfMeDIP-seq) was developed as a non-degradable liquid biopsy method to avoid the limitations of bisulfite sequencing. Bisulfite-free methods may require consistently high methylation binding specificity to detect potentially rare methylation events in circulating tumor DNA (ctDNA). Furthermore, as shown in Figure 23, with increased genomic DNA contamination, the effective cfDNA input into the cfMeDIP-seq workflow decreased, thereby reducing the probability of detecting rare ctDNA and highlighting the need for an improved workflow. gDNA contamination was measured by electrophoretic fragment size analysis (e.g., TapeStation). No correlation was observed between methylation specificity and the percentage of genomic DNA contamination. Figure 24 also shows the importance of high methylation specificity for methylation-based liquid bioassays. This figure compares two clinical samples, one with high (y-axis) methylation specificity and the other with low (x-axis) methylation specificity.CpG counts represent the counts of sequencing fragments aligned with known 0 CpG regions in the human genome. The gradient in the figure represents the number of non-CpG regions in the genome with observed counts in each sample. The deviation of no CpG reads and methylation specificity percentage from 100% indicates non-specific binding of the anti-5mC antibody during immunoprecipitation. Perfect methylation specificity (i.e., 100%) during immunoprecipitation would result in no DNA fragment counts in non-CpG regions. In real samples, generally, the lower the methylation specificity, the more counts are observed in regions without CpG. Therefore, in this embodiment, cfMeDIP-seq has been further improved with the aim of developing a robust whole-genome methylome enrichment platform for clinical use.
[0168] As shown in Figure 25, four different cfMeDIP-seq workflows were performed with cell-free DNA (cfDNA) samples to determine the methylation binding specificity of each workflow. Samples included samples derived from plasma from subjects and samples derived from genomic DNA that had been cut to mimic cfDNA. To perform workflow 1, a library was prepared from a sample of cfDNA (n=110) mixed with spike-in DNA. Next, filler DNA was added to the prepared library to generate a sample mixture. The sample mixture was heat-denatured and rapidly cooled. Then, 5-mc conjugate and magnetic beads were added sequentially to the mixture for incubation, allowing single-stranded DNA (ssDNA) to bind to the magnetic bead-5mC conjugate complex. The magnetic bead-5mC conjugate complex was separated using a magnet, and the captured ssDNA was eluted from the complex. The captured ssDNA was cleaned, and then 14 cycles of PCR amplification were performed. Another round of cleanup was performed with PCR amplicons, followed by quality control analyses, such as measuring methylation binding specificity.
[0169] Workflow 2 (n=174) was performed similarly to workflow 1; however, an additional bead-capture step was performed before adding the filler DNA. To perform bead capture, the eluent after library preparation and cleaning is further captured on a magnetic rack, discarding any carried beads (e.g., solid-phase reversible fixation (SPRI) beads) to minimize any non-specific binding with the beads. Workflow 3 (n=63) is identical to workflow 2, except that the number of cycles used in the PCR amplification step is reduced to 13 cycles.
[0170] Workflow 4 (n=189) follows a similar scheme to workflow 3, except that a pre-binding step is performed, in which the magnetic beads and 5-mC conjugate are mixed together separately to form a magnetic bead-5-mC conjugate complex. After thermal denaturation and rapid cooling of the mixture sample, the magnetic bead-5-mC conjugate complex is added. Another difference between workflow 4 and workflow 3 is that…Instead of eluting the captured ssDNA from the magnetic bead-5-mC conjugate complex and cleaning the captured ssDNA for PCR amplification, the PCR amplification step was performed on the magnetic beads.
[0171] As shown in Figure 26, for different workflows, the average methylation specificity was measured by calculating the percentage of synthetic oligonucleotides with known methylation states incorporating the sample DNA that were recovered by immunoprecipitation (e.g., recovery of exogenous methylated fragments) during the library preparation step. Workflow 4 resulted in higher methylation specificity and reduced variability between samples compared to other workflows.
[0172] The methylation specificity of the enrichment platform at different tube and sample ages was also analyzed for samples. Table 1 shows the effect of tube and sample age on methylation specificity.
[0173] Table 1 Specification 39 / 55 pages 50 CN 121464223 A
[0174] The results showed that methylation specificity was consistent regardless of tube type (Δ = 0.31%) or sample storage time (Δ < 0.26%). Example 2: Genome-wide methylome enrichment platform for early detection of multiple cancers (MCED) Genome-wide mapping of DNA methylation in circulating cell-free DNA (cfDNA) can overcome the key sensitivity problem of detecting circulating tumor DNA (ctDNA) in early cancer or low-shedding tumor subjects.
[0175] To investigate the genome-wide methylome enrichment platform for early detection of multiple cancers, a retrospective case-control study was conducted using plasma samples from commercial biobanks and the University Health Network biobank. Samples were obtained from individuals who had been diagnosed with cancer but had not yet started treatment. For non-cancer controls, samples were obtained from age- and sex-matched individuals without a known cancer diagnosis. For controls, there was confirmed cancer-free follow-up for at least 12 months after sample collection. Controls were excluded if the individual was 75 years or older at the time of sample collection, or if they were known to have multiple comorbidities.
[0176] 5–10 ng cfDNA was extracted from plasma samples and analyzed using a bisulfite-free, non-degrading whole-genome DNA methylation enrichment platform based on cell-free methylated DNA immunoprecipitation and high-throughput sequencing (cfMeDIP-seq) technology, as described in Example 1, Figure 25 (Workflow 1). Samples were divided into different sets to train and test a machine learning classifier (e.g., a linear-based method) consisting of differentially methylated regions of the whole methylation group to distinguish between cases and controls. Initial training was performed on 1,536 samples from 8 cancer types and cross-validation was performed with 100 random split iterations (80:20) to obtain the area under the receiver operating characteristic (AUC) and the 95% confidence interval (CI) of the median probability. As shown in Table 2, all cancer cases had an AUC of 0.94 (95%).The AUCs were distinguished from controls (CI: 0.93, 0.96), with individual cancer types (e.g., bladder cancer, breast cancer, colorectal cancer, head and neck cancer, lung cancer, ovarian cancer, prostate cancer, kidney cancer) ranging from 0.91 to 0.97. The AUCs for all cancers at all stages were: 0.94 for stage I / II cancer (95% CI: 0.92, 0.95) and 0.95 for stage III / IV cancer (95% CI: 0.94, 0.96). In the subgroup of low-shedding cancers (e.g., bladder cancer, breast cancer, prostate cancer, kidney cancer), the AUC was 0.92 (95% CI: 0.91, 0.94), and it showed similar performance in stages I / II (AUC 0.91; 95% CI: 0.89, 0.93) and III / IV (AUC 0.93; 95% CI: 0.91, 0.95). These results indicate that the genome-wide methylome enrichment platform can detect early-stage multi-cancer and low-shedding cancer subgroups.
[0177] Table 2. Overall AUC (95%) for all cancers and 8 cancer types. Instruction manual 40 / 55 pages 51 CN 121464223 A
[0178]
[0179] Interim training was also performed on 1,906 samples of 12 cancer types (e.g., bladder cancer, breast cancer, colorectal cancer, endometrial cancer, esophageal cancer, head and neck cancer, hepatobiliary cancer, lung cancer, ovarian cancer, pancreatic cancer, prostate cancer, and kidney cancer). Table 3 shows the characteristics between cancer cases and control cases. A machine learning classifier was used to distinguish between cancer cases and non-cancer controls. Cross-validation of the dataset was performed 5-fold for 20 iterations to obtain 95% CI for AUC and median probability.
[0180] Table 3. Interim training cohort (N=1,906) Instruction manual 41 / 55 pages 52 CN 121464223 A
[0181] As shown in Figure 3, genome-wide methylation assays distinguished cancer cases from non-cancer controls, with an overall AUC of 0.94 (95% CI: 0.93, 0.95). The AUCs for different stages were 0.92 (Stage I), 0.95 (Stage II), 0.95 (Stage III), and 0.97 (Stage IV), indicating that the AUC increased with cancer stage, but remained high even for early-stage cancers. When each cancer type (e.g., bladder cancer, breast cancer, colorectal cancer, endometrial cancer, esophageal cancer, head and neck cancer, hepatobiliary cancer, lung cancer, ovarian cancer, pancreatic cancer, prostate cancer, and kidney cancer) was evaluated separately, the AUCs ranged from 0.89 to 0.99, as shown in Figures 4A-4L. Subgroups of cancers with low shedding tumors (e.g., bladder cancer, breast cancer, endometrial cancer, and kidney cancer) were also evaluated separately. The AUC for this subgroup of low-shedding tumors was 0.91 (95% CI: 0.89, 0.97)..93), considering that only 8% are stage IV cancers (as shown in Table 4), indicating that MCED can be achieved through a whole-genome methylome enrichment platform.
[0182] Table 4. Staging distribution of all cancers, 12 cancer types and all low-shedding cancers Specification 42 / 55 pages 53 CN 121464223 A
[0183] Example 3: Evaluation of a whole-genome methylome enrichment platform for ctDNA quantification in renal cell carcinoma (RCC) Circulating tumor DNA (ctDNA) can be used to identify cancer and the presence of minimal residual disease. Using plasma-based assays to quantify ctDNA may be a useful cancer management tool for assessing prognosis; however, some methods require tumor tissue for analysis or are limited to tumor types that tend to have a larger amount of relevant ctDNA. This experiment demonstrates the feasibility of using a tumor-unaware whole-genome methylome enrichment platform to quantify plasma ctDNA and predict prognosis in RCC.
[0184] Cell-free DNA isolated from plasma (5-10 ng) was used to analyze pre-processed biobank samples (University Health Network, Ontario Tumor Biobank) from newly diagnosed stage I-IV RCC individuals using a bisulfite-free, non-degrading whole-genome DNA methylation enrichment platform, as described in Example 1, Figure 25 (Workflow 1). ctDNA was quantified based on the average normalized count of information-rich regions.
[0185] The algorithm uses size-reduced fragments of 1-150 base pairs (bp) containing at least 10 CpG to identify cancer-related methylation. Counts were aggregated in non-overlapping 300 bp windows and normalized by sequencing depth. Methylation signals were compared for each 300 bp window using a univariate method (e.g., cancer vs. non-cancer controls). For each window, a p-value between cancer and non-cancer controls was calculated using the Wilcoxin rank-sum test. The p-value was adjusted using the Benjamini-Hochberg method. Significantly differentially methylated regions were selected using an adjusted p-value threshold of less than or equal to 0.1. Differential methylation analysis identified 2027 hypermethylated regions enriched with CpG islands. ctDNA quantification scores were generated based on the average normalized counts of the 2027 regions, and these scores were adjusted for methylation specificity. An event was defined as cancer recurrence or progression. A ctDNA quantity threshold was set such that 95% of event-free samples were below the threshold (e.g., 95% specificity). The event occurrence time was compared between samples with ctDNA quantities above the threshold (pages 43 / 55 of the specification, CN 121464223 A) and samples with ctDNA quantities below the threshold. Figure 5A shows that the threshold for renal cell carcinoma was set to 0.37.
[0186] The cohort included 151 samples [64 stage I, 2 stage II, 23 stage III, 15 stage IV, and 47 with unknown or incomplete staging information]. The median follow-up time was 15.7 months, with 21 events occurring. Samples with ctDNA levels above the threshold were more likely to relapse or progress than those with ctDNA levels below the threshold [hazard ratio 13.28 (95% CI 5.47, 32.26), log-rank P < 0.001] (Figure 5B). Samples with ctDNA levels below the threshold were more likely to avoid cancer recurrence or recurrence.
[0187] This experiment demonstrates the feasibility of using a blood-based, tumor-unaware whole-genome methylome enrichment platform for ctDNA quantification and prognostic prediction in renal cell carcinoma. This is a promising demonstration of the prognostic performance of a cancer type that is often difficult to detect due to low ctDNA levels. Further evaluation of post-processing and longitudinal samples will further test the platform.
[0188] In another similar study, biobank plasma samples from newly diagnosed stage I-IV RCC individuals were included (collected between 2015 and 2021; at the Princess Margaret Cancer Center of the University Health Network and the Ontario Tumour Bank). All samples were obtained after cancer diagnosis but before surgery or other definitive treatment. Table 5 shows clinicopathological information for the 148 samples used, and Figure 8 shows the age of the individuals at the time of sample collection.
[0189] Table 5. Clinicopathological Information (N=148)
[0190] The samples used were analyzed using a bisulfite-free, non-degrading whole-genome DNA methylation enrichment platform using approximately 5-10 ng of cfDNA extracted from plasma. The whole-genome methylation assays used were based on cell-free methylated DNA immunoprecipitation and high-throughput sequencing (cfMeDIP-seq) technology, as described in Example 1, Figure 25 (Workflow 1).
[0191] For analysis, an event was defined as cancer recurrence, progression, or death due to renal cell carcinoma (based on the earliest occurrence, see specification page 44 / 55, CN 121464223 A). A threshold for ctDNA quantity was set for baseline prediction such that approximately 95% of event-free samples were below the threshold (i.e., approximately 95% specificity). Event-free survival was estimated using the Kaplan-Meier method, and differences were assessed using a log-rank test between the overall population and individuals with stage I-III cancer.
[0192] As shown in Figure 9, individuals with ctDNA quantification above the cutoff value showed significantly worse event-free survival across the entire cohort. Furthermore, as shown in Figure 10, individuals with ctDNA quantification above the cutoff value also showed significantly worse event-free survival in the stage I-III disease subgroup.The event-free survival rate was significantly worse. Individuals were stratified based on a specific cutoff value of 95% or higher for ctDNA quantification.
[0193] Thus, these data demonstrate the feasibility of using a blood-based genome-wide methylome enrichment platform for ctDNA quantification and prognostic performance in RCC. The observed performance also represents a promising indication for predicting cancer types that are often difficult to detect due to low ctDNA levels. Furthermore, the assay used here is tumor-naïve, meaning that no patient-specific tumor tissue is required to generate a custom genome for ctDNA.
[0194] Example 4: Prognostic Performance of a Genome-Wide Methylome Enrichment Platform in Head and Neck Cancer Quantifying circulating tumor DNA (ctDNA) using plasma-based assays is emerging as a promising new approach for cancer management. ctDNA quantification can be used to assess prognosis and detect minimal residual disease after initial treatment. This experiment demonstrates the feasibility of using a tumor-naïve genome-wide methylome enrichment platform for ctDNA quantification and prognostic prediction in head and neck cancer.
[0195] Using 5–10 ng of cell-free DNA isolated from plasma, pre-processed biobank samples (University Health Network) from newly diagnosed stage I–IV head and neck cancer individuals were analyzed using a bisulfite-free, non-degrading whole-genome DNA methylation enrichment platform, as described in Example 1, Figure 25 (Workflow 4). ctDNA was quantified based on the average normalized count of information-rich regions.
[0196] The algorithm uses 1–150 bp size-reduced fragments containing at least 10 CpGs to identify cancer-related methylation. Counts were aggregated in non-overlapping 300 bp windows and normalized by sequencing depth. Methylation signals were compared for each 300 bp window using a univariate approach (e.g., cancer vs. non-cancer controls). For each window, a p-value between cancer and non-cancer controls was calculated using the Wilcoxin rank-sum test. The p-value was adjusted using the Benjamini-Hochberg method. Significantly differentially methylated regions were selected using an adjusted p-value threshold less than or equal to 0.1. Differential methylation analysis identified 2027 hypermethylated regions enriched with CpG islands. ctDNA quantification scores were generated based on the average normalized counts of the 2027 regions, and these scores were adjusted for methylation specificity. An event was defined as cancer recurrence or progression. A ctDNA quantity threshold was set such that 95% of event-free samples were below the threshold (e.g., 95% specificity). The time to recurrence or progression was compared between samples with ctDNA quantities above the threshold and those with ctDNA quantities below the threshold. Figure 6A shows that the threshold for head and neck cancer was set to 3.96.
[0197] Of the 93 samples included (7 stage I, 17 stage II, 23 stage III, and 46 stage IV), the median follow-up was 50.6 months, and 25 events occurred. Samples with ctDNA levels above the threshold had a significantly higher likelihood of recurrence or progression [hazard ratio (HR) 3.18 (95% CI 1.09, 9.28), log-rank P = 0.026] (Figure 6B). In multivariate analyses considering cancer stage and clinical characteristics (sex, age, smoking history, BMI), ctDNA quantification above the threshold showed similar associations [HR 3.51 (95% CI 1.1, 11.19), P = 0.034]. This experiment demonstrates the feasibility of using treatment-naïve plasma samples for ctDNA quantification and prognostic prediction in head and neck cancer using a blood-based, tumor-unaware whole-genome methylation enrichment platform.
[0198] Furthermore, quantifying ctDNA using plasma-based assays can also be used to assess prognosis and improve post-treatment surveillance. For head and neck cancer, this provides an opportunity to inform whether chemotherapy should be administered after surgery for certain non-metastatic tumors (stages I-IVb) and to improve post-treatment surveillance by detecting minimal residual disease before clinical testing.
[0199] In another similar study, plasma samples from newly diagnosed stage I-IV head and neck cancer individuals (collected between 2008 and 2019; Princess Margaret Cancer Centre, University Health Network) were included. All samples were obtained after cancer diagnosis but before surgery or other definitive treatment. All samples were analyzed using a bisulfite-free, non-degrading whole-genome DNA methylation enrichment platform using 5-10 ng of cell-free DNA isolated from plasma. The assay was based on cell-free methylated DNA immunoprecipitation and high-throughput sequencing (cfMeDIP-seq) technology, as described in Example 1, Figure 25 (Workflow 4). Machine learning algorithms were used to quantify ctDNA in information-rich regions.
[0200] For analysis, an event was defined as cancer recurrence, progression, or death due to head and neck cancer (whichever occurred earliest). The longest follow-up period was 5 years after diagnosis. A threshold for ctDNA quantity was set such that 95% of event-free samples were below the threshold (i.e., 95% specificity). The Kaplan-Meier method was used to estimate event-free survival, and samples with ctDNA above the threshold were compared with those with ctDNA below the threshold. The difference between the two groups was assessed by a log-rank test. Multivariate Cox regression analysis was used to adjust for known prognostic covariates.
[0201] Table 6 shows the clinicopathological information of the subjects. Table 7 shows the preliminary cancer treatment profile of the subjects. As shown in Tables 6 and 7, a total of 91 samples were included, with a mean follow-up time of 50.6 months and 27 events.
[0202] Individuals with ctDNA levels above the threshold showed significantly worse event-free survival, as shown in Figure 7. The lead time (time interval from blood draw to event) ranged from 1.25 to 16.37 months, with a median lead time of 4.87 months (Figure 7). In multivariate analyses considering cancer stage and clinical characteristics, ctDNA quantification above the threshold showed similar correlations (Table 8).
[0203] Table 6. Clinicopathological Information (N=91) Instruction Manual 46 / 55 pages 57 CN 121464223 A
[0204] Table 7. Overview of Initial Cancer Treatment
[0205] Table 8. Multivariate Analysis Instruction Manual 47 / 55 pages 58 CN 121464223 A
[0206] Similar to previous experiments, this experiment also demonstrated the feasibility of using a blood-based, tumor-unaware whole-genome methylome enrichment platform for ctDNA quantification and prognostic prediction in head and neck cancer. This test uses tumor-unaware material, meaning that patient-specific tumor tissue is not required to generate a custom set for ctDNA detection. The tumor-unaware approach allows for more flexible clinical use, 1) without contact with the original tumor tissue, and 2) not limited to differentially methylated genomic regions at diagnosis, which can allow for sensitive longitudinal minimal residual disease (MRD) detection.
[0207] A cfMEDIP-seq-based tissue-agnostic, genome-wide methylome enrichment platform for head and neck cancer can also be used to predict recurrence, guide adjuvant therapy after therapeutic intent treatment, and detect early recurrence. To investigate the detection of MRD in head and neck cancer patients after therapeutic intent treatment, longitudinal data acquisition and sampling were used to analyze biobank samples (Princess Margaret Cancer Center) from individuals with stage I-IVb human papillomavirus (HPV) negative and HPV positive head and neck cancer. The entire cohort included 325 unique patients and 1,155 samples. Patients diagnosed with Epstein-Barr virus-associated nasopharyngeal carcinoma were excluded. Of the 1,155 samples in the entire cohort, 1,119 (96.9%) passed the quality threshold for cfDNA quantity and were processed by cfMeDIP. These samples were divided into different ensembles to train and test machine learning classifiers with differentially methylated regions. As shown in Figure 12, blood collection time points include at diagnosis, before therapeutic intent treatment (baseline (BL)), and approximately 3 months after therapeutic intent treatment (standard).The time points were B1, 12 months (B2), and 24 months (B3). Therapeutic intent included surgery alone, radiotherapy (RT) + / - chemotherapy, or surgery + RT + / - chemotherapy. 5–10 ng of plasma cfDNA was used for each sample and enriched using a bisulfite-free, non-degrading whole-genome methylome platform based on cfMeDIP-seq, as described in further detail in Example 1, Figure 25 (Workflow 4). MRD signals were quantified based on the average normalized count of information-rich methylated regions and binarized into positive and negative groups. Recurrence-free survival (RFS) was compared longitudinally between patients who tested positive and those who tested negative at 3 months post-therapeutic treatment.
[0208] A total of 173 samples from 52 unique patients (Stage I (33%), Stage II (17%), Stage III (23%), Stage IV (27%)) were analyzed and correlated with relapse in this interim training outcome. At the landmark time point, patients who tested positive showed significantly worse RFS than those who tested negative (hazard ratio (HR) 8.91; 95% CI, 3.14–25.26, P<0.001). Combined with a series of longitudinal samples, the RFS of patients who tested positive was statistically significantly worse than that of patients who tested negative (HR 10.47; 95% CI, 3.81–28.82, P<0.001).
[0209] Similarly, a larger total sample size of 249 samples collected from 75 patients was also analyzed. The clinicopathological information of the 75 patients is shown in Table 9. The analysis included comparing the prognoses of cancer patients with and without detected ctDNA. The response rate (RFS) was compared between the two groups using a two-tailed log-rank test and the Kaplan-Meier (KM) method. The risk factor (HR) between the two groups was estimated using a Cox proportional hazards model. These analyses were performed at the landmark time point (Figure 13A) and longitudinally (Figure 13B). As shown in Figures 13A and 13B, ctDNA positivity was found to predict survival outcomes in patients with head and neck cancer following therapeutically intended treatment. ctDNA positivity was associated with RFS at the landmark time point and longitudinally. Significant differences in RFS were observed when patients were stratified by ctDNA status, with an HR of 10.97 at the landmark time point (CI: 4.76 to 25.29; p < 0.001) and an HR of 22.83 longitudinally (CI: 2.8 to 186; p < 0.001). These interim training analyses showed that, following therapeutically intended treatment, MRD detection using a blood-based, tissue-unaware, genome-wide methylome enrichment platform was strongly correlated with RFS in HNC patients, with hazard ratios consistent with tumor-informed assays.
[0210] Furthermore, the cfMEDIP-seq-based genome-wide methylome enrichment platform was found to be able to monitor ctDNA dynamics, i.e., the quantitative changes in ctDNA levels over time. ctDNA was quantified and plotted over time for all relapsed and non-relapsed subjects to demonstrate the ability to monitor ctDNA dynamics. As shown in Figure 14, the ctDNA quantification trajectory plot shows the correlation with relapse-free and relapse-free outcomes. Estimated ctDNA quantification from pre-treatment (BL) to post-treatment (e.g., B1, B2, B3) was consistent with the expected ctDNA dynamics. ctDNA in individual patients before and after treatment with therapeutic intent was also analyzed. Figure 15 shows a representative case study of ctDNA dynamics in three individual patients (Patient A, Patient B, and Patient C). Patient A (Figure 15, top left) was diagnosed with stage III oropharyngeal HPV tumor. ctDNA was detected at diagnosis before treatment with RT and at 315 and 413 days post-diagnosis, prior to distant recurrence detected at 502 days (lead time 187 days). Patient B (Figure 15, top right) was diagnosed with stage IVA hypopharyngeal HPV+ tumor. ctDNA was detected at diagnosis prior to chemoradiotherapy (CRT), but not at 107 days post-diagnosis and after CRT completion. ctDNA was detected again at 413 days post-diagnosis, prior to distant recurrence at 458 days (lead time 45 days). Patient C (Figure 15, bottom) presented with stage I oropharyngeal HPV+ tumor. ctDNA was detected at diagnosis prior to radiotherapy (RT), but not at any time point post-treatment. Patient C remained disease-free until the last clinical follow-up (>2.5 years). These data suggest that a blood-based, tissue-unaware, genome-wide methylome enrichment platform can demonstrate robust performance in quantifying ctDNA and monitoring ctDNA dynamics in HPV+ and HPV- head and neck cancers.
[0211] Overall, these results may aid in decisions regarding treatment escalation or de-escalation after therapeutically intended treatment and can be used to monitor recurrence before clinical or radiographic presentation.
[0212] Table 9. Patient Clinical Demographic Information 49 / 55 pages 60 CN 121464223 A
[0213] Example 5: Robust Processing of Plasma-Derived Cell-Free DNA (cfDNA) Using a Full Methylome Enrichment Platform Pre-analytical Variables and Quality Control Cell-free methylated DNA immunoprecipitation and high-throughput sequencing (cfMeDIP-seq) was developed as a non-degradable liquid biopsy method to avoid the limitations of bisulfite sequencing. However, bisulfite-free methods may require consistently high methylation binding specificity to detect potential rare methylation events in circulating tumor DNA (ctDNA).cfMeDIP-seq has been improved with the aim of developing a robust whole-genome methylome enrichment platform for clinical use. In this experiment, the effects of pre-analytical variables and analytical configurations on the methylation specificity of the platform were investigated (e.g., downregulation of methylation region specificity).
[0214] cfDNA from plasma was used for the whole-genome methylome enrichment platform. The cfDNA was prepared as a standard library, combined with DNA filler, denatured, and immunoprecipitated using an anti-5-mC antibody. The captured methylated DNA was amplified and sequenced as described in Example 1, Figure 25 (Workflow 4). The methylation specificity of the immunoprecipitation step was monitored using methylated and unmethylated spike-in DNA fragments added before adaptor ligation (e.g., downregulation of methylation region specificity). Methylation specificity was assessed among several pre-analytical variables: sample collection tube type (Streck, EDTA), sample age (0–5 years, 5–10 years, ≥10 years), and genomic DNA (gDNA) contamination (1%–50% gDNA). Each pre-analytical variable was assessed as the difference between mean methylation specificity (categorical variable) or correlation (continuous variable) in a cohort of >4,000 archived plasma samples from cancer and non-cancer individuals. The immunoprecipitation step was further optimized using 20 replicates from donor-derived cfDNA to improve centrality and reduce variability in methylation specificity.
[0215] Methylation specificity (e.g., downregulated methylation region specificity) was consistent and independent of tube type (Δ = 0.29%) or sample age (Δ < 0.32%). There was no correlation between methylation specificities when comparing samples with gDNA contamination ranging from 1% to 50% (Kendall correlation = 0.0064). Modifications to the immunoprecipitation step were evaluated to improve methylation specificity, with a significant improvement observed after optimization (Kolmogorov-Smirnov p-value < 0.00001). The optimized assay had an average methylation specificity of 99.7%, with 20 / 20 samples exhibiting ≥ 99.6% specificity.
[0216] The data indicate that the whole-genome methylome enrichment platform is a powerful and versatile assay that is minimally affected by pre-analytical variables and consistently exhibits high methylation specificity.
[0217] Example 6: Analytical performance of the whole-genome methylome enrichment platform for detecting minimal residual disease from plasma-derived cell-free DNA. Tissue-unaware methods for detecting cancer signals from plasma can have significant benefits, especially when tissue is unavailable for evaluation or response speed is important. Plasma-derived cell-free DNA (cfDNA) can be used for detection.Cancer, including minimal residual disease (MRD) in patients receiving therapeutic cancer treatment. However, these tests may require highly sensitive methods for detecting cancer signals for clinical applications. Therefore, a genome-wide methylome enrichment platform using cell-free methylated DNA immunoprecipitation and high-throughput sequencing (cfMeDIP-seq) can be combined with a custom algorithm that utilizes differentially methylated regions (DMRs) found in cfDNA to distinguish between cancer and non-cancer signals. Therefore, the analytical performance metrics of an algorithm under development for detecting MRD using a methylation-based approach were investigated.
[0218] Plasma-derived cfDNA was used for the genome-wide methylome enrichment platform. As shown in Figure 16 and further detailed in Example 1 (Figure 25 – Workflow 4), a standard library was prepared from plasma-derived cfDNA (10 ng input mass) combined with spike-in DNA fragments, combined with DNA filler, denatured, and immunoprecipitated using an anti-5-mC antibody. The captured methylated DNA was amplified and sequenced. The sequencing results were then subjected to bioinformatics pipeline and algorithm classification. Candidate algorithms incorporating DMR were used to quantify cancer-specific methylation. The detection limit and accuracy were then evaluated using engineered cancer samples designed to mimic low levels of circulating tumor DNA (ctDNA) representative of MRD.
[0219] To evaluate the detection limit, samples were engineered by titrating enzymatically fragmented DNA from three immortalized tumor-derived cell lines (H-1299 lung cancer cell line, FaDu head and neck (H&N) cancer cell line, and A-253 H&N cancer cell line) into pooled non-cancer donor-derived cfDNA in titration series targeting less than 1% ctDNA levels. As shown in Figure 17, each ctDNA level was replicated using five techniques, for a total of 65 independent cfMeDIP runs. As shown in Figure 18, all non-cancer samples and engineered cancer samples met process quality control metrics, including methylation specificity greater than or equal to 98.5% and 80 million unique molecules greater than or equal to. Furthermore, as shown in Figure 19, ctDNA methylation scores at different levels of ctDNA were measured, and the detection computational limit (LoD95) at 95% sensitivity was found to be less than 0.1% using a tissue-unaware algorithm. Thresholds were set using 12 non-cancer donors and 5 replicates from a non-cancer pool to establish a 95% true negative rate.
[0220] To evaluate the accuracy of the blood-based, tissue-unaware, whole-genome methylome enrichment platform, samples were designed as follows: enzymatically fragmented DNA from two immortalized tumor-derived cell lines (A-253 H&N cancer cell line and FaDu H&N cancer cell line) was titrated to high (1.12%), intermediate (0.84%), and low (0.56%) values in 18 technical replicates.Pages 51 / 55 62 CN 121464223 A The pooled non-cancer donor-derived cfDNA at the ctDNA level is shown in Figure 20. Next, the samples were processed using a blood-based, tissue-unaware, whole-genome methylome enrichment platform and replicated with different operators, sequencing runs, and antibody reagent batches. As shown in Figure 21, the consistency with expected results and variability was assessed using ctDNA methylation scoring. The results were consistent with expected results across all levels of ctDNA. Furthermore, as shown in Figure 22, the combination of full variable analysis (operator, sequencing run, and antibody reagent batch) resulted in a variance component (CV%) of less than 40% for all levels of ctDNA.
[0221] In summary, the assessment of the detection limit and accuracy indicates that a blood-based, tissue-unaware, whole-genome methylome enrichment platform utilizing a non-degradation method, combined with specific algorithms and DMR levels suitable for MRD detection, was effective.
[0222] Example 7: Prognostic Performance of a Whole-Genome Methylome Enrichment Platform in Early-Stage Non-Small Cell Lung Cancer (NSCLC) Circulating tumor DNA (ctDNA) can be used to identify cancer and the presence of minimal residual disease (MRD). Quantification of ctDNA may be a useful cancer management tool for assessing prognosis. In this example, the feasibility of using a tumor-unknown whole-genome methylome enrichment platform to quantify plasma ctDNA and predict recurrence in early-stage non-small cell lung cancer (NSCLC) was evaluated.
[0223] In a retrospective evaluation, pre-processed archived samples from newly diagnosed stage I and II NSCLC patients (collected from 2009 to 2013, Princess Margaret Cancer Center) were used. Blood samples were obtained after cancer diagnosis and before treatment. Samples were analyzed using 5–10 ng of cell-free DNA isolated from plasma using a bisulfite-free, non-degrading whole-genome DNA methylation enrichment platform as described in Example 1, Figure 25 (Workflow 4). Samples from 41 patients were included. Table 10 describes the clinical demographic information of the patients. ctDNA was quantified based on the mean normalized count of information-rich regions. An event was defined as cancer recurrence or death from any cause, whichever occurred earlier. A threshold for ctDNA count was set, with 100% of samples from event-free patients falling below the threshold (i.e., 100% specificity). The time to recurrence or death was compared between samples with ctDNA counts above the threshold and those with counts below the threshold. Recurrence-free survival was estimated using the Kaplan-Meier method and compared between samples with ctDNA counts above and below the threshold. Differences between the two groups were assessed using a log-rank test. Multivariate Cox regression analysis was used to adjust for pre-treatment levels.Known prognostic values for ctDNA levels (adjusted for covariates identified in univariate analysis).
[0224] Twenty-seven events occurred, with a median follow-up time of 55.8 months. As shown in Figure 11, samples with ctDNA above the threshold showed significantly worse relapse-free survival [hazard ratio (HR) 2.70 (95% CI 1.26, 5.78), log-rank P = 0.008]. Multivariate analysis (Table 11) showed that even after considering histological factors (selected using univariate analysis), samples with ctDNA above the threshold still showed significantly worse relapse-free survival [HR 2.79 (95% CI 1.30, 6.02), P = 0.009].
[0225] Overall, the data suggest that this blood-based genome-wide methylome enrichment platform can be used for ctDNA quantification and prediction in early-stage NSCLC. Treatment-naïve plasma samples were used here to assess initial feasibility. The application of cancer management will be further evaluated in future studies using post-treatment samples and longitudinal samples.
[0226] Table 10. Patient Clinical Demographic Information Manual 52 / 55 pages 63 CN 121464223 A
[0227] Table 11. Multivariate Analysis Manual 53 / 55 pages 64 CN 121464223 A
[0228] Example 8: Different amounts of nucleosome-free cell-free DNA (ncfDNA) derived from FaDu cancer cell lines were generated from the cancer methylome and the fully methylated group to simulate cancer signals and diluted against a background of cfDNA collected from multiple cancer-free donors. Table 12 shows the dilution points and technical replicates (32 samples in total) created for limit of detection (LoD) assessment.
[0229] Table 12. Samples with different dilution points and technical replicates
[0230] As described in Example 1, Figure 25 (Workflow 4), 10 ng of input DNA from each of 32 samples was subjected to a whole-genome DNA methylation enrichment platform based on cell-free methylated DNA immunoprecipitation to generate enriched libraries. After cell-free methylated DNA immunoprecipitation, the enriched libraries were mixed with probes to target regions of interest, as shown in Figure 28. Two target capture pools (Table 13) were created by evenly distributing the enriched libraries generated from each dilution percentage to ensure that the number of technical replicates per dilution percentage was equal in each pool, with a total of 16 samples per pool. Each pool was subjected to a capture process using a probe set targeting the following cancer methylome target groups, followed by the Twist Target Enrichment Standard Hybridization v2 protocol to generate enriched libraries for sequencing on the Illumina next-generation sequencing system. The cancer methylome target groups consist of: 1) housekeeping methylated regions methylated in different samples regardless of cancer status.1) Regions without CpG, used to calculate endogenous binding specificity scores, which are used to quantify nonspecific methylation binding as a quality control metric; 2) Cancer-high differentially methylated regions (DMR); 3) Noise regions; 4) Regions with low counts in cancer-free samples (controls).
[0231] Table 13. Enriched libraries are divided into two target capture pools. Instructions 54 / 55 pages 65 CN 121464223 A
[0232] The target enriched samples were sequenced using an Illumina NovaSeq 6000 sequencer, with a target of approximately 26 million reads per sample. Next, the sequencing reads from the target capture samples were aligned with the hg38 human reference genome using the Bowtie2 alignment tool. A unique molecular identifier (UMI)-based approach was used to identify and remove sequence duplications that may be caused by PCR or flow cytometry optical clustering errors. DNA sequences were further filtered according to size, minimum number of CpGs, and mapping quality, and then counted in the target genome region of interest (ROI).
[0233] Cancer DNA was quantified using DNA fragment counting in the target ROI, and a score was generated by comparison with an internal baseline. Scores from the same ROI were used to compare the cancer methylome approach (which integrates additional enrichment steps with the probe) with the full methylome approach (which does not integrate additional capture of the target region with the probe). Full methylome results were based on previous cell line titration studies and the highest performance score. The same region was used to obtain the cancer methylome score. Next, the fold change was calculated by dividing the full methylome approach LoD by the cancer methylome approach LoD. In both approaches, ctDNA quantification at different dilution points was used to determine the LoD. As shown in Figure 27, the cancer methylome approach showed an improvement in LoD compared to the full methylome approach, as indicated by the fold change, suggesting that integrating additional enrichment steps with the probe can improve the LoD.
[0234] Although preferred embodiments of the invention have been described herein, those skilled in the art will understand that modifications may be made thereto without departing from the spirit of the invention or the scope of the appended claims. All documents disclosed herein, including those in the following list of references, are incorporated herein by reference. Instruction manual, pages 55 / 55; 66 CN 121464223 A, Figure 1; Instruction manual, Figure 1 / 32, page 67 CN 121464223 A, Figure 2; Figure 3; Instruction manual, Figure 2 / 32, page 68 CN 121464223 A, Figure 4A; Figure 4B; Figure 4C; Instruction manual, Figure 3 / 32, page 69 CN 121464223 A, Figure 4D; Figure 4E; Figure 4F; Instruction manual, Figure 4 / 32, page 70 CN 121464223 AFigure 4G Figure 4H Figure 4I Appendix 5 / 32 Page 71 CN 121464223 A Figure 4J Figure 4K Figure 4L Appendix 6 / 32 Page 72 CN 121464223 A Figure 5A Figure 5B Appendix 7 / 32 Page 73 CN 121464223 A Figure 6A Figure 6B Appendix 8 / 32 Page 74 CN 121464223 A Figure 7 Figure 8 Appendix 9 / 32 Page 75 CN 121464223 A Figure 9 Appendix 10 / 32 Page 76 CN 121464223 A Figure 10 Appendix 11 / 32 Page 77 CN 121464223 A Figure 11 Appendix 12 / 32 Page 78 CN 121464223 A Figure 12 Appendix 13 / 32 Page 79 CN 121464223 Figure 12 (continued) Instruction Manual Figure 14 / 32 Page 80 CN 121464223 Figure 13A Figure 13B Instruction Manual Figure 15 / 32 Page 81 CN 121464223 Figure 14 Instruction Manual Figure 16 / 32 Page 82 CN 121464223 Figure 15 Instruction Manual Figure 17 / 32 Page 83 CN 121464223 Figure 16 Instruction Manual Figure 18 / 32 Page 84 CN 121464223 Figure 16 (continued) Instruction Manual Figure 19 / 32 Page 85 CN 121464223 Figure 17 Figure 18 Instruction Manual Figure 20 / 32 Page 86 CN 121464223 Figure 19 Instruction Manual Figure 21 / 32 Page 87 CN 121464223 Figure 20 Instruction Manual Figure 22 / 32 Page 88 CN Figure 21 of the instruction manual, page 23 / 32, CN 121464223 A; Figure 22 of the instruction manual, page 24 / 32, CN 121464223 A; Figure 23 of the instruction manual, page 25 / 32, CN 121464223 A; Figure 24 of the instruction manual, page 26 / 32, CN 121464223 A; Figure 25 of the instruction manual, page 27 / 32, CN 121464223 A; Figure 25 (continued); Figure 28 of the instruction manual, page 28 / 32, CN 121464223 AFigure 26 Appendix to the Instruction Manual, Page 29 / 32, 95 CN 121464223 A Figure 27 Appendix to the Instruction Manual, Page 30 / 32, 96 CN 121464223 A Figure 28 Appendix to the Instruction Manual, Page 31 / 32, 97 CN 121464223 A Figure 28 (continued) Appendix to the Instruction Manual, Page 32 / 32, 98 CN 121464223 A
Claims
1. A method, the method comprising: (a) Obtaining the first nucleic acid molecule from a cell-free sample of the subject; (b) Generating a second group of nucleic acid molecules from the first group of nucleic acid molecules or their derivatives, wherein the second group of nucleic acid molecules is enriched at the methylation level relative to the first group of nucleic acid molecules; (c) Enrich one or more targets from the second group of nucleic acid molecules or their derivatives to obtain the third group of nucleic acid molecules; and (d) Sequencing of the third group of nucleic acid molecules or their derivatives.
2. The method of claim 1, wherein the enrichment comprises contacting the second group of nucleic acid molecules or their derivatives with one or more nucleic acid capture probes.
3. The method according to any one of claims 1 to 2, wherein the generation comprises contacting the first group of nucleic acid molecules or their derivatives with a methylated nucleic acid capture reagent.
4. The method according to claim 3, wherein the methylated nucleic acid capture reagent is formed by incubating the methylated binding molecule together with (ii) a solid matrix.
5. The method according to claim 4, wherein the solid matrix is a bead.
6. The method according to any one of claims 4 to 5, wherein the solid matrix is a magnetic solid matrix.
7. The method according to any one of claims 4 to 6, wherein the solid matrix comprises protein A.
8. The method according to any one of claims 4 to 6, wherein the solid matrix comprises streptavidin.
9. The method according to any one of claims 4 to 8, wherein the methylated binding molecule is an antibody.
10. The method according to any one of claims 4 to 9, wherein the methylated binding molecule comprises biotin.
11. The method according to any one of claims 4 to 10, wherein the methylation-binding molecule binds to methylated cytosine.
12. The method according to any one of claims 4 to 11, wherein the method further comprises amplifying the second group of molecules.
13. The method of claim 12, wherein the amplification is performed simultaneously with the binding of a subgroup of the first group of nucleic acids to the methylated nucleic acid capture reagent.
14. The method according to any one of claims 1 to 13, wherein the sequencing is performed by a synthetic reaction.
15. The method according to any one of claims 1 to 14, wherein the sequencing does not include bisulfite sequencing.
16. The method according to any one of claims 1 to 15, wherein, prior to the sequencing, the third group of nucleic acid molecules undergoes one or more library preparation reactions.
17. The method of claim 16, further comprising, after performing the one or more library preparation reactions and before sequencing, incubating the third group of nucleic acid molecules together with a plurality of magnetic beads that interact with nucleic acids.
18. The method of claim 17, further comprising, after incubating the nucleic acid sample with a plurality of magnetic beads interacting with the nucleic acid, magnetically trapping the sample to remove the plurality of magnetic beads interacting with the nucleic acid.
19. The method of claim 18, further comprising performing additional magnetic trapping.
20. The method according to any one of claims 1 to 19, the method further comprising, prior to (b), adding a certain amount of filler DNA to the first nucleic acid molecule.
21. A method, the method comprising: (a) Provide multiple nucleic acid molecules generated from a cell-free deoxyribonucleic acid (cfDNA) sample of the subject; (b) Sequencing the plurality of nucleic acid molecules or their derivatives to generate multiple sequencing reads; (c) The computer processes the multiple sequencing reads to generate methylation profiles of the multiple nucleic acid molecules; and (d) The computer processes the methylation spectrum to determine that the subject has cancer when the area under the subject operating characteristic curve (AUROC) is at least about 91%, wherein the cancer is a low-shedding cancer.
22. The method of claim 21, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from low-shedding tumors.
23. The method of claim 21, wherein the cancer is bladder cancer, breast cancer, endometrial cancer, prostate cancer, or kidney cancer.
24. The method of claim 23, wherein the cancer is endometrial cancer or prostate cancer.
25. A method, the method comprising: (a) Provide multiple nucleic acid molecules generated from a cell-free deoxyribonucleic acid (cfDNA) sample of the subject; (b) Sequencing the plurality of nucleic acid molecules or their derivatives to generate multiple sequencing reads; (c) The computer processes the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules; and (d) The computer processes the methylation spectrum to determine that the subject has cancer when the area under the subject operating characteristic curve (AUROC) is at least about 94%, and wherein the cancer is early-stage cancer.
26. The method of claim 25, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from early-stage tumors.
27. The method of claim 26, wherein the early tumor is a stage I tumor.
28. The method of claim 26, wherein the early tumor is a stage II tumor.
29. The method according to any one of claims 25 to 28, wherein the cancer is bladder cancer, breast cancer, colorectal cancer, endometrial cancer, esophageal cancer, head and neck cancer, hepatobiliary cancer, lung cancer, ovarian cancer, prostate cancer, or kidney cancer.
30. The method of claim 29, wherein the cancer is esophageal cancer, hepatobiliary cancer, or ovarian cancer.
31. A method, the method comprising: (a) Provide multiple nucleic acid molecules generated from a cell-free deoxyribonucleic acid (cfDNA) sample of the subject; (b) Sequencing the plurality of nucleic acid molecules or their derivatives without bisulfite conversion to generate multiple sequencing reads; (c) The computer processes the plurality of sequencing reads to generate methylation profiles of the plurality of nucleic acid molecules; and (d) The computer processes the methylation spectrum to determine that the subject has cancer, wherein the cancer is endometrial cancer, esophageal cancer, hepatobiliary cancer, ovarian cancer, prostate cancer, bladder cancer, breast cancer, colorectal cancer, head and neck cancer, lung cancer, pancreatic cancer, or kidney cancer.
32. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from endometrial cancer.
33. The method of claim 32, wherein the computer processes the methylation spectrum to determine that the subject has endometrial cancer when the area under the receiver operating characteristic curve (AUROC) is at least about 90%.
34. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from esophageal cancer.
35. The method of claim 34, wherein the computer processes the methylation spectrum to determine that the subject has esophageal cancer when the area under the receiver operating characteristic curve (AUROC) is at least about 99%.
36. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from hepatobiliary cancer.
37. The method of claim 36, wherein the methylation spectrum is processed by a computer to determine that the subject has hepatobiliary cancer when the area under the receiver operating characteristic curve (AUROC) is at least about 99%.
38. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from ovarian cancer.
39. The method of claim 38, wherein the methylation spectrum is processed by a computer to determine that the subject has ovarian cancer when the area under the receiver operating characteristic curve (AUROC) is at least about 97%.
40. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from prostate cancer.
41. The method of claim 40, wherein the methylation spectrum is processed by a computer to determine that the subject has prostate cancer when the area under the receiver operating characteristic curve (AUROC) is at least about 89%.
42. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from bladder cancer.
43. The method of claim 42, wherein the methylation spectrum is processed by a computer to determine that the subject has bladder cancer when the area under the receiver operating characteristic curve (AUROC) is at least about 95%.
44. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from breast cancer.
45. The method of claim 44, wherein the methylation spectrum is processed by a computer to determine that the subject has breast cancer when the area under the receiver operating characteristic curve (AUROC) is at least about 92%.
46. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from colorectal cancer.
47. The method of claim 46, wherein the computer processes the methylation spectrum to determine that the subject has colorectal cancer when the area under the receiver operating characteristic curve (AUROC) is at least about 98%.
48. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from head and neck cancer.
49. The method of claim 48, wherein the computer processes the methylation spectrum to determine that the subject has head and neck cancer when the area under the receiver operating characteristic curve (AUROC) is at least about 96%.
50. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from lung cancer.
51. The method of claim 50, wherein the computer processes the methylation spectrum to determine that the subject has lung cancer when the area under the subject operating characteristic curve (AUROC) is at least about 96%.
52. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from pancreatic cancer.
53. The method of claim 52, wherein the methylation spectrum is processed by a computer to determine that the subject has pancreatic cancer when the area under the receiver operating characteristic curve (AUROC) is at least about 99%.
54. The method of claim 31, wherein the cfDNA sample comprises circulating tumor nucleic acid molecules derived from renal cell carcinoma.
55. The method of claim 54, wherein the computer processes the methylation spectrum to determine that the subject has renal cell carcinoma when the area under the subject operating characteristic curve (AUROC) is at least about 91%.
56. The method according to any one of claims 21 to 55, the method further comprising, prior to (b), adding a group of nucleic acid molecules not from the object to the plurality of nucleic acid molecules.
57. The method according to any one of claims 21 to 56, wherein the methylation profile is genome-wide.
58. The method of claim 57, wherein the methylation spectrum comprises a fully methylated group.
59. The method according to any one of claims 21 to 58, wherein (d) comprises a supervised machine learning method, wherein the supervised machine learning method is regression, support vector machine, tree-based method, neural network or nearest neighbor method.
60. The method according to any one of claims 21 to 58, wherein (d) includes an unsupervised machine learning method, wherein the unsupervised machine learning method is clustering, neural network, principal component analysis or matrix factorization.
61. The method according to any one of claims 21 to 60, wherein the subject has previously received cancer treatment and is substantially free of said cancer, wherein (d) includes determining that the subject has a recurrence of said cancer.
62. The method according to any one of claims 21 to 61, the method further comprising, prior to (b), adding a certain amount of filler DNA to the plurality of nucleic acid molecules or derivatives thereof.
63. The method of claim 62, wherein the filler DNA comprises double-stranded DNA.
64. The method according to any one of claims 62 to 63, wherein the amount of filler DNA is about 20 nanograms (ng) to about 100 ng.
65. The method according to any one of claims 62 to 64, wherein at least a portion of the filler DNA is methylated.
66. The method of claim 65, wherein 10%-40% of the filler DNA is methylated, and the remainder is unmethylated filler DNA.
67. The method according to any one of claims 21 to 66, the method further comprising, prior to (b), contacting the cfDNA sample with a methylated nucleic acid capture reagent to generate the plurality of nucleic acid molecules, wherein the plurality of nucleic acids comprise one or more methylated regions.
68. The method of claim 67, wherein the methylated nucleic acid capture reagent comprises a conjugate and a solid matrix.
69. The method of claim 68, wherein the methylated nucleic acid capture reagent is generated by coupling the conjugate to the solid matrix by incubating the conjugate together with the solid matrix.
70. The method of claim 69, wherein the coupling of the conjugate to the solid matrix is performed prior to the contact of the cfDNA sample with the methylated nucleic acid capture reagent.
71. The method according to any one of claims 68 to 70, wherein the solid matrix is beads.
72. The method according to any one of claims 68 to 70, wherein the solid matrix is protein A beads.
73. The method according to any one of claims 68 to 72, wherein the solid matrix is a magnetic solid matrix.
74. The method according to any one of claims 68 to 73, wherein the conjugate comprises an antibody.
75. The method according to any one of claims 68 to 74, wherein the conjugate is selected from anti-5-methylcytosine antibody or a derivative thereof, anti-5-carboxycytosine antibody or a derivative thereof, anti-5-formylcytosine antibody or a derivative thereof, anti-5-hydroxymethylcytosine antibody or a derivative thereof, anti-3-methylcytosine antibody or a derivative thereof, and any combination thereof.
76. The method according to any one of claims 67 to 75, wherein one or more methylated regions of the plurality of nucleic acids are enriched with at least about 99% specificity.
77. The method according to any one of claims 21 to 76, the method further comprising, prior to (b), amplifying the plurality of nucleic acid molecules to generate an amplicon, wherein (c) includes sequencing the amplicon.
78. The method of claim 77, wherein the amplification of the plurality of nucleic acids is performed simultaneously with the binding of the plurality of nucleic acids to the solid support.
79. The method according to any one of claims 77 to 78, wherein the amplification comprises PCR amplification.
80. The method of claim 79, wherein the PCR amplification comprises at least 13 or at least 14 cycles.
81. The method according to any one of claims 21 to 80, the method further comprising, prior to (b), contacting the plurality of nucleic acid molecules with one or more nucleic acid capture probes to enrich one or more target sequences.
82. The method of claim 81, wherein the one or more target sequences comprise one or more genes.
83. A method for processing a nucleic acid sample from an object, the method comprising: (a) Generate a mixture of nucleic acid samples comprising a plurality of methylated nucleic acids and a plurality of filler nucleic acids from said object, wherein said plurality of filler nucleic acids comprise at least one methylated nucleic acid molecule; (b) Incubate (i) the methylated binding molecule with (ii) a solid matrix to form a methylated nucleic acid capture reagent; (c) The methylated nucleic acids in the nucleic acid sample mixture are enriched by adding the methylated nucleic acid capture reagent to the nucleic acid sample mixture to capture the methylated nucleic acids.
84. The method of claim 83, further comprising, after the capture, amplifying the captured methylated nucleic acid to generate amplicones of the plurality of methylated nucleic acids.
85. The method of claim 84, wherein the amplification is performed simultaneously with the binding of the plurality of methylated nucleic acids to the methylated nucleic acid capture reagent.
86. The method according to claim 84 or 85, wherein the captured methylated nucleic acid is not eluted prior to amplification.
87. The method according to any one of claims 83 to 86, the method further comprising performing a sequencing reaction on the plurality of methylated nucleic acids or their derivatives.
88. The method of claim 87, wherein the sequencing reaction is performed by a synthesis reaction.
89. The method according to claim 87 or 88, wherein the sequencing reaction does not include bisulfite sequencing.
90. The method according to any one of claims 83 to 89, wherein the plurality of methylated nucleic acids comprise cell-free nucleic acids.
91. The method according to any one of claims 83 to 90, wherein the solid matrix is beads.
92. The method according to any one of claims 83 to 91, wherein the solid matrix is a magnetic solid matrix.
93. The method according to any one of claims 83 to 92, wherein the solid matrix comprises protein A.
94. The method according to any one of claims 83 to 93, wherein the solid matrix comprises streptavidin.
95. The method according to any one of claims 83 to 94, wherein the methylation-binding molecule is an antibody.
96. The method according to any one of claims 83 to 95, wherein the methylated binding molecule comprises biotin.
97. The method according to any one of claims 83 to 96, wherein the methylation-binding molecule binds to methylated cytosine.
98. The method according to any one of claims 83 to 97, the method further comprising, prior to (a), obtaining the nucleic acid sample from the object and performing one or more library preparation reactions on the nucleic acid sample.
99. The method of claim 98, further comprising, after performing the one or more library preparation reactions and before (a), incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acid.
100. The method of claim 99, further comprising, after incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acid, magnetically trapping the sample to remove the plurality of magnetic beads that interact with the nucleic acid.
101. The method of claim 100, further comprising performing additional magnetic trapping.
102. The method according to any one of claims 83 to 101, the method further comprising denaturing the nucleic acids in the nucleic acid sample mixture after (a) and before (c).
103. The method according to any one of claims 83 to 102, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched by at least 2-fold.
104. The method according to any one of claims 83 to 103, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched by at least 100-fold.
105. The method according to any one of claims 83 to 104, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched with at least 99% specificity.
106. The method according to any one of claims 83 to 105, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched with at least 99.5% specificity.
107. The method according to any one of claims 83 to 106, the method further comprising contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences.
108. The method of claim 107, wherein the one or more target sequences comprise one or more genes.
109. A method for processing a nucleic acid sample selected from the subject, the method comprising: (a) Provide a nucleic acid sample containing multiple methylated nucleic acids; (b) Incubate (i) the methylated binding molecule with (ii) a solid matrix to form a methylated nucleic acid capture reagent; (c) The plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched by adding the methylated nucleic acid capture reagent to the nucleic acid sample mixture, wherein the plurality of methylated nucleic acids are enriched with a specificity of greater than 99%.
110. The method of claim 109, further comprising, after the capture, amplifying the captured methylated nucleic acid to generate a plurality of amplicones of methylated nucleic acids.
111. The method of claim 110, wherein the amplification is performed simultaneously with the binding of the plurality of methylated nucleic acids to the methylated nucleic acid capture reagent.
112. The method according to claim 109 or 110, wherein the captured methylated nucleic acid is not eluted prior to amplification.
113. The method according to any one of claims 109 or 112, the method further comprising performing a sequencing reaction on the plurality of methylated nucleic acids or their derivatives.
114. The method of claim 113, wherein the sequencing reaction is performed by a synthesis reaction.
115. The method according to claim 113 or 114, wherein the sequencing reaction does not include bisulfite sequencing.
116. The method according to any one of claims 109 to 115, wherein the plurality of methylated nucleic acids comprise cell-free nucleic acids.
117. The method according to any one of claims 109 to 116, wherein the solid matrix is beads.
118. The method according to any one of claims 109 to 117, wherein the solid matrix is a magnetic solid matrix.
119. The method according to any one of claims 109 to 118, wherein the solid matrix comprises protein A.
120. The method according to any one of claims 109 to 119, wherein the solid matrix comprises streptavidin.
121. The method according to any one of claims 109 to 120, wherein the methylation-binding molecule is an antibody.
122. The method according to any one of claims 109 to 121, wherein the methylated binding molecule comprises biotin.
123. The method according to any one of claims 109 to 122, wherein the methylation-binding molecule binds to methylated cytosine.
124. The method according to any one of claims 109 to 123, the method further comprising, prior to (c), performing one or more library preparation reactions on the methylated nucleic acid.
125. The method of claim 124, further comprising, after performing the one or more library preparation reactions and before (c), incubating the nucleic acid sample together with a plurality of magnetic beads that interact with the nucleic acid.
126. The method of claim 125, further comprising, after incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acid, magnetically trapping the sample to remove the plurality of magnetic beads that interact with the nucleic acid.
127. The method of claim 126, further comprising performing additional magnetic trapping.
128. The method according to any one of claims 109 to 127, the method further comprising denaturing the nucleic acid in the nucleic acid sample after (a) and before (c).
129. The method according to any one of claims 109 to 128, wherein the plurality of methylated nucleic acids in the nucleic acid sample are enriched by at least 2-fold.
130. The method according to any one of claims 109 to 129, wherein the plurality of methylated nucleic acids in the nucleic acid sample are enriched by at least 100-fold.
131. The method according to any one of claims 109 to 130, wherein the plurality of methylated nucleic acids in the nucleic acid sample are enriched with at least 99.5% specificity.
132. The method according to any one of claims 109 to 131, the method further comprising contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences.
133. The method of claim 132, wherein the one or more target sequences comprise one or more genes.
134. A method for processing a nucleic acid sample from an object, the method comprising: (a) Provide a nucleic acid sample containing multiple methylated nucleic acids; (b) Incubate (i) the methylated binding molecule with (ii) a solid matrix to form a methylated nucleic acid capture reagent; (c) The methylated nucleic acid is captured by adding the methylated nucleic acid capture reagent to the nucleic acid sample, thereby generating methylated nucleic acid bound to a solid matrix; (d) Amplify the methylated nucleic acids bound to the solid matrix to generate amplicones of the plurality of methylated nucleic acids.
135. The method of claim 134, wherein the amplification is performed simultaneously with the binding of the plurality of methylated nucleic acids to the methylated nucleic acid capture reagent.
136. The method according to claim 134 or 135, wherein the captured methylated nucleic acid is not eluted prior to amplification.
137. The method according to any one of claims 134 to 136, the method further comprising performing a sequencing reaction on the amplicon of the methylated nucleic acid.
138. The method of claim 137, wherein the sequencing reaction is performed via a synthesis reaction.
139. The method according to claim 137 or 138, wherein the sequencing reaction does not include bisulfite sequencing.
140. The method according to any one of claims 134 to 139, wherein the plurality of methylated nucleic acids comprise cell-free nucleic acids.
141. The method according to any one of claims 134 to 140, wherein the solid matrix is beads.
142. The method according to any one of claims 134 to 141, wherein the solid matrix is a magnetic solid matrix.
143. The method according to any one of claims 134 to 142, wherein the solid matrix comprises protein A.
144. The method according to any one of claims 134 to 143, wherein the solid matrix comprises streptavidin.
145. The method according to any one of claims 134 to 144, wherein the methylated binding molecule is an antibody.
146. The method according to any one of claims 134 to 145, wherein the methylated binding molecule comprises biotin.
147. The method according to any one of claims 134 to 146, wherein the methylation-binding molecule binds to methylated cytosine.
148. The method according to any one of claims 134 to 147, wherein, prior to (c), the methylated nucleic acid is subjected to one or more library preparation reactions.
149. The method of claim 148, further comprising, after performing the one or more library preparation reactions and before (c), incubating the nucleic acid sample together with a plurality of magnetic beads that interact with the nucleic acid.
150. The method of claim 149, further comprising, after incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acid, magnetically trapping the sample to remove the plurality of magnetic beads that interact with the nucleic acid.
151. The method of claim 150, further comprising performing additional magnetic trapping.
152. The method according to any one of claims 134 to 151, the method further comprising denaturing the nucleic acid in the nucleic acid sample after (a) and before (c).
153. The method according to any one of claims 134 to 152, wherein the plurality of methylated nucleic acids in the nucleic acid sample are enriched by at least 2-fold.
154. The method according to any one of claims 134 to 153, wherein the plurality of methylated nucleic acids in the nucleic acid sample are enriched by at least 100-fold.
155. The method according to any one of claims 134 to 154, wherein the plurality of methylated nucleic acids in the nucleic acid sample are enriched with at least 99% specificity.
156. The method according to any one of claims 134 to 155, wherein the plurality of methylated nucleic acids in the nucleic acid sample are enriched with at least 99.5% specificity.
157. The method according to any one of claims 134 to 156, the method further comprising contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences.
158. The method of claim 157, wherein the one or more target sequences comprise one or more genes.
159. A method for processing a nucleic acid sample from an object, the method comprising: (a) Generate a mixture of nucleic acid samples comprising a plurality of methylated nucleic acids and a plurality of filler nucleic acids from said object, wherein said plurality of filler nucleic acids comprise at least one methylated nucleic acid molecule; (b) The methylated nucleic acid is captured by adding a capture reagent containing a solid matrix to the nucleic acid sample mixture, thereby generating methylated nucleic acid bound to the solid matrix; and (c) Amplify the methylated nucleic acid bound to the solid matrix to generate an amplicon of the methylated nucleic acid.
160. The method of claim 159, wherein the captured methylated nucleic acid is not eluted prior to amplification.
161. The method according to claim 159 or 160, wherein the method further comprises performing a sequencing reaction on the amplicon of the methylated nucleic acid.
162. The method of claim 161, wherein the sequencing reaction is performed by a synthesis reaction.
163. The method according to claim 161 or 162, wherein the sequencing reaction does not include bisulfite sequencing.
164. The method according to any one of claims 159 to 163, wherein the plurality of methylated nucleic acids comprise cell-free nucleic acids.
165. The method according to any one of claims 159 to 164, wherein the solid matrix is beads.
166. The method according to any one of claims 159 to 165, wherein the solid matrix is a magnetic solid matrix.
167. The method according to any one of claims 159 to 166, wherein the solid matrix comprises protein A.
168. The method according to any one of claims 159 to 167, wherein the solid matrix comprises streptavidin.
169. The method according to any one of claims 159 to 168, wherein the capture reagent is a methylated nucleic acid capture reagent.
170. The method of claim 169, wherein the methylated nucleic acid capture reagent comprises a methylated binding molecule attached to the solid matrix.
171. The method of claim 170, wherein the methylation-binding molecule is an antibody.
172. The method of claim 170, wherein the methylated binding molecule comprises biotin.
173. The method of claim 172, wherein the methylation-binding molecule binds to methylated cytosine.
174. The method according to any one of claims 159 to 173, the method further comprising, prior to (a), obtaining the nucleic acid sample from the object and performing one or more library preparation reactions on the nucleic acid.
175. The method of claim 174, further comprising, after performing the one or more library preparation reactions and before (a), incubating the nucleic acid sample together with a plurality of magnetic beads that interact with the nucleic acid.
176. The method of claim 175, further comprising, after incubating the nucleic acid sample with a plurality of magnetic beads that interact with the nucleic acid, magnetically trapping the sample to remove the plurality of magnetic beads that interact with the nucleic acid.
177. The method of claim 176, further comprising performing additional magnetic trapping.
178. The method according to any one of claims 159 to 177, the method further comprising denaturing the nucleic acids in the nucleic acid sample mixture after (a) and before (b).
179. The method according to any one of claims 159 to 178, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched by at least 2-fold.
180. The method according to any one of claims 159 to 179, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched by at least 100-fold.
181. The method according to any one of claims 159 to 180, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched with at least 99% specificity.
182. The method according to any one of claims 159 to 181, wherein the plurality of methylated nucleic acids in the nucleic acid sample mixture are enriched with at least 99.5% specificity.
183. The method according to any one of claims 159 to 182, the method further comprising contacting the plurality of methylated nucleic acids with one or more nucleic acid capture probes to enrich one or more target sequences.
184. The method of claim 183, wherein the one or more target sequences comprise one or more genes.