detecting a cancer, a tissue of origin of a cancer, and / or a cancer cell type

CN122609688APending Publication Date: 2026-08-21GRAIL INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511463412.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-02-04
Filing Date
2020-02-05
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0009]因此,尚不能获得一种具成本效益的,通过侦测被差异地甲基化的数个区域而准确地诊断一疾病的方法

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122609688A_ABST
    Figure CN122609688A_ABST
Patent Text Reader

Abstract

The present description provides a cancer assay detection panel for targeted detection of methylation patterns specific to various cancer types. Further provided herein include methods of designing, manufacturing, and using the cancer assay detection panel for detecting cancer source tissue (e.g., type of cancer).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of application number 202080025351.1 (PCT application number PCT / US2020 / 016684), filed on February 5, 2020, entitled "Detection of Cancer, Cancer-originating Tissue and / or a Cancer Cell Type".

[0002] Cross-referencing

[0003] This application claims priority from U.S. Provisional Patent Application No. 62 / 801,556, filed February 5, 2019; U.S. Provisional Patent Application No. 62 / 801,561, filed February 5, 2019; U.S. Provisional Patent Application No. 62 / 965,327, filed January 24, 2020; U.S. Provisional Patent Application No. 62 / 965,324, filed January 24, 2020; PCT International Application No. PCT / US2020 / 015082, filed January 24, 2020; and PCT International Application No. PCT / US2020 / 016673, filed February 4, 2020, all of which are incorporated herein by reference in their entirety.

[0004] sequence list

[0005] This application includes a sequence list, which is submitted electronically in ASCII format and is incorporated herein by reference in its entirety. The aforementioned ASCII copy, created on February 3, 2020, is named 50251-852_601_SL.txt and has a size of 27,132,797 bits. Background Technology

[0006] DNA methylation plays a crucial role in regulating gene expression. Abnormal DNA methylation is associated with disease processes in many diseases, including cancer. DNA methylation analysis using methylation sequencing (e.g., whole genome bisulfite sequencing (WGBS)) is increasingly recognized as a valuable diagnostic tool for cancer detection, diagnosis, and / or monitoring. For example, specific patterns of several differentially methylated regions can be useful as molecular markers for various diseases.

[0007] However, WGBS is not ideally suited for a product detection combination. The reason is that the vast majority of the genome is either not differentially methylated in cancer, or the local CpG density is too low to provide a strong signal. Only a few percentages of the genome may be useful in classification.

[0008] Furthermore, identifying several differentially methylated regions across various diseases presents several challenges. First, determining several differentially methylated regions within a disease group is only meaningful when compared to a cohort of several control groups; thus, if the control groups are small in number, the determination loses confidence due to the small size of the control groups. Additionally, methylation states may vary within a cohort of several control groups, which can be difficult to interpret when determining whether the regions are differentially methylated within a disease group. On the other hand, methylation of a cytosine at a CpG site is strongly correlated with methylation at a subsequent CpG site. Generalizing this dependency is itself a challenge.

[0009] Therefore, a cost-effective method for accurately diagnosing a disease by detecting several differentially methylated regions is not yet available. Summary of the Invention

[0010] In this document, several compositions are described in certain embodiments, comprising: several different decoy oligonucleotides, wherein the several different decoy oligonucleotides are configured to collectively hybridize to a DNA molecule derived from at least 100 target genomic regions, and wherein each of the at least 100 target genomic regions is differentially methylated in at least one cancer type compared to another cancer type or compared to a non-cancer type. In some embodiments, the at least 100 target genomic regions include at least one, at least five, at least ten, at least twenty, at least fifty, or at least one hundred target genomic regions that are differentially methylated in at least one first cancer type compared to a second cancer type and compared to a non-cancer type. In some embodiments, the at least 100 target genomic regions include at least one target genomic region that is differentially methylated in the first cancer type compared to two or more, three or more, four or more, five or more, or ten or more, twelve or more, or fifteen or more other cancer types. In some embodiments, the at least 100 target genomic regions are paired with all possible pairs between a cancer type and at least 10, at least 12, at least 15, or at least 18 other cancer types or non-cancer types, including at least one target genomic region that is differentially methylated between pairs of several cancer types.

[0011] In some embodiments, the plurality of decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being derived from at least 20% of the plurality of target genomic regions of any of Lists 1 to 49. In some embodiments, the plurality of decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being derived from at least 20% of the plurality of target genomic regions of Lists 1 to 49. In some embodiments, the plurality of decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being derived from at least 20% or at least 40% of the plurality of target genomic regions of any of Lists 1 to 15. In some embodiments, the plurality of decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being derived from at least 20% or at least 40% of the plurality of target genomic regions of any of Lists 1 to 15. In some embodiments, the plurality of decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being derived from at least 20% of the plurality of target genomic regions of any of Lists 16 to 32. In some embodiments, the plurality of decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being derived from at least 20% of the plurality of target genomic regions listed in Lists 16 to 32. In some embodiments, the plurality of decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being derived from at least 20% of the plurality of target genomic regions listed in any of Lists 33 to 49. In some embodiments, the plurality of decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being derived from at least 20% of the plurality of target genomic regions listed in Lists 33 to 49.

[0012] In this document, several compositions are described in particular embodiments, comprising: several different decoy oligonucleotides configured to hybridize to several DNA molecules, said DNA molecules being derived from at least 20% of the several target genomic regions of any of Lists 1 to 49. In some embodiments, the several decoy oligonucleotides are configured to hybridize to several DNA molecules, said DNA molecules being derived from at least 20% of the several target genomic regions of Lists 1 to 49. In some embodiments, the several decoy oligonucleotides are configured to hybridize to several DNA molecules, said DNA molecules being derived from at least 20% or at least 40% of the several target genomic regions of any of Lists 1 to 15. In some embodiments, the several decoy oligonucleotides are configured to hybridize to several DNA molecules, said DNA molecules being derived from at least 20% or at least 40% of the several target genomic regions of Lists 1 to 15. In some embodiments, the plurality of decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of any of Lists 16 to 32. In some embodiments, the plurality of decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of any of Lists 16 to 32. In some embodiments, the plurality of decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of any of Lists 33 to 49. In some embodiments, the plurality of decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of Lists 33 to 49.

[0013] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being derived from at least 20% of the plurality of target genomic regions of List 1. In some embodiments, the plurality of DNA molecules are derived from at least 30%, 40%, 50%, 60%, 70%, or 80% of the plurality of target genomic regions of List 1.

[0014] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 2. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 2.

[0015] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 3. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 3.

[0016] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 4. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 4.

[0017] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 5. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 5.

[0018] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 6. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 6.

[0019] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 7. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 7.

[0020] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 8. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 8.

[0021] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 9. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 9.

[0022] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 10. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 10.

[0023] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 11. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 11.

[0024] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 12. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 12.

[0025] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 13. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 13.

[0026] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 14. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 14.

[0027] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 15. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 15.

[0028] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 16. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 16.

[0029] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 17. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 17.

[0030] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 18. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 18.

[0031] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 19. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 19.

[0032] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 20. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 20.

[0033] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 21. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 21.

[0034] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 22. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 22.

[0035] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 23. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 23.

[0036] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 24. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 24.

[0037] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 25. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 25.

[0038] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 26. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 26.

[0039] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 27. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 27.

[0040] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 28. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 28.

[0041] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 29. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 29.

[0042] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 30. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 30.

[0043] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 31. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 31.

[0044] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 32. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 32.

[0045] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 33. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 33.

[0046] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 34. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 34.

[0047] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 35. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 35.

[0048] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 36. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 36.

[0049] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 37. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 37.

[0050] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 38. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 38.

[0051] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 39. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 39.

[0052] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 40. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 40.

[0053] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 41. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 41.

[0054] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 42. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 42.

[0055] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 43. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 43.

[0056] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 44. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 44.

[0057] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 45. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 45.

[0058] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 46. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 46.

[0059] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 47. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 47.

[0060] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 48. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 48.

[0061] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from the plurality of target genomic regions of List 49. In some embodiments, the plurality of DNA molecules are at least 30%, 40%, 50%, 60%, 70%, or 80% derived from the plurality of target genomic regions of List 49.

[0062] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% of the plurality of target genomic regions derived from any two or more, three or more, four or more, or five or more of the lists 16 to 32.

[0063] In some embodiments, the plurality of DNA molecules are derived from at least 30%, 40%, 50%, 60%, 70%, or 80% of the plurality of target genomic regions of any two or more, three or more, four or more, or five or more, six or more, seven or more, eight or more, nine or more, or ten or more of the lists 16 to 32.

[0064] In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% of the plurality of target genomic regions derived from any two or more, three or more, four or more, or five or more of the lists 33 to 49.

[0065] In some embodiments, the plurality of DNA molecules are derived from at least 30%, 40%, 50%, 60%, 70%, or 80% of the plurality of target genomic regions of any two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more of the lists 33 to 49.

[0066] In some embodiments, the total size of the plurality of target genomic regions is less than 1100kb, less than 750kb, less than 270kb, less than 200kb, less than 150kb, less than 100kb, or less than 50kb. In some embodiments, the total number of the plurality of target genomic regions is less than 1700, less than 1300, less than 900, less than 700, or less than 400.

[0067] In some embodiments, the total size of the plurality of target genomic regions is less than 5000kb, less than 2500kb, less than 2000kb, less than 1500kb, less than 1000kb, less than 750kb, or less than 500kb. In some embodiments, the total number of the plurality of target genomic regions is less than 20000, less than 18000, less than 16000, less than 14000, less than 12000, less than 10000, less than 8000, less than 6000, less than 4000, or less than 2000.

[0068] In some embodiments, the plurality of DNA molecules are converted cfDNA fragments. In some embodiments, the plurality of target genomic regions are hypermethylated regions, hypomethylated regions, or binary regions that may be either hypermethylated or hypomethylated, as indicated in the sequence listing. In some embodiments, the plurality of decoy oligonucleotides are configured to hybridize to hypermethylated converted DNA molecules, hypomethylated converted DNA molecules, or both hypermethylated and hypomethylated converted DNA molecules derived from each target genomic region, as indicated in the sequence listing.

[0069] In some embodiments, each of the plurality of decoy oligonucleotides is bound to an affinity moiety. In some embodiments, the affinity moiety is biotin. In some embodiments, each of the plurality of decoy oligonucleotides is bound to a solid surface. In some embodiments, the solid surface is a microarray or chip.

[0070] In some embodiments, each of the plurality of decoy oligonucleotides has a length of 45 to 300 nucleotide bases, 75 to 200 nucleotide bases, 100 to 150 nucleotide bases, or about 120 nucleotide bases. In some embodiments, the plurality of decoy oligonucleotides comprises: groups of two or more decoy oligonucleotides, wherein each decoy oligonucleotide in one group of the plurality of decoy oligonucleotides is configured to bind to the same converted target genomic region or configured to bind to a nucleic acid molecule derived from said target genomic region. In some embodiments, each group of decoy oligonucleotides includes one or more pairs of a first decoy oligonucleotide and a second decoy oligonucleotide, wherein each decoy oligonucleotide includes a 5' end and a 3' end, wherein a sequence of at least X nucleotide bases at the 3' end of the first decoy oligonucleotide is identical to a sequence of X nucleotide bases at the 5' end of the second decoy oligonucleotide, and wherein X is at least 25, 30, 35, 40, 45, 50, 60, 70, 75, or 100. In some embodiments, the first decoy oligonucleotide includes a sequence of at least 31, 40, 50, or 60 nucleotide bases, the sequence of which does not overlap with a sequence of the second decoy oligonucleotide.

[0071] In some embodiments, the composition further comprises: converted cfDNA from a test subject. In some embodiments, the cfDNA from the test subject is converted by a process including treatment with bisulfite or a cytosine deaminase.

[0072] In this document, several methods for enriching cfDNA fragments that can provide information about a type of cancer are described in several specific embodiments. The methods include the steps of: contacting any of several bait oligonucleotide compositions described herein with DNA derived from a test subject; and enriching a sample of cfDNA corresponding to several genomic regions associated with the type of cancer by heterozygous capture.

[0073] In this document, several methods for obtaining sequence information that can provide information on the presence or absence of a type of cancer are described in several specific embodiments. The methods include the steps of: (a) enriching the converted DNA by contacting it with any of several bait oligonucleotide compositions described herein; and (b) sequencing the enriched converted DNA.

[0074] In this document, several methods for determining that a subject has a type of cancer are described in several specific embodiments, the methods comprising the steps of: (a) capturing several cfDNA fragments from the subject with any of several decoy oligonucleotide compositions described herein; (b) sequencing the captured several cfDNA fragments; and (c) applying a trained classifier to the several cfDNA sequences to determine that the subject has the type of cancer.

[0075] In this document, a method for determining that a subject has a type of cancer is described in several specific embodiments, the method comprising the steps of: (a) capturing several cfDNA fragments from the subject with any of several decoy oligonucleotide compositions described herein; (b) detecting the captured several cfDNA fragments by a DNA microarray; and (c) applying a trained classifier to the several DNA fragments hybridized to the DNA microarray to determine that the subject has the type of cancer.

[0076] In some embodiments, the trained classifier is a hybrid model classifier. In some embodiments, the classifier is trained on a plurality of converted DNA sequences derived from at least 1,000, at least 2,000, or at least 4,000 target genomic regions selected from any of Lists 1 to 49.

[0077] In some embodiments, the trained classifier determines the presence or absence of cancer, or a cancer type, by: (i) generating a set of multiple features for a sample, wherein each feature in the set of multiple features includes a numerical value; (ii) inputting the set of multiple features into the classifier, wherein the classifier includes a multinomial classifier; (iii) determining a set of probability scores in the classifier based on the set of multiple features, wherein the set of probability scores includes a probability score for each cancer type category and each non-cancer type category; and (iv) measuring the set of probability scores with a threshold based on one or more values ​​determined during the training of the classifier to determine a final cancer classification for the sample. In some embodiments, the set of multiple features includes a set of binary features. In some embodiments, the numerical value includes a single binary value. In some embodiments, the multinomial classifier includes a multinomial logistic regression ensemble trained to predict a source tissue for the cancer.

[0078] In some embodiments, the method further includes the steps of: determining the final cancer classification relative to a minimum value, based on the difference between the two highest probability scores, wherein the minimum value corresponds to a predefined percentage of training cancer samples, the predefined percentage of training cancer samples being assigned the correct cancer type as the highest score during the training of the classifier. In some embodiments, (i) assigning a cancer label as the final cancer classification based on determining that the difference between the two highest probability scores exceeds the minimum value, the cancer label corresponding to the highest probability score determined by the classifier; and (ii) assigning an indeterminate cancer label as the final cancer classification based on determining that the difference between the two highest probability scores does not exceed the minimum value. In some embodiments, the cancer type is selected from the group consisting of anorectal cancer, bladder cancer, bladder and urethral epithelial cancer, breast cancer, cervical cancer, colorectal cancer, head and neck cancer, hepatobiliary cancer, liver and bile duct cancer, lung cancer, melanoma, ovarian cancer, pancreatic cancer, pancreatic and gallbladder cancer, prostate cancer, kidney cancer, sarcoma, thyroid cancer, upper gastrointestinal cancer, and uterine cancer. In some embodiments, the captured cfDNA fragments are several converted cfDNA fragments.

[0079] In this document, several cancer assay combinations are described in specific embodiments, comprising: at least five pairs of probes, each of the at least five pairs comprising: two probes configured to overlap each other by an overlapping sequence comprising a sequence of at least 30 nucleotides, wherein the at least 30 nucleotides sequence is configured to hybridize to a converted cfDNA molecule, the converted cfDNA molecule corresponding to or derived from one or more genomic regions, each of the genomic regions comprising at least five methylation sites, wherein the at least five methylation sites have an aberrant methylation pattern in several first cancer samples, and wherein each of the at least five pairs of probes comprises a non-overlapping sequence of at least 31 nucleotides. In some embodiments, the cancer assay combinations comprise at least 10 pairs, at least 20 pairs, at least 30 pairs, at least 50 pairs, at least 100 pairs, at least 200 pairs, or at least 500 pairs of probes.

[0080] In some embodiments, the plurality of genomic regions are selected from a list, and: the list is list 1, and the plurality of first cancer samples are samples from subjects with bladder cancer; the list is list 2, and the plurality of first cancer samples are samples from subjects with breast cancer; the list is list 3, and the plurality of first cancer samples are samples from subjects with cervical cancer; the list is list 4, and the plurality of first cancer samples are samples from subjects with colorectal cancer; the list is list 5, and the plurality of first cancer samples are samples from subjects with head and neck cancer; the list is list 6, and the plurality of first cancer samples are samples from subjects with hepatobiliary cancer; the list is list 7, and the plurality of first cancer samples are samples from subjects with lung cancer; the list is list 8, and... The plurality of first cancer samples are samples from subjects with melanoma; the list is list 9, and the plurality of first cancer samples are samples from subjects with ovarian cancer; the list is list 10, and the plurality of first cancer samples are samples from subjects with pancreatic cancer; the list is list 11, and the plurality of first cancer samples are samples from subjects with prostate cancer; the list is list 12, and the plurality of first cancer samples are samples from subjects with kidney cancer; the list is list 13, and the plurality of first cancer samples are samples from subjects with thyroid cancer; the list is list 14, and the plurality of first cancer samples are samples from subjects with upper gastrointestinal cancer; or the list is list 15, and the plurality of first cancer samples are samples from subjects with uterine cancer.

[0081] In some embodiments, the plurality of genomic regions are selected from a list, wherein: the list is list 16 or list 33, and the plurality of first cancerous samples are samples from subjects with anorectal cancer; the list is list 17 or list 34, and the plurality of first cancerous samples are samples from subjects with bladder or urethral epithelial cancer; the list is list 18 or list 35, and the plurality of first cancerous samples are samples from subjects with breast cancer; the list is list 19 or list 36, and the plurality of first cancerous samples are samples from subjects with... Several samples from subjects with cervical cancer; the list is list 20 or list 37, and the several first cancer samples are from subjects with colorectal cancer; the list is list 21 or list 38, and the several first cancer samples are from subjects with head and neck cancer; the list is list 22 or list 39, and the several first cancer samples are from subjects with liver or bile duct cancer; the list is list 23 or list 40, and the several first cancer samples are from subjects with lung cancer; the list is a list Table 24 or List 41, wherein the plurality of first cancerous samples are from subjects with melanoma; the list is List 25 or List 42, wherein the plurality of first cancerous samples are from subjects with ovarian cancer; the list is List 26 or List 43, wherein the plurality of first cancerous samples are from subjects with pancreatic or gallbladder cancer; the list is List 27 or List 44, wherein the plurality of first cancerous samples are from subjects with prostate cancer; the list is List 28 or List 45, wherein the plurality of first cancerous samples are from subjects with prostate cancer. The sex samples are several samples from subjects with kidney cancer; or the list is list 29 or list 46, and the several first cancer samples are several samples from subjects with sarcoma; the list is list 30 or list 47, and the several first cancer samples are several samples from subjects with thyroid cancer; the list is list 31 or list 48, and the several first cancer samples are several samples from subjects with upper gastrointestinal cancer; or the list is list 32 or list 49, and the several first cancer samples are several samples from subjects with uterine cancer.

[0082] In some embodiments, the plurality of genomic regions comprises at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the plurality of genomic regions in the list. In some embodiments, the plurality of genomic regions comprises at least 30, 53, 103, 159, 160, 200, 250, 300, 400, 500, 600, 800, or 1000 genomic regions in the list. In some embodiments, the converted cfDNA molecule comprises a cfDNA molecule treated to convert unmethylated C (cytosine) to U (uracil). In some embodiments, each of the at least 5 pairs of probes is bound to a nonnucleotide affinity moiety. In some embodiments, the nonnucleotide affinity moiety is a biotin moiety. In some embodiments, the aberrant methylation pattern has at least a threshold p-value rarity in the plurality of first cancerous samples. In some embodiments, each of the plurality of probes is designed to have sequence homology or sequence complementarity with fewer than 20 off-target genomic regions. In some embodiments, the fewer than 20 off-target genomic regions are identified using a k-mer seeding strategy. In some embodiments, the fewer than 20 off-target genomic regions are identified by binding to local alignment at several seed sites using a k-mer seeding strategy. In some embodiments, each of the plurality of probes comprises at least 61, 75, 100, 120, or 121 nucleotides. In some embodiments, each of the plurality of probes comprises fewer than 300, 250, 200, 160, or 159 nucleotides. In some embodiments, each of the plurality of probes comprises 100 to 159 or 100 to 160 nucleotides. In some embodiments, each of the plurality of probes comprises fewer than 20, 15, 10, 8, or 6 methylation sites. In some embodiments, at least 80, 85, 90, 92, 95, or 98% of the at least five methylation sites are either methylated or unmethylated in the plurality of cancerous samples. In some embodiments, at least 3%, 5%, 10%, 15%, or 20% of the plurality of probes do not include G (guanine). In some embodiments, each of the plurality of probes includes multiple binding sites for the plurality of methylation sites of the converted cfDNA molecule, wherein at least 80, 85, 90, 92, 95, or 98% of the plurality of binding sites include only CpG or CpA. In some embodiments, each of the plurality of probes is configured to have sequence homology or sequence complementarity with fewer than 15, 10, or 8 off-target genomic regions.

[0083] In some embodiments, at least 30% of the plurality of genomic regions are located in exons or introns. In some embodiments, at least 15% of the plurality of genomic regions are located in exons. In some embodiments, at least 20% of the plurality of genomic regions are located in exons. In some embodiments, less than 10% of the plurality of genomic regions are located in intergenic regions. In some embodiments, the cancer testing kit includes at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1200, 1400, 1600, 1800, 2000, 2200, 2400, 2600, 2800, 3000, 3200, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 10000, 15000, or 20000 probes. In some embodiments, the at least 5 pairs of probes comprise at least 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 120,000, 140,000, 160,000, 180,000, 200,000, 240,000, and 260,000 probes in total. 1, 280,000, 300,000, 320,000, 400,000, 450,000, 500,000, 550,000, 600,000, 650,000, 700,000, 750,000, 800,000, 850,000, 900,000, 1 million, 1.5 million, 2 million, 2.5 million or 3 million nucleotides.

[0084] In this document, a method for detecting cancer and / or a tissue of cancer origin (TOO) is described in several specific embodiments, the method comprising the steps of: (a) receiving a sample comprising a plurality of cfDNA molecules; (b) processing the plurality of cfDNA molecules to convert unmethylated C (cytosine) to U (uracil), thereby obtaining a plurality of converted cfDNA molecules; (c) applying any of a plurality of cancer assay combinations described herein to the plurality of converted cfDNA molecules to enrich a subset of the plurality of converted cfDNA molecules; and (d) sequencing the enriched subset of the converted cfDNA molecules to provide a set of sequence reads.

[0085] In this document, a method for detecting cancer and / or a tissue of cancer origin (TOO) is described in several specific embodiments, the method comprising the steps of: (a) receiving a sample comprising a plurality of cfDNA molecules; (b) processing the plurality of cfDNA molecules to convert unmethylated C (cytosine) to U (uracil), thereby obtaining a plurality of converted cfDNA molecules; (c) applying any of the plurality of cancer assay combinations described herein to the plurality of converted cfDNA molecules to enrich a subset of the plurality of converted cfDNA molecules; and (d) detecting the enriched subset of the converted cfDNA molecules by hybridization to a DNA microarray.

[0086] In some embodiments, the method further includes the steps of: determining a health condition by evaluating the set of sequence readings, wherein the health condition is (a) the presence or absence of cancer; (b) a stage of cancer; (c) the presence or absence of tissue of cancer origin (TOO); (d) the presence or absence of a cancer cell type; or (e) the presence or absence of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 different types of cancer. In some embodiments, the sample comprising several cfDNA molecules is obtained from a human subject.

[0087] In this document, several methods for detecting a cancer are described in specific embodiments, the methods comprising the steps of: (a) obtaining a set of sequence reads by sequencing a set of nucleic acid fragments from a subject, wherein each of the nucleic acid fragments corresponds to or is derived from a genomic region selected from one or more lists 1 to 15; one or more lists 16 to 32; or one or more lists 33 to 49; (b) determining the methylation status at several CpG sites for each of the sequence reads; and ( c) Determining that cancer is detected in the subject by assessing the methylation status of the plurality of sequence readings, wherein the detection of cancer includes one or more of the following: (i) the presence or absence of a cancer; (ii) a stage of cancer; (iii) the presence or absence of a cancer-originating tissue (TOO); (iv) the presence or absence of a cancer cell type; or (v) the presence or absence of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 different types of cancer.

[0088] In some embodiments, (a) the plurality of genomic regions are selected from List 1, and the cancer detection includes a detection of bladder cancer; (b) the plurality of genomic regions are selected from List 2, and the cancer detection includes a detection of breast cancer; (c) the plurality of genomic regions are selected from List 3, and the cancer detection includes a detection of cervical cancer; (d) the plurality of genomic regions are selected from List 4, and the cancer detection includes a detection of rectal cancer; (e) the plurality of genomic regions are selected from List 5, and the cancer detection includes a detection of head and neck cancer; (f) the plurality of genomic regions are selected from List 6, and the cancer detection includes a detection of hepatobiliary cancer; (g) the plurality of genomic regions are selected from List 7, and the cancer detection includes a detection of lung cancer; (h) the plurality of genomic regions are selected from List 8. The cancer detection includes a detection of melanoma; (i) the plurality of genomic regions are selected from List 9, and the cancer detection includes a detection of ovarian cancer; (j) the plurality of genomic regions are selected from List 10, and the cancer detection includes a detection of pancreatic cancer; (k) the plurality of genomic regions are selected from List 11, and the cancer detection includes a detection of prostate cancer; (i) the plurality of genomic regions are selected from List 12, and the cancer detection includes a detection of kidney cancer; (m) the plurality of genomic regions are selected from List 13, and the cancer detection includes a detection of thyroid cancer; (n) the plurality of genomic regions are selected from List 14, and the cancer detection includes a detection of upper gastrointestinal cancer; or (o) the plurality of genomic regions are selected from List 15, and the cancer detection includes a detection of uterine cancer.

[0089] In some embodiments, (a) the plurality of genomic regions are selected from List 16 or List 33, and the cancer detection includes a detection of anorectal cancer; the plurality of genomic regions are selected from List 17 or List 34, and the cancer detection includes a detection of bladder or urethral epithelial cancer; the plurality of genomic regions are selected from List 18 or List 35, and the cancer detection includes a detection of breast cancer; the plurality of genomic regions are selected from List 19 or List 36, and the cancer detection includes a detection of cervical cancer. The plurality of genomic regions are selected from List 20 or List 37, and the cancer detection includes a detection for colorectal cancer; the plurality of genomic regions are selected from List 21 or List 38, and the cancer detection includes a detection for head and neck cancer; the plurality of genomic regions are selected from List 22 or List 39, and the cancer detection includes a detection for liver or bile duct cancer; the plurality of genomic regions are selected from List 23 or List 40, and the cancer detection includes a detection for lung cancer; the plurality of genomic regions are selected from List 20 or List 37. Table 24 or List 41, and the cancer detection includes a detection for melanoma; the plurality of genomic regions are selected from List 25 or List 42, and the cancer detection includes a detection for ovarian cancer; the plurality of genomic regions are selected from List 26 or List 43, and the cancer detection includes a detection for pancreatic or gallbladder cancer; the plurality of genomic regions are selected from List 27 or List 44, and the cancer detection includes a detection for prostate cancer; the plurality of genomic regions are selected from List 28 or List 45, and the cancer... The detection includes a detection of kidney cancer; the plurality of genomic regions are selected from List 29 or List 46, and the cancer detection includes a detection of sarcoma; the plurality of genomic regions are selected from List 30 or List 47, and the cancer detection includes a detection of thyroid cancer; the plurality of genomic regions are selected from List 31 or List 48, and the cancer detection includes a detection of upper gastrointestinal cancer; or the plurality of genomic regions are selected from List 32 or List 49, and the cancer detection includes a detection of uterine cancer.

[0090] In some embodiments, the plurality of genomic regions comprises at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the plurality of genomic regions in the list. In some embodiments, the plurality of genomic regions comprises at least 30, 50, 100, 150, 200, 250, or 300 genomic regions in the list. In some embodiments, the plurality of genomic regions comprises less than 90%, 80%, 70%, 60%, 50%, 40%, 30%, or 20% of the genomic regions in the list. In some embodiments, the plurality of genomic regions comprises less than 25,000, 20,000, 15,000, 10,000, 7,500, 5,000, or 2,500 genomic regions in the list. In some embodiments, the plurality of genomic regions includes fewer than 1,000, 500, 400, 300, 200, or 100 genomic regions from the list.

[0091] In this document, several cancer assay combinations comprising several probes are described in particular embodiments. Each of the probes is configured to hybridize to a converted cfDNA molecule corresponding to several genomic regions selected from one or more of Lists 1 to 15. In some embodiments, the converted cfDNA molecule comprises several cfDNA molecules processed to convert unmethylated cytosine to uracil. In some embodiments, the probes are configured to hybridize to several nucleic acid molecules corresponding to or derived from at least 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, or 100% of the several genomic regions in a list, and the list is one or more of Lists 1 to 15. In some embodiments, the plurality of probes are configured to hybridize to a plurality of nucleic acid molecules corresponding to or derived from at least 30, 50, 100, 159, 171, 200, 250, 300, 400, 500, 600, 800, or 1000 genomic regions from a list of 1 to 15. In some embodiments, at least 3%, 5%, 10%, 15%, or 20% of the plurality of probes excludes G (guanine). In some embodiments, each of the plurality of probes includes a plurality of binding sites that bind to a plurality of methylation sites of the converted cfDNA molecule, wherein at least 80%, 85%, 90%, 92%, 95%, or 98% of the plurality of binding sites comprise only CpG or CpA. In some embodiments, each of the plurality of probes binds to a nonnucleotide affinity moiety. In some embodiments, the nonnucleotide affinity moiety is a biotin moiety.

[0092] In this document, several methods for determining the presence or absence of cancer in a subject are described in several specific embodiments. The methods include the steps of: (i) capturing several cfDNA fragments from the subject with a composition comprising several different oligonucleotide decoys; (ii) sequencing the captured cfDNA fragments; and (iii) applying a trained classifier to the cfDNA sequences to determine the presence or absence of cancer. In some embodiments, the probability of a false positive determination of the presence or absence of cancer is less than 1%, and the probability of an accurate determination of the presence or absence of cancer is at least 40%. In some embodiments, the cancer is a stage I cancer, the probability of a false positive determination of the presence or absence of cancer is less than 1%, and the probability of an accurate determination of the presence or absence of cancer is at least 9%. In some embodiments, the several cfDNA fragments are converted cfDNA fragments.

[0093] In this document, a method for detecting a cancer type is described in several specific embodiments, the method comprising the steps of: (i) capturing several cfDNA fragments from a subject with a composition comprising several different oligonucleotide decoys; (ii) sequencing the captured several cfDNA fragments; and (iii) applying a trained classifier to the several cfDNA sequences to determine a cancer type; wherein the several oligonucleotide decoys are configured to hybridize to several cfDNA fragments derived from several target genomic regions; wherein the several target genomic regions are differentially methylated in one or more cancer types compared to a different cancer type or a non-cancer type; wherein the probability of a false positive determination of cancer is less than 1%; and wherein the probability of an accurate identification of a cancer type is at least 75%, at least 80%, at least 85%, or at least 89%, or at least 90%. In some embodiments, the method further comprises the step of applying a trained classifier to the several cfDNA sequences to determine the presence of cancer before determining the cancer type. In some embodiments, the plurality of cff)NA fragments are converted cff)NA fragments.

[0094] In some embodiments, the cancer type is selected from uterine cancer, upper gastrointestinal squamous cell carcinoma, all other upper gastrointestinal cancers, thyroid cancer, sarcoma, urothelial renal cell carcinoma, all other renal cancers, prostate cancer, pancreatic cancer, ovarian cancer, neuroendocrine carcinoma, multiple myeloma, melanoma, lymphoma, small cell lung cancer, lung adenocarcinoma, all other lung cancers, leukemia, hepatocellular carcinoma, hepatobiliary cancer, head and neck cancer, colorectal cancer, cervical cancer, breast cancer, bladder cancer, and anorectal cancer. In some embodiments, the cancer type is selected from anal cancer, bladder cancer, colorectal cancer, esophageal cancer, head and neck cancer, hepatobiliary / biliary duct cancer, lung cancer, lymphoma, ovarian cancer, pancreatic cancer, plasmacytoma, and gastric cancer. In some embodiments, the cancer type is selected from thyroid cancer, melanoma, sarcoma, myeloma, kidney cancer, prostate cancer, breast cancer, uterine cancer, ovarian cancer, bladder cancer, urethral cancer, cervical cancer, anorectal cancer, head and neck cancer, colorectal cancer, liver cancer, bile duct cancer, pancreatic cancer, gallbladder cancer, upper gastrointestinal cancer, multiple myeloma, lymphoma, and lung cancer.

[0095] In some embodiments, the cancer type is a stage I cancer type, and the probability of an accurate designation is at least 70% or at least 75%. In some embodiments, the cancer type is a stage II cancer type, and the probability of an accurate designation is at least 85%.

[0096] In some embodiments, the cancer type is anorectal cancer, the plurality of target genomic regions are selected from List 16 or 33, and the accuracy of detecting anorectal cancer in several samples with detected cancer is at least 80% or 88%. In some embodiments, the cancer type is stage I or stage II anorectal cancer, the plurality of target genomic regions are selected from List 16 or 33, and the accuracy of detecting stage I or stage II anorectal cancer in several samples with detected cancer is at least 75% or 85%.

[0097] In some embodiments, the cancer type is bladder and urethral epithelial carcinoma, the plurality of target genomic regions are selected from List 1, 17, or 34, and the accuracy of detecting bladder and urethral epithelial carcinoma in several samples with detected cancer is at least 80% or 90%. In some embodiments, the cancer type is stage I or stage II bladder and urethral epithelial carcinoma, the plurality of target genomic regions are selected from List 1, 17, or 34, and the accuracy of detecting stage I or stage II bladder and urethral epithelial carcinoma in several samples with detected cancer is at least 75% or 85%.

[0098] In some embodiments, the cancer type is breast cancer, the plurality of target genomic regions are selected from List 2, 18, or 35, and the accuracy of detecting breast cancer in several samples with detected cancer is at least 80% or 88%. In some embodiments, the cancer type is stage I or stage II breast cancer, the plurality of target genomic regions are selected from List 2, 18, or 35, and the accuracy of detecting stage I or stage II breast cancer in several samples with detected cancer is at least 75% or 84%.

[0099] In some embodiments, the cancer type is cervical cancer, the plurality of target genomic regions are selected from List 3, 19, or 36, and the accuracy of detecting cervical cancer in several samples with detected cancer is at least 80% or 88%. In some embodiments, the cancer type is stage I or stage II cervical cancer, the plurality of target genomic regions are selected from List 3, 19, or 36, and the accuracy of detecting stage I or stage II cervical cancer in several samples with detected cancer is at least 75% or 85%.

[0100] In some embodiments, the cancer type is colorectal cancer, the plurality of target genomic regions are selected from List 4, 20, or 37, and the accuracy of detecting colorectal cancer in several samples with detected cancer is at least 80% or 88%. In some embodiments, the cancer type is stage I or stage II colorectal cancer, the plurality of target genomic regions are selected from List 4, 20, or 37, and the accuracy of detecting stage I or stage II colorectal cancer in several samples with detected cancer is at least 75% or 85%.

[0101] In some embodiments, the cancer type is head and neck cancer, the plurality of target genomic regions are selected from List 5, 21, or 38, and the accuracy of detecting head and neck cancer in several samples with detected cancer is at least 80% or 88%. In some embodiments, the cancer type is stage I or stage II head and neck cancer, the plurality of target genomic regions are selected from List 5, 21, or 38, and the accuracy of detecting stage I or stage II head and neck cancer in several samples with detected cancer is at least 75% or 85%.

[0102] In some embodiments, the cancer type is liver and bile duct cancer, the plurality of target genomic regions are selected from List 6, 22, or 39, and the accuracy of detecting liver and bile duct cancer in several samples with detected cancer is at least 80% or 88%. In some embodiments, the cancer type is stage I or stage II liver and bile duct cancer, the plurality of target genomic regions are selected from List 6, 22, or 39, and the accuracy of detecting stage I or stage II liver and bile duct cancer in several samples with detected cancer is at least 75% or 85%.

[0103] In some embodiments, the cancer type is lung cancer, the plurality of target genomic regions are selected from List 7, 23, or 40, and the accuracy of detecting lung cancer in several samples with detected cancer is at least 80% or 88%. In some embodiments, the cancer type is stage I or stage II lung cancer, the plurality of target genomic regions are selected from List 7, 23, or 40, and the accuracy of detecting stage I or stage II lung cancer in several samples with detected cancer is at least 75% or 85%.

[0104] In some embodiments, the cancer type is melanoma, the plurality of target genomic regions are selected from List 8, 24, or 41, and the accuracy of detecting melanoma in several samples with detected cancer is at least 80% or 88%. In some embodiments, the cancer type is stage I or stage II melanoma, the plurality of target genomic regions are selected from List 8, 24, or 41, and the accuracy of detecting stage I or stage II melanoma in several samples with detected cancer is at least 75% or 84%.

[0105] In some embodiments, the cancer type is ovarian cancer, the plurality of target genomic regions are selected from List 9, 25, or 42, and the accuracy of detecting ovarian cancer in several samples with detected cancer is at least 80% or 88%. In some embodiments, the cancer type is stage I or stage II ovarian cancer, the plurality of target genomic regions are selected from List 9, 25, or 42, and the accuracy of detecting stage I or stage II ovarian cancer in several samples with detected cancer is at least 75% or 85%.

[0106] In some embodiments, the cancer type is pancreatic and gallbladder cancer, the plurality of target genomic regions are selected from List 10, 26, or 43, and the accuracy of detecting pancreatic and gallbladder cancer in several samples with detected cancer is at least 80% or 88%. In some embodiments, the cancer type is stage I or stage II pancreatic and gallbladder cancer, the plurality of target genomic regions are selected from List 10, 26, or 43, and the accuracy of detecting stage I or stage II pancreatic and gallbladder cancer in several samples with detected cancer is at least 75%, 81%, or 83%.

[0107] In some embodiments, the cancer type is prostate cancer, the plurality of target genomic regions are selected from List 11, 27, or 44, and the accuracy of detecting prostate cancer in several samples with detected cancer is at least 80% or 88%. In some embodiments, the cancer type is stage I or stage II prostate cancer, the plurality of target genomic regions are selected from List 11, 27, or 44, and the accuracy of detecting stage I or stage II prostate cancer in several samples with detected cancer is at least 75% or 83%.

[0108] In some embodiments, the cancer type is renal cell carcinoma, the plurality of target genomic regions are selected from List 12, 28, or 45, and the accuracy of detecting renal cell carcinoma in several samples with detected cancer is at least 80% or 88%. In some embodiments, the cancer type is stage I or stage II renal cell carcinoma, the plurality of target genomic regions are selected from List 12, 28, or 45, and the accuracy of detecting stage I or stage II renal cell carcinoma in several samples with detected cancer is at least 75% or 85%.

[0109] In some embodiments, the cancer type is sarcoma, the plurality of target genomic regions are selected from List 29 or 46, and the accuracy of detecting sarcoma in several samples with detected cancer is at least 80% or 88%. In some embodiments, the cancer type is stage I or stage II sarcoma, the plurality of target genomic regions are selected from List 29 or 46, and the accuracy of detecting stage I or stage II sarcoma in several samples with detected cancer is at least 75% or 83%.

[0110] In some embodiments, the cancer type is thyroid cancer, the plurality of target genomic regions are selected from List 13, 30, or 47, and the accuracy of detecting thyroid cancer in several samples with detected cancer is at least 80% or 88%. In some embodiments, the cancer type is stage I or stage II thyroid cancer, the plurality of target genomic regions are selected from List 13, 30, or 47, and the accuracy of detecting stage I or stage II thyroid cancer in several samples with detected cancer is at least 75% or 87%.

[0111] In some embodiments, the cancer type is upper gastrointestinal cancer, the plurality of target genomic regions are selected from List 14, 31, or 48, and the accuracy of detecting upper gastrointestinal cancer in several samples with detected cancer is at least 80% or 88%. In some embodiments, the cancer type is stage I or stage II upper gastrointestinal cancer, the plurality of target genomic regions are selected from List 14, 31, or 48, and the accuracy of detecting stage I or stage II upper gastrointestinal cancer in several samples with detected cancer is at least 75% or 83%.

[0112] In some embodiments, the cancer type is uterine cancer, the plurality of target genomic regions are selected from List 15, 32, or 49, and the accuracy of detecting uterine cancer in several samples with detected cancer is at least 80% or 88%. In some embodiments, the cancer type is stage I or stage II uterine cancer, the plurality of target genomic regions are selected from List 16 or 33, and the accuracy of detecting stage I or stage II uterine cancer in several samples with detected cancer is at least 75% or 85%.

[0113] In some embodiments, the cancer type is anorectal cancer, the plurality of target genomic regions are selected from List 16 or 33, and the sensitivity for anorectal cancer is at least 65% or 75%. In some embodiments, the cancer type is stage I or stage II anorectal cancer, the plurality of target genomic regions are selected from List 16 or 33, and the sensitivity for stage I or stage II anorectal cancer is at least 65% or 55%.

[0114] In some embodiments, the cancer type is bladder and urethral epithelial carcinoma, the plurality of target genomic regions are selected from List 1, 17, or 34, and the sensitivity for bladder and urethral epithelial carcinoma is at least 50% or 40%. In some embodiments, the cancer type is stage I or II bladder and urethral epithelial carcinoma, the plurality of target genomic regions are selected from List 1, 17, or 34, and the accuracy for stage I or II bladder and urethral epithelial carcinoma is at least 40% or 50%.

[0115] In some embodiments, the cancer type is breast cancer, the plurality of target genomic regions are selected from List 2, 18, or 35, and the sensitivity to breast cancer is at least 20% or 25%. In some embodiments, the cancer type is stage I or stage II breast cancer, the plurality of target genomic regions are selected from List 2, 18, or 35, and the sensitivity to stage I or stage II breast cancer is at least 15% or 18%.

[0116] In some embodiments, the cancer type is cervical cancer, the plurality of target genomic regions are selected from List 3, 19, or 36, and the sensitivity for cervical cancer is at least 25% or 35%. In some embodiments, the cancer type is stage I or stage II cervical cancer, the plurality of target genomic regions are selected from List 3, 19, or 36, and the sensitivity for stage I or stage II cervical cancer is at least 17% or 22%.

[0117] In some embodiments, the cancer type is colorectal cancer, the plurality of target genomic regions are selected from List 4, 20, or 37, and the sensitivity for colorectal cancer is at least 55% or 65%. In some embodiments, the cancer type is stage I or stage II colorectal cancer, the plurality of target genomic regions are selected from List 4, 20, or 37, and the sensitivity for stage I or stage II colorectal cancer is at least 25%, 29%, or 34%.

[0118] In some embodiments, the cancer type is head and neck cancer, the plurality of target genomic regions are selected from List 5, 21, or 38, and the sensitivity for head and neck cancer is at least 70% or 80%. In some embodiments, the cancer type is stage I or stage II head and neck cancer, the plurality of target genomic regions are selected from List 5, 21, or 38, and the sensitivity for stage I or stage II head and neck cancer is at least 70% or 79%.

[0119] In some embodiments, the cancer type is liver and bile duct cancer, the plurality of target genomic regions are selected from List 6, 22, or 39, and the sensitivity for liver and bile duct cancer is at least 75% or 85%. In some embodiments, the cancer type is stage I or stage II liver and bile duct cancer, the plurality of target genomic regions are selected from List 6, 22, or 39, and the sensitivity for stage I or stage II liver and bile duct cancer is at least 65% or 75%.

[0120] In some embodiments, the cancer type is lung cancer, the plurality of target genomic regions are selected from List 7, 23, or 40, and the sensitivity to lung cancer is at least 55% or 62%. In some embodiments, the cancer type is stage I or stage II lung cancer, the plurality of target genomic regions are selected from List 7, 23, or 40, and the sensitivity to stage I or stage II lung cancer is at least 20% or 25%.

[0121] In some embodiments, the cancer type is melanoma, the plurality of target genomic regions are selected from List 8, 24 or 41, and the sensitivity to melanoma is at least 40% or 30%.

[0122] In some embodiments, the cancer type is ovarian cancer, the target genomic regions are selected from List 9, 25 or 42, and the sensitivity to ovarian cancer is at least 70% or 80%.

[0123] In some embodiments, the cancer type is pancreatic and gallbladder cancer, the plurality of target genomic regions are selected from List 10, 26, or 43, and the sensitivity for pancreatic and gallbladder cancer is at least 60%, 70%, or 74%. In some embodiments, the cancer type is stage I or stage II pancreatic and gallbladder cancer, the plurality of target genomic regions are selected from List 10, 26, or 43, and the sensitivity for stage I or stage II pancreatic and gallbladder cancer is at least 40% or 50%.

[0124] In some embodiments, the cancer type is sarcoma, the target genomic regions are selected from List 29 or 46, and the sensitivity to sarcoma is at least 40% or 50%.

[0125] In some embodiments, the cancer type is upper gastrointestinal cancer, the plurality of target genomic regions are selected from List 14, 31, or 48, and the sensitivity for upper gastrointestinal cancer is at least 70% or 60%. In some embodiments, the cancer type is stage I or stage II upper gastrointestinal cancer, the plurality of target genomic regions are selected from List 14, 31, or 48, and the sensitivity for stage I or stage II upper gastrointestinal cancer is at least 35% or 45%.

[0126] In some embodiments, the composition comprising several oligonucleotide decoys is any of the several compositions described herein, or any of the several cancer assay combinations described herein. In some embodiments, the several genomic regions comprise no more than 1700, 1300, 900, 700, or 400 genomic regions. In some embodiments, the total size of the several genomic regions is less than 4 MB, less than 2 MB, less than 1100 kb, less than 750 kb, less than 270 kb, less than 200 kb, less than 150 kb, less than 100 kb, or less than 50 kb. In some embodiments, the subject has an increased risk of one or more cancer types. In some embodiments, the subject exhibits several symptoms associated with one or more cancer types. In some embodiments, the subject has not been diagnosed with a cancer.

[0127] In some embodiments, the classifier is trained on several converted DNA sequences derived from at least 100 subjects having a first cancer type, at least 100 subjects having a second cancer type, and at least 100 subjects not having cancer. In some embodiments, the first cancer type is ovarian cancer. In some embodiments, the first cancer type is colorectal cancer. In some embodiments, the first cancer type is selected from thyroid cancer, melanoma, sarcoma, myeloma, kidney cancer, prostate cancer, breast cancer, uterine cancer, ovarian cancer, bladder cancer, urethral cancer, cervical cancer, anorectal cancer, head and neck cancer, colorectal cancer, liver cancer, pancreatic cancer, gallbladder cancer, esophageal cancer, stomach cancer, multiple myeloma, lymphoma, lung cancer, or leukemia. In some embodiments, the classifier is trained on several converted DNA sequences derived from at least 1000, at least 2000, or at least 4000 target genomic regions selected from any of Lists 1 to 49.

[0128] In some embodiments, the trained classifier determines the presence or absence of cancer, or a cancer type, by: (i) generating a set of multiple features for the sample, wherein each feature in the set of multiple features includes a numerical value; (ii) inputting the set of multiple features into the classifier, wherein the classifier includes a multinomial classifier; (iii) determining a set of probability scores in the classifier based on the set of multiple features, wherein the set of probability scores includes a probability score for each cancer type category and each non-cancer type category; and (iv) measuring the set of probability scores with a threshold based on one or more values ​​determined during the training of the classifier to determine a final cancer classification for the sample. In some embodiments, the set of multiple features includes a set of binary features. In some embodiments, the numerical value includes a single binary value. In some embodiments, the multinomial classifier includes a multinomial logistic regression ensemble trained to predict a source tissue for the cancer. In some embodiments, the method further includes the step of: determining the final cancer classification relative to a minimum value, based on the difference between two highest probability scores, wherein the minimum value corresponds to a predefined percentage of training cancer samples, the predefined percentage of training cancer samples being assigned the correct cancer type as the highest score during the training of the classifier.

[0129] In some embodiments, (i) a cancer label is assigned as the final cancer classification based on the determination that the difference between the two highest probability scores exceeds the minimum value, the cancer label corresponding to the highest probability score determined by the classifier; and (ii) an indeterminate cancer label is assigned as the final cancer classification based on the determination that the difference between the two first two probability scores does not exceed the minimum value.

[0130] In this document, several methods for treating a type of cancer in a desired subject are described in specific embodiments, the methods comprising the steps of: (i) detecting the type of cancer by any of the methods described herein; and (ii) administering an anticancer therapeutic agent to the subject. In some embodiments, the anticancer therapeutic agent is a chemotherapeutic agent selected from the group consisting of alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, cytoskeleton disruptors (taxanes), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, and platinum-based reagents.

[0131] Incorporated by reference

[0132] All publications, patents and patent applications mentioned in this specification are incorporated herein by reference to the extent that each individual publication, patent or patent application is specifically and individually indicated to be incorporated herein by reference. Attached Figure Description

[0133] The novel features of this disclosure are set forth in detail in the appended claims. A better understanding of these features and the advantages of this disclosure will be obtained by referring to the following detailed description of several exemplary embodiments, in which the principles of this disclosure are applied, and the accompanying drawings of these embodiments:

[0134] Figure 1A The illustration depicts a 2x probe layout according to one embodiment, with three probes targeting a small target region, and each base in the target region (framed within a dashed rectangle) being covered by at least two probes.

[0135] Figure 1B The illustration depicts a 2x probe layout according to one embodiment, in which more than three probes target a large target region, and each base in the target region (framed within a dashed rectangle) is covered by at least two probes.

[0136] Figure 1C The illustration depicts probe design for several hypomethylated and / or hypermethylated fragments in several genomic regions, according to one embodiment.

[0137] Figure 2 This illustration depicts a procedure for generating a combination of cancer laboratory tests, according to one embodiment.

[0138] Figure 3AIt is a flowchart describing a procedure for creating a data structure for a control group according to one embodiment.

[0139] Figure 3B It is a flowchart describing an embodiment of the process. Figure 3A The control group performs an additional step in verifying the data structure.

[0140] Figure 4 It is a flowchart describing a procedure, according to one embodiment, for selecting several genomic regions for designing several probes for a combination of cancer assays.

[0141] Figure 5 This is a diagram illustrating the calculation of an exemplary p-value score according to one embodiment.

[0142] Figure 6A It is a flowchart describing a procedure, according to one embodiment, for training a classifier based on several hypomethylated and hypermethylated fragments that indicate cancer.

[0143] Figure 6B It is a flowchart describing a procedure, according to one embodiment, for determining several segments indicating cancer using a probability model.

[0144] Figure 7A It is a flowchart describing a procedure for sequencing a fragment of cell-free (cf) DNA according to one embodiment.

[0145] Figure 7B According to one embodiment, a methylation state vector is obtained by sequencing a fragment of cell-free (cf) DNA. Figure 7A A diagram of the procedure.

[0146] Figure 8A The chart illustrates the extent of bisulfite conversion (top chart) and average coverage / sequencing depth across various stages of cancer (bottom chart).

[0147] Figure 8B The diagram illustrates the concentration of cfDNA in each sample across various stages of cancer.

[0148] Figure 9 It is a graph showing the amount of DNA fragments hybridized to the probes based on the size of the overlap between the DNA fragments and the probes.

[0149] Figure 10A A flowchart illustrating several devices for sequencing several nucleic acid samples according to one embodiment; Figure 10B An analytical system for analyzing the methylation state of cff)NA according to one embodiment is illustrated.

[0150] Figure 11 It is a color-coded chart that shows the number of genomic regions selected to distinguish each target TOO (x-axis) from a contrasting TOO (y-axis).

[0151] Figure 12 Presents data used to validate several selected genomic regions using cfDNA and WBGgDNA. The ratio (y-axis) for correctly classifying each TOO (x-axis) is provided.

[0152] Figure 13 It is a receiver-operator curve that compares the true positive rate and false positive rate of cancer detection by applying methylation status information from several target genomic regions (optimized for lung cancer) from list 23, which is applied by a trained classifier. Detailed Implementation

[0153] definition:

[0154] Unless otherwise defined, all technical and scientific terms used herein have the meanings commonly understood by one of ordinary skill in the art to which this description pertains. As used herein, the following terms have the meanings attributed to them hereinafter.

[0155] As used herein, any reference to "an embodiment" or "an embodiment" means a particular element, feature, structure, or characteristic described in association with the described embodiment, which is included in at least one embodiment. The appearance of the term "in an embodiment" throughout the specification does not necessarily refer to the same embodiment, thereby providing a framework for the various possibilities of several described embodiments operating together.

[0156] As used herein, “comprisis,” “comrising,” “including,” “has,” “having,” or any other variation thereof are intended to cover a non-exclusive inclusion. For example, a procedure, method, article, or apparatus that includes a series of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent in such a procedure, method, article, or apparatus. Furthermore, unless expressly stated to the contrary, “or” means an inclusive or rather than an exclusive or. For example, a condition A or B is satisfied by any of the following: A is true (or exists) and B is false (or does not exist); A is false (or does not exist) and B is true (or exists); and both A and B are true (or exist).

[0157] Furthermore, the use of “a” or “an” is applied to describe elements and components of several embodiments herein. This is for convenience only and to give the general meaning of this description. This description should be read as including one or at least one, and the singular includes the plural, unless clearly otherwise implied.

[0158] As used herein, ranges and dosages may be expressed as “about” for a specific value or range. “About” also includes the precise dosage. Therefore, “about 5 micrograms” means “about 5 micrograms” and also “5 micrograms”. Generally, the term “about” includes a dosage expected to be within experimental error. In some embodiments, “about” means a indicated number or value that is “+” or “-” 20%, 10%, or 5%. Furthermore, ranges referenced herein are understood to be shorthand for all values ​​within the range, including the referenced endpoints. For example, a range of 1 to 50 is understood to include any number, combination of several numbers, or subrange of numbers that comes from the group consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, and 50.

[0159] The term “methylation,” as used herein, refers to the process by which a methyl group is added to a DNA molecule. For example, a hydrogen atom on the pyrimidine ring of a cytosine base can be converted to a methyl group, forming 5-methylcytosine. The term also refers to the process by which a hydroxymethyl group is added to a DNA molecule, for example, through the oxidation of the methyl group on the pyrimidine ring of the cytosine base. Methylation and hydroxymethylation tend to occur at the dinucleotides of cytosine and guanine, referred to herein as “CpG sites.”

[0160] The term "methylation" can also refer to the methylation state of a CpG site. A CpG site with a 5-methylcytosine moiety is methylated. A CpG site with a hydrogen atom on the pyrimidine ring of the cytosine base is unmethylated.

[0161] In several of these embodiments, as is well known in the art, the wet laboratory tests used to detect methylation may differ from those described herein.

[0162] The term “methylation site,” as used herein, refers to a region of a DNA molecule to which a methyl group can be added. “CpG” sites are the most common methylation sites, but methylation sites are not limited to CpG sites. For example, DNA methylation can occur in cytosine at CHG and CHH, where H is adenine, cytosine, or thymine. Using the methods and procedures disclosed herein, methylation of cytosine in the form of 5-hydroxymethylcytosine and its characterization can also be evaluated (see, for example, WO 2010 / 037001 and WO 2011 / 127136, which are incorporated herein by reference).

[0163] The term “CpG site” is used in this document to refer to a region in a DNA molecule in which a cytosine nucleotide is followed by a guanine nucleotide in a linear sequence of several bases along the 5′ to 3′ direction. “CpG” is a shorthand for 5′-C-phosphate-G-3′, where 5′-C-phosphate-G-3′ consists of cytosine and guanine separated by only one phosphate group. The cytosine in the CpG dinucleotide can be methylated to form 5-methylcytosine.

[0164] The term “CpG detection site” as used herein refers to a region in a probe configured to hybridize to a CpG site on a target DNA molecule. The CpG site on the target DNA molecule may comprise cytosine and guanine separated by a single phosphate group, wherein the cytosine is methylated or unmethylated. Alternatively, the CpG site on the target DNA molecule may comprise uracil and guanine separated by a single phosphate group, wherein the uracil is generated by the conversion of unmethylated cytosine.

[0165] The term "UpG" is a shorthand for 5′-U-phosphate-G-3′, which is a combination of uracil and guanine separated by only one phosphate group. UpG can be produced by bisulfite treatment of DNA, which converts unmethylated cytosine into uracil. Cytosine can also be converted into uracil by other methods known in the art, such as chemical modification, synthesis, or enzymatic conversion.

[0166] Terms such as “hypomethylation” or “hypermethylation”, as used herein, refer to the monomethylation state of a DNA molecule containing multiple (e.g., more than 3, 4, 5, 6, 7, 8, 9, 10, etc.) CpG sites, wherein a high proportion (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage in the range of 50% to 100%) of the CpG sites are either unmethylated or methylated.

[0167] The terms "methylation state vector" or "methylation states vector," as used herein, refer to a vector comprising multiple elements, each element indicating the methylation state of one methylation site in a DNA molecule containing multiple methylation sites, in the order of appearance of methylation sites from 5′ to 3′. For example, <M x M x+1 M x+2 >、 <M x M x+1 U x+2 >、...、<U x U x+1 U x+2 > It can be several methylation vectors of a DNA molecule that includes three methylation sites, where M represents a methylated methylation site and U represents an unmethylated methylation site.

[0168] The terms “abnormal methylation pattern” or “anomalous methylation pattern” as used herein refer to a methylation pattern or methylation state vector of a DNA molecule that is expected to be found less frequently in a sample than a threshold. In one embodiment provided herein, the expectedness of finding a particular methylation state vector in a health control group comprising several healthy individuals is represented by a p-value. A low p-value generally corresponds to a methylation state vector that is less expected than other methylation state vectors in samples from healthy individuals. A high p-value generally corresponds to a methylation state vector that is more expected than other methylation state vectors in samples from healthy individuals in the health control group. A methylation state vector with a p-value below a threshold (e.g., 0.1, 0.01, 0.001, 0.0001, etc.) can be defined as an abnormal / anomalous methylation pattern. Various methods known in the art can be used to calculate a p-value or predictability of a methylation pattern or methylation state vector. Exemplary methods provided herein involve using a Markov chain probability that assumes the methylation state of a CpG site depends on the methylation state of neighboring CpG sites. Alternative methods provided herein calculate the predictability of a specific methylation state vector observed in a healthy individual by applying a mixture model comprising multiple components, each a site-independent model, wherein methylation at each CpG site is assumed to be independent of methylation states at other CpG sites.

[0169] The term “cancer sample” as used herein means a sample comprising genomic DNA from an individual diagnosed with cancer. The genomic DNA may be, but is not limited to, a fragment of cfDNA or chromosomal DNA from an individual with cancer. The genomic DNA may be sequenced (or detected), and its methylation status may be evaluated by methods known in the art, such as bisulfite sequencing. When the genomic sequence is obtained from a public database (e.g., The Cancer Genome Atlas (TCGA)) or experimentally obtained by sequencing the genome of an individual diagnosed with cancer, a cancer sample may mean genomic DNA or a fragment of cfDNA containing the genomic sequence. The term “several cancer samples” as a plural means several samples comprising genomic DNA from multiple individuals, each diagnosed with cancer. In various embodiments, several cancerous samples were used from more than 100, 300, 500, 1000, 2000, 5000, 10000, 20000, 40000, 50000 or more individuals diagnosed with cancer.

[0170] The terms “non-cancerous sample” or “healthy sample,” as used herein, mean a sample comprising genomic DNA from an individual not diagnosed with cancer. The genomic DNA may be, but is not limited to, fragments of cfDNA or chromosomal DNA from an individual without cancer. The genomic DNA may be sequenced (or detected), and its methylation status may be assessed by methods known in the art, such as bisulfite sequencing. When the genomic sequence is obtained from a public database (e.g., The Cancer Genome Atlas (TCGA)) or experimentally obtained by sequencing the genome of an individual without cancer, a non-cancerous sample may mean genomic DNA or fragments of cfDNA containing the genomic sequence. The term “several non-cancerous samples” as a plural means several samples comprising genomic DNA from multiple individuals, each diagnosed as not having cancer. In various embodiments, healthy samples from more than 100, 300, 500, 1000, 2000, 5000, 10000, 20000, 40000, 50000 or more individuals diagnosed as cancer-free were used.

[0171] The term “training sample,” as used herein, means a sample used to train a classifier described herein and / or to select one or more genomic regions for cancer detection or to detect a cancer-originating tissue or cancer cell type. The training sample may include genomic DNA or modifications thereof from one or more healthy subjects or from one or more subjects having a condition (e.g., cancer, a specific type of cancer, a specific stage of cancer, etc.). The genomic DNA may be, but is not limited to, several cfDNA fragments or chromosomal DNA. The genomic DNA may be sequenced (or detected), and its methylation status may be evaluated by methods known in the art, such as bisulfite sequencing. When the genomic sequence is obtained from a public database (e.g., the Cancer Genome Atlas (TCGA)) or experimentally obtained by sequencing an individual's genome, a training sample may mean genomic DNA or cfDNA fragments having said genomic sequence.

[0172] The term “test sample,” as used herein, means a sample from an object whose health condition has been or will be detected using a classifier and / or a combination of laboratory tests described herein. The test sample may include genomic DNA or modifications thereof. The genomic DNA may be, but is not limited to, several cfDNA fragments or chromosomal DNA.

[0173] The term "target genomic region," as used herein, refers to a region of a genome selected for analysis in a test sample. An assay assay is generated having several probes designed to hybridize to (and optionally pull down) several nucleic acid fragments derived from or from the target genomic region. A nucleic acid fragment derived from the target genomic region means a nucleic acid fragment produced by degradation, cleavage, bisulfite conversion, or other treatment of DNA from the target genomic region.

[0174] Various target genomic regions are described according to their chromosomal locations in the sequence listings submitted with this document. The sequence listings include the following information: (1) the chromosome in which the region is located, and the start and stop positions of the genomic region; and (2) whether the region is hypomethylated or hypermethylated in cancer (or “binary”, if both hypomethylation and hypermethylation are provided). Chromosome numbers and start and stop positions are provided relative to a known human reference genome, hg19. The sequence of the human reference genome hg19 is available from the Genome Reference Consortium under the reference number GRCh37 / hg19, and is also available from the Genome Browser provided by the Santa Cruz Genomics Institute. Chromosomal DNA is double-stranded; therefore, a target genomic region comprises two DNA strands: one with the sequence provided in the listings, and a second strand that is the opposite complementary strand of the sequence in the listings. Probes can be designed to hybridize to one or both sequences. Optionally, the probe is hybridized to a converted sequence, which, for example, is derived from a sequence treated with sodium bisulfite.

[0175] As used herein, the term “off-target genomic region” refers to a region of a genome that was not selected for analysis in a test sample but has sufficient homology to a target genomic region and is potentially designed to link and pull down a probe targeting that target genomic region. In one embodiment, an off-target genomic region is a genomic region that is aligned with a probe along at least 45 bases with at least 90% concordance.

[0176] "Converted DNA molecule," "converted cfDNA molecule," and "treated modified fragment obtained from said cfDNA molecule" refer to DNA molecules obtained by treating a sample of DNA or cfDNA molecules to distinguish between methylated and unmethylated nucleotides in DNA or cfDNA molecules. For example, in one embodiment, the sample may be treated with bisulfite ions (e.g., using sodium bisulfite), as is well known in the art, to convert unmethylated cytosine ("C") to uracil ("U"). In another embodiment, the conversion of unmethylated cytosine to uracil is accomplished using an enzymatic conversion reaction, for example, using a cytidine deaminase (such as APOBEC). After treatment, the converted DNA molecule or cfDNA molecule includes additional uracil not present in the original cfDNA sample. The DNA strand including uracil is replicated by DNA polymerase, resulting in the addition of adenine to a new complementary strand, instead of guanine, which is normally the complement of cytosine or methylcytosine.

[0177] The terms "cell-free nucleic acid," "cell-free DNA," or "cfDNA" refer to nucleic acid fragments that circulate within an individual's body (e.g., in the bloodstream) and originate from one or more healthy cells and / or one or more cancerous cells. Furthermore, cfDNA can originate from other sources such as viruses or fetuses.

[0178] Terms such as “circulating tumor DNA” or “ctDNA” refer to nucleic acid fragments originating from tumor cells that may be released into the bloodstream as a result of biological processes, such as apoptosis or necrosis of dying cells, or actively by surviving tumor cells.

[0179] The term “fragment,” as used herein, can mean a segment of a nucleic acid molecule. For example, in one embodiment, a fragment can mean a cfDNA molecule in a blood or plasma sample, or a cfDNA molecule extracted from a blood or plasma sample. An amplified product of a cfDNA molecule can also be referred to as a “fragment.” In another embodiment, the term “fragment,” as described herein, means a sequence read, or a set of sequence reads, that has been processed for (e.g., in machine learning-based classification) subsequent analysis. For example, as is well known in the art, raw sequence reads can be aligned to a reference genome, and matched end sequence reads can be assembled into a longer fragment for subsequent analysis.

[0180] The term “individual” refers to a human being. The term “healthy individual” refers to an individual who is presumed not to have cancer or disease.

[0181] The term “subject” refers to an individual whose DNA is analyzed. A subject can be a testing subject whose DNA is evaluated using a targeted assay combination as described herein to assess whether the person has a cancer or other disease. A subject can also be a member of a control group, known not to have a cancer or other disease. A subject can also be a member of a cancer or other disease group, known to have a cancer or other disease. Control groups and cancer / disease groups can be used to assist in the design or validation of the targeted assay combination.

[0182] The term “sequence readout” as used herein refers to a nucleotide sequence readout from a sample. Sequence readouts can be obtained by various methods provided herein or known in the art.

[0183] The term “sequencing depth,” as used herein, refers to the count of the number of times a given target nucleic acid is sequenced in a sample (e.g., the count of sequence reads at a given target region). Increasing sequencing depth can reduce the amount of nucleic acid required to assess a disease state (e.g., cancer or cancer-derived tissue).

[0184] Terms such as “tissue of origin” or “TOO,” as used in this article, refer to the organ, organ group, body region, or cell type from which a cancer arises or originates. Identification of a tissue of origin or cancer cell type typically allows for the identification of the most appropriate next step in cancer care continuum for further diagnosis, staging, and treatment decisions.

[0185] "Transition" generally refers to a change in the base composition from one purine to another, or from one pyrimidine to another. For example, the following changes are transitions: C→U, U→C, G→A, A→G, C→T, and T→C.

[0186] The term "entire set of probes" or "entire set of polynucleotide-containing probes" in a detection combination or decoy set generally means all probes delivered with a particular detection combination or decoy set. For example, in some embodiments, a detection combination or decoy set may include (1) several probes having the features specified herein (e.g., several probes for linking to cell-free DNA fragments corresponding to or derived from genomic regions presented herein in one or more lists) and (2) additional probes that do not contain such features. The term "entire set of probes" in a detection combination generally means all probes delivered with the detection combination or decoy set, including probes that do not contain the specified features.

[0187] Cancer testing kit:

[0188] In one aspect, this description provides a cancer assay suite comprising a plurality of probes or probe pairs. The plurality of assay suites described herein may alternatively be referred to as a plurality of decoy sets, or as a plurality of compositions comprising a plurality of decoy oligonucleotides. The plurality of probes are specifically designed to target one or more nucleic acid molecules that correspond to, or are derived from, a plurality of genomic regions that are differentially methylated between cancer and non-cancer samples, between different tissues of origin (TOO) types, between different cancer cell types, or between samples at different stages of cancer, as identified by the methods provided herein. In some embodiments, a plurality of probes target a plurality of genomic regions (or nucleic acid molecules derived from said plurality of genomic regions) having a methylation pattern specific to a cancer type, such as (1) bladder cancer, (2) breast cancer, (3) cervical cancer, (4) colorectal cancer, (5) head and neck cancer, (6) hepatobiliary cancer, (7) lung cancer, (8) melanoma, (9) ovarian cancer, (10) pancreatic cancer, (11) prostate cancer, (12) kidney cancer, (13) thyroid cancer, (14) upper gastrointestinal cancer, or (15) uterine cancer. In some embodiments, the detection combination comprises a plurality of probes targeting a plurality of genomic regions specific to a single cancer type. In some embodiments, the detection combination comprises a plurality of probes targeting 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more cancer types. In some embodiments, subject to a size budget (determined by the sequencing budget and the desired sequencing depth), several target genomic regions are selected to maximize classification accuracy.

[0189] To design the aforementioned cancer assay suite, an analytical system can collect several samples corresponding to various considered outcomes, such as samples known to have cancer, samples considered healthy, samples from a known source tissue, etc. The source of cfDNA and / or ctDNA used to select several target genomic regions can vary depending on the purpose of the assay. For example, different sources may be desirable for an assay intended to generally diagnose cancer, diagnose a specific type of cancer, diagnose a stage of cancer, or diagnose a source tissue. These samples can be processed using one or more methods known in the art (e.g., whole-genome bisulfite sequencing (WGBS)) to determine the methylation status of several CpG sites, or the information can be obtained from a public database (e.g., TCGA). The analytical system can be any general-purpose computing system having a computer processor and a computer-readable storage medium having several instructions for executing the computer processor to perform any or all of the operations described in this disclosure.

[0190] The design and application of the cancer testing suite are generally described in Figure 2 In order to design the cancer assay suite, an analytical system collects several samples corresponding to various considered outcomes, such as samples known to have cancer, samples considered healthy, samples from a known TOO, etc. These samples may be processed (e.g., by whole-genome bisulfite sequencing (WGBS)) or obtained from a public database (e.g., TCGA). The analytical system may be any general-purpose computing system having a computer processor and a computer-readable storage medium having several instructions for executing the computer processor to perform any or all of the operations described in this disclosure. With the several samples, the analytical system determines the methylation status at several CpG sites for each fragment in the samples.

[0191] The analysis system can then select target genomic regions to be included in a cancer assay suite based on the methylation patterns of several nucleic acid fragments. One approach is to consider the pairwise discriminability between several pairs of results (e.g., one cancer type versus a second cancer type) for the selection of several target regions. Another approach is to consider the discriminability of several target genomic regions when each result is considered relative to several other results (e.g., one cancer type versus all other cancer types). From the selected target genomic regions with high discriminability power, the analysis system can design probes to target several nucleic acid fragments that include or are derived from the selected genomic regions. The analysis system can generate cancer test kits of varying sizes. For example, a small cancer test kit includes several probes targeting the most informative genomic regions; a medium-sized cancer test kit includes several probes from the small-sized kit, plus additional probes targeting second-layer informative genomic regions; and a large cancer test kit includes several probes from the small and medium-sized kits, plus more probes targeting third-layer informative genomic regions. With data obtained from such cancer test kits (e.g., methylation status of nucleic acids derived from the test kits), the analysis system can train a classifier using various classification techniques to predict the likelihood of a sample having a specific outcome or state, such as cancer, a specific cancer type, or other conditions.

[0192] Exemplary methods for designing a combination of cancer diagnostic tests are generally described in Figure 2 For example, to design a cancer assay suite, an analytical system can collect information on the methylation status of several CpG sites from several nucleic acid fragments derived from several samples corresponding to various considered outcomes, such as samples known to have cancer, samples considered healthy, samples from a known TOO, etc. These samples can be processed (e.g., with whole-genome bisulfite sequencing (WGBS)) to determine the methylation status of the CpG sites, or the information can be obtained from TCGA. The analytical system can be any general-purpose computing system having a computer processor and a computer-readable storage medium having several instructions for executing the computer processor to perform any or all of the operations described in this disclosure.

[0193] In some embodiments, the cancer assay suite includes at least 500 pairs of probes, each of the at least 500 pairs comprising two probes configured to overlap each other via an overlapping sequence comprising at least 30 nucleotides, and each probe being configured to hybridize to a converted DNA molecule (e.g., cfDNA) corresponding to one or more genomic regions. In some embodiments, each of the plurality of genomic regions includes at least five methylation sites, and wherein the at least five methylation sites have an anomalous methylation pattern in cancerous samples or have different methylation patterns between samples of different origins. For example, in one embodiment, the at least five methylation sites are differentially methylated between cancerous and non-cancerous samples, or between one or more pairs of cancer samples from tissues with different cancer origins. In some embodiments, each probe pair includes a first probe and a second probe, wherein the second probe is different from the first probe. The second probe may overlap with the first probe via an overlapping sequence that is at least 30, at least 40, at least 50, or at least 60 nucleotides long.

[0194] The target genomic regions may be selected from any of Lists 1 to 49 (Table 1). In some embodiments, the cancer assay suite includes a plurality of probes, each of which is configured to hybridize to a converted cfDNA molecule corresponding to one or more genomic regions from any of Lists 1 to 49 or any combination of the lists. In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 20% derived from any of the target genomic regions from Lists 1 to 49. In some embodiments, the plurality of different decoy oligonucleotides are configured to hybridize to a plurality of DNA molecules, the plurality of DNA molecules being at least 30%, 40%, 50%, 60%, 70%, or 80% derived from any of the target genomic regions from Lists 1 to 49.

[0195] The plurality of target genomic regions may be selected from List 1. In some embodiments, a method for detecting bladder cancer includes the step of: assessing the methylation status of a plurality of sequence reads derived from the plurality of target genomic regions of List 1. The plurality of target genomic regions may be selected from List 2. In some embodiments, a method for detecting breast cancer includes the step of: assessing the methylation status of a plurality of sequence reads derived from the plurality of target genomic regions of List 2. The plurality of target genomic regions may be selected from List 3. In some embodiments, a method for detecting cervical cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from the plurality of target genomic regions of List 3. The plurality of target genomic regions may be selected from List 4. In some embodiments, a method for detecting colorectal cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from the plurality of target genomic regions of List 4. The plurality of genomic regions may be selected from List 5. In some embodiments, a method for detecting head and neck cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from the plurality of target genomic regions of List 5. The target genomic regions may be selected from List 6. In some embodiments, a method for detecting hepatobiliary cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from the plurality of target genomic regions in List 6. The plurality of target genomic regions may be selected from List 7. In some embodiments, a method for detecting lung cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from the plurality of target genomic regions in List 7. The plurality of target genomic regions may be selected from List 8. In some embodiments, a method for detecting melanoma includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from the plurality of target genomic regions in List 8. The plurality of target genomic regions may be selected from List 9. In some embodiments, a method for detecting ovarian cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from the plurality of target genomic regions in List 9. The plurality of target genomic regions may be selected from List 10. In some embodiments, a method for detecting pancreatic cancer includes the step of: assessing the methylation status of a plurality of sequence reads, said plurality of sequence reads being derived from a plurality of target genomic regions of List 10. The plurality of target genomic regions may be selected from List 11. In some embodiments, a method for detecting prostate cancer includes the step of: assessing the methylation status of a plurality of sequence reads, said plurality of sequence reads being derived from a plurality of target genomic regions of List 11. The plurality of target genomic regions may be selected from List 12. In some embodiments, a method for detecting renal cell carcinoma includes the step of: assessing the methylation status of a plurality of sequence reads, said plurality of sequence reads being derived from a plurality of target genomic regions of List 12.The target genomic regions may be selected from List 13. In some embodiments, a method for detecting thyroid cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from a plurality of target genomic regions in List 13. The plurality of target genomic regions may be selected from List 14. In some embodiments, a method for detecting upper gastrointestinal cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from a plurality of target genomic regions in List 14. The plurality of target genomic regions may be selected from List 15. In some embodiments, a method for detecting uterine cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from a plurality of target genomic regions in List 15.

[0196] The target genomic regions may be selected from List 16. In some embodiments, a method for detecting anorectal cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from a plurality of target genomic regions in List 16. The plurality of target genomic regions may be selected from List 17. In some embodiments, a method for detecting bladder and urethral epithelial cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from a plurality of target genomic regions in List 17. The plurality of target genomic regions may be selected from List 18. In some embodiments, a method for detecting breast cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from a plurality of target genomic regions in List 18. The plurality of target genomic regions may be selected from List 19. In some embodiments, a method for detecting cervical cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from a plurality of target genomic regions in List 19. The plurality of target genomic regions may be selected from List 20. In some embodiments, a method for detecting colorectal cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being a plurality of target genomic regions derived from List 20. The plurality of target genomic regions may be selected from List 21. In some embodiments, a method for detecting head and neck cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being a plurality of target genomic regions derived from List 21. The plurality of target genomic regions may be selected from List 22. In some embodiments, a method for detecting liver and bile duct cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being a plurality of target genomic regions derived from List 22. The plurality of target genomic regions may be selected from List 23. In some embodiments, a method for detecting lung cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being a plurality of target genomic regions derived from List 23. The plurality of target genomic regions may be selected from List 24. In some embodiments, a method for detecting melanoma includes the step of: assessing the methylation status of a plurality of sequence reads, said plurality of sequence reads being derived from a plurality of target genomic regions of List 24. The plurality of target genomic regions may be selected from List 25. In some embodiments, a method for detecting ovarian cancer includes the step of: assessing the methylation status of a plurality of sequence reads, said plurality of sequence reads being derived from a plurality of target genomic regions of List 25. The plurality of target genomic regions may be selected from List 26. In some embodiments, a method for detecting pancreatic and gallbladder cancer includes the step of: assessing the methylation status of a plurality of sequence reads, said plurality of sequence reads being derived from a plurality of target genomic regions of List 26. The plurality of target genomic regions may be selected from List 27.In some embodiments, a method for detecting prostate cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being a plurality of target genomic regions derived from List 27. The plurality of target genomic regions may be selected from List 28. In some embodiments, a method for detecting renal cell carcinoma includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being a plurality of target genomic regions derived from List 28. The plurality of target genomic regions may be selected from List 29. In some embodiments, a method for detecting sarcoma includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being a plurality of target genomic regions derived from List 29. The plurality of target genomic regions may be selected from List 30. In some embodiments, a method for detecting thyroid cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being a plurality of target genomic regions derived from List 30. The plurality of target genomic regions may be selected from List 31. In some embodiments, a method for detecting upper gastrointestinal cancer includes the step of: assessing the methylation status of a plurality of sequence reads, said plurality of sequence reads being derived from a plurality of target genomic regions of List 31. The plurality of target genomic regions may be selected from List 32. In some embodiments, a method for detecting uterine cancer includes the step of: assessing the methylation status of a plurality of sequence reads, said plurality of sequence reads being derived from a plurality of target genomic regions of List 32.

[0197] The target genomic regions may be selected from List 33. In some embodiments, a method for detecting anorectal cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from a plurality of target genomic regions in List 33. The plurality of target genomic regions may be selected from List 34. In some embodiments, a method for detecting bladder and urethral epithelial cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from a plurality of target genomic regions in List 34. The plurality of target genomic regions may be selected from List 35. In some embodiments, a method for detecting breast cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from a plurality of target genomic regions in List 35. The plurality of target genomic regions may be selected from List 36. In some embodiments, a method for detecting cervical cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being derived from a plurality of target genomic regions in List 36. The plurality of target genomic regions may be selected from List 37. In some embodiments, a method for detecting colorectal cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being a plurality of target genomic regions derived from List 37. The plurality of target genomic regions may be selected from List 38. In some embodiments, a method for detecting head and neck cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being a plurality of target genomic regions derived from List 38. The plurality of target genomic regions may be selected from List 39. In some embodiments, a method for detecting liver and bile duct cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being a plurality of target genomic regions derived from List 39. The plurality of target genomic regions may be selected from List 40. In some embodiments, a method for detecting lung cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being a plurality of target genomic regions derived from List 40. The plurality of target genomic regions may be selected from List 41. In some embodiments, a method for detecting melanoma includes the step of: assessing the methylation status of a plurality of sequence reads, said plurality of sequence reads being derived from a plurality of target genomic regions of List 41. The plurality of target genomic regions may be selected from List 42. In some embodiments, a method for detecting ovarian cancer includes the step of: assessing the methylation status of a plurality of sequence reads, said plurality of sequence reads being derived from a plurality of target genomic regions of List 42. The plurality of target genomic regions may be selected from List 43. In some embodiments, a method for detecting pancreatic and gallbladder cancer includes the step of: assessing the methylation status of a plurality of sequence reads, said plurality of sequence reads being derived from a plurality of target genomic regions of List 43. The plurality of target genomic regions may be selected from List 44.In some embodiments, a method for detecting prostate cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being a plurality of target genomic regions derived from List 44. The plurality of target genomic regions may be selected from List 45. In some embodiments, a method for detecting renal cell carcinoma includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being a plurality of target genomic regions derived from List 45. The plurality of target genomic regions may be selected from List 46. In some embodiments, a method for detecting sarcoma includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being a plurality of target genomic regions derived from List 46. The plurality of target genomic regions may be selected from List 47. In some embodiments, a method for detecting thyroid cancer includes the step of: assessing the methylation status of a plurality of sequence reads, the plurality of sequence reads being a plurality of target genomic regions derived from List 47. The plurality of target genomic regions may be selected from List 48. In some embodiments, a method for detecting upper gastrointestinal cancer includes the step of: assessing the methylation status of a plurality of sequence reads, said plurality of sequence reads being derived from a plurality of target genomic regions of List 48. The plurality of target genomic regions may be selected from List 49. In some embodiments, a method for detecting uterine cancer includes the step of: assessing the methylation status of a plurality of sequence reads, said plurality of sequence reads being derived from a plurality of target genomic regions of List 49.

[0198] Because the probes are configured to hybridize to a converted DNA or cfDNA molecule corresponding to or derived from one or more genomic regions, the probes may have a sequence different from the target genomic region. For example, a DNA containing an unmethylated CpG site will be converted to include UpG instead of CpG because unmethylated cytosine is converted to uracil via a conversion reaction (e.g., bisulfite treatment). As a result, a probe is configured to hybridize to a sequence including UpG, rather than naturally occurring unmethylated CpG. Therefore, a complementary site to the unmethylated site in the probe may include CpA instead of CpG, and some probes targeting a hypomethylated site where all methylated sites are unmethylated may not have a guanine (G) base. In some embodiments, at least 3%, 5%, 10%, 15%, or 20% of the probes do not have a CpG sequence.

[0199] The cancer assay suite can be used to detect the overall presence or absence of cancer and / or provide a cancer classification, such as a cancer type, a cancer stage such as stage I, II, III, or IV, or to provide a TOO (Total Origin) from which the cancer is believed to originate. The assay suite may include several probes targeting several target genomic regions (e.g., several lung cancer-specific targets) that are differentially methylated between overall cancerous (multi-cancer) samples and non-cancer samples, or only in cancerous samples with a specific cancer type. For example, in some embodiments, a cancer assay suite is designed to include several differentially methylated genomic regions based on bisulfite sequencing data generated from cfDNA from both cancerous and non-cancer individuals.

[0200] Each of the probes (or probe pairs) is designed to target one or more target genomic regions. These target genomic regions are selected based on several criteria designed to increase the selective enrichment of informative cfDNA fragments while reducing noise and nonspecific binding.

[0201] In one embodiment, a detection suite may include a plurality of probes that selectively bind to and optionally enrich several cfDNA fragments that are differently methylated in a cancerous sample. In this case, sequences from the several enriched fragments can provide information relevant to cancer detection. Further, the plurality of probes are designed to target several genomic regions identified as having an abnormal methylation pattern in a cancer sample, or in samples from a specific tissue type or cell type. In one embodiment, the plurality of probes are designed to target several genomic regions identified as hypermethylated or hypomethylated in a specific cancer, or in cancer-derived tissue, to provide additional selectivity and specificity for detection. In some embodiments, a detection suite includes several probes targeting several hypomethylated fragments. In some embodiments, a detection suite includes several probes targeting several hypermethylated fragments. In some embodiments, a detection suite includes several probes targeting a first group of several hypermethylated fragments and several probes targeting a second group of several hypomethylated fragments. Figure 1CIn some embodiments, the ratio (overmethylation:hypomethylation ratio) between the probes of the first group targeting several overmethylated fragments and the probes of the second group targeting several hypomethylated fragments is between 0.4 and 2, between 0.5 and 1.8, between 0.5 and 1.6, between 1.4 and 1.6, between 1.2 and 1.4, between 1 and 1.2, between 0.8 and 1, between 0.6 and 0.8, or between 0.4 and 0.6. Methods for identifying several genomic regions (i.e., genomic regions that produce differently methylated or anomalously methylated DNA molecules between cancer and non-cancer samples, between different types of tissue of origin (TOO), between different types of cancer cells, or between samples from cancer at different stages) are provided in detail herein, and methods for identifying several anomalously methylated DNA molecules or fragments identified as indicators of cancer are also provided in detail herein.

[0202] In a second example, several genomic regions may be selected when said several genomic regions produce anomalously methylated DNA molecules in cancer samples or samples with known tissue of origin (TOO) types. For example, as described herein, a Markov model trained on a set of non-cancerous samples can be used to identify several genomic regions that produce anomalously methylated DNA molecules (i.e., DNA molecules with a monomethylation pattern below a p-value threshold).

[0203] Each of the plurality of probes may target a genomic region comprising at least 30 bp (base pairs), 35 bp, 40 bp, 45 bp, 50 bp, 60 bp, 70 bp, 80 bp, 90 bp, 100 bp, or more. In some embodiments, the plurality of genomic regions may be selected to have fewer than 30, 25, 20, 15, 12, 10, 8, or 6 methylation sites.

[0204] The plurality of genomic regions may be selected when at least 80, 85, 90, 92, 95, or 98% of the at least five methylation (e.g., CpG) sites in the regions are methylated or unmethylated in non-cancerous or cancerous samples, or in cancer samples from a cancer-originating tissue (TOO).

[0205] Several genomic regions can be further filtered based on their methylation patterns to select only a few genomic regions that may provide information. For example, CpG sites that are differently methylated between cancerous and non-cancer samples (e.g., abnormally methylated or unmethylated in cancer compared to non-cancer samples), CpG sites that are differently methylated between cancerous samples of one TOO and cancerous samples of a different TOO, and CpG sites that are differently methylated only in cancerous samples of a specific TOO. For this selection, calculations can be performed for each CpG or several CpG sites. For example, a first count is determined, which is the number of cancerous samples including a fragment overlapping with the CpG (cancer_count), and a second count is determined, which is the total number of samples including the fragment overlapping with the CpG site (sum). Several genomic regions can be selected based on criteria that are positively correlated with the number of cancer-containing samples (cancer count) including cancer-indicating fragments overlapping with the CpG site, and negatively correlated with the total number of samples (total number) including cancer-indicating fragments overlapping with the CpG site. In one embodiment, the number of non-cancer samples (n) having a fragment overlapping with a CpG site is also considered. 非-癌症 ) and the number of cancer samples (n 癌症 The probability of a sample being cancer is then calculated, for example, as (n) 癌症 +1) / (n 癌症 +n 非-癌症 +2).

[0206] Several CpG sites scored by this metric are ranked and greedily added to a test suite until the test suite size budget is exhausted. The procedure for selecting several genomic regions indicating cancer is further detailed herein. In some embodiments, a test suite for detecting a specific cancer type can be designed using a similar procedure, depending on whether the test is intended as a multi-cancer test or a single-cancer test, or depending on the desired flexibility in selecting which CpG sites contribute to the test suite. In this embodiment, for each cancer type and for each CpG site, information gain is calculated to determine whether a probe targeting that CpG site should be included. The information gain can be calculated for several samples of a given cancer type with a TOO, compared to all other samples. For example, consider two random variables, “AF” and “CT”. “AF” is a binary variable indicating whether there is an anomalous fragment (yes or no) in a particular sample that overlaps with a particular CpG site. “CT” is a binary random variable indicating whether the cancer is a specific type (e.g., lung cancer or a cancer different from lung cancer). Mutual information about a given “AF” can be calculated with respect to “CT”. That is, knowing whether an anomalous fragment overlaps with a specific CpG site, how many bits of information about the cancer type (e.g., lung cancer versus non-lung cancer in this example) will be obtained. This can be used to rank several CpGs based on how lung-specific they are. This procedure is repeated for several cancer types. If a particular region is generally differentially methylated only in lung cancer (and not in other cancer types or non-cancer), several CpGs in that region will tend to have high information gain for lung cancer. For each cancer type, several CpG sites are ranked using this information gain metric and then greedily added to a detection pool until the size budget for that cancer type is exhausted.

[0207] Further filtering can be performed to select several probes that exhibit high specificity (i.e., high binding efficiency) for enrichment of nucleic acids derived from several target genomic regions. The probes can be filtered to reduce non-specific binding (or off-target binding) to nucleic acids derived from non-target genomic regions. For example, the probes can be filtered to select only those probes with fewer than a set threshold of off-target binding events. In one embodiment, the probes can be aligned to a reference genome (e.g., a human reference genome) to select several probes that are aligned across the genome to fewer than a set threshold of regions. For example, the probes can be selected to be aligned across the reference genome to fewer than 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, or 8 off-target regions. In other cases, filtering is performed to remove said several genomic regions when the sequences of said several target genomic regions appear more than 5, 10, 15, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 times in a genome. When a probe sequence or a set of probe sequences that is homologous to several target genomic regions at 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% occurs less than 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, or 8 times in a reference genome, further filtering can be performed to select several target genomic regions, or when designed to enrich... When the probe sequence or a set of probe sequences of the target genomic region is 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% homologous to the target genomic regions, and occurs more than 5, 10, 15, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 times in a reference genome, the target genomic regions are removed. This is to exclude duplicate probes that may pull out several off-target fragments, which are undesirable and could impact assay efficiency.

[0208] In some embodiments, a fragment-probe overlap of at least 45 bp is demonstrated to be effective for achieving a non-negligible pull-down as provided in Example 1 (although those skilled in the art will understand that this number is variable). In some embodiments, a mismatch of more than 10% in the overlapping region between the probe and several fragment sequences is sufficient to significantly disrupt the connection and thus impair pull-down efficiency. Therefore, several sequences aligned to the probe with at least 45 bp at a pairing rate of at least 90% may be candidates for off-target pull-down. Thus, in one embodiment, the number of such regions is scored. The optimal probes have a score of 1, meaning they pair only in one place (in the intended target region). Several probes with a middle score (i.e., less than 5 or 10) may be acceptable in some cases, and in some cases, any probe with a score higher than a certain threshold is discarded. Other cutoff values ​​may be used for specific samples.

[0209] Once the plurality of probes hybridize and capture several DNA fragments corresponding to or derived from a target genomic region, the hybridized probe-DNA fragment intermediates are pulled down (or separated), the target DNA is amplified, and the methylation state of the target DNA is determined, for example, by sequencing or hybridization to a microarray. The sequence reads provide information associated with cancer detection. For this purpose, a detection array is designed to include several probes that can capture several fragments that collectively provide information associated with cancer detection. In some embodiments, a detection combination includes at least 5, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1200, 1400, 1600, 1800, 2000, 2200, 2400, 2600, 2800, 3000, 3200, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, or 10000 pairs of probes. In other embodiments, a detection combination includes at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1200, 1400, 1600, 1800, 2000, 2200, 2400, 2600, 2800, 3000, 3200, 4000, 4500, 5000, 5500, 6000, 6500, 7000, 7500, 8000, 8500, 9000, 10000, 15000, or 20000 pairs of probes. The total number of probes may include at least 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 120,000, 140,000, 160,000, 180,000, 200,000, 240,000, 260,000, and 2... 80,000, 300,000, 320,000, 400,000, 450,000, 500,000, 550,000, 600,000, 650,000, 700,000, 750,000, 800,000, 850,000, 900,000, 1 million, 1.5 million, 2 million, 2.5 million, or 3 million nucleotides.

[0210] The selected genomic regions can be located at various positions within a genome, including but not limited to exons, introns, intergenic regions, and other parts. In some embodiments, several probes targeting non-human genomic regions, such as several probes targeting viral genomic regions, may be added.

[0211] In some cases, primers can be used (e.g., via PCR) to specifically amplify several target / biomarkers of interest, thereby enriching the sample with the desired several target / biomarkers (optionally without heterozygous capture). For example, forward and reVerse primers can be prepared for each genomic region of interest and used to amplify several fragments corresponding to or derived from the desired genomic region. Therefore, while this disclosure focuses particularly on cancer assay suites and decoy sets for heterozygous capture, it is broad enough to encompass other methods for enriching cell-free DNA. Thus, a skilled person will recognize, with the aid of this disclosure, several methods similar to those described herein in connection with heterozygous capture, can alternatively be accomplished by replacing heterozygous capture with other enrichment strategies, such as PCR amplification of cell-free DNA fragments corresponding to several genomic regions of interest. In some embodiments, bisulfite padlock probe capture is used to enrich several regions of interest, as described in the patent of Zhang et al. (US2016 / 0340740). In some embodiments, additional or alternative methods are used for enrichment (e.g., non-targeted enrichment), such as reduced representation bisulfite sequencing, methylation restriction enzyme sequencing, methylated DNA immunoprecipitation sequencing, methyl CpG binding domain protein sequencing, methylated DNA capture sequencing, or droplet PCR.

[0212] Probe:

[0213] The cancer assay suite provided herein is an assay suite comprising a set of hybrid probes (also referred to herein as “probes”) designed to target and pull down several nucleic acid fragments of interest upon enrichment for the assay. In some embodiments, the probes are designed to hybridize and enrich DNA or cfDNA molecules from several cancerous samples, processed to convert unmethylated cytosine (C) to uracil (U). In other embodiments, the probes are designed to hybridize and enrich several cancerous samples from a TOO, processed to convert unmethylated cytosine (C) to uracil (U) DNA or cfDNA molecules. The probes may be designed to anneal (or hybridize) to a target (complementary) strand of DNA or RNA. The target strand may be a “positive” strand (e.g., a strand transcribed into mRNA and subsequently translated into protein) or a complementary “negative” strand. In one particular embodiment, a cancer assay suite may include an array of two probes, one probe targeting the positive strand and the other probe targeting the negative strand of a target genomic region.

[0214] For each target genomic region, four possible probe sequences can be designed. The DNA molecules corresponding to or derived from each target region are double-stranded; therefore, a probe or probe set can target a "positive" or forward strand, or its opposite complement (the "negative" strand). Furthermore, in some embodiments, the probes or probe sets are designed to enrich several DNA molecules or fragments that have been treated to convert unmethylated cytosine (C) to uracil (U). Because the probes or probe sets are designed to enrich the converted DNA molecules corresponding to or derived from the target regions, the sequences of the probes can be designed (by applying A at the position of G in the DNA molecules or fragments corresponding to or derived from the target regions where unmethylated cytosine is present) to enrich several DNA molecules in which unmethylated C has been converted to U. In one embodiment, several probes are designed to bind to or hybridize to several DNA molecules or fragments (e.g., hypermethylated or hypomethylated DNA molecules) from several genomic regions known to contain cancer-specific methylation patterns, thereby enriching (or detecting) cancer-specific DNA molecules or fragments. Targeting several genomic regions or several cancer-specific methylation patterns can advantageously allow for the specific enrichment of DNA molecules or fragments identified as providing information about cancer or cancer TOO, and thus reduce detection requirements and costs (e.g., reduce sequencing costs). In other embodiments, two probe sequences (one probe per DNA strand) may be designed for each target genomic region. In yet another example, several probes are designed to enrich all DNA molecules or fragments corresponding to or derived from a target region (i.e., regardless of strand or methylation state). This could be because the cancer methylation state is neither highly methylated nor unmethylated, or because the probes are designed to target several small mutations or other variations, rather than methylation changes, that similarly indicate the presence or absence of a cancer, or the presence or absence of one or more TOOs of a cancer. In this case, all four possible probe sequences could be included for each target genomic region.

[0215] The plurality of probes may be 10, 100, 200, or 300 base pairs in length. The plurality of probes may comprise at least 50, 75, 100, or 120 nucleotides. The plurality of probes may comprise less than 300, 250, 200, or 150 nucleotides. In one embodiment, the plurality of probes comprises 100 to 150 nucleotides. In a particular embodiment, the plurality of probes comprises 120 nucleotides.

[0216] In some embodiments, the plurality of probes are designed in a “2x tiled” manner to cover several overlapping portions of a target region. Each probe may optionally overlap at least partially with another probe in the library in terms of coverage. In such embodiments, the detection combination contains multiple pairs of probes, each of which overlaps with the other by at least 25, 30, 35, 40, 45, 50, 60, 70, 75, or 100 nucleotides. In some embodiments, the overlapping sequence may be designed to be complementary to a target genomic region (or a cff)NA derived from the target genomic region), or to a sequence homologous to a target region or cff)NA. Thus, in some embodiments, at least two probes are complementary to the same sequence in a target genomic region, and a nucleotide fragment corresponding to or derived from the target genomic region may be ligated and pulled down by at least one of the plurality of probes. Other tiled levels are possible, such as 3x tiled, 4x tiled, etc., in which each nucleotide in a target region may bind to more than two probes.

[0217] In one embodiment, each base in a target genomic region is formed by exactly two overlapping probes, as illustrated in the diagram. Figure 1A In the middle ground, if the overlap between the two probes is longer than the target genomic region and extends beyond both ends of the target genomic region, then a single pair of probes is sufficient to pull down a genomic region. In some cases, even relatively small target regions can be targeted by three probes (see [link to relevant documentation]). Figure 1A A probe set comprising three or more probes can optionally be used to capture a large genomic region (see [link]). Figure 1B In some embodiments, several sub-combinations of several probes will collectively extend across an entire genomic region (e.g., may be complementary to several unconverted or converted fragments from said genomic region). A set of probes may optionally include several probes collectively comprising at least two probes that overlap with each nucleotide in said genomic region. This is done to ensure that several cfDNAs comprising a small portion of a target genomic region at one end will have a substantial overlap with at least one probe extending into an adjacent non-target genomic region to provide effective capture.

[0218] For example, a 100 bp cfDNA fragment comprising a 30 nt target genomic region can be ensured to overlap at least 65 bp with at least one of several overlapping probes. Other levels of plating are possible. For example, to increase the target size and add more probes to a detection ensemble, several probes can be designed to expand a 30 bp target region by at least 70 bp, 65 bp, 60 bp, 55 bp, or 50 bp. To capture any fragments that overlap with the target region even slightly (even by only 1 bp), the probes can be designed to extend beyond the ends of the target region on both sides.

[0219] The probes are designed to analyze the methylation status of several target genomic regions (e.g., in humans or another organism) that are suspected of being associated with the presence or absence of cancer in general, the presence or absence of a specific type of cancer, a stage of cancer, or the presence or absence of other types of disease.

[0220] Further, the plurality of probes are designed to efficiently hybridize to and optionally pull down several cfDNA fragments containing a target genomic region. In some embodiments, the plurality of probes are designed to cover several overlapping portions of a target region, such that each probe is “laid out” in coverage, while each probe at least partially overlaps with another probe in the library in terms of coverage. In such embodiments, the detection combination comprises multiple pairs of probes, and each pair of probes comprises at least two probes that overlap each other by an overlapping sequence of at least 25, 30, 35, 40, 45, 50, 60, 70, 75, or 100 nucleotides. In some embodiments, the overlapping sequence may be designed to have a sequence complementary to a target genomic region (or a converted version of a target genomic region), so that a nucleotide fragment derived from or containing the target genomic region can be linked to and optionally pulled down by at least one of the plurality of probes.

[0221] In one embodiment, the minimum target genomic region is 30 bp. When a new target region (based on greedy selection as described above) is added to the detection ensemble, the 30 bp new target region can be centered at a specific CpG site of interest. The new target region is then examined to see if each edge of this new target is close enough to several other targets to allow them to be fused. This is based on a "fusion distance" parameter, which can default to 200 bp but can be adjusted. This allows several close but separate target regions to be enriched with several overlapping probes. Depending on whether there are targets sufficiently close to the left or right of the new target, the new target may not fuse with anything (increasing the number of targets in the detection ensemble by one), fuse with only one target, fuse to the left or right (without changing the number of targets in the detection ensemble), or fuse with existing targets on the left and right (reducing the number of targets in the detection ensemble by one).

[0222] Methods for selecting several target genomic regions:

[0223] In another aspect, several methods are provided for selecting several target genomic regions for detecting cancer and / or a TOO. These target genomic regions can be used to design and manufacture several probes for a cancer assay suite. The methylation status of DNA or cfDNA molecules corresponding to or derived from the several target genomic regions can be screened using the cancer assay suite. Alternative methods, such as WGBS or other methods known in the art, can also be applied to detect the methylation status of several DNA molecules or fragments corresponding to or derived from the several target genomic regions.

[0224] Sample processing:

[0225] Figure 7A This is a flowchart of a procedure 100 for processing a nucleic acid sample and generating several methylation state vectors for several DNA fragments, according to one embodiment. While this disclosure focuses particularly on sequencing-based methods for detecting nucleic acids and determining methylation state, this disclosure is broad enough to encompass other methods for determining the methylation state of several nucleic acid sequences (such as methylation-aware sequencing approaches described in WO2014 / 043763, which is incorporated herein by reference). Figure 7A As described herein, the method includes, but is not limited to, the following steps. For example, any step of the method may include a quantitative sub-step for quality control, or other laboratory testing procedures known to those skilled in the art.

[0226] In step 105, a nucleic acid sample (DNA or RNA) is extracted from a subject. In this disclosure, unless otherwise indicated, DNA and RNA may be used interchangeably. That is, the several embodiments described herein can be applied to both DNA and RNA-type nucleic acid sequences. However, the examples described herein focus on DNA for the purposes of brevity and explanation. The sample can be any combination of the human genome, including the whole genome. The sample may include blood, plasma, serum, urine, feces, saliva, other types of bodily fluids, or any combination thereof. In some embodiments, a method for obtaining a blood sample (e.g., a syringe or finger prick) may be less invasive than a procedure for obtaining a tissue biopsy that may require surgery. The extracted sample may include cfDNA and / or ctDNA. In healthy individuals, the body naturally removes cfDNA and other cellular debris. If a subject has a cancer or disease, the cfDNA and / or ctDNA in the extracted sample may be present at a detectable level for detecting the cancer or disease.

[0227] In step 110, the plurality of cfDNA fragments are processed to convert unmethylated cytosine to uracil. In one embodiment, the method uses a bisulfite treatment of DNA, which converts unmethylated cytosine to uracil without converting methylated cytosine. For example, a commercial kit such as EZ DNA methylation... TM -Gold (EZ DNA Methylation) TM -Gold) kit, EZ DNA methylation TM -Direction (EZ DNAMethylation) TM -Direct) kit or EZ DNA methylation TM -Lightning Set (EZ DNA Methylation) TM The Lightning kit (available from Zymo Research Corp., Irvine, California) is used for the bisulfite conversion. In another embodiment, the conversion of unmethylated cytosine to uracil is achieved using an enzymatic reaction. For example, the conversion can be performed using a commercially available kit for the conversion of unmethylated cytosine to uracil, such as APOBEC-Seq (NEBiolabs, Ipswich, MA).

[0228] In step 115, a sequenced library is prepared. In a first step, an ssDNA adapter is added to the 3′-OH end of a bisulfite-converted ssDNA molecule using an ssDNA ligation reaction. In one embodiment, the ssDNA ligation reaction uses CircLigase II (Epicentre) to ligate the ssDNA adapter to the 3′-OH end of a bisulfite-converted ssDNA molecule, wherein the 5′ end of the adapter is phosphorylated and the bisulfite-converted ssDNA is dephosphorylated (i.e., the 3′ end has a hydroxyl group). In another embodiment, the ssDNA ligation reaction uses a thermostable 5′ AppDNA / RNA ligase (available from New England Biolabs (Ipswich, MA)) to ligate the ssDNA adapter to the 3′-OH end of a bisulfite-converted ssDNA molecule. In this example, the first UMI adapter is adenylated at the 5′ end and blocked at the 3′ end. In another embodiment, the ssDNA ligation reaction uses T4 RNA ligase (available from New England Biolabs) to ligate the ssDNA translocator to the 3′-OH end of a bisulfite-converted ssDNA molecule. In a second step, a second strand of DNA is synthesized in an extension reaction. For example, an extension primer heterozygous for a primer sequence included in the ssDNA translocator is used in a primer extension reaction to form a double-stranded bisulfite-converted DNA molecule. Optionally, in one embodiment, the extension reaction uses an enzyme capable of reading through several uracil residues in the bisulfite-converted template strand. Optionally, in a third step, a dsDNA translocator is added to the double-stranded bisulfite-converted DNA molecule. Finally, the double-stranded bisulfite-converted DNA is amplified to add several sequence translocators. For example, PCR amplification using a forward primer including a P5 sequence and a reverse primer including a P7 sequence is employed to add the P5 and P7 sequences to the bisulfite-converted DNA. Optionally, during library preparation, unique molecular identifiers (UMIs) can be added to the plurality of nucleic acid molecules (e.g., DNA molecules) via transposon ligation. The plurality of UMIs are short nucleic acid sequences (e.g., 4 to 10 base pairs) added to the ends of the plurality of DNA fragments during transposon ligation. In some embodiments, the UMI is a plurality of degenerate base pairs that serve as a unique tag that can be used to identify several sequence reads originating from a specific DNA fragment. In PCR amplification following transposon ligation, the plurality of UMIs, along with the ligated DNA fragment, are replicated, providing a method for identifying several sequence reads from the same original fragment in downstream analysis.

[0229] In step 120, several target DNA sequences may be enriched from the library. This is exemplified when a target detection combinatorial assay is performed on several samples. During enrichment, several hybrid probes (also referred to herein as “probes”) are used to target and pull down several nucleic acid fragments that provide information about the presence or absence of cancer (or disease), cancer status, or a cancer classification (e.g., cancer type or tissue of origin). For a given workflow, the probes may be designed to adhere (or hybridize) to a target (complementary) strand of DNA or RNA. The target strand may be a “positive” strand (e.g., a strand transcribed into mRNA and subsequently translated into a protein) or a complementary “negative” strand. The length of the probes may be in the range of 10S, 100S, or 1000S base pairs. Furthermore, the probes may cover several overlapping portions of a target region.

[0230] Following a hybridization step 120, the hybridized nucleic acid fragments are captured and can be amplified using PCR (enrichment 125). For example, the target sequences can be enriched to obtain enriched sequences, which can then be sequenced. Generally, any method known in the art can be used to isolate and enrich probe-hybridized target nucleic acids. For example, as is well known in the art, a biotinylate portion can be added to the 5′ end of the probes (i.e., biotinylation) using a streptavidin-coated surface (e.g., streptavidin-coated beads) to facilitate the isolation of target nucleic acids hybridized to the probes.

[0231] In step 130, several sequence reads are generated from the several enriched DNA sequences, for example, several enriched sequences. Sequence data can be obtained from the several enriched DNA sequences using methods known in the art. For example, the methods may include next-generation sequencing (NGS) technologies, including synthetic sequencing (Illumina), pyrosequencing (454 Life Sciences), ion semiconductor technology (Ion Torrent sequencing), single-molecule real-time sequencing (Pacific Biosciences), linker-based sequencing (SOLiD sequencing), nanopore sequencing (Oxford Nanopore Technologies), or paired-end sequencing. In some embodiments, massively parallel sequencing is performed using synthetic sequencing with reversible dye terminators. In other embodiments, any known methods for detecting nucleic acids and determining methylation status, as will be readily understood by those skilled in the art, may be used. For example, using known methylation detection sequencing (see, for example, WO2014 / 043763), a DNA microarray (e.g., with several labeled probes adhered to or bound to a solid surface or DNA array wafer), several sequences can be detected and the methylation state determined.

[0232] In step 140, several methylation state vectors are generated from the several sequence reads. To do this, a sequence read is aligned to a reference genome. The reference genome helps provide context regarding the location of the cfDNA fragment within a human genome. In a simplified example, the sequence reads are aligned so that three CpG sites are associated with CpG sites 23, 24, and 25 (an arbitrary reference identifier used for ease of description). After alignment, information is available regarding the methylation state of all CpG sites on the cfDNA fragment and the location within the human genome to which these CpG sites are mapped. With the methylation state and location, a methylation state vector can be generated for the cfDNA fragment.

[0233] The emergence of data structures:

[0234] Figure 3AThis is a flowchart describing a procedure 300 for generating a data structure for a health control group according to one embodiment. To create a health control group data structure, the analysis system obtains information about the methylation status of several CpG sites on several sequence reads, said several sequence reads being derived from several DNA molecules or fragments from several healthy subjects. The method provided herein for creating a health control group data structure can similarly be performed on several subjects having cancer, several subjects having a TOO type of cancer, several subjects having a known cancer type, or several subjects having another known disease state. A methylation status vector is generated for each DNA molecule or fragment, for example, via said procedure 100.

[0235] Using the methylation state vector of each fragment, the analysis system subdivides the methylation state vector 310 into several strings of several CpG sites. In one embodiment, the analysis system subdivides the methylation state vector 310 such that the resulting strings are all less than a given length. For example, a methylation state vector of length 11 can be subdivided into several strings, with lengths less than or equal to 3 resulting in 9 strings of length 3, 10 strings of length 2, and 11 strings of length 1. In another example, a methylation state vector of length 7 subdivided into strings of length less than or equal to 4 will result in 4 strings of length 4, 5 strings of length 3, 6 strings of length 2, and 7 strings of length 1. If a methylation state vector is shorter than or equal to the length of a specific string, the methylation state vector can be converted into a single string containing all CpG sites of the vector.

[0236] The analysis system records 320 the number of possible strings in the control group that have the specific CpG site as the first CpG site in the string and have the methylation state for each possible CpG site and methylation state in the vector. For example, at a given CpG site, and considering a string length of 3, there are 2^3 or 8 possible string configurations. At that given CpG site, for each of the 8 possible string configurations, the analysis system records 320 how many times each methylation state vector might occur in the control group. Continuing this example, this might involve recording the following quantities for each starting CpG site x in the reference genome: <M x M x+1 M x+2 >、 <M x M x+1 U x+2 >、...、<U x U x+1 Ux+2 The analysis system creates a 330 data structure that stores the recorded count for each starting CpG site and the likelihood of a string.

[0237] Setting an upper limit on string length has several benefits. First, depending on the maximum string length, the size of the data structure created by the analysis system can increase significantly. For example, a maximum string length of 4 means that for each CpG site with several strings of length 4, there are at least 2^4 numbers to record. Increasing the maximum string length to 5 means that each CpG site has an additional 2^4 or 16 numbers to record, doubling the number of numbers to be recorded (and the required computer memory) compared to the previous string length. Reducing the string size helps to keep the data structure creation and performance (e.g., for subsequent evaluation as described below) computationally and in storage reasonable. Second, a statistical consideration for limiting the maximum string length is to avoid overfitting several downstream models that use the string counts. If several long strings of several CpG sites are not biologically effective in predicting outcomes (e.g., the presence of cancer-predicting anomalies), calculating several probabilities based on these large strings can be problematic. This is because it requires potentially vast amounts of unavailable data and would therefore be far too sparse for a model to function properly. For example, calculating a probability of an anomaly / cancer conditioned on the previous 100 CpG sites would require counting several strings of length 100 in a data structure, ideally some of which exactly match the previous 100 methylation states. If only sparse counts of these 100-length strings are available, the data would be insufficient to determine whether a given 100-length string in a test sample is an anomaly.

[0238] Data structure verification:

[0239] Once the data structure has been created, the analysis system can seek to validate the data structure and / or any downstream models using the data structure. One type of validation checks the consistency within the data structure of the control group. For example, if there are any outlier objects, samples, and / or fragments in a control group, the analysis system can perform various calculations to determine whether to remove any fragments from one of these categories. In a representative example, the healthy control group may include a sample that is undiagnosed but cancerous, causing the sample to include several abnormal methylation fragments. This first type of validation ensures that potentially cancerous samples are removed from the healthy control group without affecting the purity of the control group.

[0240] A second type of verification checks the probability model used to calculate the p-value using several counts from the data structure itself (i.e., from the health control group). A procedure for calculating the p-value is described below along with... Figure 5 The analysis system, once generating a p-value for several methylated state vectors in the validation set, constructs a cumulative density function (CDF) based on these p-values. The analysis system can then perform various calculations on the CDF to validate the data structure of the control set. A test utilizes the fact that the CDF should ideally be at or below an identity function such that CDF(x) ≤ x. Conversely, above the identity function reveals some flaws in the probabilistic model used for the data structure of the control set. For example, if a fragment of 1 / 100 has a p-value fraction of 1 / 1000, meaning CDF(1 / 1000) = 1 / 100 > 1 / 1000, then the second type of validation fails, indicating a problem with the probabilistic model.

[0241] A third type of validation uses a healthy group of validation samples, separated from the validation samples used to construct the data structure. This third type of validation tests whether the data structure is properly constructed and whether the model functions. An exemplary procedure for performing this type of validation is described below along with... Figure 3B The third type of validation quantifies how well the health control group generalizes the distribution of several healthy samples. If the third type of validation fails, the health control group does not generalize well to the distribution of health.

[0242] A fourth type of validation is performed using several samples from a non-healthy validation group. The analysis system calculates several p-values ​​for the non-healthy validation group and constructs the CDF. For a non-healthy validation group, the analysis system expects to see that for at least some samples, CDF(x) > x. Or, in other words, the opposite of what is expected for the healthy control group and the healthy validation group in the second and third types of validation. If the fourth type of validation fails, it indicates that the model has not properly identified the anomalies that the model was designed to identify.

[0243] Figure 3B It is a flowchart describing an embodiment of the process. Figure 3AThe control group verifies the data structure in an additional step 340. In this embodiment of step 340, the analysis system performs a fourth type of verification test as described above, which applies a verification group having a composition of objects, samples, and / or fragments assumed to be similar to the control group. For example, if the analysis system selects several healthy subjects without cancer as the control group, the analysis system also uses several healthy subjects without cancer in the verification group.

[0244] The analysis system takes the verification group, and as in Figure 3A The system generates methylation state vectors in groups of 100, as described above. The analysis system performs a p-value calculation for each methylation state vector from the validation group. The p-value calculation procedure will be combined with... Figures 4 to 5 And further described. For each possibility of a methylation state vector, the analysis system calculates a probability from the data structure of the control group. Once the plurality of probabilities for the plurality of possibilities of the plurality of methylation state vectors have been calculated, the analysis system calculates a p-score for that methylation state vector based on the plurality of calculated probabilities. The p-score represents an expectation of finding that particular methylation state vector in the control group and other possible methylation state vectors with even lower probabilities. Thus, a low p-score generally corresponds to a methylation state vector that is less expected relative to other methylation state vectors in the control group, while a high p-score generally corresponds to a methylation state vector that is more expected relative to other methylation state vectors found in the control group. Once the analysis system produces a p-score for the plurality of methylation state vectors in the validation group, the analysis system constructs a cumulative density function (CDF) from the plurality of p-scores from the validation group. The analysis system verifies the consistency of the CDF in the fourth type of validation test as described above.

[0245] Aberrantly methylated fragments:

[0246] According to Figure 4 In one embodiment outlined herein, several aberrant methylation fragments with abnormal methylation patterns in cancer patient samples, objects with a TOO cancer, several objects with a known cancer type, or objects with another known disease state are selected as target genomic regions. An exemplary procedure 440 for selecting aberrant methylation fragments is visually illustrated in… Figure 5 In, and in Figure 4The description is further elaborated below. In procedure 400, the analysis system generates 100 methylation state vectors from several cfDNA fragments of the sample. The analysis system processes each methylation state vector as follows.

[0247] For a given methylation state vector, the analysis system enumerates 410 all possibilities of methylation state vectors that have the same starting CpG site and the same length (i.e., the set of CpG sites). Since each methylation state can be methylated or unmethylated, there are only two possible states at each CpG site, and therefore the unique count of possible methylation state vectors depends on powers of 2, such that a methylation state vector of length n will be associated with 2^n possibilities of methylation state vectors.

[0248] The analysis system evaluates the health control group data structure and calculates 420 probabilities of observing each possible methylation state vector for each of the identified initiation CpG sites / methylation state vector lengths. In one embodiment, the calculation of the probability of observing a given possible probability uses a Markov chain probability to model the joint probability calculation, which will be referenced below. Figure 5 And described in more detail. In other embodiments, a method for calculating probabilities other than those of Markov chains is used to determine each possible probability of observing a methylation state vector.

[0249] The analysis system uses each of the several possible probabilities to calculate a p-value score for the methylation state vector. In one embodiment, this includes identifying the possible calculated probabilities that correspond to the considered methylation state vector. Specifically, this is the probability of having the same set of CpG sites as the methylation state vector, or similarly having the same starting CpG sites and length. The analysis system sums the several calculated probabilities to produce the p-value score. The several calculated probabilities are several possible calculated probabilities, and the several possibilities may have any probability that is less than or equal to the identified probability.

[0250] This p-value represents the probability that the methylation state vector of the fragment or other, even less likely, methylation state vectors are observed in the healthy control group. Thus, a low p-value score generally corresponds to a methylation state vector that is rare in a healthy individual, and causes the fragment to be labeled as anomalously methylated relative to the healthy control group. A high p-value score is generally associated with a methylation state vector that is expected to be present in a healthy subject in a relative sense. For example, if the healthy control group is a non-cancer group, a low p-value indicates that the fragment is anomalously methylated relative to the non-cancer group, and therefore may indicate the presence of cancer in the tested subject.

[0251] As described above, the analysis system calculates a p-value score for each of several methylation state vectors, each representing a cfDNA fragment in the test sample. To identify which of the fragments is aberrantly methylated, the analysis system can filter the set of several methylation state vectors based on their p-value scores. In one embodiment, filtering is performed by comparing the p-value scores to a threshold and retaining only those fragments below the threshold. This threshold p-value score can be on the order of 0.1, 0.01, 0.001, 0.0001, or similar.

[0252] P-value score calculation:

[0253] Figure 5 This is a diagram 500 illustrating the calculation of an exemplary p-value score according to one embodiment. To calculate a p-value score for a given detected methylation state vector 505, the analysis system takes the detected methylation state vector 505 and several possible enumeration 410 methylation state vectors. In this exemplary example, the detected methylation state vector 505 is <M 23 M 24 M 25 U 26 Since the length of the detected methylation state vector 505 is 4, there are 2^4 possible methylation state vectors containing CpG sites 23 to 26. In a general example, the possible number of methylation state vectors is 2^n, where n is the length of the detected methylation state vector or, alternatively, the length of the sliding window (described further below).

[0254] The analysis system calculates the listed probabilities of 420 methylation state vectors. Because methylation conditionally depends on the methylation state of nearby CpG sites, one method for calculating the probability of observing a given methylation state vector is to use a Markov chain model. Generally, a methylation state vector, such as <S1, S2, ..., S...>, is considered as a given methylation state vector. n >(where S represents the methylation state, or methylated (represented by M), unmethylated (represented by U), or indeterminate (represented by I)) has a joint probability, which can be expanded using the chain rule of probabilities as follows:

[0255] P( <S1,S2,...,S n >)=P(S n |S1,...,S n-1 )*P(S n-1 |S1,...,S n-2 (1)

[0256] ...*P(S2|S1)*P(S1).

[0257] Markov chain models can be used to make the calculation of each possible conditional probability more efficient. In one embodiment, the analysis system selects a Markov chain level k, which corresponds to how many previous CpG sites in the vector (or window) are to be considered in the conditional probability calculation, such that the conditional probability is modeled as P(S n |S1,...,S n-1 )~P(S n |S n-k-2 S n-1 ).

[0258] To calculate the probability of each possible Markov-modeled methylation state vector, the analysis system accesses the data structure of the control group, specifically the counts of various strings of several CpG sites and states. To calculate P(M... n |S n-k-2 S n -1), the analysis system conforms to <S n-k-2 S n-1 M n The data structure takes the number of stored counts of several strings, divided by the number of strings conforming to <S>. n-k-2 S n-1 M n > and <S n-k-2 S n-1 Un The ratio of the number of strings in the data structure to the sum of the stored counts. Therefore, P(M) n |S n-k-2 S n-1 () is a calculated ratio, having the following form:

[0259]

[0260] The calculation can be further smoothed by applying a prior distribution. In one embodiment, the prior distribution is a uniform prior, as in Laplace smoothing. As an example, a constant is added to the numerator of the equation and another constant (e.g., twice the constant in the numerator) is added to the denominator. In other embodiments, an algorithmic technique, such as Knesser-Ney smoothing, is used.

[0261] In the illustration, the formula described above is applied to the detection methylation state vector 505 covering sites 23 to 26. Once the calculated probability 515 is completed, the analysis system calculates a p-value score 525, which is added to a total number of probabilities, the total number of probabilities being less than or equal to the possible probability of a methylation state vector matching the detection methylation state vector 505.

[0262] In one embodiment, the computational burden of computational probability and / or p-value scores can be further reduced by caching at least some computations. For example, the analysis system can cache the calculation of the probabilities of several methylation state vectors (or windows thereof) in temporary or permanent memory. If other fragments have the same CpG site, caching the probabilities allows for efficient calculation of p-value scores without recalculating the potential probabilities. Ultimately, the analysis system can calculate p-value scores for each of the probabilities of several methylation state vectors associated with a set of CpG sites from a vector (or a window thereof). The analysis system can cache these p-value scores for use in determining the p-value scores of other fragments that include the same CpG site. Generally, the possible p-value scores of methylation state vectors having the same CpG site can be used to determine the p-value score of the possible different ones from the same set of CpG sites.

[0263] Sliding window:

[0264] In one embodiment, the analysis system uses a 435-sliding window to determine the possibilities of the methylation state vector and calculate p-values. The analysis system lists possibilities and calculates p-values ​​only for a window of a consecutive number of CpG sites, rather than for the entire methylation state vector, wherein the window is shorter in length (of CpG sites) than at least some segments (otherwise, the window would be ineffective). The window length can be static, user-determined, dynamic, or otherwise selected.

[0265] When calculating the p-value for a methylation state vector larger than the window, the window, starting from the first CpG site in the vector, identifies a consecutive set of CpG sites from the vector within the window. The analysis system calculates a p-value score for the window including the first CpG site. The analysis system then "slides" the window to the second CpG site in the vector and calculates another p-value score for the second window. Therefore, for a window of size l and a methylation vector length m, each methylation state vector will generate m-l+1 p-value scores. After completing the p-value calculation for each portion of the vector, the lowest p-value score from all sliding windows is taken as the overall p-value score for the methylation state vector. In another embodiment, the analysis system sums the p-value scores of the plurality of methylation state vectors to generate an overall p-value score.

[0266] Using the sliding window helps reduce the number of possible methylation state vectors that can be enumerated, and the corresponding probability calculations that would otherwise need to be performed. An example probability calculation is shown in... Figure 5 In general, however, the number of possible methylation state vectors increases exponentially with the size of the methylation state vector. To give a realistic example, a fragment may have more than 54 CpG sites. The analysis system could use a window of size 5 for that fragment, resulting in 50 p-value calculations performed on each of the 50 windows of the methylation state vector, instead of calculating 2^54 (approximately 1.8 × 10^16) possible probabilities to produce a single p-value score. Each of the 50 calculations enumerates 2^5 (32) possibilities for the methylation state vector, resulting in a total of 50 × 2^5 (1.6 × 10^3) ​​probability calculations. This leads to a significant reduction in the number of calculations to be performed that lack meaningful hits for accurate identification of anomalous fragments. This additional step could also be applied when validating the control group 340 with several methylation state vectors of the validation group.

[0267] Identify segments that indicate cancer:

[0268] The analysis system identifies 450 DNA fragments that indicate cancer from a filtered group of aberrantly methylated fragments.

[0269] Hypomethylated and hypermethylated fragments:

[0270] According to a first method, the analysis system can identify several DNA fragments considered hypomethylated or hypermethylated from the filtered group of aberrantly methylated fragments as indicators of cancer. The several hypomethylated or hypermethylated fragments can be defined as several fragments of a specific length (e.g., more than 3, 4, 5, 6, 7, 8, 9, 10, etc.) with several CpG sites, the fragments having a high percentage of methylated CpG sites (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage in the range of 50% to 100%) or a high percentage of unmethylated CpG sites (e.g., more than 80%, 85%, 90%, or 95%, or any other percentage in the range of 50% to 100%).

[0271] Probability model:

[0272] According to a method described herein, the analysis system applies a probabilistic model fitted to methylation patterns for each cancer type and non-cancer type to identify several fragments indicating cancer. The analysis system uses several DNA fragments from the several genomic regions to calculate a log-probability ratio for a sample, considering various cancer types using the fitted probabilistic model for each cancer type and non-cancer type. The analysis system can determine that a DNA fragment indicates cancer based on whether at least one of the several log-probability ratios considered relative to the various cancer types is above a threshold.

[0273] In one embodiment of genome segmentation, the analysis system divides the genome into several regions through several stages. In a first stage, the analysis system separates the genome into several blocks of several CpG sites. Each block is defined when there is an interval between two adjacent CpG sites exceeding some threshold, for example, greater than 200bp, 300bp, 400bp, 500bp, 600bp, 700bp, 800bp, 900bp, or 1000bp. From each block, the analysis system further subdivides each block into several regions of a specific length, for example, 500bp, 600bp, 700bp, 800bp, 900bp, 1000bp, 1100bp, 1200bp, 1300bp, 1400bp, or 1500bp, in a second stage. The analysis system may further overlap with several neighboring regions at a percentage of the length, for example, 10%, 20%, 30%, 40%, 50%, or 60%.

[0274] The analysis system analyzes several sequence reads derived from several DNA fragments for each region. The system can process several samples from tissues and / or high-signal cfDNA. High-signal cfDNA samples can be determined by a binary classification model based on cancer stage or by other metrics.

[0275] For each cancer type and non-cancer, the analysis system fits a separate probability model to several fragments. In one embodiment, each probability model is a mixture model comprising a combination of several mixture components, and each mixture component is an independent site model, wherein methylation at each CpG site is assumed to be independent of the methylation state at other CpG sites.

[0276] In several alternative embodiments, calculations are performed for each CpG site. Specifically, a first count is determined, which is the number of cancerous samples (cancer_count) that include an aberrantly methylated DNA fragment overlapping with the CpG, and a second count is determined, which is the total number (sum) of samples in the group containing a fragment overlapping with the CpG. Several genomic regions may be selected based on said several counts, for example, based on a criterion that is positively correlated with the number of cancerous samples (cancer_count) that include a DNA fragment overlapping with the CpG and negatively correlated with the total number (sum) of samples in the group containing a fragment overlapping with the CpG.

[0277] Various types of cancer with different TOOs can be selected from the following groups: breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, renal pelvis and urethral epithelial carcinoma, renal cancer other than urethral epithelial carcinoma, prostate cancer, anorectal cancer, anal cancer, colorectal cancer, hepatobiliary cancer originating from hepatocytes, hepatobiliary cancer originating from cells other than hepatocytes, liver / cholecystitis, esophageal cancer, pancreatic cancer, upper gastrointestinal squamous cell carcinoma, upper gastrointestinal cancer other than squamous cells, head and neck cancer, lung cancer, lung adenocarcinoma, small cell lung cancer, squamous cell lung cancer and cancer other than lung adenocarcinoma or small cell lung cancer, neuroendocrine carcinoma, melanoma, thyroid cancer, sarcoma, plasmacytoma, multiple myeloma, myeloma, lymphoma and leukemia.

[0278] In some embodiments, various cancer types can be classified and labeled using classification methods available in the art, such as the International Classification of Cancer Diseases (ICD-O-3) (codes.iarc.fr) or the Surveillance, Epidemiology and End Outcomes (SEER) program (seer.cancer.gov). In other embodiments, cancer types are classified using three orthogonal codes: (i) location code, (ii) morphology code, or (iii) behavior code. Under the behavior code, benign tumors are 0, indeterminate behavior is 1, carcinoma in situ is 2, malignant, primary site is 3 and malignant, and metastatic site is 6.

[0279] In some embodiments, a cancer TOO may be selected from a group defined by guidelines that will be used to stage a detected cancer. For example, the following references identify several groups of different cancers that are co-staged following standard guidelines: Amin, MB, Edge, S., Greene, F., Byrd, DR, Brookland, RK, Washington, MK, Gershenwald, JE, Compton, CC, Hess, KR, Sullivan, DC, Jessup, JM, Brierley, JD, Gaspar, LE, Schilsky, RL, Balch, CM, Winchester, DP, Asare, EA, Madera, M., Gress, DM, Meyer, LR. (Editors), AJCC Cancer Staging Guidelines, 8th Edition, Springer, 2017. Typically, such staging is the next step in cancer management after cancer detection and diagnosis.

[0280] The analysis system can further consider various cancer types with a fitted probability model for each cancer type and non-cancer type, or for a cancer TOO, calculating a log-probability ratio (“R”) for a segment indicating the likelihood that the segment is cancerous. The two probabilities can be derived from probability models fitted for each cancer type and non-cancer type, said probability models being defined to calculate the likelihood of observing a methylation pattern on a segment given each of the said cancer types and non-cancer types. For example, said probability models can be defined to be fitted for each of the said cancer types and non-cancer types.

[0281] Selection of genomic regions that indicate cancer:

[0282] The analysis system identifies 460 genomic regions that indicate cancer. To identify these information-providing regions, the analysis system calculates an information gain for each genomic region, or more specifically for each CpG site, which describes the ability to distinguish between various outcomes.

[0283] A method for identifying several genomic regions that can distinguish between cancer and non-cancer types involves applying a trained classification model. This trained model can be applied to groups of aberrantly methylated DNA molecules or fragments corresponding to or derived from a cancerous or non-cancer group. The trained classification model can be trained to identify any cases of interest that can be identified from the several methylation state vectors.

[0284] In one embodiment, the trained classification model is a binary classifier trained based on the methylation status of several cfDNA fragments or several genomic sequences obtained from a population of subjects with cancer or a cancer TOO, and a population of healthy subjects without cancer. The binary classifier is then used to classify the probability that a detected subject has cancer, a cancer TOO, or does not have cancer based on several anomalous methylation status vectors. In other embodiments, different classifiers may be trained using populations of subjects known to have specific cancers (e.g., breast cancer, lung cancer, prostate cancer, etc.); cancers known to have specific TOOs believed to originate from those TOOs; or specific cancers known to be at different stages (e.g., breast cancer, lung cancer, prostate cancer, etc.). In these embodiments, several different classifiers may be trained using several sequence reads obtained from samples rich in tumor cells. These samples are from populations of subjects known to have specific cancers (e.g., breast cancer, lung cancer, prostate cancer, etc.). The ability of each genomic region to distinguish between cancer and non-cancer types in the classification model is used to rank the genomic regions by classification performance, from most informative to least informative. The analysis system can identify genomic regions from this ranking, which is based on the information gain of the classification between non-cancer and cancer types.

[0285] Calculate the information gain from hypomethylated and hypermethylated fragments that indicate cancer:

[0286] According to one embodiment, using several fragments indicating cancer, the analysis system can be based on a plot shown in [the diagram / image / etc.]. Figure 6AA procedure 600 trains a classifier. The procedure 600 accesses two training groups of samples: a non-cancer group and a cancer group, and, for example, via step 440 from procedure 400, obtains 605 a non-cancer group and a cancer group comprising a number of methylation state vectors, including a number of anomalous methylation fragments.

[0287] For each methylation state vector, the analysis system determines 610 whether the methylation state vector indicates cancer. Here, if at least a number of CpG sites have a specific state (methylated or unmethylated, respectively) and / or sites with a threshold ratio are in the specific state (again, methylated or unmethylated, respectively), several fragments indicating cancer can be defined as overmethylated or hypomethylated fragments. In one embodiment, if several cfDNA fragments overlap with at least 5 CpG sites, and at least 80%, 90%, or 100% of the CpG sites of the several cfDNA fragments are methylated, or at least 80%, 90%, or 100% of the CpG sites of the several cfDNA fragments are unmethylated, then the several cfDNA fragments are identified as hypomethylated or overmethylated, respectively.

[0288] In an alternative embodiment, the procedure considers several portions of the methylation state vector and determines whether a portion is hypomethylated or hypermethylated, and can distinguish between hypomethylated and hypermethylated portions. This alternative solves the problem of missing several large methylation state vectors containing at least one densely hypomethylated or hypermethylated region. This procedure for defining hypomethylation and hypermethylation can be performed in... Figure 4 The steps in step 450 are applied. In another embodiment, the plurality of segments indicating cancer can be defined based on a plurality of probabilities output from a plurality of trained probability models.

[0289] In one embodiment, the analysis system generates a hypomethylation fraction (P) of 620 for each CpG site in the genome. 低 ) and permethylation fraction (P 过To generate two scores at a given CpG site, the classifier takes four counts at that CpG site: (1) a count of several (methylation state) vectors of the cancer group that overlap with the CpG site and are labeled as hypomethylated; (2) a count of several vectors of the cancer group that overlap with the CpG site and are labeled as hypermethylated; (3) a count of several vectors of the non-cancer group that overlap with the CpG site and are labeled as hypomethylated; and (4) a count of several vectors of the non-cancer group that overlap with the CpG site and are labeled as hypermethylated. Furthermore, the procedure can normalize these counts for each group to account for group size differences between the non-cancer and cancer groups. In several alternative embodiments where the multiple segments indicating cancer are used more generally, the multiple scores can be more broadly defined as a count of multiple segments indicating cancer at each genomic region and / or CpG site.

[0290] In one embodiment, to generate a hypomethylation fraction of 620 at a given CpG site, the procedure takes (1) and divides it by a ratio of the sum of (1) and (3). Similarly, the hypermethylation fraction is calculated by taking (2) and dividing it by a ratio of (2) and (4). Furthermore, these ratios can be calculated using additional smoothing techniques as discussed above. Given the presence of hypomethylation or hypermethylation in several fragments from the cancer group, the hypomethylation fraction and the hypermethylation fraction are associated with an estimate of the cancer probability.

[0291] The analysis system generates a total hypomethylation score and a total hypermethylation score for each anomalous methylation state vector. The total hypermethylation and hypomethylation scores are determined based on the hypermethylation and hypomethylation scores of the plurality of CpG sites in the methylation state vector. In one embodiment, the total hypermethylation and hypomethylation scores are respectively recorded as the maximum hypermethylation and hypomethylation scores of the plurality of sites in each state vector. However, in several alternative embodiments, the plurality of total scores may be based on the mean, median, or other calculations using the plurality of hypermethylation / hypomethylation scores of the plurality of sites in each vector.

[0292] The analysis system then ranks 640 objects based on all their methylation state vectors, resulting in two rankings for each object. The ranking 640 is based on a total low-methylation score and a total overmethylation score based on the plurality of methylation state vectors. The process selects a plurality of total low-methylation scores from the low-methylation rankings and a plurality of total overmethylation scores from the overmethylation rankings. Using the selected scores, the classifier generates a single feature vector for each object. In one embodiment, the scores selected from the two rankings are chosen in a fixed order, which is the same for each generated feature vector of each object in each of the plurality of training groups. As an example, in one embodiment, the classifier selects the first, second, fourth, and eighth total overmethylation scores from each ranking, and the same for each total low-methylation score, and writes these scores into the feature vector of the object.

[0293] The analysis system trains a binary classifier to distinguish feature vectors from the cancer and non-cancer training groups. Generally, any of several classification techniques can be used. In one embodiment, the classifier is a non-linear classifier. In a particular embodiment, the classifier is a non-linear classifier applying L2-regularized kernel logistic regression with a Gaussian radial basis function (RBF) kernel.

[0294] Specifically, in one embodiment, the number (n) of non-cancer samples or (several) different cancer types having an aberrant methylation fragment overlapping a CpG site. 其它 and the number of cancer samples or (several) cancer types (n) 癌症 The number of samples is counted. Then, the probability that a sample is cancerous is estimated by a score (“S”), which is correlated with n. 癌症 Positively correlated with n 其它 They are negatively correlated. The scores can be expressed using the equation: (n 癌症 +1) / (n 癌症 +n 其它 +2) or (n 癌症 ) / (n 癌症 +n 其它The analysis system calculates an information gain of 670 for each cancer type and for each genomic region or CpG site to determine whether the genomic region or CpG site indicates cancer. The information gain is calculated for several training samples with a given cancer type, compared to all other samples. For example, two random variables are used: “Anomalous Fragment” (“AF”) and “Cancer Type” (“CT”). In one embodiment, AF, as determined by the anomalous score / eigenvector above, is a binary variable indicating whether an anomalous fragment overlaps with a given CpG site in a given sample. CT is a random variable indicating whether the cancer belongs to a specific type. The analysis system calculates mutual information about CT given AF. That is, how many bits of information about the cancer type are obtained if it is known whether an anomalous fragment overlaps with a specific CpG site.

[0295] For a given cancer type, the analysis system uses this information to rank several CpG sites based on how cancer-specific they are. This procedure is repeated for all cancer types considered. If a particular region is generally aberrantly methylated in several training samples for a given cancer, but not generally aberrantly methylated in several training samples for other cancer types or in several healthy training samples, then several CpG sites overlapping these aberrant fragments will tend to have high information gain for the given cancer type. For each ranked CpG site for each cancer type, based on their ranking, they are greedily added (selected) to a selected group of CpG sites for use in the cancer classifier.

[0296] Pairwise information gain is calculated from the fragments indicating cancer identified by the probabilistic model:

[0297] Several fragments indicating cancer, identified according to a method described herein, can be analyzed based on... Figure 6BThe program 680 in the analysis system identifies several genomic regions. The analysis system defines a feature vector 690 for each sample, each region, and each cancer type, defined by a count of several DNA fragments having a probability higher than several thresholds, the fragments indicating a calculated log-probability ratio for the cancer, where each count is a value in the feature vector. In one embodiment, the analysis system counts the number of fragments present in a sample for a region for each cancer type having a log-probability ratio higher than one or several possible thresholds. The analysis system defines a feature vector for each sample by counting several DNA fragments for each genomic region for each cancer type providing a calculated log-probability ratio higher than several thresholds for the fragments, where each count is a value in the feature vector. The analysis system uses the defined feature vectors to calculate an information score for each genomic region, the information score describing the ability of the genomic region to distinguish between each pair of cancer types. For each pair of cancer types, the analysis system ranks several regions based on the information scores. The analysis system may select several regions based on the ranking of the information scores.

[0298] The analysis system calculates an information score of 695 for each region, which describes the region's ability to distinguish between each pair of cancer types. For each different pair of cancer types, the analysis system can designate one type as a positive type and the other as a negative type. In one embodiment, the ability of a region to distinguish between the positive and negative types is based on mutual information, using the estimated fractions of at least one fragment of the layer from the positive and negative type cfDNA samples that are expected to be non-zero in the final assay using the feature. These fractions are estimated using the observed rate of occurrence of the feature in healthy cfDNA, in high-signal cfDNA, and / or in tumor samples of each cancer type. For example, if a feature occurs frequently in healthy cfDNA, then the feature will also be expected to occur frequently in cfDNA of any cancer type and may result in a low information score. The analysis system can select a specific number of regions from the ranking for each pair of cancer types, for example, 1024.

[0299] In several additional embodiments, the analysis system further identifies predominantly hypermethylated or hypomethylated regions from the ranking of several regions. The analysis system may load the group of several fragments into the positive type(s) for a region identified as providing information. The analysis system evaluates whether the several loaded fragments are predominantly hypermethylated or hypomethylated from the loaded fragments. If the several loaded fragments are predominantly hypermethylated or hypomethylated, the analysis system may select several probes corresponding to the predominant methylation pattern. If the several loaded fragments are not predominantly hypermethylated or hypomethylated, the analysis system may use a mixture of several probes to target both hypermethylation and hypomethylation. The analysis system may further identify a minimum group of CpG sites that overlap with some ratios of the several fragments.

[0300] In other embodiments, the analysis system, after ranking the regions based on several information scores, labels each region with the lowest information ranking across all cancer type pairs. For example, if a region is the 10th most informative region for distinguishing between breast and lung cancer, and the 5th most informative region for distinguishing between breast and colorectal cancer, then that region will be assigned an overall label of "5". The analysis system can design several probes starting from the lowest-labeled regions and add several regions to the detection ensemble, for example, until the size budget of the detection ensemble is exhausted.

[0301] Off-target genomic regions:

[0302] In some embodiments, several probes targeting several selected genomic regions are further filtered based on the number of their off-target regions.475 This is to screen for probes that pull down too many cfDNA fragments corresponding to or derived from off-target genomic regions. Excluding probes with many off-target regions is valuable by reducing the off-target rate and increasing the target coverage of a given amount of sequencing.

[0303] An off-target genomic region is a genomic region with sufficient homology to a target genomic region such that DNA molecules or fragments derived from several off-target genomic regions heterozygosize to a probe designed to heterozygosize to a target genomic region and are pulled down by the probe. An off-target genomic region may be aligned to a probe along at least 35 bp, 40 bp, 45 bp, 50 bp, 60 bp, 70 bp, or 80 bp with at least 80%, 85%, 90%, 95%, or 97% congruence. In one embodiment, an off-target genomic region is a genomic region (or a transformed sequence of the same region) aligned to a probe along at least 45 bp with at least 90% congruence. Various methods known in the art can be employed to screen for several off-target genomic regions.

[0304] Thoroughly searching the genome to find all off-target genomic regions can be computationally challenging. In one embodiment, a k-mer seeding strategy (which may allow one or more mismatches) is incorporated into the local alignment at the seed site. In this case, a thorough search with good alignment can be guaranteed based on the k-mer length, the number of allowed mismatches, and the number of k-mer seed hits at a particular location. This requires dynamically programmed local alignment at a large number of locations, making this approach highly suitable for use with vector CPU instructions (e.g., AVX2, AVX512) and parallelizable across many cores of a single machine and across many machines connected by a network. Those skilled in the art will recognize that modifications and variations of this approach can be applied for the purpose of identifying several off-target genomic regions.

[0305] In some embodiments, probes having sequences homologous to several off-target genomic regions, or comprising more than a threshold number of DNA molecules corresponding to or derived from several off-target genomic regions, are excluded (or filtered) from the detection combination. For example, probes having sequences homologous to several off-target genomic regions, or corresponding to or derived from DNA molecules of off-target genomic regions derived from more than 30, more than 25, more than 20, more than 18, more than 15, more than 12, more than 10, or more than 5 off-target regions, are excluded.

[0306] In some embodiments, depending on the number of off-target regions, several probes are divided into 2, 3, 4, 5, 6, or more separate groups. For example, several probes that are not sequence homologous to DNA molecules corresponding to or derived from several off-target regions are assigned to a high-quality group; several probes that are sequence homologous to DNA molecules corresponding to or derived from 1 to 18 off-target regions are assigned to a low-quality group; and several probes that are sequence homologous to DNA molecules corresponding to or derived from more than 19 off-target regions are assigned to a poor-quality group. Other cutoff values ​​may be used for grouping.

[0307] In some embodiments, several probes in the lowest quality group are excluded. In some embodiments, several probes in several groups different from the highest quality group are excluded. In some embodiments, separate detection combinations are created for the probes in each group. In some embodiments, all probes are placed on the same detection combination, but separate analyses are performed based on the assigned group.

[0308] In some embodiments, a detection combination includes a larger number of high-quality probes than the number of probes in a lower group. In some embodiments, a detection combination includes a smaller number of poor-quality probes than the number of probes in other groups. In some embodiments, more than 95%, 90%, 85%, 80%, 75%, or 70% of the probes in a detection combination are high-quality probes. In some embodiments, less than 35%, 30%, 20%, 10%, 5%, 4%, 3%, 2%, or 1% of the probes in a detection combination are low-quality probes. In some embodiments, less than 5%, 4%, 3%, 2%, or 1% of the probes in a detection combination are poor-quality probes. In some embodiments, no poor-quality probes are included in a detection combination.

[0309] In some embodiments, probes having less than 50%, less than 40%, less than 30%, less than 20%, less than 10%, or less than 5% are removed. In some embodiments, probes having more than 30%, more than 40%, more than 50%, more than 60%, more than 70%, more than 80%, or more than 90% are selectively included in a detection combination.

[0310] Using a combination of cancer testing methods:

[0311] In another aspect, a method for using a cancer detection kit is provided. The method may include the steps of: (e.g., using bisulfite treatment) treating several DNA molecules or fragments to convert unmethylated cytosine to uracil; applying a cancer detection kit (as described herein) to the converted DNA molecules or fragments; enriching a sub-combination of the converted DNA molecules or fragments hybridized (or bound) to the probes in the detection kit; and, for example, detecting the nucleic acid sequence and determining the methylation state of the nucleic acid sequence by sequencing the enriched cfDNA fragment. In some embodiments, the several sequence readings may be compared with a reference genome (e.g., a human reference genome) to allow identification of the methylation state at several CpG sites in the DNA molecules or fragments, and thus providing information about cancer detection. While this disclosure focuses particularly on sequencing-based methods for detecting nucleic acids and determining their methylation status (via several sequence reads), it is broad enough to encompass other methods for detecting nucleic acids and determining their methylation status, such as (as described in WO2014 / 043763, which is incorporated herein by reference) other methylation detection sequencing methods, DNA microarrays (e.g., with several labeled probes adhered to or bonded to a solid surface or DNA array wafer), etc.

[0312] Analysis of sequence readings:

[0313] In some embodiments, the plurality of sequence reads may be aligned to a reference genome using methods known in the art to determine alignment position information. The alignment position information may indicate a start position and an end position in the reference genome corresponding to a start nucleotide base and an end nucleotide base of a given sequence read. The alignment position information may also include the sequence read length, which may be determined from the start and end positions. A region in the reference genome may be associated with a gene or a segment of a gene.

[0314] In various embodiments, a sequence read comprises a read pair denoted as R1 and R2. For example, the first read R1 may be sequenced from a first end of a nucleic acid fragment, and the second read R2 may be sequenced from a second end of the nucleic acid fragment. Therefore, several nucleotide base pairs of the first read R1 and the second read R2 may be aligned consistently (e.g., in opposite directions) with the nucleotide bases of the reference genome. Alignment information derived from the read pair R1 and R2 may include a start position in the reference genome corresponding to an end of a first read (e.g., R1) and an end position in the reference genome corresponding to an end of a second read (e.g., R2). In other words, the start position and the end position in the reference genome represent possible positions in the reference genome corresponding to the nucleic acid fragment. An output file in SAM (Sequence Alignment Map) or BAM (Binary Alignment Map) format may be generated and output for further analysis.

[0315] From the plurality of sequence reads, the location and methylation state of each CpG site can be determined based on alignment to a reference genome. Further, a methylation state vector for each fragment can be generated specifying a location of the fragment in a reference genome (e.g., specified by the location of the first CpG site in each fragment or other similar metrics), the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment—either methylated (e.g., denoted M), unmethylated (e.g., denoted U), or intermediate (e.g., denoted I). The plurality of methylation state vectors can be stored in temporary or permanent computer memory for later use and processing. Further, multiple copies of reads or copy methylation state vectors from a single object can be removed. In an additional embodiment, a particular fragment can be determined to have one or more CpG sites with an intermediate methylation state. Such fragments can be excluded from subsequent processing or selectively included when such intermediate methylation states are incorporated into downstream data models.

[0316] Figure 7BAccording to one embodiment, sequencing a cfDNA fragment to obtain a methylation state vector. Figure 7A An illustration of procedure 100. As an example, the analysis system takes a cfDNA fragment 112. In this example, the cfDNA fragment 112 includes three CpG sites. As shown, the first and third CpG sites of the cfDNA fragment 112 are methylated 114. In the processing step 120, the cfDNA fragment 112 is converted to produce a converted cfDNA fragment 122. In the processing 120, the cytosine at the unmethylated second CpG site is converted to uracil. However, the first and third CpG sites are not converted.

[0317] After conversion, a sequence library 130 is prepared and sequenced 140, producing a sequence readout 142. The analysis system aligns the sequence readout 142 150 to a reference genome 144. The reference genome 144 provides background information on where the fragment cfDNA originates in a human genome. In this simplified example, the analysis system aligns the sequence readout 150, associating three CpG sites with CpG sites 23, 24, and 25 (arbitrary reference markers are used for ease of description). The analysis system thus produces information on the methylation status of all CpG sites on the cfDNA fragment 112 and where these CpG sites are mapped to in the human genome. As shown, the methylated CpG sites on sequence readout 142 are read as cytosine. In this example, cytosine appears only at the first and third CpG sites in sequence readout 142, allowing the inference that the first and third CpG sites in the original cfDNA fragment are methylated. The second CpG site is read as a thymine (U is converted to T in the sequencing process), and therefore, it can be inferred that the second CpG site is unmethylated in the original cfDNA fragment. Using both methylation state and location information, the analysis system generates a methylation state vector 152 for the cfDNA fragment 112. In this example, the resulting methylation state vector 152 is <M 23 U 24 M 25 > where M corresponds to a methylated CpG site, U corresponds to an unmethylated CpG site, and the subscript number corresponds to the position of each CpG site in the reference genome.

[0318] Figure 8A and 8BThree graphs are presented to demonstrate the consistency of sequencing data from a control group. The first graph, 170, shows the conversion accuracy of unmethylated cytosine to uracil (step 120) on a test sample obtained from several patients across different stages of cancer: stage 0, stage 1, stage 2, stage 3, and stage 4, as well as non-cancer patients. As shown, there is a consistent accuracy in converting unmethylated cytosine to uracil on cfDNA fragments. There is an overall accuracy of 99.47%, with a precision of ±0.024%. The second graph, 180, compares the coverage (sequencing depth) across various stages of cancer. Counting only a few sequence reads confidently labeled to a reference genome, the average coverage across all groups is approximately 34. The third graph, 190, shows the concentration of cfDNA in each sample across various stages of cancer.

[0319] Cancer detection:

[0320] The sequence reads obtained by the methods provided herein can be further processed by automated algorithms. For example, the analysis system is used to receive sequence data from a sequencer and perform various aspects of the processing described herein. The analysis system can be a PC, desktop computer, laptop computer, notbook, tablet PC, or mobile device. A computing device can be communicatively coupled to the sequencer via a wireless, wired, or a combination of wireless and wired communication technologies. Generally, the computing device is configured with a processor and a memory storing several computer instructions. When executed by the processor, the several computer instructions cause the processor to perform several steps as described in the remainder of this document. Generally, the amount of genetic data and data derived from said genetic data is large enough, and the required computing power is so great, that it is impossible to perform this solely on paper or by human thought.

[0321] Clinical interpretation of several methylation states of several target genomic regions is a procedure that includes classifying the clinical effects of each or a combination of the several methylation states and reporting the results in a manner meaningful to a medical professional. The clinical interpretation may be based on comparisons of the several sequence reads with databases specific to cancer or non-cancer subjects, and / or on the number and type of cfDNA fragments with cancer-specific methylation patterns identified in a single sample. In some embodiments, the several target genomic regions are ranked or classified based on their likelihood of being differentially methylated in several cancer samples, and the ranking or classification is used in the interpretation process. The ranking and classification may include (1) the type of clinical effect, (2) the strength of evidence for the effect, and (3) the magnitude of the effect. Various clinical analysis and genomic data interpretation methods can be used for the analysis of the several sequence reads. In some other embodiments, the clinical interpretation of the methylation states of such differentially methylated regions can be based on machine learning, which interprets a current sample using a classification or regression method trained on samples from cancer and non-cancer patients with known cancer status, cancer type, cancer stage, TOO, etc.

[0322] Clinically significant information may include the presence or absence of cancer broadly, the presence or absence of a specific type of cancer, the stage of cancer, or the presence or absence of other types of disease. In some embodiments, the information relates to the presence or absence of one or more cancer types selected from the group consisting of: breast cancer, endometrial cancer, cervical cancer, ovarian cancer, bladder cancer, urethral carcinoma of the renal pelvis, renal cell carcinoma, prostate cancer, anorectal cancer, colorectal cancer, hepatocellular carcinoma, bile duct cancer and hepatocellular carcinoma, pancreatic cancer, upper gastrointestinal adenocarcinoma, esophageal squamous cell carcinoma, head and neck cancer, squamous cell lung cancer, lung adenocarcinoma, small cell lung cancer, neuroendocrine carcinoma, melanoma, thyroid cancer, sarcoma, multiple myeloma, myeloma, lymphoma, and leukemia. In some embodiments, the samples are not cancerous and are derived from subjects with clonal expansion of white blood cells or without cancer.

[0323] Cancer classifier:

[0324] In some embodiments, the laboratory test combination described herein can be used with a cancer type classifier that predicts a disease state for a sample, such as a cancer or non-cancer prediction, a tissue of origin prediction, and / or an intermediate prediction. In some examples, the cancer type classifier can generate several features based on several sequence reads by incorporating several methylated and unmethylated fragments of DNA located in a specific genomic region of interest. For example, if the cancer type classifier determines that a methylation pattern at a fragment resembles a methylation pattern of a specific cancer type, the cancer type classifier can set a feature of that fragment to 1, and if no such fragment exists, the feature can be set to 0. In this way, the cancer type classifier can generate a set of binary features for each sample (30,000 features for example only). Further, in some examples, all or part of the set of binary features for a sample can be input into the cancer type classifier to provide a set of probability scores, such as one probability score for each cancer type category and one non-cancer type category. Furthermore, in some examples, the cancer type classifier may integrate or be used in conjunction with a threshold to determine whether a sample should be called as cancerous or non-cancer, and / or as an intermediate threshold to reflect the confidence level for a particular TOO call. Such methods are further described below.

[0325] To train the cancer type classifier, the analysis system (e.g., analysis system 800) may obtain a set of training samples. In some examples, each training sample includes fragment files (e.g., several files containing sequence read data), a label corresponding to a type of cancer (TOO) or non-cancer state of the sample, and / or the sex of the individual in the sample. The analysis system may apply the training set to train the cancer type classifier to predict the disease state of the sample.

[0326] In some embodiments, for training purposes, the analysis system divides the genome (e.g., the whole genome) or a single combination of the genome (e.g., several targeted methylation regions) into several regions. By way of example only, several portions of the genome may be divided into several "blocks" of CpGs, with a new block starting when the distance between the nearest adjacent CpGs is at least a minimum separation distance (e.g., at least 500 bp). Further, in some examples, each block may be divided into several 1000 bp regions and positioned such that several adjacent regions have a certain amount of overlap (e.g., 50% or 500 bp).

[0327] Furthermore, in some examples, the analysis system can divide the training set into K sub-combinations or folds, which will be used in a K-fold cross-validation. In some examples, the folds can be balanced based on cancer / non-cancer status, tissue of origin, cancer stage, age (e.g., grouped in buckets of 10 years), and / or smoking status. In some examples, the training set is divided into 5 folds, thereby training five separate classifiers, in each case trained on 4 / 5 of the training samples and using the remaining 1 / 5 for validation.

[0328] During training with the training set, the analysis system can fit a probabilistic model to several fragments derived from samples of each cancer type (and for healthy cfDNA). As used herein, a “probabilistic model” is any mathematical model that assigns a probability to a sequence read based on the methylation status at one or more sites on the read. During training, the analysis system fits several sequence reads derived from one or more samples from several subjects with a known disease and can be used to determine the likelihood of several sequence reads indicating a disease state by applying methylation information or several methylation status vectors. In particular, in some cases, the analysis system determines the observed methylation ratio for each CpG site in a sequence read. The methylation ratio represents the proportion or percentage of base pairs methylated at a CpG site. The trained probabilistic model can be parameterized by the product of the several methylation ratios. Generally, any known probabilistic model for assigning several probabilistics to several sequence reads from a sample can be used. For example, the probability model can be a binary model in which a methylation probability is assigned to each site (e.g., a CpG site) on a nucleic acid fragment, or it can be an independent site model in which methylation of each CpG is assigned by a different methylation probability, and methylation at a site is assumed to be independent of methylation at one or more other sites on the nucleic acid fragment.

[0329] In some embodiments, the probability model is a Markov model in which the probability of methylation at each CpG site depends on the methylation state at some number of preceding CpG sites in the sequence read, or in nucleic acid molecules from which the sequence read is derived. See, for example, U.S. Patent Application No. 16 / 352,602, entitled “Anomalous Fragment Detection and Classification,” filed May 13, 2019, which is incorporated herein by reference in its entirety and can be used in various embodiments.

[0330] In some examples, the probability model is a "mixture model" fitted using a mixture of components from several lower-level models. For example, in some embodiments, the several mixture components can be determined using multiple independent site models, where methylation at each CpG site (e.g., the methylation rate) is assumed to be independent of methylation at other CpG sites. Applying an independent site model, a probability assigned to a sequence read, or to the nucleic acid molecule from which the sequence read is derived, is the product of the methylation probability of each CpG site where the sequence read is methylated and a subtracted methylation probability of each CpG site where the sequence read is not methylated. According to this example, the analytical system determines the methylation rate of each of the several mixture components. The mixture model is parameterized by a sum of the several mixture components, each of which is associated with a product of the several methylation rates. A probability model Pr for n mixture components can be represented by the following:

[0331]

[0332] For an input segment, m i ∈{0,1} represents the observed methylation state of the fragment at position i in a reference genome, where 0 indicates unmethylated and 1 indicates methylated. The fractional assignment value for each mixture component k is f. k , where f k ≥0 months The methylation probability at position i in the CpG site of the mixed component k is β. ki Therefore, the probability of unmethylation is 1-β. ki The number n of the mixed components can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, etc.

[0333] In some examples, the analysis system uses a maximum probability estimate to fit the probability model in order to identify a set of parameters {β}. ki f k}, the set of parameters {β ki f k Maximize the logarithmic probability of all fragments derived from a disease state, subject to a regularization penalty imposed on each methylation probability with a regularization strength r. The maximized magnitude of the N total fragments can be expressed as:

[0334]

[0335] In some examples, the analysis system performs fitting separately for each cancer type and for healthy cfDNA. As those skilled in the art will understand, other methods can be used to fit the plurality of probabilistic models or to identify a plurality of parameters that maximize the logarithmic probability of all sequence readings derived from the plurality of reference samples. For example, in some examples, Bayesian fitting (using, for example, Markov chain Monte Carlo) is used, where each parameter is not assigned a single value but is associated with a distribution. In some examples, gradient-based optimization is used, where gradients of the probabilities (or log probabilities) of the plurality of parameter values ​​are used to progressively approach an optimum through the parameter space. In other examples, predictability maximization is used, where a set of latent parameters (e.g., the identity of the mixture components derived from each fragment) are set to their predictability under the previous plurality of model parameters, and then the model parameters are specified to maximize the probability under the assumptions of these latent variables. The two-step procedure is then repeated until convergence.

[0336] Furthermore, in some examples, the analysis system can generate several features for each sample in the training group. For example, for each sample (regardless of label), for each region, for each cancer type, for each fragment, the analysis system can evaluate the log-probability ratio R according to several fitted probability models based on the following formula:

[0337]

[0338] Next, for each sample, for each region, for each cancer type, and for each group of "tier" values, the analysis system can count those with R 癌症类型 The number of segments in the layers, and assigning those counts as non-negative integer values ​​as features. For example, the layers include thresholds of 1, 2, 3, 4, 5, 6, 7, 8, and 9, resulting in 9 features for each region for each cancer type.

[0339] In some examples, the analysis system may select specific features to include in a feature vector for each sample. For example, for each pair of different cancer types, the analysis system may designate one type as "positive" and the other as "negative," and rank the features based on their ability to distinguish these types. In some cases, the ranking is based on mutual information calculated by the analysis system. For example, the mutual information may be calculated using estimated proportions of several samples of the positive and negative types (e.g., cancer types A and B), for which the feature is expected to be non-zero in an outcome test. For example, if a feature occurs frequently in healthy cfDNA, the analysis system determines that the feature is unlikely to occur frequently in cfDNA associated with various types of cancer. Therefore, the feature may be a weak criterion when distinguishing between several disease states. When calculating mutual information I, variable X is a specific feature (e.g., a binary feature) and variable Y represents a disease state, such as cancer type A or B:

[0340]

[0341] p(1|A)=f A +f H -f H f A

[0342] The joint probability mass function for X and Y is p(x, y), and the marginal probability mass functions are p(x) and p(y). The analytical system can a priori assume that feature loss is uninformative and that each disease state is equally likely; for example, p(Y=A) = p(Y=B) = 0.5. (For example, in cfDNA) the probability of observing a given binary feature of cancer type A is represented by p(1|A), while f... A It is the probability of observing the aforementioned feature in ctDNA samples (or high-signal cfDNA samples) from tumors associated with cancer type A, and f H It is the probability of observing the aforementioned feature in a healthy or non-cancer cfDNA sample.

[0343] In some examples, only features corresponding to the positive type are included in the ranking, and only if the predicted incidence rate of these features is higher in the positive type than in the negative type. For example, if "liver" is the positive type and "breast" is the negative type, only the "liver_x" feature is considered, and only if their predicted incidence in liver cfDNA is greater than their predicted incidence in breast cfDNA. Further, in some examples, for each region, for each cancer type pair (including non-cancer types that are negative), the analysis system retains only the best-performing layer. Further, in some examples, the analysis system transforms several feature values ​​through binary transformation, such that any feature value greater than 0 is set to 1, while all features are either 0 or 1.

[0344] In some examples, the analysis system trains a multinomial logistic regression classifier on a fold of training data and generates predictions for the excluded data. For example, for each of the K folds, a logistic regression can be trained for each combination of several hyperparameters. Such hyperparameters may include L2 penalty and / or topK (e.g., the number of high-ranking regions retained for each tissue type pair (including non-cancer), ranked as outlined in the mutual information procedure above). For each pair of hyperparameters, performance is evaluated on cross-validation predictions on the full training set, and the hyperparameter set with the best performance is selected for retraining on the full training set. In some examples, the analysis system uses log loss as a performance metric, which is calculated by taking the negative logarithm of the prediction for the correct label for each sample and then summing it across several samples (i.e., a perfect prediction of 1.0 for the correct label would give a log loss of 0).

[0345] To generate a prediction for a new sample, several feature values ​​are computed using the same method described above, but narrowed down to a few features (region / positive class combinations) selected at a chosen topK value. The generated features are then used to create a prediction using the logistic regression model trained above.

[0346] In some examples, the analysis trains a two-stage classifier. For instance, the analysis system trains a binary cancer classifier based on the feature vectors of the training samples to distinguish between several labels, cancer, and non-cancer. In this case, the binary classifier outputs a prediction score indicating the probability of cancer's presence or absence. In another example, the analysis system trains a multi-class cancer classifier to distinguish between many cancer types. In this multi-class cancer classifier, the cancer classifier is trained to determine a cancer prediction, which includes a predicted value for each of the several cancer types it is classified as. The several predicted values ​​may correspond to a probability that a given sample has each of the several cancer types. For example, the cancer classifier returns a cancer prediction that includes a prediction of breast cancer, lung cancer, and no cancer. For example, the cancer classifier may return a cancer prediction for a test sample that includes a prediction score for breast cancer, lung cancer, and / or no cancer.

[0347] The analysis system can train the cancer classifier according to any of several methods. As an example, the binary cancer classifier can be an L2-normalized logistic regression classifier trained using a logarithmic loss function. As another example, the multiple cancer (TOO) classifier can be a multinomial logistic regression. In application, both types of cancer classifiers can be trained using other techniques. Numerous techniques exist, including potential applications of kernel methods, machine learning algorithms such as multilayer neural networks, etc. In particular, methods described in PCT / US2019 / 022122 and U.S. Patent Application No. 16 / 352,602, which are incorporated herein by reference in their entirety, can be used in various embodiments. Furthermore, in some examples, the TOO classifier is trained only on several cancer samples that the binary classifier successfully identifies as cancer, thereby ensuring sufficient cancer signal in the cancer samples. On the other hand, in some examples, the binary classifier is trained on several training samples regardless of TOO.

[0348] Exemplary sequencers and analysis systems:

[0349] Figure 10A This is a flowchart of several systems and apparatuses for sequencing several nucleic acid samples according to one embodiment. This illustrative flowchart includes several apparatuses, such as a sequencer 820 and an analysis system 800. The sequencer 820 and the analysis system 800 can work together to perform one or more steps of the procedure described herein.

[0350] In various embodiments, the sequencer 820 receives an enriched nucleic acid sample 810. (As shown in...) Figure 10A The sequencer 820 may include a graphical user interface 825 and one or more loading stations 830. The graphical user interface 825 allows user interaction for specific tasks (e.g., initiating or terminating sequencing). The one or more loading stations 830 are used to load a sequencing cartridge, which includes several enriched fragment samples and / or buffer solutions necessary for performing the sequencing assay. Therefore, once a user of the sequencer 820 provides the necessary reaction reagents and sequencing cartridge to the loading station 830 of the sequencer 820, the user can initiate sequencing by interacting with the graphical user interface 825 of the sequencer 820. Once initiated, the sequencer 820 performs sequencing and outputs sequence readings of the several enriched fragments from the nucleic acid sample 810.

[0351] In some embodiments, the sequencer 820 is communicatively coupled to the analysis system 800. The analysis system 800 includes a number of computing devices for processing the plurality of sequence readings for various applications, such as evaluating methylation status at one or more CpG sites, variable call, or quality control. The sequencer 820 can provide the plurality of sequence readings in BAM file format to the analysis system 800. The analysis system 800 can be coupled to the sequencer 820 via a wireless, wired, or a combination of both communication technologies. Generally, the analysis system 800 is configured with a processor and a non-transitory computer-readable storage medium storing a plurality of computer instructions. When executed by the processor, the plurality of computer instructions cause the processor to process the plurality of sequence readings or perform one or more steps of any of the methods or procedures disclosed herein.

[0352] In some embodiments, the plurality of sequence reads may be aligned to a reference genome using methods known in the art to determine alignment position information. The alignment position may broadly describe a start and end position of a region in the reference genome corresponding to a start and end nucleotide base of a given sequence read. Corresponding to methylation sequencing, the alignment position information may be summarized to indicate a first CpG site and a last CpG site included in the sequence read, based on alignment to the reference genome. The alignment position information may further indicate the methylation status and location of all CpG sites in a given sequence read. A region in the reference genome may be associated with a gene or a segment of a gene. Thus, the analysis system 800 may align the sequence read to one or more gene-labeled sequences. In one embodiment, fragment length (or size) is determined from the start and end positions.

[0353] In various embodiments, for example, when a pairing end sequencing procedure is used, a sequence read includes a read pair denoted as R_1 and R_2. For example, the first read R_1 may be sequenced from a first end of a double-stranded DNA (dsDNA) molecule, and the second read R_2 may be sequenced from a second end of the double-stranded DNA (dsDNA). Therefore, several nucleotide base pairs of the first read R_1 and the second read R_2 may be consistently (e.g., in opposite orientations) aligned with the nucleotide bases of the reference genome. Alignment information derived from the read pair R_1 and R_2 may include a start position in the reference genome corresponding to one end of a first read (e.g., R_1) and an end position in the reference genome corresponding to one end of a second read (e.g., R_2). In other words, the start and end positions in the reference genome represent possible positions in the reference genome corresponding to the nucleotide fragment. In one embodiment, the readings R_1 and R_2 may be combined into a segment, and the segment may be used for subsequent analysis and / or classification. An output file in SAM (Order Aligned Map) or BAM (Binary Map) format may be generated and output for further analysis.

[0354] Now for reference Figure 10B , Figure 10BThis is a block diagram of an analysis system 800 for processing DNA samples according to one embodiment. The analysis system utilizes one or more computing devices for analyzing multiple DNA samples. The analysis system 800 includes a sequence processor 840, a sequence database 845, a model database 855, several models 850, a parameter database 865, and a scoring engine 860. In some embodiments, the analysis system 800 performs... Figure 3A Program 300 Figure 3B Program 340 Figure 4 Program 400 Figure 5 Program 500 Figure 6A Program 600 or Figure 6B The procedure 680 and one or more steps of other procedures described herein.

[0355] The sequence processor 840 generates several methylation state vectors for several fragments from a sample. At each CpG site on a fragment, the sequence processor 840 via... Figure 3A The program 300 generates a methylation state vector for each fragment, the methylation state vector specifying the location of the fragment in the reference genome, the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment as methylated, unmethylated, or intermediate. The sequence processor 840 can store the methylation state vectors of several fragments in the sequence database 845. The data in the sequence database 845 can be organized such that the several methylation state vectors from a single sample are correlated with each other.

[0356] Furthermore, multiple different models 850 can be stored in the model database 855 or recycled for use on several detection samples. In one example, a model is a trained cancer classifier used to determine a cancer prediction for a detection sample using a feature vector derived from several anomalous fragments. The training and use of the cancer classifier are discussed elsewhere herein. The analysis system 800 can train one or more models 850 and store various training parameters in the parameter database 865, and the analysis system 800 stores the multiple models 850 along with several functions in the model database 855.

[0357] During inference, the scoring engine 860 uses one or more models 850 to return an output. The scoring engine 860 accesses the plurality of models 850 in the model database 855 along with a plurality of trained parameters from the parameter database 865. For each model, the parameter engine receives an input appropriate for that model and computes an output based on the received input, the plurality of parameters and a function of each model relating the input and the output. In some use cases, the scoring engine 860 further computes a plurality of metrics associated with a confidence level of the computed output from the model. In other use cases, the scoring engine 860 computes additional intermediate values ​​for the model.

[0358] application:

[0359] In some embodiments, the methods, analytical systems, and / or classifiers of the present invention can be used to detect the presence (or absence) of cancer, monitor cancer progression or recurrence, monitor treatment response or effectiveness, determine the presence of or monitor a minimum residual disease (MRD), or for any combination thereof. In some embodiments, the analytical systems and / or classifiers can be used to identify the tissue or origin of a cancer. For example, the aforementioned systems and / or classifiers can be used to identify a cancer as any of the following cancer types: breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, renal pelvis urethral epithelial carcinoma, renal cancer other than urethral epithelial carcinoma, prostate cancer, anorectal cancer, anal cancer, colorectal cancer, hepatobiliary cancer originating from hepatocytes, hepatobiliary cancer originating from cells other than hepatocytes, hepatic / cholecystitis, esophageal cancer, pancreatic cancer, upper gastrointestinal squamous cell carcinoma, upper gastrointestinal cancer other than squamous cell carcinoma, head and neck cancer, lung cancer, lung adenocarcinoma, small cell lung cancer, squamous cell lung cancer and cancers other than lung adenocarcinoma or small cell lung cancer, neuroendocrine carcinoma, melanoma, thyroid cancer, sarcoma, plasmacytoma, multiple myeloma, myeloma, lymphoma, and leukemia. For example, as described herein, a classifier can be used to generate a sample feature vector that is a probability or probability score (e.g., from 0 to 100) from an object having cancer. In some embodiments, the probability score is compared to a threshold probability to determine whether the subject has cancer. In other embodiments, the probability score may be assessed at different time points (e.g., before or after treatment) to monitor disease progression or treatment effectiveness (e.g., treatment outcome). In still other embodiments, the probability score may be used to make or influence a clinical decision (e.g., cancer detection, treatment selection, assessment of treatment effectiveness, etc.). For example, in one embodiment, if the probability score exceeds a threshold, a physician may formulate an appropriate treatment.

[0360] Cancer detection:

[0361] In some embodiments, the methods and / or classifiers of the present invention are used to detect a cancer type in an object suspected of having cancer. For example, a classifier (as described herein) may be used to determine a probability or likelihood score that a sample feature vector comes from an object having a cancer type.

[0362] In one embodiment, a probability score greater than or equal to 60 may indicate that the subject has the type of cancer. In other embodiments, a probability score greater than or equal to 65, 70, 75, 80, 85, 90, or 95 indicates that the subject has the type of cancer. In other embodiments, a probability score may indicate the severity of the disease. For example, a probability score of 80 compared to a score below 80 (e.g., a score of 70) may indicate a more severe form of cancer or a later stage. Similarly, an increase in the probability score over time (e.g., at a second, later point in time) may indicate disease progression, or a decrease in the probability score over time (e.g., at a second, later point in time) may indicate successful treatment.

[0363] In another embodiment, a cancer log odds ratio can be calculated for a test subject by taking the logarithm of the probability of having a cancer type divided by the probability of not having that cancer type (i.e., a minus the probability of having the cancer type), as described herein. According to this embodiment, a cancer log odds ratio greater than 1 can indicate that the subject has a cancer type. In yet another embodiment, a cancer type log odds ratio greater than 1.2, 1.3, 1.4, 1.5, 1.7, 2, 2.5, 3, 3.5, or 4 indicates that the subject has that cancer type. In other embodiments, a cancer log odds ratio can indicate the severity of the disease. For example, a cancer log odds ratio greater than 2 compared to a score less than 2 (e.g., a score of 1) can indicate a more severe form of cancer or a later stage of a cancer type. Similarly, an increase in the cancer log odds ratio over time (e.g., at a second, later time point) can indicate disease progression, or a decrease in the cancer log odds ratio over time (e.g., at a second, later time point) can indicate successful treatment.

[0364] According to several aspects of the present invention, the methods and systems described herein can be trained to detect or classify multiple cancer indicators. For example, the methods, systems, and classifiers of the present invention can be used to detect the presence of one or more, two or more, three or more, five or more, or ten or more different types of cancer.

[0365] In some embodiments, the cancer is one or more of head and neck cancer, liver / cholecystitis, upper gastrointestinal cancer, pancreatic / gallbladder cancer, colorectal cancer, ovarian cancer, lung cancer, multiple myeloma, lymphoma, melanoma, sarcoma, breast cancer, and uterine cancer. In some embodiments, the cancer is one or more of anorectal cancer, bladder or urethral epithelial cancer, or cervical cancer. In some embodiments, the cancer is breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, urethral cancer of the renal pelvis, renal cancer other than urethral epithelium, prostate cancer, anorectal cancer, anal cancer, colorectal cancer, hepatobiliary cancer originating from hepatocytes, hepatobiliary cancer originating from cells other than hepatocytes, liver / cholecystitis, esophageal cancer, pancreatic cancer, upper gastrointestinal squamous cell carcinoma, upper gastrointestinal cancer other than squamous cells, head and neck cancer, lung cancer, lung adenocarcinoma, small cell lung cancer, squamous cell lung cancer and cancer other than lung adenocarcinoma or small cell lung cancer, neuroendocrine carcinoma, melanoma, thyroid cancer, sarcoma, plasmacytoma, multiple myeloma, myeloma, lymphoma and leukemia.

[0366] In some embodiments, the probability or chance score may be assessed at different time points (e.g., before or after treatment) to monitor disease progression or treatment effectiveness (e.g., treatment outcome). For example, this disclosure provides several methods involving obtaining a first sample (e.g., a first plasma cfDNA sample) from a cancer patient at a first time point, determining a first probability or chance score from the first sample (as described herein), obtaining a second sample (e.g., a second plasma cfDNA sample) from the cancer patient at a second time point, and determining a second probability or chance score from the second sample (as described herein).

[0367] treat:

[0368] In other embodiments, information obtained from any of the methods described herein (e.g., the probability or chance score) can be used to make or influence a clinical decision (e.g., cancer detection, treatment selection, assessment of treatment effectiveness, etc.). For example, in one embodiment, if the probability or chance score exceeds a threshold, a physician can formulate an appropriate treatment (e.g., a surgical resection, radiation therapy, chemotherapy, and / or immunotherapy). In some embodiments, information, such as a probability or chance score, can be provided to a physician or subject as a readout.

[0369] A classifier (as described herein) can be used to determine a probability or likelihood score that a sample feature vector comes from an object having cancer or a specific type of cancer (e.g., tissue of origin). In one embodiment, an appropriate treatment (e.g., surgical resection or therapy) is prescribed when the probability or likelihood exceeds a threshold. For example, in one embodiment, one or more appropriate treatments are prescribed if the probability or likelihood score is greater than or equal to 60. In another embodiment, one or more appropriate treatments are prescribed if the probability or likelihood score is greater than or equal to 65, greater than or equal to 70, greater than or equal to 75, greater than or equal to 80, greater than or equal to 85, greater than or equal to 90, or greater than or equal to 95. In other embodiments, a cancer log odds ratio can indicate the effectiveness of a cancer treatment. For example, an increase in the cancer log odds ratio (e.g., after a second treatment) can indicate that the treatment is ineffective. Similarly, a decrease in the cancer log odds ratio (e.g., after a second treatment) can indicate successful treatment. In another embodiment, if the log odds ratio of the cancer is greater than 1, greater than 1.5, greater than 2, greater than 2.5, greater than 3, greater than 3.5, or greater than 4, one or more appropriate treatments are formulated.

[0370] In some embodiments, the treatment is one or more cancer therapeutic agents selected from the group consisting of: a chemotherapy agent, a targeted cancer therapeutic agent, a differentiation therapy agent, a hormone therapy agent, and an immunotherapy agent. For example, the treatment may be selected from the group consisting of alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, cytoskeleton disruptors (taxanes), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based agents, and any combination thereof. In some embodiments, the treatment is one or more targeted cancer therapeutic agents selected from the group consisting of signal transduction inhibitors (e.g., tyrosine kinase and growth factor receptor inhibitors), histone deacetylase (HDAC) inhibitors, retinoic acid receptor agonists, proteosome inhibitors, angiogenesis inhibitors, and monoclonal antibody conjugates. In some embodiments, the treatment is one or more differentiation therapy agents, including retinoids such as tretinoin, alitretinoin, and bexarotene. In some embodiments, the treatment is one or more hormonal therapy agents selected from the group consisting of anti-estrogens, aromatase inhibitors, progestins, estrogens, anti-androgens, and GnRH agonists or analogues. In one embodiment, the treatment is one or more immunotherapy agents selected from the group consisting of monoclonal antibody therapy agents such as rituximab and alemtuzumab, nonspecific immunotherapeutic agents and adjuvants such as BCG, interleukin-2 (IL-2), and interferon-α, immunomodulatory drugs such as thalidomide and lenalidomide. The selection of an appropriate cancer therapy agent based on characteristics such as tumor type, cancer stage, prior exposure to cancer treatment or agents, and other characteristics of the cancer is within the capabilities of a skilled physician or oncologist.

[0371] Example

[0372] The following examples are presented to provide a complete disclosure and description of how a person skilled in the art can make and use this disclosure, and are not intended to limit the scope of what the inventors consider to be described, nor are they intended to represent that all or only the experiments described below were performed. Efforts have been made to ensure the accuracy of the figures used (e.g., dosage, temperature, etc.), but some experimental errors and biases should be taken into account.

[0373] Example 1: Analysis of probe measurements

[0374] To test the required overlap between a cfDNA fragment and a probe to achieve a non-negligible pull-down, overlaps of various lengths were tested using a assay combination designed to include three different types of probes (VID3, VID4, VIE2). These three different types of probes had varying degrees of overlap with several 175bp target DNA fragments specific to each probe. The tested overlap ranged from 0bp to 120bp. Several samples including the 175bp target DNA fragments were applied to the assay combination and washed, followed by the collection of several DNA fragments linked to the probes. The amount of the collected DNA fragments was measured, and this amount was plotted as density against the size of the overlap, as shown in [the image / image / etc.]. Figure 9 Provided by [the source].

[0375] When there is less than 45 bp overlap, there is no significant binding or pull-down of the target DNA fragment. These results suggest that at least 45 bp of fragment-probe overlap is generally required to achieve a non-negligible pull-down, although this number may vary depending on the assay conditions.

[0376] Furthermore, it has been shown that a mismatch rate of more than 10% between the probe and fragment sequences in the overlapping region is sufficient to significantly interfere with binding and thus with pull-down efficiency. Therefore, several sequences that can be aligned to the probe along at least 45 bp with a pairing rate of at least 90% are candidates for off-target pull-down.

[0377] Therefore, we performed an exhaustive search of all genomic regions (i.e., off-target regions) with 45 bp alignments at a pairing rate of 90%+ for each probe. Specifically, we combined a k-mer seeding strategy (which allows one or more mismatches) with local alignment at several seed locations. This guarantees that no good alignments are missed based on the k-mer length, the allowed number of mismatches, and the number of k-mer seed hits at a particular location. This involves performing dynamically programmed local alignment at a large number of locations, making this approach suitable for parallelization using vector CPU instructions (e.g., AVX2, AVX512) and across many cores in a single machine, and across many machines connected by a network. This allows for valuable exhaustive searches when designing a high-performance detection combo (i.e., low off-target rate and high target coverage for a given amount of sequencing).

[0378] After the exhaustive search, each probe is scored based on the number of off-target regions. The best probes have a score of 1, meaning they only hit one region (high Q). Several probes with a low score between 2 and 19 hits (low Q) are accepted, but several probes with a difference score of more than 20 hits (difference Q) are discarded. Other cutoff values ​​may be used for specific samples.

[0379] The number of high-quality, low-quality, and poor-quality probes was then counted among several probes targeting several hypermethylated genomic regions or several hypomethylated genomic regions.

[0380] Example 2: A combination of cancer tests for detecting specific cancer types

[0381] Cancer Types: Several cancer-specific test combinations were designed to detect (15) different cancer types and / or cancer-derived tissues. The 15 cancer types include (1) bladder cancer, (2) breast cancer, (3) cervical cancer, (4) colorectal cancer, (5) head and neck cancer, (6) hepatobiliary cancer, (7) lung cancer, (8) melanoma, (9) ovarian cancer, (10) pancreatic cancer, (11) prostate cancer, (12) kidney cancer, (13) thyroid cancer, (14) upper gastrointestinal cancer, and (15) uterine cancer (see Lists 1 through 15). Cancer-specific classifications were applied to the samples for relevant classification and labeling.

[0382] Samples used for genomic region selection: DNA samples used for this work came from various sources.

[0383] The Cell-Free Circulating Genome Mapping Study (“CCGA”; ClinicalTrial.gov ID NCT02889978) was a long-term, prospective, case-controlled observational study. Deidentified biosamples were collected from approximately 15,000 participants at 142 locations. Several samples were selected to ensure a pre-specific distribution of cancer type and non-cancer characteristics across several locations in each cohort, and several cancer and non-cancer samples were frequency-age-matched by sex.

[0384] The Cancer Genome Atlas (“TCGA”; ClinicalTrial.gov identifier NCT02889978) is a public resource developed through a collaboration between the National Cancer Institute (NCI) and the National Human Genome Research Institute (NHGRI).

[0385] Disseminated tumor cells (DTCs) were obtained from Conversant.

[0386] Non-cancer cells were provided by Yuval Dor and Ben Glaser (Hebrew University) and were derived from human tissue obtained through standard clinical procedures. For example, mammary ductal and basal epithelial cells were obtained from chest reduction surgery; colonic epithelial cells were obtained from tissue near the site of reimplantation after partial resection of the colon following pathological surgery; bone marrow cells were obtained from joint replacement surgery; vascular and arterial endothelial cells were obtained from vascular surgery; and head and neck epithelium were obtained from tonsillectomy.

[0387] WGBS was performed on over 1000 genomic DNA samples collected from healthy individuals and several individuals diagnosed with cancer at various stages and from tissues of origin. These samples included: formaldehyde-fixed, paraffin-embedded (FFPE) tissue blocks; disseminated tumor cells (DTCs) from several different TOOs of cancer; bone marrow mononuclear cells (BMMCs); white blood cells (WBCs); and peripheral blood mononuclear cells (PBMCs). The DTCs underwent negative selection using a Miltenyi negative selection kit to remove WBCs, fibroblasts, and endothelial cells prior to gDNA separation. This negative selection yielded purified tumor cells, allowing for clearer identification of several differentially methylated regions.

[0388] The TCGA data were collected by heterozygous bisulfite-converted DNA fragments from 8809 samples into a methylation-sensitive oligonucleotide array. Several β values ​​from this study represent the relative abundance of methylation at 480,000 individual CpG sites. After excluding CpG sites (360,000) from high-noise genomic regions and CpG sites with cross-hybrid probes (45,000), 75,000 of these CpG sites were analyzed. The TCGA data were analyzed using different algorithms because the TCGA data describes the methylation of individual CpG sites, while the WGBS data reveals methylation patterns of several strings of adjacent CpG sites on several DNA fragments.

[0389] Source Tissue Category: Each sample was classified into one of twenty-five (25) different Source Tissue (TOO) categories: breast cancer, uterine cancer, cervical cancer, ovarian cancer, bladder cancer, renal pelvis urethral epithelial carcinoma, renal cancer other than urethral epithelial carcinoma, prostate cancer, anorectal cancer, colorectal cancer, hepatobiliary cancer originating from hepatocytes, hepatobiliary cancer originating from cells other than hepatocytes, pancreatic cancer, upper gastrointestinal squamous cell carcinoma, upper gastrointestinal cancer other than squamous cells, head and neck cancer, lung adenocarcinoma, small cell lung cancer, squamous cell lung cancer and cancer other than lung adenocarcinoma or small cell lung cancer, neuroendocrine carcinoma, melanoma, thyroid cancer, sarcoma, multiple myeloma, lymphoma, and leukemia. After filtering out liquid cancer, brain cancer, small bowel cancer, vaginal + vulvar cancer, and penile + testicular cancer, these TOO categories covered 97% of the cancer incidence reported by the Surveillance, Epidemiology, and End Results Program (SEER; seer.cancer.gov). Rare cancers such as sarcomas and neuroendocrine carcinomas were aggregated to prevent misclassification. International Classification of Diseases of Cancer (ICD-O-3) site codes, morphology codes, and behavioral codes, as well as World Health Organization (WHO) site names, were used to categorize several individual samples into the aforementioned TOO categories. For example, as shown in Table 1, the 34 TCGA studies were assigned to the aforementioned TOO categories. The TOO classification was iteratively refined based on observed classification performance.

[0390] Table 1: Source Organization (TOO) Classification of Several TCGA Types

[0391]

[0392]

[0393] Region Selection: For region selection, several fragments with anomalous methylation patterns in the cancer sample were selected using one or more methods described herein. The use of these methods allowed for the identification of several low-noise regions as presumed targets. Among these low-noise regions, several fragments that were most informative in distinguishing cancer types were ranked and selected.

[0394] Specifically, in some embodiments, when WGBS data is used, several fragment sequences in the database are filtered using a non-cancer distribution based on p-values, and only fragments with p-values ​​less than 0.001 are retained, as described herein. In some cases, the selected cfDNAs are further filtered, retaining only cfDNAs that are at least 90% methylated or 90% unmethylated. Next, for each CpG site in the selected fragments, the number of cancer or non-cancer samples, including several fragments overlapping with that CpG site, is counted. Specifically, P(cancer|overlapping fragment) is calculated for each CpG, and several genomic sites with high p-values ​​are selected as general cancer targets. By design, the selected fragments have very low noise (i.e., few non-cancer fragment overlaps).

[0395] To identify several cancer type-specific targets, a similar selection process is performed. Several CpG sites are ranked based on their information gain, including (i) the information gain of the number of samples or other samples of a particular TOO, including non-cancer samples and samples of a different TOO; (ii) the information gain of the number of samples or non-cancer samples of a particular TOO; and / or (iii) the information gain of the number of samples of a particular TOO or a different TOO that overlaps with that CpG site. The procedure is applied to each of the 25 TOOs, and the comparison is performed for all pairwise combinations of the 25 TOOs. For example, P(cancer|overlapping fragment of a TOO) is calculated and then compared with P(cancer|overlapping fragment of a different TOO). An outlier fragment in each TOO, which is much more likely to be a target of the TOO than a cancer fragment in a different TOO, is selected as a target of that TOO. Therefore, the selected genomic regions through the pairwise comparisons include several differentially methylated genomic regions to separate a target TOO and a contrast TOO. The number of genomic regions used to distinguish each target TOO (x-axis) from a contrast TOO (y-axis) is... Figure 11 It was provided in China.

[0396] When TCGA data were used, CpG site β values, indicating methylation density, were used to identify several target genomic regions. This is because the array data are not at the CpG site level and are therefore prone to false positives. To avoid false positives, several CpG sites spanning the genome were converted into several 350 bp bins. The β value of each bin was calculated as the average of the CpG β values ​​within that bin. Bins with fewer than 2 CpGs were excluded from the analysis. Subsequently, several bins were selected based on (i) several samples of a specific TOO and other samples, including non-cancer samples and samples of a different TOO, (ii) several samples of a specific TOO and non-cancer samples, and / or (iii) a specific TOO and samples including a different TOO overlapping with that CpG site with a β difference greater than 0.95.

[0397] The selected genomic regions, as described above, were then filtered based on the number of their off-target genomic regions, as detailed herein. Specifically, the number of genomic locations with at least 90% equal alignment of 45 bp or more was counted as the number of off-target genomic regions. Genomic regions with more than 20 off-target genomic regions were discarded.

[0398] The various lists of target genomic regions selected as described in this section are identified in Table 2 (see lists 1 through 15).

[0399] Table 2: Summary of Lists 1 to 15

[0400] For each list, the table identifies the detected cancer type, the total number of target genomic regions in the list, a series of SEQ ID NOs corresponding to all target genomic regions in the list to be found in the sequence listing submitted with this application, and the detection combination size (the sum of the lengths of all target genomic regions in the list). The sequence listing identifies the chromosomal location of each target genomic region, whether the cfDNA fragment enriched in the region is hypermethylated or hypomethylated, and the sequence of a DNA strand of the target genomic region. Chromosome numbers and start and stop positions are provided relative to the known human reference genome hg19. The sequence of the human reference genome hg19 can be obtained from the Genome Reference Consortium under the reference number GRCh37 / hg19, and is also available from the Genome Browser provided by the Santa Cruz Genomics Institute.

[0401]

[0402] Example 3: A combination of cancer tests for diagnosing specific cancer types

[0403] Additional cancer testing combinations were designed to identify specific cancer types in a manner similar to that proposed in Example 2. Various lists of target genomic regions selected as described in this paragraph are identified in Table 3 (see Lists 16 to 49). The target genomic regions in Lists 16 to 32 each contain sub-combinations of several methylation sites from the target genomic regions in Lists 33 to 49.

[0404] Table 3: Summary of Lists 16 to 49

[0405] For each list, the table identifies the detected cancer type, the total number of target genomic regions in the list, a series of SEQ ID NOs corresponding to all target genomic regions in the list to be found in the sequence listing submitted with this application, and the detection combination size (the sum of the lengths of all target genomic regions in the list). The sequence listing identifies the chromosomal location of each target genomic region, whether the cfDNA fragment enriched in the region is hypermethylated or hypomethylated, and the sequence of a DNA strand of the target genomic region. Chromosome numbers and start and stop positions are provided relative to the known human reference genome hg19. The sequence of the human reference genome hg19 can be obtained from the Genome Reference Consortium under the reference number GRCh37 / hg19, and is also available from the Genome Browser provided by the Santa Cruz Genomics Institute.

[0406]

[0407] Example 4: Generation of a hybrid model classifier

[0408] To maximize performance, the predictive cancer model described in this example was trained using sequence data obtained from: several samples of known cancer types and non-cancer samples from CCGA sub-studies (CCGA1 and CCGA22); several tissue samples of several known cancers obtained from CCGA1; and several non-cancer samples from the STRIVE study (see ClinicalTrial.gov ID: NCT03085888 ( / / clinicaltrials.gov / ct2 / show / NCT03085888)). The STRIVE study is a prospective, multicenter observational cohort study to validate a test for early detection of breast cancer and other aggressive cancers. Additional non-cancer training samples were obtained from the study to train the classifier described herein. The known cancer types included in the CCGA sample group include the following: breast cancer, lung cancer, prostate cancer, colorectal cancer, kidney cancer, uterine cancer, pancreatic cancer, esophageal cancer, lymphoma, head and neck cancer, ovarian cancer, hepatobiliary cancer, melanoma, cervical cancer, multiple myeloma, leukemia, thyroid cancer, bladder cancer, stomach cancer, and anorectal cancer. Thus, a model can be a multi-cancer model (or a multi-cancer classifier) ​​for detecting one or more, two or more, three or more, four or more, five or more, ten or more, or 20 or more different types of cancer.

[0409] The classifier performance data presented below are reported for a locked classifier trained on cancer and non-cancer samples obtained from CCGA2, a CCGA sub-study, and on non-cancer samples from STRIVE. Several individuals in the CCGA2 sub-study differed from several individuals in the CCGA1 sub-study, in which several target genomes were selected. From the CCGA2 study, several blood samples were collected from several individuals diagnosed with untreated cancer (including 20 tumor types and all cancer stages) and several healthy individuals without a cancer diagnosis (control group). For STRIVE, several blood samples were collected from several women within 28 days of their screening mammograms. Cell-free DNA (cfDNA) was extracted from each sample and treated with bisulfite to convert unmethylated cytosine to uracil. The bisulfite-treated cfDNA was enriched using several heterozygous probes designed to enrich bisulfite-converted nucleic acids derived from several target genomic regions in an assay suite comprising all genomic regions listed in 1 to 16. The enriched bisulfite-converted nucleic acid molecules were sequenced using paired-end sequencing on an Illumina platform (San Diego, California) to obtain a set of sequence reads for each of the several training samples. The resulting read pairs were aligned to the reference genome, combined into several fragments, and methylated and unmethylated CpG sites were identified.

[0410] Characterization based on hybrid models:

[0411] For each cancer type (including non-cancer), a probability mixture model is trained and applied to assign a probability to each fragment from each cancer and non-cancer sample based on how likely a fragment is to be observed in a given sample type.

[0412] Fragment-level analysis:

[0413] In short, for each sample type (cancer and non-cancer samples), for each region (where each region is used as is if it is less than 1 kb, otherwise it is subdivided into several 1 kb-long regions with 50% overlap between adjacent regions (e.g., 500 base overlap), for each type of cancer and non-cancer, a probabilistic model is fitted to the several segments derived from the several training samples. The probabilistic model trained for each sample type is a mixture model, where each of the three mixture components is an independent site model, in which methylation at each CpG is assumed to be independent of methylation at other CpGs. Several fragments are excluded from the model if: the fragments have a p-value greater than 0.01 (from a non-cancer Markov model), are identified as duplicate fragments, (only for the target methylated sample) the fragments have a bag size greater than 1, do not cover at least one CpG site, or the fragment length is greater than 1000 bases. If the retained training fragments overlap with at least one CpG from a region, the training fragments are assigned to that region. If a fragment overlaps with several CpGs in multiple regions, the fragment is assigned to all of the multiple regions.

[0414] Local source model:

[0415] Each probability model is fitted using a maximum probability estimate to identify a set of parameters that maximizes the log probability of all segments derived from each sample type, subject to a normalization penalty.

[0416] Specifically, within each classification region, a set of probabilistic models are trained, each model used for one training label (i.e., one for each cancer type and one for each non-cancer type). Each model takes the form of a Bernoulli mixture model with three components. Mathematically:

[0417] (1)

[0418] Where n is the quantity of the mixture components, set to 3 and m i ∈{0,1} is the observed methylation at position i of the fragment, f k It is the fractional value assigned to component k (f) k ≥0 and ∑f k =1) and β ki This represents the methylation percentage of component k at CpGi. The product over i includes only a few positions, for which the monomethylation state can be identified from the sequence. The parameters {f} for each model... k ,β kiThe maximum probabilities of a given number of fragments bearing a training label are maximized using the rprop algorithm (e.g., the rprop algorithm described in Riedmiller M, Braun H, RPROP: A Fast Adaptive Learning Algorithm, Proceedings of the International Symposium on Computer and Information Sciences VII, 1992), subject to a β-distribution prior. ki The total logarithmic probability of a normalized penalty is estimated. Mathematically, the maximum value of this maximum is:

[0419] (2)∑ j ln(Pr(fragment) j |{β ki f k}))+∑ k,i r ln(β ki (1-β ki ))

[0420] Where r is the normalization strength, which is set to 1.

[0421] Characterization:

[0422] Once the several probability models are trained, a set of numerical features is computed for each sample. Specifically, in each region, for each cancer type and non-cancer sample, several features are extracted for each fragment from each training sample. The extracted features are records of several outlier fragments (i.e., fragments that are aberrantly methylated), defined as fragments whose log probability under a first cancer model exceeds their log probability under a second cancer model or a non-cancer model by at least one tier value. Several outlier fragments are recorded for each genomic region, sample model (i.e., cancer type), and tier (for tiers 1, 2, 3, 4, 5, 6, 7, 8, and 9), yielding nine features per region for each sample type. In this way, each feature is defined by three properties: a genomic region, a “positive” cancer type label (excluding non-cancer), and a tier value selected from the group {1, 2, 3, 4, 5, 6, 7, 8, 9}. The numerical value of each feature is defined as the number of segments in that region, and thus:

[0423] (3)

[0424] The plurality of probabilities are defined by equation (1) using the values ​​of the plurality of maximum probability estimates corresponding to the “positive” cancer type (in the numerator of the logarithm) or corresponding to non-cancer (in the denominator).

[0425] Feature ranking:

[0426] For each pair of features, the features are ranked using mutual information based on their ability to distinguish between the first cancer type (which defines the log-probability model from which the features are derived) and the second cancer type or non-cancer. Specifically, two ranked lists of features are compiled for each unique pair of class labels: one list has a first label designated as “positive” and a second label designated as “negative”, and another list has alternating positive / negative designations (except for the “non-cancer” label, which is only permitted as a negative label). For each of these ranked lists, only features whose positive cancer type label (as in Equation (3)) matches the considered positive label are included in the ranking. For each such feature, the proportion of training samples with non-zero feature values ​​is calculated separately for the positive and negative labels. Features with a larger proportion in the positive labels are ranked based on their mutual information relative to the pair of class labels.

[0427] The top 256 features from each pairwise comparison were identified and added to the final feature set for each cancer type and non-cancer sample. To avoid redundancy, if more than one feature was selected from the same positive type and genomic region (i.e., for multiple negative types), only the feature with the lowest (most informative) ranking assigned to its cancer type pair was retained, breaking down several layers by selecting higher-level values. The features in the final feature set for each sample (cancer type and non-cancer sample) were binarized (any feature value greater than 0 was set to 1, making all features either 0 or 1).

[0428] Classifier training:

[0429] The training samples were then divided into different 5-fold cross-validation training groups, and a two-stage classifier was trained on each fold, with 4 / 5 of the training samples trained in each case and the remaining 1 / 5 used for validation.

[0430] In the first phase of training, a binary (two-class) logistic regression model is trained to detect the presence of cancer among several cancer samples (regardless of whether they are cancer-positive or not) from among non-cancer samples. While training this binary classifier, a sample weight is assigned to male non-cancer samples to counteract the sex imbalance in the training group. For each sample, the binary classifier outputs a prediction score indicating the probability of the presence or absence of cancer.

[0431] In the second phase of training, a parallel multi-class logistic regression model for determining the tissue of origin of cancer is trained with TOO as the target label. Only cancer samples that received a score higher than the 95th percentile of the non-cancer samples in the first-phase classifier are included in the training of this multi-class classifier. For each cancer sample used in training the multi-class classifier, the multi-class classification outputs several predictions for the classified cancer type, where each prediction is a probability that a given sample has a specific cancer type. For example, the cancer classifier may return a cancer prediction for a tested sample, including a predicted score for breast cancer, a predicted score for lung cancer, and / or a predicted score for no cancer.

[0432] Both the binary and multi-class classifiers are trained using mini-batch stochastic gradient descent, and in each case, training is prematurely stopped when performance on the validation fold (evaluated by cross-entropy loss) begins to degrade. For predictions on samples outside the training set, in each phase, the scores assigned by the five cross-validation classifiers are averaged. Scores assigned to gender-inappropriate cancer types are set to zero, and the remaining values ​​are renormalized to sum to one.

[0433] Several scores from the validation cross-validations assigned to the training group are retained for use when specifying cutoff values ​​(thresholds) for a particular performance metric. Specifically, the probability scores from the non-cancer samples assigned to the training group are used to define several thresholds corresponding to a specific level of specificity. For example, for a desired specificity target of 99.4%, the threshold is set at the 99.4 percentile of the cancer detection probability scores from the cross-validations of the non-cancer samples assigned to the training group. Several training samples with probability scores exceeding a threshold are called positive for cancer.

[0434] Subsequently, for each training sample determined to be positive for cancer, a TOO or cancer type assessment is made by the multi-class classifier. First, the multi-class logistic regression classifier assigns a set of probability scores to each sample, one probability score for each expected cancer type. Next, the confidence of these scores is evaluated as the difference between the highest and second-highest scores assigned to each sample by the multi-class classifier. Then, cross-validated training set scores are used to identify a minimum threshold such that 90% of cancer samples in the training set where the difference between the top two scores exceeds the threshold are assigned the correct TOO label as their highest score. In this way, the scores assigned to the validation folds during training are further used to determine a second threshold, which is used to distinguish between confident and uncertain TOO calls.

[0435] During prediction, samples receiving a score below a predetermined threshold from the binary (first-stage) classifier are assigned a "non-cancer" label. For the remaining samples, those from the second-stage classifier whose difference between the first two TOO scores is below a second predetermined threshold are assigned a "uncertain cancer" label. The remaining samples are assigned the cancer label based on the highest score assigned by the TOO classifier.

[0436] Example 5: Classification of target genomic regions using lists 16 to 32

[0437] The distinguishing values ​​of the target genomic regions listed in Lists 16 to 32 were evaluated by testing the ability of a cancer classifier to detect cancer and any of 20 different cancer types based on the methylation status of these target genomic regions. As shown in Table 4, performance was evaluated across 1532 cancer samples and 1521 non-cancer samples not used to train the classifier. For each sample, differently methylated cfDNA was enriched using a decoy set comprising all the target genomic regions listed in Lists 16 to 32. The classifier was then narrowed down to provide cancer determination solely based on the methylation status of the target genomic regions from the evaluated list.

[0438] Table 4

[0439] cfDNA was used to validate individual cancer diagnoses using a classifier.

[0440]

[0441] The results of the performance analysis of classifiers listed in lists 16 to 32 are presented in Tables 5 to 8. An exemplary receiver operator curve (ROC) generated by a trained classifier is shown in... Figure 13 The ROC curves show true positive and false positive results for a determination of cancer or non-cancer based on the methylation status of several target genomic regions in List 23 optimized for lung cancer. The asymmetric shape of the ROC curves illustrates that the classifier is designed to minimize false positives. Except for List 28 (renal cancer), the areas under the curves are closely clustered between 0.77 and 0.80, as shown in Table 5. These results indicate that a single detection of cancer is not significantly impaired by using several test combinations optimized for individual cancer types. Furthermore, the classifier performance was tested on 50% of randomly selected combinations of several target genomic regions in List 20 (colorectal cancer), List 23 (lung cancer), and List 26 (pancreatic and gallbladder cancer). The areas under the ROC curves for these subcombinations of several target genomic regions also clustered closely between 0.77 and 0.80, indicating that the determination of cancer was not detectably impaired by using smaller test combinations of fewer than 400 to 700 target genomic regions with a total test combination size of less than 75 to 140 kb.

[0442] Once a cancer diagnosis is made, the classifier assigns the cancer to one of twenty different cancer types. The accuracy of these diagnoses, with a specificity of 0.990, is presented in various formats. Table 5 shows true positives, false positives, and false negatives, scored based on the methylation status of each list of several target genomic regions optimized for detecting a specific cancer type. A true positive occurs when the presence of cancer is detected and the cancer type is accurately determined. A false positive occurs when cancer is detected and an inaccurate cancer type is scored for samples from individuals diagnosed with the cancer type for which the list is optimized. A false negative occurs when the presence of cancer is detected and the cancer type is inaccurately recorded as the cancer type for which the list is optimized, for samples from several individuals diagnosed with a cancer type different from the cancer type for which the list is optimized.

[0443] Table 5

[0444] Cancer detection and cancer type determination using data from a list of target genomic regions optimized for detection of specific cancer types.

[0445]

[0446]

[0447] The accuracy of cancer detection by a trained classifier, based on the methylation status of several target genomic regions selected for several specific cancer types, is presented for the various cancer types listed in Table 6. When cancer is detected, a cancer type is assigned from one of twenty possible categories of cancer types. The accuracy of cancer type determination is presented in Table 7. The cancer type determination results are the accuracy for determining all twenty cancer types, although the several lists of target genomic regions are optimized to detect a single cancer type.

[0448] The results in Tables 6 and 7 show the separation of various cancer stages. Cancer detection and cancer type determination are more accurate for samples from individuals diagnosed with advanced cancer. This is expected, as advanced tumors shed more cfDNA. Nevertheless, for early-stage cancers, the accuracy in detecting cancer and identifying a specific cancer type is very high. Furthermore, randomly removing 50% of List 20 (colorectal cancer), List 23 (lung cancer), and List 26 (pancreatic and gallbladder cancer) has virtually no impact on classifier accuracy.

[0449] The sensitivity for detecting stage I to IV cancers of various cancer types, with a specificity of 0.990, is shown in Table 8 by a classifier acting on the methylation status of several target genomic regions from a selected list of specific cancer types to be detected. For example, this assumes the false positive rate for cancer detection is limited to 1%. Considering the methylation status of the target genomic regions in List 16, a classifier accurately detected anorectal cancer in 50% (2 out of 4) of samples collected from individuals diagnosed with stage I anorectal cancer. An overall sensitivity of greater than 70% was achieved for all cancer stages, including anorectal cancer, head and neck cancer, liver and bile duct cancer, ovarian cancer, pancreatic and gallbladder cancer, and upper gastrointestinal cancer. Sensitivity for detecting stage I + II cancers was greater than 50% for anorectal cancer, bladder and urethral cancer, head and neck cancer, liver and bile duct cancer, and pancreatic and gallbladder cancer. The sensitivity based on 50% methylation status of randomly selected target genomic regions for colorectal cancer, lung cancer, or pancreatic and gallbladder cancer is substantially the same as the sensitivity using the same number of corresponding target genomic regions. Table 6: Cancer detection accuracy with 99.0% specificity using only a classifier targeting several target genomic regions for the indicated cancer type.

[0450]

[0451]

[0452]

[0453] Table 6 (continued)

[0454]

[0455] Table 7: Accuracy of cancer type determination with 99.0% specificity by using only a classifier targeting several genomic regions for the indicated cancer type.

[0456]

[0457] Table 7 (continued)

[0458]

[0459]

[0460] Table 8: Sensitivity to the indicated cancer type with 99.0% specificity by using only a classifier targeting several genomic regions for the indicated cancer type.

[0461]

[0462]

[0463]

[0464] Table 8 (continued)

[0465]

[0466] Example 6: Cancer detection using a combination of cancer testing tests

[0467] Several blood samples were collected from a group of individuals previously diagnosed with one TOO cancer (“Test Group”) and another group of individuals without cancer or diagnosed with a different type of cancer (“Other Group”). cfDNA fragments were extracted from the blood samples and treated with bisulfite to convert unmethylated cytosine to uracil. The cancer assay suite described herein was applied to the bisulfite-treated samples. Unligated cfDNA fragments were washed, and cfDNA fragments ligated to the probes were collected. The collected cfDNA fragments were amplified and sequenced. The sequence readings confirmed that the probes specifically enriched cfDNA fragments with methylation patterns indicative of one TOO cancer, compared to cfDNA fragments from the Test Group, which had significantly more differentially methylated cfDNA fragments compared to the Other Group.

[0468] While several preferred embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many variations, alterations, and substitutions will now be conceived by those skilled in the art without departing from the present disclosure. It should be understood that various substitutions for the embodiments of the present disclosure described herein can be applied in implementing the present disclosure. The following claims are intended to define the scope of the present disclosure, and the methods and structures within the scope of these claims, and their equivalents, are covered by those claims.

Claims

1. A composition, characterized in that: The composition comprises: several different bait oligonucleotides, (a) each of the several different bait oligonucleotides having a length of at least 45 nucleotides; (b) For each of at least 10 cancer types, the different plurality of decoy oligonucleotides comprise a different set of decoy oligonucleotides; (c) Each group of decoy oligonucleotides is collectively heterozygous into a DNA molecule derived from at least 100 target genomic regions, which are differentially methylated relative to different cancer types or relative to non-cancer types in their respective cancer types, and (d) The total size of the target genome region is 50kb to 5MB.

2. The composition according to claim 1, characterized in that: (a) Each group of decoy oligonucleotides is collectively heterozygous for at least 300 target genomic regions, which are differentially methylated relative to different cancer types or relative to non-cancer types in the corresponding cancer types; or (b) For each group of decoy oligonucleotides, for all possible pairings between the corresponding cancer type and at least 10 other cancer types, the at least 100 target genomic regions include at least one target genomic region, which is differentially methylated between cancer type pairings.

3. The composition according to claim 1, characterized in that, The target genomic regions include: (a) At least 20% of the target genome region or its complement from any one of Lists 1 to 49; (b) At least 20% of the target genome region or its complement from any one of Lists 1 to 15; (c) At least 20% of the target genome regions or their complements as listed in Lists 1 to 15; (d) At least 20% of the target genome region or its complement from any one of Lists 16 to 32; (e) At least 20% of the target genome regions or their complements as listed in Lists 16 to 32; (f) At least 20% of the target genome region or its complement from any one of lists 33 to 49; or (g) At least 20% of the target genome regions or their complements in Lists 33 to 49.

4. The composition according to claim 1, characterized in that: (a) The total size of the target genomic regions is less than 1100 kb; (b) The total number of the target genomic regions is less than 10,000; (c) The DNA molecule is a converted cfDNA fragment; or (d) Each of the several bait oligonucleotides has a length of 45 to 300 nucleotide bases.

5. The composition according to claim 1, characterized in that: (a) Each group of bait oligonucleotides consists of paired bait oligonucleotides; (b) Each pair of bait oligonucleotides includes a first bait oligonucleotide and a second bait oligonucleotide; (c) Each bait oligonucleotide includes a 5′ end and a 3′ end; (d) For each pair of bait oligonucleotides, the sequence of at least X nucleotide bases at the 3' end of the first bait oligonucleotide is identical to the sequence of X nucleotide bases at the 5' end of the second bait oligonucleotide; and (e) where X is at least 25, 30, 35, 40, 45, 50, 60, 70, 75 or 100.

6. The composition of claim 5, characterized in that, The first decoy oligonucleotide comprises a sequence of at least 31, 40, 50, or 60 nucleotide bases, the sequence of which does not overlap with a sequence of the second decoy oligonucleotide.

7. A method for enriching converted cfDNA fragments, said converted cfDNA fragments providing information on a type of cancer, characterized in that: The method includes the following steps: Contact the bait oligonucleotide composition as described in claim 1 with converted cfDNA from a subject, and Samples of cfDNA corresponding to several target genomic regions were enriched through heterozygous capture.

8. The composition according to claim 1, characterized in that: (a) The target genomic regions are several human sequences, and each decoy oligonucleotide is designed to have sequence homology or sequence complementarity with fewer than 20 off-target human genomic regions; (b) Each bait oligonucleotide is at least 61 nucleotides in length; (c) Each bait oligonucleotide is less than 300 nucleotides in length; (d) Each target genomic region includes at least five methylation sites; (e) At least 3% of the bait oligonucleotides do not contain guanine G; or (f) Each bait oligonucleotide includes several binding sites that bind to several methylation sites of the converted cfDNA molecule, wherein at least 83% of the several binding sites consist only of CpG or CpA.

9. A method for detecting several cells of a cancer type, characterized in that, The method includes the following steps: (a) Treating cfDNA from a biological sample with a deamination agent to produce a cfDNA sample comprising several deamination nucleotides; (b) Enriching the cfDNA sample or its amplified product to produce several enriched DNA molecules, wherein: (i) The enrichment includes contacting the cfDNA sample or its amplified product with a composition comprising several different bait oligonucleotides; (ii) Each of the several different bait oligonucleotides is at least 45 nucleotides in length; and (iii) The different bait oligonucleotides are collectively hybridized to at least 100 target genomic regions or their complements from each of the several lists 33 to 49; (c) Sequencing the enriched DNA molecules to generate a set of sequencing reads; and (d) Detecting several sequencing reads of cfDNA molecules from the several cells of the cancer type, thereby detecting the several cells of the cancer type.

10. The method as described in claim 9, characterized in that, (a) The plurality of target genomic regions include target genomic regions selected from List 1 or their complements, and the cancer type is bladder cancer; (b) The plurality of target genomic regions include target genomic regions selected from List 2 or their complements, and the cancer type is breast cancer; (c) The plurality of target genomic regions include target genomic regions selected from List 3 or their complements, and the cancer type is cervical cancer; (d) The plurality of target genomic regions include target genomic regions selected from List 4 or their complements, and the cancer type is arthrorectal cancer; (e) The plurality of target genomic regions include target genomic regions selected from List 5 or their complements, and the cancer type is head and neck cancer; (f) The plurality of target genomic regions include target genomic regions selected from List 6 or their complements, and the cancer type is hepatobiliary cancer; (g) The plurality of target genomic regions include target genomic regions selected from List 7 or their complements, and the cancer type is lung cancer; (h) The plurality of target genomic regions include target genomic regions selected from List 8 or their complements, and the cancer type is melanoma; (i) The plurality of target genomic regions include target genomic regions selected from List 9 or their complements, and the cancer type is ovarian cancer; (j) The plurality of target genomic regions include target genomic regions selected from List 10 or their complements, and the cancer type is pancreatic cancer; (k) The plurality of target genomic regions include target genomic regions selected from List 11 or their complements, and the cancer type is prostate cancer; (l) The plurality of target genomic regions include target genomic regions selected from List 12 or their complements, and the cancer type is kidney cancer; (m) The plurality of target genomic regions include target genomic regions selected from List 13 or their complements, and the cancer type is thyroid cancer; (n) The plurality of target genomic regions include target genomic regions selected from List 14 or their complements, and the cancer type is upper gastrointestinal cancer; or (o) Several target genomic regions include target genomic regions selected from List 15 or their complements, and the cancer type is uterine cancer.

11. The method as described in claim 9, characterized in that, (a) The plurality of target genomic regions include target genomic regions selected from List 16 or List 33 or their complements, and the detection of cancer includes the detection of anorectal cancer; (b) The target genomic regions include target genomic regions selected from List 17 or List 34 or their complements, and the detection of cancer includes the detection of bladder or urethral epithelial cancer; (c) The plurality of target genomic regions include target genomic regions selected from List 18 or List 35 or their complements, and the cancer type is breast cancer; (d) The plurality of target genomic regions include target genomic regions selected from List 19 or List 36 or their complements, and the cancer type is cancer; (e) The plurality of target genomic regions include target genomic regions selected from List 20 or List 37 or their complements, and the cancer type is cervical cancer; (f) The plurality of target genomic regions include target genomic regions selected from List 21 or List 38 or their complements, and the cancer type is head and neck cancer; (g) The plurality of target genomic regions include target genomic regions selected from List 22 or List 39 or their complements, and the cancer type is liver or bile duct cancer; (h) The plurality of target genomic regions include target genomic regions selected from List 23 or List 40 or their complements, and the cancer type is lung cancer; (i) The plurality of target genomic regions include target genomic regions selected from List 24 or List 41 or their complements, and the cancer type is melanoma; (j) The plurality of target genomic regions include target genomic regions selected from List 25 or List 42 or their complements, and the cancer type is ovarian cancer; (k) The plurality of target genomic regions include target genomic regions selected from List 26 or List 43 or their complements, and the cancer type is pancreatic or gallbladder cancer; (l) The plurality of target genomic regions include target genomic regions selected from List 27 or List 44 or their complements, and the cancer type is prostate cancer; (m) The plurality of target genomic regions include target genomic regions selected from List 28 or List 45 or their complements, and the cancer type is kidney cancer; (n) The plurality of target genomic regions include target genomic regions selected from List 29 or List 46 or their complements, and the cancer type is sarcoma; (o) The plurality of target genomic regions include target genomic regions selected from List 30 or List 47 or their complements, and the cancer type is thyroid cancer; (p) The plurality of target genomic regions includes target genomic regions selected from List 31 or List 48 or their complements, and the cancer type is upper gastrointestinal cancer; or (q) Several target genomic regions include target genomic regions selected from List 32 or List 49 or their complements, and the cancer type is uterine cancer.

12. The method as described in claim 9, characterized in that: (a) The plurality of target genome regions include at least 20% of the target genome regions or their complements in each of the corresponding lists; (b) The target genomic regions include genomic regions that comprise less than 90% of each corresponding list or their complements; (c) The plurality of target genomic regions include at least 100 target genomic regions or their complements from each of Lists 33 to 49; (d) The plurality of target genomic regions includes at least 100 target genomic regions or their complements from each of Lists 16 to 32; or (e) The plurality of target regions includes all target regions or their complements from each of Lists 1 to 15.

13. A method for detecting several cells of a certain type of cancer in a subject, characterized in that, The method includes the following steps: (i) Capturing cfDNA fragments or their amplified products from said object using a composition comprising several different bait oligonucleotides, wherein: (a) Each of the several different bait nucleotides is at least 45 nucleotides in length; (b) For each of at least 10 cancer types, the different plurality of bait oligonucleotides comprise different groups of bait oligonucleotides; (c) Each group of decoy oligonucleotides collectively heterozygous for at least 100 target genomic regions, which are differentially methylated relative to different cancer types or relative to non-cancer types in their respective cancer types; and (d) The capture step includes separating the bait-bound DNA from the unbound DNA; (ii) Sequencing the captured cfDNA fragment or its amplified product to generate several sequencing reads; and (iii) For each of the at least 10 cancer types, a trained classifier is applied to the plurality of sequencing reads, wherein the classifier: (a) narrows down the at least 100 target genomic regions of the decoy oligonucleotides to the group for the corresponding cancer type; and (b) assigns a score for each of the at least ten cancer types; and (c) detects the plurality of cells of the cancer type as the cancer type assigned the highest score.

14. The method as described in claim 13, characterized in that, The probability of a false positive detection for the aforementioned number of cells of the cancer type is less than 1%, and the probability of an accurate detection for the aforementioned number of cells of the cancer type is at least 40%.

15. The method as described in claim 13, characterized in that, The cfDNA fragment is a converted cfDNA fragment.

16. The method as described in claim 13, characterized in that, The at least 10 cancers are selected from the following combinations: thyroid cancer, melanoma, sarcoma, kidney cancer, prostate cancer, breast cancer, uterine cancer, ovarian cancer, bladder cancer, urothelial carcinoma, cervical cancer, anorectal cancer, head and neck cancer, colorectal cancer, liver cancer, bile duct cancer, pancreatic cancer, gallbladder cancer, upper gastrointestinal cancer, and lung cancer.

17. The method as described in claim 16, characterized in that: (a) The cancer type is a stage I cancer type, and the probability of specifying an accurate cancer type is at least 70%; (b) The cancer type is a stage II cancer type, and the probability of specifying an accurate cancer type is at least 85%; (c) The cancer type is a stage I or a stage II cancer type, and the probability of specifying an accurate cancer type is at least 75%; or (d) An accurate cancer type is specified at least 80%.

18. The method as described in claim 13, characterized in that: (a) The type of cancer is anorectal cancer, and the sensitivity to anorectal cancer is at least 65% or 75%; (b) The cancer type is bladder and urothelial carcinoma, with a sensitivity of at least 40% for bladder and urothelial carcinoma; (c) The type of cancer is breast cancer, and the sensitivity to breast cancer is at least 20%; (d) The type of cancer is cervical cancer, and the sensitivity to cervical cancer is at least 25%; (e) The type of cancer is colorectal cancer, and the sensitivity to colorectal cancer is at least 55%; (f) The type of cancer is head and neck cancer, and the sensitivity to head and neck cancer is at least 70%; (g) The type of cancer is hepatobiliary cancer, and the sensitivity to hepatobiliary cancer is at least 75%; (h) The type of cancer is lung cancer, and the sensitivity to lung cancer is at least 55%; (i) The type of cancer is melanoma, and the sensitivity to melanoma is at least 30%; (j) The type of cancer is ovarian cancer, and the sensitivity to ovarian cancer is at least 70%; (k) The cancer type is pancreatic and gallbladder cancer, and the sensitivity for pancreatic and gallbladder cancer is at least 60%; (l) The cancer type is sarcoma, and the sensitivity to sarcoma is at least 40%; or (m) The type of cancer is upper gastrointestinal cancer, and the sensitivity to upper gastrointestinal cancer is at least 60%.

19. The method as described in claim 13, characterized in that, Each group of decoy oligonucleotides collectively hybridizes to at least 300 target genomic regions, which are differentially methylated relative to different cancer types or relative to non-cancer types in their respective cancer types.

20. The method as described in claim 13, characterized in that, The total size of the target genome region is between 50kb and 4MB.

21. The method as described in claim 13, characterized in that: (a) The subject has an increased risk of one or more types of cancer; (b) The subject exhibits several symptoms associated with one or more types of cancer; or (c) The subject has not been diagnosed with cancer.

22. The method as described in claim 13, characterized in that, The classifier is trained on several transformed DNA sequences derived from at least 100 subjects with a first cancer type, at least 100 subjects with a second cancer type, and at least 100 subjects without cancer; and wherein the first cancer type, the second cancer type, and the third cancer type are selected from the at least 10 cancer types.

23. The method as described in claim 13, characterized in that, The classifier is trained on several transformed DNA sequences derived from the several target genomic regions, and the several different decoy oligonucleotides are collectively hybridized to at least 100 target genomic regions or their complements from each of Lists 33 to 49.

24. The method as described in claim 23, characterized in that, The trained classifier detects the number of cells of the cancer type in the following manner: (i) Generate a set of several features for a sample, wherein each feature in the set of several features includes a numerical value; (ii) Input several features of the group into the classifier, wherein the classifier includes a multinomial classifier; (iii) Based on several features of the set, a set of probability scores is determined in the classifier, wherein the set of probability scores includes a probability score for each cancer type category and a probability score for each non-cancer type category; and (iv) The set of probability scores is measured by a threshold based on one or more values ​​determined during the training of the classifier.

25. The method as described in claim 24, characterized in that: (a) The set of features includes a set of binary features; (b) The numerical value includes a single binary value; (c) The multinomial classifier includes a multinomial logistic regression ensemble, trained to predict the origin of the cancer; or (d) The method further includes determining the final cancer classification relative to a minimum value, based on the difference between two highest probability scores, wherein the minimum value corresponds to a predefined percentage of training cancer samples, the predefined percentage of training cancer samples being assigned the correct cancer type as the highest score during the training of the classifier.

26. The method as described in claim 13, characterized in that, The method further includes administering a chemotherapeutic agent to the subject; optionally, the chemotherapeutic agent is selected from the group consisting of alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, cytoskeleton disruptors (taxanes), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, and platinum-based reagents.

Citation Information

Patent Citations

  • Anomalous fragment detection and classification

    US12027237B2

  • Methylation haplotyping for non-invasive diagnosis (MONOD)

    US20160340740A1

  • Selective oxidation of 5-methylcytosine by TET-family proteins

    WO2010037001A2

  • Composition and methods related to modification of 5-hydroxymethylcytosine (5-HMC)

    WO2011127136A1

  • Non-invasive determination of methylome of fetus or tumor from plasma

    WO2014043763A1