Sample barcodes in multiplex sample sequencing
The use of unique sample barcodes and UMIs in next-generation sequencing improves cancer detection by accurately assigning reads and filtering out contamination, addressing inefficiencies in current methods and enabling earlier, more precise cancer classification.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-13
- Publication Date
- 2026-03-25
AI Technical Summary
Current cancer detection methods are cancer type-specific, inefficient, and prone to high false-positive rates, with low detection rates for early-stage cancers, and multiplex sequencing techniques suffer from sample contamination and index hopping events leading to biased analysis.
A system utilizing unique sample barcodes and unique molecular identifiers (UMIs) is employed to accurately assign sequence reads to samples, detect contamination events, and improve next-generation sequencing by filtering out index hopping errors.
Enhances the accuracy of cancer classification by reducing contamination and index hopping, enabling earlier and more precise detection of cancer types.
Smart Images

Figure 2026509734000001_ABST
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims the benefits and priority thereto of U.S. Provisional Patent Application No. 63 / 489,848, filed March 13, 2023, which is incorporated by reference. [Background technology]
[0002] Cancer is the leading cause of death worldwide. Cancer mortality rates are exacerbated by the fact that cancer is typically detected late, limiting the effectiveness of treatment options for long-term survival. Current detection methods are generally cancer type-specific, meaning each cancer type is screened individually. Each screening process is tailored to the specific cancer type. For example, mammography is used to detect breast cancer, while colonoscopy or stool tests help detect colorectal cancer. Each of these various screening methods is not necessarily applicable to other cancer types. For instance, to screen an individual for three different possible cancer types, a healthcare provider would need to perform, or instruct, three different screening processes. Each of these screening processes may require a combination of invasive and / or non-invasive procedures to identify tumor growth, collect biopsies of the growths, and perform analysis of some tissue biopsies.
[0003] Furthermore, current screening methods are hampered by low detection rates or high false-positive rates. Low detection rates often fail to detect cancer in its early stages. High positive rates misdiagnose cancer-free individuals as positive for cancer status. As a result, most screening tests are only practical when used to test individuals at high risk of developing the cancer being screened, and have limited ability to detect cancer in the general population.
[0004] Novel research suggests that abnormal DNA methylation is involved in many disease processes, including cancer. DNA methylation plays a role in regulating gene expression. Therefore, abnormal DNA methylation can cause problems in normal gene expression pathways, leading to cancer or other diseases. For example, specific patterns of differentially methylated regions may be useful as molecular markers for various disease states. Nevertheless, such models face many problems. Early cancer detection is particularly difficult because the ratio of tumor cells to non-cancerous cells in the study is extremely small. The extremely small ratio can be on the order of 1:1000, 1:10,000, or even 1:100,000. This presents the challenge of detecting small amounts of cancer signaling surrounded by healthy signals. Furthermore, DNA containing age-related genetic mutations that often resemble cancerous abnormal methylation can be released by blood cells. These informative methylated fragments released from blood cells often artificially amplify cancer signaling.
[0005] In one or more implementations, the early cancer detection workflow involves running a next-generation sequencer capable of high-throughput sequencing. In such implementations, these sequencing devices may include multiplex sequencing capabilities to sequence molecules from multiple samples. To assist multiplex sequencing, a library of indices is used to tag and identify sequence reads across various samples. However, index hopping events in the sequencing process can lead to misassignment of sequence reads to the wrong sample. Such misassignments of reads can bias downstream analyses, for example, inaccurately detecting cancer in one sample based on misassigned sequence reads from another sample. [Overview of the project] [Problems that the invention aims to solve]
[0006] This disclosure aims to address the aforementioned issues. The background information provided herein is intended to provide a general overview of the context of this disclosure. Unless otherwise stated herein, the elements described in this section are not prior art to the claims of this application, nor are they incorporated into this section construed as prior art or an indication of prior art. [Means for solving the problem]
[0007] Early detection of disease conditions (such as cancer) in subjects is crucial because it allows for earlier treatment and, consequently, a higher chance of survival. Sequencing of DNA fragments in cell-free (cf) DNA samples can identify features that can be used for disease classification. For example, in cancer assessment, features based on cell-free DNA from blood samples (presence or absence of somatic variants, methylation status, or other genetic abnormalities) can provide insight into whether a subject may have cancer and further insight into what type of cancer they may have. To achieve this objective, this specification includes systems and methods for detecting contamination of nucleic acid samples for the purpose of analyzing sequencing data for cancer assessment.
[0008] This disclosure addresses the problems identified above by providing an improved system and method for detecting sample contamination of contaminated fragments for cancer classification. The system utilizes a substantially unique sample barcode sequence ligated to nucleic acid (NA) molecules derived from a single sample. In addition to one or more indices in a sequencing library, the sample barcode sequence facilitates the precise assignment of sequence reads to the appropriate sample, thereby improving next-generation sequencing techniques and, more generally, the physical assignment process.
[0009] The analysis system first evaluates the index of sequence reads from the multiplex sequencing process. Sequence reads with similar indices are matched against each other. Sequence reads with mismatched indices (two indices that are not paired in the sequencing library) may be determined to be a single-index hopping event. Sequence reads with matching indices but different sample barcodes are determined to be a double-index hopping event. If the index and / or sample barcode sequences are different but close, the analysis system may determine and identify a sequencing error.
[0010] In one or more embodiments, NA molecules in the sample are also ligated with a unique molecular identifier (UMI), which facilitates accurate duplication of sequence reads belonging to the same original NA molecule. Various configurations of sample barcodes, UMIs, and indices can be performed. In one or more configurations, the sample barcode and UMI on each NA molecule are located at opposite ends of the NA molecule. In other configurations, the sample barcode and UMI are ligated to the same ends of the NA molecule. In further configurations, single-index sequencing libraries or double-index sequencing libraries are used instead.
[0011] Clause 1. A method comprising: ligating one of a plurality of molecular identifiers (MIs) to the first end of each nucleic acid (NA) fragment of a first sample, wherein at least two of the plurality of MIs are different from each other; amplifying the NA fragments to produce amplified NA fragments containing one or more copies of each NA fragment; ligating a first sample barcode to the amplified NA fragments of the first sample; indexing the amplified NA fragments to produce indexed NA fragments containing a first index and a second index, respectively; sequencing the indexed NA fragments to produce sequence reads for each indexed NA fragment; collecting sequence reads of indexed NA fragments containing a first index sequence read and a second index sequence read, respectively, within a group; and identifying contamination events of the first sequence reads within a group by identifying a second sample barcode different from the first sample barcode on the sequence reads.
[0012] Clause 2. The method according to Clause 1 or any of the provisions subordinate thereto, wherein each amplified NA fragment having a first sample barcode includes a target NA region derived from the sample, and each amplified NA fragment having a second sample barcode includes a target NA region derived from a second sample different from the first sample.
[0013] Clause 3. The method of Clause 1 or any of the provisions subordinate thereto, further comprising removing a first sequence read having an index-hopping event from a first group of samples.
[0014] Clause 4. The method of Clause 1 or any of the provisions subordinate thereto, further comprising creating a binary cancer prediction between the presence and absence of cancer based on sequence reads in the group from which the first sequence read has been excluded.
[0015] Clause 5. The method of Clause 1 or any of the provisions subordinate thereto, further comprising creating a multiclass cancer prediction among multiple cancer types based on sequence reads in a group from which a first sequence read has been excluded.
[0016] Clause 6. One of the plurality of MIs and the NA fragment are single-stranded during ligation, the method described in Clause 1 or any clause subordinate thereto.
[0017] Clause 7. The indexed NA fragment contains a first index at the first end and a second index at the second end, respectively, the method described in Clause 1 or any clause subordinate thereto.
[0018] Clause 8. The first end is the 3' end, the method described in Clause 1 or any clause subordinate thereto.
[0019] Clause 9. The first end is the 5' end, the method described in Clause 1 or any clause subordinate thereto.
[0020] Clause 10. The first sample barcode is ligated to the second end of the amplified NA fragment on the opposite side of the first end, the method described in Clause 1 or any clause subordinate thereto.
[0021] Clause 11. The first sample barcode is ligated to the first end of the amplified NA fragment adjacent to one of the plurality of MIs, the method described in Clause 1 or any clause subordinate thereto.
[0022] Clause 12. The first index is ligated to the first end, and the second index is ligated to the second end on the opposite side of the first end, the method described in Clause 1 or any clause subordinate thereto.
[0023] Clause 13. The first index and the second index are ligated to the first end of the amplified NA fragment, the method described in Clause 1 or any clause subordinate thereto.
[0024] Clause 14. The first index and the second index are ligated to the second end of the amplified NA fragment on the opposite side of the first end, the method described in Clause 1 or any clause subordinate thereto.
[0025] Clause 15. The sequencing is multiplexed with multiple samples across multiple flow cells, and the first and second indices are used for the first column in which the first sample resides, as described in Clause 1 or any of the provisions subordinate thereto.
[0026] Clause 16. A third index and a fourth index, different from the first and second indexes, are used for the second column as described in Clause 1 or any of the clauses subordinate thereto.
[0027] Clause 17. Sequencing is a method of Clause 1 or any of the provisions relating thereto, including targeted methylation sequencing or whole-genome bisulfite sequencing.
[0028] Clause 18. Each MI has a length selected from the range of 3 nucleic acid bases to 20 nucleic acid bases, as described in Clause 1 or any of the clauses subordinate thereto.
[0029] Clause 19. The first sample barcode has a length selected from the range of 3 nucleic acid bases to 20 nucleic acid bases, as described in Clause 1 or any of the clauses subordinate thereto.
[0030] Clause 20. The first sample barcode is substantially unique among multiple sample barcodes, as described in Clause 1 or any of the provisions subordinate thereto.
[0031] Clause 21. The method according to Clause 1 or any of the sub-clauses, wherein the NA fragment is a cell-free deoxyribonucleic acid (cfDNA) fragment.
[0032] Clause 22. Amplification of an NA fragment is performed in accordance with the method of Clause 1 or any of the provisions subordinate thereto, including performing linear amplification.
[0033] Clause 23. A method for calling a contamination event, comprising receiving a plurality of sequence reads of an amplified fragment, each of which includes a first index, a second index, and a sample barcode; matching the sequence reads of the amplified fragment having the first index and the second index in a first bag of a first sample, the first sample barcode being assigned to the first sample; and calling a contamination event of the first sequence read in the first bag based on the matching and identifying a second sample barcode on the sequence read that is different from the first sample barcode.
[0034] Clause 24. The method of Clause 23 or any of the subclased Clauses, wherein the first index is at the first end of the amplified segment, and the second index is at the second end of the amplified segment.
[0035] Clause 25. The first and second indices are the methods described in Clause 23 or any of the subclaves thereof, located at the first end of the amplified fragment.
[0036] Clause 26. The method according to Clause 23 or any of the provisions subordinate thereto, wherein each amplified fragment further comprises a molecular identifier (MI), an amplified fragment derived from a first nucleic acid fragment of the sample comprises a first MI, and an amplified fragment derived from a second nucleic acid fragment of the sample comprises a second MI different from the first MI.
[0037] Clause 27. MI is located at the first end of the amplified fragment, and the sample barcode is ligated on the first end, as described in Clause 23 or any of the clauses subordinate thereto.
[0038] The method of Clause 28. MI is located at the first end of the amplified fragment, and the sample barcode is ligated onto the second end opposite the first end, as described in Clause 23 or any of the subclaves thereof.
[0039] Clause 29. The method of Clause 23 or any of the provisions subordinate thereto, further comprising removing a first sequence read having a contamination event from a first bag of the first sample.
[0040] Clause 30. The method according to Clause 29, further comprising making a cancer prediction for a first sample based on the sequence reads in a first bag from which the first sequence reads have been removed.
[0041] Clause 31. Cancer prediction is a binary prediction between the presence and absence of cancer, as described in Clause 30.
[0042] Clause 32. Cancer prediction is a multi-class prediction across multiple cancer types, as described in Clause 30 or any of the clauses subordinate thereto.
[0043] Clause 33. Developing a cancer prediction comprises dividing sequence reads into distinct sets of sequence reads, determining an informative score for each distinct sequence read, the informative score indicating the likelihood of observing the distinct sequence read in a healthy population of the sample, and identifying a set of informative fragments by comparing the informative score of the distinct sequence reads to an informative score threshold, and the cancer prediction is based on the set of informative fragments, as described in Clause 30 or any of the provisions subordinate thereto.
[0044] The method according to Clause 33, further comprising: creating a feature vector for a first sample based on a set of informative fragments; and inputting the feature vector into a cancer classifier to determine a cancer prediction, wherein the cancer classifier is trained on at least a first cohort of cancer samples and a second cohort of non-cancer samples.
[0045] Clause 35. The first sample barcode has a length selected from the range of 3 nucleic acid bases to 20 nucleic acid bases, as described in Clause 23 or any of the clauses subordinate thereto.
[0046] Clause 36. A method for processing sequencing data, comprising receiving sequencing data comprising a set of sequence reads generated from multiplex sequencing of multiple biological samples, each containing nucleic acids, wherein the sequencing data comprises data generated from single-index hopping and double-index hopping events occurring during multiplex sequencing, and filtering the sequencing data to exclude data corresponding to single-index hopping events, wherein the filtering comprises identifying one or more reads in the set of sequence reads that have index mismatch pairs, where the index mismatch pairs are from two different biological samples. A method comprising identifying and filtering sequencing data to exclude data corresponding to double index hopping events, wherein the filtering includes identifying one or more pad-hopping duplicate reads in a set of sequence reads, each pad-hopping duplicate read having a duplicate sequence coexisting in a flow cell used during multiplex sequencing, and subsequently identifying one or more singletons in a set of sequence reads, each singleton containing a unique sequence read in the set of sequence reads.
[0047] Clause 37. The method of Clause 36 or any subordinate Clause, comprising removing one or more identified pad-hopping duplicate reads from a set of sequencing reads, and, following the removal of one or more identified pad-hopping duplicate reads, identifying one or more singletons in the remaining set of sequencing reads.
[0048] Clause 38. Filtering is the method of Clause 36 or any of its subordinate clauses, including flagging in the sequencing data one or more identified reads having a mismatch index, one or more identified pad-hopping duplicate reads, or one or more identified singletons.
[0049] Clause 39. Filtering is performed in the manner of Clause 36 or any of the provisions subordinate thereto, including removing from sequencing data one or more identified reads having mismatch indexes, one or more identified pad-hopping duplicate reads, or one or more identified singletons.
[0050] Clause 40. The method of Clause 36 or any subordinate Clause, wherein a flow cell comprises a plurality of physically separated lanes, each lane comprising a plurality of columns, each column comprising a plurality of tiles, and each lane further defining a surface having a plurality of wells arranged thereon.
[0051] Clause 41. Sequence data includes spatial information indicating the position of each sequence read on a flow cell, wherein the spatial information includes at least one of a lane ID, column ID, tile ID, or xy coordinate pair, as described in Clause 36 or any of its subordinate clauses.
[0052] The method of identifying padhopping duplicate reads as described in Clause 42 or any of the Clauses subordinating thereto, which includes identifying groups of identical or nearly identical sequence reads, determining whether grouping reads coexist based on sequence data, where grouping reads coexist if at least one of the following spatial relationships is met: grouping reads share a common tile, grouping reads are located on neighboring tiles, grouping reads are located on different tiles within a common column, grouping reads are located within threshold x-distances and threshold y-distances from each other on a flow cell, or grouping reads are located within a predefined boundary region, and identifying grouping reads as padhopping duplicate reads in accordance with the determination that grouping reads coexist.
[0053] Clause 43. The method of Clause 42 or any of its subordinate clauses, which includes identifying multiple groups of sequence reads as pad-hopping duplicate reads, each group containing coexisting identical or nearly identical sequence reads.
[0054] Clause 44. A predefined boundary area includes a geometric shape having an x-distance of 7,500 flow cell position units and a y-distance of 100,000 flow cell position units, as described in Clause 42 or any of its subordinate clauses.
[0055] Clause 45. The method of Clause 42 or any of the subclaves thereof, wherein the geometric shape includes a rectangle, and the longer side of the rectangle extends along the lane in the y-direction of the flow cell in the longitudinal direction.
[0056] Clause 46. The method of Clause 42 or any subordinate clause, wherein the threshold x distance between grouped leads is in the range of 0 to 50 mm, and the threshold y distance between grouped leads is in the range of 0 to 50 mm.
[0057] Clause 47. The method of Clause 36 or any of its subordinate clauses, including removing pad-hopping duplicate reads if the expected error rate associated with multiplex sequencing exceeds a threshold error rate.
[0058] Clause 48. The method of Clause 36 or any subordinate clause, comprising providing filtered sequencing data for analysis using a statistical model, wherein the detection limits associated with the filtered sequencing data are lower than the detection limits associated with the unfiltered sequencing data.
[0059] Clause 49. Nucleic acids are those described in Clause 36 or any of the provisions relating thereto, including cell-free DNA (cfDNA) or cell-free RNA (cfRNA).
[0060] Clause 50. Nucleic acids include genomic DNA (gDNA) as defined in Clause 36 or any of the provisions relating thereto.
[0061] Clause 51. The method of Clause 36 or any of the Clauses subordinating thereto, comprising: fragmenting nucleic acids extracted from multiple biological samples into genomic fragments; ligating unique dual index pairs to the terminal portions of the genomic fragments to generate multiple library fragments, each unique dual index pair identifying an individual biological sample in the multiple biological samples; enriching and amplifying the library fragments by capturing a specific library fragment with a targeted probe and amplifying the captured library fragment in multiple wells on a flow cell, each well configured to hold a clonal cluster of amplified fragments resulting from a single library fragment; and sequencing the enriched fragments to generate sequencing data comprising a set of sequence reads, each sequence read comprising multiple nucleotide base calls; and demultiplexing the set of sequence reads based on the unique dual index pairs to determine the original biological sample for each sequence read.
[0062] Clause 52. The method according to Clause 51, further comprising filtering the sequencing data and then demultiplexing the sequencing data to remove data corresponding to single index hopping events and double index hopping events.
[0063] Clause 53. A method for processing sequencing data, comprising: receiving sequencing data comprising a set of sequence reads generated from multiplex sequencing of multiple biological samples containing nucleic acids, wherein the sequencing data comprises data generated from single-index hopping and double-index hopping events occurring during multiplex sequencing; filtering the sequencing data to exclude data corresponding to single-index hopping events, wherein the filtering comprises identifying one or more reads in the set of sequence reads that have index mismatch pairs, where each index mismatch pair comprises two unique indices corresponding to two different biological samples; and filtering the sequencing data to exclude data corresponding to double-index hopping events, wherein the filtering comprises identifying one or more singletons in the set of sequence reads, where each singleton comprises a unique sequence read in the set of sequence reads.
[0064] Clause 54. A method for training a cancer classifier, comprising: obtaining a plurality of sequence reads of amplified fragments by performing next-generation multiplex sequencing on a set of training samples, each having a known cancer state, each having a ligated index pair on a target region of a nucleic acid fragment and a sample barcode; matching the sequence reads into a plurality of bags based on the index pairs of the sequence reads, each bag having a sequence read having a common index pair; detecting a sample cross-contamination event in a first bag for a first sequence read having a first sample barcode different from the other sequence reads; removing the first sequence read from the first bag; assigning the remaining sequence reads in each bag to a single training sample; determining a feature vector for each training sample based on the sequence reads in the corresponding bag; and training a cancer classifier with the feature vectors of the training samples, the trained cancer classifier being configured to predict the likelihood of the presence of cancer based on input feature vectors derived from sequence reads in a test sample.
[0065] Clause 55. The method according to Clause 54 or any of the preceding clauses, wherein each sequence read further comprises a molecular identifier (MI) ligated onto a corresponding nucleic acid fragment, and the method further comprises separating the sequence reads in each bag into separate sequence reads based on the molecular identifiers, determining sequence reads having overlapping and matching molecular identifiers at similar genomic locations and being read onto the same nucleic acid fragment.
[0066] Clause 56. The method of Clause 54 or any of the provisions relating thereto, further comprising assigning a first sample barcode to a second bag, wherein the sequence reads in the second bag have the first sequence reads.
[0067] Clause 57. A cancer classifier is a machine learning model, as described in Clause 54 or any of the clauses relating thereto.
[0068] Clause 58. A cancer classifier is trained to predict a binary prediction between the presence and absence of cancer, as described in Clause 54 or any of the provisions thereof.
[0069] Clause 59. Each training sample is known to have one of several cancer types, as described in Clause 54 or any of the clauses subordinate thereto.
[0070] Clause 60. The method according to Clause 59, wherein the cancer classifier is trained to predict multiclass predictions as the likelihood of the presence of one of multiple cancer types.
[0071] Clause 61. The sequence reads of each sample are methylated sequence reads, and the feature vector is based on the methylated sequence reads having an informative methylation pattern, as described in Clause 54 or any of the subclaves thereof.
[0072] Clause 62. A non-temporary computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system, wherein one or more programs include instructions for carrying out the methods described in any of Clauses 1 to 61.
[0073] Clause 63. A computer system comprising one or more processors and a non-temporary computer-readable storage medium as described in Clause 62.
[0074] Clause 64. A processing kit comprising a collection container for collecting a biological sample from a subject, optionally one or more reagents for isolating DNA fragments in the biological sample, optionally a library of paired indices, optionally multiple copies of a first sample barcode for ligating onto DNA fragments of the biological sample, and optionally a non-temporary computer-readable storage medium as described in Clause 62.
[0075] Clause 65. A set of nucleic acid (NA) constructs comprising a first NA construct comprising a first sample-specific barcode comprising a first target NA region, a first barcode sequence, and a first index, and a second NA construct comprising a second sample-specific barcode comprising a second target NA region, a second barcode sequence, and a first index, wherein the first NA construct is identified as originating from a first sample based on the first sample-specific barcode, and the second NA construct is identified as originating from a second sample based on the second sample-specific barcode.
[0076] Clause 66. The first sample-specific barcode and the first index are located at the first end of the first NA construct, in the set of NA constructs described in Clause 65 or any of the subordinating clauses.
[0077] Clause 67. A set of NA constructs as described in Clause 65 or any of its subordinate clauses, wherein the first sample-specific barcode is located at the first end of the first NA construct, and the first index is located at the second end of the first NA construct opposite the first end.
[0078] Clause 68. The first NA construct and the second NA construct are placed together in a bag based on identifying the first index, as part of the set of NA constructs described in Clause 65 or any of the clauses subordinate thereto.
[0079] Clause 69. A set of NA constructs described in Clause 65 or any subordinating Clause, where the first NA construct includes the second index, and the second NA construct includes the second index.
[0080] Clause 70. The first and second indices are sets of NA constructs described in Clause 65 or any of the subordinating Clauses, which are located at the first end of the first NA construct.
[0081] Clause 71. A set of NA structures described in Clause 65 or any of its subordinate clauses, wherein the first index is located at the first end of the first NA structure, and the second index is located at the second end of the first NA structure opposite the first end.
[0082] Clause 72. The first NA construct and the second NA construct are a set of NA constructs described in Clause 65 or any of the subordinating Clauses, which are bagged together based on identifying the first index and the second index.
[0083] Clause 73. A set of NA constructs described in Clause 65 or any of its subordinate clauses, wherein the first barcode is located at the first end of the first NA construct, and the first NA construct further includes a molecular identifier (MI) at the second end of the first NA construct opposite the first end.
[0084] Clause 74. A set of NA constructs described in Clause 65 or any of the subordinating Clauses, where at least one of the first NA construct and the second NA construct is an amplification construct.
[0085] Clause 75. A set of NA constructs described in Clause 65 or any of the subordinating clauses, wherein the first sample-specific barcode is substantially unique with respect to the second sample-specific barcode.
[0086] Clause 76. A set of NA constructs as described in Clause 65 or any of its subordinate clauses, wherein the first NA construct is constructed on the first column, the second NA construct is constructed on the second column, and the first index is assigned to the first column but index-hops to the second column. [Brief explanation of the drawing]
[0087] [Figure 1] An exemplary flowchart illustrating the overall workflow for cancer classification of a sample, according to one or more embodiments, is shown. [Figure 2A]An exemplary flowchart illustrating the process of sequencing cell-free (cf) DNA fragments to obtain methylation state vectors, according to one or more embodiments, is shown. [Figure 2B] Figure 3A illustrates an exemplary process of sequencing cell-free (cf) DNA fragments to obtain a methylation state vector, according to one or more embodiments. [Figure 3A] An exemplary flowchart illustrating the sample processing process using a sample barcode, according to one or more embodiments, is shown. [Figure 3B] An exemplary flowchart illustrating a process for detecting contamination events using sample barcodes in part, according to one or more embodiments, is shown. [Figure 4A] Exemplary illustrations of sample processing according to one or more embodiments are shown. [Figure 4B] Figure 4A illustrates an exemplary sample sequencing process following sample processing, according to one or more embodiments. [Figure 5A] Exemplary illustrations of contamination detection according to one or more embodiments are shown. [Figure 5B] Exemplary illustrations of nucleic acid constructs incorporating sample barcodes into patterned flow cells, according to one or more embodiments, are shown. [Figure 6] Exemplary illustrations of various configurations of sample barcodes and unique molecular identifiers according to one or more embodiments are shown. [Figure 7] An exemplary flowchart is shown illustrating a process 700 for identifying index-hopping events, according to one or more embodiments. [Figure 8A] An exemplary flowchart illustrating the process of creating a control group data structure for determining informative methylation fragments, according to one or more embodiments, is shown. [Figure 8B] An exemplary flowchart illustrating the process of determining an informative methylation fragment based on a control group data structure, according to one or more embodiments, is shown. [Figure 9A]An exemplary flowchart illustrating the process of training a cancer classifier according to one or more embodiments is shown. [Figure 9B] Examples of creating feature vectors used to train a cancer classifier are illustrated using one or more embodiments. [Figure 10A] An exemplary flowchart of a device for sequencing nucleic acid samples, according to one or more embodiments, is shown. [Figure 10B] An exemplary block diagram of an analysis system according to one or more embodiments is shown. [Modes for carrying out the invention]
[0088] These figures illustrate various embodiments for illustrative purposes. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods illustrated herein may be used without departing from the principles described herein.
[0089] I. Overview Early detection and classification of cancer are crucial skills. Detecting cancer before symptoms appear benefits all parties involved, including patients, physicians, and their families. For patients, early cancer detection increases the likelihood of a favorable outcome; for physicians, it opens up more treatment options that can lead to favorable outcomes; and for family and friends, early cancer detection increases the chances of avoiding the loss of loved ones to the disease.
[0090] In recent years, early cancer detection technologies have advanced in the direction of analyzing gene fragments (e.g., DNA) in a person's blood and determining whether a portion of those gene fragments originates from cancer cells. These new technologies allow physicians to detect the presence of cancer in patients that might otherwise be undetectable by conventional screening processes. Consider, for example, a person at high risk of breast cancer. Traditionally, such individuals undergo regular doctor visits for mammography, during which images of breast tissue are created (e.g., X-ray images) which are then used by the physician to identify cancerous tissue. Unfortunately, even with the highest resolution mammography, physicians cannot detect tumors until they reach approximately 1 millimeter in size. This means that the cancer has been present in the person for some time, undiagnosed, and untreated. Such visual determination is common for most cancers, meaning they are only identifiable after they have grown to a sufficient size, and even then, only using certain imaging techniques.
[0091] This problem can be improved by detecting cancer using analysis of gene fragments in a patient's blood, for example. For instance, cancer cells begin releasing DNA fragments into the human bloodstream as soon as they form. This occurs when even a small number of cancer cells are present and before they can be visually detected by imaging techniques. Using appropriate methods, a system that analyzes DNA fragments in the bloodstream can identify the presence of cancer in a person based on the released cancer DNA fragments, and more importantly, it can do so before cancer can be identified using conventional cancer detection techniques.
[0092] Cancer detection based on the analysis of DNA fragments is made possible by next-generation sequencing (NGS) technology. In a broad sense, NGS refers to a group of technologies that enable high-throughput sequencing of genetic material. As described in detail herein, NGS mainly consists of (1) sample preparation, (2) DNA sequencing, and (3) data analysis. Sample preparation is the laboratory technique required to prepare DNA fragments for sequencing, sequencing is the process of reading sequenced nucleotides in a sample, and data analysis is the process of processing and analyzing the genetic information in the sequencing data to identify the presence of cancer.
[0093] While these processes in NGS can facilitate early cancer detection, they also introduce their own complex and detrimental problems to cancer detection. Therefore, cancer detection techniques can be improved and early cancer detection can become more common through sample preparation, DNA sequencing and / or data analysis, such as pretreatment, algorithmic processing and summarization or presentation of predictions or conclusions.
[0094] For example, (1) problems that arise in sample preparation include DNA sample quality, sample contamination, fragmentation bias, and accurate indexing. Improving these issues would result in better genetic data for cancer detection. Similarly, (2) problems that arise in sequencing include, for example, errors in accurate transcription of fragments (e.g., reading "A" instead of "C"), inaccurate or difficult fragment assembly and overlap, varying coverage uniformity, insufficient sequencing depth vs. cost vs. specificity, and insufficient sequencing length. Improving any of these issues would also lead to improved genetic data for cancer detection.
[0095] (3) The problems in data analysis are the most difficult and complex. The challenges that arise stem from the enormous amount of data generated by NGS sequencing technology. The resulting genetic datasets are typically terabytes in scale, and analyzing such a volume of data effectively and efficiently is extremely difficult, both procedurally and computationally. For example, the analysis of NGS sequencing involves several baseline processing steps, such as read alignment, read-to-reference genome alignment and mapping, variant gene identification and calling, abnormal methylation gene identification and calling, and functional annotation. Performing any of these processes on terabytes of genetic data is computationally expensive even with the most powerful computer architectures, and is completely impossible for the human brain. In addition, genetic sequencing data obtained from error-prone processes of sample preparation and sequence reading may result in a large portion of the resulting genetic data being of low quality or unusable for cancer detection. For example, large amounts of genetic data may contain contaminated samples, transcription errors, mismatched regions, and over-recurring regions, making them unsuitable for high-precision cancer detection. Identifying and taking into account low-quality genetic data across the vast amount of genetic data obtained from NGS sequencing requires rigorous procedures and calculations, which is virtually impossible for the human brain to achieve. Overall, creating processes that make the processing of large sequence sequencing data more efficient would lead to improvements in cancer detection using NGS sequencing.
[0096] Lastly, and perhaps most importantly, accurately identifying informative DNA from NGS data to pinpoint the presence of cancer is also difficult (even more so in relation to early cancer detection). To be effective, algorithms must compensate for errors arising from, for example, sample preparation and sequencing, and overcome the problems of large-scale data analysis associated with NGS technology. In other words, the design of machine learning models or other computer processing algorithms that enable early cancer detection based on next-generation sequencing technology must be configured to highlight the problems posed by those technologies. Some of these technologies and models are discussed below, and specific improvements to existing technologies and models are further considered.
[0097] This disclosure addresses some of the problems specified above by providing an improved system and method for detecting sample contamination of contaminated fragments for cancer classification. The system utilizes a substantially unique sample barcode sequence ligated to nucleic acid (NA) molecules derived from a single sample. The sample barcode sequence, in addition to one or more indices in the sequencing library, facilitates the accurate assignment of sequence reads to appropriate samples, thereby resulting in improvements to next-generation sequencing techniques, and more generally, improvements to physical assay processes. Ligation of sample barcodes upstream in multiplex sequencing provides a basis for increasing the number of multiplexes, i.e., increasing the number of samples sequenced together in a single flow cell. In particular, the additional layer of contamination detection with sample barcodes identifies and detects a higher frequency of contamination (e.g., index-hopping events), preventing biased analysis from contaminated samples.
[0098] When contaminated samples are detected, the analysis system may implement corrective measures. For example, the analysis system may remove contaminated samples from use in training the cancer classification model, thereby improving the accuracy and precision of the model. More specifically, the analysis system may perform contamination detection on the initial set of training samples to identify any contaminated samples and remove them from the set. The filtered set of training samples (free of contamination) can then be used to train the model. As another example, the analysis system may subtract contaminated test samples from the cancer classification, thereby reducing the false positive rate of cancer predictions.
[0099] When contaminated samples are detected, further corrective measures may be taken to address the contamination. Other corrective measures may include identifying and eliminating the source of contamination by conducting tests under various sample preparation conditions to accurately identify the conditions that contribute to the contamination. Eliminating the source of contamination improves the physical assay process and sample handling. When identifying the source of contamination, measures are taken to remove the source or minimize contamination in the sample handling workflow. Another corrective measure may be to physically discard the contaminated sample or, optionally, to collect a new sample from the individual.
[0100] The training of machine learning models described herein (e.g., cancer classifiers, neural network structures, and other models referred to herein) involves the performance of one or more non-mathematical tasks or the realization of at least a partial non-mathematical function by a machine or computing system, including, but not limited to, data loading, data storage, data toggling or modification, modification of non-temporary computer-readable storage media, metadata removal or data cleansing, data compression, protein structure modification, image modification, noise application, noise reduction, etc. Therefore, the training of machine learning models described herein may be based on or include mathematical concepts, but is not limited to mathematical calculations, the performance of mathematical tasks, or the act of calculating variables or numbers using mathematical methods.
[0101] Similarly, it should be noted that training these models described herein is practically impossible to do in the human mind. The models are inherently complex, involving one or more complex functions and a vast number of associated weights and parameters. Training and / or deploying such models requires so much work that it cannot be feasibly done by human intellect alone, or even with the aid of pen and paper. In such embodiments, the work can reach hundreds, thousands, tens of thousands, hundreds of thousands, millions, billions, or trillions. Furthermore, the training data can contain hundreds, thousands, tens of thousands, hundreds of thousands, millions, billions, or trillions of sequence reads, each sequence read potentially containing hundreds or thousands of nucleotides. Therefore, such models are inevitably anchored in computer technology for their implementation and use.
[0102] Overview of the IA cancer classification workflow Figure 1 is an exemplary flowchart illustrating an overall workflow 100 for cancer classification of a sample, according to one or more embodiments. Workflow 100 is performed by one or more entities, such as healthcare professionals, sequencing devices, and analysis systems. The purpose of the workflow includes cancer detection and / or monitoring in an individual. From a medical perspective, workflow 100 can complement other existing cancer diagnostic tools. Workflow 100 may help provide early cancer detection and / or routine cancer monitoring to appropriately inform treatment plans for individuals diagnosed with cancer. The overall workflow 100 may include more / fewer steps than those shown in Figure 1.
[0103] A healthcare professional performs sample collection 110. The individual undergoing cancer classification visits the healthcare professional. The healthcare professional collects a sample for cancer classification. Examples of biological samples, but not limited to, include tissue biopsy, blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or ascites. The sample contains genetic material belonging to the individual, which can be extracted and sequenced for cancer classification. After the sample is collected, it is fed into a sequencing device. Along with the sample, the healthcare professional may collect other information about the individual, such as biological sex, age, ethnicity, smoking status, and past diagnostic history.
[0104] A sequencing device performs sample sequencing 120. A laboratory clinician may perform one or more processing steps on the sample during sequencing preparation. Once preparation is complete, the clinician loads the sample into the sequencing device. An example of a device used for sequencing is described further in conjunction with Figures 10A and 10B. Sequencing devices generally extract and separate nucleic acid fragments and subject them to sequencing to determine the sequences of nucleic acid bases corresponding to these fragments. Sample sequencing includes sample processing during sequencing preparation of fragments in the sample. Sample processing may include one or more ligation steps and amplification of the nucleic acid material. In one or more embodiments, sample processing includes ligation of a sample barcode and a unique molecular identifier that can be used for contamination detection. The sample barcode is a substantially unique polynucleotide sequence for each sample. The sample barcode is ligated to each fragment in the sample before indexing and sequencing. A unique molecular identifier is, for example, a polynucleotide sequence ligated to each sample-derived fragment before amplification. The unique molecular identifier may be used for deduplication of sequence reads to identify unique fragments from the sample. A detailed description of sample sequencing 120 is shown in Figures 2-7.
[0105] Various sequencing methods include Sanger sequencing, fragment analysis, and next-generation sequencing. Sequencing can be either whole-genome sequencing or targeted sequencing using a targeted panel. In relation to DNA methylation, bisulfite sequencing (as further described, e.g., in Figures 2A and 2B) can determine the methylation state by bisulfite conversion of unmethylated cytosine at CpG sites. Sample sequencing 120 yields sequences of multiple nucleic acid fragments in the sample. In one or more embodiments, the sequences may include methylation state vectors, each methylation state vector representing the methylation state of CpG sites on the fragment.
[0106] The analysis system performs a pre-analysis process 130. An example of the analysis system is illustrated in Figure 10B. The pre-analysis process 130 includes, but is not limited to, demultiplexing, deduplicate removal of sequence reads, determination of coverage metrics, identification of contamination events, determination of whether the sample is contaminated, corrective actions for contamination events, detection of sequencing errors, and implementation of corrective actions. As a result of the pre-analysis process 130, the analysis system collects one set of sequence reads related to the sample that can be used for analysis 140.
[0107] The analysis system performs one or more analyses 140. These analyses are either statistical analyses or apply one or more trained models to predict the cancer status of at least the individual from which the sample originates. Various genetic features, such as CpG site methylation, single nucleotide polymorphisms (SNPs), insertions or deletions (indels), and other types of gene mutations, may be evaluated and considered. Contamination detection as an example of one analysis is further described in Figures 3A-7. In relation to methylation, analysis 140 may include informative methylation identification 142 (e.g., further described in Figures 8A and 8B), feature extraction 144 (e.g., further described in Figures 9A and 9B), and application of a cancer classifier 146 to determine cancer prediction (e.g., further described in Figures 9A and 9B). In one or more embodiments of feature extraction, the analysis system may utilize one or more age covariate predictive models to generate one or more age covariate residuals as features of the cancer classification. The cancer classifier 146 takes the extracted features as input and determines the cancer prediction. The cancer prediction can be a label or a value. A label indicates a specific cancer state; for example, a binary label may indicate the presence or absence of cancer, and a multiclass label may indicate one or more cancer types from a group of cancer types being screened. The value may indicate the likelihood of a specific cancer state, for example, the likelihood of cancer and / or the likelihood of a specific cancer type.
[0108] The analysis system returns a prediction of 150 to the healthcare provider. The healthcare provider can then establish or adjust the treatment plan based on the cancer prediction. Treatment optimization is further described in section IVC. Therapy.
[0109] Overview of IB Methylation As described herein, individual-derived cfDNA fragments are processed, for example, by converting unmethylated cytosine to uracil, and then sequenced. The sequenced reads are then compared to a reference genome to identify the methylation status at specific CpG sites within the DNA fragment. Each CpG site may be either methylated or unmethylated. Identifying informatively methylated fragments compared to healthy individuals can provide insights into the state of the target cancer. As is well known, abnormal DNA methylation (compared to a healthy control group) can cause various effects, which may contribute to cancer. Identifying informatively methylated cfDNA fragments presents several challenges. Firstly, determining whether a DNA fragment is informally methylated requires comparison with a control group. Therefore, if the number of control groups is small, the reliability of the determination decreases due to statistical variability within the small control group. Furthermore, differences in methylation status may occur among individuals in the control group, and it can be difficult to account for these differences when determining whether the target DNA fragment is informally methylated. Cytosine methylation at CpG sites may, in turn, influence subsequent methylation at other CpG sites. Simplifying this dependency can itself become another challenge.
[0110] Methylation typically occurs in deoxyribonucleic acid (DNA) when a hydrogen atom on the pyrimidine ring of a cytosine base is replaced by a methyl group to form 5-methylcytosine. In particular, methylation can occur at cytosine-guanine dinucleotides, which are referred to herein as “CpG sites.” In other cases, methylation may occur at cytosine or other nucleotides that do not belong to a CpG site, but this is rare. For the sake of clarity, methylation in this disclosure is described in relation to CpG sites. Informative DNA methylation is identified as hypermethylation or hypomethylation, both of which may indicate cancerous conditions. Throughout this disclosure, hypermethylation and hypomethylation can be characterized for a DNA fragment if it contains a number of CpG sites exceeding a threshold, and the proportion of those CpG sites exceeding the threshold is methylated or unmethylated.
[0111] The principles described herein are equally applicable to the detection of non-CpG-related methylation, including non-cytosine methylation. In such embodiments, the wet lab assay used to detect methylation may differ from that described herein. Furthermore, the methylation state vector described herein may generally include elements that are methylated or unmethylated sites (even if these sites are not specifically CpG sites). The remainder of the process described herein, including substitutions, is the same, and therefore the concepts of the invention described herein are applicable to other forms of methylation.
[0112] IC definition The term "cell-free nucleic acid" or "cfNA" refers to nucleic acid fragments circulating in an individual's body (e.g., blood), which originate from one or more healthy cells and / or one or more unhealthy cells (e.g., cancer cells). The term "cell-free DNA" or "cfDNA" refers to deoxyribonucleic acid fragments circulating in an individual's body (e.g., blood). Furthermore, cfNA or cfDNA in an individual's body may also originate from other non-human sources.
[0113] The terms “genomic nucleic acid,” “genomic DNA,” or “gDNA” refer to nucleic acid molecules or deoxyribonucleic acid molecules obtained from one or more cells. In various embodiments, gDNA can be extracted from healthy cells (e.g., non-tumor cells) or tumor cells (e.g., biopsy samples). In some embodiments, gDNA can be extracted from cells of blood cell lineage, such as leukocytes.
[0114] The term “circulating tumor DNA” or “ctDNA” refers to nucleic acid fragments derived from tumor cells or other types of cancer cells, which may be released into the body fluids of an individual (e.g., blood, sweat, urine, or saliva) as a result of biological processes such as apoptosis or necrosis of dead cells, or may be actively released by living tumor cells.
[0115] The terms "DNA fragment," "fragment," or "DNA molecule" generally refer to deoxyribonucleic acid (DNA) fragments, such as cfDNA, gDNA, and ctDNA.
[0116] The terms "NA fragment" or "NA molecule" generally refer to any nucleic acid molecule, including DNA molecules and ribonucleic acid (RNA) molecules.
[0117] The term "amplicon" can generally refer to nucleic acid molecules resulting from amplification processes, i.e., molecules derived from samples taken from molecules individually and / or synthetically produced as copies of the original molecules.
[0118] The term "sample barcode" can generally refer to a nucleotide sequence that is assigned to a sample and ligated onto a sequence read for the purpose of accurately assigning the sequence read belonging to the sample.
[0119] The term “molecular identifier” or “MI” can generally mean a nucleotide sequence ligated onto an original NA molecule derived from a sample for the purpose of identifying a distinct original NA molecule. The term “unique molecular identifier” or “UMI” can generally mean a molecular identifier that is substantially unique to other UMIs.
[0120] The term "contamination event" generally refers to an instance of contamination in a sample. Examples of contamination events, though not limited to them, include single-index hopping events, double-index hopping events, sequencing errors, and sample replacement events. A "single-index hopping event" may refer to an instance where a single index differs from what was expected. A "double-index hopping event" may refer to an instance where both indices differ from what was expected. A "sequencing error" may refer to an instance where one or more nucleotides differ in one amplicon compared to another, due to errors in the sequencing process, which may include errors introduced during ligation, amplification, etc. A "sample mix-up event" may refer to an instance where one sample is mistakenly labeled as belonging to another sample.
[0121] The terms “informative fragment,” “informative methylated fragment,” or “fragment with an informative methylation pattern” refer to a fragment that exhibits abnormal methylation of CpG sites compared to the methylation status of a population of healthy subjects. Abnormal methylation of a fragment may be determined using a probabilistic model to identify the unexpectedness of the observed methylation pattern of the fragment in a control group of healthy subjects.
[0122] The terms “UFXM” or “UfxM” refer to hypomethylated or hypermethylated fragments. Hypomethylated and hypermethylated fragments refer to fragments having at least a certain number (e.g., 5) of CpG sites that exceed some threshold percentage (e.g., 90%) of methylation or demethylation, respectively.
[0123] The term "informative score" refers to a score for a CpG site based on the number of informative fragments (or UFXMs in some embodiments) from the sample that overlap with that CpG site. The informative score is used in connection with the featureization of the sample for classification.
[0124] In this specification, the terms “about” or “approximately” may mean within a tolerance range of a particular value as determined by those skilled in the art, but the range may vary in part depending on the method of measuring or determining the value, for example, the limits of the measuring system. For example, “about” may mean within one or more standard deviations, according to the practice of the art. “About” may mean within ±20%, ±10%, ±5%, or ±1% of a given value. The terms “about” or “approximately” may mean within one decimal place, five times, or twice a value. Where a particular value is described in this application and claims, unless otherwise specified, the term “about” should be understood to mean within a tolerance range of that particular value. The term “about” may have a meaning that is generally understood by those skilled in the art. The term “about” may mean ±10%. The term “about” may mean ±5%.
[0125] In this specification, the terms “biological sample,” “patient sample,” or “sample” refer to any sample taken from a subject, which may represent the subject’s biological state and may include cell-free DNA. Examples of biological samples include, but are not limited to, the subject’s blood, whole blood, plasma, serum, urine, cerebrospinal fluid, feces, saliva, sweat, tears, pleural fluid, pericardial fluid, or ascites. Biological samples may include tissues or materials obtained from a living or deceased subject. Biological samples may be cell-free samples. Biological samples may include nucleic acids (e.g., DNA or RNA) or fragments thereof. The term “nucleic acid” may mean deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or hybrids or fragments thereof. Nucleic acids in a sample may be cell-free nucleic acids. Samples may be liquid samples or solid samples (e.g., cell or tissue samples). Biological samples may include blood, plasma, serum, urine, vaginal secretions, fluid from hydrops (e.g., testicular), vaginal lavage fluid, pleural fluid, ascites, cerebrospinal fluid, saliva, sweat, tears, sputum, bronchoalveolar lavage fluid, nipple secretions, and aspirates from various parts of the body (e.g., thyroid gland, breast). Biological samples may also include fecal samples. In various embodiments, the majority of the DNA in a biological sample enriched with cell-free DNA (e.g., plasma samples obtained by a centrifugation protocol) may be cell-free (e.g., 50%, 60%, 70%, 80%, 90%, 95%, or more than 99% of the DNA may be cell-free). Biological samples can be subjected to processing (e.g., centrifugation and / or cell lysis) to physically destroy tissues or cellular structures, thereby releasing intracellular components into a solution further containing enzymes, buffers, salts, surfactants, etc., which can then be used to prepare samples for analysis.
[0126] In this specification, the terms “control,” “control sample,” “reference,” “reference sample,” “normal,” and “normal sample” mean samples from subjects that are disease-free or otherwise healthy. For example, the methods disclosed herein may be performed on a subject with a tumor, and the reference sample would be a sample taken from healthy tissue of that subject. Reference samples may be obtained from subjects or databases. A reference may be, for example, a reference genome used to map nucleic acid fragment sequences obtained by sequencing of a sample of a subject. A reference genome may refer to a haploid or diploid genome for aligning and comparing nucleic acid fragment sequences from biological samples and constituent samples. An example of a constituent sample may be DNA from leukocytes obtained from a subject. In the case of a haploid genome, there may be only one nucleotide at each locus. In the case of a diploid genome, heterozygous loci can be identified, and each heterozygous locus may have two alleles, but either allele may enable matching in alignment with the locus.
[0127] As used herein, “cancer” or “tumor” means a mass of abnormal tissue whose growth exceeds and is not in harmony with the growth of normal tissue.
[0128] As used herein, the term “healthy” refers to an object in good health. A healthy object may present the absence of malignant or non-malignant disease. A “healthy individual” may have other diseases or conditions unrelated to the disease being examined that would not normally be considered “healthy.”
[0129] As used herein, the term “methylation” refers to the modification of deoxyribonucleic acid (DNA) in which a hydrogen atom on the pyrimidine ring of a cytosine base is converted to a methyl group, forming 5-methylcytosine. In particular, methylation tends to occur at cytosine and guanine dinucleotides, which are referred to herein as “CpG sites.” In other cases, methylation may occur at cytosine or other nucleotides that do not belong to CpG sites, but such methylation is rare. Informative cfDNA methylation can be identified as hypermethylation or hypomethylation, both of which may indicate cancerous conditions. Abnormal DNA methylation (compared to a healthy control group) can cause a variety of effects, which may contribute to cancer. The principles described herein apply equally to the detection of CpG-related and non-CpG-related methylation, including methylation of non-cytosine nucleotides. Furthermore, methylation state vectors may generally include elements that are vectors of sites where methylation has occurred or not (even if such sites are not specifically CpG sites).
[0130] Where used interchangeably herein, the terms “methylated fragment” or “nucleic acid methylated fragment” refer to a sequence of methylation states of each CpG site within a plurality of CpG sites, determined by methylation sequencing of nucleic acids (e.g., nucleic acid molecules and / or nucleic acid fragments). In a methylated fragment, the location and methylation state of each CpG site within the nucleic acid fragment are determined based on the alignment of the sequence read (e.g., obtained from nucleic acid sequencing) with a reference genome. A nucleic acid methylated fragment includes the methylation states of each CpG site within a plurality of CpG sites (e.g., a methylation state vector), thereby specifying the location of the nucleic acid fragment in the reference genome (e.g., indicated by the location of the first CpG site within the nucleic acid fragment using a CpG index or similar indicator) and the number of CpG sites within the nucleic acid fragment. Alignment of sequence reads with a reference genome based on methylation sequencing of nucleic acid molecules can be performed using a CpG index. In this specification, the term “CpG index” refers to a list of each of several CpG sites (e.g., CpG 1, CpG 2, CpG 3, etc.) within a reference genome (e.g., the human reference genome), which may be in electronic format. The CpG index further includes, for each CpG site in the CpG index, the corresponding genomic location within the corresponding reference genome. Thus, each CpG site within each nucleic acid methylation fragment is indexed to a specific location within its respective reference genome, which can be determined using the CpG index.
[0131] As used herein, the term “true positive” (TP) refers to a subject having a disease. A “true positive” may refer to a subject having a tumor, cancer, a precancerous condition (e.g., a precancerous lesion), a localized or metastatic cancer, or a non-malignant disease. A “true positive” refers to a subject having a disease, which is identified as having the disease by the assay or method of this disclosure. As used herein, the term “true negative” (TN) refers to a subject that is free from disease or does not have a detectable disease. A “true negative” may refer to a subject that is free from disease or a detectable disease, such as a tumor, cancer, a precancerous condition (e.g., a precancerous lesion), a localized or metastatic cancer, a non-malignant disease, or a otherwise healthy subject. A “true negative” refers to a subject that is free from disease or does not have a detectable disease, or a subject identified as not having the disease by the assay or method of this disclosure.
[0132] As used herein, the term “reference genome” refers to a known, sequenced or characterized genome (whether partial or complete) of any organism or virus, which can be used to reference identified sequences from a subject. Exemplary reference genomes used for human subjects and many other species are available in online genome browsers hosted by the National Center for Biotechnology Information (NCBI) or the University of California, Santa Cruz (UCSC). “Genome” refers to the complete genetic information of an organism or virus represented by nucleic acid sequences. As used herein, a reference sequence or reference genome refers to an assembled or partially assembled genome sequence from one or more individuals. In some embodiments, a reference genome refers to an assembled or partially assembled genome sequence from one or more individuals. A reference genome can be thought of as a representative example of a certain set of genes. In some embodiments, a reference genome includes sequences assigned to chromosomes. Exemplary human reference genomes include, but are not limited to, NCBI build 34 (equivalent to UCSC:hg16), NCBI build 35 (equivalent to UCSC:hg17), NCBI build 36.1 (equivalent to UCSC:hg18), GRCh37 (equivalent to UCSC:hg19), and GRCh38 (equivalent to UCSC:hg38).
[0133] As used herein, “sequence read” or “read” means a nucleotide sequence produced by any sequencing process described herein or known to those skilled in the art. Reads may be produced from one end of a nucleic acid fragment (“single-ended read”) and, optionally, from both ends of a nucleotide (e.g., paired-ended read, double-ended read). In some embodiments, sequence reads (e.g., single-ended or paired-ended reads) may be produced from one or both strands of a target nucleic acid fragment. The length of a sequence read is often related to the particular sequencing technique. High-throughput methods may yield sequence reads that vary in size from, for example, tens to hundreds of base pairs (bp). In some embodiments, the sequence reads have an average, median, or mean length of approximately 15 bp to 900 bp (for example, approximately 20 bp, 25 bp, 30 bp, 35 bp, 40 bp, 45 bp, 50 bp, 55 bp, 60 bp, 65 bp, 70 bp, 75 bp, 80 bp, 85 bp, 90 bp, 95 bp, 100 bp, 110 bp, 120 bp, 130, 140 bp, 150 bp, 200 bp, 450 bp, 300 bp, 350 bp, 400 bp, 450 bp, or approximately 500 bp). In some embodiments, the sequence reads have an average, median, or mean length of approximately 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or more. For example, nanopore sequencing can provide sequence reads that can vary in size from tens to hundreds or thousands of base pairs. Illumina parallel sequencing yields sequence reads with less size variation; for example, most sequence reads may be less than 200 bp. A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g., a string of nucleotides). For example, a sequence read may correspond to a string of nucleotides consisting of a portion of a nucleic acid fragment (e.g., about 20 to about 150), to a string of nucleotides at one or both ends of a nucleic acid fragment, or to nucleotides in the entire nucleic acid fragment.Sequence reads can be obtained in various ways, for example, using sequencing techniques, or amplification techniques such as probes or capture probes in hybridization arrays, polymerase chain reaction (PCR), or linear or isothermal amplification using a single primer.
[0134] As used herein, “sequencing” and similar terms generally refer to any biochemical process that can be used to determine the order of biological macromolecules such as nucleic acids or proteins. For example, sequencing data may include all or some of the nucleotide bases in a nucleic acid molecule, such as a DNA fragment.
[0135] As used herein, the term “sequencing depth” is used interchangeably with the term “coverage” and refers to the number of times a locus is covered by census sequence reads corresponding to unique nucleic acid target molecules aligned with the locus. For example, sequencing depth is equal to the number of unique nucleic acid target molecules that cover the locus. A locus can be as small as a single nucleotide, or as large as a chromosome arm or an entire genome. Sequencing depth can be expressed as “Y×”, for example, 50×, 100×, etc., where “Y” refers to the number of times a particular locus is covered by sequences corresponding to nucleic acid targets, for example, the number of times independent sequence information covering that particular locus is obtained. In some embodiments, sequencing depth refers to the number of genomes that have been sequenced. Sequencing depth may apply to multiple loci or an entire genome, where “Y” refers to the average number of times a locus, haploid genome, or whole genome has been sequenced, respectively. When average depth is cited, the actual depth of different loci included in the dataset may range within a specific range. Ultradeep sequencing can refer to a sequencing depth of at least 100x at a given gene locus.
[0136] As used herein, the term “bag” refers to a method of grouping sequence reads together. For example, in demultiplexing, bags may be used to separate sequence reads belonging to a particular sample. As another example, in deduplication, bags may be used to identify sequence reads related to amplicons of the same original DNA fragment in a sample.
[0137] As used herein, the terms “sensitivity” or “true positive rate” (TPR) refer to the number of true positives divided by the sum of the number of true positives and the number of false negatives. Sensitivity can represent the ability of an assay or method to correctly determine the proportion of a population that truly has the disease. For example, sensitivity can represent the ability of a method to correctly determine the number of subjects in a population that have cancer. In another example, sensitivity can represent the ability of a method to correctly identify one or more markers indicating cancer.
[0138] As used herein, the terms “specificity” or “true negative rate” (TNR) refer to the number of true negatives divided by the sum of the number of true negatives and the number of false positives. Specificity can represent the ability of an assay or method to correctly determine the proportion of a population that does not truly have the disease. For example, specificity can represent the ability of a method to correctly determine the number of subjects in a population that do not have cancer. In another example, specificity can represent the ability of a method to correctly identify one or more markers indicating cancer.
[0139] As used herein, the term “Subject” means any living or non-living thing, including but not limited to humans (e.g., males, females, fetuses, pregnant women, children, etc.), non-human animals, plants, bacteria, fungi, or protists. Humans or non-human animals may be used as subjects, but are not limited to mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, cattle (e.g., cattle), equids (e.g., horses), goats and sheep (e.g., sheep, goats), pigs (e.g., pigs), camelids (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), bears (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. In some embodiments, the subject is a male or female at any stage (e.g., male, female, or child). The subjects from whom samples are taken or treated by any of the methods or compositions described herein may be of any age and may be adults, infants, or children.
[0140] As used herein, the term “tissue” may refer to a group of cells that function together as a single unit. A single tissue may contain more than one type of cells. Various types of tissue may consist of various types of cells (e.g., hepatocytes, alveolar cells, or blood cells), or they may be tissues from different organisms (mother versus fetus) or healthy cells versus tumor cells. The term “tissue” may generally refer to any group of cells present in the human body (e.g., cardiac tissue, lung tissue, kidney tissue, nasopharyngeal tissue, oropharyngeal tissue). In some embodiments, the terms “tissue” or “tissue type” may be used to refer to the tissue from which cell-free nucleic acids originate. In one example, viral nucleic acid fragments may originate from blood tissue. In another example, viral nucleic acid fragments may originate from tumor tissue.
[0141] As used herein, the term “genome” refers to the characteristics of the genome of an organism. Examples of genomic characteristics include, but are not limited to,: characteristics related to major nucleic acid sequences of all or part of the genome (e.g., presence or absence of nucleotide polymorphisms, indels, sequence rearrangements, mutation frequency, etc.); copy number of one or more specific nucleotide sequences within the genome (e.g., copy number, allele frequency, ploidy of a single chromosome or the entire genome, etc.); epigenetic status of the entire genome or part thereof (e.g., covalent nucleic acid modifications such as methylation, histone modification, nucleosome arrangement, etc.); and the expression profile of the genome of an organism (e.g., gene expression levels, isotype expression levels, gene expression ratios, etc.).
[0142] The terms used herein are intended solely to describe specific embodiments and are not intended to limit them. Where used herein, the singular forms “a,” “an,” and “it” are intended to encompass the plural form unless the context clearly indicates otherwise. Furthermore, where “contains,” “includes,” “has,” “possesses,” “with,” or variations thereof are used in the detailed description and / or claims, these terms are intended to be as comprehensive as the term “contains.”
[0143] Example of an ID analysis system Figure 10A is an exemplary flowchart of a device for sequencing nucleic acid samples according to one or more embodiments. This exemplary flowchart includes devices such as a sequencer 1020 and an analysis system 1000. The sequencer 1020 and the analysis system 1000 can work together to perform one or more steps in process 300 in Figure 3A, process 400 in Figure 4A, process 420 in Figure 4B, and other processes described herein.
[0144] In various embodiments, the sequencer 1020 receives a concentrated nucleic acid sample 1010. As shown in Figure 10A, the sequencer 1020 may include a graphical user interface 1025 that enables interaction between the user and a specific task (e.g., starting or ending sequencing), as well as one or more loading stations 1030 for loading a sequencing cartridge containing the concentrated fragment sample and / or buffers necessary to perform the sequencing assay. Thus, once the user of the sequencer 1020 has supplied the necessary reagents and sequencing cartridge to the loading station 1030 of the sequencer 1020, the user can start sequencing by interacting with the graphical user interface 1025 of the sequencer 1020. Once started, the sequencer 1020 performs sequencing and outputs sequence reads of the concentrated fragment from the nucleic acid sample 1010.
[0145] In some embodiments, the sequencer 1020 is communicatively connected to the analysis system 1000. The analysis system 1000 includes multiple computing devices used for processing sequence reads for various applications, such as evaluation of methylation status at one or more CpG sites, variant calling, or quality control. The sequencer 1020 can provide sequence read data to the analysis system 1000 in BAM file format. The analysis system 1000 is communicatively connected to the sequencer 1020 by wireless, wired, or a combination of wireless and wired communication technologies. Generally, the analysis system 1000 comprises a processor and a non-temporary computer-readable storage medium for storing computer instructions, which, when executed by the processor, cause the processor to process sequence reads or to perform one or more steps of any of the methods or processes disclosed herein.
[0146] In some embodiments, the sequence read is aligned with a reference genome using methods known in the art to determine alignment position information, for example, via step 340 of process 300 in Figure 3A. The alignment position generally describes the start and end positions of a region in the reference genome, which correspond to the start and end nucleotides of a given sequence read. Corresponding to methylation sequencing, the alignment position information can be summarized to indicate the first and last CpG sites contained in the sequence read, according to the alignment with the reference genome. The alignment position information can also further indicate the methylation status and location of all CpG sites within a given sequence read. The region in the reference genome can be associated with a gene or a segment of a gene, so that the analysis system 1000 can label the sequence read with one or more genes to align with the sequence read. In one embodiment, the length (or size) of the fragment is determined from the start and end positions.
[0147] In various embodiments, for example, when a paired-end sequencing process is used, the sequenced reads consist of a read pair designated as R_1 and R_2. For example, the first read R_1 may be sequenced from the first end of a double-stranded DNA (dsDNA) molecule, while the second read R_2 may be sequenced from the second end of the double-stranded DNA (dsDNA) molecule. Thus, the nucleotide base pairs of the first read R_1 and the second read R_2 can be aligned consistently (for example, in the reverse direction) with the nucleotide bases of the reference genome. The alignment position information obtained from the read pair R_1 and R_2 may include the start position of the reference genome (corresponding to the end of the first read (e.g., R_1)) and the end position of the reference genome (corresponding to the end of the second read (e.g., R_2)). That is, the start and end positions in the reference genome may indicate the positions in the reference genome to which nucleic acid fragments are likely to correspond. It can generate output files in SAM (Sequence Alignment Map) format or BAM (Binary) format, which can then be output for further analysis.
[0148] Referring to Figure 10B, Figure 10B is a block diagram of an analysis system 1000 for processing a DNA sample according to one embodiment. The analysis system implements one or more computing devices used for analyzing the DNA sample. The analysis system 1000 includes a sequence processor 1040, a sequence database 1045, a model database 1055, a model 1050, a parameter database 1065, and a score engine 1060. In some embodiments, the analysis system 1000 performs some or all of process 300 in Figure 3A and process 400 in Figure 4A.
[0149] The sequencing processor 1040 generates methylation state vectors for fragments derived from the sample. For each CpG site in a fragment, the sequencing processor 1040 generates a methylation state vector for each fragment via process 300 (Figure 3A) that specifies the fragment's location in the reference genome, the number of CpG sites in the fragment, and the methylation state of each CpG site in the fragment (methylated, unmethylated, or indeterminate). The sequencing processor 1040 stores the methylation state vectors of the fragments in the sequencing database 1045. The data in the sequencing database 1045 is organized so that the methylation state vectors from the sample are correlated with each other.
[0150] Furthermore, multiple different models 1050 can be stored in the model database 1055 or retrieved for use with test samples. For example, a model is a trained cancer classifier for determining cancer predictions in test samples using feature vectors derived from informative fragments. The training and use of cancer classifiers are discussed further in conjunction with Section III, Cancer Classifiers for Cancer Determination. The analysis system 1000 trains one or more models 1050 and stores the various trained parameters in the parameter database 1065. The analysis system 1000 also stores the models 1050 along with their functions in the model database 1055.
[0151] During estimation, the score engine 1060 uses one or more models 1050 to return the output. The score engine 1060 accesses the models 1050 in the model database 1055 along with the trained parameters from the parameter database 1065. Depending on the model, the score engine receives inputs appropriate for the model and calculates the output based on the received inputs, parameters, and the function of each model that relates the inputs to the output. In some use cases, the score engine 1060 further calculates metrics that correlate with the confidence of the calculated output from the model. In other use cases, the score engine 1060 calculates other intermediate values for use with the model.
[0152] II. Sequencing and Processing of Samples II.A. Generation of Methylation State Vectors for DNA Fragments Figure 2A is an illustrative flowchart illustrating a process 200 for sequencing a sample containing nucleic acid (NA) fragments according to one or more embodiments. To analyze NA methylation, the analysis system first obtains a sample from an individual containing multiple NA molecules 205. Process 200 can be applied to sequencing various types of NA molecules, such as DNA molecules, RNA molecules, extracellular DNA molecules, circulating tumor DNA molecules, tissue DNA molecules, and other types of NA molecules. Process 200 is one embodiment of the sample sequencing 120 in Figure 1.
[0153] Generally, the sample sequencing process 200 includes at least three steps. The analysis system obtains a sample from an object containing NA molecules 205 and isolates the NA molecules. The sample can be any type of biological sample containing NA molecules derived from an individual. For example, the sample may be a blood sample, a urine sample, a tissue sample, or another type of biological sample. The analysis system prepares a sequencing library 215 for preparing NA molecules for sequencing. The preparation of the sequencing library may include one or more ligation steps to add additional molecules to be used for sequencing the NA molecules, amplification of the NA molecules to produce amplified molecules to ensure capture and sequencing of all NA molecules in the sample, enrichment of the sample by targeting a specific genomic region using a targeting probe, and ligation of one or more indices and one or more adapters to the NA molecules. The analysis system then sequences the NA molecules 225 and obtains sequence reads. The sequence reads may include forward reads and reverse reads.
[0154] In one or more embodiments of methylation sequencing, the analysis system obtains a sample containing DNA fragments (e.g., cfDNA) and separates each DNA fragment. The DNA fragments can be treated before sequencing to convert unmethylated cytosine to uracil. In one embodiment, this method uses bisulfite treatment of the DNA to convert unmethylated cytosine to uracil without converting methylated cytosine. Commercially available kits such as EZ DNA Methylation®-Gold, EZ DNA Methylation®-Direct, or EZ DNA Methylation®-Lightning kits (available from Zymo Research Corp (Irvine, CA)) can be used for bisulfite conversion. In another embodiment, an enzymatic reaction is used to achieve the conversion of unmethylated cytosine to uracil. For example, commercially available kits such as APOBEC-Seq (NEBiolabs, Ipswich, MA) can be used for the conversion of unmethylated cytosine to uracil.
[0155] A sequencing library can be prepared from the converted DNA fragments.215 During library preparation, adapter ligation can add a unique molecular identifier (UMI) to nucleic acid molecules (e.g., DNA molecules). The UMI is a short nucleic acid sequence (e.g., 4-10 base pairs) added to the end of a DNA fragment (e.g., a DNA molecule fragmented by physical cutting, enzymatic digestion, and / or chemical fragmentation) during adapter ligation. The UMI can be a degenerate base pair that functions as a unique tag that can be used to identify sequence reads derived from a particular DNA fragment. During PCR amplification after adapter ligation, the UMI can be replicated together with the ligated DNA fragment. This provides a method for identifying sequence reads derived from the same initial fragment in downstream analysis.
[0156] Optionally, the sequencing library can also be enriched for DNA fragments or genomic regions indicating cancerous status using multiple hybridization probes. Hybridization probes are short oligonucleotides that can hybridize to specifically designated DNA fragments or target regions and have the ability to enrich those fragments or regions for subsequent sequencing and analysis. Hybridization probes can be used to target, high-depth analysis of a designated set of CpG sites of interest to the researcher. Hybridization probes can be tiled across one or more target sequences with coverage of 1×, 2×, 3×, 4×, 5×, 6×, 7×, 8×, 9×, 10×, or more than 10×. For example, hybridization probes tiled with 2× coverage include probes that overlap so that each portion of the target sequence hybridizes with two independent probes. Hybridization probes can be tiled across one or more target sequences with coverage of less than 1×.
[0157] In one embodiment, a hybridization probe is designed to enrich DNA molecules that have been treated (e.g., with a bisulfite) to convert unmethylated cytosine to uracil. During enrichment, the hybridization probe (also referred to herein as the “probe”) is used to target and pull down nucleic acid fragments that provide information about the presence or absence of cancer (or disease), cancer status, or cancer classification (e.g., cancer class or tissue origin). The probe is designed to anneal to (or hybridize) a target (complementary) strand of DNA. The target strand may be a “plus” strand (e.g., the strand transcribed into mRNA and subsequently translated into protein) or a complementary “minus” strand. The probe may be in the length range of 10 base pairs, 100 base pairs, or 1,000 base pairs. The probe may be designed based on a methylation site panel. The probe may be designed based on a targeted gene panel to analyze specific mutations or targeted regions of the genome (e.g., human or other organism) suspected to correspond to a particular cancer or other disease. Furthermore, the probe can be designed to cover overlapping areas of the target region.
[0158] Once a sequencing library or a portion thereof is prepared, it can be subjected to sequencing to obtain multiple sequence reads. These sequence reads may be in a computer-readable digital format for processing and interpretation by computer software. The sequence reads are aligned to a reference genome to determine alignment position information. Alignment position information indicates the start and end positions of a region in the reference genome, which correspond to the start and end nucleotide bases of a given sequence read. Alignment position information may also include the length of the sequence read, which can be determined from the start and end positions. The region in the reference genome can be associated with a gene or a segment of a gene. The sequence reads may consist of read pairs denoted as R1 and R2. For example, the first read R1 may be sequenced from the first end of a nucleic acid fragment, while the second read R2 may be sequenced from the second end of the nucleic acid fragment. Thus, the nucleotide base pairs of the first read R1 and the second read R2 can be aligned consistently (e.g., in the reverse direction) with the nucleotide bases of the reference genome. The alignment position information derived from read pairs R1 and R2 may include a start position corresponding to the end of the first read (e.g., R1) in the reference genome and an end position corresponding to the end of the second read (e.g., R2) in the reference genome. That is, the start and end positions in the reference genome can indicate the possible corresponding positions of nucleic acid fragments within the reference genome. Output files in SAM (Sequence Alignment Map) or BAM (Binary) format can be generated and output for further analysis, such as determination of methylation status.
[0159] From the sequenced reads, the analysis system determines the location and methylation status of each CpG site based on alignment to the reference genome. For each fragment, the analysis system generates a methylation status vector indicating the location of the fragment in the reference genome (e.g., the location of the first CpG site in each fragment or a similar indicator), the number of CpG sites in the fragment, and the methylation status of each CpG site in the fragment (methylated (e.g., represented by M), unmethylated (e.g., represented by U), or uncertain (e.g., represented by I)).235 Observed states are methylated and unmethylated states, while unobserved states are uncertain. Uncertain methylation states may result from sequencing errors and / or mismatches in methylation status between complementary strands of DNA fragments. The methylation status vectors can be stored in temporary or persistent computer memory for later use and processing. Furthermore, the analysis system can remove duplicate reads or duplicate methylation status vectors from a single sample. The analysis system may determine that a particular fragment containing one or more CpG sites has an uncertain methylation state exceeding a threshold number or percentage, and may exclude such fragments or selectively include such fragments, but may construct a model that takes such uncertain methylation states into account. Figure 7 further illustrates process 200 in several methylation sequencing embodiments.
[0160] II.B. Methylation Sequencing Figure 2B is an illustrative diagram illustrating the methylation sequencing of a cfDNA molecule to obtain a methylation state vector according to one or more embodiments. For example, the analysis system receives a cfDNA molecule 242 containing three CpG sites in this example. As shown in the figure, the first and third CpG sites of cfDNA molecule 242 are methylated 244. During processing step 250, cfDNA molecule 242 is transformed to produce a transformed cfDNA molecule 252. During processing step 250, the second CpG site, which was not methylated, has its cytosine converted to uracil. However, the first and third CpG sites remain untransformed.
[0161] After conversion, a sequencing library is prepared, and the molecule is sequenced to generate sequence reads 262. The analysis system aligns sequence reads 262 with a reference genome 264. The reference genome 264 provides context indicating the location in the human genome from which the cfDNA fragment originates. In this simplified example, the analysis system aligns sequence reads 262 such that three CpG sites correspond to CpG sites 23, 24, and 25 (arbitrary reference identifiers used for convenience of explanation) 270. The analysis system can thus generate information on both the methylation status of all CpG sites on the cfDNA molecule 242 and the locations in the human genome to which these CpG sites map. As shown in the figure, methylated CpG sites on sequence read 262 are read as cytosine. In this example, cytosine appears only at the first and third CpG sites in sequence read 262, so it can be inferred that the first and third CpG sites in the original cfDNA molecule are methylated. In contrast, the second CpG site can be read as thymine (U is converted to T during the sequencing process), so it can be inferred that the second CpG site in the original cfDNA molecule is not methylated. Based on these two pieces of information (methylation state and location), the analysis system generates a methylation state vector 272 for fragment cfDNA 242. In this example, the obtained methylation state vector 272 is: <M 23 ,U 24 M 25 >Here, M corresponds to a methylated CpG site, U represents a non-methylated CpG site, and the subscript indicates the location of each CpG site in the reference genome.
[0162] One or more alternative sequencing methods can be used to obtain sequence reads from nucleic acids in biological samples. These methods may include, but are not limited to, any form of sequencing available for obtaining the number of sequence reads measured from nucleic acids (e.g., cell-free nucleic acids), including: the Roche 454 platform, the Applied Biosystems SOLID platform, Helicos True Single Molecule DNA sequencing technology, Affymetrix Inc.'s hybridization sequencing platform, Pacific Biosciences' single-molecule, real-time (SMRT) technology, 454 Life Sciences, Illumina / Solexa and Helicos Biosciences' synthesis sequencing platforms, and Applied Biosystems' ligation sequencing platform. Life Technologies' ION TORRENT technology and nanopore sequencing can also be used to obtain sequence reads from nucleic acids (e.g., cell-free nucleic acids) in biological samples. To form genotype datasets, sequence reads can be obtained from cell-free nucleic acids obtained from training biological samples using synthetic sequencing and reversible terminator-based sequencing (e.g., Illumina's Genome Analyzer; Genome Analyzer II; HISEQ 2000; HISEQ 4500 (Illumina, San Diego Calif.)). Millions of cell-free nucleic acid (e.g., DNA) fragments can be sequenced in parallel. An example of this type of sequencing technique uses a flow cell containing optically transparent slides with eight separate lanes, to which oligonucleotide anchors (e.g., adapter primers) are bound to the slide surface. Cell-free nucleic acid samples may contain signals or tags to facilitate detection.Acquisition of sequence reads from cell-free nucleic acids obtained from biological samples may include obtaining quantitative information of signals or tags using various techniques such as flow cytometry, quantitative polymerase chain reaction (qPCR), gel electrophoresis, gene chip analysis, microarrays, mass spectrometry, cytoflowmetry, fluorescence microscopy, confocal laser scanning microscopy, laser scanning cytometry, affinity chromatography, manual batch mode separation, field suspension, sequencing, and combinations thereof.
[0163] One or more sequencing methods may include whole-genome sequencing assays. Whole-genome sequencing assays may include physical assays that generate sequencing reads for the entire genome or a substantial portion of the entire genome and can be used to determine large variations such as copy number variability or copy number anomalies. Such physical assays may use whole-genome sequencing techniques or whole-exome sequencing techniques. Whole-genome sequencing assays may have an average sequencing depth of at least 1×, 2×, 3×, 4×, 5×, 6×, 7×, 8×, 9×, 10×, at least 20×, at least 30×, or at least 40× across the entire genome under test. In some embodiments, the sequencing depth is about 30,000 times. One or more sequencing methods may include targeted panel sequencing assays. A targeted panel sequencing assay may have an average sequencing depth of at least 50,000×, at least 55,000×, at least 60,000×, or at least 70,000× for the targeted gene panel. The targeted gene panel may contain 450 to 500 genes. The targeted gene panel may contain genes in the range of 500±5, 500±10, or 500±25.
[0164] One or more sequencing methods may include paired-end sequencing. One or more sequencing methods may generate multiple sequence reads. Multiple sequence reads may have average lengths in the ranges of 10–700, 50–400, or 100–300. One or more sequencing methods may include methylation sequencing assays. Methylation sequencing is i) whole-genome methylation sequencing or ii) targeted DNA methylation sequencing using multiple nucleic acid probes. For example, methylation sequencing is whole-genome bisulfite sequencing (e.g., WGBS). Methylation sequencing may be targeted DNA methylation sequencing using multiple nucleic acid probes targeting the most informative regions of the methylome, a unique methylation database, and conventional whole-genome and targeted sequencing assays.
[0165] Methylation sequencing can detect one or more 5-methylcytosine (5mC) and / or 5-hydroxymethylcytosine (5hmC) in each nucleic acid methylation fragment. Methylation sequencing may include a process of converting one or more unmethylated cytosines or one or more methylated cytosines in each nucleic acid methylation fragment to the corresponding one or more uracils. During methylation sequencing, one or more uracils may be detected as the corresponding one or more thymines. The conversion of one or more unmethylated cytosines or one or more methylated cytosines may include chemical conversion, enzymatic conversion, or a combination thereof.
[0166] For example, bisulfite conversion is the process of converting cytosine to uracil while leaving methylated cytosine (e.g., 5-methylcytosine or 5-mC) intact. In some DNAs, approximately 95% of the cytosine in the DNA may be unmethylated, and the resulting DNA fragments may contain a large amount of uracil, which is represented as thymine. Nucleic acids can be treated using enzymatic conversion processes before sequencing, and this can be carried out in various ways. An example of bisulfite-free conversion is bisulfite-free and base-resolution sequencing, TET-assisted pyridine borane sequencing (TAPS), for the non-destructive and direct detection of 5-methylcytosine and 5-hydroxymethylcytosine without affecting unmodified cytosine. In each nucleic acid methylation fragment, the methylation status of the corresponding CpG sites can be determined as methylated if the CpG site is determined to be methylated by methylation sequencing, and as unmethylated if the CpG site is determined to be unmethylated by methylation sequencing.
[0167] Methylation sequencing analysis (e.g., WGBS and / or targeted methylation sequencing) may have average sequencing depths of 1,000×, 2,000×, 3,000×, 5,000×, 10,000×, 15,000×, 20,000×, or 30,000×. Methylation sequencing may have sequencing depths exceeding 30,000×, e.g., at least 40,000× or 50,000×. Whole-genome bisulfite sequencing may have an average sequencing depth of 20× to 50×, and targeted methylation sequencing may have an average effective depth of 100× to 1,000×, where effective depth refers to whole-genome bisulfite sequencing coverage equivalent to obtaining the same number of sequence reads as obtained by targeted methylation sequencing.
[0168] For further details regarding methylation sequencing (e.g., WGBS and / or targeted methylation sequencing), see, for example, U.S. Patent Application No. 16 / 352,602, entitled “Methylation Fragment Anomaly Detection,” filed March 13, 2019, and U.S. Patent Application No. 16 / 719,902, entitled “Systems and Methods for Estimating Cell Source Fractions Using Methylation Information,” filed December 18, 2019 (both incorporated herein by reference). Fragment methylation patterns can be obtained using other methods of methylation sequencing, including those disclosed herein and / or modifications, substitutions, or combinations thereof. Methylation sequencing can be used to identify one or more methylation state vectors, for example, in accordance with the techniques disclosed in U.S. Patent Application No. 16 / 352,602, entitled “Anomalous Fragment Detection and Classification,” filed March 13, 2019, or in U.S. Patent Application No. 15 / 931,022, entitled “Model-Based Featuration and Classification,” filed May 13, 2020, both of which are incorporated herein by reference.
[0169] Multiple nucleic acid methylation fragments can be obtained using nucleic acid methylation sequencing and the resulting one or more methylation state vectors. Each corresponding set of multiple nucleic acid methylation fragments (e.g., for each genotype dataset) may contain more than 100 nucleic acid methylation fragments. The average number of nucleic acid methylation fragments across all corresponding sets of multiple nucleic acid methylation fragments may include more than 1,000, more than 5,000, more than 10,000, more than 20,000, or more than 30,000 nucleic acid methylation fragments. The average number of nucleic acid methylation fragments across all corresponding sets of multiple nucleic acid methylation fragments may range from 10,000 to 50,000. The corresponding multiple nucleic acid methylation fragments may include 1,000 or more, 10,000 or more, 100,000 or more, 1,000,000 or more, 10,000,000 or more, 500,000,000 or more, 1,000,000 or more, 2,000,000 or more, 3,000,000 or more, 4,000,000 or more, 5,000,000 or more, 6,000,000 or more, 7,000,000 or more, 8,000,000 or more, 9,000,000 or more, or 10,000,000 or more nucleic acid methylation fragments. The average length of the corresponding multiple nucleic acid methylation fragments may be between 140 and 480 nucleotides.
[0170] Further details regarding nucleic acid sequencing methods and methylation sequencing data are disclosed in U.S. Patent Application No. 17 / 191,914, filed on March 4, 2021, entitled "Systems and Methods for Cancer Condition Determination Using Autoencoders," which is incorporated herein by reference in its entirety.
[0171] II.C. Sample contamination detection using sample barcodes Figures 3A and 3B show flowcharts of one or more embodiments utilizing sample barcodes and unique molecular identifiers (UMIs) for detecting contamination events. In particular, Figure 3A shows an exemplary flowchart of DNA fragment isolation, sequencing library preparation, and DNA fragment sequencing, and Figure 3B shows an exemplary flowchart of detecting contamination events from sequence reads based on sample barcodes and UMIs. In one or more embodiments, a sequencer (e.g., sequencer 1020 in Figure 10) performs process 300 in Figure 3A, and an analysis system performs process 335 in Figure 3B. In other embodiments, any other combination of the devices shown in Figures 10A and 10B may perform the steps shown in Figures 3A and 3B. In additional embodiments, parts of the steps, such as ligation, amplification, indexing, etc., in the sequence library may be performed by a laboratory technician or scientist.
[0172] Figure 3A shows an exemplary process for sequencing DNA fragments using sample barcodes and UMI.300 The sequencer isolates DNA fragments from each of multiple samples.305 The sequencer (or clinician) may perform a series of steps on the DNA fragments. For example, DNA isolation may include lysing the sample, clarifying the lysate, binding it across the matrix to isolate the DNA fragments, and washing away other contaminants. Additional steps specific to each type of sample may be performed. For example, plasma separation may be used for blood samples.
[0173] The sequencer performs the first ligation of UMIs to the isolated DNA fragment.310 In this step, the DNA fragment present in the sample is all the original molecules derived from the individual. A set of UMIs is used. Each UMI is a polynucleotide sequence that is substantially unique to the other UMIs in the set. The length of a UMI can be 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 base pairs. In one embodiment, all UMIs in the set are the same length. In other embodiments, at least two UMIs are of different lengths. Substantial uniqueness indicates some threshold difference between two UMIs. The threshold difference can be at least two different nucleotides, at least three different nucleotides, etc. The difference between UMIs is based on various matrices, e.g., nucleotides, nucleotide lengths, nucleotide sequence diversity, etc. A UMI can be ligated to one of the two ends of a DNA fragment. For example, UMI may be ligated to the 3' end, or alternatively, to the 5' end. In one or more embodiments, a ligase is added to the reaction to facilitate initial ligation to the DNA fragment.
[0174] The number of UMIs in a UMI set can be adjusted to optimize the balance between mismatches and false-positive characteristic fragments. Mismatches occur when multiple fragments are joined into multiple copies of the same UMI. A false-positive characteristic fragment occurs when a sequencing error in the UMI causes the sequencer to consider two sequence reads as a characteristic fragment for one of two sequence reads of the same fragment. A smaller set of UMIs may reduce the rate of false-positive characteristic fragments but increase the mismatch rate. A larger set of UMIs may reduce the mismatch rate but increase the rate of false-positive characteristic fragments. The size of a UMI set can be 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100.
[0175] The sequencer performs amplification of DNA fragments with UMIs attached.315 The amplification results in multiple copies of each DNA fragment. These copies are synthetic, i.e., not derived from the sample, but are copies based on the DNA fragments in the sample. In one or more embodiments, amplification may be achieved by polymerase chain reaction (PCR) amplification, linear amplification, clonal amplification, isothermal amplification, loop-mediated isothermal amplification, nucleic acid sequence-based amplification, strand displacement amplification, rolling circle amplification, ligase chain reaction, branch extension amplification, and other types of amplification techniques. Each copy ("amplicon") generally contains the UMI sequence attached to each original DNA fragment (unless amplification errors during amplification affect the UMI sequence).
[0176] The sequencer performs a second ligation of a unique sample barcode to all fragments in a single sample.320 As described, the amplicons likely contain the UMI sequence added in the first ligation to the original DNA fragment. The second ligation involves ligating a unique sample barcode to all amplicons belonging to the sample. The second ligation, performed after amplification, reduces the probability of amplification errors, affecting the sample barcode sequence added to the amplicons. The sample barcode is a polynucleotide sequence. The sample barcode may have a length of 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 base pairs. All amplicons in a single sample are ligated with the same sample barcode. Each sample barcode is substantially unique to other sample barcodes. Substantially unique means there is some threshold difference between any two sample barcodes. The differences are based on different nucleotide sequences, nucleotide sequence lengths, nucleotide sequence diversity, etc. The sample barcode can be ligated to one of the two ends of the DNA fragment. For example, the sample barcode may be ligated to the 3' end, or conversely, to the 5' end. In one or more embodiments, a ligase is added to the reaction to facilitate the first ligation of the UMI to the DNA fragment. The composition of the UMI and sample barcode is further illustrated in Figure 6.
[0177] The sequencer indexes fragments using a sequencing library specific to the wells of the flow cell.325 In one or more embodiments, the sequencer is configured for multiplex sequencing, for example, each flow cell containing multiple wells, each well configured to sequence DNA fragments belonging to one sample. The sequencer's sequencing library contains an index designated for each well in the flow cell. The index is a polynucleotide sequence used by the sequencer to anneal a primer for DNA polymerase and to add a fluorescent deoxyribonucleotide triphosphate (dNTP) used for sequencing the DNA fragment. The index may have a length of 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 base pairs. In some embodiments, an adapter sequence is included on the index and is a polynucleotide that binds to an oligo covering the flow cell. In one or more embodiments, two unique indices are used for each cell in the flow cell. One index is ligated to one end of the amplicon to which the UMI and sample barcode are attached, and the other index is ligated to the opposite end of the amplicon. In other embodiments, a single index is used for each cell. The single index may be ligated to one of the ends, either the 3' end or the 5' end.
[0178] The sequencer sequences the DNA fragments.330 In this process, the DNA fragments are amplicons having at least a sample barcode, UMI and one or more attached indices. The sequencer performs a series of steps to detect the forward and reverse sequence reads of each amplicon.
[0179] In one example of synthetic sequencing, a DNA fragment is attached to a flow cell by a series of amplification steps (e.g., under step 315). While attached, the DNA fragment is a single-stranded fragment. The reverse DNA fragment is cleaved, leaving only the forward DNA fragment attached to the flow cell. A first primer is introduced into the flow cell and anneals to the remaining forward DNA fragment. A fluorescent dNTP is introduced into the flow cell, and DNA polymerase proceeds with the synthesis of complementary polynucleotides initiated by the annealed primer. As the dNTP is added to the growing polynucleotide chain, a fluorescent tag is excited using light, which can then be detected by an imaging device. Each fluorescent tag corresponds to a different nucleotide. This process continues until all forward sequence reads are obtained. The sequencer amplifies again to obtain both the forward DNA fragment and the reverse DNA fragment attached to the flow cell. The forward DNA fragment is cleaved, leaving only the reverse DNA fragment. A similar process is performed, and a second primer is introduced and anneals to the reverse DNA fragment. DNA polymerase sequentially grows the strand of the reverse DNA fragment, while simultaneously exciting fluorescent dNTPs, causing them to fluoresce, thereby yielding reverse sequence reads.
[0180] Figure 3B is an exemplary flowchart of process 340 for detecting contamination events from sequence reads based on sample barcodes and UMI. Using forward and reverse sequence reads, the analysis system identifies contamination events based on the sequence reads. The analysis system can identify these as contamination events, single-index hopping events, double-index hopping events, sequencing errors, or any combination thereof.
[0181] In detecting single-index hopping, the analysis system identifies sequence reads with mismatched double indices as single-index hopping events.350 For example, one uniquely tagged library uses a first index at the 3' end and a second index at the 5' end, and another uniquely tagged library uses a third index at the 3' end and a fourth index at the 5' end. If a sequence read contains nucleotide sequences corresponding to one index from the first library and another index from the second library, these are mismatched indices. For example, the sequence read contains nucleotide sequences corresponding to the first and fourth indices, or the sequence read contains nucleotide sequences corresponding to the third and second indices.
[0182] If a sequence read is not identical to the predicted index, but is not a significant mismatch, the analysis system may then identify one or more sequencing errors in the sequence read.355 For example, if a fragment has an index containing one nucleotide different from the predicted index sequence, while all indices are at least 3 (i.e., at least 3 different nucleotides) apart, then, as a possibility, the single different nucleotide is attributable to a sequencing error. In some embodiments, if the index sequence of a sequence read does not exactly match any of the indices in the various libraries, the analysis system may then determine which index of the index sequence in the sequence read is closest. The analysis system may calculate the distance from each of the indices in the library. The distance is obtained based on the number of nucleotide differences; for example, if there are two nucleotides in the sequence read that are different from the index, the distance is 2. The index with the shortest distance from the sequence read is considered the closest index. Different nucleotides may be determined to be sequencing errors. In some embodiments, the closest index is determined to be a suitable index that matches the other indices.
[0183] The analysis system matches the sequence read against a matching double index.360. The analysis system may include sequence reads that have one or more different nucleotides that have been determined to be sequencing errors. For example, the first index is paired with the second index in the sequencing library. An example sequence read contains the sequences of the second index and the first index, which have the determined sequencing errors, and the sequence read is then determined to contain both the first index and the second index.
[0184] The analysis system identifies sequence reads with different sample barcodes as a double index hopping event 365. The analysis system compares the sample barcodes of the sequence reads matched in step 360. The analysis system can distinguish the sample barcodes assigned to the matched index pair. The analysis system identifies any sample barcode that is different from the assigned sample barcode as a double index hopping event. For example, the analysis system identifies any sequence reads that matched in step 360 that <atgctacgc>It is predicted that the sample barcode has a sequence (from the 5' end to the 3' end). If nothing other than the sample barcode is identified by the analysis system, the analysis system can determine that the sequence read is a double index hopping event. For example, the sample barcode sequence <caccgtaa>Sequence reads that have a 5' end to a 3' end are considered double index hopping events.
[0185] The analysis system may identify sequencing errors in the sample barcode.370 If it is determined that the sample barcode in step 365 is different from the assigned sample barcode, the analysis system may consider whether the different nucleotides are due to sequencing errors in the correct sample barcode. To recall this example, the analysis system may consider whether the sequence is due to sequencing errors in the correct sample barcode. <atgctacgc>If the sample barcode (from 5' end to 3' end) is present, the sequence read matched in step 360 is predicted. The analysis system predicts the sequence read of the sample barcode sequence. <atgcgacgc>Identify a different sequence read (from the 5' end to the 3' end). The nucleotides in bold and underlined differ from the assigned sample barcode. In step 365, the analysis system may have determined that the sequence read is a double index hopping event. In step 370, the analysis system further evaluates whether the different nucleotides are due to a sequencing error. In some embodiments, the analysis system allows for a certain error tolerance, for example, if a single nucleotide is different. In other embodiments, the analysis system may identify the closest sample barcode from multiple sample barcodes used for the sample. The analysis system may determine the distance between the sample barcodes of the sequence read from the sample barcode used in the sequencing process (for example, as shown in Figure 3A). The distance may be measured based on the number of different nucleotides, for example, four different nucleotides being assigned a distance of 4. The distance may also be based on other metrics, such as sequence length or nucleotide diversity in the sequence, but is not limited to these. The sample barcode used in the sequencing process that has the shortest distance from the sample barcode of the sequence read is determined to be the closest sample barcode.
[0186] The analysis system implements corrective actions in step 375. In step 375, the analysis system identifies a contamination event (e.g., single-index hopping, double-index hopping, sequencing error, or any combination thereof). The analysis system may implement one or more of the following corrective actions.
[0187] In one example of an improvement, the analysis system may call a sample contaminated. To call a sample contaminated, the analysis system may utilize one or more rules. For example, one rule may be based on a threshold for the total number of detected contamination events (such as single-index hopping, double-index hopping, sequencing errors, or any combination thereof). Another or more rules may be based on a threshold number of hits for a specific contamination event (single-index hopping, double-index hopping, or sequencing errors). The analysis system may utilize rules together or independently (whenever the rules are met, the sample is determined to be contaminated). The analysis system may further normalize based on sample coverage, for example, based on sequencing depth. The analysis system may, for example, notify healthcare providers of contaminated samples to prompt the collection of new samples.
[0188] In another example of an improvement, the analysis system may remove contaminated sequence reads.385 The analysis system may eliminate contaminated sequence reads and proceed with downstream analysis on the remaining uncontaminated sequence reads.
[0189] In another example of an improved solution, the analysis system may return a sequence read that has experienced index hopping to the true sample.390 Because the sample barcode functions to identify which sample the sequence read belongs to, the sequence read retains the sample barcode as an identifier for the sample to which it belongs, despite the presence of an index hopping event. Thus, the analysis system may return the sequence read to the true sample. For example, a first sequence read is matched into a first bag based on double index matching (step 360). The first bag corresponds to a first sample, but the first sequence read has a sample barcode assigned to a second sample. The analysis system may determine that the first sequence read is a double index hopping event and then place the first sequence read together with a second bag of the second sample.
[0190] The use of sample barcodes in contamination detection process 340 provides accuracy in identifying specific sequence reads as contaminated. In other contamination detection workflows, the analysis system is not informed which sequence reads are contaminated; rather, it is generally determined that the sample is contaminated. In such workflows, the only available means is to collect a new sample. In contrast, contamination detection process 340 enables the identification of specific sequence reads contaminated during the sequencing process (e.g., single-index hopping, double-index hopping, sequencing errors, etc.), expanding the available means beyond discarding contaminated samples and collecting new samples.
[0191] Furthermore, the use of sample barcodes in contamination detection process 340 can improve the calculation of methyl variant allele rates. MVAF (methyl variant allele rate) is an important metric for post-diagnostic products (e.g., MRD, prognosis, etc.). Compared to deduplication workflows based on shared endpoints of sequence reads, these workflows can perform excessive deduplication because original different cfDNA fragments may share similar endpoints. Collision rates increase with increasing coverage. At high collision rates, MVAF may be inaccurate because total coverage and / or abnormal coverage in the MVAF target is underestimated. In workflows that do not use sample barcodes, singletons (bag size 1 or fragments without duplicates) are removed from downstream analysis. Fragment recovery rates differ between abnormal and background fragments due to probe design. Singleton duplicate rates or percentages also differ depending on the input level. The linearity and accuracy of MVAF are limited by these constraints. Implementing sample barcodes improves fragment recovery rates, thereby providing more accurate MVAF calculations.
[0192] Figures 4A and 4B are illustrative descriptions of the sequencing process 300 of Figure 3A, using a sample barcode and a unique molecular identifier. In particular, Figure 4A is an illustrative description of sample preparation according to one or more embodiments, and Figure 4B is an illustrative description of sample sequencing after sample preparation in Figure 4A, according to one or more embodiments. As an example, the sample includes the original fragment, fragment 1 410, and fragment 2 420.
[0193] The first ligation 430 is performed to ligate a unique molecular identifier (UMI) to each of the original fragments. As shown in Figure 4A, UMI1 415 is ligated onto fragment 1 410 and UMI2 425 is ligated onto fragment 2 420. UMI1 415 and UMI2 425 are substantially different. The diagonally shaded boxes represent nucleotides on the non-ligated ends of the UMIs. The diagonally shaded boxes may be primers that bind the enzyme to the fragments. The UMIs are ligated onto the 3' ends of the original fragments.
[0194] Amplification 440 is performed to amplify the original fragment ligated with UMI. The resulting amplicon may contain the synthetically generated fragment and / or the original fragment. For example, fragment 1 410 is copied, resulting in three amplicons, fragment 1A 412, fragment 1B 412, and fragment 1C 412. Each of the amplicons of fragment 1 410 contains UMI1 415 ligated on the original fragment 1 410. Fragment 2 420 is amplified, resulting in two amplicons, fragment 2A 422 and fragment 2B 422, each containing UMI2 425.
[0195] In the second ligation 450, the sample barcode is attached to the amplicon. Specifically, all amplicon receive the sample barcode (SB) 405 at the 5' end opposite the UMI. SB 405 may also contain one or more nucleotides at the non-ligated end to protect the sample barcode.
[0196] In indexing 460 (see Figure 4B), the fragments containing SB405 and UMI are indexed for sequencing. Indexing involves ligating an index primer at either end of the fragment. In this example, i50x 462, which has a P5 nucleotide sequence, is ligated to the 5' end adjacent to SB405 (optionally having a protective nucleotide sequence), and i70x 464, which has a P7 nucleotide sequence, is ligated to the 3' end adjacent to UMI (optionally having a protective nucleotide sequence). The P5 and P7 nucleotide sequences are bound to complementary oligos on the sequencer flow cell, i.e., the fragments are attached to the flow cell.
[0197] In sequencing 470, the sequencer sequences the fragments to obtain forward and reverse sequence reads for each fragment. The sequence reads include an index, sample barcode, fragment sequence, and UMI. The sequence reads may omit the P5 and P7 nucleotide sequences. Thus, fragment 1A 412 yields the fragment 1A forward read 471 and the fragment 1A reverse read 472; fragment 1B 412 yields the fragment 1B forward read 473 and the fragment 1B reverse read 474; fragment 1C 412 yields the fragment 1C forward read 475 and the fragment 1C reverse read 476; fragment 2A 422 yields the fragment 2A forward read 477 and the fragment 2A reverse read 478; and fragment 2B 422 yields the fragment 2B forward read 479 and the fragment 2B reverse read 480.
[0198] Figure 5A is an exemplary description of contamination detection according to one or more embodiments. Figure 5A illustrates various sequence reads having different combinations of sample barcodes, UMIs, indices, etc. Using the sequence reads, contamination events (e.g., single-index hopping, double-index hopping, sequencing errors, etc.) are identified. For example, a predicted sequence read 500 (as a result of the sample processing process described and explained in Figures 3A, 4A, and 4B). The predicted read 500 has a first index as i50x 508, a sample barcode (SB506), a fragment sequence 502, a UMI 504, and a second index as i70x 509, from the 5' end to the 3' end. In this example, in a unique sequencing library, the index is used for the 5' end i50x 508, together with the index for the 3' end i70x 509.
[0199] The first sequence read 510 has a first index as i50x 508 from its 5' end to its 3' end, and a second index as the unknown fragment sequence 512, UMI 514, and i70y 519. The first index of the first sequence read 510 (i50x 508) matches the first index of the predicted read 500 (i50x 508). The second index of the first sequence read 510 (i70y 519) does not match the second index of the predicted read 500 (i70x 509). With respect to mismatched indices (having two indices from two different sequencing libraries), the analysis system will determine that the first sequence read 510 is a single-index hopping event. After determining that the indices are mismatched, the analysis system may proceed to evaluate the sample barcode. In this case, the sample barcode matches the predicted sample barcode SB506. Despite a single-index hopping event, the analysis system may determine that the first sequence read 510 belongs to the sample coupled with sample barcode SB506. In another embodiment, the second index matches, but the first index does not, and it remains a mismatch pair. The analysis system may similarly determine that the sequence read is a single-index hopping event.
[0200] The second sequence read 520 has a first index as i50y 528, a second index as SB X526, an unknown fragment sequence 522, UMI 524, and i70x 509, from its 5' end to its 3' end. A mismatch exists in the index pair for the second sequence read 520. The first index as i50y 528 differs from the first index of i50x 508 on the predicted read 500, while the second index of i70x 509 matches the second index of i70x509 on the predicted read 500. Similar to the first sequence read 510, the analysis system can determine that the mismatch index for the second sequence read 520 is a single-index hopping event. For the second sequence read 520, the analysis system evaluates the sample barcode SB X526. The analysis system may determine that SB X526 does not match the assigned SB506 and assign the second sequence read as belonging to the sample coupled with SB X526. As an improvement, the analysis system may return the second sequence read 520 to the sample coupled with sample barcode SB X 526.
[0201] The third sequence read 530 has a first index as i50x 508, a sample barcode as SB Y536, an unknown fragment sequence 532, UMI 534, and a second index as i70x 509, from its 5' end to its 3' end. The analysis system evaluates the indices. Both indices match the predicted read 500. The indices also match each other, i.e., the two indices form a pair in one sequencing library. Next, the analysis system evaluates the sample barcode. The analysis system will determine that the sample barcode for SB Y536 is different from that of SB 506. Therefore, the analysis system will determine that the third sequence read 530 is a double index hopping event. As a remedy, the analysis system may return the third sequence read 530 to the sample coupled with the sample barcode for SB Y536.
[0202] The fourth sequence read 540 contains, from its 5' end to its 3' end, the first index as i50z 548, the sample barcode as SB506, an unknown fragment sequence 542, UMI 544, and i70z 549. The analysis system evaluates the index. None of the indexes match the predicted read 500. The analysis system may have matched the fourth sequence read 540 with other sequence reads that have the indices i50z 548 and i70z 549. The analysis system may determine that the fourth sequence read 540 is a double index hopping event. The analysis system then evaluates the sample barcode SB506 and it matches the sample barcode of the predicted read 500. In this case, the analysis system may return the fourth sequence read 540 to the bag of sample reads coupled with the sample barcode SB506.
[0203] The fifth sequence read 550 contains, from its 5' end to its 3' end, a first index as i50x 508, SB506, an unknown fragment sequence 552, UMI 554, and a second index as i70x 509. The analysis system evaluates the indices. The fifth sequence read 550 has both indices that match the predicted read 500. The analysis system evaluates the sample barcodes. With respect to the fifth sequence read 550, the sample barcode for SB506 matches the barcode for the predicted read 500. The analysis system determines that the fifth sequence read 550 does not contain any contamination events.
[0204] The analysis system uses UMI sequences in sequence reads to identify unique original fragments. Each original fragment has a UMI ligated onto it. The analysis system can identify original fragments based at least on UMI. In some embodiments, the analysis system also relies on identification at the endpoints of sequence reads, for example, when aligned to a reference genome. For example, the analysis system may identify multiple sequence reads that have the same UMI sequence and similar endpoints as amplicons of the same original fragment. In another example, the analysis system identifies two sequence reads that have different UMI sequences originating from two different original fragments. In yet another example, the analysis system identifies two sequence reads that have the same UMI sequence but different endpoints as two different original fragments. By utilizing UMI and endpoint analysis, the likelihood of collisions is reduced, i.e., the two original fragments are determined to be the same original fragment.
[0205] Figure 5B is an exemplary description of a nucleic acid construct incorporating a sample barcode into a patterned flow cell 560 according to one or more embodiments. The patterned flow cell 560 is a slide that holds the sample when sequenced by the sequencer 1020. In one or more embodiments, the patterned flow cell 560 is a substrate having one or more columns in the substrate. Each column is separated from the other columns. Each column contains a number of nanowells covered with a lone of oligonucleotides that bind to a complementary adapter on the nucleic acid construct (e.g., P5 and P7 shown in other figures). A sequencing library containing one or two indices is assigned to each sample sequenced on the patterned flow cell 560. In one or more embodiments, a single sample is assigned to each column. In the example of Figure 5B, the patterned flow cell 560 consists of two columns, a first column 570 and a second column 580. Two nucleic acid (NA) samples are prepared, for example, according to process 300 in Figure 3A, and then multiplex-sequenced in two flow cells.
[0206] In this process, the analysis system can identify sequence reads belonging to each sample based on the index sequence on the sequence read and the sample barcode.
[0207] Sequence reads belonging to the first sample, sequenced in the first column 570, are predicted to have a double index sequence, shown as two black boxes on the nucleic acid (NA) construct in the first column 570. Furthermore, sequence reads belonging to the first sample are predicted to have a white box as the sample barcode. As an example, NA construct 572 shows the predicted double black box index sequence and the white box sample barcode sequence. The analysis system would determine that NA construct 572 does not contain any contamination events. NA construct 574 is shown to have a mismatch double index sequence with the left outer box shown in white, but NA construct 574 has a predicted sample barcode sequence for the first sample. The analysis system would likely determine that NA construct 574 has a single index hopping event, but the analysis system may assign NA construct 574 to the first sample based on the appropriate sample barcode sequence.
[0208] Sequence reads belonging to the second sample, sequenced in the second column 580, are predicted to have a double-index sequence, indicated as two white boxes on the nucleic acid (NA) construct in the second column 580. Furthermore, sequence reads belonging to the second sample are predicted to have a diagonal box as the sample barcode sequence. As another example, NA construct 582, sequenced in the second column 580, does not contain a contamination event. NA construct 582 has a predicted double-index sequence and a predicted sample barcode for the second sample. NA construct 584 has a double-index sequence assigned to the first sample, a black double outer box, but also a predicted sample barcode for the second sample. In the case of NA construct 584, the analysis system would determine that NA construct 584 is a double-index hopping event.
[0209] Figure 6 illustrates various arrangements of sample barcodes and unique molecular identifiers according to one or more embodiments. The various arrangements utilize fragment 600, which is sequenced using a double-indexed library.
[0210] In the first configuration 610, UMI 612 is ligated onto the 3' end of fragment 600, and sample barcode SB614 is ligated onto the 5' end of fragment 600. Two indices are ligated after the sample barcode and UMI. Specifically, the 5' index 602 is ligated after the sample barcode 614, and the 3' index 604 is ligated after UMI 612. The fragment also has oligos P5 and P7 after the indices, with P5 ligated onto the 5' end and P3 ligated onto the m3' end.
[0211] In the second configuration 620, UMI 622 is ligated to the 5' end of fragment 600, and sample barcode SB624 is ligated to the 3' end of fragment 600. Two indices are ligated after the sample barcode and UMI. Specifically, the 5' index 602 is ligated after the sample barcode 614, and the 3' index 604 is ligated after UMI 612. The fragment also has oligos P5 and P7 after the indices, with P5 ligated to the 5' end and P3 ligated to the 3' end.
[0212] In the third, fourth, fifth, and sixth arrangements, the UMI and sample barcode are ligated to the same side of the fragment.
[0213] In the third configuration 630, UMI 632 is ligated onto fragment 600 before the sample barcode SB634 on the 3' end of fragment 600. Then, two indices are ligated after the sample barcode and UMI. In particular, the 5' index 602 is ligated after the sample barcode 614, and the 3' index 604 is ligated after UMI 612. The fragment also has oligos P5 and P7 after the indices, with P5 ligated on the 5' end and P3 ligated on the 3' end.
[0214] In the fourth configuration 640, the sample barcode SB644 is ligated onto fragment 600 before the UMI 642 on the 3' end of fragment 600. Then, two indices are ligated after the sample barcode and UMI. In particular, the 5' index 602 is ligated after the sample barcode 614, and the 3' index 604 is ligated after the UMI 612. The fragment also has oligos P5 and P7 after the indices, with P5 ligated on the 5' end and P3 ligated on the 3' end.
[0215] In the fifth configuration 650, UMI 652 is ligated before the sample barcode 654 on the 5' end of fragment 600. Then, two indices are ligated after the sample barcode and UMI. In particular, the 5' index 602 is ligated after the sample barcode 614, and the 3' index 604 is ligated after UMI 612. The fragment also has oligos P5 and P7 after the indices, with P5 ligated on the 5' end and P3 ligated on the 3' end.
[0216] In the sixth configuration 660, the sample barcode 664 is ligated before the UMI 662 on the 5' end of fragment 600. Then, two indices are ligated after the sample barcode and UMI. In particular, the 5' index 602 is ligated after the sample barcode 614, and the 3' index 604 is ligated after the UMI 612. The fragment also has oligos P5 and P7 after the indices, with P5 ligated on the 5' end and P3 ligated on the 3' end.
[0217] II.D. Detection of other contamination In some embodiments, the analysis system can identify sequencing errors occurring in singleton sequence reads. A singleton is a read that has a sequence (unique / isolated) that exists only once between reads. Singletons do not have PCR duplication. Furthermore, a singleton with UMI is a single sequence read that has UMI and is free from other PCR duplications. Contaminated singleton reads can result from single index hopping or double index hopping events. Contaminated singleton reads can lead to read misalignment, inaccurate sequencing results, inaccurate assumptions in downstream analysis, or any combination thereof. Index hopping can result from contamination from free adapters / primers that were not properly removed after adapter ligation during the library cleaning process, which is typically cleaned by a bead-based or gel purification process to remove the free adapters or primers.
[0218] Other problems can occur during the sequencing process. One such problem is pad hopping, a form of contamination of adjacent wells in a flow cell. Pad hopping causes mixed clusters or identical sequences in adjacent wells. Pad hopping duplicate reads are PCR duplicates and coexisting sequence reads. Optical duplication can also result from pad hopping; for example, if insufficient DNA loading occurs, DNA strands "jump" from already seeded wells to adjacent wells, resulting in identical clusters. Clustering duplication can also occur as a result of pad hopping. Clustering duplication occurs when two or more adjacent wells are filled with the same original fragment.
[0219] The sequencer can be configured for multiplex sequencing. For complementary sequences, the analysis system can be configured for demultiplexing. Multiplexing allows for the simultaneous pooling and sequencing of multiple libraries during a single sequencing run by adding a unique index sequence to each DNA fragment during library preparation. Sequencing reads are sorted for each sample during demultiplexing, enabling proper alignment.
[0220] In some other examples, the method involves deciding whether to remove singletons. If so, existing singletons are not removed by default. For example, when WGBS is used to generate sequencing data, WGBS samples have almost all singletons, so singletons are not removed by default.
[0221] On the other hand, when multiplex sequencing is used to generate sequencing data, singletons may be removed by default. This is because double index hopping during enrichment of library multiplexing can likely result in contaminated singleton reads. In such cases, double index hopping can be used as evidence that a singleton read is contaminated (e.g., a contaminated singleton) and is eligible for filtering removal. For example, a method may involve deciding to remove singletons if they are identified as contaminated singletons, where contaminated singletons are characterized by double index hopping (i.e., any index of the singleton is swapped with an index of another sample).
[0222] Furthermore, in some cases, whether a singleton is removed may depend on the sequencing prediction error rate. For example, if a high prediction error rate exists (e.g., the prediction error rate exceeds the threshold error rate), the singleton will be removed. The singleton read error rate becomes more uniform over a high allele frequency (MAF) range.
[0223] Methods for identifying double-index-hopping events may include determining whether the duplication is sequencing duplication or PCR duplication resulting from PCR amplification. Double-index-hopping events are recognized when a single copy (i.e., a singleton) of a sequence read exists or when signature duplication exists, indicated by the physical coordination of neighbors in the flow cell. For example, if a duplicate read is found to originate from an adjacent cluster on the flow cell, it is then determined that the duplication occurred during sequencing and is a pad-hopping duplicate read, which is then filtered out. PCR duplication, on the other hand, does not have a spatial relationship on the flow cell.
[0224] Figure 7 shows an exemplary flowchart illustrating a process 700 for identifying index-hopping events in one or more embodiments. Process 700 may be performed by an analysis system in a demultiplex process. Process 700 is applied to double index sequencing, i.e., using two indices ligated to both ends of an NA fragment.
[0225] The analysis system filters out single-index hopping events.710 The analysis system identifies mismatched index sequences in sequence reads. In a dual-indexing system, each sequencing library contains index pairs. Paired indices are substantially unique to other indices used in other sequencing libraries. If one index sequence corresponds to a third index that is not part of an index pair in the sequencing library, the analysis system identifies a mismatch in that index sequence. For example, the first sequencing library contains a first index at the 5' end and a second index at the 3' end. Sequence reads that have an index sequence for the first index at the 5' end, excluding a third index that is different from the second index at the 3' end, would be determined to be single-index hopping events.
[0226] The analysis system matches the matched index with the sequence read. To demultiplex the sequence reads, the analysis system matches the matched index with the sequence read into individual bags. For example, if multiple mixed sequence reads belong to 12 different samples, each with a different index pair, the analysis system looks at the index sequences in the sequence reads, and the 12 different index pairs are identified. Sequence reads with the same index pair are grouped into bags.
[0227] The analysis system removes pad-hopping duplicate reads. Pad-hopping duplicate reads are identified as being relatively close to each other in a patterned flow cell. Each sequence read may contain spatial information indicating the location of each sequence read on the flow cell, which includes at least one of a lane ID, column ID, tile ID, or xy coordinate pair. To identify pad-hopping duplicate reads, the analysis system may identify groups of identical or nearly identical sequence reads. With respect to these identical or nearly identical sequence reads, the analysis system determines whether the reads coexist.
[0228] Coexistent leads can occur if at least one of the following positional relationships is met: the grouped leads share a common tile, the grouped leads are located on neighboring tiles, the grouped leads are located on different tiles within a common column, the grouped leads are located within threshold x-distances and threshold y-distances from each other on the flow cell, and the grouped leads are located within a predefined boundary region. Coexisting identical or nearly identical leads are determined to be padhopping duplicate leads. The predefined boundary region includes a geometric shape with an x-distance of 7,500 flow cell position units and a y-distance of 100,000 flow cell position units. The geometric shape includes a rectangle, with the longer side of the rectangle extending along the lane in the y-direction of the flow cell. The threshold x-distance between grouped leads is in the range of 0 to 50 mm, and the threshold y-distance between grouped leads is in the range of 0 to 50 mm. The analysis system may remove padhopping duplicate leads if the expected error rate associated with multiplex sequencing exceeds the threshold error rate. The analysis system can calculate the expected error rate by spiking the flow cell with non-human sequences before multiplex sequencing and determining the proportion of sequence reads corresponding to non-human-based sequences that are mixed with sequences corresponding to multiple biological samples.
[0229] The analysis system identifies double-index-hopping events as singletons with unique sequences among all other sequence reads matched based on the index.740 The analysis system can calculate the distance between two sequence reads based on nucleotide sequence, alignment, sequence read length, sequence diversity, etc. As an efficient approach, the analysis system can first align the sequence reads to each other and then calculate the distance between overlapping sequence reads. The analysis system determines a sequence read to be a singleton if its distance from all other sequence reads matched to the sample exceeds a threshold distance. As an example of duplicate exclusion, the analysis system creates a new bag for each different sequence read. If a sequence read is longer than the threshold distance from an existing bag, the analysis system creates a new bag for that sequence read. After classifying all sequence reads, a singleton is a sequence read in its own bag of bag size 1 that does not contain any other duplicate sequence reads. The analysis system can then proceed with improvements, for example, by removing sequence reads with double-index-hopping events from downstream analyses.
[0230] tube. A classifier for diagnosing cancer. Cancer classification involves the process of extracting genetic features and applying one or more models to the extracted features to determine cancer predictions. The extracted features provide a feature vector for the test sample, and cancer predictions are determined based on the input feature vector. Cancer predictions may include labels and / or values. Labels may be binary, indicating the presence or absence of cancer in the subject, and / or multiclass, indicating one or more specific cancer types from multiple screened cancer types. In particular, a cancer classifier may be a machine learning model that includes multiple classification parameters and a function that represents the relationship between the feature vector as input and the cancer prediction as output. Cancer predictions are obtained by inputting the feature vector along with the classification parameters into the function. In one or more embodiments, an age-covariate prediction model is used to predict the age of the test sample based on methylation features. The residual between the predicted age and the reported age of the subject is used as a feature of the cancer classifier. In one or more embodiments, the input of feature vectors to the cancer classifier is based on a set of informative fragments (also called "extremely methylated anomalous fragments" (UFXM)) determined from the test sample. Before deploying the cancer classifiers, the analysis system trains them. Furthermore, prior to training, the analysis system can apply a sample contamination detection workflow to detect and remove contamination events from the multiplex sequencing data.
[0231] III.A. Identification of Informative Fragments The analysis system can identify informative fragments within a sample using the methylation state vector of the sample. For each fragment in the sample, the analysis system can determine whether or not it is an informative fragment using the methylation state vector corresponding to that fragment. In some embodiments, the analysis system calculates a p-value score for each methylation state vector, representing the probability of observing that methylation state vector or other methylation state vectors with an even lower probability in a healthy control group. The process of calculating the p-value score is described in more detail in Section III.Ci, “p-value filtering,” below. The analysis system determines fragments with methylation state vectors that do not meet a threshold p-value score as informative fragments. In some embodiments, the analysis system further labels fragments with CpG sites where methylation or demethylation exceeds a threshold percentage as highly methylated fragments and low-methylated fragments, respectively. Highly methylated or low-methylated fragments may also be called anomalous fragments with extreme methylation (UFXM). In other embodiments, the analysis system may implement various other probabilistic models for determining informative fragments. Other examples of probabilistic models include mixed models and deep probabilistic models. In some embodiments, the analysis system can identify informative fragments using any combination of the processes described below. Using the identified informative fragments, the analysis system can filter a set of methylation state vectors of the sample for use in other processes, such as training and introducing cancer classifiers.
[0232] III. AIP Value Filtering In some embodiments, the analysis system calculates a p-value score for each methylation state vector by comparing it with methylation state vectors obtained from a healthy control group of fragments. The p-value score can describe the probability of observing a methylation state that matches that methylation state vector or other methylation state vectors with even lower probabilities in the healthy control group. To determine whether a DNA fragment is informatively methylated, the analysis system can use a healthy control group in which normally methylated fragments constitute the majority. When this probabilistic analysis is performed to determine an informative fragment, the determination may place emphasis on comparison with a control group that constitutes the healthy control group. To ensure the robustness of the healthy control group, the analysis system can select some threshold number of healthy individuals from whom the sample containing the DNA fragment is sourced. Figure 8A illustrates how to generate a data construct for the healthy control group, and using this structure, the analysis system can calculate the p-value score. Figure 8B illustrates how to calculate the p-value score using the generated data construct.
[0233] Figure 8A is a flowchart illustrating a process 800 for generating a data construct for a healthy control group according to an embodiment. To generate a data construct for a healthy control group, the analysis system can receive multiple DNA fragments (e.g., cfDNA) from multiple healthy individuals. For example, a methylation state vector can be identified for each fragment via process 300.
[0234] Using the methylation state vector of each fragment, the analysis system can subdivide the methylation state vector into a string of CpG sites 205. In some embodiments, the analysis system subdivides the methylation state vector such that all of the resulting strings are less than a given length. For example, subdividing a methylation state vector of length 11 into strings of length 3 or less results in 9 strings of length 3, 10 strings of length 2, and 11 strings of length 1. In another example, subdividing a methylation state vector of length 7 into strings of length 4 or less results in 4 strings of length 4, 5 strings of length 3, 6 strings of length 2, and 7 strings of length 1. If the methylation state vector is shorter than or the same length as the specified string length, the methylation state vector can be converted into a single string that includes all of the CpG sites of the vector.
[0235] For each possible CpG site in the vector and for each methylation state possibility, the analysis system aggregates the strings by counting the number of strings in the control group that have the specified CpG site as the first CpG site in the string and that have that methylation state possibility 810. For example, for a given CpG site, considering a string length of 3, there are 2^3, or 8, possible string configurations. For each of the 8 possible string configurations for that given CpG site, the analysis system aggregates how many times each methylation state vector possibility appears in the control group 810. Continuing with this example, for each starting CpG site x in the reference genome, the following quantities: <M x ,M x+1 ,M x+2 , >, <M x ,M x+1 ,U x+2 , >,... <U x ,U x+1 ,U x+2 , > can be aggregated. The analysis system forms a data structure that stores the counts aggregated for each starting CpG site and each string possibility 815.
[0236] Setting an upper limit on string length offers several advantages. First, depending on the maximum string length, the size of the data construct created by the analysis system can increase significantly. For example, a maximum length of ×4 means that every CpG site has at least 2^4 numbers to aggregate for a string of length ×4. Increasing the maximum string length to 5 means that every CpG site has an additional 2^4, or 16, numbers to aggregate, doubling the number to aggregate (and the computer memory required) compared to the previous string length. Reducing string size can help maintain reasonable computation and storage while forming and performing data structures (e.g., when used for subsequent access as described below). Secondly, a statistical consideration for limiting the maximum string length is to avoid overfitting in downstream models that use string counting. If long strings of CpG sites do not have a significant biological impact on the outcome (e.g., predicting the informativeness of predicting the presence of cancer), calculating probabilities based on long strings of CpG sites can be problematic because it uses a considerable amount of data that may become unavailable, thus potentially over-diluting the model for it to function properly. For example, when calculating the informativeness / cancer probability conditioned on the previous 100 CpG sites, one could use counts of 100 strings in the data structure, some of which ideally closely match the previous 100 methylation states. If only dilute counts of 100 strings are available, there may be insufficient data to determine whether the 100 strings in the test sample are informative or not.
[0237] Figure 8B is a flowchart illustrating a process 820 for identifying informative methylation fragments from an individual, according to one embodiment. In process 820, the analysis system generates methylation state vectors 200 from the target cfDNA fragment. The analysis system can process each methylation state vector as follows.
[0238] For a given methylation state vector, the analysis system enumerates all 830 possibilities of methylation state vectors having the same starting CpG site and the same length (i.e., set of CpG sites). Since each methylation state is generally either methylated or unmethylated, there are substantially two possible states for each CpG site, and therefore the count of different possibilities for a methylation state vector depends on a power of 2, thereby a methylation state vector of length n has 2 possible states. n This is associated with the possibility that, if the methylation state vector contains uncertain states at one or more CpG sites, the analysis system can list 830 possible methylation state vectors, considering only the CpG sites that have the observed states.
[0239] The analysis system accesses the data structure of a healthy control group to calculate the probability of observing each possibility of a methylation state vector for an identified starting CpG site and methylation state vector length.840 In some embodiments, the calculation of the probability of observing a given possibility models a co-probability calculation using Markov chain probabilities. The Markov model is at least partially trained on an assessment of the methylation state of each CpG site in the corresponding CpG sites of each fragment (e.g., a nucleic acid methylation fragment) across a healthy non-cancer cohort with a total of corresponding CpG sites. For example, a Markov model (e.g., a Hidden Markov Model (HMM)) is used to determine the probability that a sequence of methylation states (e.g., including "M" or "U") may be observed for a nucleic acid methylation fragment among a total of nucleic acid methylation fragments, by assuming a set of probabilities for each state of the sequence that determines the probability of observing the next state of the sequence. The set of probabilities is obtained by training the HMM. Such training may involve calculating statistical parameters (e.g., the probability of transitioning from the first state to the second state (transition probability) and / or the probability of observing a given methylation state for each CpG site (release probability)) taking into account an initial training dataset of methylation state sequences (e.g., methylation patterns). The HMM can be trained using supervised training (e.g., with samples where the basic sequence and observed states are known) and / or unsupervised training (e.g., Viterbi learning, maximum likelihood estimation, expectation maximization training and / or Baum-Welch training). In other embodiments, calculation methods other than Markov chain probabilities are used to determine the probability of observing each possibility of the methylation state vector. For example, such calculation methods may include learned representations. The p-value threshold may be 0.01 to 0.10 or 0.03 to 0.06. The p-value threshold may be 0.05. The p-value threshold may be less than 0.01, less than 0.001, or less than 0.0001.
[0240] The analysis system calculates a p-value score for a methylation state vector using the probabilities calculated for each possibility.850 In some embodiments, this includes identifying the calculated probabilities corresponding to the possibilities of matching the given methylation state vector. Specifically, this may be the possibility of having the same set of CpG sites as the methylation state vector, or similarly, the possibility of having the same starting CpG sites and length. The analysis system can generate a p-value score by summing the calculated probabilities of all possibilities that have a probability less than or equal to the identified probability.
[0241] This p-value may represent the probability of observing the methylation state vector of the fragment, or other methylation state vectors, with an even lower probability in a healthy control group. Therefore, a low p-value score generally corresponds to a methylation state vector that is rare in healthy individuals, thereby labeling the fragment with informative methylation compared to a healthy control group. A high p-value score generally corresponds to a methylation state vector that is expected to be present in a healthy individual in a relative sense. If the healthy control group is non-cancer, for example, a low p-value may indicate the presence of cancer in the subject because the fragment may show informative methylation compared to the non-cancer group.
[0242] As described above, the analysis system can calculate p-value scores for multiple methylation state vectors, each representing a cfDNA fragment in the test sample. To identify which fragments exhibit informative methylation, the analysis system can filter the set of methylation state vectors based on the p-value scores.860 In some embodiments, filtering is performed by comparing the p-value scores to a threshold and retaining only fragments with p-values below the threshold. This threshold p-value score may be 0.1, 0.01, 0.001, 0.0001, or similar values.
[0243] Following the exemplary results from Process 800, the analysis system may yield a median (range) of 2,800 (1,500–12,000) fragments showing informative methylation patterns in participants without cancer during training, and a median (range) of 3,000 (1,200–420,000) fragments showing informative methylation patterns in participants with cancer during training. A filtered set of fragments with informative methylation patterns can be used for the downstream analysis described in Section III below.
[0244] In some embodiments, the analysis system uses a sliding window to determine the possibilities of the methylation state vector and calculate the p-values.855 Rather than enumerating possibilities for the entire methylation state vector and calculating the p-values, the analysis system may enumerate possibilities only for a window of consecutive CpG sites and calculate the p-values, with the length of this window (number of CpG sites) being shorter than at least some fragments (otherwise the window serves no purpose). The window length may be static, user-specified, dynamic, or otherwise selectable.
[0245] When calculating the p-value for a methylation state vector larger than the window, the window can identify a contiguous set of CpG sites within the window, starting from the first CpG site in the vector. The analysis system can calculate the p-value score for the window containing the first CpG site. The analysis system can then "slide" the window to the second CpG site in the vector and calculate another p-value score for the second window. Thus, for a window size l and a methylation vector length m, each methylation state vector can generate m-l+1 p-value scores. After the p-value calculations for each part of the vector are complete, the lowest p-value score from all sliding windows can be adopted as the overall p-value score for the methylation state vector. In other embodiments, the analysis system aggregates the p-value scores of the methylation state vectors to generate an overall p-value score.
[0246] Using a sliding window can help reduce the number of possible methylation state vectors to enumerate and the number of corresponding probability calculations that would otherwise need to be performed. For example, a fragment may have more than 54 CpG sites. Instead of calculating probabilities for 2^54 (approximately 1.8 × 10^16) possibilities to produce a single p-score, the analysis system uses a window of size 5 (for example), resulting in 50 p-value calculations for each of the 50 windows of the fragment's methylation state vectors. Each of the 50 calculations can enumerate 2^5 (32) possible methylation state vectors, totaling 50 × 2^5 (1.6 × 10^3) probability calculations. This significantly reduces the calculations that need to be performed without significantly impacting the accurate identification of informative fragments.
[0247] In embodiments with uncertain states, the analysis system can calculate a p-value score by summing the CpG sites with uncertain states in the methylation state vector of the fragment. The analysis system can identify all possibilities that match all methylation states (excluding uncertain states) in the methylation state vector. The analysis system can assign probabilities to the methylation state vector as the sum of the probabilities of the identified possibilities. As an example, the analysis system observes that the methylation states of CpG sites 1 and 3 are consistent with the methylation states of the fragment at CpG sites 1 and 3,<M1,I2,U3> The probability of the methylation state vector is<M1,M2,U3> and<M1,U2,U3> The probability of a methylation state vector can be calculated as the sum of the probabilities of its possible states. The above method for summing CpG sites with uncertain states can use probability calculations up to 2^i, where i is the number of uncertain states in the methylation state vector. In another embodiment, a dynamic programming algorithm can also be implemented to calculate the probability of a methylation state vector having one or more uncertain states. Advantageously, the dynamic programming algorithm operates in linear computation time.
[0248] In some embodiments, the computational load associated with calculating probabilities and / or p-value scores can be further reduced by caching at least some of the calculations. For example, the analysis system can cach the calculation of probabilities for methylation state vectors (or their windows) in temporary or persistent memory. If other fragments have the same CpG sites, caching the probability of probabilities allows for efficient calculation of p-score values without the need to recalculate the basic probability of probabilities. Similarly, the analysis system can calculate a p-value score from the vector (or its window) for each possibility of a methylation state vector associated with a set of CpG sites. The analysis system can cach the p-value scores to determine the p-value scores of other fragments containing the same CpG sites. In general, the p-value scores of possibilities of methylation state vectors having the same CpG sites can be used to determine the p-value scores of different possibilities from the same set of CpG sites.
[0249] One or more nucleic acid methylation fragments can be filtered before training a region model or cancer classifier. Filtering nucleic acid methylation fragments may involve removing each nucleic acid methylation fragment from a group of corresponding nucleic acid methylation fragments that does not meet one or more selection criteria (e.g., less than or greater than one selection criterion). One or more selection criteria may include p-value thresholds. The output p-value for each nucleic acid methylation fragment is determined, at least in part, based on a comparison of the corresponding methylation pattern of each nucleic acid methylation fragment with the corresponding distribution of methylation patterns of such nucleic acid methylation fragments in a healthy non-cancer cohort dataset containing the corresponding multiple CpG sites of each nucleic acid methylation fragment.
[0250] Filtering multiple nucleic acid methylation fragments may involve removing each nucleic acid methylation fragment that does not satisfy a p-value threshold. The filter can be applied to the methylation pattern of each nucleic acid methylation fragment using the methylation pattern observed across the first set of nucleic acid methylation fragments. Each methylation pattern of each nucleic acid methylation fragment (e.g., fragment 1, ..., fragment N) may include one or more corresponding methylation sites (e.g., CpG sites) identified by methylation site identifiers, and the corresponding methylation pattern (represented by sequences of 1s and 0s, where each "1" indicates a methylated CpG site at one or more CpG sites, and each "0" indicates an unmethylated CpG site at one or more CpG sites). Using the methylation pattern observed across the first set of nucleic acid methylation fragments, a methylation state distribution of CpG site states (e.g., CpG site A, CpG site B, ..., CpG site ZZZ) collectively represented by the first set of nucleic acid methylation fragments can be constructed. Further details relating to the processing of nucleic acid methylated fragments are described in U.S. Provisional Patent Application No. 17 / 191,914, entitled "Systems and Methods for Cancer Condition Determination Using Autoencoders," filed on March x4, 2021, which is incorporated in its entirety by reference in this application.
[0251] Each nucleic acid methylation fragment may not meet one or more selection criteria if it has an informative methylation score that falls below an informative methylation score threshold. In this situation, the informative methylation score can be determined by a mixed model. For example, a mixed model can detect an informative methylation pattern in a nucleic acid methylation fragment by determining the likelihood of methylation state vectors (e.g., methylation patterns) for each nucleic acid methylation fragment based on the number of possible methylation state vectors of the same length and at the same corresponding genomic location. This can be done by generating multiple possible methylation states for vectors of specified length at each genomic location in the reference genome. Using the multiple possible methylation states, the total number of possible methylation states, followed by the probability of each predicted methylation state at the genomic location, can be determined. Then, the likelihood of a sample nucleic acid methylation fragment corresponding to a genomic location in the reference genome can be determined by matching the sample nucleic acid methylation fragment to the predicted (e.g., possible) methylation states and obtaining the calculated probability of the predicted methylation states. Next, an informative methylation score can be calculated based on the probability of nucleic acid methylation fragments in the sample.
[0252] Each nucleic acid methylation fragment may fail to meet one or more selection criteria if it has a number of residues below the threshold. The threshold number of residues may be 10–50, 50–100, 100–150, or greater than 150. The threshold number of residues may be a fixed value of 20–90. Each nucleic acid methylation fragment may fail to meet one or more selection criteria if it has a number of CpG sites below the threshold. The threshold number of CpG sites may be 8, 5, 6, 7, 8, 9, or 10. Each nucleic acid methylation fragment may fail to meet one or more selection criteria if its genome start and end locations indicate that it represents a number of nucleotides below the threshold in the human genome reference sequence.
[0253] Filtering can remove nucleic acid methylation fragments from among multiple corresponding nucleic acid methylation fragments that have the same corresponding methylation pattern and the same corresponding genome start and end locations as another nucleic acid methylation fragment in the same set of multiple nucleic acid methylation fragments. This filtering step can remove exact duplicate fragments (including PCR duplicate fragments, if applicable). Filtering can remove nucleic acid methylation fragments from among multiple corresponding nucleic acid methylation fragments that have the same corresponding genome start and end locations as another nucleic acid methylation fragment and have a number of different methylation states less than a threshold. The threshold number of different methylation states used for retaining nucleic acid methylation fragments can be 1, 2, 3, 8, 5, or greater than 5. For example, even if a first nucleic acid methylation fragment has the same corresponding genome start and end locations as a second nucleic acid methylation fragment, it will be retained if each CpG site (e.g., aligned with a reference genome) has at least one, at least two, at least three, at least eight, or at least five different methylation states. As another example, a first nucleic acid methylation fragment is also retained, which has the same methylation state vector (e.g., methylation pattern) as the second nucleic acid methylation fragment, but with different corresponding genome start and end positions.
[0254] Filtering can remove assay artifacts in multiple nucleic acid methylation fragments. Removal of assay artifacts may include removing sequence reads obtained from sequenced hybridization probes and / or sequence reads obtained from sequences that failed to undergo conversion during bisulfite conversion. Filtering can remove contaminants (e.g., those from sequencing, nucleic acid separation, and / or sample preparation).
[0255] Filtering allows for the removal of subsets of methylation fragments from multiple methylation fragments based on cross-information filtering of each methylation fragment for cancer status across multiple training subjects. For example, cross-information provides a measure of interdependence between two simultaneously sampled subject conditions. Cross-information is determined by selecting an independent set of CpG sites (e.g., in all or part of nucleic acid methylation fragments) from one or more datasets and comparing the probabilities of methylation states of the CpG site set between two sample groups (e.g., genotype datasets, biological samples and / or subsets and / or groups of subjects). The cross-information score can indicate the probabilities of the methylation patterns of the first and second conditions in each region within each frame of a sliding window, and thus indicate the discriminative ability of each region. The cross-information score can be calculated in a similar manner for each region within each frame as the sliding window progresses across the selected set of CpG sites and / or selected genomic regions. Further details regarding cross-information filtering are disclosed in U.S. Patent Application No. 17 / 119,606, filed on 11 December 2020, entitled "Cancer Classification using Patch Convolutional Neural Networks," which is incorporated in its entirety by reference in this application.
[0256] III.A.II. Highly Methylated and Lowly Methylated Fragments In some embodiments, the analysis system determines an informative fragment as one in which the number of CpG sites exceeds a threshold and the methylation of CpG sites exceeds a threshold percentage, or the unmethylated CpG sites exceed a threshold percentage, and the analysis system identifies such fragments as highly methylated or low-methylated fragments. Examples of thresholds for fragment length (or number of CpG sites) include values exceeding 3, 4, 5, 6, 7, 8, 9, 10, etc. Examples of percentage thresholds for methylation or unmethylation include values exceeding 80%, 85%, 90%, or 95%, or any other percentage within the range of 50% to 100%.
[0257] III.B. Training of cancer classifiers Figure 9A is a flowchart illustrating the training process 900 of a cancer classifier according to one embodiment. The analysis system acquires a set of informative fragments and a number of training samples having a label for a certain cancer type (910). The number of training samples includes any combination of samples from healthy individuals with the general label "non-cancerous" and samples from subjects with the general or specific label "cancer" (e.g., "breast cancer," "lung cancer," etc.). Training samples from subjects of a certain cancer type are called a cancer type cohort or cancer type cohort.
[0258] The analysis system determines feature vectors for each training sample based on the set of informative fragments of that training sample.920 The analysis system can calculate an informative score for each CpG site within the initial set of CpG sites. The initial set of CpG sites can be all or part of all CpG sites in the human genome, approximately 10 4 , 10 5 , 10 6 , 10 7 , 10 8 These are some possible approaches. In one embodiment, the analysis system defines an informative score for a feature vector using binary scoring based on whether an informative fragment exists in a set of informative fragments containing CpG sites. In another embodiment, the analysis system defines an informative score based on the count of informative fragments that overlap with CpG sites. For example, the analysis system may use ternary scoring, assigning a first score if no informative fragments are present, a second score if a small number of informative fragments are present, and a third score if more than a small number of informative fragments are present. For example, the analysis system counts 5x informative fragments that overlap with CpG sites in a sample and calculates an informative score based on the 5x count.
[0259] Once all informative scores for the training sample have been determined, the analysis system can determine a feature vector for each element as a vector of elements containing one informative score associated with one of the CpG sites in the initial set. The analysis system can normalize the informative scores of the feature vector based on the sample coverage, where coverage refers to the median or mean of the sequencing depth across all CpG sites covered by the initial set of CpG sites used by the classifier, or it may be a value based on a set of informative fragments of a particular training sample.
[0260] As an example, see Figure 9B illustrating the matrix of training feature vectors 922. In this example, the analysis system identifies CpG sites [K]926 to consider in generating feature vectors for the cancer classifier. The analysis system selects training sample [N]924. The analysis system determines a first informative score 928 for the first arbitrary CpG site [k1] used in the feature vectors of training sample [n1]. The analysis system examines each informative fragment within the set of informative fragments. Once the analysis system identifies at least one informative fragment containing the first CpG site, the analysis system determines the first informative score 928 for the first CpG site to be 1, as shown in Figure 9B. Considering a second arbitrary CpG site [k2], the analysis system similarly checks the set of informative fragments for at least one fragment containing the second CpG site [k2]. If the analysis system does not detect such informative fragments containing the second CpG site, the analysis system determines the second informative score 929 for the second CpG site [k2] to be 0, as shown in Figure 9B. Once the analysis system has determined all informative scores for the initial set of CpG sites, it determines a feature vector containing the informative scores for the first training sample [n1], which includes the first informative score 928:1 for the first CpG site [k1], the second informative score 929:0 for the second CpG site [k2], and subsequent informative scores, thus forming the feature vector [1,0,...].
[0261] Further methods for feature-finding samples are described in U.S. Patent Application No. 15 / 931,022, entitled "Model-Based Featurization and Classification"; U.S. Patent Application No. 16 / 579,805, entitled "Mixture Model for Targeted Sequencing"; U.S. Patent Application No. 16 / 352,602, entitled "Anomalous Fragment Detection and Classification"; and U.S. Patent Application No. 16 / 723,716, entitled "Source of Origin Deconvolution Based on Methylation Fragments in Cell-Free DNA Samples," all of which are incorporated herein by reference in their entirety.
[0262] The analysis system may further restrict the range of CpG sites that are considered for use in cancer classifiers. The analysis system calculates the information gain for each CpG site in the initial set of CpG sites based on the feature vectors of the training sample 930. From step 920, each training sample has feature vectors that may contain informative scores for all CpG sites in the initial set of CpG sites (and possibly all CpG sites in the human genome). However, some CpG sites in the initial set may not be as informative as others in distinguishing cancer types, or they may overlap with other CpG sites.
[0263] In one embodiment, the analysis system calculates the information gain for each cancer type and each CpG site in the initial set and determines whether that CpG site should be included in the classifier. The information gain is calculated by comparing a training sample with a given cancer type with all other samples. For example, two random variables, “informative fragment” (“IF”) and “cancer type” (“CT”), are used. In one embodiment, IF is a binary variable indicating whether an informative fragment exists in a given sample that overlaps with a given CpG site, for informative score / feature vector. CT is a random variable indicating whether the cancer is of a particular type. The analysis system calculates the mutual information about CT assuming AF. This is the number of bits of information obtained regarding the cancer type, given that it is known whether a specific CpG site overlapping with an informative fragment exists. In practice, for the first cancer type, the analysis system calculates the pairwise mutual information gain with each of the other cancer types and then sums the mutual information gains for all other cancer types.
[0264] For a given cancer type, the analysis system can use this information to rank CpG sites based on their cancer-specificity. This procedure can be repeated for all cancer types considered. If a particular region is commonly informatively methylated in training samples of a given cancer, but not in training samples of other cancer types or healthy training samples, CpG sites overlapping with that informative fragment may have a high information gain for the given cancer type. The ranked CpG sites for each cancer type are greedily added (selected) to a set of CpG sites selected based on their rank for use in the cancer classifier.
[0265] In additional embodiments, the analysis system may consider other selection criteria for selecting informative CpG sites to be used in the cancer classifier. One selection criterion may be that the selected CpG sites are beyond a threshold interval from other selected CpG sites. For example, the selected CpG sites are more than a threshold number of base pairs (e.g., 100 base pairs) away from other selected CpG sites, thereby ensuring that CpG sites within the threshold interval are not selected for consideration in the cancer classifier.
[0266] In one embodiment, based on a set of CpG sites selected from an initial set, the analysis system can modify the feature vectors of the training sample as needed.950 For example, the analysis system can truncate feature vectors to remove informative scores corresponding to CpG sites that are not in the selected set of CpG sites.
[0267] Using the feature vectors of the training samples, the analysis system can train a cancer classifier in one of several ways. The feature vectors may correspond to an initial set of CpG sites from step 920 or a selected set of CpG sites from step 950. In one embodiment, the analysis system trains a binary cancer classifier to distinguish between cancerous and non-cancerous samples based on the feature vectors of the training samples 960. In this way, the analysis system uses training samples that include non-cancerous samples from healthy individuals and cancerous samples from subjects. Each training sample may have one of two labels: "cancer" or "non-cancerous". In this embodiment, the classifier outputs a cancer prediction indicating the likelihood of the presence or absence of cancer.
[0268] In another embodiment, the analysis system trains a multi-class cancer classifier to distinguish between multiple cancer types (also called primary tissue (TOO) markers).970 A cancer type may include one or more cancers and may also include non-cancerous types (including any other diseases or genetic disorders, etc.). To do this, the analysis system uses a cancer type cohort, which may or may not include a non-cancerous type cohort. In this multi-cancerous embodiment, the cancer classifier is trained to determine a cancer prediction (or more specifically a TOO prediction) including a predicted value for each of the cancer types to be classified. The predicted value may correspond to the likelihood that a given training sample (test sample in estimation time) has each of the cancer types. In one embodiment, the predicted value is scored between 0 and 100, and the cumulative sum of the predicted values is equal to 100. For example, the cancer classifier returns a cancer prediction including predicted values for breast cancer, lung cancer, and non-cancerous cancer. For example, the classifier might return a cancer prediction that the test sample has a 65% likelihood of breast cancer, a 25% likelihood of lung cancer, and a 10% likelihood of being non-cancerous. The analysis system can further evaluate the predictions to generate a prediction of the presence of one or more cancers in the sample, also known as a TOO prediction, which indicates one or more TOO labels, for example, a first TOO label with the highest prediction value, a second TOO label with the second highest prediction value, and so on. Continuing the above example and considering the percentages, in this example, the system might determine that the sample has breast cancer, considering that breast cancer has the highest likelihood.
[0269] In both embodiments, the analysis system trains the cancer classifier by inputting a set of training samples and their feature vectors, and adjusting the classification parameters, so that the classifier's function accurately correlates the training feature vectors with their corresponding labels. For iterative batch training of the cancer classifier, the analysis system can classify the training samples into one or more sets of training samples. After inputting all sets of training samples, including their training feature vectors, and adjusting the classification parameters, the cancer classifier is sufficiently trained to label the test samples according to their feature vectors within a reasonable margin of error. The analysis system may train the cancer classifier according to one of several methods. For example, a binary cancer classifier may be an L2 regularized logistic regression classifier trained using a log loss function. As another example, a multi-cancer classifier may be a multinomial logistic regression. In fact, any type of cancer classifier may be trained using other techniques. There are many such techniques, including the potential use of machine learning algorithms such as kernel methods, random forest classifiers, mixture models, autoencoder models, and multilayer neural networks.
[0270] Examples of classifiers include logistic regression algorithms, neural network algorithms, support vector machine algorithms, naive Bayes algorithms, k-nearest neighbor algorithms, boosted tree algorithms, random forest algorithms, decision tree algorithms, multinomial logistic regression algorithms, linear models, or linear regression algorithms.
[0271] III.C. Arrangement of Cancer Classifiers During the use of the cancer classifier, the analysis system may take test samples from subjects of an unknown cancer type. The analysis system can process the test samples containing DNA molecules and generate a set of informative fragments by any combination of processes 300, 400, and 420. The analysis system can determine the test feature vectors used by the cancer classifier, following a similar principle to that described in process 500. The analysis system can calculate an informative score for each CpG site within a set of multiple CpG sites used by the cancer classifier. For example, the cancer classifier receives a feature vector containing informative scores for 1,000 selected CpG sites as input. The analysis system can thus determine the test feature vector containing informative scores for the 1,000 selected CpG sites based on the set of informative fragments. The analysis system can calculate the informative scores in a similar manner to that for training samples. In some embodiments, the analysis system explicitly expresses the informative scores as binary scores based on whether or not highly methylated or hypomethylated fragments are present in the set of informative fragments containing CpG sites.
[0272] Next, the analysis system can input the test feature vector into the cancer classifier. The function of the cancer classifier is to generate a cancer prediction based on the classification parameters trained in process 500 and the test feature vector. In the first method, the cancer prediction is binary and is selected from the group consisting of "cancer" or "non-cancer". In the second method, the cancer prediction is selected from the group of many cancer types and "non-cancer". In yet another embodiment, the cancer prediction has a predicted value for each of the many cancer types. Further, the analysis system can determine that its test sample is most likely to be one of the cancer types. According to the above example where the cancer prediction of the test sample has a likelihood of 65% for breast cancer, 25% for lung cancer, and 10% for non-cancer, the analysis system may determine that the test sample is most likely to have breast cancer. In another example where the cancer prediction is binary with a likelihood of 60% for non-cancer and 40% for cancer, the analysis system determines that the test sample is most likely not to be cancer. In yet another embodiment, the cancer prediction with the highest probability can be further compared with a threshold (e.g., 40%, 50%, 60%, 70%) in order to call the subject as a person having that cancer type. If the cancer prediction with the highest probability does not exceed the threshold, the analysis system may return a result of uncertainty.
[0273] In another embodiment, the analysis system combines a cancer classifier trained in step 960 of process 900 with another cancer classifier trained in step 970 of process 900. The analysis system can input a test feature vector into a cancer classifier trained as a binary classifier in step 960 of process 900. The analysis system can receive an output of cancer prediction. The cancer prediction can be a binary prediction indicating whether the subject is likely to have cancer or not likely to have cancer. In other embodiments, the cancer prediction includes prediction values representing the likelihood of cancer and the likelihood of non-cancer. For example, the cancer prediction has an 85% cancer prediction value and a 15% non-cancer prediction value. The analysis system can determine that the subject is likely to have cancer. If the analysis system determines that the subject is likely to have cancer, the analysis system can input the test feature vector into a multi-class cancer classifier trained to distinguish various cancer types. The multi-class cancer classifier receives the test feature vector and returns a cancer prediction of a cancer type among a plurality of cancer types. For example, the multi-class cancer classifier provides a cancer prediction specifying that the subject is most likely to have ovarian cancer. In another embodiment, the multi-class cancer classifier provides prediction values for each cancer type of a plurality of cancer types. For example, the cancer prediction can include a 40% breast cancer type prediction value, a 15% colorectal cancer type prediction value, and a 45% liver cancer type prediction value.
[0274] According to a general embodiment of binary cancer classification, the analysis system can determine a cancer score of a test sample based on sequencing data of the test sample (e.g., methylation sequencing data, SNP sequencing data, other DNA sequencing data, RNA sequencing data, etc.). The analysis system compares the cancer score of the test sample with a binary threshold cutoff and predicts whether the test sample may have cancer. This binary threshold cutoff can be adjusted using a TOO threshold process based on one or more TOO subtype classes. The analysis system further generates a feature vector of the test sample for use in a multi-class cancer classifier and determines a cancer prediction indicating one or more likely cancer types.
[0275] A classifier can be used to determine the disease status of a subject, for example, a subject whose disease status is unknown. This method includes obtaining an electronic test genome data construct (e.g., single-time period test data) containing values for each of several genomic features of several corresponding nucleic acid fragments in a biological sample obtained from the subject. Next, this method includes applying the test genome data construct to a test classifier to determine the disease status of the subject. The subject does not need to have been previously diagnosed with a disease status.
[0276] The classifier may be a time-series classifier that uses at least (i) a first test genome data construct generated from a first biological sample obtained from the subject at a first time point, and (ii) a second test genome data construct generated from a second biological sample obtained from the subject at a second time point.
[0277] A trained classifier can be used to determine the disease status of a subject, for example, a subject whose disease status is unknown. In this case, the method includes the step of obtaining a test time-series dataset electronically for the subject, which includes corresponding genotype data for each time point at each point in time, containing values for the genotype characteristics of corresponding nucleic acid fragments in corresponding biological samples obtained from the subject at each point in time, and information indicating the time interval between each pair of consecutive time points at each point in time. Next, the method includes the step of applying the test genotype data construct to the test classifier to determine the disease status of the subject. The subject does not need to have been previously diagnosed with a disease status.
[0278] IV.Applications In some embodiments, the methods, analysis systems, and / or classifiers of the present invention can be used for detecting the presence of cancer, monitoring cancer progression or recurrence, monitoring treatment response or effectiveness, determining or monitoring the presence of minimal residual disease (MRD), or any combination thereof. For example, as described herein, the classifiers can be used to generate a probability score (e.g., 0 to 100) that represents the likelihood that a test feature vector originates from a subject with cancer. In some embodiments, the probability score can be compared to a threshold probability to determine whether or not a subject has cancer. In other embodiments, the probability score can be evaluated at several different time points (e.g., before or after treatment) to monitor disease progression or treatment effectiveness (e.g., treatment effect). In yet another embodiment, the likelihood or probability score can be used to make or influence clinical decisions (e.g., diagnosis of cancer, selection of treatment, evaluation of treatment effectiveness, etc.). For example, in one embodiment, if the probability score exceeds a threshold, a physician can prescribe appropriate treatment.
[0279] IV.A. Early detection of cancer The method and / or classifier of the present invention is used to detect the presence or absence of cancer in subjects suspected of having cancer. For example, using the classifier (e.g., as described in Section III and illustrated in Section V), a cancer prediction can be determined, which represents the likelihood that the test feature vector originates from a subject having cancer.
[0280] In one embodiment, cancer prediction is the likelihood (e.g., a score between 0 and 100) of whether a test sample has cancer or not (i.e., a binary classification). Thus, the analysis system can determine a threshold for determining whether a subject has cancer. For example, a cancer prediction of 60 or higher may indicate that the subject has cancer. In yet another embodiment, a cancer prediction of 65 or higher, 70 or higher, 75 or higher, 80 or higher, 85 or higher, 90 or higher, or 95 or higher indicates that the subject has cancer. In yet another embodiment, cancer prediction indicates the severity of the disease. For example, a cancer prediction of 80 may indicate a more severe form or a later stage of cancer compared to a cancer prediction of less than 80 (e.g., a probability score of 70). Similarly, an increase in cancer prediction over time (e.g., determined by classifying the test feature vectors of multiple samples taken from the same subject at two or more time points) may indicate disease progression, and a decrease in cancer prediction over time may indicate successful treatment.
[0281] In another embodiment, cancer prediction includes multiple prediction values, with each of the multiple cancer types being classified (i.e., multi-class classification) being assigned a prediction value (e.g., a score from 0 to 100). These prediction values may correspond to the likelihood that a given training sample (and training sample under prediction) has each of the cancer types. The analysis system identifies the cancer type with the highest prediction value, indicating that the subject is likely to have that cancer type. In another embodiment, the analysis system compares the highest prediction value to a threshold (e.g., 50, 55, 60, 65, 70, 75, 80, 85, etc.) to further determine that the subject is likely to have that cancer type. In another embodiment, the prediction values may also indicate disease severity. For example, a prediction value above 80 may indicate a more severe form or a later stage of cancer compared to a prediction value of 60. Similarly, an increase in the prediction value over time (e.g., determined by the classification of test feature vectors from multiple samples taken from the same subject at two or more time points) may indicate disease progression, and a decrease in the prediction value over time may indicate treatment success.
[0282] According to several aspects of the present invention, the methods and systems of the present invention can be trained to detect or classify multiple cancer indicators. For example, the methods, systems and classifiers of the present invention can be used to detect one or more, two or more, three or more, five or more, ten or more, fifteen or more, or twenty or more different types of cancer.
[0283] Examples of cancers that can be detected using the methods, systems, and classifiers of the present invention include cancer, lymphoma, blastoma, sarcoma, and leukemia or malignant lymphoma. More specific examples of such cancers include, but are not limited to, squamous cell carcinoma (e.g., epithelial squamous cell carcinoma), skin cancer, melanoma, lung cancer (small cell lung cancer, non-small cell lung cancer ("NSCLC"), lung adenocarcinoma and lung squamous cell carcinoma, etc.), peritoneal cancer, gastric cancer (e.g., gastrointestinal cancer), pancreatic cancer (e.g., pancreatic ductal adenocarcinoma), cervical cancer, ovarian cancer (e.g., high-grade serous ovarian cancer), liver cancer (e.g., hepatocellular carcinoma (HCC)), liver tumors, liver cancer, and bladder cancer. (e.g., urothelial bladder cancer), testicular cancer (germ cell tumors), breast cancer (e.g., HER2-positive, HER2-negative, and triple-negative breast cancer), brain tumors (e.g., astrocytoma, glioma (e.g., glioblastoma)), colon cancer, rectal cancer, colorectal cancer, endometrial cancer or uterine cancer, salivary gland cancer, kidney cancer or renal cancer (e.g., renal cell carcinoma, nephroblastoma, or Wilms' tumor), prostate cancer, vulvar cancer, thyroid cancer, anal cancer, penile cancer, head and neck cancer, esophageal cancer, and nasopharyngeal cancer (NPC). Further examples of cancers include, but are not limited to, retinoblastoma, theca cell tumor, masculinizing cell tumor, hematological malignancies (but are not limited to, non-Hodgkin lymphoma (NHL), multiple myeloma, and acute hematological malignancies), endometriosis, fibrosarcoma, choriocarcinoma, laryngeal cancer, Kaposi's sarcoma, Schwann cell tumor, oligodendroglioma, neuroblastoma, rhabdomyosarcoma, osteosarcoma, leiomyosarcoma, and urinary tract cancer.
[0284] In some embodiments, cancer is anorectal cancer, bladder cancer, breast cancer, cervical cancer, colorectal cancer, esophageal cancer, gastric cancer, head and neck cancer, hepatobiliary tract cancer, leukemia, lung cancer, lymphoma, melanoma, multiple myeloma, ovarian cancer, pancreatic cancer, prostate cancer, kidney cancer, thyroid cancer, uterine cancer, or any combination thereof.
[0285] In some embodiments, one or more cancers may be “high-signal” cancers (defined as cancers with a 5-year cancer-specific mortality rate greater than 50%), such as anorectal cancer, colorectal cancer, esophageal cancer, head and neck cancer, hepatobiliary tract cancer, lung cancer, ovarian cancer, and pancreatic cancer, as well as lymphoma and multiple myeloma. High-signal cancers are more aggressive and typically have above-average cell-free nucleic acid concentrations in test samples taken from patients.
[0286] IV.B. Monitoring of Cancer and Treatment In some embodiments, cancer predictions can be evaluated at multiple different time points (e.g., before or after treatment) to monitor disease progression or treatment effectiveness (e.g., treatment efficacy). For example, the present invention includes a method comprising the steps of: collecting a first sample (e.g., a first plasma cfDNA sample) from a cancer patient at a first time point; determining a first cancer prediction from the sample; and collecting a second test sample (e.g., a second plasma cfDNA sample) from a cancer patient at a second time point; and determining a second cancer prediction (as described herein) from the sample.
[0287] In certain embodiments, the first time point is before cancer treatment (e.g., before surgical resection or therapeutic intervention), and the second time point is after cancer treatment (e.g., after surgical resection or therapeutic intervention), and the effectiveness of the treatment is monitored using a classifier. For example, if the second cancer prediction is lower than the first cancer prediction, the treatment is considered successful. However, if the second cancer prediction is higher than the first cancer prediction, the treatment is considered unsuccessful. In other embodiments, both the first and second time points are before cancer treatment (e.g., before surgical resection or therapeutic intervention). In other embodiments, both the first and second time points are after cancer treatment (e.g., after surgical resection or therapeutic intervention). In yet another embodiment, cfDNA samples are collected from cancer patients at the first and second time points and analyzed to, for example, monitor cancer progression, determine whether the cancer is in remission (e.g., after treatment), monitor or detect residual lesions or disease recurrence, or monitor the effectiveness of the treatment (e.g., therapy).
[0288] Those skilled in the art will readily understand that the cancerous state of a cancer patient can be monitored by taking test samples from the patient at any desired time point and analyzing them according to the method of the present invention. In some embodiments, the time between the first and second time points is in the range of about 15 minutes to about 30 years, for example, about 30 minutes, for example, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 or about 24 hours, for example, about 1, 2, 3, 4, 5, 10, 15, 20, 25 or about 50 days, or for example, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12 months, or for example, about 1, 1.5, 2, 2.5, 3 , 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5, 10, 10.5, 11, 11.5, 12, 12.5, 13, 13.5, 14, 14.5, 15, 15.5, 16, 16.5, 17, 17.5, 18, 18.5, 19, 19.5, 20, 20.5, 21, 21.5, 22, 22.5, 23, 23.5, 24, 24.5, 25, 25.5, 26, 26.5, 27, 27.5, 28, 28.5, 29, 29.5, or leave an interval of approximately 30 years. In other embodiments, test samples can be collected from patients at least once every 5 months, at least once every 6 months, at least once a year, at least once every 2 years, at least once every 3 years, at least once every 4 years, or at least once every 5 years.
[0289] IV.C. Treatment In yet another embodiment, cancer prediction can be used to make or influence clinical decisions (e.g., cancer diagnosis, treatment selection, evaluation of treatment effectiveness, etc.). For example, in one embodiment, if cancer prediction (e.g., with respect to cancer or a particular type of cancer) exceeds a threshold, a physician may prescribe appropriate treatment (e.g., surgical resection, radiotherapy, chemotherapy, and / or immunotherapy).
[0290] Using the classifiers described herein, it is possible to determine whether the feature vectors of a sample originate from a subject with cancer. In one embodiment, if the cancer prediction exceeds a threshold, an appropriate treatment (e.g., surgical resection or therapy) is prescribed. For example, in one embodiment, if the cancer prediction is 60 or higher, one or more appropriate treatments are prescribed. In another embodiment, if the cancer prediction is 65 or higher, 70 or higher, 75 or higher, 80 or higher, 85 or higher, 90 or higher, one or more appropriate treatments are prescribed. In yet another embodiment, the cancer prediction may indicate the severity of the disease. An appropriate treatment can be prescribed according to the severity of the disease.
[0291] In some embodiments, the treatment is one or more cancer therapeutic agents selected from the group consisting of chemotherapeutic agents, cancer-targeted therapeutic agents, differentiation-inducing therapeutic agents, hormone-based therapeutic agents, and immunotherapeutic agents. For example, the treatment may be one or more chemotherapeutic agents selected from the group consisting of alkylating agents, antimetabolites, anthracyclines, antitumor antibiotics, cytoskeletal disruptors (taxanes), topoisomerase inhibitors, mitotic inhibitors, corticosteroids, kinase inhibitors, nucleotide analogs, platinum-based drugs, and any combination thereof. In some embodiments, the treatment is one or more targeted cancer therapeutic agents selected from the group consisting of signaling inhibitors (e.g., tyrosine kinase inhibitors and growth factor receptor inhibitors), histone deacetylase (HDAC) inhibitors, retinoic acid receptor agonists, proteasome inhibitors, angiogenesis inhibitors, and monoclonal antibody conjugates. In some embodiments, the treatment is one or more differentiation-inducing therapeutic agents, such as retinoids such as tretinoin, alitretinoin, and bexarotene. In some embodiments, the treatment is one or more hormonal therapies selected from the group consisting of anti-estrogens, aromatase inhibitors, progestins, estrogens, anti-androgens, and GnRH agonists or analogs. In one embodiment, the treatment is one or more immunotherapies selected from the group consisting of monoclonal antibody therapies such as rituximab (RITUXAN) and alemtuzumab (CAMPATH), nonspecific immunotherapies and adjuvants (e.g., BCG, interleukin-2 (IL-2) and interferon-alpha), and immunomodulators, e.g., thalidomide and lenalidomide (REVLIMID). Selecting an appropriate cancer drug based on tumor type, cancer stage, history of previous cancer treatment or exposure to the drug, and other characteristics of the cancer is within the capabilities of a skilled physician or oncologist.
[0292] VD kit implementation configuration This specification also discloses kits for carrying out the methods described above, including methods relating to cancer classifiers. These kits may include one or more collection containers for collecting samples containing genetic material from an individual. Samples may include blood, plasma, serum, urine, feces, saliva, other types of body fluids, or combinations thereof. Such kits may include reagents for isolating nucleic acids from the samples. These reagents may further include reagents for sequencing nucleic acids, such as buffers and detection agents. In one or more embodiments, the kit may include one or more sequencing panels containing probes that target specific genomic regions, specific mutations, specific gene variants, or any combination thereof. In other embodiments, samples collected using the kit are provided to a sequencing laboratory where nucleic acids in the samples can be sequenced using the sequencing panels. WBC contamination detection can be applied to various configurations of the kit to minimize potential WBC contamination resulting from the kit's components. For example, experiments may be carried out comparing types of collection containers. WBC contamination can be evaluated and compared across different types of collection containers to identify the optimal type that minimizes WBC contamination.
[0293] The kit may further include instructions on how to use the reagents included in the kit. For example, the kit may include instructions on sample collection and extraction of nucleic acids from the test sample. Exemplary instructions may include the order of reagent addition, the centrifugation rate to be used to isolate nucleic acids from the test sample, the nucleic acid amplification method, the nucleic acid sequencing method, or any combination thereof. These instructions may further clarify how to operate the computer as the analysis system 200 for the purpose of carrying out any step of the described method.
[0294] In addition to the above elements, the kit may include a computer-readable storage medium storing computer software for performing the various methods described throughout this disclosure. These instructions may exist as information printed on a suitable medium or support (e.g., the packaging of the kit, as a package insert, a piece of paper on which the information is printed, etc.). Yet another means may be a computer-readable medium in which the instructions are stored in the form of computer code, such as a floppy disk, CD, hard drive, network data storage. Yet another means that may exist is a website address or QR code, which can be used via the Internet to access information about the removed site.
[0295] V.F. Detection and mitigation of sources of contamination In some embodiments, sample contamination is detected using the methods and / or classifiers of the present invention, such as by the various methodologies described in Sections II.C and II.D.
[0296] The analysis system may utilize a WBC contamination workflow to identify the source of contamination. To identify the source of contamination, the analysis system may isolate one or more variables of the sample processing workflow. The analysis system may process a first set of samples in a first sample processing workflow and a second set of samples in a second sample processing workflow, where the second sample workflow has different one or more variables to be evaluated. For example, the second sample processing workflow may include a different protocol, a different clinical product, or a different sequencing device. The protocol may include steps addressed in the processing of the sample, such as centrifugation, storage temperature, storage period, etc. A clinical product is a manufactured product used in the sample processing workflow and may include, for example, any container, chemical, compound, buffer, solution, enzyme, or any other product used in the workflow. A sequencing device generally includes a sequencer but may also include other devices related to the sequencing process. The sample processing workflow may further include other laboratory devices, such as centrifuges, storage devices, and other laboratory devices for sample processing.
[0297] The analysis system applies a contamination model to samples from both sample processing workflows. Based on the contamination results for each sample set, the analysis system can determine aggregate metrics for each sample processing workflow. If there is a significant difference in the aggregate metrics between the sample processing workflows, the analysis system can then identify the source of contamination based on which variables differed between the first and second sample processing workflows. Corrective measures may also be implemented.
[0298] Once the source of contamination is identified, the analysis system can determine the optimal sample handling workflow to mitigate WBC contamination. For example, repeated testing may determine a set of protocols that minimize contamination, a set of clinical products that minimize WBC contamination, a set of one or more sequencing devices, or any combination thereof, that minimizes WBC contamination. The optimal sample handling workflow can then be applied to subsequent samples.
[0299] V. Exemplary Results Collection and processing of VA samples Study Design and Samples: CCGA (NCT02889978) is a prospective, multicenter, case-control, observational study with long-term follow-up. De-identified biological samples were collected from approximately 15,000 participants from 342 centers. Samples were divided into training (1,785) sets and study (1,015) sets. Sample selection was made to ensure a pre-directed distribution of cancer and non-cancerous cells across the centers in each cohort, and cancer and non-cancerous samples were frequency-age matched by sex.
[0300] Whole-genome bisulfite sequencing: cfDNA was isolated from plasma, and whole-genome bisulfite sequencing (WGBS; 30x depth) was used for cfDNA analysis. Using the improved QIAamp Circulating nucleic acid kit (Qiagen; Germantown, MD), cfDNA was extracted from two plasma tubes (up to a combined volume of 10 ml) per patient. Up to 75 ng of plasma cfDNA was bisulfite converted using the EZ-96 DNA methylation kit (Zymo Research, D5003). Using the converted cfDNA, a dual-index sequencing library was prepared using the Accel-NGS Methyl-Seq DNA Library Preparation Kit (Swift BioSciences; Ann Arbor, MI), and the prepared library was quantified using the KAPA Library Quantification Kit for Illumina Platforms (Kapa Biosystems; Wilmington, MA). Four libraries were pooled together, including a 10% PhiX v3 library (Illumina, FC-110-3001), and clustered on an Illumina NovaSeq 7000S2 flow cell, followed by 150bp paired-end sequencing (30x).
[0301] For each sample, the WGBS fragment set was reduced to a small subset of fragments with informative methylation patterns. Furthermore, hypermethylated or hypomethylated cfDNA fragments were selected. cfDNA fragments were selected for having an informative methylation pattern and being hypermethylated or hypomethylated, i.e., UFXM. Fragments with methylation that occurs frequently or is unstable in cancer-free individuals are unlikely to provide highly discriminatory features in the classification of cancer status. Therefore, we generated statistical models and data structures of typical fragments from the CCGA study using an independent reference set (i.e., reference genome) of 108 cancer-free, non-smoking participants (age: 58 ± 14 years, 79 women [73%]). Using these samples, a Markov chain model (3-order) was trained to estimate the likelihood of given sequences of CpG methylation status within the fragments described above in Section II.C. This model was demonstrated to be calibrated within the normal fragment range (p-value > 0.001), and the model was used to exclude fragments with p-values >= 0.001 from the Markov model as insufficiently anomalous.
[0302] As described above, in the further data reduction step, only fragments that covered at least five CpGs and had an average methylation >0.9 (hypermethylation) or <0.1 (hypomethylation) were selected. This procedure yielded a median (range) of 2,800 (1,500–12,000) UFXM fragments for participants without cancer in training and a median (range) of 3,000 (1,200–420,000) UFXM fragments for participants with cancer in training. Since this data reduction procedure used only reference set data, this step only needed to be applied once to each sample.
[0303] VB sample contamination detection results Following one exemplary experiment, the contamination detection process was able to determine that 88% of singleton reads were dual-index-hopping events. This experiment was performed on 36 individual cfDNA samples using 36 unique sample barcodes. While all dual-index-hopping fragments are expected to be singletons, not all singletons are generated by double-index-hopping. The table below shows the number of sequence reads positively identified as being due to double-index-hopping.
[0304] [Table 1]
[0305] In total, there were 90 singleton sequence reads and 56 non-singleton sequence reads that were found to be contaminated. Contamination could be attributed to double index hopping, sequencing errors, or sample contamination during collection or processing. For singleton sequence reads, the contamination detection process using sample barcodes was able to positively identify 79 of the 90 contaminated reads as dual index hopping events. Using sample barcodes, the analysis system was able to return these reads to the appropriate samples.
[0306] VI. Additional Considerations The descriptions of embodiments detailed above are based on the drawings accompanying this specification, which illustrate specific embodiments as described herein. Other embodiments having different structures and operations do not depart from the scope of this disclosure. The terms “the present invention” or similar terms are used in reference to specific examples of many alternative aspects or embodiments of the applicant’s invention described herein, and neither use nor non-use is intended to limit the scope of the applicant’s invention or claims.
[0307] Embodiments of the present invention also relate to apparatus for carrying out the operations described herein. Such apparatus may include a general-purpose computer device that can be manufactured for a specific purpose and / or selectively operated or reconfigured by a computer program stored on the computer. Such a computer program may be stored on a non-temporary tangible computer-readable storage medium or any kind of medium suitable for storing electronic instructions, which may be connected to a computer system bus. Furthermore, any computing system referred to herein may include a single processor or may be an architecture that uses a multi-processor design for increased computing power.
[0308] The steps, operations, or processes described herein as being performed by an analysis system may be performed or implemented by one or more hardware or software modules of the device alone or in combination with other computing devices. In one embodiment, the software module is implemented in a computer program product including a computer-readable medium containing computer program code executable by a computer processor in order to perform all or part of the steps, operations, or processes described herein.< / atgcgacgc> < / atgctacgc> < / caccgtaa> < / atgctacgc>
Claims
1. Ligate one of a plurality of molecular identifiers (MIs) to the first end of each nucleic acid (NA) fragment of a first sample, wherein at least two of the plurality of MIs are different from each other. The NA fragment is amplified to generate an amplified NA fragment containing one or more copies of each NA fragment, Ligate the first sample barcode to the amplified NA fragment of the first sample, The process involves adding an index to the amplified NA fragment to generate an indexed NA fragment containing a first index and a second index, respectively. The indexed NA fragments are sequenced to generate sequence reads for each indexed NA fragment, To collect sequence reads of indexed NA fragments containing the first index sequence read and the second index sequence read within the group, By identifying a second sample barcode on the sequence read that is different from the first sample barcode, the contamination event of the first sequence read within the group can be identified. A method that includes this.
2. The method according to claim 1 or any dependent claim, wherein each amplified NA fragment having the first sample barcode includes a target NA region derived from the sample, and each amplified NA fragment having the second sample barcode includes a target NA region derived from a second sample different from the first sample.
3. The method according to claim 1 or any dependent claim, further comprising removing the first sequence reads having an index-hopping event from the group of the first sample.
4. The method according to claim 1 or any dependent claim, further comprising creating a binary cancer prediction between the presence and absence of cancer based on the sequence reads in the group from which the first sequence reads have been excluded.
5. The method according to claim 1 or any dependent claim, further comprising creating a multiclass cancer prediction among multiple cancer types based on the sequence reads in the group from which the first sequence reads have been excluded.
6. The method according to claim 1 or any dependent claim, wherein one of the plurality of MIs and the NA fragment are single-stranded during ligation.
7. The method according to claim 1 or any dependent claim, wherein the indexed NA fragment comprises the first index at the first end and the second index at the second end.
8. The method according to claim 1 or any dependent claim, wherein the first end is a 3' end.
9. The method according to claim 1 or any dependent claim, wherein the first end is a 5' end.
10. The method according to claim 1 or any dependent claim, wherein the first sample barcode is ligated to the second end of the amplified NA fragment opposite to the first end.
11. The method according to claim 1 or any dependent claim, wherein the first sample barcode is ligated to the first end of the amplified NA fragment adjacent to one of the plurality of MIs.
12. The method according to claim 1 or any dependent claim, wherein the first index is ligated to the first end, and the second index is ligated to the second end opposite to the first end.
13. The method according to claim 1 or any of the claims thereof, wherein the first index and the second index are ligated to the first end of the amplified NA fragment.
14. The method according to claim 1 or any dependent claim, wherein the first index and the second index are ligated to the second end of the amplified NA fragment opposite to the first end.
15. The method according to claim 1 or any dependent claim, wherein the sequencing is multiplexed with multiple samples across multiple flow cells, and the first index and the second index are used for a first column in which the first sample is present.
16. The method according to claim 1 or any dependent claim, wherein the third and fourth indexes, which are different from the first and second indexes, are used for the second column.
17. The method according to claim 1 or any dependent claim, wherein the sequencing comprises targeted methylation sequencing or whole-genome bisulfite sequencing.
18. The method according to claim 1 or any dependent claim, wherein each MI has a length selected from the range of 3 nucleic acid bases to 20 nucleic acid bases.
19. The method according to claim 1 or any dependent claim, wherein the first sample barcode has a length selected from the range of 3 nucleic acid bases to 20 nucleic acid bases.
20. The method according to claim 1 or any dependent claim, wherein the first sample barcode is substantially unique among a plurality of sample barcodes.
21. The method according to claim 1 or any dependent claim, wherein the NA fragment is a cell-free deoxyribonucleic acid (cfDNA) fragment.
22. The method according to claim 1 or any dependent claim, wherein the amplification of the NA fragment is performed by carrying out linear amplification.
23. A method for calling a contamination event, Receiving multiple sequence reads of amplified fragments, each of which includes a first index, a second index, and a sample barcode, Matching sequence reads of amplified fragments having a first index and a second index within a first bag of a first sample, wherein the first sample barcode is assigned to the first sample, and matching Based on identifying a second sample barcode on the sequence lead that is different from the first sample barcode, the contamination event of the first sequence lead in the first bag is called. A method that includes this.
24. The method according to claim 23 or any dependent claim, wherein the first index is located at the first end of the amplified segment, and the second index is located at the second end of the amplified segment.
25. The method according to claim 23 or any dependent claim, wherein the first index and the second index are located at the first end of the amplified fragment.
26. The method according to claim 23 or any dependent claim, wherein each of the amplified fragments further comprises a molecular identifier (MI), an amplified fragment derived from a first nucleic acid fragment of a sample comprises a first MI, and an amplified fragment derived from a second nucleic acid fragment of the sample comprises a second MI different from the first MI.
27. The method according to claim 23 or any dependent claim, wherein the MI is located at the first end of the amplified fragment, and the sample barcode is ligated onto the first end.
28. The method according to claim 23 or any dependent claim, wherein the MI is located at the first end of the amplified fragment, and the sample barcode is ligated onto the second end opposite to the first end.
29. The method according to claim 23 or any dependent claim, further comprising removing the first sequence read having the contamination event from the first bag of the first sample.
30. The method according to claim 29, further comprising making a cancer prediction for the first sample based on the sequence reads in the first bag from which the first sequence reads have been removed.
31. The method according to claim 30, wherein the cancer prediction is a binary prediction between the presence or absence of cancer.
32. The method according to claim 30 or any dependent claim, wherein the cancer prediction is a multi-class prediction among multiple cancer types.
33. The creation of the aforementioned cancer prediction is The aforementioned sequence reads are divided into separate sets of sequence reads, The objective is to determine the informative score for each individual sequence read, wherein the informative score indicates the likelihood of observing the individual sequence read in a healthy population of the sample. By comparing the informative score of the separate sequence reads against an informative score threshold, a set of informative fragments can be identified. The method according to claim 30 or any dependent claim, wherein the cancer prediction is based on the set of informative fragments.
34. The creation of the aforementioned cancer prediction is Based on the set of informative fragments, a feature vector of the first sample is created, The method involves inputting the feature vectors into a cancer classifier to determine the cancer prediction, wherein the cancer classifier is trained on at least a first cohort of cancer samples and a second cohort of non-cancer samples. The method according to claim 33, further comprising:
35. The method according to claim 23 or any dependent claim, wherein the first sample barcode has a length selected from the range of 3 nucleic acid bases to 20 nucleic acid bases.
36. A method for processing sequencing data, Receiving sequencing data comprising a set of sequence reads generated from multiplex sequencing of multiple biological samples, each containing nucleic acids, wherein the sequencing data includes data generated from single-index hopping and double-index hopping events occurring during the multiplex sequencing. The sequencing data is filtered to remove data corresponding to the single-index hopping event, and the filtering is performed by: Identifying one or more reads in the aforementioned set of sequence reads that have an index mismatch pair. The mismatch pairs of the index include two unique indices corresponding to two different biological samples, and the exclusion of such mismatch pairs includes two unique indices corresponding to two different biological samples. The sequencing data is filtered to remove data corresponding to the double-index hopping event, and the filtering is performed by: Identifying one or more pad-hopping duplicate reads in the set of sequence reads, wherein the pad-hopping duplicate reads have duplicate sequences that coexist in the flow cell used during the multiplex sequencing. Identifying one or more pad-hopping duplicate reads, followed by identifying one or more singletons in the set of sequence reads, wherein each singleton includes a unique sequence read from the set of sequence reads. Including the exclusion of A method that includes this.
37. Removing one or more identified pad-hopping duplicate leads from the set of sequencing leads, Following the removal of the identified one or more pad-hopping duplicate reads, the one or more singletons in the remaining set of sequencing reads are identified. The method according to claim 36 or any dependent claim, including the method described herein.
38. Filtering is, In the sequencing data, flag one or more identified reads having a mismatch index, one or more identified pad-hopping duplicate reads, or one or more identified singletons. The method according to claim 36 or any dependent claim, including the method described herein.
39. Filtering is, From the sequencing data, remove the one or more identified reads having a mismatch index, the one or more identified padhopping duplicate reads, or the one or more identified singletons. The method according to claim 36 or any dependent claim, including the method described herein.
40. The method according to claim 36 or any dependent claim, wherein the flow cell comprises a plurality of physically separated lanes, each lane comprising a plurality of columns, each column comprising a plurality of tiles, and each lane further defining a surface having a plurality of wells arranged thereon.
41. The method according to claim 36 or any dependent claim, wherein the sequence data includes spatial information indicating the position of each sequence read on the flow cell, and the spatial information includes at least one of a lane ID, a column ID, a tile ID, or an x-y coordinate pair.
42. Identifying pad-hopping duplicate reads is Identifying groups of identical or nearly identical sequence reads, Based on the sequence data, it is determined whether the grouped reads coexist, wherein the grouped reads have the following positional relationship: The grouped leads share a common tile. The grouped leads are located in a nearby tile. The grouped reads are located in different tiles within a common column. The grouped leads are located within a threshold x-distance and a threshold y-distance from each other on the flow cell, and The grouped leads are located within a predefined boundary region. They coexist if at least one of the following conditions is met: In accordance with the decision that the grouped reads coexist, the grouped reads are identified as pad-hopping duplicate reads. The method according to claim 36 or any dependent claim, including the method described herein.
43. The method according to claim 42 or any dependent claim, comprising identifying a plurality of groups of sequence reads as pad-hopping duplicate reads, each group comprising coexisting identical or substantially identical sequence reads.
44. The method according to claim 42 or any dependent claim, wherein the predefined boundary region includes a geometric shape having an x-distance of 7,500 flow cell position units and a y-distance of 100,000 flow cell position units.
45. The method according to claim 42 or any dependent claim, wherein the geometric shape includes a rectangle, and the longer side of the rectangle extends along the lane in the y-direction of the flow cell in the longitudinal direction.
46. The method according to claim 42 or any dependent claim, wherein the threshold x distance between the grouped leads is in the range of 0 to 50 mm, and the threshold y distance between the grouped leads is in the range of 0 to 50 mm.
47. The method according to claim 36 or any dependent claim, comprising removing the pad-hopping duplicate reads if the expected error rate associated with the multiplex sequencing exceeds a threshold error rate.
48. Provide the filtered sequencing data for analysis using a statistical model. The method according to claim 36 or any dependent claim, wherein the detection limit associated with the filtered sequencing data is lower than the detection limit associated with the unfiltered sequencing data.
49. The method according to claim 36 or any of the claims relating thereto, wherein the nucleic acid comprises cell-free DNA (cfDNA) or cell-free RNA (cfRNA).
50. The method according to claim 36 or any of the claims relating thereto, wherein the nucleic acid includes genomic DNA (gDNA).
51. The nucleic acids extracted from the aforementioned multiple biological samples are fragmented into genome fragments, The process involves ligating a unique dual index pair to the terminal portion of the genome fragment to generate multiple library fragments, wherein each unique dual index pair identifies and generates individual biological samples in the multiple biological samples. Enriching and amplifying a library fragment by capturing it with a targeted probe, and amplifying the captured library fragment in a plurality of wells on the flow cell, wherein each well is configured to hold a clone cluster of amplified fragments resulting from a single library fragment. The enriched fragment is sequenced to generate sequencing data, which includes a set of sequence reads, each of which includes a plurality of nucleotide base calls. Based on the aforementioned unique dual index pairs, the set of sequence reads is demultiplexed to determine the original biological sample for each sequence read. The method according to claim 36 or any dependent claim, including the method described herein.
52. The method according to claim 51, further comprising filtering the sequencing data, then demultiplexing the sequencing data to remove the data corresponding to the single-index-hopping event and the double-index-hopping event.
53. A method for processing sequencing data, Receiving sequencing data comprising a set of sequence reads generated from multiplex sequencing of multiple biological samples containing nucleic acids, wherein the sequencing data includes data generated from single-index hopping and double-index hopping events occurring during the multiplex sequencing. The sequencing data is filtered to remove data corresponding to the single-index hopping event, and the filtering is performed by: Identifying one or more reads in the aforementioned set of sequence reads that have an index mismatch pair. The mismatch pairs of the index include two unique indices corresponding to two different biological samples, and the exclusion of such mismatch pairs includes two unique indices corresponding to two different biological samples. The sequencing data is filtered to remove data corresponding to the double-index hopping event, and the filtering is performed by: Identifying one or more singletons in the aforementioned set of sequence reads. Each singleton includes a unique sequence read from the set of sequence reads, and excludes A method that includes this.
54. A method for training a cancer classifier, The method involves performing next-generation multiplex sequencing on a set of training samples, each possessing a known cancerous state, to obtain multiple sequence reads of amplified fragments, each sequence read containing a pair of ligated indices on a target region of the nucleic acid fragment and a sample barcode. Matching the sequence reads into a plurality of bags based on the pairs of indices of the sequence reads, wherein each bag contains sequence reads having a common pair of indices, In the first bag, the first sequence read having a first sample barcode different from other sequence reads is detected as a sample cross-contamination event, Removing the first array reads from the first bag, Assign the remaining sequence reads in each bag to one training sample, Based on the sequence reads in the corresponding bag, the feature vectors of each training sample are determined. Training a cancer classifier with the feature vectors of the training sample, wherein the trained cancer classifier is configured to predict the likelihood of cancer presence based on input feature vectors derived from sequence reads in the test sample. A method that includes this.
55. Each sequence read further comprises a molecular identifier (MI) ligated onto the corresponding nucleic acid fragment, and the method is Based on the aforementioned molecular identifier, the sequence reads in each bag are separated into individual sequence reads. The method according to claim 54 or any dependent claim, further comprising determining sequence reads having overlapping and matching molecular identifiers at similar genomic locations and reading them on the same nucleic acid fragment.
56. The method according to claim 54 or any dependent claim, further comprising assigning the first sequence reads to a second bag, wherein the sequence reads in the second bag have the first sample barcode.
57. The method according to claim 54 or any of the claims thereof, wherein the cancer classifier is a machine learning model.
58. The method according to claim 54 or any of the claims thereof, wherein the cancer classifier is trained to predict a binary prediction between the presence and absence of cancer.
59. The method according to claim 54 or any dependent claim, wherein each training sample is known to have one of several cancer types.
60. The method according to claim 59, wherein the cancer classifier is trained to predict multiclass predictions as the likelihood of the presence of one of the plurality of cancer types.
61. The method according to claim 54 or any dependent claim, wherein the sequence reads of each sample are methylated sequence reads, and the feature vectors are based on methylated sequence reads having an informative methylation pattern.
62. A non-temporary computer-readable storage medium for storing one or more programs configured to be executed by one or more processors of a computer system, wherein the one or more programs include instructions for carrying out the method according to any one of claims 1 to 61.
63. One or more processors, The non-temporary computer-readable storage medium according to claim 62 and A computer system including a computer system.
64. A collection container for collecting biological samples from the subject, Optionally, one or more reagents for isolating DNA fragments from the biological sample, Optionally, a library of paired indexes, Multiple copies of a first sample barcode for ligating onto the DNA fragment of the biological sample, Optionally, the non-temporary computer-readable storage medium described in claim 62 and A processing kit that includes this.
65. The first target NA region, A first sample-specific barcode including a first barcode sequence, and First Index A first NA structure including, Second target NA region, A second sample-specific barcode including a second barcode sequence, and The first index The second NA structure, which includes A set of nucleic acid (NA) constructs including, The first NA construct is identified as originating from the first sample based on the first sample-specific barcode, The second NA construct is a set of nucleic acid (NA) constructs identified as originating from the second sample based on the second sample-specific barcode.
66. The set of NA constructs according to claim 65 or any dependent claim, wherein the first sample-specific barcode and the first index are located at the first end of the first NA construct.
67. A set of NA constructs according to claim 65 or any dependent claim, wherein the first sample-specific barcode is located at a first end of the first NA construct, and the first index is located at a second end of the first NA construct opposite to the first end.
68. The set of NA constructs according to claim 65 or any dependent claim, wherein the first NA construct and the second NA construct are placed together in a bag based on identifying the first index.
69. A set of NA constructs according to claim 65 or any of the claims dependent thereon, wherein the first NA construct includes a second index, and the second NA construct includes the second index.
70. The set of NA structures according to claim 65 or any of the claims subordinating thereto, wherein the first index and the second index are located at the first end of the first NA structure.
71. A set of NA structures according to claim 65 or any dependent claim, wherein the first index is located at a first end of the first NA structure, and the second index is located at a second end of the first NA structure opposite to the first end.
72. The set of NA structures according to claim 65 or any of the claims thereof, wherein the first NA structure and the second NA structure are placed together in a bag based on identifying the first index and the second index.
73. A set of NA constructs according to claim 65 or any dependent claim, wherein the first barcode is located at a first end of the first NA construct, and the first NA construct further includes a molecular identifier (MI) at a second end of the first NA construct opposite to the first end.
74. A set of NA structures according to claim 65 or any of the claims relating thereto, wherein at least one of the first NA structure and the second NA structure is an amplification structure.
75. A set of NA constructs according to claim 65 or any dependent claim, wherein the first sample-specific barcode is substantially unique with respect to the second sample-specific barcode.
76. A set of NA constructs according to claim 65 or any of the claims dependent thereon, wherein the first NA construct is constructed in a first column, the second NA construct is constructed in a second column, and the first index is assigned to the first column but index-hopped to the second column.