Method and system for detecting contamination between samples

By receiving and processing the sequencing reads of tagged polynucleotides generated by the nucleic acid sequencer, using computer comparison and grouping, family identifiers are generated and pollution conditions are determined, the accuracy and false positive problems of sample pollution detection in the prior art are solved, and more efficient pollution detection is achieved.

CN120158499APending Publication Date: 2025-06-17GUARDANT HEALTH INC
View PDF 26 Cites 0 Cited by

Patent Information

Application Number
CN202510314430.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2018-08-30
Filing Date
2019-08-30
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect contamination between samples, especially in nucleic acid sequencing. Due to the small and variable molecular weight of the contamination, false positive results are easily generated.

Method used

By receiving more than one sequencing read of the set of tagged polynucleotides generated by the nucleic acid sequencer, the computer is used to compare and group, family identifiers are generated, and the sample is judged based on quantitative measurements.

Benefits of technology

It improves the accuracy of detection of sample contamination, reduces the occurrence of false positive results, and can effectively distinguish between pollution and non-pollution under high pollution rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120158499A_ABST
    Figure CN120158499A_ABST
Patent Text Reader

Abstract

The invention relates to a method and system for detecting contamination between samples. Provided herein are various methods and related systems for detecting the presence / absence of contamination of a first sample with a second sample. For example, in some embodiments, the method comprises (a) sequencing a set of polynucleotides to produce more than one sequenced read, (b) comparing the more than one sequenced read to a reference sequence, (c) grouping the more than one sequenced read into more than one family, (d) generating a family identifier for the more than one family, (e) screening out a set of common family identifiers, (f) determining a quantitative measure of the set of common family identifiers, and (g) classifying the first sample as contaminated or uncontaminated by the second sample based on the quantitative measure of common family identifiers.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of an application with an application date of August 30, 2019, an application number of 201980072064.3, and an invention title of "Methods and Systems for Detecting Contamination between Samples".

[0002] Cross - References

[0003] This application claims the benefit and priority of U.S. Provisional Application No. 62 / 724,622, filed on August 30, 2018, which is hereby incorporated by reference in its entirety. Background

[0004] Cancer is typically caused by the accumulation of mutations within the normal cells of an individual, at least some of which result in improper regulation of cell division. Such mutations typically include single - nucleotide variants (SNVs), gene fusions, insertions and deletions (indels), transversions, translocations, and inversions.

[0005] Cancer is typically detected by tissue biopsy of a tumor followed by analysis of cell pathology, biomarkers extracted from the cells, or DNA. However, it has recently been proposed that cancer can also be detected by cell - free nucleic acids (e.g., circulating nucleic acids, circulating tumor nucleic acids, exosomes, nucleic acids from apoptotic and / or necrotic cells) in body fluids such as blood or urine (see, e.g., Siravegna et al., Nature Reviews, 14:531 - 548 (2017)). Such tests have the advantage that they are non - invasive and can be performed without a biopsy to identify suspicious cancer cells and sample nucleic acids from all parts of the cancer. However, due to the fact that the amount of nucleic acid released into body fluids is low and variable, such tests are complex, as is the recovery of nucleic acids from such body fluids in an analyzable form. These tests are designed such that they can detect very low - frequency sequences, represented by as few as 1 in 1000 molecules at a given locus. Thus, such tests may be prone to false - positive results based on low - level molecular contamination from other samples.

[0006] Samples can be contaminated by various sources such as, but not limited to: physical carryover of liquid between samples (e.g., pipetting, automated liquid handling via sample prep or sequencer, manipulation of amplified material); demultiplexing artifacts (e.g., base call errors that confound sample indices with limited pairwise Hamming distance; insertions / deletions that confound sample indices with limited pairwise edit distance); and reagent impurities (e.g., sample index oligonucleotides with a certain degree of loss of oligonucleotides synthesized in the same batch; sample index oligonucleotides contaminated with oligonucleotides containing another sample index (by legacy of synthesis error)).

[0007] Overview

[0008] This application discloses methods and systems for detecting contamination between two samples. Previous sample contamination detection methods were based on the detection of certain molecules that could only be present at high abundance or not at all in an uncontaminated sample, but if observed at low abundance, indicated contamination. Two such types of molecules are molecules carrying common germline single nucleotide polymorphisms (SNPs) or Y chromosome molecules. These methods are limited by the fact that the above molecules are usually only a small fraction of all contaminating molecules, and their amounts may not be sufficient for detection in the presence of sequencing errors and sampling errors. In addition, in cases of high contamination rates, contaminant germline SNVs may be indistinguishable from the germline SNVs native to the contaminated sample. Since Y chromosome molecules are naturally only present in male patients, using Y chromosome molecules as a detection mechanism is further limited to cases where female patient samples are contaminated by male patient samples. In addition to physical contamination, digital cross-contamination can occur when sample indices are easily converted to another index that is subsequently misassigned algorithmically. This problem can be mitigated by dual indexing, but this method has its own drawbacks.

[0009] The present disclosure provides methods, compositions, and systems for detecting the presence or absence of contamination of a first sample by a second sample.

[0010] In one aspect, the present disclosure provides a system for detecting the presence or absence of contamination of a first sample by a second sample, the system comprising: a communication interface that receives, via a communication network, more than one sequencing read of a set of tagged polynucleotides from a sample generated by a nucleic acid sequencer, wherein the sequencing read comprises a tag sequence and a sequence derived from the polynucleotide; and a computer in communication with the communication interface, wherein the computer comprises one or more computer processors and a computer-readable medium comprising machine-executable code that, when executed by the one or more computer processors, implements a method that includes: (a) receiving, via the communication network, more than one sequencing read of a set of tagged polynucleotides from a sample generated by a nucleic acid sequencer; (b) aligning the more than one sequencing read with a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping the more than one sequencing read into more than one family based on grouping features that include at least one of the following (i), (ii), (iii), and (iv): (i) a tag, (ii) a start region, (iii) an end region, and (iv) the length of the polynucleotide, wherein each family in the sample comprises sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the set of polynucleotides in the sample; (d) generating family identifiers for the more than one family; (e) screening out a set of common family identifiers, wherein a given common family identifier is a family identifier of the first sample that is the same as or substantially the same as a family identifier of the second sample; (f) determining a quantitative measure of the set of common family identifiers; and (g) classifying the first sample as contaminated by the second sample if the quantitative measure of the set of common family identifiers is higher than a predetermined threshold, or classifying the first sample as not contaminated by the second sample if the quantitative measure of the set of common family identifiers is at or below the predetermined threshold.

[0011] In another aspect, the present disclosure provides a system that includes a controller, the controller including or having access to a computer-readable medium that includes non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform a method that includes: (a) sequencing a collection of polynucleotides from a first sample and a second sample to produce more than one sequencing read; (b) aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping the more than one sequencing read into more than one family based on grouping characteristics that include at least one of the following (i), (ii), and (iii): (i) the start region, (ii) the end region, and (iii) the length of the polynucleotide, wherein each family in a sample includes sequencing reads of daughter polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides in the sample; (d) generating family identifiers for the more than one family; (e) screening for a set of common family identifiers, wherein a given common family identifier is a family identifier of the first sample that is the same as or substantially the same as a family identifier of the second sample; (f) determining a quantitative measure of the set of common family identifiers; and (g) classifying the first sample as contaminated by the second sample if the quantitative measure of the set of common family identifiers is above a predetermined threshold, or classifying the first sample as not contaminated by the second sample if the quantitative measure of the set of common family identifiers is at or below the predetermined threshold.

[0012] In another aspect, the present disclosure provides a system that includes a controller that includes or has access to a computer-readable medium that includes non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform a method that includes: (a) sequencing a collection of polynucleotides from a sample to produce more than one sequencing read; (b) aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) grouping the more than one sequencing read from two samples together into more than one family based on grouping characteristics that include at least one of (i), (ii), and (iii) below: (i) the start region, (ii) the end region, and (iii) the length of the polynucleotide, wherein each family in a sample includes sequencing reads of progeny polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides in the sample; (d) screening the more than one family to identify a set of common families, wherein a given common family includes at least one sequencing read from a first sample and at least one sequencing read from a second sample; (e) determining a quantitative measure derived from the set of common families; (f) classifying the first sample as contaminated by the second sample if the quantitative measure of the set of common families is above a predetermined threshold, or classifying the first sample as not contaminated by the second sample if the quantitative measure of the set of common families is at or below the predetermined threshold.

[0013] In another aspect, the present disclosure provides a system that includes a controller that includes or has access to a computer-readable medium that includes non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform a method that includes: (a) sequencing a collection of tagged polynucleotides from a sample to produce more than one sequencing read, wherein each tagged polynucleotide includes a tag and a polynucleotide; (b) aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping the more than one sequencing read into more than one family based on grouping characteristics that include the tag, wherein each family in a sample includes sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the collection of tagged polynucleotides in the sample; (e) screening out a set of common family identifiers, wherein a given common family identifier is a family identifier of a first sample that is the same as or substantially the same as a family identifier of a second sample; (f) determining a quantitative measure of the set of common family identifiers; and (g) classifying the first sample as contaminated by the second sample if the quantitative measure of the set of common family identifiers is above a predetermined threshold, or classifying the first sample as not contaminated by the second sample if the quantitative measure of the common family identifier is at or below the predetermined threshold.

[0014] In another aspect, the present disclosure provides a system that includes a controller that includes or has access to a computer-readable medium containing non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform a method that includes: (a) sequencing a collection of polynucleotides from a sample to generate more than one sequencing read; (b) aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping the more than one sequencing read into more than one family based on grouping characteristics that include at least one of the following (i), (ii), and (iii): (i) the start region, (ii) the end region, and (iii) the length of the polynucleotide, wherein each family in the sample includes sequencing reads of daughter polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides in the sample; (d) screening the more than one family to identify a set of common families, wherein a given common family is a family in a first sample that has grouping characteristics that are the same as or substantially the same as the grouping characteristics of a family in a second sample; (e) determining a quantitative measure of the set of common families in the first sample; and (f) classifying the first sample as contaminated by the second sample if the quantitative measure of the set of common families is above a predetermined threshold, or classifying the first sample as not contaminated by the second sample if the quantitative measure of the set of common families is at or below the predetermined threshold.

[0015] In some embodiments, the sequencing reads include (i) a tag sequence and (ii) a sequence derived from a polynucleotide. In some embodiments, the system further includes, for each sample, grouping the more than one sequencing read into more than one family based on information from at least one of the following (i), (ii), (iii), and (iv): (i) tags, (ii) start regions, (iii) end regions, and (iv) the length of the polynucleotide, wherein each family in the sample includes sequencing reads of daughter polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides in the sample.

[0016] In another aspect, the present disclosure provides a system that includes a controller that includes or has access to a computer-readable medium that contains non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform a method that includes: (a) sequencing a collection of tagged polynucleotides from a sample to produce more than one sequencing read, where each tagged polynucleotide includes a tag and a polynucleotide; (b) aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping the more than one sequencing read into more than one family based on grouping characteristics that include the tag, where each family in a sample includes sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the collection of tagged polynucleotides from the sample; (d) screening the more than one family to identify a set of common families, where a given common family is a family in a first sample that has grouping characteristics that are the same or substantially the same as the grouping characteristics of a family in a second sample; (e) determining a quantitative measure of the set of common families in the first sample; and (f) classifying the first sample as contaminated by the second sample if the quantitative measure of the set of common families is above a predetermined threshold, or classifying the first sample as not contaminated by the second sample if the quantitative measure of the set of common families is at or below the predetermined threshold.

[0017] In another aspect, the present disclosure provides a system that includes a controller that includes or has access to a computer-readable medium that contains non-transitory computer-executable instructions that, when executed by at least one electronic processor, perform a method that includes: (a) sequencing a collection of tagged polynucleotides from a sample to produce more than one sequencing read, where each tagged polynucleotide includes a tag and a polynucleotide; (b) aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) grouping the more than one sequencing read from two samples together into more than one family based on grouping characteristics that include the tag, where each family in a sample includes sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the collection of tagged polynucleotides from the sample; (d) screening the more than one family to identify a set of common families, where a given common family includes at least one sequencing read from a first sample and at least one sequencing read from a second sample; (e) determining a quantitative measure of the set of common families; (f) classifying the first sample as contaminated by the second sample if the quantitative measure of the set of common families is above a predetermined threshold, or classifying the first sample as not contaminated by the second sample if the quantitative measure of the common families is at or below the predetermined threshold.

[0018] In some embodiments, the system further includes detecting somatic genetic variations of polynucleotides in a first sample by excluding sequencing reads of a shared family of the first sample, wherein the first sample is classified as being contaminated by a second sample.

[0019] In some embodiments, the system further includes generating a report, which optionally includes information about the contamination status of the sample and / or information derived from the contamination status of the sample.

[0020] In some embodiments, the system further includes transmitting the report to a third party, such as a subject from whom the sample is derived or a health care practitioner.

[0021] In another aspect, the present disclosure provides a method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising: (a) sequencing a collection of polynucleotides from the sample to produce more than one sequencing read; (b) aligning the more than one sequencing read with a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping the more than one sequencing read into more than one family based on grouping features, the grouping features including at least one of the following (i), (ii), and (iii): (i) the start region, (ii) the end region, and (iii) the length of the polynucleotide, wherein each family in the sample includes sequencing reads of progeny polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides in the sample; (d) generating family identifiers for the more than one family; (e) screening out a set of shared family identifiers, wherein a given shared family identifier is a family identifier of the first sample that is the same as or substantially the same as the family identifier of the second sample; (f) determining a quantitative measure of the set of shared family identifiers; and (g) classifying the first sample as being contaminated by the second sample if the quantitative measure of the set of shared family identifiers is higher than a predetermined threshold, or classifying the first sample as not being contaminated by the second sample if the quantitative measure of the set of shared family identifiers is at or below the predetermined threshold.

[0022] In another aspect, the present disclosure provides a method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising: (a) accessing, by a computer system, sequence information comprising more than one sequencing read from the first sample and the second sample; (b) aligning, by the computer system, the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping, by the computer system, the more than one sequencing read into more than one family based on grouping features, the grouping features including at least one of the following (i), (ii), and (iii): (i) start region, (ii) end region, and (iii) length of the polynucleotide, wherein each family in the sample comprises sequencing reads of progeny polynucleotides amplified from a unique polynucleotide in the set of polynucleotides in the sample; (d) generating, by the computer system, family identifiers for the more than one family; (e) screening, by the computer system, for a set of common family identifiers, wherein a given common family identifier is a family identifier of the first sample that is the same as or substantially the same as the family identifier of the second sample; (f) determining, by the computer system, a quantitative measure of the set of common family identifiers; and (g) classifying, by the computer system, the first sample as contaminated by the second sample if the quantitative measure of the set of common family identifiers is higher than a predetermined threshold, or classifying, by the computer system, the first sample as not contaminated by the second sample if the quantitative measure of the set of common family identifiers is at or below the predetermined threshold.

[0023] In another aspect, the present disclosure provides a method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising: (a) obtaining sequence information comprising more than one sequencing read from the first sample and the second sample; (b) aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping the more than one sequencing read into more than one family based on grouping features, the grouping features including at least one of the following (i), (ii), and (iii): (i) start region, (ii) end region, and (iii) length of the polynucleotide, wherein each family in the sample comprises sequencing reads of progeny polynucleotides amplified from a unique polynucleotide in the set of polynucleotides in the sample; (d) generating family identifiers for the more than one family; (e) screening for a set of common family identifiers, wherein a given common family identifier is a family identifier of the first sample that is the same as or substantially the same as the family identifier of the second sample; (f) determining a quantitative measure of the set of common family identifiers; and (g) classifying the first sample as contaminated by the second sample if the quantitative measure of the set of common family identifiers is higher than a predetermined threshold, or classifying the first sample as not contaminated by the second sample if the quantitative measure of the set of common family identifiers is at or below the predetermined threshold.

[0024] In some embodiments, the method further comprises, prior to a), tagging a collection of polynucleotides to produce tagged polynucleotides, wherein each tagged polynucleotide comprises a tag and a polynucleotide. In some embodiments, the method further comprises, for each sample, grouping more than one sequencing read into more than one family based on grouping features, the grouping features comprising at least one of the following (i), (ii), (iii), and (iv): (i) a tag, (ii) a start region, (iii) an end region, and (iv) the length of the polynucleotide, wherein each family in the sample comprises sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides in the sample.

[0025] In another aspect, the present disclosure provides a method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising: (a) sequencing a collection of tagged polynucleotides or polynucleotides from a sample to produce more than one sequencing read, wherein each tagged polynucleotide comprises a tag and a polynucleotide; (b) aligning the more than one sequencing read with a reference sequence, thereby determining an alignment start region and an alignment end region; (c) for each sample, grouping more than one sequencing read into more than one family based on grouping features, the grouping features comprising a tag, wherein each family in the sample comprises sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the collection of tagged polynucleotides in the sample; (d) generating family identifiers for the more than one family; (e) screening out a set of common family identifiers, wherein a given common family identifier is a family identifier of the first sample that is the same as or substantially the same as a family identifier of the second sample; (f) determining a quantitative measure of the set of common family identifiers; and (g) classifying the first sample as contaminated by the second sample if the quantitative measure of the set of common family identifiers is higher than a predetermined threshold, or classifying the first sample as not contaminated by the second sample if the quantitative measure of the set of common family identifiers is at or below the predetermined threshold.

[0026] In some embodiments, a quantitative measure of the set of shared family identifiers is the number of shared family identifiers in the first sample. In some embodiments, a quantitative measure of the set of shared family identifiers includes the ratio of the number of shared family identifiers in the first sample to the total number of family identifiers in the first sample. In some embodiments, a quantitative measure of the set of shared family identifiers does not include the following shared identifiers in the first sample: those shared family identifiers in the family of the first sample for which the number of sequencing reads is greater than the number of sequencing reads in the corresponding family of the second sample. In some embodiments, a quantitative measure of the set of shared family identifiers in the first sample does not include shared family identifiers at over-represented genomic start and genomic end position pairs. In some embodiments, the total number of family identifiers in the first sample does not include family identifiers at over-represented genomic start and genomic end position pairs.

[0027] In some embodiments, over-represented genomic start and genomic end position pairs are determined by: (a) providing more than one sample, where the more than one sample includes a distribution of genomic start and genomic end positions that is the same as or substantially the same as that of the first sample and / or the second sample; (b) determining the family identifiers in the more than one sample; (c) quantifying the number of family identifiers that share a pair of genomic start and genomic end positions in the more than one sample; and (d) classifying the pair of genomic start and genomic end positions as over-represented if the number of family identifiers exceeds a set threshold. In some embodiments, the more than one sample does not include the first sample or the second sample. In some embodiments, the more than one sample does not include the first sample and the second sample. In some embodiments, the more than one sample includes samples processed in the same flow cell as the first sample. In some embodiments, the more than one sample includes training samples. In some embodiments, the set threshold is at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, or at least 60 families.

[0028] In another aspect, the present disclosure provides a method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising: (a) sequencing a collection of polynucleotides from the sample to generate more than one sequencing read; (b) aligning the more than one sequencing read with a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping the more than one sequencing read into more than one family based on grouping features, the grouping features including at least one of the following (i), (ii), and (iii): (i) the start region, (ii) the end region, and (iii) the length of the polynucleotide, wherein each family in the sample includes sequencing reads of daughter polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides from the sample; (d) screening the more than one family to identify a set of common families, wherein a given common family is a family in the first sample having grouping features that are the same or substantially the same as the grouping features of a family in the second sample; (e) determining a quantitative measure of the set of common families in the first sample; and (f) classifying the first sample as contaminated by the second sample if the quantitative measure of the set of common families is above a predetermined threshold, or classifying the first sample as not contaminated by the second sample if the quantitative measure of the set of common families is at or below the predetermined threshold.

[0029] In another aspect, the present disclosure provides a method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising: (a) sequencing a collection of polynucleotides from the sample to generate more than one sequencing read; (b) aligning the more than one sequencing read with a reference sequence, thereby determining a start region and an end region of the alignment; (c) grouping the more than one sequencing read from the two samples together into more than one family based on grouping features, the grouping features including at least one of the following (i), (ii), and (iii): (i) the start region, (ii) the end region, and (iii) the length of the polynucleotide, wherein each family includes sequencing reads of daughter polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides; (d) screening the more than one family to identify a set of common families; wherein the common families include at least one sequencing read from the first sample and at least one sequencing read from the second sample; (e) determining a quantitative measure of the set of common families; and (f) classifying the first sample as contaminated by the second sample if the quantitative measure of the set of common families is above a predetermined threshold, or classifying the first sample as not contaminated by the second sample if the quantitative measure of the set of common families is at or below the predetermined threshold.

[0030] In some embodiments, the method further comprises, prior to sequencing, tagging the collection of polynucleotides to produce tagged polynucleotides, wherein each tagged polynucleotide comprises a tag and a polynucleotide.

[0031] In some embodiments, the method includes, for each sample, grouping more than one sequencing read into more than one family based on grouping features, the grouping features including at least one of the following (i), (ii), (iii), and (iv): (i) a tag, (ii) a start region, (iii) an end region, and (iv) the length of a polynucleotide, wherein each family in the sample includes sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the set of polynucleotides in the sample.

[0032] In another aspect, the present disclosure provides a method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising: (a) sequencing a set of tagged polynucleotides from the sample to produce more than one sequencing read, wherein each tagged polynucleotide comprises a tag and a polynucleotide; (b) aligning the more than one sequencing read with a reference sequence, thereby determining an alignment start region and an alignment end region; (c) for each sample, grouping more than one sequencing read into more than one family based on grouping features, the grouping features including a tag, wherein a family in the sample includes sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the set of tagged polynucleotides in the sample; (d) screening the more than one family to identify a set of common families, wherein a given common family is a family of the first sample having grouping features that are the same as or substantially the same as the grouping features of a family of the second sample; (e) determining a quantitative measure of the set of common families of the first sample; and (f) classifying the first sample as contaminated by the second sample if the quantitative measure of the set of common families is above a predetermined threshold, or classifying the first sample as not contaminated by the second sample if the quantitative measure of the set of common families is at or below the predetermined threshold.

[0033] In another aspect, the present disclosure provides a method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising: (a) sequencing a collection of tagged polynucleotides from the sample to generate more than one sequencing read, wherein each tagged polynucleotide comprises a tag and a polynucleotide; (b) aligning the more than one sequencing read with a reference sequence, thereby determining a start region and an end region of the alignment; (c) grouping the more than one sequencing read of the two samples together into more than one family based on grouping features, the grouping features including tags, wherein each family in the sample comprises sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the collection of tagged polynucleotides from the sample; (d) screening the more than one family to identify a set of common families, wherein a given common family comprises at least one sequencing read from the first sample and at least one sequencing read from the second sample; (e) determining a quantitative measure derived from the set of common families; and (f) classifying the first sample as contaminated by the second sample if the quantitative measure of the set of common families is higher than a predetermined threshold, or classifying the first sample as not contaminated by the second sample if the quantitative measure of the set of common families is at or below the predetermined threshold.

[0034] In some embodiments, the quantitative measure includes the number of shared families in a first sample. In some embodiments, the quantitative measure includes the ratio of the number of sequencing reads of the first sample to the number of sequencing reads of the second sample in a shared family. In some embodiments, the quantitative measure includes the ratio of the number of shared families in a first sample to the total number of families in the first sample. In some embodiments, the quantitative measure of the set of shared families does not include the following shared families in the first sample: those shared families in the families of the first sample whose number of sequencing reads is greater than the number of sequencing reads of the corresponding family in the second sample. In some embodiments, the quantitative measure of the set of shared families in the first sample does not include the shared families at over-represented genomic start and genomic end position pairs. In some embodiments, the total number of families in the first sample does not include the families at over-represented genomic start and genomic end position pairs. In some embodiments, over-represented genomic start and genomic end position pairs are determined by: (a) providing more than one sample, where the more than one sample includes a distribution of genomic start and genomic end positions that is the same as or substantially the same as that of the first sample and / or the second sample; (b) determining the families in the more than one sample; (c) quantifying the number of families in the more than one sample that share a pair of genomic start and genomic end positions; and (d) classifying the pair of genomic start and genomic end positions as over-represented if the number of families exceeds a set threshold. In some embodiments, the more than one sample does not include the first sample or the second sample. In some embodiments, the more than one sample does not include the first sample and the second sample. In some embodiments, the more than one sample includes samples processed in the same flow cell as the first sample. In some embodiments, the more than one sample includes training samples. In some embodiments, the set threshold is at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, or at least 60 families. In some embodiments, the set threshold is about 5 families. In some embodiments, the set threshold is about 10 families. In some embodiments, the set threshold is about 15 families. In some embodiments, the set threshold is about 20 families. In some embodiments, the set threshold is about 30 families. In some embodiments, the set threshold is about 40 families. In some embodiments, the set threshold is about 50 families. In some embodiments, the set threshold can be at least 10 of the total families observed in the more than one sample -3 , at least 10 -4 , at least 10 -5 , at least 10 -6 , at least 10-7 、 at least 10 -8 or at least 10 -9 。 In some embodiments, the set threshold can be about 10 of the total families observed in more than one sample -4 。 In some embodiments, the set threshold can be about 10 of the total families observed in more than one sample -5 。 In some embodiments, the set threshold can be about 10 of the total families observed in more than one sample -6 。 In some embodiments, the set threshold can be about 10 of the total families observed in more than one sample -7 。 In some embodiments, the set threshold can be about 10 of the total families observed in more than one sample -8 。

[0035] In some embodiments, the start region includes the genomic start position of the sequencing read, at which the 5' end of the sequencing read is determined to start aligning with the reference sequence, and the end region includes the genomic end position of the sequencing read, at which the 3' end of the sequencing read is determined to stop aligning with the reference sequence. In some embodiments, the start region includes the first 1, first 2, first 5, first 10, first 15, first 20, first 25, first 30, or at least the first 30 base positions of the 5' end of the sequencing read aligned with the reference sequence. In some embodiments, the end region includes the last 1, last 2, last 5, last 10, last 15, last 20, last 25, last 30, or at least the last 30 base positions of the 3' end of the sequencing read aligned with the reference sequence.

[0036] In some embodiments, the tag includes one or more molecular barcodes attached to the end of the polynucleotide. In some embodiments, the length of one or more molecular barcodes is at least 2, at least 4, at least 5, at least 6, at least 8, at least 10, at least 15, or at least 20 nucleotides. In some embodiments, one or more molecular barcodes attached to the polynucleotide of the first sample are different from one or more molecular barcodes attached to the polynucleotide of the second sample. In some embodiments, the polynucleotide of the sample is tagged with at least 5, at least 10, at least 15, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10,000, at least 50,000, or at least 100,000 different molecular barcodes.

[0037] In some embodiments, the first sample and the second sample are sequenced in the same flow cell. In some embodiments, the second sample is sequenced in a flow cell different from the first sample. In some embodiments, the second sample is processed on the same day as the first sample, but at a different time from the first sample. In some embodiments, the second sample is processed at least 1 minute, at least 30 minutes, at least 1 hour, at least 2 hours, at least 3 hours, or at least 4 hours after processing the first sample. In some embodiments, the first sample and the second sample are processed on different dates. In some embodiments, the first sample and the second sample are in the same sample batch. In some embodiments, the second sample is processed with the same batch of reagents as the first sample. In some embodiments, the first sample and the second sample are processed at different geographical locations.

[0038] In some embodiments, the set of tagged polynucleotides of the sample is uniquely tagged. In some embodiments, the set of tagged polynucleotides of the sample is non-uniquely tagged. In some embodiments, the first sample is obtained from the body fluid of one subject, and the second sample is obtained from the body fluid of another subject.

[0039] In some embodiments, the polynucleotide is a cell-free polynucleotide. In some embodiments, the cell-free polynucleotide is cell-free DNA. In some embodiments, at least one of the subjects has a disease. In some embodiments, the disease is cancer.

[0040] In some embodiments, a collection of polynucleotides of a sample is amplified prior to sequencing, thereby generating amplified progeny polynucleotides. In some embodiments, the method further comprises selectively enriching at least a portion of the amplified progeny polynucleotides from regions of the genome or transcriptome of a subject prior to sequencing. In some embodiments, the method further comprises attaching one or more sample indices to one or both ends of the amplified progeny polynucleotides prior to sequencing, wherein the sample indices distinguish a first sample from a second sample. In some embodiments, the predetermined threshold is at least 0.001%, at least 0.005%, at least 0.01%, at least 0.05%, at least 0.1%, at least 0.5%, at least 1%, at least 2%, at least 5%, or at least 10% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 0.01% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 0.05% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 0.1% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 0.5% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 1% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 2% of the total number of families in the first sample.

[0041] In some embodiments, the method further comprises detecting somatic genetic variations of the polynucleotides of the first sample by excluding sequencing reads of the shared family identifiers of the first sample, wherein the first sample is classified as being contaminated with the second sample. In some embodiments, the method further comprises detecting somatic genetic variations of the polynucleotides of the first sample by excluding sequencing reads of the shared family of the first sample, wherein the first sample is classified as being contaminated with the second sample.

[0042] In some embodiments, the method further comprises generating a report, which optionally includes information about the contamination status of the sample and / or information derived from the contamination status of the sample. In some embodiments, the method includes transmitting the report to a third party, such as the subject from whom the sample is derived or a health care practitioner.

[0043] The embodiments described herein can be used in or applied to both the methods and systems described herein.

[0044] In some embodiments, the results of the systems and / or methods disclosed herein are used as input to generate a report. The report can be in paper format or electronic format. For example, information regarding the contamination status of a first sample as determined by the methods or systems disclosed herein and / or information derived from the contamination status of the first sample can be presented in such a report. The methods or systems disclosed herein can also include the step of transmitting the report to a third party, such as the subject from whom the sample was derived or a health care practitioner.

[0045] The various steps of the methods disclosed herein, or the steps performed by the systems disclosed herein, can be performed at the same time or different times, and / or at the same geographical location or different geographical locations (e.g., countries). The various steps of the methods disclosed herein can be performed by the same person or different persons.

[0046] In certain aspects, the present disclosure provides a non-transitory computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, can perform one or more of the steps or methods described herein.

[0047] In another aspect, the present disclosure provides a non-transitory computer-readable medium comprising non-transitory computer-executable instructions that, when executed by at least one electronic processor, can perform at least the following: (a) obtaining more than one sequencing read from a set of tagged polynucleotides from a sample generated by a nucleic acid sequencer; (b) aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping the more than one sequencing read into more than one family based on grouping characteristics, the grouping characteristics including at least one of the following (i), (ii), (iii), and (iv): (i) a tag, (ii) a start region, (iii) an end region, and (iv) the length of the polynucleotide, wherein each family in the sample includes sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the set of polynucleotides in the sample; (d) generating family identifiers for the more than one family; (e) screening out a set of common family identifiers, wherein a given common family identifier is a family identifier of a first sample that is the same or substantially the same as a family identifier of a second sample; (f) determining a quantitative measure of the set of common family identifiers; and (g) classifying the first sample as contaminated by the second sample if the quantitative measure of the set of common family identifiers is above a predetermined threshold, or classifying the first sample as not contaminated by the second sample if the quantitative measure of the set of common family identifiers is at or below the predetermined threshold.

[0048] In some aspects, the methods, systems, and / or computer-readable media described herein can be used as quality control metrics for determining performance and / or for assessing the quality of the obtained sequencing data to ensure reliable detection of somatic variants in a sample.

[0049] This application provides the following:

[0050] 1. A method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising:

[0051] (a) Sequencing a collection of polynucleotides from the sample to generate more than one sequencing read;

[0052] (b) Aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment;

[0053] (c) For each sample, grouping the more than one sequencing read into more than one family based on grouping features, the grouping features including at least one of the following (i), (ii), and (iii): (i) the start region, (ii) the end region, and (iii) the length of the polynucleotide, wherein each family in the sample includes sequencing reads of progeny polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides from the sample;

[0054] (d) Generating family identifiers for the more than one family;

[0055] (e) Screening out a set of common family identifiers, wherein a given common family identifier is a family identifier of the first sample that is the same as or substantially the same as a family identifier of the second sample;

[0056] (f) Determining a quantitative measure of the set of common family identifiers; and

[0057] (g) If the quantitative measure of the set of common family identifiers is higher than a predetermined threshold, classifying the first sample as contaminated by the second sample, or if the quantitative measure of the set of common family identifiers is at or below the predetermined threshold, classifying the first sample as not contaminated by the second sample.

[0058] 2. A method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising:

[0059] (a) Accessing, by a computer system, sequence information comprising more than one sequencing read from the first sample and the second sample;

[0060] (b) Align the more than one sequencing reads with a reference sequence by means of the computer system, thereby determining a start region and an end region of the alignment;

[0061] (c) For each sample, group the more than one sequencing reads into more than one family by means of the computer system based on grouping features, the grouping features including at least one of the following (i), (ii), and (iii): (i) the start region, (ii) the end region, and (iii) the length of the polynucleotide, wherein each family in the sample includes sequencing reads of daughter polynucleotides amplified from a unique polynucleotide in the set of polynucleotides of the sample

[0062] (d) Generate family identifiers for the more than one family by means of the computer system;

[0063] (e) Screen out a set of common family identifiers by means of the computer system, wherein a given common family identifier is a family identifier of the first sample that is the same as or substantially the same as the family identifier of the second sample;

[0064] (f) Determine a quantitative measure of the set of common family identifiers by means of the computer system; and

[0065] (g) If the quantitative measure of the set of common family identifiers is higher than a predetermined threshold, classify the first sample as contaminated by the second sample by means of the computer system, or if the quantitative measure of the set of common family identifiers is at or below the predetermined threshold, classify the first sample as not contaminated by the second sample by means of the computer system.

[0066] 3. A method for detecting the presence or absence of contamination of a first sample by a second sample, comprising:

[0067] (a) Obtain sequence information comprising more than one sequencing read from the first sample and the second sample;

[0068] (b) Align the more than one sequencing reads with a reference sequence, thereby determining a start region and an end region of the alignment;

[0069] (c) For each sample, group the more than one sequencing reads into more than one family based on grouping features, the grouping features including at least one of the following (i), (ii), and (iii): (i) the start region, (ii) the end region, and (iii) the length of the polynucleotide, wherein each family in the sample includes sequencing reads of daughter polynucleotides amplified from a unique polynucleotide in the set of polynucleotides of the sample;

[0070] (d) Generate family identifiers for said more than one family;

[0071] (e) Screen out a set of common family identifiers, where a given common family identifier is a family identifier of said first sample that is the same as or substantially the same as the family identifier of said second sample;

[0072] (f) Determine a quantitative measure of said set of common family identifiers; and

[0073] (g) If the quantitative measure of said set of common family identifiers is higher than a predetermined threshold, classify said first sample as contaminated by said second sample, or if the quantitative measure of said set of common family identifiers is at or below said predetermined threshold, classify said first sample as not contaminated by said second sample.

[0074] 4. The method according to any one of items 1 - 3, the method further comprising, before a), tagging said set of polynucleotides to produce tagged polynucleotides, where each tagged polynucleotide comprises a tag and a polynucleotide.

[0075] 5. The method according to item 4, where for each sample, said more than one sequencing read is grouped into more than one family based on grouping features, said grouping features including at least one of the following (i), (ii), (iii), and (iv): (i) said tag, (ii) said start region, (iii) said end region, and (iv) the length of the polynucleotide, where each family in said sample comprises sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in said set of polynucleotides in said sample.

[0076] 6. A method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising:

[0077] (a) Sequencing said tagged polynucleotides or said set of polynucleotides from said sample to produce more than one sequencing read, where each tagged polynucleotide comprises a tag and a polynucleotide;

[0078] (b) Aligning said more than one sequencing read with a reference sequence, thereby determining the start region and end region of said alignment;

[0079] (c) For each sample, grouping said more than one sequencing read into more than one family based on grouping features, said grouping features including said tag, where each family in said sample comprises sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in said set of tagged polynucleotides in said sample;

[0080] (d) Generate family identifiers for the more than one family;

[0081] (e) Screen out a set of common family identifiers, where a given common family identifier is a family identifier of the first sample that is the same as or substantially the same as the family identifier of the second sample;

[0082] (f) Determine a quantitative measure of the set of the common family identifiers; and

[0083] (g) If the quantitative measure of the common family identifiers is higher than a predetermined threshold, classify the first sample as contaminated by the second sample, or if the quantitative measure of the set of the common family identifiers is at or below the predetermined threshold, classify the first sample as not contaminated by the second sample.

[0084] 7. The method according to any one of the above items, wherein the quantitative measure of the set of the common family identifiers is the number of common family identifiers in the first sample.

[0085] 8. The method according to any one of the above items, wherein the quantitative measure of the set of the common family identifiers includes the ratio of the number of common family identifiers in the first sample to the total number of family identifiers in the first sample.

[0086] 9. The method according to any one of the above items, wherein the quantitative measure of the set of the common family identifiers does not include the following common family identifiers in the first sample: those common family identifiers in the family of the first sample whose number of sequencing reads is greater than the number of sequencing reads in the corresponding family of the second sample.

[0087] 10. The method according to item 4 or 6, wherein the quantitative measure of the set of the common family identifiers in the first sample does not include the common family identifiers at overrepresented genomic start and genomic end position pairs.

[0088] 11. The method according to item 10, wherein the overrepresented genomic start and genomic end position pairs are determined by:

[0089] (a) Provide more than one sample, wherein the more than one sample includes a distribution of genomic start and genomic end positions that is the same as or substantially the same as that of the first sample and / or the second sample;

[0090] (b) Determine the family identifiers in the more than one sample;

[0091] (c) Quantifying the number of family identifiers having a common pair of genomic start positions and genomic end positions in the more than one sample; and

[0092] (d) Classifying the pair of genomic start positions and genomic end positions as over-represented if the number of family identifiers exceeds a set threshold.

[0093] 12. The method according to item 11, wherein the more than one sample does not include the first sample or the second sample.

[0094] 13. The method according to item 11, wherein the more than one sample does not include the first sample and the second sample.

[0095] 14. The method according to item 11, wherein the more than one sample includes samples processed in the same flow cell as the first sample.

[0096] 15. The method according to item 11, wherein the more than one sample includes training samples.

[0097] 16. The method according to item 11, wherein the set threshold is at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, or at least 60 families.

[0098] 17. A method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising:

[0099] (a) Sequencing a collection of polynucleotides from the sample to generate more than one sequencing read;

[0100] (b) Aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment;

[0101] (c) For each sample, grouping the more than one sequencing read into more than one family based on grouping features, the grouping features including at least one of the following (i), (ii), and (iii): (i) the start region, (ii) the end region, and (iii) the length of the polynucleotide, wherein each family in the sample includes sequencing reads of daughter polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides from the sample;

[0102] (d) Screen the more than one family to identify a set of common families, where a given common family is a family of the first sample having grouping characteristics that are the same as or substantially the same as the grouping characteristics of the family of the second sample;

[0103] (e) Determine a quantitative measure of the set of the common families of the first sample; and

[0104] (f) If the quantitative measure of the set of the common families is higher than a predetermined threshold, classify the first sample as contaminated by the second sample, or if the quantitative measure of the set of the common families is at or below the predetermined threshold, classify the first sample as not contaminated by the second sample.

[0105] 18. A method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising:

[0106] (a) Sequencing a collection of polynucleotides from the sample to generate more than one sequencing read;

[0107] (b) Align the more than one sequencing read with a reference sequence, thereby determining a start region and an end region of the alignment;

[0108] (c) Group the more than one sequencing read of the two samples together into more than one family based on grouping characteristics, the grouping characteristics including at least one of the following (i), (ii), and (iii): (i) the start region, (ii) the end region, and (iii) the length of the polynucleotide, where each family includes sequencing reads of progeny polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides;

[0109] (d) Screen the more than one family to identify a set of common families, where the common families include at least one sequencing read from the first sample and at least one sequencing read from the second sample;

[0110] (e) Determine a quantitative measure derived from the set of the common families; and

[0111] (f) If the quantitative measure of the set of the common families is higher than a predetermined threshold, classify the first sample as contaminated by the second sample, or if the quantitative measure of the set of the common families is at or below the predetermined threshold, classify the first sample as not contaminated by the second sample.

[0112] 19. The method according to item 17 or 18, the method further comprising, prior to the sequencing, tagging a collection of polynucleotides to produce tagged polynucleotides, wherein each tagged polynucleotide comprises a tag and a polynucleotide.

[0113] 20. The method according to item 19, wherein for each sample, the more than one sequencing reads are grouped into more than one family based on grouping features, the grouping features comprising at least one of the following (i), (ii), (iii) and (iv): (i) the tag, (ii) the start region, (iii) the end region, and (iv) the length of the polynucleotide, wherein each family in the sample comprises sequencing reads of progeny polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides in the sample.

[0114] 21. A method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising:

[0115] (a) Sequencing a collection of tagged polynucleotides from the sample to produce more than one sequencing read, wherein each tagged polynucleotide comprises a tag and a polynucleotide;

[0116] (b) Aligning the more than one sequencing read with a reference sequence, thereby determining the start region and end region of the alignment;

[0117] (c) For each sample, grouping the more than one sequencing read into more than one family based on grouping features, the grouping features comprising the tag, wherein the families in the sample comprise sequencing reads of tagged progeny polynucleotides amplified from unique polynucleotides in the collection of tagged polynucleotides in the sample;

[0118] (d) Screening the more than one family to identify a set of common families, wherein a given common family is a family of the first sample having grouping features that are the same or substantially the same as the grouping features of the families of the second sample;

[0119] (e) Determining a quantitative measure of the set of common families of the first sample; and

[0120] (f) If the quantitative measure of the set of common families is above a predetermined threshold, classifying the first sample as contaminated by the second sample, or if the quantitative measure of the set of common families is at or below the predetermined threshold, classifying the first sample as not contaminated by the second sample.

[0121] 22. A method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising:

[0122] (a) Sequence a collection of tagged polynucleotides from the sample to generate more than one sequencing read, where each tagged polynucleotide comprises a tag and a polynucleotide;

[0123] (b) Align the more than one sequencing read with a reference sequence, thereby determining a start region and an end region of the alignment;

[0124] (c) Group the more than one sequencing read of two samples together into more than one family based on grouping features, the grouping features including the tag, where each family in the sample comprises sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the collection of tagged polynucleotides in the sample;

[0125] (d) Screen the more than one family to identify a set of common families, where a given common family comprises at least one sequencing read from the first sample and at least one sequencing read from the second sample;

[0126] (e) Determine a quantitative measure derived from the set of common families; and

[0127] (f) If the quantitative measure of the set of common families is higher than a predetermined threshold, classify the first sample as contaminated by the second sample, or if the quantitative measure of the set of common families is at or below the predetermined threshold, classify the first sample as not contaminated by the second sample.

[0128] 23. The method according to any one of items 17 - 22, wherein the quantitative measure comprises the number of common families in the first sample.

[0129] 24. The method according to item 18 or 22, wherein the quantitative measure comprises the ratio of the number of sequencing reads of the first sample to the number of sequencing reads of the second sample in the common family.

[0130] 25. The method according to any one of the above items, wherein the quantitative measure comprises the ratio of the number of common families in the first sample to the total number of families in the first sample.

[0131] 26. The method according to any one of the above items, wherein the quantitative measure of the set of common families does not include the following common families in the first sample: those common families in the families of the first sample whose number of sequencing reads is greater than the number of sequencing reads of the corresponding families in the second sample.

[0132] 27. The method according to any one of items 19 - 22, wherein the quantitative measure of the set of common families in the first sample does not include common families at over - represented genomic start and genomic end position pairs.

[0133] 28. The method according to item 27, wherein the over - represented genomic start and genomic end position pairs are determined by:

[0134] (a) providing more than one sample, wherein the more than one sample includes a distribution of genomic start and genomic end positions that is the same as or substantially the same as that of the first sample and / or the second sample;

[0135] (b) determining the families in the more than one sample;

[0136] (c) quantifying the number of families that share a pair of genomic start and genomic end positions in the more than one sample; and

[0137] (d) classifying the genomic start and genomic end position pair as over - represented if the number of families exceeds a set threshold.

[0138] 29. The method according to item 28, wherein the more than one sample does not include the first sample or the second sample.

[0139] 30. The method according to item 28, wherein the more than one sample does not include the first sample and the second sample.

[0140] 31. The method according to item 28, wherein the more than one sample includes samples processed in the same flow cell as the first sample.

[0141] 32. The method according to item 28, wherein the more than one sample includes training samples.

[0142] 33. The method according to item 28, wherein the set threshold is at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, or at least 60 families.

[0143] 34. The method according to any one of the above items, wherein the start region includes the genomic start position of the sequencing read, at which the 5' end of the sequencing read is determined to start aligning with the reference sequence, and the end region includes the genomic end position of the sequencing read, at which the 3' end of the sequencing read is determined to end aligning with the reference sequence.

[0144] 35. The method according to item 34, wherein the start region includes the first 1, first 2, first 5, first 10, first 15, first 20, first 25, first 30, or at least the first 30 base positions of the 5' end of the sequencing read aligned with the reference sequence.

[0145] 36. The method according to item 34, wherein the end region includes the last 1, last 2, last 5, last 10, last 15, last 20, last 25, last 30, or at least the last 30 base positions of the 3' end of the sequencing read aligned with the reference sequence.

[0146] 37. The method according to any one of the above items, wherein the tag includes one or more molecular barcodes attached to the ends of the polynucleotide.

[0147] 38. The method according to item 37, wherein the length of the one or more molecular barcodes is at least 2, at least 4, at least 5, at least 6, at least 8, at least 10, at least 15, or at least 20 nucleotides.

[0148] 39. The method according to item 37, wherein the one or more molecular barcodes attached to the polynucleotide of the first sample are different from the one or more molecular barcodes attached to the polynucleotide of the second sample.

[0149] 40. The method according to any one of the above items, wherein the polynucleotide of the sample is tagged with at least 5, at least 10, at least 15, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10,000, at least 50,000, or at least 100,000 different molecular barcodes.

[0150] 41. The method according to any one of the above items, wherein the first sample and the second sample are sequenced in the same flow cell.

[0151] 42. The method according to any one of the above items, wherein the second sample is sequenced in a flow cell different from the first sample.

[0152] 43. The method according to any one of the above items, wherein the second sample is on the same day as the first sample, but is processed at a different time from the first sample.

[0153] 44. The method according to item 43, wherein the second sample is processed at least 1 minute, at least 30 minutes, at least 1 hour, at least 2 hours, at least 3 hours, or at least 4 hours after processing the first sample.

[0154] 45. The method according to any one of the above items, wherein the first sample and the second sample are processed on different dates.

[0155] 46. The method according to any one of the above items, wherein the first sample and the second sample are in the same sample batch.

[0156] 47. The method according to any one of the above items, wherein the second sample and the first sample are processed with the same batch of reagents.

[0157] 48. The method according to item 47, wherein the first sample and the second sample are processed at different geographical locations.

[0158] 49. The method according to any one of the above items, wherein the set of tagged polynucleotides of the sample is uniquely tagged.

[0159] 50. The method according to any one of the above items, wherein the set of tagged polynucleotides of the sample is non-uniquely tagged.

[0160] 51. The method according to any one of the above items, wherein the first sample is obtained from the body fluid of one subject, and the second sample is obtained from the body fluid of another subject.

[0161] 52. The method according to any one of the above items, wherein the polynucleotide is a cell-free polynucleotide.

[0162] 53. The method according to item 52, wherein the cell-free polynucleotide is cell-free DNA.

[0163] 54. The method according to item 51, wherein at least one of the subjects has a disease.

[0164] 55. The method according to item 54, wherein the disease is cancer.

[0165] 56. The method according to any one of the above items, wherein the set of polynucleotides of the sample is amplified prior to sequencing, thereby generating amplified progeny polynucleotides.

[0166] 57. The method according to item 56, the method further comprising selectively enriching at least a portion of the amplified progeny polynucleotides from regions of the genome or transcriptome of the subject prior to the sequencing.

[0167] 58. The method according to item 57, the method further comprising attaching one or more sample indices to one end or both ends of the amplified progeny polynucleotides prior to sequencing, wherein the sample indices distinguish the first sample and the second sample.

[0168] 59. The method according to any one of the above items, wherein the predetermined threshold is at least 0.001%, at least 0.005%, at least 0.01%, at least 0.05%, at least 0.1%, at least 0.5%, at least 1%, at least 2%, at least 5% or at least 10% of the total number of families in the first sample.

[0169] 60. The method according to any one of the above items, the method further comprising detecting somatic genetic variations of the polynucleotides of the first sample by excluding the sequencing reads of the shared families of the first sample, wherein the first sample is classified as being contaminated by the second sample.

[0170] 61. The method according to any one of the foregoing items, the method further comprising generating a report, the report optionally including information about the contamination status of the sample and / or information derived from the contamination status of the sample.

[0171] 62. The method according to item 61, the method further comprising transmitting the report to a third party, such as the subject from whom the sample is derived or a health care practitioner.

[0172] From the following detailed description, additional aspects and advantages of the present disclosure will become readily apparent to those skilled in the art. Only illustrative embodiments of the present disclosure are shown and described in the following detailed description. As will be appreciated, the present disclosure is capable of other and different embodiments, and several details thereof can be modified in various obvious aspects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not restrictive. Brief Description of the Drawings

[0173] The accompanying drawings illustrate certain embodiments and, together with the written description, are used to explain certain principles of the methods, computer-readable media, and systems disclosed herein. The drawings are incorporated in and constitute a part of this specification. The description provided herein is better understood when read in conjunction with the drawings, which are included by way of example and not limitation. It will be understood that, unless the context otherwise indicates, like reference numerals identify like elements throughout the drawings. It will also be understood that some or all of the drawings may be schematic representations for illustrative purposes and do not necessarily depict the actual relative sizes or positions of the elements shown.

[0174] Figure 1 is a flow chart representation of a method for detecting the presence or absence of contamination between two samples according to an embodiment of the present disclosure.

[0175] Figure 2 is a flow chart representation of a method for detecting the presence or absence of contamination between two samples according to an embodiment of the present disclosure.

[0176] Figure 3 is a schematic diagram illustrating grouping sequencing reads into families and thereby detecting the presence or absence of contamination between two samples according to an embodiment of the present disclosure.

[0177] Figure 4 is a schematic diagram of an exemplary system suitable for some embodiments of the present disclosure.

[0178] Definition of Terms

[0179] Although various embodiments of the present disclosure have been shown and described herein, those skilled in the art will understand that such embodiments are provided by way of example only. Many variations, changes, and substitutions will occur to those skilled in the art without departing from the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be employed.

[0180] To more readily understand the present disclosure, certain terms are first defined below. Additional definitions of these and other terms may be set forth throughout the specification. If the definitions of the terms set forth below are inconsistent with the definitions in the applications or patents incorporated by reference, the definitions set forth in this application shall be used to understand the meaning of the term.

[0181] As used in this specification and the appended claims, unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" include plural referents. Thus, for example, reference to "a method" includes one or more methods, and / or types of steps, etc. described herein and / or that will become apparent to one of ordinary skill in the art upon reading the present disclosure.

[0182] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. Additionally, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. When describing and claiming methods, computer-readable media, and systems, the following terms and their grammatical variations will be used in accordance with the definitions set forth below.

[0183] About: As used herein, "about" or "approximately", when applied to one or more values or elements of interest, refers to a value or element that is similar to the stated reference value or element. In certain embodiments, the term "about" or "approximately" refers to a range of values or elements that fall within 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1% or less in either direction (greater than or less than) of the stated reference value or element, unless otherwise stated or otherwise apparent from the context (except when such numbers would exceed 100% of the possible value or element).

[0184] Adapter: As used herein, an "adapter" refers to a short nucleic acid (e.g., less than about 500 nucleotides in length, less than about 100 nucleotides in length, or less than about 50 nucleotides in length), which is typically at least partially double-stranded and is used to ligate either one or both ends of a nucleic acid molecule of a given sample. An adapter can include a nucleic acid primer binding site and / or a sequencing primer binding site, where the nucleic acid primer binding site allows for the amplification of a nucleic acid molecule flanked by adapters at both ends, and / or the sequencing primer binding site includes a primer binding site for sequencing applications such as various next-generation sequencing (NGS) applications. An adapter can also include a binding site for capture probes such as oligonucleotides attached to a flow cell support. An adapter can also include a nucleic acid tag as described herein. The nucleic acid tag is typically positioned relative to the binding sites of the amplification primers and the sequencing primers such that the nucleic acid tag is included in the amplicon and the sequencing reads of a given nucleic acid molecule. The same or different adapters can be ligated to the respective ends of a nucleic acid molecule. In some embodiments, adapters of the same sequence except for the nucleic acid tag are ligated to the respective ends of a nucleic acid molecule. In some embodiments, the adapter is a Y-shaped adapter, where one end is blunt-ended or tailed as described herein for ligating a nucleic acid molecule that is also blunt-ended or tailed with one or more complementary nucleotides. In still other exemplary embodiments, the adapter is a bell-shaped adapter that includes a blunt-ended or tailed end for ligation to a nucleic acid molecule to be analyzed. Other examples of adapters include T-tailed and C-tailed adapters.

[0185] Amplify: As used herein, "amplify" or "amplification" in the context of nucleic acids refers to the process of typically generating multiple copies of a polynucleotide or a portion of the polynucleotide starting from a small number of polynucleotides (e.g., a single polynucleotide molecule), where the amplification product or amplicon is typically detectable. Amplification of polynucleotides encompasses a variety of chemical and enzymatic processes.

[0186] Barcode: As used herein, "barcode" or "molecular barcode" in the context of nucleic acids refers to a nucleic acid molecule that contains a sequence that can be used as a molecular identifier. For example, during next-generation sequencing (NGS) library preparation, individual "barcode" sequences are typically added to each DNA fragment such that each read can be identified and sorted prior to final data analysis.

[0187] Cancer type: As used herein, "cancer type" refers to the type or subtype of cancer defined, for example, by histopathology. Cancer types can be defined by any conventional criteria, such as based on occurrence in a given tissue (e.g., blood cancer, central nervous system (CNS) cancer, brain cancer, lung cancer (small cell and non-small cell), skin cancer, nasal cancer, laryngeal cancer, liver cancer, bone cancer, lymphoma, pancreatic cancer, bowel cancer, rectal cancer, thyroid cancer, bladder cancer, kidney cancer, oral cancer, stomach cancer, breast cancer, prostate cancer, ovarian cancer, lung cancer, small intestine cancer, soft tissue cancer, neuroendocrine cancer, gastroesophageal cancer, head and neck cancer, gynecological cancer, colorectal cancer, urothelial cancer, solid state cancers, heterogeneous cancer, homogeneous cancer), unknown primary source, etc., and / or having the same cell lineage (e.g., carcinoma, sarcoma, lymphoma, cholangiocarcinoma, leukemia, mesothelioma, melanoma or glioblastoma) and / or showing cancer markers such as Her2, CA15-3, CA19-9, CA-125, CEA, AFP, PSA, HCG, hormone receptor and NMP-22. Cancers can also be classified by stage (e.g., stage 1, stage 2, stage 3 or stage 4) and whether they are of primary or secondary origin.

[0188] Cell-free nucleic acid: As used herein, "cell-free nucleic acid" refers to nucleic acids that are not contained within cells or otherwise associated with cells, or in some embodiments, nucleic acids that remain in a sample after removal of intact cells. Cell-free nucleic acids can include, for example, all unencapsulated nucleic acids derived from body fluids of a subject (e.g., blood, plasma, serum, urine, cerebrospinal fluid (CSF), etc.). Cell-free nucleic acids include DNA (cfDNA), RNA (cfRNA), and their hybrids, including genomic DNA, mitochondrial DNA, circulating DNA, siRNA, miRNA, circular RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), and / or fragments of any of these. Cell-free nucleic acids can be double-stranded, single-stranded, or their hybrids. Cell-free nucleic acids can be released into body fluids by secretion or cell death processes, such as necrosis, apoptosis, etc. Some cell-free nucleic acids are released from cancer cells into body fluids, e.g., circulating tumor DNA (ctDNA). Others are released from healthy cells. CtDNA can be fragmented tumor-derived DNA that is unencapsulated. Another example of cell-free nucleic acid is fetal DNA that circulates freely in the maternal bloodstream, also known as cell-free fetal DNA (cffDNA). Cell-free nucleic acids can have one or more epigenetic modifications, e.g., cell-free nucleic acids can be acetylated, 5-methylated, ubiquitinated, phosphorylated, sumoylated, ribosylated, and / or citrullinated.

[0189] Cellular nucleic acid: As used herein, "cellular nucleic acid" means nucleic acids that are at least within one or more cells that give rise to the nucleic acids at the point of obtaining or collecting a sample from a subject, even if these nucleic acids are subsequently removed (e.g., via cell lysis) as part of a given analytical process.

[0190] Contamination of a sample: As used herein, the term "contamination" or "contamination of a sample" refers to any chemical or digital contamination of one sample by another sample. Contamination can be due to a variety of sources, such as but not limited to: physical remnants of liquids between samples (e.g., pipetting, automated liquid handling via sample preparation or sequencer systems, manipulation of amplified material); demultiplexing artifacts (e.g., base calling errors that confound sample indices with limited pairwise Hamming distance; insertions / deletions that confound sample indices with limited pairwise edit distance); and reagent impurities (e.g., sample index oligonucleotides contaminated (by remnants of synthesis errors) with oligonucleotides containing another sample index).

[0191] Deoxyribonucleic acid or ribonucleic acid: As used herein, "deoxyribonucleic acid" or "DNA" refers to natural or modified nucleotides having a hydrogen group at the 2'-position of the sugar moiety. DNA typically comprises a nucleotide chain containing the following four types of nucleobases: adenine (A), thymine (T), cytosine (C), and guanine (G). As used herein, "ribonucleic acid" or "RNA" refers to natural or modified nucleotides having a hydroxyl group at the 2'-position of the sugar moiety. RNA typically comprises a nucleotide chain containing the following four types of nucleobases: A, uracil (U), G, and C. As used herein, the term "nucleotide" refers to natural nucleotides or modified nucleotides. Certain nucleotide pairs specifically bind to each other in a complementary manner (referred to as complementary base pairing). In DNA, adenine (A) pairs with thymine (T) and cytosine (C) pairs with guanine (G). In RNA, adenine (A) pairs with uracil (U) and cytosine (C) pairs with guanine (G). When a first nucleic acid strand binds to a second nucleic acid strand composed of nucleotides complementary to those in the first strand, the two strands bind to form a double strand. As used herein, "nucleic acid sequencing data", "nucleic acid sequencing information", "sequence information", "nucleic acid sequence", "nucleotide sequence", "genomic sequence", "gene sequence", or "fragment sequence", or "nucleic acid sequencing reads" refers to any information or data indicating the order and identity of nucleobases (e.g., adenine, guanine, cytosine, and thymine or uracil) in a molecule of nucleic acid such as DNA or RNA (e.g., whole genome, whole transcriptome, exome, oligonucleotide, polynucleotide, or fragment). It should be understood that the present teachings contemplate sequence information obtained using all available various techniques, platforms, or technologies, including but not limited to the following: capillary electrophoresis, microarrays, ligation-based systems, polymerase-based systems, hybridization-based systems, direct or indirect nucleotide identification systems, pyrosequencing, ion- or pH-based detection systems, and electronic signature-based systems.

[0192] Family: As used herein, the term "family" refers to one or more sequencing reads derived from a single polynucleotide molecule. Bioinformatically, one or more sequencing reads derived from a single polynucleotide molecule will have the same or substantially the same grouping features, where the grouping features include at least one of the following: (i) a tag (i.e., a molecular barcode), (ii) the start region of the alignment, (iii) the end region of the alignment, and (iv) the length of the polynucleotide. Those sequencing reads having the same or substantially the same grouping features can be grouped together into a family. In some embodiments, although with low probability, at least two molecules can have the same grouping features, and thus sequencing reads derived from at least two molecules can be grouped into a single family.

[0193] In some embodiments, sequencing reads derived from a single polynucleotide molecule are detected only in a single sample. In some embodiments, in the presence of contamination of at least two samples, then sequencing reads derived from a single polynucleotide molecule (of a single sample) can be detected in at least two samples. In these embodiments, in the case of grouping sequencing reads independently for each sample, then the sequencing reads derived from a single polynucleotide molecule detected in each sample will be grouped as separate families in that sample. In other embodiments, in the case of grouping sequencing reads together for all at least two samples, then the sequencing reads derived from a single polynucleotide molecule detected in at least two samples will be grouped into a single family.

[0194] The grouping features of a family represent the grouping features of the sequencing reads in that family. In some embodiments, if a family includes sequencing reads with the same grouping features, then the grouping features of any sequencing read are the grouping features of the family. In other embodiments, if a family includes sequencing reads with the same and substantially the same grouping features, then the grouping features of the family can be one of but not limited to the following or a combination of but not limited to the following: (i) the most frequently represented grouping feature of the sequencing reads; (ii) the average value of the grouping features of the sequencing reads; (iii) the most frequently represented nucleotide base in the molecular barcode; (iv) the maximum likelihood value of the molecular barcode and / or the start region and / or the end region of the sequencing reads.

[0195] In some embodiments, a family includes at least two sequencing reads derived from a single polynucleotide molecule. In some embodiments, a family can include sequence reads from a single strand of a double-stranded polynucleotide molecule. In some embodiments, a family includes sequence reads from both strands (sense and antisense strands) of a double-stranded polynucleotide molecule. In an example, a molecular barcode, a genomic start position, and a genomic end position are considered grouping features of the family. In this example, if a family has 10 sequence reads and all the sequence reads have the same molecular barcode and genomic start position, but the genomic end positions are different, then the molecular barcode and genomic start position become the grouping features of the family, and for the genomic end position, the genomic end position represented by most of the sequencing reads in the family will be considered the genomic end position of the family (which is part of the grouping features of the family).

[0196] Family identifier: As used herein, the term "family identifier" refers to an identifier that uniquely identifies each family, and it includes grouping features and / or information derived from the grouping features of the family. In some embodiments, a family identifier can include an integer, a letter, or a combination of both. In some embodiments, a family identifier is assigned to the sequencing reads in a family.

[0197] Germline mutation: As used herein, the terms "germline mutation" or "germline variant" are used interchangeably and refer to a genetic mutation (i.e., a mutation that does not occur after conception). Germline mutations may be the only mutations that can be passed on to offspring and may be present in every somatic and germline cell in the offspring.

[0198] Indel: As used herein, "Indel" refers to a mutation involving an insertion or deletion of nucleotides in the genome of a subject.

[0199] Mutant Allele Fraction: As used herein, "Mutant Allele Fraction", "mutation dose", or "MAF" refers to the fraction of nucleic acid molecules carrying an allelic change or mutation at a given genomic position / locus in a given sample. MAF is typically expressed as a fraction or a percentage. For example, the MAF of a somatic variant may be less than 0.15.

[0200] Mutation: As used herein, "mutation" refers to a variation from a known reference sequence and includes mutations such as, for example, single nucleotide variants (SNVs) and insertions or deletions (Indels). Mutations can be germline mutations or somatic mutations. In some embodiments, the reference sequence for comparison purposes is the wild-type genomic sequence of the species of the subject providing the test sample, typically the human genome.

[0201] Neoplasm: As used herein, the terms "neoplasm" and "tumor" are used interchangeably. They refer to the abnormal growth of cells in a subject. A neoplasm or tumor can be benign, potentially malignant, or malignant. A malignant neoplasm refers to a cancer or cancerous tumor.

[0202] Next-generation sequencing: As used herein, "next-generation sequencing" or "NGS" refers to sequencing technologies that have increased throughput compared to traditional Sanger and capillary electrophoresis-based methods, e.g., sequencing technologies that have the ability to generate hundreds of thousands of relatively small sequence reads at one time. Some examples of next-generation sequencing technologies include, but are not limited to, sequencing by synthesis, sequencing by ligation, and sequencing by hybridization.

[0203] Nucleic acid tag: As used herein, a "nucleic acid tag" refers to a short nucleic acid (e.g., less than about 500 nucleotides in length, about 100 nucleotides, about 50 nucleotides, or about 10 nucleotides) that is used to distinguish nucleic acids from different samples (e.g., representing a sample index), or different nucleic acid molecules of different types or that have undergone different treatments in the same sample (e.g., representing a molecular barcode). The nucleic acid tag contains a predefined, fixed, non-random, random, or semi-random oligonucleotide sequence. Such nucleic acid tags can be used to label different nucleic acid molecules or different nucleic acid samples or subsamples. The nucleic acid tag can be single-stranded, double-stranded, or at least partially double-stranded. The nucleic acid tags optionally have the same length or different lengths. The nucleic acid tag can also include a double-stranded molecule having one or more blunt ends, including 5' or 3' single-stranded regions (e.g., overhangs), and / or one or more other single-stranded regions at other positions within a given molecule. The nucleic acid tag can be attached to one end or both ends of other nucleic acids (e.g., sample nucleic acids to be amplified and / or sequenced). The nucleic acid tag can be decoded to reveal information such as the sample origin, form, or treatment performed on a given nucleic acid. For example, nucleic acid tags can also be used to enable pooling and / or parallel processing of multiple samples containing nucleic acids with different molecular barcodes and / or sample indices, where the nucleic acids are subsequently deconvolved by detecting (e.g., reading) the nucleic acid tags. The nucleic acid tag can also be referred to as an identifier (e.g., a molecular identifier, a sample identifier). Additionally or alternatively, the nucleic acid tag can be used as a molecular barcode (e.g., to distinguish different molecules or amplicons of different parental molecules in the same sample or subsample). This includes, for example, uniquely tagging different nucleic acid molecules in a given sample, or non-uniquely tagging such molecules. In the case of non-uniquely tagging applications, a limited number of tags (e.g., molecular barcodes) can be used to tag nucleic acid molecules such that different molecules can be distinguished based on a combination of their endogenous sequence information (e.g., the start and / or end positions where it maps to a selected reference genome, a subsequence at one or both ends of the sequence, and / or the length of the sequence) and at least one molecular barcode. Generally, a sufficient number of different molecular barcodes are used such that the probability is low (e.g., less than about 10%, less than about 5%, less than about 1%, or less than about 0.1% probability) that any two molecules are likely to have the same endogenous sequence information (e.g., start and / or end positions, a subsequence at one or both ends of the sequence, and / or length) and also have the same molecular barcode.

[0204] Over-represented genomic start and genomic end position pairs: As used herein, the term "over-represented genomic start and genomic end position pair" or "over-represented pair" refers to a genomic start and genomic end position pair in which the number or frequency of families in which the genomic start and genomic end position pair is common in more than one sample exceeds a set threshold. In some embodiments, more than one sample includes samples run in a flow cell, and a first sample and a second sample are run in the flow cell. For example, more than one sample can be a training sample or a sample processed in a particular flow cell of a nucleic acid sequencer related to the first sample and / or the second sample being analyzed. In some embodiments, more than one sample does not include the first sample and / or the second sample. In some embodiments, the set threshold can be any value between 2 and 100. In some embodiments, the set threshold can be 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, at least 21, at least 25, at least 30, at least 35, at least 40, or at least 50. In some embodiments, the set threshold can be 5. In some embodiments, the set threshold can be 10. In some embodiments, the set threshold can be 15. In some embodiments, the set threshold can be 20. In some embodiments, the set threshold can be at least 10 -3 , at least 10 -4 , at least 10 -5 , at least 10 -6 , at least 10 -7 , at least 10 -8 or at least 10 -9 . In some embodiments, the set threshold can be 10 of the total families observed in more than one sample -4 . In some embodiments, the set threshold can be 10 of the total families observed in more than one sample -5 . In some embodiments, the set threshold can be 10 of the total families observed in more than one sample -6 . In some embodiments, the set threshold can be 10 of the total families observed in more than one sample -7 . In some embodiments, the set threshold can be 10 of the total families observed in more than one sample -8 .

[0205] Polynucleotide: As used herein, "polynucleotide", "nucleic acid", "nucleic acid molecule", or "oligonucleotide" refers to a linear polymer of nucleosides (including deoxyribonucleosides, ribonucleosides, or analogs thereof) linked by internucleoside linkages. Typically, a polynucleotide contains at least three nucleosides. The size range of oligonucleotides generally ranges from a few monomer units (e.g., 3 - 4) to several hundred monomer units. Whenever a polynucleotide is represented by a string of letters such as "ATGCCTG", it will be understood that these nucleotides are in the 5'→3' order from left to right, and in the case of DNA, "A" represents deoxyadenosine, "C" represents deoxycytidine, "G" represents deoxyguanosine, and "T" represents deoxythymidine, unless otherwise stated. As is standard in the art, the letters A, C, G, and T can be used to refer to the base itself, the nucleoside, or the nucleotide containing these bases.

[0206] Reference sequence: As used herein, "reference sequence" refers to a known sequence used for the purpose of comparison with an experimentally determined sequence. For example, the known sequence can be an entire genome, a chromosome, or any segment thereof. A reference sequence typically includes at least about 20, at least about 50, at least about 100, at least about 200, at least about 250, at least about 300, at least about 350, at least about 400, at least about 450, at least about 500, at least about 1000, or more than 1000 nucleotides. The reference sequence can be aligned with a single continuous sequence of a genome or chromosome, or can include non - continuous segments aligned with different regions of a genome or chromosome. Examples of reference sequences include, for example, the human genome, such as, hG19 and hG38.

[0207] Sample: As used herein, "sample" means anything that can be analyzed by the methods and / or systems disclosed herein.

[0208] Sequencing: As used herein, "sequencing" refers to any of a number of techniques for determining the sequence (e.g., the identity and order of monomeric units) of a biomolecule such as a nucleic acid, such as DNA or RNA. Examples of sequencing methods include, but are not limited to, targeted sequencing, single molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxy chain termination sequencing, whole genome sequencing, sequencing by hybridization, pyrosequencing, capillary electrophoresis, duplex sequencing, cycle sequencing, single base extension sequencing, solid phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification at lower denaturation temperature PCR (COLD-PCR), multiplex PCR, reversible dye terminator sequencing, paired-end sequencing, near-term sequencing, exonuclease sequencing, ligation sequencing, short read sequencing, single molecule sequencing, sequencing by synthesis, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa genome analyzer sequencing, SOLiD TM sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, sequencing can be performed by a genetic analyzer, such as, for example, a commercially available genetic analyzer from many other companies such as Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific.

[0209] Sequence information: As used herein, "sequence information" in the context of a nucleic acid polymer means the order and identity of the monomeric units (e.g., nucleotides, etc.) in the polymer.

[0210] Consensus family: If the grouping of sequencing reads into families is performed independently for a first sample and a second sample, the term "consensus family" refers to a family in the first sample whose grouping characteristics are the same or substantially the same as the grouping characteristics of a family in the second sample. Alternatively, if the grouping of sequencing reads into families is performed together for both the first sample and the second sample, the term "consensus family" refers to a family that includes at least one sequencing read from the first sample and at least one sequencing read from the second sample.

[0211] In some embodiments, in the presence of contamination of at least two samples, sequencing reads derived from a single polynucleotide molecule (from a single sample) can then be detected in at least two samples. In these embodiments, when the sequencing reads are grouped independently for each sample, then the sequencing reads derived from a single polynucleotide molecule detected in each sample are grouped into separate families in that sample. In these embodiments, a shared family refers to a family in the first sample whose grouping characteristics are the same or substantially the same as those of a family in the second sample.

[0212] Optionally, in other embodiments, when the sequencing reads are grouped together for all at least two samples, then the sequencing reads derived from a single polynucleotide molecule detected in at least two samples are grouped into a single family. In these embodiments, a shared family refers to a family that has at least one sequencing read from at least two samples.

[0213] In some embodiments, the first sample and the second sample can be in the same flow cell or in different flow cells.

[0214] Shared family identifier: As used herein, the term "shared family identifier" refers to a family identifier in the first sample that is the same or substantially the same as the family identifier of a family in the second sample, i.e., a grouping characteristic in the first sample that is the same or substantially the same as the grouping characteristic of a family in the second sample. In some embodiments, the first sample and the second sample can be in the same flow cell or in different flow cells.

[0215] Single nucleotide polymorphism: As used herein, the terms "single nucleotide polymorphism" or "SNP" are used interchangeably. They refer to a variation of a single nucleotide at a specific position in the genome, where each variation exists in the population at a certain assessable frequency (e.g., greater than about 1%).

[0216] Single nucleotide variant: As used herein, "single nucleotide variant" or "SNV" means a mutation or variation of a single nucleotide that occurs at a specific position in the genome.

[0217] Somatic mutation: As used herein, the terms "somatic mutation" or "somatic variant" are used interchangeably. They refer to mutations in the genome that occur after conception. Somatic mutations can occur in any cell of the body except germ cells and are therefore not passed on to offspring.

[0218] Subject: As used herein, "subject" refers to an animal, such as a mammalian species (e.g., human), or an avian (e.g., bird) species, or other organism, such as a plant. More specifically, a subject can be a vertebrate, e.g., a mammal, such as a mouse, a primate, an ape, or a human. Animals include farm animals (e.g., production cattle, dairy cows, poultry, horses, pigs, etc.), sport animals, and companion animals (e.g., pets or support animals). A subject can be a healthy individual, an individual having or suspected of having a disease or predisposed to having the disease, or an individual in need of or suspected of needing treatment. The terms "individual" or "patient" are intended to be interchangeable with "subject".

[0219] For example, a subject can be an individual who has been diagnosed with cancer, is going to receive cancer treatment, and / or has received at least one cancer treatment. A subject can be in cancer remission. As another example, a subject can be an individual diagnosed with an autoimmune disease. As another example, a subject can be a female individual who is pregnant or planning to become pregnant and who may have been diagnosed with or suspected of having a disease, such as cancer, an autoimmune disease.

[0220] Substantially identical: As used herein, the term "substantially identical" refers to two different entities that are 99.9% identical, at least 95% identical, at least 90% identical, at least 85% identical, at least 80% identical, at least 75% identical, at least 70% identical, at least 60% identical, or at least 50% identical. For example, when a family in a first sample is substantially identical to a family in a second sample, the grouping characteristics of the family in the first sample and the grouping characteristics of the family in the second sample are 99.9% identical, at least 95% identical, at least 90% identical, at least 85% identical, at least 80% identical, at least 75% identical, at least 70% identical, at least 60% identical, or at least 50% identical. In the case where the entity is a molecular barcode, the term "substantially identical" refers to two different molecular barcodes having a Hamming distance or edit distance of less than 1, less than 2, less than 3, less than 4, less than 5, less than 6, less than 7, or less than 8. In the case where the entity is a start region or an end region, the term "substantially identical" refers to two different regions that are within 1 bp, within 2 bp, within 3 bp, within 4 bp, within 5 bp, within 6 bp, within 7 bp, within 8 bp, within 9 bp, within 10 bp, within 11 bp, within 15 bp, within 20 bp, or within 25 bp. In the case where the entity is the length of a polynucleotide, the term "substantially identical" refers to two different lengths that are within 1 bp, within 2 bp, within 3 bp, within 4 bp, within 5 bp, within 6 bp, within 7 bp, within 8 bp, within 9 bp, within 10 bp, within 11 bp, within 15 bp, within 20 bp, within 25 bp, within 30 bp, within 40 bp, or within 50 bp.

[0221] Threshold: As used herein, "threshold" refers to a predetermined value that is used to characterize an experimentally determined value of the same parameter for different samples, depending on their relationship to the threshold. For example, a threshold for a p-value can refer to any predetermined value between 0 and 1 and is used to identify the source of a nucleic acid variant.

[0222] Training sample: As used herein, "training sample" refers to a set of samples having properties, parameters, and / or composition similar to the first sample and / or the second sample, where the first sample and / or the second sample are analyzed for the presence or absence of contamination.

[0223] Variation: As used herein, "variation" can relate to an allele. Depending on whether the allele is heterozygous or homozygous, the variation is typically present at a frequency of 50% (0.5) or 100% (1). For example, germline variations are inherited and typically have a frequency of 0.5 or 1. However, somatic variations are acquired variations and typically have a frequency of less than about 0.5. The major and minor alleles of a genetic locus refer to the nucleic acids containing that locus where the locus is occupied by a nucleotide of the reference sequence and a variant nucleotide different from the reference sequence, respectively. Measurements at a locus can take the form of an allele fraction (AF), which measures the frequency of the allele observed in a sample.

[0224] Detailed description

[0225] I. Overview

[0226] When processing samples for analysis, false positive results may be introduced by chemical or digital cross - contamination of samples processed in the same batch or in a temporally and spatially proximate manner by spreading molecules present in one sample into another. In the case of assaying cell - free nucleic acids from a sample containing contaminants or a second genome (i.e., a genome other than the genome of the subject and genomes arising from, for example, grafts, transfusions, or fetuses), the sample may require additional manual inspection or even additional sequencing runs.

[0227] The present disclosure provides methods and systems for detecting the presence or absence of contamination of a first sample by a second sample.

[0228] In one aspect, the present disclosure provides a method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising: (a) accessing, by a computer system, sequence information comprising more than one sequencing read from the first sample and the second sample; (b) aligning, by the computer system, the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping, by the computer system, the more than one sequencing read into more than one family based on grouping features, the grouping features comprising at least one of the following (i), (ii), and (iii): (i) the start region, (ii) the end region, and (iii) the length of the sequence read, wherein each family in the sample comprises sequencing reads of progeny polynucleotides amplified from a unique polynucleotide in a collection of polynucleotides in the sample; (d) generating, by the computer system, family identifiers for the more than one family; (e) screening, by the computer system, for a set of common family identifiers, wherein the common family identifiers are family identifiers of the first sample that are the same as or substantially the same as the family identifiers of the second sample; (f) determining, by the computer system, a quantitative measure of the set of common family identifiers; and (g) if the quantitative measure of the set of common family identifiers is higher than a predetermined threshold, classifying, by the computer system, the first sample as contaminated by the second sample, or if the quantitative measure of the common family identifiers is at or below the predetermined threshold, classifying, by the computer system, the first sample as not contaminated.

[0229] In another aspect, the present disclosure provides a method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising: (a) obtaining sequence information comprising more than one sequencing read from the first sample and the second sample; (b) aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping the more than one sequencing read into more than one family based on grouping features, the grouping features comprising at least one of the following (i), (ii), and (iii): (i) the start region, (ii) the end region, and (iii) the length of the sequence read, wherein each family in the sample comprises sequencing reads of progeny polynucleotides amplified from a unique polynucleotide in a collection of polynucleotides in the sample; (d) generating family identifiers for the more than one family; (e) screening for a set of common family identifiers, wherein the common family identifiers are family identifiers of the first sample that are the same as or substantially the same as the family identifiers of the second sample; (f) determining a quantitative measure of the set of common family identifiers; and (g) if the quantitative measure of the set of common family identifiers is higher than a predetermined threshold, classifying the first sample as contaminated by the second sample, or if the quantitative measure of the common family identifiers is at or below the predetermined threshold, classifying the first sample as not contaminated.

[0230] In another aspect, the present disclosure provides a method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising: (a) sequencing a collection of polynucleotides from the sample to generate more than one sequencing read; (b) aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping the more than one sequencing read into more than one family based on grouping features, the grouping features including at least one of the following (i), (ii), and (iii): (i) the start region, (ii) the end region, and (iii) the length of the polynucleotide, wherein each family in the sample includes sequencing reads of daughter polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides from the sample; (d) generating family identifiers for the more than one family; (e) screening out a set of common family identifiers, wherein a common family identifier is a family identifier of the first sample that is the same as or substantially the same as the family identifier of the second sample; (f) determining a quantitative measure of the set of common family identifiers; and (g) classifying the first sample as contaminated by the second sample if the quantitative measure of the set of common family identifiers is higher than a predetermined threshold, or classifying the first sample as not contaminated if the quantitative measure of the common family identifier is at or below the predetermined threshold.

[0231] In some embodiments, prior to sequencing or prior to accessing / obtaining sequence information, the collection of polynucleotides is tagged to produce tagged polynucleotides, wherein each tagged polynucleotide comprises a tag and a polynucleotide. In these embodiments, for each sample, the more than one sequencing read is grouped into more than one family based on grouping features, the grouping features including at least one of the following (i), (ii), (iii), and (iv): (i) the tag, (ii) the start region, (iii) the end region, and (iv) the length of the polynucleotide, wherein each family in the sample includes sequencing reads of daughter polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides from the sample.

[0232] In another aspect, the present disclosure provides a method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising: (a) sequencing a collection of tagged polynucleotides or polynucleotides from a sample to generate more than one sequencing read, wherein each tagged polynucleotide comprises a tag and a polynucleotide; (b) aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping the more than one sequencing read into more than one family based on grouping features, the grouping features including tags, wherein each family in a sample comprises sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the collection of tagged polynucleotides from the sample; (d) generating family identifiers for the more than one family; (e) screening out a set of common family identifiers, wherein a common family identifier is a family identifier of the first sample that is the same as or substantially the same as a family identifier of the second sample; (f) determining a quantitative measure of the set of common family identifiers; and (g) classifying the first sample as contaminated by the second sample if the quantitative measure of the common family identifiers is above a predetermined threshold, or classifying the first sample as uncontaminated if the quantitative measure of the common family identifiers is at or below the predetermined threshold.

[0233] Figure 1 FIG. 4 is a flow chart representation of a method for detecting the presence or absence of contamination between two samples obtained from two different subjects, according to an embodiment of the present disclosure. The grouping features of the sequencing reads, and thus the grouping features of the families, are used to determine the presence or absence of contamination between the two samples. The grouping features of the sequencing reads generally include at least one of the following (i), (ii), (iii), and (iv): (i) tags, (ii) start regions, (iii) end regions, and (iv) lengths of polynucleotides. In 101, a collection of polynucleotides from a sample (i.e., a first sample and a second sample) is sequenced to generate more than one sequencing read. In some embodiments, the first sample and the second sample are sequenced in the same flow cell. In some embodiments, the second sample is sequenced in a flow cell different from the first sample. In some embodiments, the first sample is processed at a different time from the second sample. For example, the second sample is processed at least 1 minute, at least 30 minutes, at least 1 hour, at least 2 hours, at least 3 hours, or at least 4 hours after processing the first sample. In some embodiments, the first sample and the second sample are processed on different dates. In some embodiments, the first sample and the second sample are in the same sample batch. In some embodiments, the second sample is processed with the same batch of reagents as the first sample. In some embodiments, the first sample and the second sample are processed by the same liquid handling robot. In some embodiments, the first sample and the second sample are processed by the same laboratory personnel.

[0234] In some embodiments, the first sample and the second sample are processed at different geographical locations. In some embodiments, the first sample is obtained from the body fluid of one subject, and the second sample is obtained from the body fluid of another subject. In some embodiments, the sample is blood. In some embodiments, the sample is plasma. In some embodiments, the sample is serum. In some embodiments, the polynucleotide is a cell-free polynucleotide. In some embodiments, the cell-free polynucleotide is cell-free DNA. In some embodiments, at least one of the subjects has a disease, such as cancer.

[0235] In some embodiments, the set of polynucleotides undergoes a series of library preparation steps prior to sequencing. The library preparation steps include end repair, ligation of adapters (including tags - i.e., molecular barcodes), amplification of the tagged polynucleotides, and / or selective enrichment of at least a portion of the amplified progeny polynucleotides from regions of the genome or transcriptome of the subject. In some embodiments, the first sample and the second sample are tagged with tags including molecular barcodes to produce a set of tagged polynucleotides. In some embodiments, the set of tagged polynucleotides of the sample is uniquely tagged. In some embodiments, the set of tagged polynucleotides of the sample is non-uniquely tagged. In some embodiments, the method further includes attaching one or more sample indices to one or both ends of the amplified progeny polynucleotides prior to sequencing, wherein the sample indices distinguish the first sample from the second sample.

[0236] To determine a start region, an end region, and / or the length of a polynucleotide, in 102, typically more than one sequencing read is aligned to a reference sequence. The reference sequence can be the human genome. In 103, more than one sequencing read in each sample is grouped into more than one family based on grouping features including at least one of the following (i), (ii), (iii), and (iv): (i) a tag (if the polynucleotide is tagged), (ii) a start region, (iii) an end region, and (iv) the length of the polynucleotide, wherein each family in the sample includes sequencing reads of daughter polynucleotides amplified from a unique polynucleotide in the set of polynucleotides in the sample or tagged daughter polynucleotides (in the case where the polynucleotide is tagged with a molecular barcode). In some embodiments, the start region includes the genomic start position of the sequencing read at which the 5' end of the sequencing read is determined to begin alignment with the reference sequence, and the end region includes the genomic end position of the sequencing read at which the 3' end of the sequencing read is determined to end alignment with the reference sequence. In some embodiments, the start region includes the first 1, first 2, first 5, first 10, first 15, first 20, first 25, first 30, or at least the first 30 base positions of the 5' end of the sequencing read aligned with the reference sequence. In some embodiments, the end region includes the last 1, last 2, last 5, last 10, last 15, last 20, last 25, last 30, or at least the last 30 base positions of the 3' end of the sequencing read aligned with the reference sequence. In some embodiments, the tag includes one or more molecular barcodes attached to both ends of the polynucleotide molecule. In some embodiments, the length of one or more molecular barcodes is at least 2, at least 4, at least 5, at least 6, at least 8, at least 10, at least 15, or at least 20 nucleotides. In some embodiments, the polynucleotides of the sample are tagged with at least 5, at least 10, at least 15, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10,000, at least 50,000, or at least 100,000 different tags / molecular barcodes.

[0237] In 104, family identifiers for more than one family are generated based on the grouping features. In 105, a set of common family identifiers of the family identifiers is filtered out, wherein a common family identifier is a family identifier of a family in a first sample that is the same or substantially the same as the family identifier of a family in a second sample - that is, the grouping features of the family in the first sample are the same or substantially the same as the grouping features of the family in the second sample.

[0238] In 106, a quantitative measure of the set of common family identifiers is determined to classify a sample as being contaminated by another sample or not. In some embodiments, the quantitative measure of the set of common family identifiers is the number of common family identifiers in the first sample. In some embodiments, the quantitative measure of the set of common family identifiers includes the ratio of the number of common family identifiers in the first sample to the total number of family identifiers in the first sample. In some embodiments, the quantitative measure of the set of common family identifiers does not include the following common family identifiers in the first sample: those common family identifiers in the family of the first sample whose number of sequencing reads is greater than the number of sequencing reads of the corresponding family in the second sample. In some embodiments, the quantitative measure of the set of common family identifiers in the first sample does not include the common family identifiers at over-represented genomic start and genomic end position pairs. In some embodiments, the total number of family identifiers in the first sample does not include the family identifiers at over-represented genomic start and genomic end position pairs. In some embodiments, over-represented genomic start and genomic end position pairs are determined by: (a) providing more than one sample, where the more than one sample includes a distribution of genomic start and genomic end positions that is the same as or substantially the same as that of the first sample and / or the second sample; (b) determining the family identifiers in the more than one sample; (c) quantifying the number of family identifiers in the more than one sample that share a pair of genomic start and genomic end positions; and (d) classifying the pair of genomic start and genomic end positions as over-represented if the number of family identifiers exceeds a set threshold. In some embodiments, the more than one sample does not include the first sample or the second sample. In some embodiments, the more than one sample does not include the first sample and the second sample. In some embodiments, the more than one sample includes samples processed in the same flow cell as the first sample. In some embodiments, the more than one sample includes training samples. In some embodiments, the set threshold is at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, or at least 60 families. In some embodiments, the set threshold is about 5 families. In some embodiments, the set threshold is about 10 families. In some embodiments, the set threshold is about 15 families. In some embodiments, the set threshold is about 20 families. In some embodiments, the set threshold is about 30 families. In some embodiments, the set threshold is about 40 families. In some embodiments, the set threshold is about 50 families. In some embodiments, the set threshold can be at least 10 of the total families observed in the more than one sample -3 、at least 10-4 At least 10 -5 At least 10 -6 At least 10 -7 At least 10 -8 Or at least 10 -9 In some embodiments, the set threshold can be about 10 of the total families observed in more than one sample -4 In some embodiments, the set threshold can be about 10 of the total families observed in more than one sample -5 In some embodiments, the set threshold can be about 10 of the total families observed in more than one sample -6 In some embodiments, the set threshold can be about 10 of the total families observed in more than one sample -7 In some embodiments, the set threshold can be about 10 of the total families observed in more than one sample -8 .

[0239] In 107, if the quantitative measure of the shared family identifier is above a predetermined threshold, the first sample is classified as contaminated by the second sample, or if the quantitative measure of the shared family identifier is at or below the predetermined threshold, the first sample is classified as not contaminated. In some embodiments, the predetermined threshold is at least 0.001%, at least 0.005%, at least 0.01%, at least 0.05%, at least 0.1%, at least 0.5%, at least 1%, at least 2%, at least 5% or at least 10% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 0.01% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 0.05% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 0.1% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 0.5% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 1% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 2% of the total number of families in the first sample.

[0240] In some embodiments, even if the first sample is classified as contaminated by the second sample, the method can also allow for reliable detection of at least one somatic variation of the polynucleotide of the first sample by excluding sequencing reads of the shared family identifier of the first sample prior to detecting somatic variations.

[0241] In another aspect, the present disclosure provides a method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising: (a) sequencing a collection of polynucleotides from the sample to generate more than one sequencing read; (b) aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping the more than one sequencing read into more than one family based on information from at least one of (i), (ii), and (iii) below: (i) the start region, (ii) the end region, and (iii) the length of the polynucleotide, wherein each family in the sample comprises sequencing reads of daughter polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides in the sample; (d) screening the more than one family to identify a set of common families, wherein a common family is a family of the first sample that is the same or substantially the same as a family of the second sample; (e) determining a quantitative measure of the set of common families of the first sample; and (f) classifying the first sample as contaminated by the second sample if the quantitative measure of the common families is above a predetermined threshold, or classifying the first sample as not contaminated if the quantitative measure of the common families is at or below the predetermined threshold.

[0242] In another aspect, the present disclosure provides a method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising: (a) sequencing a collection of polynucleotides from the sample to generate more than one sequencing read; (b) aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) grouping the more than one sequencing read of two samples into more than one family based on grouping characteristics, the grouping characteristics comprising at least one of (i), (ii), and (iii) below: (i) the start region, (ii) the end region, and (iii) the length of the polynucleotide, wherein each family in the sample comprises sequencing reads of daughter polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides in the sample; (d) screening the more than one family to identify a set of common families, wherein the common families comprise sequencing reads from the first sample and the second sample; (e) determining a quantitative measure derived from the set of common families; and (f) classifying the first sample as contaminated by the second sample if the quantitative measure of the common families is above a predetermined threshold, or classifying the first sample as not contaminated if the quantitative measure of the common families is at or below the predetermined threshold.

[0243] In some embodiments, prior to sequencing, a collection of polynucleotides can be tagged to produce tagged polynucleotides, where each tagged polynucleotide comprises a tag and a polynucleotide. In these embodiments, for each sample, more than one sequencing read is grouped into more than one family based on grouping features, the grouping features including at least one of the following (i), (ii), (iii), and (iv): (i) the tag, (ii) the start region, (iii) the end region, and (iv) the length of the polynucleotide, where each family in the sample comprises sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the collection of polynucleotides in the sample.

[0244] In another aspect, the present disclosure provides a method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising: (a) sequencing a collection of tagged polynucleotides from a sample to produce more than one sequencing read, where each tagged polynucleotide comprises a tag and a polynucleotide; (b) aligning the more than one sequencing read with a reference sequence, thereby determining an alignment start region and an alignment end region; (c) for each sample, grouping the more than one sequencing read into more than one family based on grouping features, the grouping features including the tag, where each family in the sample comprises sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the collection of tagged polynucleotides in the sample; (d) screening the more than one family to identify a set of common families, where a common family is a family of the first sample that is the same or substantially the same as a family of the second sample; (e) determining a quantitative measure of the set of common families of the first sample; and (f) classifying the first sample as contaminated by the second sample if the quantitative measure of the common families is higher than a predetermined threshold, or classifying the first sample as uncontaminated if the quantitative measure of the common families is at or below the predetermined threshold.

[0245] In another aspect, the present disclosure provides a method for detecting the presence or absence of contamination of a first sample by a second sample, the method comprising: (a) sequencing a collection of tagged polynucleotides from the sample to generate more than one sequencing read, wherein each tagged polynucleotide comprises a tag and a polynucleotide; (b) aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) grouping the more than one sequencing read from the two samples into more than one family based on information from the tags, wherein each family in the sample comprises sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the collection of tagged polynucleotides from the sample; (d) screening the more than one family to identify a set of common families, wherein the common families comprise sequencing reads from the first sample and the second sample; (e) determining a quantitative measure derived from the set of common families; and (f) classifying the first sample as contaminated by the second sample if the quantitative measure of the common families is above a predetermined threshold, or classifying the first sample as uncontaminated if the quantitative measure of the common families is at or below the predetermined threshold.

[0246] Figure 2 is a flow chart representation of a method for detecting the presence or absence of contamination between two samples obtained from two different subjects, according to an embodiment of the present disclosure. The grouping characteristics of the sequencing reads, and thus the grouping characteristics of the families, are used to determine the presence or absence of contamination between the two samples. The grouping characteristics of the sequencing reads generally include at least one of the following (i), (ii), (iii), and (iv): (i) tags, (ii) start regions, (iii) end regions, and (iv) lengths of polynucleotides. In 201, a collection of polynucleotides from the samples (i.e., the first sample and the second sample) is sequenced to generate more than one sequencing read. In some embodiments, the first sample and the second sample are sequenced in the same flow cell. In some embodiments, the second sample is sequenced in a flow cell different from the first sample. In some embodiments, the first sample is processed at a different time from the second sample. For example, the second sample is processed at least 1 minute, at least 30 minutes, at least 1 hour, at least 2 hours, at least 3 hours, or at least 4 hours after processing the first sample. In some embodiments, the first sample and the second sample are processed on different dates. In some embodiments, the first sample and the second sample are in the same sample batch. In some embodiments, the second sample is processed with the same batch of reagents as the first sample.

[0247] In some embodiments, the first sample and the second sample are processed at different geographical locations. In some embodiments, the first sample is obtained from the body fluid of one subject, and the second sample is obtained from the body fluid of another subject. In some embodiments, the sample is blood. In some embodiments, the sample is plasma. In some embodiments, the sample is serum. In some embodiments, the polynucleotide is a cell-free polynucleotide. In some embodiments, the cell-free polynucleotide is cell-free DNA. In some embodiments, at least one of the subjects has a disease, such as cancer.

[0248] In some embodiments, the set of polynucleotides undergoes a series of library preparation steps prior to sequencing. The library preparation steps include end repair, ligation of adapters (including tags - i.e., molecular barcodes), amplification of the tagged polynucleotides, and / or selective enrichment of at least a portion of the amplified progeny polynucleotides from regions of the genome or transcriptome of the subject. In some embodiments, the first sample and the second sample are tagged with tags comprising molecular barcodes to produce a set of tagged polynucleotides. In some embodiments, the set of tagged polynucleotides of the sample is uniquely tagged. In some embodiments, the set of tagged polynucleotides of the sample is non-uniquely tagged. In some embodiments, the method further comprises attaching one or more sample indices to one or both ends of the amplified progeny polynucleotides prior to sequencing, wherein the sample indices distinguish the first sample from the second sample.

[0249] To determine the start region, end region, and / or length of a polynucleotide, in 202, more than one sequencing read is aligned to a reference sequence. The reference sequence can be the human genome (e.g., hg18, hg19). In 203, more than one sequencing read in each sample is grouped into more than one family based on grouping features, where the grouping features include at least one of the following (i), (ii), (iii), and (iv): (i) a tag (if the polynucleotide is tagged), (ii) the start region, (iii) the end region, and (iv) the length of the polynucleotide, where each family in the sample includes sequencing reads of progeny polynucleotides or tagged progeny polynucleotides (in the case where the polynucleotide is tagged with molecular barcodes) amplified from a unique polynucleotide in the set of polynucleotides in the sample. In some embodiments, the start region includes the genomic start position of the sequencing read at which the 5' end of the sequencing read is determined to start aligning with the reference sequence, and the end region includes the genomic end position of the sequencing read at which the 3' end of the sequencing read is determined to stop aligning with the reference sequence. In some embodiments, the start region includes the first 1, first 2, first 5, first 10, first 15, first 20, first 25, first 30, or at least the first 30 base positions of the 5' end of the sequencing read aligned with the reference sequence. In some embodiments, the end region includes the last 1, last 2, last 5, last 10, last 15, last 20, last 25, last 30, or at least the last 30 base positions of the 3' end of the sequencing read aligned with the reference sequence. In some embodiments, the tag includes one or more molecular barcodes attached to both ends of the polynucleotide molecule. In some embodiments, the length of one or more molecular barcodes is at least 2, at least 4, at least 5, at least 6, at least 8, at least 10, at least 15, or at least 20 nucleotides. In some embodiments, the polynucleotides of the sample are tagged with at least 5, at least 10, at least 15, at least 20, at least 50, at least 100, at least 500, at least 1000, at least 5000, at least 10,000, at least 50,000, or at least 100,000 different tags / molecular barcodes.

[0250] In 204, a set of common families of more than one family is selected based on the grouping features, where a common family is a family in a first sample that is the same or substantially the same as a family in a second sample, i.e., the grouping features of the family in the first sample are the same or substantially the same as the grouping features of the family in the second sample.

[0251] In 205, a quantitative measure of the set of shared families is determined to classify a sample as being contaminated by another sample or not. In some embodiments, the quantitative measure of the set of shared families is the number of shared families in the first sample. In some embodiments, the quantitative measure of the set of shared families includes the ratio of the number of shared families in the first sample to the total number of families in the first sample. In some embodiments, the quantitative measure of the set of shared families in the first sample does not include the following shared families in the first sample: those shared families in the families of the first sample whose number of sequencing reads is greater than the number of sequencing reads of the corresponding families in the second sample. In some embodiments, the quantitative measure of the set of shared families in the first sample does not include the shared families at over-represented genomic start and genomic end position pairs. In some embodiments, the total number of families in the first sample does not include the families at over-represented genomic start and genomic end position pairs. In some embodiments, the over-represented genomic start and genomic end position pairs are determined by: (a) providing more than one sample, where the more than one sample includes a distribution of genomic start and genomic end positions that is the same or substantially the same as that of the first sample and / or the second sample; (b) determining the families in the more than one sample; (c) quantifying the number of families in the more than one sample that share a pair of genomic start and genomic end positions; and (d) classifying the pair of genomic start and genomic end positions as over-represented if the number of families exceeds a set threshold. In some embodiments, the more than one sample does not include the first sample or the second sample. In some embodiments, the more than one sample does not include the first sample and the second sample. In some embodiments, the more than one sample includes samples processed in the same flow cell as the first sample. In some embodiments, the more than one sample includes training samples. In some embodiments, the set threshold is at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, or at least 60 families. In some embodiments, the set threshold is about 5 families. In some embodiments, the set threshold is about 10 families. In some embodiments, the set threshold is about 15 families. In some embodiments, the set threshold is about 20 families. In some embodiments, the set threshold is about 30 families. In some embodiments, the set threshold is about 40 families. In some embodiments, the set threshold is about 50 families. In some embodiments, the set threshold can be at least 10 -3 %, at least 10 -4 %, at least 10 -5 %, at least 10 -6, at least 10 -7 , at least 10 -8 or at least 10 -9 . In some embodiments, the set threshold can be about 10 of the total families observed in more than one sample -4 . In some embodiments, the set threshold can be about 10 of the total families observed in more than one sample -5 . In some embodiments, the set threshold can be about 10 of the total families observed in more than one sample -6 . In some embodiments, the set threshold can be about 10 of the total families observed in more than one sample -7 . In some embodiments, the set threshold can be about 10 of the total families observed in more than one sample -8 .

[0252] In 206, if the quantitative measure of the shared family identifier is above a predetermined threshold, the first sample is classified as contaminated by the second sample, or if the quantitative measure of the shared family identifier is at or below the predetermined threshold, the first sample is classified as not contaminated. In some embodiments, the predetermined threshold is at least 0.001%, at least 0.005%, at least 0.01%, at least 0.05%, at least 0.1%, at least 0.5%, at least 1%, at least 2%, at least 5% or at least 10% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 0.01% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 0.05% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 0.1% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 0.5% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 1% of the total number of families in the first sample. In some embodiments, the predetermined threshold is about 2% of the total number of families in the first sample

[0253] In some embodiments, even if the first sample is classified as contaminated by the second sample, the method can still detect at least one somatic genetic variation of the polynucleotides of the first sample by excluding the sequencing reads of the shared families of the first sample, wherein the first sample is classified as contaminated by the second sample

[0254] Figure 3FIG. 0 is a schematic diagram showing the grouping of sequencing reads into families according to an embodiment of the present disclosure and thereby detecting the presence or absence of contamination between two samples (Sample 1 and Sample 2). 301 represents a reference sequence (e.g., hG18 or hG19) to which the sequencing reads of Sample 1 and Sample 2 are aligned. For ease of illustration, Read 1 and Read 2 in the sequencing reads generated from the sequencer by paired-end sequencing are shown as a single paired-end sequencing read, where the Read 1 and Read 2 sequence reads are combined together. The lines with patterned filled boxes at both ends of the line represent paired-end sequencing reads (Read 1 + Read 2). The filled boxes with patterns represent molecular barcodes that have been attached to both ends of the polynucleotide. Each different pattern represents a different molecular barcode sequence. Based on the grouping characteristics, the paired-end sequencing reads are grouped into families. In this embodiment, the grouping characteristics are (i) the tag (i.e., the molecular barcode); (ii) the starting position of the polynucleotide and (iii) the ending position.

[0255] 302A, 303A, 304A, and 305A are common families of Sample 1 because the grouping characteristics of these families are the same or substantially the same as those of families 302B, 303B, 304B, and 305B of Sample 2, respectively. Similarly, 302B, 303B, 304B, and 305B are common families of Sample 2 because the grouping characteristics of these families are the same or substantially the same as those of families 302A, 303A, 304A, and 305A of Sample 1, respectively. 306 represents a pair of genomic starting position and genomic ending position. At 306, Sample 1 has three families and Sample 2 has four families, and thus the total number of families at 306 is seven. In this embodiment, to determine whether a particular pair of genomic starting position and genomic ending position is an over-represented pair, the set threshold value is 6. Since the total number of families at 306 (i.e., 7) is higher than the set threshold, 306 is an over-represented pair of genomic starting position and genomic ending position.

[0256] Scenario I : Determine whether Sample 1 is contaminated by Sample 2.

[0257] The number of common families in Sample 1 is four (302A, 303A, 304A, and 305A), where two families, 302A and 303A, are in over-represented genome start and genome end position pairs. In this embodiment, to determine the quantitative measure of the common families in Sample 1, the common families at the over-represented genome start and genome end position pairs in Sample 1 are excluded. Since 306 is an over-represented pair, two families (302A and 303A) are excluded when calculating the quantitative measure of the common families. Thus, the quantitative measure of the common families in Sample 1 is 2. In this embodiment, the quantitative measure also excludes the following common families in Sample 1: common families in the families of Sample 1 whose number of sequencing reads is greater than the number of sequencing reads in the corresponding families of Sample 2. In this embodiment, the common families (304A and 305A) in Sample 1 each have three paired-end sequencing reads (i.e., six sequencing reads), while the corresponding families (304B and 305B) in Sample 2 each have one paired-end sequencing read (i.e., two sequencing reads). Thus, the common families 304A and 305A are excluded when calculating the quantitative measure. Thus, the quantitative measure of the common families in Sample 1 is zero. To classify Sample 1 as contaminated by Sample 2, the quantitative measure of the common families should be higher than a predetermined threshold. In this embodiment, the predetermined threshold is 0.5% of the total families. Since the quantitative measure (i.e., zero for the first sample) is lower than the predetermined threshold, Sample 1 is determined not to be contaminated by Sample 2.

[0258] Scenario II : Determine whether Sample 2 is contaminated by Sample 1

[0259] The number of common families in Sample 2 is four (302B, 303B, 304B, and 305B), and two of the families, 302B and 303B, are in over-represented genome start and genome end position pairs. In this embodiment, to determine the quantitative measure of the common families in Sample 2, the common families at the over-represented genome start and genome end position pairs in Sample 2 are excluded. Since 306 is an over-represented pair, two families (302B and 303B) are excluded when determining the quantitative measure of the common families. Thus, the quantitative measure of the common families in Sample 2 is 2. In this embodiment, the quantitative measure also excludes the following common families in Sample 2: the common families in the families of Sample 2 whose number of sequencing reads is greater than the number of sequencing reads in the corresponding families of Sample 1. In this embodiment, each of the common families (304B and 305B) in Sample 2 has one paired-end sequencing read (i.e., two sequencing reads), while each of the corresponding families (304A and 305A) in Sample 1 has three paired-end sequencing reads (i.e., six sequencing reads). Thus, the common families 304B and 305B are not excluded when calculating the quantitative measure. Thus, the quantitative measure of the common families in Sample 2 is 2. To classify Sample 2 as contaminated by Sample 1, the quantitative measure of the common families in Sample 2 should be higher than a predetermined threshold. In this embodiment, the predetermined threshold is 0.5% of the total families. For Sample 2, the total number of families is 21. In this embodiment, the families at the over-represented genome start and genome start position pairs are not included in the total number of families. The number of families at the over-represented genome start and genome end position pair 306 is 4. Thus, after excluding the families at the over-represented pairs, the total number of families in Sample 2 is 17. Further, in this embodiment, the quantitative measure of the common families is the percentage of the common families in Sample 2 in the total families, and this percentage is equal to 11.765% (100 * 2 / 17), and is higher than the predetermined threshold. Thus, Sample 2 is determined to be contaminated by Sample 1.

[0260] The various steps of the method can be performed at the same or different times, in the same or different geographical locations (e.g., countries), and by the same or different people or entities.

[0261] II. General Features of the Method

[0262] A. Sample

[0263] The sample can be any biological sample isolated from a subject. The sample can include body tissue, whole blood, platelets, serum, plasma, feces, red blood cells, white blood cells (leucocytes), endothelial cells, tissue biopsies (e.g., biopsies from known or suspected solid tumors), cerebrospinal fluid, synovial fluid, lymphatic fluid, ascites, interstitial fluid or extracellular fluid (e.g., fluid from the interstitial space), gingival fluid, sulcular fluid, bone marrow, pleural effusion, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, urine. The sample is preferably a body fluid, particularly blood and its fractions, and urine. Such samples include nucleic acids shed from tumors. The nucleic acids can include DNA and RNA and can be in double-stranded and single-stranded forms. The sample can be in the form initially isolated from the subject or can be subjected to further processing to remove or add components such as cells, enrich one component relative to another, or convert one form of nucleic acid to another, such as converting RNA to DNA or single-stranded nucleic acid to double-stranded. Thus, for example, the body fluid for analysis is plasma or serum containing cell-free nucleic acids such as cell-free DNA (cfDNA). In some embodiments, the method includes obtaining a sample from a subject. Optionally, substantially any sample type can be used. In certain embodiments, for example, the sample is tissue, blood, plasma, serum, sputum, urine, semen, vaginal fluid, feces, synovial fluid, spinal fluid, saliva, and / or the like. Generally, the subject is a mammalian subject (e.g., a human subject). In some embodiments, the sample is blood. In some embodiments, the sample is plasma. In some embodiments, the sample is serum.

[0264] In some embodiments, the volume of the sample of body fluid taken from the subject depends on the desired read depth of the sequencing region. Exemplary volumes are about 0.4 ml - 40 ml, about 5 ml - 20 ml, about 10 ml - 20 ml. For example, the volume can be about 0.5 ml, about 1 ml, about 5 ml, about 10 ml, about 20 ml, about 30 ml, about 40 ml or more milliliters. The volume of plasma sampled is typically between about 5 ml and about 20 ml.

[0265] The sample can contain various amounts of nucleic acids. Generally, the amount of nucleic acid in a given sample is equivalent to a plurality of genome equivalents. For example, a sample of about 30 ng of DNA can contain about 10,000 (10 4 ) haploid human genome equivalents, and in the case of cfDNA, can contain about 200 billion (2×10 11 ) individual polynucleotide molecules. Similarly, a sample of about 100 ng of DNA can contain about 30,000 haploid human genome equivalents, and in the case of cfDNA, can contain about 600 billion individual molecules.

[0266] In some embodiments, the sample comprises nucleic acids from different sources, e.g., from a cellular source and from a cell-free source (e.g., a blood sample, etc.). Generally, the sample comprises nucleic acids carrying mutations. For example, the sample optionally comprises DNA carrying germline mutations and / or somatic mutations. Generally, the sample comprises DNA carrying cancer-related mutations (e.g., cancer-related somatic mutations). In some embodiments, the sample comprises cell-free DNA (i.e., a cfDNA sample). In some embodiments, the cfDNA sample comprises circulating tumor nucleic acids.

[0267] Exemplary amounts of cell-free nucleic acids in a sample prior to amplification typically range from about 1 femtogram (fg) to about 1 microgram (μg), e.g., from about 1 picogram (pg) to about 200 nanograms (ng), from about 1 ng to about 100 ng, from about 10 ng to about 1000 ng. In some embodiments, the sample comprises up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. Optionally, the amount is at least about 1 fg, at least about 10 fg, at least about 100 fg, at least about 1 pg, at least about 10 pg, at least about 100 pg, at least about 1 ng, at least about 10 ng, at least about 100 ng, at least about 150 ng, or at least about 200 ng of cell-free nucleic acid molecules. In certain embodiments, the amount is up to about 1 fg, about 10 fg, about 100 fg, about 1 pg, about 10 pg, about 100 pg, about 1 ng, about 10 ng, about 100 ng, about 150 ng, or about 200 ng of cell-free nucleic acid molecules. In some embodiments, the method comprises obtaining from the sample between about 1 fg and about 200 ng of cell-free nucleic acid molecules. In certain embodiments, the method comprises obtaining from the sample between about 5 ng and about 30 ng of cell-free nucleic acid molecules. In certain embodiments, the method comprises obtaining from the sample between about 5 ng and about 100 ng of cell-free nucleic acid molecules. In certain embodiments, the method comprises obtaining from the sample between about 5 ng and about 150 ng of cell-free nucleic acid molecules. In certain embodiments, the method comprises obtaining from the sample between about 5 ng and about 200 ng of cell-free nucleic acid molecules. In some embodiments, the amount is up to about 100 ng of cell-free nucleic acid molecules from the sample. In some embodiments, the amount is up to about 150 ng of cell-free nucleic acid molecules from the sample. In some embodiments, the amount is up to about 200 ng of cell-free nucleic acid molecules from the sample. In some embodiments, the amount is up to about 250 ng of cell-free nucleic acid molecules from the sample. In some embodiments, the amount is up to about 300 ng of cell-free nucleic acid molecules from the sample. In some embodiments, the method comprises obtaining from the sample between about 1 fg and about 200 ng of cell-free nucleic acid molecules.

[0268] Cell-free nucleic acids generally have a size distribution between about 100 nucleotides in length and about 500 nucleotides in length, where molecules between about 110 nucleotides in length and about 230 nucleotides in length represent about 90% of the molecules in the sample, where the mode is about 168 nucleotides in length, and a second minor peak is in the range between about 240 nucleotides and about 440 nucleotides in length. In certain embodiments, the cell-free nucleic acids are about 160 nucleotides to about 180 nucleotides in length, or about 320 nucleotides to about 360 nucleotides in length, or about 440 nucleotides to about 480 nucleotides in length.

[0269] In some embodiments, cell-free nucleic acids are isolated from a body fluid by a partitioning step in which cell-free nucleic acids found in solution are separated from intact cells and other insoluble components in the body fluid. In some of these embodiments, partitioning includes techniques such as centrifugation or filtration. Alternatively, the cells in the body fluid are lysed and the cell-free nucleic acids and cellular nucleic acids are processed together. Generally, after adding a buffer and a washing step, the cell-free nucleic acids are precipitated with, for example, alcohol. In certain embodiments, additional clean up steps such as silica-based columns are used to remove contaminants or salts. For example, non-specific bulk carrier nucleic acids are optionally added throughout the reaction to optimize certain aspects of the exemplary procedure such as yield. After such processing, the sample generally contains nucleic acids in various forms, including double-stranded DNA, single-stranded DNA, and / or single-stranded RNA. Optionally, the single-stranded DNA and / or single-stranded RNA are converted to a double-stranded form such that they are included in subsequent processing and analysis steps.

[0270] B. Nucleic Acid Tags

[0271] In some embodiments, nucleic acid molecules (from a sample of polynucleotides) can be tagged with sample indices and / or molecular barcodes (commonly referred to as "tags"). The tags can be incorporated into adapters or otherwise linked to adapters by chemical synthesis, ligation (e.g., blunt-end ligation or sticky-end ligation), or overlap extension polymerase chain reaction (PCR) and other methods. Such adapters can ultimately be ligated to target nucleic acid molecules. In other embodiments, typically one or more rounds of amplification cycles (e.g., PCR amplification) are applied to introduce sample indices into nucleic acid molecules using conventional nucleic acid amplification methods. The amplification can be performed in one or more reaction mixtures (e.g., more than one microwell in an array). The molecular barcodes and / or sample indices can be introduced simultaneously or in any order. In some embodiments, the molecular barcodes and / or sample indices are introduced before and / or after a sequence capture step. In some embodiments, only the molecular barcode is introduced before probe capture, and the sample index is introduced after the sequence capture step. In some embodiments, both the molecular barcode and the sample index are introduced before a probe-based capture step. In some embodiments, the sample index is introduced after the sequence capture step. In some embodiments, the molecular barcode is incorporated into nucleic acid molecules (e.g., cfDNA molecules) in the sample via ligation (e.g., blunt-end ligation or sticky-end ligation) through an adapter. In some embodiments, the sample index is incorporated into nucleic acid molecules (e.g., cfDNA molecules) in the sample by overlap extension polymerase chain reaction (PCR). Typically, a sequence capture protocol involves introducing single-stranded nucleic acid molecules complementary to the nucleic acid sequences being targeted, such as the coding sequences of genomic regions, and mutations in such regions are associated with cancer types.

[0272] In some embodiments, the tags can be located at one or both ends of the sample nucleic acid molecules. In some embodiments, the tags are predetermined or random or semi-random sequence oligonucleotides. In some embodiments, the length of the tags can be less than about 500, 200, 100, 50, 20, 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 nucleotide. The tags can be randomly or non-randomly linked to the sample nucleic acid.

[0273] In some embodiments, each sample is uniquely tagged with a sample index or a combination of sample indexes. In some embodiments, each nucleic acid molecule of a sample or subsample is uniquely tagged with a molecular barcode or a combination of molecular barcodes. In other embodiments, more than one molecular barcode can be used such that the molecular barcodes are not necessarily unique to each other among the more than one molecular barcodes (e.g., non-unique molecular barcodes). In these embodiments, the molecular barcodes are typically attached (e.g., by ligation) to individual molecules such that the combination of the molecular barcode and the sequence to which it can be attached results in a unique sequence that can be individually traced. Detection of the combination of non-uniquely tagged molecular barcodes with endogenous sequence information (e.g., corresponding to the start (initiation) and / or end (termination) portions of the original nucleic acid molecule sequence in the sample, subsequences of sequence reads at one or both ends, the length of the sequence reads, and / or the length of the original nucleic acid molecule in the sample) generally allows assignment of a unique identity to a particular molecule. The length or number of base pairs of an individual sequence read is also optionally used to assign a unique identity to a given molecule. As described herein, fragments of single strands of nucleic acids that have been assigned a unique identity can thus allow subsequent identification of fragments from the parental strand and / or complementary strand.

[0274] In some embodiments, the molecular barcodes are introduced at a ratio to the molecules in the sample of a set of desired identifiers (e.g., a combination of unique molecular barcodes or non-unique molecular barcodes). One example form uses from about 2 to about 1,000,000 different molecular barcodes, or from about 5 to about 150 different molecular barcodes, or from about 20 to about 50 different molecular barcodes, that are ligated to both ends of a target molecule. Alternatively, from about 25 to about 1,000,000 different molecular barcodes can be used. For example, 20 - 50 × 20 - 50 molecular barcodes can be used. In some embodiments, from 20 to 50 different molecular barcodes can be used. In some embodiments, from 5 to 100 different molecular barcodes can be used. In some embodiments, from 5 to 150 molecular barcodes can be used. In some embodiments, from 5 to 200 different molecular barcodes can be used. Such a number of identifiers is generally sufficient to give different molecules with the same start and end points a high probability (e.g., at least 94%, 99.5%, 99.99%, or 99.999%) of receiving different combinations of identifiers. In some embodiments, about 80%, about 90%, about 95%, or about 99% of the molecules have the same molecular barcode combination.

[0275] In some embodiments, the assignment of unique or non-unique molecular barcodes in a reaction is performed using methods and systems described, for example, in U.S. Patent Application Nos. 20010053519, 20030152490, and 20110160078, and U.S. Patents Nos. 6,582,908, 7,537,898, 9,598,731, and 9,902,992, each of which is incorporated herein by reference in its entirety. Optionally, in some embodiments, only endogenous sequence information (e.g., start position and / or end position, subsequences at one or both ends of the sequence, and / or length) can be used to identify different nucleic acid molecules of a sample.

[0276] C. Amplification

[0277] Sample nucleic acids flanked by adapters are typically amplified by PCR and other amplification methods using nucleic acid primers that bind to primer binding sites in the adapters flanking the DNA molecule to be amplified. In some embodiments, the amplification method involves cycles of extension, denaturation, and annealing generated by thermal cycling, or can be isothermal, for example, in transcription-mediated amplification. Other exemplary amplification methods optionally used include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and self-sustained sequence replication, and other methods.

[0278] Typically, one or more rounds of amplification cycles are applied to introduce molecular barcodes and / or sample indices into nucleic acid molecules using conventional nucleic acid amplification methods. Amplification is typically performed in one or more reaction mixtures. The molecular barcodes and sample indices are optionally introduced simultaneously or in any order. In some embodiments, the molecular barcodes and sample indices are introduced before and / or after the sequence capture step. In some embodiments, only the molecular barcode is introduced before probe capture, and the sample index is introduced after the sequence capture step. In certain embodiments, both the molecular barcode and the sample index are introduced before the probe-based capture step. In some embodiments, the sample index is introduced after the sequence capture step. Typically, the sequence capture protocol involves introducing single-stranded nucleic acid molecules complementary to the targeted nucleic acid sequences, such as the coding sequences of genomic regions, and mutations in such regions are associated with cancer types. Typically, the amplification reaction produces more than one non-unique or uniquely tagged nucleic acid amplicon with molecular barcodes and sample indices, and the molecular barcodes and sample indices range in size from about 200 nucleotides (nt) to about 700 nt, 250 nt to about 350 nt, or about 320 nt to about 550 nt. In some embodiments, the amplicon has a size of about 300 nt. In some embodiments, the amplicon has a size of about 500 nt.

[0279] D. Enrichment

[0280] Sequences can be enriched prior to sequencing. Enrichment can be performed for specific target regions or non-specifically (“target sequences”). In some embodiments, target regions of interest can be enriched using capture probes (“baits”) selected against one or more baitset panels using differential tiling and capture schemes. Differential tiling and capture schemes use different relative concentrations of baitset panels to differentially tile across genomic regions associated with the baits (e.g., at different “resolutions”), subject to a set of constraints (e.g., sequencer constraints such as sequencing throughput, utility of each bait, etc.), and capture them at levels desired for downstream sequencing. These targeted genomic regions of interest can include native nucleotide sequences or synthetic nucleotide sequences of nucleic acid constructs. In some embodiments, biotinylated beads with probes for one or more regions of interest can be used to capture target sequences, optionally followed by amplification of these regions to enrich the regions of interest.

[0281] Sequence capture can include using oligonucleotide probes that hybridize to target sequences. Probe set strategies can involve tiling probes across regions of interest. Such probes can be, for example, about 60 to 120 bases in length. The set can have a depth of about 2x, 3x, 4x, 5x, 6x, 8x, 9x, 10x, 15x, 20x, 50x, or more than 50x. The effectiveness of sequence capture depends in part on the length of the sequences in the target molecule that are complementary (or nearly complementary) to the sequences of the probes.

[0282] In some embodiments, more than one genomic region contains genetic variants found in COSMIC, The Cancer Genome Atlas (TCGA), or the Exome Aggregation Consortium (ExAC). In some cases, the genetic variants may belong to a predefined set of clinically actionable variants. For example, such variants can be found in various variant databases, the presence of which in a subject's sample has been shown to be associated with or indicative of a disease or disorder (e.g., cancer) in the subject. Such variant databases can include, for example, the Catalogue of Somatic Mutations in Cancer (COSMIC), The Cancer Genome Atlas (TCGA), and the Exome Aggregation Consortium (ExAC). A predefined set of such classified variants can be designated for further bioinformatics analysis due to their relevance to clinical decision-making (e.g., diagnosis, prognosis, treatment selection, targeted therapy, treatment monitoring, recurrence monitoring, etc.). Such a predefined set can be determined based on, for example, the analysis of clinical samples (e.g., clinical samples of patient cohorts known to have or not have a disease or disorder) and annotation information from public databases and clinical literature.

[0283] E. Sequencing

[0284] Sample nucleic acids flanked by adaptors, with or without pre-amplification, can be subjected to sequencing. Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, sequencing by synthesis, single molecule sequencing, nanopore sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, RNA-Seq (Illumina), digital gene expression (Helicos), next-generation sequencing, single molecule synthesis sequencing (SMSS) (Helicos), massively parallel sequencing, clonal single molecule array (Solexa), shotgun sequencing, Ion Torrent, Oxford nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, sequencing using PacBio, SOLiD, Ion Torrent, or nanopore platforms. The sequencing reaction can be carried out in a number of sample processing units, which can include multiple lane, multi-channel, multi-well, or other means for processing multiple groups of samples substantially simultaneously. The sample processing unit can also include multiple sample chambers to enable simultaneous processing of multiple runs.

[0285] Sequencing reactions can be performed on one or more nucleic acid fragment types or regions known to contain markers for cancer or other diseases. Sequencing reactions can also be performed on any nucleic acid fragments present in a sample. Sequencing reactions can be performed on at least about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of the genome. In other cases, sequencing reactions can be performed on less than about 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9% or 100% of the genome.

[0286] Simultaneous sequencing reactions can be performed using multiplex sequencing techniques. In some cases, cell-free polynucleotides can be sequenced with at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000 or 100,000 sequencing reactions. In other cases, cell-free polynucleotides can be sequenced with fewer than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000 or 100,000 sequencing reactions. Sequencing reactions can be performed sequentially or simultaneously. Subsequent data analysis can be performed on all or a portion of the sequencing reactions. In some cases, data analysis can be performed on at least 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000 or 100,000 sequencing reactions. In other cases, data analysis can be performed on fewer than 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 50000 or 100,000 sequencing reactions. An exemplary read depth is 1000 - 50000 reads / locus (base). In some embodiments, the read depth can be greater than 50000 reads / locus (base).

[0287] F. Analysis

[0288] Sequencing according to embodiments of the present invention generates more than one sequencing read or read. Sequencing reads or reads according to the present invention typically comprise sequences of nucleotide data having a length less than about 150 bases or a length less than about 90 bases. In certain embodiments, the length of the read is between about 80 bases and about 90 bases, such as about 85 bases. In some embodiments, the methods of the present invention are applied to very short reads, i.e., having a length less than about 50 bases or about 30 bases. Sequencing read data can include sequence data as well as meta-information. Sequence read data can be stored in any suitable file format, including, for example, VCF files, FASTA files, or FASTQ files.

[0289] FASTA was originally a computer program for retrieving sequence databases, and the name FASTA also refers to the standard file format. See Pearson & Lipman, 1988, Improved tools for biological sequence comparison, PNAS 85:2444 - 2448. Sequences in FASTA format begin with a single-line description, followed by lines of sequence data. The description line is distinguished from the sequence data by a greater-than (“>”) symbol in the first column. The word following the “>” symbol is the identifier of the sequence, and the remainder of the line is a description (both are optional). There should be no space between the “>” and the first letter of the identifier. It is recommended that all lines of text be shorter than 80 characters. If another line starting with “>” appears, the sequence ends; this indicates the start of another sequence.

[0290] The FASTQ format is a text-based format for storing biological sequences (usually nucleotide sequences) and their corresponding quality scores. The FASTQ format is similar to the FASTA format, but has quality scores following the sequence data. For brevity, both the sequence letters and the quality scores are encoded using a single ASCII character. The FASTQ format is the de facto standard for storing the output of high-throughput sequencing instruments such as Illumina genome analyzers, as described, for example, by Cock et al. (“The Sanger FASTQ file format for sequences with quality scores, and the Solexa / Illumina FASTQ variants,” Nucleic Acids Res 38(6):1767 - 1771, 2009), which is hereby incorporated by reference in its entirety.

[0291] For FASTA and FASTQ files, the meta-information includes the description lines and not the sequence data lines. In some embodiments, for FASTQ files, the meta-information includes quality scores. For FASTA and FASTQ files, the sequence data begins after the description line and is typically presented using some subset of the IUPAC ambiguity codes optionally with "-". In a preferred embodiment, the sequence data will use the A, T, C, G, and N characters, optionally including "-" as needed or including U (e.g., to represent a gap or uracil).

[0292] In some embodiments, at least one primary sequence read file and the output file are stored as plain text files (e.g., using an encoding such as ASCII; ISO / IEC 646; EBCDIC; UTF-8; or UTF-16). The computer system provided by the present invention may include a text editor program capable of opening plain text files. A text editor program may refer to a computer program capable of presenting the content of a text file (such as a plain text file) on a computer screen and allowing a person to edit the text (e.g., using a screen, keyboard, and mouse). Exemplary text editors include, but are not limited to, Microsoft Word, emacs, pico, vi, BBEdit, and TextWrangler. Preferably, the text editor program is capable of displaying the plain text file on a computer screen, displaying the meta-information and sequence reads in a human-readable format (e.g., not binary encoded but using alphanumeric characters since alphanumeric characters can be used to print human writing).

[0293] Although the methods have been discussed with reference to FASTA or FASTQ files, the methods and systems of the present invention can be used to compress any suitable sequence file format, including, for example, files in the Variant Call Format (VCF) format. A typical VCF file includes a header section and a data section. The header contains any number of meta-information lines, each starting with the character '##', and a TAB-delimited field definition line starting with a single '#' character. The field definition line names eight required columns, and the body section contains data lines that populate the columns defined by the field definition line. The VCF format is described by Danecek et al. ("The variant call format and VCFtools," Bioinformatics 27(15):2156-2158, 2011), which is hereby incorporated by reference in its entirety. The header section can be considered the meta-information to be written to the compressed file, and the data section can be considered lines, where each line is stored in the main file only if it is unique.

[0294] Certain embodiments of the present invention provide for the assembly of sequencing reads. For example, in assembly by alignment, sequencing reads are aligned to each other or to a reference sequence. By aligning each read, and then to a reference genome, all reads are positioned relative to each other to produce an assembly. In addition, aligning or mapping sequencing reads to a reference sequence can also be used to identify variant sequences in the sequencing reads. Identifying variant sequences can be used in combination with the methods and systems described herein to further assist in the diagnosis or prognosis of a disease or condition or for guiding treatment decisions.

[0295] In some embodiments, any or all steps are automated. Optionally, the methods of the present invention can be implemented in whole or in part in one or more dedicated programs, for example each dedicated program optionally written in a compiled language such as C++ and then compiled and distributed in binary. The methods of the present invention can be implemented in whole or in part as a module within an existing sequence analysis platform or by invoking functions within an existing sequence analysis platform. In certain embodiments, the methods of the present invention include multiple steps that are automatically invoked in response to a single starting cue (e.g., one event or combination of events from a triggering event originating from human activity, another computer program, or a machine). Thus, the present invention provides methods in which any one or any combination of steps can occur automatically in response to a signal. Automatically generally means without intervening human input, influence, or interaction (i.e., only in response to original or pre-cued human activity).

[0296] The system also includes outputs in various forms, which include accurate and sensitive interpretation of the subject nucleic acid. The output of the retrieval can be provided in the format of a computer file. In certain embodiments, the output is a FASTA file, a FASTQ file, or a VCF file. The output can be processed to generate a text file or an XML file containing sequence data (such as the sequence of a nucleic acid aligned to a reference genome). In other embodiments, the processing generates an output containing coordinates or strings that describe one or more mutations in the subject nucleic acid relative to the reference genome. Alignment strings can include Simple UnGapped Alignment Report (SUGAR), Verbose Useful Labeled Gapped Alignment Report (VALGAR), and Compact Idiosyncratic Gapped Alignment Report (CIGAR) (Ning et al., Genome Research 11(10):1725-9, 2001, which is hereby incorporated by reference in its entirety). These strings are implemented, for example, in the Exonerate sequence alignment software from the European Bioinformatics Institute (Hinxton, UK).

[0297] In some embodiments, a sequence alignment containing a CIGAR string is generated—such as, for example, a Sequence Alignment / Map (SAM) or Binary Alignment / Map (BAM) file (the SAM format is described, for example, by Li et al., “The Sequence Alignment / Map format and SAMtools,” Bioinformatics, 25(16):2078-9, 2009, which is incorporated by reference in its entirety). In some embodiments, the CIGAR shows or includes one gap alignment per line. CIGAR is a compressed pairwise alignment format reported as a CIGAR string. The CIGAR string can be used to present long (e.g., genomic) pairwise alignments. The CIGAR string is used in the SAM format to represent the alignment of a read to a reference genome sequence.

[0298] The CIGAR string follows an established motif. Each character is preceded by a number giving the base count of the event. The characters used can include M, I, D, N, and S (M = match; I = insertion; D = deletion; N = gap; S = substitution). The CIGAR string defines the sequence of matches / mismatches and deletions (or gaps). For example, the CIGAR string 2MD3M2D2M would mean that the alignment contains 2 matches, 1 deletion (the number 1 is omitted for some space savings), 3 matches, 2 deletions, and 2 matches.

[0299] In some embodiments, nucleic acid populations for sequencing are prepared by enzymatically forming blunt ends on double-stranded nucleic acids having single-stranded overhangs at one or both ends. In these embodiments, the population is typically treated with an enzyme having 5'-3' DNA polymerase activity and 3'-5' exonuclease activity in the presence of nucleotides in the form of dNTPs (e.g., A, C, G, and T or U). Exemplary enzymes or catalytic fragments thereof that may optionally be used include the Klenow fragment and T4 polymerase. At a 5' overhang, the enzyme typically extends the recessed 3' end on the opposing strand until it is flush with the 5' end to produce a blunt end. At a 3' overhang, the enzyme typically digests from the 3' end up to and sometimes beyond the 5' end of the opposing strand. If the digestion proceeds beyond the 5' end of the opposing strand, the gap can be filled with an enzyme having the same polymerase activity as that used for the 5' overhang. The formation of blunt ends on double-stranded nucleic acids facilitates, for example, the attachment of adapters and subsequent amplification.

[0300] In some embodiments, the nucleic acid population is subjected to additional processing, such as converting single-stranded nucleic acids to double-stranded nucleic acids and / or converting RNA to DNA. These forms of nucleic acids are also optionally ligated to adapters and amplified.

[0301] With or without prior amplification, the nucleic acids that have been subjected to the blunt-end forming process described above and optionally other nucleic acids in the sample can be sequenced to produce sequenced nucleic acids. A sequenced nucleic acid can refer to the sequence of a nucleic acid (i.e., sequence information) or a nucleic acid whose sequence has been determined. Sequencing can be performed so as to provide sequence data for individual nucleic acid molecules in the sample either directly or indirectly from the consensus sequence of the amplification products of the individual nucleic acid molecules in the sample.

[0302] In some embodiments, double-stranded nucleic acids in the sample having single-stranded overhangs are ligated at both ends to adapters containing molecular barcodes after blunt-end formation, and sequencing determines the nucleic acid sequence as well as the molecular barcodes introduced by the adapters. Blunt-end DNA molecules are optionally ligated to the blunt ends of at least partially double-stranded adapters (e.g., Y-shaped adapters or bell-shaped adapters). Alternatively, the blunt ends of the sample nucleic acids and adapters can be tailed with complementary nucleotides to facilitate ligation (e.g., sticky-end ligation).

[0303] A nucleic acid sample is typically contacted with a sufficient number of adaptors such that the probability that any two copies of the same nucleic acid receive the same adaptor barcode combination from the adaptors ligated at both ends is low (e.g., less than <1% or <0.1%). Using adaptors in this way allows for the identification of families of nucleic acid sequences that have the same starting and ending points on a reference nucleic acid and are ligated to the same molecular barcode combination. Such families represent the amplified product sequences of the nucleic acids in the sample prior to amplification. The sequences of the family members can be assembled to obtain one or more consensus nucleotides or a complete consensus sequence of the nucleic acid molecules in the original sample, which nucleic acid molecules have been modified by blunt-end formation and adaptor attachment. In other words, the nucleotide occupying a particular position in the nucleic acids of the sample is determined to be the consensus nucleotide of the nucleotides occupying the corresponding position in the family member sequences. A family can include the sequences of one or both strands of a double-stranded nucleic acid. If the members of a family include the sequences of both strands of a double-stranded nucleic acid, for the purpose of assembling all of the sequences to obtain one or more consensus nucleotides or sequences, the sequences of one strand can be converted to their complementary sequences. Some families contain only a single member sequence. In this case, the sequence can be considered to be the sequence of the nucleic acid in the sample prior to amplification. Optionally, families with only a single member sequence can be excluded from subsequent analysis.

[0304] By comparing the sequenced nucleic acid with a reference sequence, nucleotide variations in the sequenced nucleic acid can be determined. The reference sequence is typically a known sequence, for example, a known whole or partial genomic sequence from a subject (e.g., the whole genome sequence of a human subject). The reference sequence can be, for example, hG19 or hG38. As described above, the sequenced nucleic acid can represent the sequence of the nucleic acid in a directly determined sample or the consensus sequence of the amplification products of such nucleic acids. The comparison can be made at one or more designated positions on the reference sequence. When the corresponding sequences are maximally aligned, a subset of the sequenced nucleic acids can be identified that includes the positions corresponding to the designated positions on the reference sequence. In such a subset, it can be determined which (if any) of the sequenced nucleic acids include nucleotide variations at the designated positions and optionally which (if any) include the reference nucleotides (i.e., the same as in the reference sequence). If the number of sequenced nucleic acids in the subset that include nucleotide variations exceeds a selected threshold, the variant nucleotides can be identified at the designated positions. The threshold can be a simple number, such as at least 1, 2, 3, 4, 5, 6, 7, 9, or 10 sequenced nucleic acids in the subset that include nucleotide variations, or the threshold can be a ratio, such as at least 0.5%, 1%, 2%, 3%, 4%, 5%, 10%, 15%, or 20% of the sequenced nucleic acids in the subset that include nucleotide variations, and other possibilities. The repeated comparison can be made for any designated position of interest in the reference sequence. Sometimes the comparison can be made for designated positions that occupy at least about 20, 100, 200, or 300 consecutive positions on the reference sequence, for example, about 20 - 500 or about 50 - 300 consecutive positions.

[0305] Additional details regarding nucleic acid sequencing, including the forms and applications described herein, are also provided in the following references: e.g., Levy et al., Annual Review of Genomics and Human Genetics, 17:95-115 (2016); Liu et al., J. of Biomedicine and Biotechnology, 2012, Article ID 251364:1-11 (2012); Voelkerding et al., Clinical Chem., 55:641-658 (2009); MacLean et al., Nature Rev. Microbiol., 7:287-296 (2009); Astier et al., J Am Chem Soc., 128(5):1705-10 (2006); U.S. Patent No. 6,210,891, U.S. Patent No. 6,258,568, U.S. Patent No. 6,833,246, U.S. Patent No. 7,115,400, U.S. Patent No. 6,969,488, U.S. Patent No. 5,912,148, U.S. Patent No. 6,130,073, U.S. Patent No. 7,169,560, U.S. Patent No. 7,282,337, U.S. Patent No. 7,482,120, U.S. Patent No. 7,501,245, U.S. Patent No. 6,818,395, U.S. Patent No. 6,911,345, U.S. Patent No. 7,501,245, U.S. Patent No. 7,329,492, U.S. Patent No. 7,170,050, U.S. Patent No. 7,302,146, U.S. Patent No. 7,313,308, and U.S. Patent No. 7,476,503, each of which is incorporated herein by reference in its entirety.

[0306] III. Computer System

[0307] The methods of the present disclosure can be implemented using or by means of a computer system. For example, such a method can include (a) obtaining more than one sequencing read from a set of tagged polynucleotides from a first sample and a second sample generated by a nucleic acid sequencer, wherein the sequencing read includes a tag sequence and a sequence derived from the polynucleotide; (b) aligning the more than one sequencing read with a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping the more than one sequencing read into more than one family based on grouping features, the grouping features including at least one of the following (i), (ii), (iii), and (iv): (i) a tag, (ii) a start region, (iii) an end region, and (iv) the length of the polynucleotide, wherein each family in the sample includes sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the set of polynucleotides in the sample; (d) generating family identifiers for the more than one family; (e) screening out a set of common family identifiers, wherein the common family identifiers are family identifiers of the first sample that are the same as or substantially the same as the family identifiers of the second sample; (f) determining a quantitative measure of the set of common family identifiers; and (g) classifying the first sample as contaminated by the second sample if the quantitative measure of the common family identifiers is higher than a predetermined threshold, or classifying the first sample as uncontaminated if the quantitative measure of the common family identifiers is at or below the predetermined threshold, and the method can be executed by a computer processor.

[0308] Figure 4 Shown is a computer system 401 programmed or otherwise configured to implement the methods of the present disclosure. The computer system 401 can regulate various aspects of sample preparation, sequencing, and / or analysis. In some instances, the computer system 401 is configured to perform sample preparation and sample analysis, including nucleic acid sequencing.

[0309] The computer system 401 includes a central processing unit (CPU, also referred to herein as "processor" and "computer processor") 405, which can be a single-core processor, a multi-core processor, or more than one processor for parallel processing. The computer system 401 also includes a memory or memory location 410 (e.g., random access memory, read-only memory, flash memory), an electronic storage unit 415 (e.g., hard disk), a communication interface 420 for communicating with one or more other systems (e.g., network adapter), and peripheral devices 425, such as a cache, other memories, data storage, and / or an electronic display adapter. The memory 410, storage unit 415, interface 420, and peripheral devices 425 communicate with the CPU 405 via a communication network or bus (solid lines), such as a motherboard. The storage unit 415 can be a data storage unit (or data repository) for storing data. The computer system 401 can be operatively coupled to a computer network 430 via the communication interface 420. The computer network 430 can be the Internet, an intranet, and / or an extranet, or an intranet and / or an extranet that communicates with the Internet. In some cases, the computer network 430 is a telecommunications and / or data network. The computer network 430 can include one or more computer servers, which can enable distributed computing, such as cloud computing. In some cases, via the computer system 401, the computer network 430 can implement a peer-to-peer network, which can enable devices coupled to the computer system 401 to operate as clients or servers.

[0310] The CPU 405 can execute a series of machine-readable instructions, which can be implemented as a program or software. The instructions can be stored in a memory location, such as in the memory 410. Examples of operations performed by the CPU 405 can include reading, decoding, executing, and writing back.

[0311] The storage unit 415 can store files, such as drivers, libraries, and saved programs. The storage unit 415 can store user-generated programs and recorded sessions, as well as one or more outputs related to the programs. The storage unit 415 can store user data, e.g., user preferences and user programs. In some cases, the computer system 401 can include one or more additional data storage units, which are external to the computer system 401, such as on a remote server that communicates with the computer system 401 via an intranet or the Internet. Data can be transferred from one location to another using, for example, a communication network or a physical data transporter (e.g., using a hard disk drive, a thumb drive, or other data storage mechanisms).

[0312] The computer system 401 can communicate with one or more remote computer systems via the network 430. For example, the computer system 401 can communicate with a remote computer system of a user (e.g., an operator). Examples of remote computer systems include personal computers (e.g., portable PCs), slate or tablet PCs (e.g., iPad, Galaxy Tab), telephones, smart phones (e.g., iPhone, Android-supported devices, ) or personal digital assistants. The user can access the computer system 401 via the network 430.

[0313] The methods described herein can be implemented in the form of machine (e.g., computer processor) executable code that is stored at an electronic storage location of the computer system 401, such as, for example, on the memory 410 or the electronic storage unit 415. The machine executable code or machine readable code can be provided in the form of software. During use, the code can be executed by the processor 405. In some cases, the code can be retrieved from the storage unit 415 and stored on the memory 410 for ready access by the processor 405. In some cases, the electronic storage unit 415 may not be included and the machine executable instructions are stored on the memory 410.

[0314] In one aspect, the present disclosure provides a non-transitory computer-readable medium comprising computer-executable instructions that, when executed by at least one electronic processor, perform a method comprising: (a) obtaining more than one sequencing read from a set of tagged polynucleotides from a first sample and a second sample generated by a nucleic acid sequencer, wherein the sequencing reads comprise a tag sequence and a sequence derived from the polynucleotide; (b) aligning the more than one sequencing read to a reference sequence, thereby determining a start region and an end region of the alignment; (c) for each sample, grouping the more than one sequencing read into more than one family based on grouping features, the grouping features comprising at least one of the following (i), (ii), (iii), and (iv): (i) a tag, (ii) a start region, (iii) an end region, and (iv) the length of the polynucleotide, wherein each family in a sample comprises sequencing reads of tagged progeny polynucleotides amplified from a unique polynucleotide in the set of polynucleotides in the sample; (d) generating family identifiers for the more than one family; (e) screening out a set of common family identifiers, wherein a common family identifier is a family identifier of the first sample that is the same as or substantially the same as the family identifier of the second sample; (f) determining a quantitative measure of the set of common family identifiers; and (g) classifying the first sample as contaminated by the second sample if the quantitative measure of the common family identifiers is higher than a predetermined threshold, or classifying the first sample as not contaminated if the quantitative measure of the common family identifiers is at or below the predetermined threshold.

[0315] The code can be pre-compiled and configured to be used with a machine having a processor suitable for executing the code or can be compiled during run-time. The code can be provided in a programming language that can be selected such that the code can be executed in a pre-compiled or as-compiled manner.

[0316] Aspects of the systems and methods provided herein, such as computer system 401, can be implemented in programming. Various aspects of the technology can be considered to be in the form of a “product” or “articles of manufacture” typically carried on or implemented in a type of machine-readable medium or a type of machine-readable medium as machine (or processor)-executable code and / or associated data. The machine-executable code can be stored in an electronic storage unit such as a memory (e.g., read-only memory, random access memory, flash memory) or a hard disk. A “storage” type medium can include any one or all of tangible memories such as computers, processors, etc. or associated modules thereof, such as various semiconductor memories, tape drives, disk drives, etc., which can provide non-transitory storage for software programming at any time.

[0317] All or part of the software can sometimes communicate via the Internet or various other telecommunications networks. For example, such communication can enable the loading of software from one computer or processor to another, such as from a management server or host to the computer platform of an application server. Thus, another type of medium that can carry software elements includes, for example, light waves, radio waves, and electromagnetic waves used across physical interfaces between local devices, over wired and fiber optic landline networks, and over various air-links. Physical elements that carry such waves, such as wired or wireless links, optical links, etc., can also be considered media that carry software. As used herein, unless restricted to non-transitory, tangible "storage" media, terms such as computer or machine "readable media" refer to any medium that participates in providing instructions to a processor for execution.

[0318] Thus, machine-readable media, such as computer-executable code, can take many forms, including but not limited to tangible storage media, carrier media, or physical transmission media. Non-volatile storage media includes, for example, optical or magnetic disks, such as any storage device in any computer, etc., as shown in the figures, such as can be used to implement a database, etc. Volatile storage media includes dynamic memory, such as the main memory of such a computer platform. Tangible transmission media includes coaxial cables; copper wire and fiber optics, including the wires that make up the bus within a computer system. Carrier transmission media can take the form of electrical or electromagnetic signals, or acoustic or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communication. Thus, common forms of computer-readable media include, for example: floppy disk, flexible disk, hard disk, magnetic tape, any other magnetic medium, CD-ROM, DVD or DVD-ROM, any other optical medium, punched cards, paper tape, any other physical storage medium with a hole pattern, RAM, ROM, PROM, and EPROM, FLASH-EPROM, any other memory chip or cartridge, a carrier wave that transports data or instructions, a cable or link that transports such a carrier wave, or any other medium from which a computer can read programming code and / or data. Many of these forms of computer-readable media can participate in transporting one or more strings of one or more instructions to a processor for execution.

[0319] The computer system 401 can include or communicate with an electronic display that includes a user interface (UI) to provide, for example, one or more results of a sample analysis. Examples of UIs include, but are not limited to, graphical user interfaces (GUIs) and web-based user interfaces.

[0320] Additional details regarding computer systems and networks, databases, and computer program products are provided in the following literature: e.g., Peterson, Computer Networks: A Systems Approach, Morgan Kaufmann, 5th Edition (2011); Kurose, Computer Networking: A Top-Down Approach, Pearson, 7th Edition (2016); Elmasri, Fundamentals of Database Systems, Addison Wesley, 6th Edition (2010); Coronel, Database Systems: Design, Implementation, & Management, Cengage Learning, 11th Edition (2014); Tucker, Programming Languages, McGraw-Hill Science / Engineering / Math, 2nd Edition (2006); and Rhoton, Cloud Computing Architected: Solution Design Handbook, Recursive Press (2011), each of which is hereby incorporated by reference in its entirety into this document.

[0321] Applications

[0322] Cancer and other diseases

[0323] Typically, the disease under consideration is a type of cancer. Non-limiting examples of such cancers include biliary tract cancer, bladder cancer, transitional cell carcinoma, urothelial carcinoma, brain cancer, glioma, astrocytoma, breast cancer, metaplastic carcinoma, cervical cancer, cervical squamous cell carcinoma, rectal cancer, colorectal cancer, colon cancer, hereditary non-polyposis colorectal cancer, colorectal adenocarcinoma, gastrointestinal stromal tumor (GIST), endometrial cancer, endometrial stromal sarcoma, esophageal cancer, esophageal squamous cell carcinoma, esophageal adenocarcinoma, ocular melanoma, uveal melanoma, gallbladder cancer, gallbladder adenocarcinoma, renal cell carcinoma, clear cell renal cell carcinoma, transitional cell carcinoma, urothelial carcinoma, Wilms tumor, leukemia, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), chronic lymphocytic leukemia (CLL), chronic myeloid leukemia (CML), chronic myelomonocytic leukemia (CMML), liver cancer, liver carcinoma, hepatoma, hepatocellular carcinoma, cholangiocarcinoma, hepatoblastoma, lung cancer, non-small cell lung cancer (NSCLC), mesothelioma, B-cell lymphoma, non-Hodgkin lymphoma, diffuse large B-cell lymphoma, mantle cell lymphoma, T-cell lymphoma, non-Hodgkin lymphoma, precursor T-lymphoblastic lymphoma / leukemia, peripheral T-cell lymphoma, multiple myeloma, nasopharyngeal carcinoma (NPC), neuroblastoma, oropharyngeal cancer, oral squamous cell carcinoma, osteosarcoma, ovarian cancer, pancreatic cancer, pancreatic ductal adenocarcinoma, solid pseudopapillary tumor, acinar cell carcinoma, prostate cancer, prostatic adenocarcinoma, skin cancer, melanoma, malignant melanoma, cutaneous melanoma, small intestine cancer, stomach cancer, gastric carcinoma, gastrointestinal stromal tumor (GIST), uterine cancer or uterine sarcoma.

[0324] Non-limiting examples of other genetically-based diseases, disorders or conditions optionally evaluated using the methods and systems disclosed herein include achondroplasia, alpha-1 antitrypsin deficiency, antiphospholipid syndrome, autism, autosomal dominant polycystic kidney disease, Charcot-Marie-Tooth (CMT), cri du chat syndrome, Crohn's disease, cystic fibrosis, Dercum disease, Down syndrome, Duane syndrome, Duchenne muscular dystrophy, factor V Leiden thrombophilia, familial hypercholesterolemia, familial Mediterranean fever, fragile X syndrome, Gaucher disease, hemochromatosis, hemophilia, holoprosencephaly, Huntington's disease, Klinefelter syndrome, Marfan syndrome, myotonic dystrophy, neurofibromatosis, Noonan syndrome, osteogenesis imperfecta, Parkinson's disease, phenylketonuria, Poland anomaly, porphyria, progeria, retinitis pigmentosa, severe combined immunodeficiency (scid), sickle cell disease, spinal muscular atrophy, Tay-Sachs, thalassemia, trimethylaminuria, Turner syndrome, velocardiofacial syndrome, WAGR syndrome, Wilson's disease, and the like.

[0325] Although the description has been described with reference to its specific embodiments, these specific embodiments are illustrative only and not restrictive. The concepts shown in the examples can be applied to other embodiments and implementations.

[0326] As the liquid biopsy assay is changed (e.g., changes in sequencing depth and common SNP panels), the methods and systems of the present disclosure can be retrained as needed to obtain a set of applicable thresholds (e.g., one or more criteria / thresholds to detect the presence or absence of contamination in a sample). Examples

[0327] Example 1: Determining contamination of a sample according to an embodiment of the present disclosure

[0328] A set of patient samples was analyzed using a blood-based cfDNA assay at Guardant Health (Redwood City, CA, USA). To examine the quality of assay performance and determine if there was any contamination of the samples, the set of samples was analyzed according to embodiments of the present disclosure. The analysis of two samples (Sample 1 and Sample 2) from the set of samples is described in this example. The total number of families in Sample 1 and Sample 2 was 7,811,148 and 7,141,008, respectively. In this embodiment, families at over-represented genomic start and genomic end position pairs were not included in the analysis, and the set threshold for classifying a pair of genomic start and genomic end positions as an over-represented pair was 10 families. Thus, the total number of families in Sample 1 and Sample 2 was 6,452,057 and 6,039,099, respectively.

[0329] I: Determine whether sample 1 is contaminated by sample 2

[0330] Among the 6,452,057 families in Sample 1, 54,212 families were common families (shared with Sample 2). Among the 54,212 common families: (i) 9362 common families had the same number of sequencing reads in the families of both Sample 1 and Sample 2; and (ii) 1647 common families had a greater number of sequencing reads in the families of Sample 1 than in the corresponding families of Sample 2. In this embodiment, common families having a greater number of sequencing reads in the families of Sample 1 than in the corresponding families of Sample 2 were not included when determining the quantitative measure of common families. Additionally, in this embodiment, the quantitative measure of common families was the percentage of common families in Sample 1 out of the total families, which was equal to 0.815% (100*(54212 - 1647) / 6452057). In this embodiment, the predetermined threshold for classifying a sample as contaminated was 0.5%. Since the quantitative measure of common families in Sample 1 was greater than 0.5%, Sample 1 was determined to be contaminated by Sample 2.

[0331] II: Determine whether sample 2 is contaminated by sample 1

[0332] Among the 6,039,099 families of sample 2, 54,212 families are common families (shared with sample 1). Among the 54,212 common families: (i) 9362 common families have the same number of sequencing reads in the families of both sample 1 and sample 2; and (ii) 43,203 common families have a greater number of sequencing reads in the families of sample 2 than in the corresponding families of sample 1. Excluding the common families that have a greater number of sequencing reads in the families of sample 2 than in the corresponding families of sample 1, the quantitative measure of the common families of sample 2 is equal to 0.182% (100*(54212 - 43203) / 6039099). Since the quantitative measure of the common families of sample 2 is lower than the predetermined threshold (0.5%), sample 2 is determined to be not contaminated by sample 1.

[0333] Although the preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. It is not intended to limit the present invention to the specific examples provided in this specification. While the present invention has been described with reference to the foregoing specification, the description and illustration of the embodiments herein are not intended to be construed in a limiting sense. Many variations, changes and substitutions will now occur to those skilled in the art without departing from the present invention. In addition, it should be understood that all aspects of the present invention are not limited to the specific depictions, configurations or relative proportions set forth herein that depend on various conditions and variables. It should be understood that various alternatives of the embodiments of the present invention described herein may be employed in practicing the present invention. Accordingly, it is contemplated that the present invention should also cover any such alternatives, modifications, variations or equivalent forms. The appended claims are intended to define the scope of the present invention and thereby cover the methods and structures within the scope of these claims and their equivalents.

[0334] Although for purposes of clarity and understanding, the foregoing disclosure has been described in some detail by way of illustration and example, it will be apparent to those of ordinary skill in the art upon reading this disclosure that various changes in form and detail may be made without departing from the actual scope of the disclosure and that the disclosure may be practiced within the scope of the appended claims. For example, all method, system, computer-readable medium, and / or component features, steps, elements or other aspects may be used in various combinations.

[0335] All patents, patent applications, websites, other publications or documents, accession numbers, etc., cited herein are incorporated by reference in their entirety for all purposes to the same extent as if each individual item were specifically and individually indicated to be incorporated by reference. If different versions of a sequence are associated with an accession number at different times, the version associated with the accession number on the effective filing date of the present application is meant. If applicable, the effective filing date means the earlier of the actual filing date or the filing date of the priority application that mentions the accession number. Similarly, if different versions of a publication, website, etc. are published at different times, the version most recently published on the effective filing date of the present application is meant, unless otherwise indicated.

Claims

1. A method for detecting the presence or absence of contamination of a first sample by a second sample among more than one sample, the method comprising: (a) Process the first sample and the second sample, wherein the processing comprises: i. Tagging a collection of cell-free nucleic acid molecules in each sample with a collection of molecular barcodes to produce tagged polynucleotides, wherein the collection of molecular barcodes comprises 5 - 200 different molecular barcode sequences; ii. Amplifying a portion of the tagged polynucleotides to produce progeny polynucleotides; (b) For each of the first sample and the second sample, sequencing a portion of the progeny polynucleotides to produce sequencing reads; (c) Aligning more than one sequencing read from the first sample and the second sample to a reference sequence, thereby determining the genomic start position and the genomic end position of the cell-free nucleic acid molecules from the alignment; (d) For each of the first sample and the second sample, grouping the more than one sequencing read into more than one family based on grouping features, the grouping features including at least one of the following (i), (ii), and (iii): (i) one or more molecular barcodes attached to the cell-free nucleic acid molecules in the sample, (ii) the genomic start position of the cell-free nucleic acid molecules, and (iii) the genomic end position, wherein each family in the sample comprises sequencing reads of progeny polynucleotides amplified from a unique cell-free nucleic acid molecule from the collection of cell-free nucleic acid molecules in the sample; (e) Generating family identifiers for the more than one family; (f) Screening out a set of common family identifiers, wherein a given common family identifier is a family identifier of the first sample that is the same as or substantially the same as the family identifier of the second sample; (g) Determining a quantitative measure of the set of common family identifiers; and (h) If the quantitative measure of the set of common family identifiers is higher than a predetermined threshold, classifying the first sample as contaminated by the second sample, or if the quantitative measure of the set of common family identifiers is at or below the predetermined threshold, classifying the first sample as not contaminated by the second sample, thereby detecting the presence or absence of contamination.

2. The method according to claim 1, wherein the quantitative measure of the set of common family identifiers is the number of common family identifiers in the first sample.

3. The method according to claim 1, wherein the quantitative measure of the set of common family identifiers does not include the following common family identifiers in the first sample: common family identifiers in the family of the first sample whose number of sequencing reads is greater than the number of sequencing reads in the corresponding family of the second sample.

4. The method according to claim 1, wherein the quantitative measure of the set of common family identifiers in the first sample does not include common family identifiers at over-represented genomic start and genomic end position pairs.

5. The method according to claim 4, wherein the over-represented genomic start and genomic end position pairs are determined by: (a) providing a set of sequencing reads from the more than one sample, wherein the set of sequencing reads includes a distribution of genomic start and genomic end positions that are the same as or substantially the same as those of the first sample; (b) determining family identifiers in the set of sequencing reads; (c) quantifying the number of family identifiers in the set of sequencing reads that share a pair of genomic start and genomic end positions; and (d) classifying the genomic start and genomic end position pair as over-represented if the number of family identifiers exceeds a set threshold.

6. The method according to claim 1, wherein one or more sample indices are attached to one end or both ends of the progeny polynucleotide before the sequencing, wherein the one or more sample indices distinguish the first sample and the second sample.

7. The method according to claim 5, wherein the set threshold is at least 5, at least 10, at least 15, or at least 20.

8. The method according to claim 1, wherein the one or more molecular barcodes are attached to both ends of the cell-free nucleic acid molecule.

9. The method according to claim 6, wherein the first sample and the second sample are sequenced in the same flow cell.

10. The method according to claim 1, wherein the processing further comprises enriching a portion of the progeny polynucleotides for a specific region of interest to produce enriched molecules.

11. The method according to claim 6, wherein the more than one sample comprises samples processed in the same flow cell as the first sample.

12. The method according to claim 1, wherein the predetermined threshold is at least 0.5% or at least 1% of the total number of families in the first sample.

13. The method according to claim 1, wherein the first sample is obtained from a body fluid of one subject and the second sample is obtained from a body fluid of another subject.

14. The method according to claim 13, wherein the body fluid is plasma.

15. A computer-implemented method for detecting the presence or absence of contamination of a first sample by a second sample among more than one sample, comprising: (a) Obtaining sequence information of more than one sequencing read comprising a collection of cell-free nucleic acid molecules derived from the first sample and another collection of cell-free nucleic acid molecules from the second sample; (b) Aligning the more than one sequencing read to a reference sequence, thereby determining the genomic start position and the genomic end position of the cell-free nucleic acid molecules from the alignment; (c) For each of the first sample and the second sample, group the more than one sequencing reads into more than one family based on grouping features, the grouping features including at least one of the following (i), (ii), and (iii): (i) one or more molecular barcodes attached to cell-free nucleic acid molecules in the sample, (ii) genomic start positions of the cell-free nucleic acid molecules, and (iii) genomic end positions, wherein each family in the sample includes sequencing reads of daughter polynucleotides amplified from unique cell-free nucleic acid molecules in the set of cell-free nucleic acid molecules in the sample; (d) Generate family identifiers for the more than one family; (e) Screen out a set of common family identifiers, wherein a given common family identifier is a family identifier of the first sample that is the same as or substantially the same as the family identifier of the second sample; (f) Determine a quantitative measure of the set of common family identifiers; and (g) If the quantitative measure of the set of common family identifiers is higher than a predetermined threshold, classify the first sample as contaminated by the second sample, or if the quantitative measure of the set of common family identifiers is at or below the predetermined threshold, classify the first sample as not contaminated by the second sample, thereby detecting the presence or absence of contamination.

16. The method according to claim 15, wherein the quantitative measure of the set of common family identifiers is the number of common family identifiers in the first sample.

17. The method according to claim 15, wherein the quantitative measure of the set of common family identifiers does not include the following common family identifiers in the first sample: common family identifiers in the families of the first sample for which the number of sequencing reads is greater than the number of sequencing reads in the corresponding families of the second sample.

18. The method according to claim 15, wherein the quantitative measure of the set of common family identifiers in the first sample does not include common family identifiers at over-represented genomic start and genomic end position pairs.

19. The method according to claim 18, wherein the over-represented genomic start and genomic end positions are determined by: (a) providing a set of sequencing reads from the more than one sample, wherein the set of sequencing reads includes a distribution of genomic start and genomic end positions that are the same as or substantially the same as those of the first sample; (b) determining family identifiers in the set of sequencing reads; (c) quantifying the number of family identifiers in the set of sequencing reads that share a pair of genomic start and genomic end positions; and (d) classifying the pair of genomic start and genomic end positions as over-represented if the number of family identifiers exceeds a set threshold.

20. The method according to claim 19, wherein the more than one sample includes samples containing sample indices processed in the same flow cell as the first sample.

Citation Information

Patent Citations

  • Improvement in lanterns

    US170510A

  • Oligonucleotides

    US20010053519A1

  • Method and apparatus for imaging a sample on a device

    US20030152490A1

  • Digital Counting of Individual Molecules by Stochastic Attachment of Diverse Labels

    US20110160078A1

  • Coupled amplification and ligation method

    US5912148A