Systems and methods for contamination detection in next-generation sequencing samples

CN114730609BActive Publication Date: 2026-08-11F HOFFMANN LA ROCHE & CO AG
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-20
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,在现有技术中描述的用于检测样品间污染的方法很少

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114730609B_ABST
    Figure CN114730609B_ABST
Patent Text Reader

Abstract

This paper describes a statistical method based on β-mixture modeling for detecting contamination and reporting contamination levels as both point estimates and confidence intervals for liquid biopsy samples. We validate our method using both computer simulations and in vitro contamination-spiked samples. Although we focus on liquid biopsy samples, the same strategy, with minor modifications, is applicable to any general NGS application. For example, tissue samples from biopsies can be used according to the system and method described herein.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] none.

[0003] Incorporate by reference

[0004] All publications and patent applications mentioned in this specification are incorporated herein by reference to the extent that each individual publication or patent application is specifically and individually indicated to be incorporated herein by reference. Technical Field

[0005] Embodiments of the present invention generally relate to systems and methods for next-generation sequencing, and more specifically to systems and methods for detecting contamination in next-generation sequencing samples. Background Technology

[0006] Next-generation sequencing (NGS) technology has evolved to be used for diagnostic analysis in a variety of applications, such as vector screening, infectious diseases, and cancer detection and / or analysis. In some applications, sequencing targets (which may be mutations (i.e., variants) in cancer, for example) may be present in very low amounts in the sample. When targets are present in such low amounts in the sample, the risk of false positives increases.

[0007] One potential source of false positives is sequencing errors. For example, a sequencer with an initial read accuracy of 90% is expected to produce many errors for a sequence identified in a single read, making it difficult to distinguish between errors and actual mutations. One way to reduce such sequencing errors is to determine the common sequence by sequencing the target multiple times, thereby achieving the desired common sequence accuracy (i.e., 99%, 99.9%, or 99.99%).

[0008] Another source of false positives could be sample contamination (i.e., cross-contamination between samples). However, few methods for detecting cross-contamination between samples are described in the prior art. Therefore, it is desirable for systems and methods to detect contamination in NGS samples to reduce the risk of false positives. Summary of the Invention

[0009] This invention relates generally to systems and methods for next-generation sequencing, and more specifically to systems and methods for detecting contamination in next-generation sequencing samples.

[0010] In some embodiments, a method for detecting contamination is provided. The method may include receiving an electronic file comprising a list of variants of a sequenced sample from a subject; calculating a set of alternative allele frequencies for a set of variants within a frequency range; determining whether the sample is contaminated based on analysis of the set of alternative allele frequencies; if the sample is not contaminated, administering a drug at least partially based on the variant list; and if the sample is contaminated, obtaining an uncontaminated sequenced sample comprising a second variant list, and administering a drug at least partially based on the second variant list.

[0011] In some embodiments, the frequency range is between 0 and 0.25. In some embodiments, the frequency range is between 0 and 0.1.

[0012] In some embodiments, the analysis of this set of alternative allele frequencies includes fitting the alternative allele frequencies to a clustering model. In some embodiments, the clustering model is a mixture model. In some embodiments, the mixture model is a β-mixture model.

[0013] In some embodiments, the method further includes determining whether any of the alternative allele frequencies is an outlier and removing any outliers from the set of alternative allele frequencies prior to analysis of the alternative allele frequencies.

[0014] In some embodiments, the step of determining whether any of the alternative allele frequencies is an outlier includes calculating a local outlier factor.

[0015] In some embodiments, the method further includes determining the level of contamination by analyzing the frequencies of alternative alleles.

[0016] In some embodiments, the step of determining the pollution level includes fitting alternative allele frequencies to a hybrid model.

[0017] In some embodiments, the method further includes determining a confidence level surrounding the pollution level.

[0018] In some embodiments, determining the confidence level includes bootstrap variants and corresponding alternative allele frequencies.

[0019] In some embodiments, the drug is a cancer drug.

[0020] In some embodiments, the cancer drug performs better in patients with a specific variant than in patients without that specific variant.

[0021] In some embodiments, the sample is sequenced to an average sequencing depth of at least 1000x. In some embodiments, the sample is sequenced to an average sequencing depth of at least 2000x.

[0022] In some embodiments, a system for detecting contamination is provided. The system includes a processor programmed to perform the steps listed in any of the methods described herein.

[0023] In some embodiments, a method for detecting contamination is provided. The method may include receiving an electronic file comprising a list of variants of a sequenced sample from a subject; calculating a set of alternative allele frequencies for a set of variants within a frequency range; and determining, based on analysis of the set of alternative allele frequencies, whether the sample is contaminated.

[0024] In some embodiments, a computer product is provided. The computer product includes a computer-readable medium storing a plurality of instructions for controlling a computer system to perform any of the methods listed above.

[0025] In some embodiments, a system is provided that includes the computer product described above; and one or more processors for executing instructions stored on a computer-readable medium. Attached Figure Description

[0026] The novel features of the present invention are specifically set forth in the appended claims. The features and advantages of the invention can be better understood by referring to the following detailed description, taken in conjunction with the accompanying drawings, which illustrate illustrative embodiments utilizing the principles of the invention, as shown in the drawings:

[0027] Figure 1 is a block diagram illustrating an embodiment of a computer system configured to implement one or more aspects of the present invention.

[0028] Figures 2A and 2B show that the frequencies of alternative alleles in pure, uncontaminated samples tend to concentrate around three levels: 0.0 (AA), 0.5 (Aa), and 1.0 (aa).

[0029] Figures 3A to 3C show the frequency distribution of alternative allele frequencies in the low range of pure samples.

[0030] Figures 4A to 4D show the frequency distribution of alternative allele frequencies in the low range of the contaminated samples.

[0031] Figures 5A to 5C show the simulation results for the targeted kit group; Figures 5D to 5F show the simulation results for the expanded kit group; and Figures 5G to 5I show the simulation results for the monitoring kit group.

[0032] Figure 6 shows a histogram illustrating the distribution of alternative allele frequencies in the low range of the contaminated sample.

[0033] Figure 7A shows the average of the predicted pollution levels.

[0034] Figure 7B shows the confidence level using 1000 bootstrapping.

[0035] Figure 8A shows the performance of 1000 synthesized samples with a single source of contamination and sequenced to typical coverage depth.

[0036] Figures 8B to 8F show the effect of increasing sequencing coverage of 10,000 samples with low contamination levels (less than 1% contamination).

[0037] Figure 9A shows the performance of the method using five contamination sources sequenced to typical coverage depths.

[0038] Figures 9B to 9F show the effects of varying sequencing depth on 10,000 synthetic samples with five contamination sources, resulting in a low contamination rate of less than 1%.

[0039] Figure 10 shows the predicted pollution level as comparable to the specified pollution level, with a Pearson correlation coefficient of 0.94. Detailed Implementation

[0040] Identifying cross-contamination is crucial in all next-generation sequencing (NGS) applications, especially those designed to detect somatic mutations or any other variants present at low frequencies in samples, such as liquid biopsies used for various applications, including cancer detection and / or analysis. In such applications, contamination can lead to false-positive results for critical mutations and harm patients, for example, by resulting in unnecessary treatment or prescriptions for ineffective or suboptimal medications. However, very few methods for detecting cross-contamination are available in the public domain. This paper describes a statistical method based on β-mixture modeling for detecting contamination and reporting contamination levels as both point estimates and confidence intervals for liquid biopsy samples. We... computer Simulation and in vitro Both contaminated spiked samples validated our method. Although we focused on liquid biopsy samples, the same strategy, with minor modifications, is applicable to any general NGS application. For example, tissue samples from biopsies can be used according to the system and method described herein.

[0041] I. Next-generation sequencing technology

[0042] As indicated above, sequencing analysis is performed as part of a procedure to determine sequencing reads from multiple microsatellite loci, sequencing the prepared target nucleic acid molecules (e.g., sequencing libraries). A variety of sequencing technologies or sequencing analyses can be used. As used herein, the term "next-generation sequencing (NGS)" refers to a sequencing method that allows for the massive parallel sequencing of cloned amplified molecules and individual nucleic acid molecules.

[0043] Non-limiting examples of sequencing methods suitable for use with the methods disclosed herein include nanopore sequencing (US Patent Publications 2013 / 0244340, 2013 / 0264207, 2014 / 0134616, 2015 / 0119259, and 2015 / 0337366), Sanger sequencing, capillary array sequencing, thermal cycling sequencing (Sears et al., Biotechniques 13:626-633 (1992)), solid-phase sequencing (Zimmerman et al., Methods in Molecular Cell Biology 3:39-42 (1992)), mass spectrometry sequencing using methods such as matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF / MS; Fu et al., Nature Biotech 16:381-384 (1998)), and sequencing by hybridization (Drmanac). (E. et al., Nature Biotechnology 16:54-58 (1998)) and NGS methods, including but not limited to sequencing by synthesis (e.g., HiSeq™, MiSeq™ or Genome Analyzer, all available from Illumina), sequencing by ligation (e.g., SOLiD™, Life Technologies), ion semiconductor sequencing (e.g., Ion Torrent™, Life Technologies) and SMRT® sequencing (e.g., PacificBiosciences).

[0044] Commercially available sequencing technologies include hybridization-while-sequencing (HMS) from Affymetrix Ltd. (Sunnyvale, California), synthesis-while-sequencing (SMS) from Illumina / Solexa (San Diego, California) and Helicos Biosciences (Cambridge, Massachusetts), and ligation-while-sequencing (LSMS) from Applied Biosystems (Foster City, California). Other sequencing technologies include, but are not limited to, Ion Torrent technology (ThermoFisherScientific) and nanopore sequencing (Genia Technology of Roche Sequencing Solutions, Santa Clara, California); and Oxford Nanopore Technologies (Oxford, UK).

[0045] II. Exemplary computer system for implementing the algorithm

[0046] The algorithms described herein can be implemented on a computer system. For example, Figure 1 is a block diagram illustrating one embodiment of a computer system 100 configured to implement one or more aspects of the present invention. As shown, the computer system 100 includes, but is not limited to, a central processing unit (CPU) 102 and a system memory 104, which is coupled to a parallel processing subsystem 112 via a memory bridge 105 and a communication path 113. The memory bridge 105 is further coupled to an I / O (input / output) bridge 107 via a communication path 106, and the I / O bridge 107 is in turn coupled to a switch 116.

[0047] In operation, I / O bridge 107 is configured to receive user input from input device 108 (e.g., keyboard, mouse, video / image capture device, etc.) and forward that input to CPU 102 for processing via communication path 106 and memory bridge 105. In some embodiments, the input is a real-time feed from a camera / image capture device or video data stored on a digital storage medium on which object detection operations are performed. Switch 116 is configured to provide connectivity between I / O bridge 107 and other components of computer system 100, such as network adapter 118 and various expansion cards 120 and 121.

[0048] As shown in the figure, I / O bridge 107 is coupled to system disk 114, which can be configured to store content and applications, as well as data for use by CPU 102 and parallel processing subsystem 112. Generally, system disk 114 provides non-volatile storage for applications and data and may include fixed or removable hard disk drives, flash memory devices, and CD-ROMs (optical disc read-only memory), DVD-ROMs (digital versatile optical discs), Blu-ray discs, HD-DVDs (high-definition DVDs), or other magnetic, optical, or solid-state storage devices. Finally, although not explicitly shown, other components such as universal serial buses or other port connectors, optical disc drives, digital versatile optical disc drives, video recording devices, etc., may also be connected to I / O bridge 107.

[0049] In various embodiments, memory bridge 105 may be a northbridge chip, and I / O bridge 107 may be a southbridge chip. Furthermore, communication paths 106 and 113, as well as other communication paths within computer system 100, may be implemented using any technically suitable protocol, including but not limited to AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol known in the art.

[0050] In some embodiments, the parallel processing subsystem 112 includes a graphics subsystem that transmits pixels to a display device 110, which may be any conventional cathode ray tube, liquid crystal display, light-emitting diode display, etc. In such embodiments, the parallel processing subsystem 112 incorporates circuitry optimized for graphics and video processing, including, for example, video output circuitry. Such circuitry may be incorporated across one or more parallel processing units (PPUs) included in the parallel processing subsystem 112. In other embodiments, the parallel processing subsystem 112 incorporates circuitry optimized for general and / or computational processing. Again, such circuitry may be incorporated across one or more PPUs included in the parallel processing subsystem 112, which are configured to perform such general and / or computational operations. In still other embodiments, the one or more PPUs included in the parallel processing subsystem 112 may be configured to perform graphics processing, general processing, and computational processing operations. System memory 104 includes at least one device driver 103 configured to manage the processing operations of one or more PPUs within parallel processing subsystem 112. System memory 104 also includes a software application 125 executing on CPU 102 and capable of issuing commands to control the operation of the PPUs.

[0051] In various embodiments, the parallel processing subsystem 112 may be integrated with one or more other elements of FIG1 to form a single system. For example, the parallel processing subsystem 112 may be integrated with a CPU 102 and other interconnect circuitry on a single chip to form a system-on-a-chip (SoC).

[0052] It should be understood that the system shown herein is illustrative, and variations and modifications are possible. The connection topology can be modified as needed, including the number and arrangement of bridges, the number of CPUs 102, and the number of parallel processing subsystems 112. For example, in some embodiments, system memory 104 may be connected directly to CPU 102 instead of via memory bridge 105, and other devices will communicate with system memory 104 via memory bridge 105 and CPU 102. In other alternative topologies, parallel processing subsystems 112 may be connected to I / O bridge 107 or directly to CPU 102 instead of via memory bridge 105. In still other embodiments, I / O bridge 107 and memory bridge 105 may be integrated into a single chip instead of existing as one or more discrete devices. Finally, in some embodiments, one or more components shown in FIG. 1 may be absent. For example, switch 116 may be removed, and network adapter 118 and expansion cards 120, 121 will be directly connected to I / O bridge 107.

[0053] III. Pollution Detection

[0054] In some embodiments, sample contamination can be detected by: (1) identifying a set of common single nucleotide variants (SNVs) for variant-based analysis; (2) calculating alternative allele frequencies for the identified set of common SNVs; (3) removing anomalous sites with local anomalous factors (LOFs) (e.g., sequencing errors and somatic mutations); (4) fitting a clustering model (i.e., a β-mixture model) on the alternative allele frequencies to model background samples and foreground contamination; and (5) inferring point estimates of contamination levels from the fitted clustering model and obtaining confidence intervals from nonparametric bootstrapping.

[0055] Although the results presented in this paper use a beta mixture model, other clustering models can be used in a similar manner. For example, connectivity models, centroid models, distribution models (i.e., mixture models), subspace models, group models, graph-based models, signed graph models, and neural models can also be used.

[0056] Because humans are diploid, the frequencies of alternative alleles in pure, uncontaminated samples tend to concentrate around three levels: 0.0 (AA), 0.5 (Aa), and 1.0 (aa), as shown in Figures 2A and 2B. When a sample from a target subject is contaminated by a sample from a non-target subject (i.e., a contaminated subject), for the AA locus with an expected level of 0.0, the Aa and aa SNVs from the contaminated subject will shift the frequency of the alternative allele in the target subject above 0. In short, these deviations from 0 allow us to detect contamination.

[0057] In some embodiments, the low range of alternative allele frequencies (i.e., less than about 25%, 20%, 15%, 10%, or 5%) can be analyzed to identify pure, uncontaminated samples from contaminated samples, as shown in Figures 3A-3C and 4A-4D. Figures 3A-3C show the frequency distribution in the low range of alternative allele frequencies for pure samples. As shown in Figures 3A-3C, very few SNVs in the low frequency range deviate from the expected 0.0 value in pure samples. Conversely, as shown in Figures 4A-4D for contaminated samples, significantly more SNVs in the low frequency range deviate from the expected 0.0 value.

[0058] In other embodiments, a high range of substitution allele frequencies (i.e., greater than about 75%, 80%, 85%, 90%, or 95%) can be analyzed. In other embodiments, an intermediate range of substitution allele frequencies (i.e., between about 25% and 75%, 30% and 70%, 35% and 65%, 40% and 60%, or 45% and 55%) can be analyzed. In other embodiments, any combination of low, intermediate, and high range substitution allele frequencies can be analyzed to identify pure, uncontaminated samples from contaminated samples. In other words, in some embodiments, the full frequency range from 0 to 100%, or any portion or combination of portions of the full frequency range, can be used.

[0059] To construct the model, we used 1000 Genome Project (TGP) data and identified and selected common SNVs with a population substitution allele frequency exceeding 0.5% but below 99.5%. Therefore, we selected 285, 868, and 707 common SNVs within the Avenio® ctDNA Assay Kit Liquid Biopsy group for targeting, expansion, and surveillance. 625 relevant subjects were selected from the TGP for this model, and the selection reflected the US population: White (68.2%), Hispanic or Latino (15.4%), Black (11.9%), and Asian (4.5%). In other embodiments, the population selected for the model may represent populations within a country, state, county, province, region, continent, etc. For example, factors considered in selection may include race, ethnicity, sex, age, and / or location. In some embodiments, at least 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 SNVs are selected for the model. In some embodiments, at most 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 SNVs are selected for the model. In some embodiments, a larger group providing more sequencing data covering more genes results in more SNVs that can be selected for the model. In some embodiments, the group covers at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 genes. In some embodiments, the groups cover up to 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 genes. In some embodiments, larger groups allow for more stringent selection criteria for SNVs. For example, common SNVs with population substitution allele frequencies between 5% and 95%, 10% and 90%, 15% and 85%, 20% and 80%, 25% and 75%, 30% and 70%, 35% and 65%, 40% and 60%, or 45% and 55% can be selected for the model and generate a sufficient number of SNVs. In some embodiments, the population substitution frequency range can be determined based on the group size.

[0060] In some embodiments, the method uses a low range (≤ 25%) of alternative allele frequencies detected in liquid biopsy samples to model informative loci. Based on 10,000 simulations for each group, with randomly selected target subjects and randomized contamination subjects, we show that the expected average informative SNVs in samples with a single contamination source treated using each of the three groups are 21 (target kit), 70 (expansion kit), and 54 (monitor kit). Figures 5A to 5C show the simulation results for the target kit group; Figure 5D Figures 5F to 5I show simulation results for the expanded kit group; and Figures 5G to 5I show simulation results for the monitoring kit group. In some embodiments, at least 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 informative SNVs are selected from the total number of SNVs selected for analysis of samples with a single source of contamination.

[0061] We then determined the alternative allele frequencies of informative SNVs and modeled them using a three-component β-mixture model for contamination detection. The probability density function of the general β-mixture model has the following form:

[0062] f ( x │ π 1, π 2, α 0, β 0, α 1, β 1, α 2, β 2) = (1 − π 1 − π 2)β( x │ α 0, β 0) + π 1 β( x │ α 1, β 1) + π 2β( x │ α 2, β 2)

[0063] For pure samples, the histogram of the low substitution allele frequency range has only a single peak at 0.0 (not shown), while assuming a single source of contamination, contaminated samples will have two to three peaks, including the peak at 0.0 (AA), as shown in Figure 6. Assuming that the SNV is independent and has only one source of contamination, the model will have a total of three β components: (1) background AA + contamination (AA); (2) background AA + contamination Aa; and (3) background AA + contamination aa.

[0064] We assume that each SNV is independent and that each SNV is at the same depth ( The average sequencing depth of the selected site was calculated, and only one contamination source was present. Therefore, the following simplification can be used to reduce the number of parameters in the model: (1) As shown in Figure 6, the first β component at 0.0 substitution allele frequency (i.e., background AA of contaminating AA) is used α 0=1 β 0=10 4 Direct parameterization. We expect to have sharp spikes ( β 0=10 4 The strictly decreasing distribution of ) α 0=1), as an alternative form of the Dirac delta distribution, but also allowing for sequencing errors. (2) The second β component corresponds to background Aa and contamination Aa. The expected number of alternative alleles for each SNV belonging to this component is parameterized as α 1, and the total number of alleles at each locus is set to the average coverage depth of all selected SNPs. Each SNV was sequenced at the same depth. Equivalently, α1 + β1 = α2 + β2 = N, where N is calculated as the average coverage depth of all selected SNVs for each sample. (3) The third β component corresponds to background AA and contamination aa. Since aa has two alternative alleles, the expected number of alternative alleles for each SNP belonging to this component is The total number of remaining alleles at each locus is Maximum likelihood estimation was used to fit a β-mixture model on the alternative allele frequencies. In other words, the center of the homozygous contamination distribution (aa) is twice the center of the heterozygous contamination distribution (Aa), or due to the two alternative alleles. α 2=2 α 1.

[0065] Through the above simplification, the number of free parameters is reduced from eight to three, and the general β mixture model is simplified to the following equation:

[0066] f ( x │ π 1, π 2, α 1) = (1− π 1− π 2)β( x │1, 10 4 ) + π 1 β( x │ α 1, N - α 1)+ π 2β( x │2 α 1, N -2 α 1) f( x | π 1, π 2, α 1) = (1− π 1− π 2)β( x │1, 10 4 ) + π 1 β( x | α 1, N - α 1) + π 2β( x │2 α 1, N -2 α 1)

[0067] The likelihood function ℒ can be used. π 1,2, α 1│ x )= The quasi-Newton method (i.e., the Broyden–Fletcher–Goldfarb–Shannon algorithm with boundary constraints) is used to estimate or determine the parameters of the β mixture model with maximum likelihood. π 1. π 2. α 1. Multiple initializations can be used to avoid local maxima.

[0068] Because maximum likelihood estimation methods are sensitive to outliers (e.g., sequencing errors and somatic mutations), we use a local outlier factor (LOF) to remove outliers before model fitting, resulting in a more robust estimate of contamination levels. LOF measures the local density (lrd) of a point compared to its k nearest neighbors.

[0069]

[0070] LOF > 1 means that the local density of point A is less than the average local density of its neighboring points, indicating that A may be an outlier. We use k=5 and the cutoff value of LOF = 3.

[0071] In some embodiments, other methods of outlier detection are possible. For example, in some embodiments, outlier detection methods may be univariate or multivariate. In some embodiments, outlier detection methods may be parametric or nonparametric. In some embodiments, outlier detection methods may be z-score or extreme value analysis (parametric), probabilistic and statistical modeling (parametric), linear regression models, proximity-based models, information-theoretic models, high-dimensional outlier detection methods, neural networks, Bayesian networks, hidden Markov models, fuzzy logic-based methods, and / or ensemble techniques.

[0072] Then, the point estimate of the pollution level is estimated as 2α 1 / N (i.e., the average of the third β component, homozygous contamination distribution aa). Figure 7A shows that, for this example, the specified contamination level (0.5%) is consistent with the predicted contamination level (0.54%). Confidence intervals for the contamination level were constructed by bootstrapping SNV sites (i.e., nonparametric bootstrapping) and their corresponding alternative allele frequencies. Figure 7B shows that for 1000 bootstrappings, the 90% confidence interval for the contamination level in this example is (0.49%, 0.63%). In the presence of sequencing errors, the theoretical detection limit of our method is at least... (when Or, if there is at most one sequencing error per site. This limit increases as the number of sequencing errors per site increases (i.e., the sequencing error rate increases). To obtain more conservative and fewer false positive results, we only report predicted contamination levels greater than [a certain value]. In some embodiments, confidence intervals are generated using at least 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200, 1300, 1400, 1500, 1600, 1700, 1800, 1900, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000 bootstrapping.

[0073] For 30,000 from TGP computer Comprehensive simulation studies of the synthesized samples showed that the performance of this method improved with increasing sequencing depth, and for depth-sequencing samples from the extended Avenio kit, the false positive rate and detection limit were as low as 0.02% and 0.09%, respectively (Table 1). Target individuals and contaminant individuals (one contaminant source) were again derived from samples selected from a population of 1000 genome projects. We synthesized multiple samples for each SNV using a binomial distribution. (Number) replacement alleles.

[0074]

[0075] in It is the average sequencing depth across all sites, and This is based on the theoretical substitution allele frequency at specific sites and pollution levels. For or We adjusted it to or This is used to illustrate sequencing errors. The frequency of alternative alleles at each locus is then calculated. More generally, to illustrate sequencing errors, p can be adjusted by a quantity corresponding to the expected magnitude of sequencing errors in the system. In some embodiments, this could be common sequencing errors.

[0076] In some embodiments, the sequencing depth coverage is less than 25x, 50x, 75x, 100x, 200x, 300x, 400x, 500x, 600x, 700x, 800x, 900x, 1000x, 2000x, 3000x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, or 10000x. In some embodiments, the coverage is at least 25x, 50x, 75x, 100x, 200x, 300x, 400x, 500x, 600x, 700x, 800x, 900x, 1000x, 2000x, 3000x, 4000x, 5000x, 6000x, 7000x, 8000x, 9000x, or 10000x.

[0077] In some embodiments, the synthetic data includes samples with various contamination levels between 0% and 50%. Above 50% contamination, background contamination becomes the foreground, so it is not necessary to simulate contamination levels above 50%. In some embodiments, the simulated contamination levels are less than 50%, 40%, 30%, 20%, 10%, 5%, 4%, 3%, 2%, or 1%. Figure 8A shows the performance of 1000 synthetic samples with a single contamination source, sequenced to typical coverage depth.

[0078] When the contamination level is low (i.e., less than 1%), performance begins to decline at sequencing coverage of less than approximately 2000x, as shown in Figures 8B to 8F. Figures 8B to 8F illustrate the effect of increasing sequencing coverage of 10,000 samples with low contamination levels (less than 1% contamination).

[0079] In the case of multiple contamination sources, the method was still able to label contaminated samples with lower than expected contamination levels, despite violating one of the model's assumptions (i.e., a single contamination source). To evaluate the performance of the method with multiple contamination sources, data from 1000 synthetic samples with five contamination sources were generated, where the contamination level from each source was randomized, and the total contamination level from the five sources ranged from 0% to 50%. Figure 9A shows the performance of the method using five contamination sources sequenced to typical coverage depth. Figures 9B through 9F show the effect of varying sequencing depth on 10,000 synthetic samples with five contamination sources, resulting in a total low contamination level of less than 1%. Similarly, performance began to decline when sequencing coverage was less than 2000x.

[0080] Our proposed method further uses 103 in vitro Spiked plasma samples (one source of contamination, contamination levels ranging from 0% to 8%) were evaluated using the extended Avenio kit from the CLIA laboratory in San Jose. The predicted contamination levels were comparable to the specified contamination levels, with a Pearson correlation coefficient of 0.94, as shown in Figure 10.

[0081] Table 1. 30,000 units of the expanded Avenio kit computer The performance of the recommended method for sample contamination is presented in the synthetic data. FPR – False Positive Rate; LOD – Limit of Detection; CI – Clopper-Pearson exact confidence interval. LOD is determined using probability unit regression with a 99.9% detection rate.

[0082] Average coverage depth FPR 1 source LOD (95% CI) 5 source LODs (95% CI) < 1000x 0.135% 0.97% (0.89% – 1.09%) 2.12% (1.97% – 2.35%) 1000x – 2000x 0.056% 0.36% (0.33% – 0.40%) 1.27% (1.18% – 1.40%) 2000x – 3000x 0.035% 0.12% (0.11% – 0.13%) 0.80% (0.74% – 0.89%) 3000x – 4000x 0.025% 0.09% (0.09% – 0.10%) 0.57% (0.52% – 0.64%) > 4000x 0.020% 0.09% (0.09% – 0.09%) 0.41% (0.36% – 0.48%)

[0083] The statistical methods described herein can be integrated into any pipeline that generates alternative allele frequencies for a large number of loci within a sequencing group. This facilitates the identification and proper handling of contaminated samples, thereby significantly improving the reliability of analytical results.

[0084] V. Treatment Methods

[0085] In some embodiments, the systems and methods described herein can be used to guide patient treatment based on the correct identification of variants rather than on variants arising from sample contamination. For example, cancer therapies, such as the administration of cancer drugs, can be selected based on the identified variants. In some embodiments, certain treatments may be excluded because variants have been identified as originating from or potentially originating from sample contamination. In some embodiments, when sample contamination is detected, the patient is retested (i.e., the sample is resequencing), and an appropriate therapy is selected and administered only after retesting and / or confirmation that the sample is not contaminated.

[0086] When a feature or element is referred to herein as being “on” another feature or element, it may be directly on the other feature or element, or there may be intermediate features and / or elements present. Conversely, when a feature or element is referred to as being “directly on” another feature or element, there are no intermediate features or elements present. It will also be understood that when a feature or element is referred to as being “connected,” “attached,” or “coupled” to another feature or element, it may be directly connected, attached, or coupled to the other feature or element, or there may be intermediate features or elements present. Conversely, when a feature or element is referred to as being “directly connected,” “directly attached,” or “directly coupled” to another feature or element, there are no intermediate features or elements present. Although one embodiment has been described or illustrated, the features and elements thus described or illustrated can be applied to other embodiments. Those skilled in the art will also recognize that a structure or feature referred to as being “adjacent” to another feature may have portions that overlap with or are located below the adjacent feature.

[0087] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. For example, as used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “including” are used in this specification, they specify the presence of the defined features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items and may be abbreviated to “ / ”.

[0088] For ease of description, spatially relative terms such as “below,” “below,” “lower than,” “above,” “over,” etc., are used herein to describe the relationship of one element or feature to another, as illustrated in the accompanying drawings. It should be understood that, in addition to the orientations depicted in the drawings, spatially relative terms are also intended to cover different orientations of the device in use or operation. For example, if the device in the drawings is inverted, an element described as “below” or “under” other elements or features would then be oriented “above” other elements or features. Thus, the exemplary term “below” can encompass both the orientations of “above” and “below.” The device may be oriented in other ways (rotated 90 degrees or in other orientations), and the spatially relative descriptive terms used herein shall be interpreted accordingly. Similarly, unless otherwise specifically indicated, terms such as “up,” “down,” “vertical,” “horizontal,” etc., are used herein for illustrative purposes only.

[0089] Although the terms “first” and “second” may be used herein to describe various features / elements (including steps), these features / elements should not be limited by these terms unless the context otherwise indicates. These terms can be used to distinguish one feature / element from another. Therefore, without departing from the teachings of the invention, the first feature / element discussed below may be referred to as the second feature / element, and similarly, the second feature / element discussed below may be referred to as the first feature / element.

[0090] Throughout this specification and the following claims, unless the context otherwise requires, the word “comprising” and variations such as “including” and “containing” mean that various components may be used together in methods and articles (e.g., compositions and devices that include means and methods). For example, the term “comprising” will be understood to imply the inclusion of any of the specified elements or steps, but does not exclude any other elements or steps.

[0091] As used herein in the specification and claims, including as in the examples, and unless otherwise expressly specified, all figures may be interpreted as if preceded by the words “about” or “approximately,” even if the term is not explicitly stated. When describing magnitude and / or location, the phrase “about” or “approximately” may be used to indicate that the described value and / or location is within a reasonably expected range of value and / or location. For example, numerical values ​​may have values ​​of + / - 0.1% of a specified value (or range of values), + / - 1% of a specified value (or range of values), + / - 2% of a specified value (or range of values), + / - 5% of a specified value (or range of values), + / - 10% of a specified value (or range of values), etc. Unless the context otherwise indicates, any numerical value given herein should also be understood to include about or approximately that value. For example, if the value “10” is disclosed, “about 10” is also disclosed. Any numerical ranges described herein are intended to include all subranges contained therein. It should also be understood that, as those skilled in the art would appropriately understand, when a value is disclosed, the terms "less than or equal to" that value, "greater than or equal to" that value, and possible ranges between values ​​are also disclosed. For example, if the value "X" is disclosed, "less than or equal to X" and "greater than or equal to X" are also disclosed (e.g., in the case where X is a numerical value). It should also be understood that throughout this application, data is provided in a variety of different formats, and the data represents a range of endpoints and start points, as well as any combination of data points. For example, if specific data point "10" and specific data point "15" are disclosed, it should be understood that values ​​greater than, greater than or equal to, less than, less than or equal to, equal to 10 and 15, and values ​​between 10 and 15 are considered disclosed. It should also be understood that each unit between two specific units is also disclosed. For example, if 10 and 15 are disclosed, 11, 12, 13, and 14 are also disclosed.

[0092] Although various illustrative embodiments have been described above, any of a variety of changes may be made to the various embodiments without departing from the scope of the invention as set forth in the claims. For example, in alternative embodiments, the order in which the various method steps described are performed may often be changed, while in other alternative embodiments, one or more method steps may be skipped entirely. In some embodiments, optional features of various apparatus and system embodiments may be included, while in others they may not be included. Therefore, the foregoing description is provided primarily for illustrative purposes and should not be construed as limiting the scope of the invention as set forth in the claims.

[0093] The examples and illustrations included herein are shown in an illustrative and not limiting manner, specific embodiments in which the subject matter can be practiced. As mentioned, other embodiments can be utilized and derived from them, allowing for structural and logical substitutions and changes without departing from the scope of this disclosure. These embodiments of the subject matter of the invention may be referred to herein individually or collectively by the term "invention," merely for convenience and not intended to actively limit the scope of this application to any single inventive concept, given that more than one inventive concept has actually been disclosed. Therefore, although specific embodiments have been shown and described herein, any arrangement calculated to achieve the same purpose may replace the specific embodiments shown. This disclosure is intended to cover any and all modifications or variations of the various embodiments. After reviewing the foregoing description, combinations of the above embodiments, as well as other embodiments not explicitly described herein, will be apparent to those skilled in the art.

Claims

1. A method for detecting contamination, the method comprising: Receive an electronic file containing a list of variants of the sequenced samples from the subjects; Calculate the frequencies of a set of alternative alleles for a set of variants within a frequency range; Determining whether any of the alternative allele frequencies is an outlier, and removing any outliers from the set of alternative allele frequencies before analyzing the alternative allele frequencies, wherein the step of determining whether any of the alternative allele frequencies is an outlier includes calculating a local outlier factor, wherein the local outlier factor measures the local density of a point compared to its k nearest neighbors, where k is the number of neighbors, wherein the analysis of the alternative allele frequencies includes fitting the alternative allele frequencies to a β-mixture model; The substitution allele frequencies are fitted to a β-mixture model, and point estimates of the contamination level are inferred from the fitted β-mixture model to determine whether the sample is contaminated.

2. The method of claim 1, wherein the frequency range is between 0 and 0.

25.

3. The method according to claim 1, wherein the frequency range is between 0 and 0.

1.

4. The method of claim 1, further comprising determining a confidence level surrounding the level of said contamination.

5. The method of claim 4, wherein determining the confidence level includes bootstrapping the variant and the corresponding alternative allele frequency.

6. The method of claim 1, wherein the sample is sequenced to an average sequencing depth of at least 1000x.

7. The method of claim 1, wherein the sample is sequenced to an average sequencing depth of at least 2000x.

8. A computer product comprising a computer-readable medium storing a plurality of instructions for controlling a computer system to perform the method of any one of claims 1-7.

9. A system comprising: The computer product according to claim 8; and One or more processors, the one or more processors being configured to execute instructions stored on the computer-readable medium.

Citation Information

Patent Citations

  • Nanopore Based Molecular Detection and Sequencing

    US20130244340A1

  • DNA sequencing by synthesis using modified nucleotides and nanopore detection

    US20130264207A1

  • Nucleic acid sequencing using tags

    US20140134616A1

  • Nucleic acid sequencing by nanopore detection of tag molecules

    US20150119259A1

  • Methods for creating bilayers for use with nanopore sensors

    US20150337366A1