A method and device for detecting sample cross-contamination

By screening specific SNP sites and calculating sample contamination index, the problem of low-sensitivity sample cross-contamination detection in NGS is solved, and high-precision sample contamination identification in early tumor diagnosis is achieved, reducing the risk of misdiagnosis.

CN115985389BActive Publication Date: 2025-07-18GUANGZHOU BURNING ROCK DX CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211679939.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2025-07-18
Estimated Expiration
2042-12-26

Smart Images

  • Figure BDA0004018520620000041
    Figure BDA0004018520620000041
  • Figure BDA0004018520620000051
    Figure BDA0004018520620000051
  • Figure BDA0004018520620000061
    Figure BDA0004018520620000061
Patent Text Reader

Abstract

The present invention provides a method for screening single nucleotide polymorphism (SNP) sites for detecting sample contamination in methylation sequencing, as well as a method and device for detecting sample contamination. Among them, the screening method includes the following steps: S1: Select SNP sites with frequencies between 0.3 and 0.7 in a preset population; S2: Select SNP sites with a mutation direction from adenine (A) to thymine (T) or from thymine (T) to adenine (A); S3: Select SNP sites outside repetitive regions; S4: Select SNP sites with a physical distance greater than a preset length from each other. The present invention can achieve low-cost and high-precision detection of sample contamination in methylation sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biotechnology, and particularly to a method for detecting sample cross - contamination, which monitors the single - nucleotide polymorphism allele frequencies of DNA methylation next - generation sequencing samples to provide a judgment on whether the sample is cross - contaminated. Background Art

[0002] Bisulfite sequencing (BS - seq) is the gold standard for methylation sequencing. With its single - base resolution and high - throughput characteristics, its role in cancer screening, diagnosis, and monitoring is increasingly recognized. In high - throughput next - generation sequencing (NGS) detection, since multiple samples are processed in parallel, the risk of cross - contamination of heterologous DNA between adjacent samples during sample storage, preparation, etc. is difficult to eliminate. And this risk is even more serious in the early diagnosis and screening of tumors. Because in the blood samples of early - stage tumors, the tumor component usually accounts for a very low proportion (<0.001), trace contamination of blood samples can cause screening or diagnostic results to be incorrect. And the current NGS contamination detection methods often cannot achieve the detection sensitivity for a contamination ratio of <0.001. Moreover, in the current common NGS sample contamination determination, positive and negative reference samples are usually designed in each batch of samples. However, in real - world clinical practice, due to cost control and considerations, the setting of reference samples is ignored, which also increases the risk that cross - contamination of samples cannot be accurately identified. Summary of the Invention

[0003] To overcome the deficiencies in the prior art, the present invention provides a method and device for detecting sample cross - contamination, which can judge whether there is contamination from other samples in cell - free DNA samples of blood at low cost and high precision.

[0004] In one aspect, the present invention provides a method for screening single - nucleotide polymorphism (SNP) sites for detecting sample contamination in methylation sequencing, comprising the following steps:

[0005] S1: Select SNP sites with frequencies between 0.3 and 0.7 in a preset population;

[0006] S2: Select SNP sites with a mutation direction from adenine (A) to thymine (T) or from thymine (T) to adenine (A);

[0007] S3: Select SNP sites outside the repetitive regions;

[0008] S4: Select SNP sites with a physical distance greater than a preset length from each other;

[0009] Optionally, the order of S2 and S3 is interchanged.

[0010] On the other hand, the present invention provides a method for detecting sample contamination in methylation sequencing, including the following steps:

[0011] (1) Obtain the sequencing information obtained after performing methylation sequencing on the sample to be tested;

[0012] (2) Determine the sample contamination status according to the SNP sites screened by the above method for detecting sample contamination in methylation sequencing.

[0013] On the other hand, the present invention provides a device for detecting sample contamination in methylation sequencing, including:

[0014] A sequencing information acquisition module configured to obtain the sequencing information obtained after performing methylation sequencing on the sample to be tested;

[0015] A sample status determination module configured to determine the sample contamination status according to the SNP sites screened by the above method for detecting sample contamination in methylation sequencing. Description of the Drawings

[0016] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with this specification, and are used together with the specification to explain the principles of this specification.

[0017] Figure 1 It shows the sample contamination index (SCS) corresponding to different threshold conditions when the simulated contaminated samples take different ranges of SNP sites under different numbers of SNPs scenarios. For a single SCS threshold, the simulated contamination doping gradients are respectively: 0005pct (0.005%), 001pct (0.01%), 005pct (0.05%), 001pct (0.01%), 05pct (0.5%) and 1pct (1%). Among them, Figure 1 a is the case of 50 SNPs; Figure 1 b is the case of 100 SNPs; Figure 1 c is the case of 200 SNPs; Figure 1 d is the case of 300 SNPs; Figure 1 e is the case of 400 SNPs; Figure 1 f is the case of 500 SNPs; Figure 1 g is the case of 600 SNPs.

[0018] Figure 2Shows the sample contamination index of simulated contaminated samples with various mixing ratios under different numbers of SNPs when taking SNP sites in the same range (20% before the AR value fluctuates).

[0019] Figure 3 shows the AR frequency distribution diagram of contaminated samples under different cfDNA mixing ratios, where the abscissa is the AR value and the ordinate is the number of SNP sites. Among them, Figure 3a Is the background sample without mixing ratio; Figure 3b Is the contaminated sample with a mixing ratio of one in ten thousand; Figure 3c Is the contaminated sample with a mixing ratio of five in ten thousand; Figure 3d Is the contaminated sample with a mixing ratio of one in a thousand; Figure 3e Is the contaminated sample with a mixing ratio of five in a thousand.

[0020] Figure 4 Shows a reference example. A double-stranded DNA fragment expected to be subjected to methylation detection is shown above, sorted in the direction of the arrow. It contains the original upper strand (CCGGCATGTTTAAACGCT) and the original lower strand (AGCGTTTAAACATGCCGG). Among them, it is assumed that all cytosines (C) in CpG have been methylated and are marked with -mC. After the above double-stranded DNA fragment is denatured and unwound into a single-stranded form, it is subjected to bisulfite conversion treatment. The C that is not methylated (-mC) in the original upper strand and the original lower strand is converted into uracil (U), while the methylated C remains C. In the subsequent PCR amplification process, since uracil (U) pairs with adenine (A), and the base paired with adenine (A) introduced in the PCR amplification of DNA is thymine (T). In PCR amplification, first, a target upper strand complementary strand (CTOT) complementary to the original upper strand with uracil (U) after bisulfite treatment is formed, and a target lower strand complementary strand (CTOB) complementary to the original lower strand with uracil (U) after bisulfite treatment is formed. In the subsequent PCR amplification process, a target upper strand (OT) complementary to CTOT converted from the original upper strand and a target lower strand (OB) complementary to CTOB converted from the original lower strand are formed. By comparison, it can be seen that the C that is not methylated in the original upper strand and the original lower strand is replaced by T in the target upper strand and the target lower strand, while the methylated C (marked with an underline) remains unchanged. According to this feature, the number and position of methylated C can be identified by measuring the C after bisulfite conversion treatment, so as to achieve the purpose of DNA methylation detection. Specific Embodiments

[0021] I. Definitions

[0022] In the present invention, unless otherwise specified, scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Moreover, the terms related to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, immunology and the laboratory operation procedures used herein are all widely used terms and conventional procedures in the relevant fields. Meanwhile, for a better understanding of the present invention, the definitions and explanations of relevant terms are provided below.

[0023] The term "SNP" (Single Nucleotide Polymorphism) mainly refers to the DNA (Deoxyribo Nucleic Acid) sequence polymorphism caused by the variation of a single nucleotide at the genomic level. The polymorphism exhibited by SNP sites only involves the variation of a single base, which can be caused by the transition or transversion of a single base, or by the insertion or deletion of a base.

[0024] The term "homozygous SNP site" refers to a type of SNP site where, at this site, the base at this site on all sequences aligned with the reference genome shows the same base, and this base is different from the base at this site on the reference genome sequence. For example, if the base at a certain SNP site on the reference genome sequence is G, and the base at this SNP site on all sequences aligned with the reference genome is A, then this SNP site is called a homozygous SNP site.

[0025] The term "allele" refers to genes located at the same position on a pair of homologous chromosomes that control different forms of the same trait. When an organism carries a pair of identical alleles, then the organism is homozygous for this gene; conversely, if a pair of alleles are different, then the organism is heterozygous for this gene. Alleles each encode a protein product, determine a certain trait, and can lose their function due to mutation.

[0026] The term "allelic ratio" (AR) refers to the ratio of the mutant allele to the wild-type allele, and in NGS, it is equivalent to the ratio of the number of mutant sequences to the number of wild-type sequences.

[0027] The term "wild type" refers to the common or non-mutant form of a gene or organism in nature. That is, it refers to the phenotype with the highest frequency observed in the wild population, or the system, organism or gene with this phenotype.

[0028] The term "mutation" refers to a process in which the structure of a gene changes, resulting in a stable and heritable change in the genotype of a cell, virus, or microorganism.

[0029] The term "mutant" refers to a change in the structure of a gene caused by an addition, deletion, or change of a base pair in a DNA molecule.

[0030] The term "sample contamination score" (SCS) refers to a value used to characterize the cross-contamination of a sample to be tested by other samples.

[0031] II. Detailed description of specific embodiments

[0032] In one aspect, the present invention provides a method for screening single nucleotide polymorphism (SNP) sites for detecting sample contamination in methylation sequencing, comprising the following steps:

[0033] S1: Select SNP sites with frequencies between 0.3 and 0.7 in a preset population;

[0034] S2: Select SNP sites with a mutation direction from adenine (A) to thymine (T) or from thymine (T) to adenine (A);

[0035] S3: Select SNP sites outside repetitive regions;

[0036] S4: Select SNP sites with a physical distance greater than a preset length from each other.

[0037] In some optional embodiments, the order of S2 and S3 is interchanged.

[0038] In some embodiments, the SNP sites in S1 are selected from a preset database, and the preset database is selected from one or more of the gnomAD database, the 1000Genome Project database, the HapMap database, and the dbSNP database.

[0039] In some preferred embodiments, the preset database is the gnomAD database;

[0040] Wherein, the reference genome is the GRCh37 / hg19 human reference genome.

[0041] In some embodiments, the preset population is selected from East Asian populations, African / Afro-American populations, Latin American populations, non-Finnish European populations, Finnish European populations, Ashkenazi Jewish populations, or West Asian populations.

[0042] In some preferred embodiments, the preset population is East Asian populations.

[0043] In some embodiments, the preset length described in S4 is selected from 0.4 to 1 Mb.

[0044] In some preferred embodiments, the preset length is 1 Mb.

[0045] On the other hand, the present invention provides a method for detecting sample contamination in methylation sequencing, comprising the following steps:

[0046] (1) Obtaining sequencing information obtained by performing methylation sequencing on a sample to be tested;

[0047] (2) Determining the sample contamination status according to the SNP sites screened by the above method for detecting sample contamination in methylation sequencing.

[0048] In some embodiments, step (2) includes:

[0049] Determining the homozygous SNP sites in the sample to be tested corresponding to the SNP sites for detecting sample contamination in methylation sequencing.

[0050] In some preferred embodiments, the determination method is to calculate the mutant allele ratio (AR) of the SNP sites for detecting sample contamination in methylation sequencing.

[0051] In some more preferred embodiments, the AR value of the homozygous SNP sites is less than 0.25 or greater than 0.75.

[0052] In some embodiments, step (2) includes:

[0053] Calculating the sample contamination index of the homozygous SNP sites, wherein the sample contamination index is the median of the standard AR values of the homozygous SNP sites;

[0054] Wherein, the standard AR value is obtained by normalizing the AR value of the homozygous SNP sites, including:

[0055] When the AR value is less than or equal to 0.5, the standard AR value of this SNP site is equal to the AR value;

[0056] When the AR value is greater than 0.5, the standard AR value of this SNP site is the difference between 1 and the AR value.

[0057] In some embodiments, the standard AR values of the homozygous SNP sites are sorted, and the median is calculated by selecting the standard AR values of some homozygous SNP sites.

[0058] In some preferred embodiments, the sorting is in descending order.

[0059] In some preferred embodiments, the standard AR values of the homozygous SNP sites ranked in the top 5%, 10%, 15%, 20% or 25% after sorting from large to small are selected to calculate the sample contamination index.

[0060] In some more preferred embodiments, the standard AR values of the homozygous SNP sites ranked in the top 20% after sorting from large to small are selected to calculate the sample contamination index.

[0061] In some embodiments, when the sample contamination index is greater than a preset threshold, it is determined that the sample to be tested is cross-contaminated by other samples.

[0062] In some preferred embodiments, the preset threshold is selected from 0.001 to 0.01.

[0063] In some more preferred embodiments, the preset threshold is 0.001.

[0064] In some embodiments, determining whether the sample to be tested is cross-contaminated by other samples further includes the following steps:

[0065] Before comparing the sample contamination index with the preset threshold, first compare the sample contamination index with the background noise after methylated sequencing of the sample to be tested. When the sample contamination index is greater than the background noise, then compare the sample contamination index with the preset threshold.

[0066] On the other hand, the present invention provides a device for detecting sample contamination in methylated sequencing, including:

[0067] A sequencing information acquisition module configured to acquire the sequencing information obtained by performing methylated sequencing on the sample to be tested;

[0068] A sample status determination module configured to determine the sample contamination status according to the SNP sites screened by the above method for detecting sample contamination in methylated sequencing.

[0069] For the purpose of clear and concise description, features are described herein as part of the same or separate embodiments. However, it will be understood that the scope of the present invention may include some embodiments having combinations of all or some of the described features.

[0070] Examples

[0071] Data preparation: SNP screening

[0072] The first step in evaluating whether a sample is cross - contaminated is to screen suitable SNP sites from the mutation database. First, according to the results of the hg19 gnomAD database, select single - nucleotide polymorphisms (SNPs) with frequencies between 0.3 and 0.7 in the eastern population. Too low or too high population frequencies are not suitable as evaluation sites. Since the application scenario of the present invention is methylation sequencing data, after DNA is treated with bisulfite, unmethylated C on the original upper strand and original lower strand will be converted to T, and the corresponding complementary strand G is converted to A. Based on this, the mutation direction of the SNPs screened in this step needs to satisfy A->T or T->A. If the SNP is located in a repetitive region, this site is more likely to be mis - measured during sequencing, so SNPs located in repetitive regions are also filtered out. Finally, the selected SNPs are physically more than 0.4M apart from each other.

[0073] Algorithm for building based on monitored SNP sites

[0074] 1. Calculation of allelic ratio (AR) of monitored SNP sites

[0075] For samples actually used for contamination monitoring, calculate the AR values of all SNPs screened in Example 1. The calculation method of AR is the number of reads containing the mutant SNP divided by the number of all reads covering this SNP. Theoretically, the AR of a homozygous wild - type SNP site is 0, the AR of a homozygous mutant SNP site is 1, and the AR of a heterozygous SNP site is 0.5. In real experiments, due to a series of errors such as sequencing errors, in fact, the AR of homozygous SNPs will show slight fluctuations near 0 or 1, and the AR of heterozygous SNPs will fluctuate near 0.5. According to the sequencing results of the sample, use the software samtools mpileup module to detect SNPs and perform AR statistics on all SNPs screened in Example 1. When a sample is contaminated by less than 50%, the AR values of SNPs fluctuate in the range of <0.25 or >0.75. If the AR of a certain SNP is between 0.25 and 0.75, then this SNP is considered a heterozygous SNP in this sample and is not included in the subsequent calculation. At the same time, the SNP sites used for subsequent calculation satisfy a sequencing depth greater than 50X. Retain all sites considered to be homozygous SNPs for the next step of calculation.

[0076] 2. Evaluation of the possibility of overall contamination of the sample according to the allelic frequency of SNP sites

[0077] Standardize the AR of the homozygous SNPs retained in Step 1 above, with the a value being the standardized AR value:

[0078] 1) If the SNP is homozygous wild type, i.e., AR ≤ 0.5, then a = AR;

[0079] 2) If the SNP is homozygous mutant, i.e., AR > 0.5, then a = 1 - AR. Sort all the a values from largest to smallest, and select the top 20% of the a values to calculate the median, which is set as the sample contamination score (SCS, sample contamination score). That is, select the 20% of SNPs with the largest AR value fluctuations to calculate the SCS. The formula is as follows:

[0080] SCS = median(r) -

[0081] r is the list of a values of the top 20% of a values among homozygous SNPs.

[0082] Under ideal circumstances, if a sample is not cross - contaminated by other samples, its SCS is 0. However, due to the existence of background noise such as sequencing errors, the SCS will show slight fluctuations around 0. If the SCS value of a sample is greater than the threshold of background noise, then it is judged that this sample is very likely to be cross - contaminated by other samples; otherwise, the probability that this sample is contaminated by other samples is considered very low. Considering sequencing errors, experimental errors, etc., the threshold of background noise is set to 0.001.

[0083] Example 1: Verification with simulated data

[0084] To verify the feasibility of this methodology, the SNP sites of real sample A and real sample B were artificially modified to generate multiple differentially homozygous sites between samples A and B. For example, if the genotype of SNP_1 site in sample A is ref / ref and the genotype of the same site in sample B is modified to alt / alt, then SNP_1 site is considered a differentially homozygous site between sample A and B. After modification, to simulate a real application scenario, 0.04% of sequencing background noise was retained. Multiple gradients of read number ratios were constructed, and the read fragments of sample A were incorporated into sample B. The gradients were 5 / 100000, 1 / 10000, 5 / 10000, 1 / 1000, 5 / 1000, and 1 / 100, corresponding to Figure 1 abscissas 0005pct, 001pct, 005pct, 001pct, 05pct, and 1pct respectively. Each gradient ratio was simulated 5 times, and each time reads were randomly selected. At the same time, seven gradients of the number of differentially homozygous sites were constructed, corresponding to modifying 50, 100, 200, 300, 400, 500, and 600 SNPs respectively. These SNPs were all randomly selected from 1000 SNP sites that met the screening conditions of the present invention (see Table 3).

[0085] Calculate the a value of the obtained SNPs according to the steps in Example 2.

[0086] To determine the best range of SNP sites required for calculating SCS, for seven scenarios with 50, 100, 200, 300, 400, 500, and 600 SNPs constructed respectively, sort the calculated a values from largest to smallest, and select the a values of the SNP sites ranked in the top 5%, 10%, 15%, 20%, and 25% respectively, corresponding Figure 1 to the abscissas 0.05, 0.1, 0.15, 0.2, and 0.25. The dotted line in the figure corresponds to the background baseline, that is, the SCS value of the background sample when the doping ratio is 0%. From Figure 1 it can be seen that in the application scenarios with different numbers of SNPs, when taking the a values of the top 20% of the SNP sites, the SCS values of the simulated doping ratio results can all be clearly separated from the background. Therefore, the best range of SNP sites for calculating SCS should be the top 20%, that is, select the sites with the top 20% of AR fluctuations for SCS calculation.

[0087] When taking the sites with the top 20% of AR fluctuations, the box plots of the SCS values calculated according to different doping ratio gradients and different numbers of differential SNPs used are as Figure 2 shown. It can be inferred from the doping ratio results that the fewer the number of SNPs available for calculation, the greater the SCS fluctuation and the more unstable the result. When the application scenario is 600 SNPs, the SCS values of different doping ratio samples can all be clearly separated from the baseline, and when the doping ratio gradient is higher, the SCS value increases more significantly. Therefore, in this example, 600 SNPs are selected for the next simulation.

[0088] When taking the sites with the top 20% of AR fluctuations, the SCS values calculated at different doping ratio gradients in the application scenario with 600 SNPs are shown in Table 1 below. The SCS value of the background sample (without doping) is 0. The designed 5 replicate groups were constructed by randomly extracting reads fragments of sample A and doping them into sample B. From the simulation results, it can be concluded that when determining the threshold with the purpose of being able to clearly separate from the background sample, when the threshold cutoff of the SCS value is set at 0.001, samples contaminated with a doping ratio of one in ten thousand (0.01%) can be detected. Therefore, when the SCS value of the sample to be tested is greater than 0.001, it can be determined that there is contamination with a doping ratio of one in ten thousand or more in the sample to be tested.

[0089] Table 1: SCS values in the application scenario of 600 SNPs

[0090]

[0091] Example 2: Experimental verification with real samples

[0092] During the experiment, two cfDNA blood samples from different donors were intermingled at different gradients. The gradient mixing ratios were one in ten thousand, five in ten thousand, one in a thousand, and five in a thousand. Among them, sample one was the sample to be incorporated, and sample two was the sample to be incorporated into. Sample one was incorporated into sample two at different mixing ratio gradients (0, 0.01%, 0.05%, 0.1%, 0.5%). In this example, the number of available SNP sites detected was 1000. From the results in Table 2, it can be seen that when the gradient mixing ratio was five in ten thousand (0.05%), a relatively significant change in the SCS value was observed compared to the background sample (i.e., the mixing ratio gradient was 0). Based on the SCS value threshold cutoff = 0.001 determined from the simulation data, it was verified from the experimental results that this methodology could distinguish a five-in-ten-thousand sample cross-contamination under the condition of detecting 1000 SNPs (see Table 3).

[0093] According to the different mixing ratios of the experimental cfDNA, background (no mixing) samples, samples with mixing ratios of one in ten thousand, five in ten thousand, one in a thousand, and five in a thousand were selected to make AR frequency distribution plots, as shown in Figure 3. For the background sample, the AR values of most sites were concentrated at both ends. As the mixing ratio increased, the distribution of AR showed a trend of spreading towards the middle. This was in line with the evaluation of the AR value model for the test sample after being cross-contaminated by other samples during the algorithm construction process in the aforementioned Example 2.

[0094] Table 2: Detection of different experimental cfDNA mixing ratios

[0095] Sample Name Mixing Ratio Gradient SCS Sample 1-1 0 0.0000 Sample 1-2 0 0.0000 Sample 2-1 0 0.0000 Sample 2-2 0 0.0005 Sample 2-p001-1 0.01% 0.0000 Sample 2-p001-2 0.01% 0.0000 Sample 2-p005-1 0.05% 0.0088 Sample 2-p005-2 0.05% 0.0117 Sample 2-p01-1 0.1% 0.0113 Sample 2-p01-2 0.1% 0.0107 Sample 2-p05-1 0.5% 0.0379 Sample 2-p05-2 0.5% 0.0361

[0096] Table 3: List of 1000 SNP sites

[0097] Among them: CHROM = chromosome number; POS = position; REF = reference sequence base (i.e., the wild-type base when the SNP site has not mutated);

[0098] ALT = variant sequence base (i.e., the variant base after the SNP site has mutated)

[0099]

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107]

Claims

1. A method for screening single nucleotide polymorphism (SNP) sites for detecting sample contamination in methylation sequencing, comprising the following steps: S1: Select SNP sites with frequencies between 0.3 and 0.7 in a preset population; S2: Select SNP sites with a mutation direction from adenine (A) to thymine (T) or from thymine (T) to adenine (A); S3: Select SNP sites outside repetitive regions; S4: Select SNP sites with a physical distance greater than 0.4 Mb from each other; Optionally, the order of S2 and S3 is interchanged.

2. The method according to claim 1, wherein, The SNP sites in S1 are selected from a preset database, and the preset database is selected from one or more of the gnomAD database, the 1000Genome Project database, the HapMap database, and the dbSNP database.

3. The method according to claim 2, wherein The preset database is the gnomAD database.

4. The method according to claim 1, wherein In S1, the GRCh37 / hg19 or GRCh38 / hg38 human reference genome is used as the reference genome.

5. The method according to any one of claims 1-4, wherein, The preset population is selected from East Asian population, African / Afro-American population, Latin American population, non-Finnish European population, Finnish European population, Ashkenazi Jewish population, or West Asian population.

6. The method according to claim 5, wherein The preset population is the East Asian population.

7. The method according to any one of claims 1-4 or 6, wherein In S4, select SNP sites with a physical distance greater than 1 Mb from each other.

8. The method according to claim 5, wherein In S4, select SNP sites with a physical distance greater than 1 Mb from each other.

9. A method for detecting sample contamination in methylation sequencing, comprising the following steps: (1) Obtain sequencing information obtained by performing methylation sequencing on a sample to be tested; (2) Determine the sample contamination status according to the SNP sites for detecting sample contamination in methylation sequencing screened by the method according to any one of claims 1-8; Step (2) includes: Determine the homozygous SNP sites in the sample to be tested corresponding to the SNP sites for detecting sample contamination in methylation sequencing; Calculate the sample contamination index of the homozygous SNP sites, wherein the sample contamination index is the median of the standard AR values of the homozygous SNP sites; When the sample contamination index is greater than a preset threshold, it is determined that the sample to be tested is cross-contaminated by other samples, and the preset threshold is selected from 0.001 to 0.01; Wherein, the standard AR value is obtained by normalizing the AR value of the homozygous SNP site, including: When the AR value is less than or equal to 0.5, the standard AR value of this SNP site is equal to the AR value; When the AR value is greater than 0.5, the standard AR value of this SNP site is the difference between 1 and the AR value; The method for determining the homozygous SNP sites includes calculating the allele ratio (AR) of the mutated alleles of the SNP sites for detecting sample contamination in methylation sequencing, and the AR value of the homozygous SNP sites is less than 0.25 or greater than 0.75; Step (2) further includes sorting the standard AR values of the homozygous SNP sites from large to small, and selecting the standard AR values of the homozygous SNP sites located in the top 20% after sorting from large to small to calculate the sample contamination index.

10. The method according to claim 9, wherein the standard AR value of the homozygous SNP sites ranked in the top 15% in descending order is selected to calculate the sample contamination index.

11. The method according to claim 9, wherein the standard AR value of the homozygous SNP sites ranked in the top 10% in descending order is selected to calculate the sample contamination index.

12. The method according to claim 9, wherein the standard AR value of the homozygous SNP sites ranked in the top 5% in descending order is selected to calculate the sample contamination index.

13. The method according to any one of claims 9-12, wherein, The preset threshold is 0.

001.

14. The method according to any one of claims 9-12, further comprising the following steps for determining whether the sample to be tested is cross-contaminated by other samples: Before comparing the sample contamination index with the preset threshold, first compare the sample contamination index with the background noise after methylated sequencing of the sample to be tested. When the sample contamination index is greater than the background noise, then compare the sample contamination index with the preset threshold; wherein the background noise is 0.

001.

15. The method according to claim 13, further comprising the following steps for determining whether the sample to be tested is cross-contaminated by other samples: Before comparing the sample contamination index with the preset threshold, first compare the sample contamination index with the background noise after methylated sequencing of the sample to be tested. When the sample contamination index is greater than the background noise, then compare the sample contamination index with the preset threshold; wherein the background noise is 0.

001.

16. A device for detecting sample contamination in methylated sequencing, comprising: a sequencing information acquisition module configured to acquire the sequencing information obtained after methylated sequencing of a sample to be tested; a sample status determination module configured to determine the sample contamination status according to the method described in any one of claims 9-15 for the SNP sites selected for detecting sample contamination in methylated sequencing according to the method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Method for screening SNP (Single Nucleotide Polymorphism) sites and application thereof

    CN114517223A

  • Screening method of SNP (Single Nucleotide Polymorphism) sites for detecting pollution level of sample and detection method of pollution level of sample

    CN114530198A