A screening method for detecting a SNP site of a sample contamination level and a method for detecting a sample contamination level

By screening and selecting SNP sites with specific frequencies and balance as marker sites, the problems of false positives and sensitivity in tumor sample contamination detection are solved, achieving lower cost and higher accuracy in predicting contamination levels.

CN114530198BActive Publication Date: 2026-01-13BERRY ONCOLOGY CO LTD +1
View PDF 2 Cites -1 Cited by

Patent Information

Application Number
CN202011321699.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-23
Publication Date
2026-01-13
Estimated Expiration
2040-11-23

AI Technical Summary

Technical Problem

Existing methods for detecting contamination in tumor samples suffer from numerous false positive mutations and low detection sensitivity. Furthermore, the high cost of sequencing limits the depth of sample sequencing, and there is a lack of effective panel-based methods for predicting contamination at marker sites.

Method used

SNPs with a mutation frequency of 30%–70% in the population were selected as candidate marker sites. The selected regions were divided into regions with a length of 0.7–1.3 Mb. Sites with an allele frequency of 40%–60% and that best conformed to Hardy-Wenge equilibrium were selected as marker sites. Other candidate marker sites were removed. These marker sites were used to detect the level of contamination in the samples.

Benefits of technology

It reduced testing costs, improved testing accuracy, and enabled more efficient prediction of sample contamination levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114530198B_ABST
    Figure CN114530198B_ABST
Patent Text Reader

Abstract

The application discloses a screening method of SNP sites for detecting sample pollution level and a detection method of sample pollution level, and relates to the technical field of biological sequencing. The screening method comprises the following steps: obtaining SNP sites with a human population mutation frequency of 30%-70% in a target region as candidate marker sites; dividing a region between a starting site and a terminal site of the candidate marker sites existing on a single chromosome into multiple selection regions, and the length of the selection region is 0.7-1.3 Mb; if two or more candidate marker sites exist in the selection region, selecting a site with an allele frequency of 40%-60% and different genotypes most conforming to Hardy-Weinberg equilibrium in the selection region as a marker site, and removing other candidate markers in the region. The marker sites screened based on the above screening method can be used for detecting the pollution level of a sample, and compared with the prior art, the detection cost is lower, and the detection accuracy is higher.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biological sequencing technology, in particular, to a screening method for detecting SNP sites of sample contamination level and a method for detecting sample contamination level. BACKGROUND

[0002] The second-generation sequencing (NGS) technology has been rapidly developed and applied in the genomic sequencing of tumor clinical samples due to its high-throughput and low-cost characteristics. However, the parallel library sequencing of multiple samples leads to the problem of data contamination between samples. The tumor clinical samples are usually designed in pairs of tumor and control, and the control samples are used to filter the germline mutations of the corresponding samples in the tumor samples. Generally, the contamination of the tumor samples will cause a large number of heterologous contaminated germline mutations to be incorrectly determined as somatic mutations, thereby resulting in a large number of false positive mutations in the detected somatic mutations. At the same time, the heterologous contamination will reduce the purity of the tumor samples, thereby reducing the detection sensitivity of the somatic mutations in the tumor samples. Therefore, accurately detecting the contamination level of the tumor clinical samples is an indispensable quality control step.

[0003] Whole genome sequencing (WGS) and whole exome sequencing (WES) can cover a certain amount of human population polymorphism single nucleotide polymorphism (SNP) sites. Based on the priori population frequency and the wild type and mutant information of the SNP sites in the tumor sample and the control sample in the paired samples, the sample contamination level estimation has been successfully applied in whole genome sequencing and whole exome sequencing. The existing representative methods include ConEst, Conpair and VerifyBamID, which all use Bayesian methods, use a large number of marker sites covered by WGS / WES. Compared with VerifyBamID using all covered SNP sites, ConEst and Conpair both use homozygous sites. Due to the universality of gene copy number variation (CNV) in tumor samples, the heterozygous sites of tumor samples are affected by CNV and deviate from the mutation abundance (Variation Alle Fraction, VAF) in paired samples, thereby causing the contamination level estimation value of VerifyBamID for the sample to be possibly too high.

[0004] Although WGS and WES help to more comprehensively detect and understand the overall picture of tumor gene mutations, the high sequencing cost of WGS and WES leads to limited sequencing depth of samples, and there is no report on an effective panel-based marker site contamination prediction method. As a necessary quality control module of tumor gene detection, the contamination level prediction is an important guarantee for the reliability of the tumor sample gene detection results.

[0005] Therefore, it is necessary and urgent to develop a set of marker site screening and contamination level prediction methods suitable for different panels. If all markers in the WES range are simply added to the large panel, for example, the Conpair algorithm needs to cover 7387 sites in the WES range. If a 120bp probe is designed for each marker, at least 0.886Mb of interval size needs to be covered, which will obviously greatly increase the panel size and cost. At the same time, when the number of marker sites is too small, the Conpair algorithm performs poorly. The contamination prediction algorithm of the present application can accurately predict the contamination level of the sample only by relying on the marker sites covered at the beginning of the design of the large panel.

[0006] In view of this, the present application is proposed. SUMMARY

[0007] The present application aims to provide a screening method of SNP sites for detecting sample contamination level and a method for detecting sample contamination level.

[0008] The present application is implemented as follows:

[0009] In a first aspect, the embodiments provide a screening method of SNP sites for detecting sample contamination level, obtaining SNP sites with a population mutation frequency of 30%-70% in a target region as candidate marker sites; dividing a region between a starting site and a terminal site of the candidate marker sites existing on a single chromosome into a plurality of selection regions, the length of the selection region being 0.7-1.3Mb; if there are two or more candidate marker sites in the selection region, selecting a site with an allele frequency of 40%-60% and different genotypes most consistent with Hardy-Weinberg equilibrium in the selection region as a marker site, and removing other candidate marker sites in the selection region.

[0010] In a second aspect, the embodiments provide a method for detecting sample contamination level, comprising: using the screening method of SNP sites for detecting sample contamination level as described in the preceding embodiments to screen the contamination marker sites.

[0011] In a third aspect, the embodiments provide a kit for detecting sample contamination level, comprising reagents for detecting target SNP sites, the target SNP sites being marker sites screened by the screening method of SNP sites for detecting sample contamination level as described in the preceding embodiments.

[0012] In a fourth aspect, the embodiments provide an electronic device comprising a memory and a processor, wherein the processor, when running a computer program in the memory, performs the screening method for detecting a SNP locus of a sample contamination level or the method for detecting a sample contamination level according to the foregoing embodiments.

[0013] The present application has the following beneficial effects:

[0014] The embodiments of the present application provide a screening method for detecting a SNP locus of a sample contamination level and a method for detecting a sample contamination level. The screening method comprises obtaining SNP loci with a human population mutation frequency of 30% to 70% in a target region as candidate marker loci; dividing a region between a start locus and an end locus of the candidate marker loci existing on a single chromosome into a plurality of selection regions, the length of the selection region being 0.7 to 1.3 Mb; if there are two or more candidate marker loci in the selection region, selecting a locus with an allele frequency of 40% to 60% and different genotypes most consistent with Hardy-Weinberg equilibrium in the selection region as a marker locus, and removing other candidate marker loci in the selection region. The marker locus screened based on the screening method can be used to detect the contamination level of a sample, and compared with the prior art, the detection cost is lower and the detection accuracy is higher. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced as follows. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor.

[0016] Figure 1 Flowchart for the method for detecting a sample contamination level in Example 3;

[0017] Figure 2 VAF difference distribution graph under different contamination levels in Test Example 1;

[0018] Figure 3 Number distribution graph of homozygous marker loci in Test Example 1;

[0019] Figure 4 Performance result graph of the detection performance of the method of Example 3 in Test Example 1;

[0020] Figure 5 Performance performance result graph of the detection methods of Example 4 and Example 5 in Test Example 2. DETAILED DESCRIPTION

[0021] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below. The specific conditions not noted in the embodiments are carried out according to the conventional conditions or the conditions suggested by the manufacturers. The reagents or instruments not noted by the manufacturers are all conventional products that can be obtained by market purchase.

[0022] The features and performances of the present application will be further described in detail below in combination with the embodiments.

[0023] Name definition

[0024] The "SNP" in the present application refers to single nucleotide polymorphism, mainly refers to the DNA sequence polymorphism caused by the variation of a single nucleotide at the genome level, and is the most common one of the variations that can be inherited by human beings.

[0025] The "Hardy-Weinberg equilibrium" in the present application refers to the ideal state in which the frequency of each allele is stable and unchangeable in inheritance, that is, the genetic balance is maintained.

[0026] The "wild type" in the present application can refer to the unmutated genotype; the "mutant heterozygote" refers to one of the pair of alleles, one of which is a mutant, and the other is a wild type; and the "mutant homozygote" can refer to the pair of alleles both of which are mutated.

[0027] The "mutation abundance" in the present application refers to VAF, Variant allele fraction, also known as Variant allele frequency, which can refer to the proportion of mutant reads in total reads in the sequencing process, that is, the calculation formula can be:

[0028] VAF = Allele Depth / Total Depth. Wherein, Allele Depth is the reads coverage depth of each site of the genome supporting the mutant genotype, and Total Depth is the total reads coverage depth of the site.

[0029] The "rank sum test" in the present application is also known as Wilcoxon rank sum test or rank sum test, which is a nonparametric test, a statistical method that does not depend on the type of population distribution and does not make statistical inference on the parameters of population.

[0030] Technical solutions

[0031] First, the embodiment of the present application provides a screening method for SNP sites for detecting sample contamination level, which can be applied to an electronic device for performing the following steps: obtaining SNP sites with a mutation frequency of 30%-70% in a target region as candidate marker sites;

[0032] The region between the start site and the end site in the candidate marker sites present on a single chromosome is divided into multiple selection regions, and the length of the selection region is 0.7-1.3 Mb; if there are two or more candidate marker sites in the selection region, the site with an allele frequency of 40%-60% and the most consistent Hardy-Weinberg equilibrium is selected as a marker site, and other candidate marker sites in the selection region are removed.

[0033] Through a series of studies, it is found that based on SNP candidate sites covering N non-contaminated negative sample genomes, removing sites with different genotypes significantly deviating from Hardy-Weinberg equilibrium and possibly linked adjacent sites, contamination marker sites for detecting sample contamination level can be obtained.

[0034] It should be noted that "extracting SNP sites with a mutation frequency of 30%-70% in a target region" refers to extracting all or part of SNP sites with a mutation frequency falling within the range of 30%-70%, for example, extracting all or part of sites with a mutation frequency of 40%-60%, or extracting all or part of sites with a mutation frequency of 45%-55%, or extracting all or part of sites with a mutation frequency of 50%. In some embodiments, the allele frequency can be any one of 40%, 45%, 50%, 55%, and 60%, and preferably 50%.

[0035] The "target region" is any region that is expected to be characterized by the contamination level detection method of the present application, including but not limited to the target region of sequencing and the coverage region of the panel on the genome, etc.

[0036] The information of "SNP sites" and their mutation frequency in the population can be obtained by sequencing or public or commercial databases. The database is not specifically limited, and an existing genetic database can be selected. In some embodiments, the genetic database can be selected from at least one of the gnomAD, 1000G, ExAC, and popfreq_max_20150413 databases.

[0037] In some embodiments, the length of the selected region is 0.7-1.3 Mb. On one hand, this length can effectively avoid the occurrence of adjacent sites in linkage between candidate sites; on the other hand, it can reduce the number of marker sites without affecting the detection results, so as to achieve lower cost and higher detection effectiveness. In some embodiments, the length of the selected region can be any one of 0.7 Mb, 0.8 Mb, 0.9 Mb, 1.0 Mb, 1.1 Mb, 1.2 Mb and 1.3 Mb.

[0038] In some embodiments, the mutation heterozygosity ratio of the candidate marker site is 30%-70%, and the mutation heterozygosity ratio can be any one of 31%, 35%, 40%, 45%, 50%, 55%, 60%, 65% and 70%. The mutation frequency of the site with the above mutation heterozygosity ratio is more stable, and using it as a candidate marker site can increase the effectiveness of subsequent pollution level prediction.

[0039] Preferably, the SNP candidate site is an SNP site with a mutation frequency of 30%-70% in the population based on the genomes of N non-polluted negative samples. The negative sample can be understood as a wild-type sample, and can refer to a negative control sample corresponding to the sample to be detected and without pollution. For example, when the sample to be detected is a tumor sample, the negative sample is a sample of a healthy individual corresponding to the tumor sample and without pollution.

[0040] Preferably, N≥50; preferably, N≥100; preferably, N≥500.

[0041] In some preferred embodiments, when there are multiple sites that best fit the Hardy-Weinberg equilibrium in each of the selected regions, any one of them is selected to reduce or avoid the occurrence of sites in linkage in the pollution marker site.

[0042] In some preferred embodiments, the site with the different genotypes best fitting the Hardy-Weinberg equilibrium refers to the site with the minimum site imbalance coefficient U in the selected region, and the calculation formula of the site imbalance coefficient U is as follows:

[0043] U=(S0-0.25) 2 +(S1-0.5) 2 +(S2-0.25) 2 ; wherein S0, S1 and S2 are the population occurrence frequencies of wild type, mutation heterozygosity and mutation homozygosity in the target region. It should be noted that "wild type" refers to a genotype without mutation.

[0044] The embodiment of the present application further provides an electronic device comprising a memory and a processor, wherein the processor executes a computer program in the memory to perform the screening method for detecting a SNP site of a sample pollution level according to any of the foregoing embodiments.

[0045] The electronic device can comprise a memory, a processor, a bus and a communication interface, which are electrically connected with each other directly or indirectly to realize the transmission or interaction of data. For example, the elements can be electrically connected with each other through one or more buses or signal lines. The processor can process information and / or data related to target recognition to perform one or more functions described in the present application.

[0046] Specifically, the memory can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM) and the like.

[0047] The processor can be an integrated circuit chip having a signal processing capability. The processor can be a general purpose processor, including a central processing unit (CPU), a network processor (NP) and the like; or can be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component.

[0048] The components in the electronic device can be implemented in hardware, software, or a combination thereof. In actual applications, the electronic device can be a server, a cloud platform, a mobile phone, a tablet computer, a notebook computer, an ultra-mobile personal computer (UMPC), a handheld computer, a netbook, a personal digital assistant (PDA), a wearable electronic device, a virtual reality device, or the like, and thus the embodiments of the present application do not limit the type of the electronic device.

[0049] The embodiments of the present application also provide a method for detecting a sample pollution level, which comprises: screening a pollution marker site by using the screening method for screening a SNP site for detecting a sample pollution level according to any one of the preceding embodiments.

[0050] By confirming and detecting the pollution marker site screened by the screening method, the pollution level of the sample can be predicted, and compared with the prior art, the number of marker sites to be detected is smaller, the cost is lower, and the prediction result of the sample pollution level is more accurate.

[0051] In some embodiments, the method further comprises detecting the pollution marker site. Specifically, the way of detecting the site is not limited in any way, for example, any one of a primer, a probe, and a chip can be used, as long as the sample pollution level is detected by detecting the pollution marker site.

[0052] Preferably, the method further comprises: taking the pollution marker site with a homozygous genotype in the target region of the sample to be detected and / or the uncontaminated control sample as a homozygous marker site; and performing rank sum tests on the VAF difference of the mutation abundance of the sample to be detected and the VAF difference distribution of the pollution sample under different pollution levels, respectively, to obtain P values under different pollution levels.

[0053] The VAF difference of the mutation abundance of the sample to be detected is the absolute value of the VAF difference of the homozygous marker site between the sample to be detected and the uncontaminated negative control sample. Specifically, the sample to be detected can be a sample that can be contaminated or has been contaminated with a pollution source (heterogeneous data), which can be a tumor sample, and the negative control sample thereof is a sample of a healthy individual without pollution.

[0054] The VAF difference distribution of the pollution sample under different pollution levels is the absolute value distribution of the VAF difference of the homozygous marker site between the pollution sample with different levels of pollution source and the uncontaminated negative control sample.

[0055] Specifically, the "selecting the contamination marker sites with homozygous genotype in the target region of the sample to be tested and / or the uncontaminated control sample as the homozygous marker sites" refers to selecting the contamination marker sites with wild type or mutant homozygous genotype as the homozygous marker sites based on the genotype information of each contamination marker site present in the target region of the sample to be tested and / or the uncontaminated control sample. The genotype information of each contamination marker site present in the target region of the sample to be tested and / or the uncontaminated control sample can be obtained by means of existing gene detection.

[0056] Preferably, the contamination marker sites with homozygous genotype in the target region of the uncontaminated control sample are selected as the homozygous marker sites.

[0057] Preferably, the contamination marker sites with homozygous genotype in the target region of the sample to be tested and the uncontaminated control sample are selected as the homozygous marker sites. The "contamination marker sites with homozygous genotype in the target region of the sample to be tested and the uncontaminated control sample are selected as the homozygous marker sites" refers to combining the contamination marker sites with homozygous genotype in the target region of the sample to be tested with the contamination marker sites with homozygous genotype in the target region of the uncontaminated control sample, and the combination is selected as the homozygous marker sites. The uncontaminated control sample can include but is not limited to a white blood cell sample.

[0058] In some embodiments, the method further comprises constructing contamination samples with different contamination levels: dividing the samples into samples to be contaminated and uncontaminated control samples, and adding different contamination proportions of the contamination source to the samples to be contaminated to obtain contamination samples with different contamination proportions (mass ratio). The samples to be contaminated or the control samples can be selected from white blood cell samples or existing cell line samples. The contamination source in the contamination sample can be selected from other white blood cell or cell line samples different from the "samples to be contaminated".

[0059] Preferably, the contamination proportions of the contamination samples with different contamination levels are selected from 0.01% to 99.9%. Specifically, the different contamination levels can be any one of 0.01%, 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, and 99.9%.

[0060] In detection, the span range of the different contamination levels can be 0-1%, with a set of contamination samples set at every 0.1% interval (denoted as 0-1%, every 0.1%); the span range of the different contamination levels can be 1%-10%, with a set of contamination samples set at every 0.5% interval; the span range of the different contamination levels can also be 10%-100%, with a set of contamination samples set at every 1% interval, and so on.

[0061] In some embodiments, the statistical analysis is performed by using SPSS 10.0 statistical analysis software. All statistical tests are two-tailed tests.

[0062] Preferably, the method comprises determining the maximum value of P value under different pollution levels, and determining the pollution level corresponding to the maximum value of P value as the pollution level of the sample to be detected. The P value can reflect the degree of coincidence between the VAF difference of the sample and the theoretical distribution, and the greater the P value, the more coincident.

[0063] Preferably, the pollution marker site comprises at least one of (1)-(60) in Table 1.

[0064] Table 1: Pollution marker site

[0065]

[0066]

[0067] Preferably, the pollution marker site comprises at least 10 of (1)-(60) in Table 1; preferably, the pollution marker site comprises at least 20 of (1)-(60) in Table 1.

[0068] In some embodiments, before the marker site is screened, the method further comprises a library construction step and / or a sequencing step of the sample. The library construction and sequencing steps of the sample can be performed according to the existing library construction and sequencing steps, respectively.

[0069] The embodiments of the present application also provide an electronic device comprising a memory and a processor, wherein the processor executes a computer program in the memory to perform the method for detecting the pollution level of a sample according to any of the preceding embodiments.

[0070] In addition, the embodiments of the present application also provide a kit for detecting the pollution level of a sample, which comprises a reagent for detecting a target SNP site, wherein the target SNP site is a marker site screened by the screening method according to any of the preceding embodiments.

[0071] Preferably, the pollution marker site comprises at least one of (1)-(60) in Table 1.

[0072] Preferably, the pollution marker site comprises at least 10 of (1)-(60) in Table 1.

[0073] Preferably, the pollution marker site comprises at least 20 of (1)-(60) in Table 1.

[0074] Preferably, the type of the reagent is selected from at least one of a probe, a primer and a chip.

[0075] Embodiment 1

[0076] A screening method for detecting SNP sites for sample contamination level, specifically comprising the following steps:

[0077] (1) Using the popfreq_max_20150413 database, extracting sites with a population frequency of 40-60% as SNP candidate sites.

[0078] (2) Extracting genotype information of SNP candidate sites in N (500) white blood cell samples (negative samples without contamination). Select sites with a population mutation frequency of 40-60% as candidate marker sites.

[0079] At the same time, the region between the starting site and the ending site in the candidate site existing on each chromosome is divided into multiple selection regions, each with a length of 1 Mb; if there are more than 2 candidate marker sites in the selection region, select the site with an allele frequency of 50% and the smallest site imbalance coefficient U (if there are multiple minimum sites, select one of them) as the contamination marker site, and remove other candidate marker sites in the selection region. Combine all the selected contamination marker sites on all chromosomes to obtain the final contamination marker site collection.

[0080] The calculation formula of the above site imbalance coefficient U is as follows:

[0081] U = (S0-0.25) 2 +(S1-0.5) 2 +(S2-0.25) 2 ; Wherein S0, S1 and S2 are the proportions of wild type, mutant heterozygous type and mutant homozygous type in the genomes of 500 white blood cell samples.

[0082] Embodiment 2

[0083] A set of SNP sites (Panel) for detecting sample contamination level, which is screened by the screening method of embodiment 1, specifically comprising 60 SNP sites in Table 1.

[0084] Embodiment 3

[0085] A method for detecting sample contamination level, the flow chart is shown in the attached Figure 1 , specifically comprising the following steps.

[0086] (1) Construction of VAF difference distribution of contaminated samples under different contamination levels:

[0087] 1.1 Sample Library Construction: One sample from 500 tumor samples was randomly selected as the contaminated sample, and another sample was selected as the contamination source. Different levels of contamination (heterogeneous data) were simulated and introduced into the tumor samples, with each level of contamination repeated 500 times. Contaminated samples at different levels of contamination were obtained. 50 ng of contaminated samples at different levels and control samples (uncontaminated) were prepared for subsequent library construction experiments. The library construction mainly included the following steps:

[0088] a) Fragment and end-repair the sample; b) Ligate the repaired DNA fragments with adapters; c) Perform PCR amplification on the ligated products to obtain sufficient DNA fragments with adapters, which is the pre-library library; d) Purify the pre-library library with magnetic beads, and perform concentration determination and fragment quality control; e) Hybridize the pre-library library with probes; f) Capture the probe-bound samples using streptavidin magnetic beads; g) Perform PCR amplification on the DNA fragments captured by the magnetic beads to obtain sufficient tagged DNA fragments, which is the final library; h) Purify the final library with magnetic beads, and perform concentration determination and fragment quality control, and quantify using qPCR.

[0089] 1.2 Collection of next-generation sequencing data: Based on the established library, multi-exon probes were used to capture the target region. The sequencing was performed using a gene sequencer in 150bp Pair-End mode (Read1:151; Read2:151; Index1:8, Index2:8) according to the instrument's standard operating procedures. The resulting Fastq format next-generation sequencing data was used as raw data.

[0090] 1.3 Data Splitting and Quality Control: bcl2fastq was used for data splitting, and fastp was used for data quality control to obtain high-quality paired sample data (clean data).

[0091] 1.4 Data alignment: The clean data was aligned to the reference genome hg19 sequence using bwa software to obtain the alignment information for each sequencing segment (read). Then, the gencore alignment results were used for deduplication and base correction.

[0092] 1.5 Screening of marker sites: The contaminated marker sites were selected from those provided in Example 2.

[0093] 1.6 Calculation of VAF difference for homozygous marker sites:

[0094] The method for obtaining homozygous marker sites is as follows: based on the genotype information of each contamination marker site on the obtained genome of the sample to be tested and the uncontaminated control sample, the contamination marker sites with wild type or mutant homozygous genotype on the genome of the sample to be tested are combined with the contamination marker sites with wild type or mutant homozygous genotype on the genome of the uncontaminated control sample, and the combined sites are used as the homozygous marker sites.

[0095] The VAF values of the sample to be tested and its control sample at each contamination marker site are obtained by the pysam software package, and the absolute value of the VAF difference between the tumor sample (the sample to be tested) and the control sample at the homozygous marker sites is calculated. The reference distribution of the absolute value of the VAF difference between the contaminated sample and the control sample at the homozygous marker sites under the corresponding contamination level is obtained.

[0096] (2) Rank sum test: the VAF difference between the sample to be tested and the control sample at the homozygous marker sites and the reference distribution of the absolute value of the VAF difference between the contaminated sample and its control sample at the homozygous sites under different contamination levels are subjected to rank sum test, and the P-value (two-tailed test) is calculated.

[0097] (3) Contamination level prediction:

[0098] According to the largest P-value, the contamination level of the sample to be tested is determined.

[0099] Example 4

[0100] A method for detecting the contamination level of a sample, which is substantially the same as the method provided in Example 3, except that 30 contamination marker sites are selected, all of which are selected from the contamination marker sites in Table 1, and the information of the sites is shown in Table 2.

[0101] Table 2 Contamination marker sites

[0102]

[0103]

[0104] Example 5

[0105] A method for detecting the contamination level of a sample, which is substantially the same as the method provided in Example 3, except that 60 contamination marker sites are selected, of which 30 are selected from Table 1 and the other 30 are different from the sites in Table 1 (the bolded sites are marker sites different from Table 1, which are other sites with a mutant heterozygous proportion of 40% to 60%), and the specific sites are shown in Table 3.

[0106] Table 3 Contamination marker sites

[0107]

[0108]

[0109]

[0110] Test Example 1

[0111] The detection effect of the method of Example 3 was compared with that of the prior art Conpair.

[0112] First, a simulated contaminated sample was constructed, and the method of Example 3 was used to detect the contamination level of the simulated sample. The steps of constructing the reference distribution included: randomly selecting 1 sample from 500 cell line samples as a contaminated sample (a sample known to be uncontaminated), randomly selecting other white blood cell or cell line samples as a contamination source, and mixing them at different contamination ratios according to the mass ratio to obtain contaminated samples at different contamination levels. Randomly selecting 500 contaminated samples and their control samples (samples without contamination source) at each contamination level, and obtaining the VAF difference distribution, please refer to Figure 2 .

[0113] Then, real samples at different contamination levels were constructed. The real samples were clinical samples that might have been contaminated. Other cell line samples were selected as contamination sources and mixed into the real samples to obtain real samples at different contamination levels. The method of Example 3 and the existing Conpair method were used to detect the contamination level of the real samples at different contamination levels. The marker site information of the two methods is shown in Table 4.

[0114] Table 4 Detection site

[0115] Method Number of markers Conpair 7387 Method of the invention 60

[0116] The number distribution of homozygous marker sites covered by the detection method provided in Example 3 is shown in the accompanying Figure 2 . Among them, real data (real data) refers to the distribution of homozygous marker sites obtained when detecting real samples; simulate data: simulated data, which is a binomial distribution random number (simulated parameter n is 60, and p is 0.5 binomial distribution B(60, 0.5)). From the results, it can be seen that the number distribution of the screened homozygous marker sites is consistent with the theoretical distribution, indicating that the screening condition of the contamination marker site is effective and close to the ideal situation.

[0117] The detection method of Example 3 and part of the detection results of the sample contamination level of Conpair are shown in Table 5.

[0118] Table 5 Comparison of detection results

[0119]

[0120] From the results, it can be seen that the detection method of Example 3 has higher sensitivity and specificity than the existing Conpair method, and can effectively detect the contamination level of the sample. Figure 3As shown in Table 5, the method of Embodiment 3 of the present invention requires fewer marker sites to be covered. Compared with the limitation of existing tools to more than 1,000 marker sites, it can be flexibly screened or embedded into different panel products, making it widely applicable.

[0121] For the performance of the detection method in Example 3, please refer to the appendix. Figure 4 . Specifically, Figure 4 In Figure A, the performance results of the method in Example 3 in the detection of simulated contaminated samples are shown. Figure 4 Figure B shows the performance results of the method in Example 3 in the detection of real samples under different polluted water conditions.

[0122] The results show that in simulated contaminated samples, the correlation coefficient between the predicted and theoretical values ​​is 0.9984, while in real samples, the correlation coefficient is 0.9987. Conpair performs poorly in real sample detection. Compared to Conpair, the method in Example 3 of this invention predicts values ​​closer to the actual values.

[0123] Experimental Example 2

[0124] The detection effects of the methods provided in Examples 3, 4 and 5 are compared.

[0125] Detection methods

[0126] The methods provided in Examples 3, 4 and 5 were used to test 61 samples with known pollution levels. The test results are shown in Table 6.

[0127] Table 6 Test Results

[0128]

[0129]

[0130] For the performance of the detection method in Example 4, please refer to the appendix. Figure 5 For the performance of the detection method provided in Example 5 (see section B), please refer to [reference needed]. Figure 5 A.

[0131] The results show that the correlation coefficient between the predicted and theoretical values ​​analyzed by the detection method in Example 4 is 0.9937, and the correlation coefficient between the predicted and theoretical values ​​analyzed by the detection method in Example 5 is 0.9993. It can be seen that the performance phenotype changes after randomly replacing 30 markers at the marker site (Example 5) are small (0.9987 vs 0.9993), indicating that as long as the screening conditions for the marker site are met, the performance of any 60rs marker site is stable.

[0132] When the number of markers is reduced to 30 (Example 4), the prediction performance is decreased, the determination coefficient is decreased to 0.9937, indicating that the prediction performance will be decreased by reducing the number of markers, but the 30-marker prediction of pollution is still feasible.

[0133] The above only describes the preferred embodiments of the present application and is not used to limit the present application. The present application can have various modifications and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A screening method for SNP loci for detecting sample contamination levels, characterized in that, SNP sites with a mutation frequency of 30% to 70% in the target region of the population are obtained as candidate marker sites; The region between the start site and the end site in the candidate marker sites present on a single chromosome is divided into multiple selection regions, and the length of the selection region is 0.7 to 1.3 Mb; If there are two or more candidate marker sites in the selection region, the site with an allele frequency of 40% to 60% and different genotypes most consistent with Hardy-Weinberg equilibrium in the selection region is selected as a pollution marker site, and other candidate marker sites in the selection region are removed. The site with different genotypes most consistent with Hardy-Weinberg equilibrium refers to the site with the smallest site imbalance coefficient U in the selection region, and the calculation formula of the site imbalance coefficient U is as follows: U = (S0 - 0.25) + (S1 - 0.5) + (S2 - 0.25) 2 where S0, S1 and S2 are the frequencies of wild type, mutant heterozygote and mutant homozygote in the target region, respectively. 2 where S0, S1 and S2 are the frequencies of wild type, mutant heterozygote and mutant homozygote in the target region, respectively. 2 where S0, S1 and S2 are the frequencies of wild type, mutant heterozygote and mutant homozygote in the target region, respectively.

2. The screening method for SNP loci for detecting a sample contamination level according to claim 1, characterized by, When there are multiple sites most consistent with Hardy-Weinberg equilibrium in a selection region, any one of them is selected.

3. The screening method for SNP loci for detecting sample contamination level according to claim 1, characterized by, The SNP candidate sites are SNP sites with a mutation frequency of 30% to 70% in the genome of N uncontaminated negative samples.

4. The screening method for SNP loci for detecting sample contamination level according to claim 3, characterized by, N≥50.

5. The screening method for SNP loci for detecting sample contamination level according to claim 3, wherein, N≥100.

6. The screening method for a SNP site for detecting a contamination level of a sample according to any one of claims 1 to 5, characterized by, The SNP candidate sites are sites with a frequency of 40% to 60% in the gene database.

7. A method for detecting a level of sample contamination, characterized in that, It comprises: The pollution marker site obtained by screening using the screening method for SNP sites for detecting sample pollution level according to any one of claims 1 to 6.

8. The method for detecting a level of sample contamination of claim 7, wherein, The method comprises: taking the pollution marker site with a homozygous genotype in the target region of the sample to be tested and / or the uncontaminated control sample as a homozygous marker site; and performing rank sum tests on the mutation abundance VAF difference value of the sample to be tested and the VAF difference value distribution of the pollution sample under different pollution levels, respectively, to obtain P values under different pollution levels. The mutation abundance VAF difference value of the sample to be tested is the absolute value of the VAF difference between the sample to be tested and the uncontaminated negative control sample at the homozygous marker site. The VAF difference value distribution of the pollution sample under different pollution levels is the absolute value distribution of the VAF difference between the pollution sample with different levels of pollution source and the uncontaminated negative control sample at the homozygous marker site.

9. The method for detecting a level of sample contamination of claim 8, wherein, The method comprises: determining the maximum value of the P values under different pollution levels, and determining the pollution level corresponding to the maximum value of the P values as the pollution level of the sample to be tested.

10. The method for detecting a level of sample contamination of claim 9, wherein, The pollution proportion of different pollution levels is selected from 0.01% to 99.9%.

11. The method for detecting a level of sample contamination of claim 8, wherein, The pollution marker site is 60 sites or 30 sites, the 60 sites are shown in Table 1 or Table 3, and the 30 sites are shown in Table 2. Table 1 Pollution marker site Table 2 Pollution marker site Table 3 Pollution marker site 。 12. A kit for detecting a level of sample contamination, characterized in that, It comprises reagents for detecting target SNP sites, and the target SNP sites are 30 sites or 60 sites, the 60 sites are shown in Table 1 or Table 3 of claim 11, and the 30 sites are shown in Table 2 of claim 11.

13. An electronic device, comprising: The computer program product comprises a computer readable storage medium having program instructions stored on it. The program instructions can comprise one or more of the steps of the method of any one of claims 1 to 6 or the method of any one of claims 7 to 11. The computer program product comprises a computer readable storage medium having program instructions stored on it. The program instructions can comprise one or more of the steps of the method of any one of claims 1 to 6 or the method of any one of claims 7 to 11. The computer program product comprises a computer readable storage medium having program instructions stored on it. The program instructions can comprise one or more of the steps of the method of any one of claims 1 to 6 or the method of any one of claims 7 to

Citation Information

Patent Citations

  • Method for analyzing gene mutation based on second-generation sequencing data

    CN108304694A

  • Biological information quality control method and device based on next-generation sequencing and storage medium

    CN110444255A