Method and apparatus for detecting sample contamination and identifying sample mismatches
By screening mutation sites and constructing indicators such as correlation level, homozygosity ratio, and mutation abundance, the problem of sample mismatch and contamination identification in high-throughput sequencing was solved, and rapid and accurate sample contamination and mismatch detection was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU BURNING ROCK DX CO LTD
- Filing Date
- 2023-03-09
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies cannot effectively identify sample mismatches and contamination in high-throughput sequencing, leading to erroneous variant detection results, and strict quality control standards may miss the opportunity to identify sample mismatches.
By screening mutation sites and constructing indicators such as correlation level, homozygosity ratio and mutation abundance, sample contamination and mismatch can be quickly and accurately identified, and computer programs can be used to achieve automated judgment.
It enables low-cost, rapid, and accurate identification of sample mismatches and contamination, improving the accuracy and efficiency of detection and reducing the possibility of false detections.
Smart Images

Figure CN116312779B_ABST
Abstract
Description
Technical Field
[0001] This application discloses a method and apparatus for detecting contamination and mismatch in test samples during high-throughput sequencing. This application also provides a system, device, and computer-readable medium for assessing the level of sample contamination. Background Technology
[0002] In high-throughput assays based on paired samples, sample mismatches or contamination are prone to occur during experimental procedures because both the test sample and the paired sample need to be sequenced simultaneously. Sample mismatches and contamination typically lead to erroneous variant detection results; therefore, identifying sample mismatches and contamination is a necessary quality control step in high-throughput assays. Normally, the test sample and the paired sample should originate from the same individual. However, sample mismatches can occur due to errors in manual labeling during the assay process. Sample mismatch refers to the test sample and the paired sample originating from different individuals. Sample contamination usually arises from the sample preparation process, where DNA from other individuals is mixed into the test sample slice.
[0003] Existing technologies typically cannot directly determine mismatches or contamination between test samples and paired samples. Sample contamination and mismatches are often impossible to completely eliminate during manual sample processing. Furthermore, if all samples failing to meet quality control standards are classified as contaminated, the possibility of identifying sample mismatches and thus easily and quickly re-pairing samples and conducting subsequent experiments is lost. Therefore, existing technologies lack a method that can easily detect contamination while effectively identifying sample mismatches. This application provides a mutation abundance-based detection method for high-throughput sequencing, enabling low-cost, rapid, and accurate identification of sample mismatches and contamination. Summary of the Invention
[0004] This application relates to a method, apparatus, device, and storage medium for detecting contamination in a test sample and identifying sample mismatches. It enables low-cost, rapid, and accurate identification of mismatches and contamination in test samples.
[0005] On the one hand, this application provides a method for detecting contamination in a sample to be tested, wherein the method includes the following steps:
[0006] Step 1: Screening for mutation sites used to identify contamination in the sample to be tested;
[0007] Step 2: Based on the mutation sites of the test sample and the paired sample, construct indicators to determine the contamination of the test sample. The above indicators include any one or more of the average values of correlation level, homozygosity ratio and mutation abundance.
[0008] Step 3: Identify and determine the contamination of the sample to be tested based on at least one of the judgment indicators constructed in Step 2.
[0009] On the other hand, this application provides a method for identifying sample mismatches, wherein the method includes the following steps:
[0010] Step 1: Screening for mutation sites to identify sample mismatches;
[0011] Step 2: Based on the mutation sites of the test sample and the paired sample, construct indicators to determine sample mismatch. The above indicators include any one or more of the following: correlation level, homozygous ratio, average mutation abundance, and paired homozygous variation indicators.
[0012] Step 3: Identify and determine sample mismatches based on at least one of the judgment indicators constructed in Step 2.
[0013] On the other hand, this application provides a method for identifying sample mismatches, wherein the method includes:
[0014] Perform one or more of the methods described above on the sample to detect whether the sample is contaminated; and
[0015] If contamination is detected in the sample to be tested, one or more of the methods described above will be performed on the sample to be tested to further identify whether there is a mismatch between the sample to be tested and the paired sample.
[0016] On the other hand, this application provides an apparatus for detecting contamination in a test sample and / or identifying sample mismatch, comprising:
[0017] The screening module is configured to screen for detecting contamination in the sample to be tested and / or identifying mutation sites of sample mismatch;
[0018] The building module is configured to construct indicators for determining contamination and / or identifying sample mismatches based on the mutation sites of the test sample and the paired sample. The determination indicators include any one or more of the following: correlation level, homozygous ratio, and average homozygous mutation abundance of the sample.
[0019] The determination module is configured to identify and determine contamination and / or mismatch of the test sample based on at least one determination index constructed in step two.
[0020] On the other hand, this application provides an apparatus for detecting contamination in a test sample and / or identifying sample mismatch, comprising:
[0021] One or more processors;
[0022] Storage device, on which one or more programs are stored,
[0023] When the above one or more programs are executed by the above one or more processors, the above one or more processors shall implement the above method.
[0024] On the other hand, this application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the above-described method when executed by one or more processors. Attached Figure Description
[0025] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this specification and, together with the description, serve to explain the principles of this specification.
[0026] Figure 1 The mutation abundance distribution of the unmismatched and uncontaminated sample S1 is shown, where the x-axis represents the abundance of mutations in the paired samples and the y-axis represents the abundance of mutations in the sample to be tested.
[0027] Figure 2 The mutation abundance distribution of the contaminated test sample S2 is shown, where the x-axis represents the abundance of mutations in paired samples and the y-axis represents the abundance of mutations in the test sample.
[0028] Figure 3 The mutation abundance distribution of the contaminated test sample S3 is shown, where the x-axis represents the abundance of mutations in paired samples and the y-axis represents the abundance of mutations in the test sample.
[0029] Figure 4 The mutation abundance distribution of the mismatched test sample S4 is shown, where the x-axis represents the abundance of mutations in the paired samples and the y-axis represents the abundance of mutations in the test sample. Detailed Implementation
[0030] I. Definition
[0031] In this application, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the terms and laboratory procedures related to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, and immunology used herein are all widely used terms and routine procedures in their respective fields. To better understand this application, definitions and explanations of relevant terms are provided below.
[0032] As used herein, the term “sample contamination” refers to a situation where a sample to be tested is contaminated with samples from other individuals during preparation or other processing, such as the contamination of nucleic acids from other individuals during the preparation of a sample to be sequenced.
[0033] As used herein, the term "sample mismatch" refers to a situation where the test sample and the paired sample come from different individuals. However, it should be understood that in practical applications, when a mismatch occurs between the test sample and the paired sample, the resulting parameter results are similar to extreme contamination. Therefore, in this application, when the test sample is determined to be severely contaminated (e.g., when it performs poorly according to the indicators provided in this application), it is necessary to further identify whether the test sample is a sample mismatch.
[0034] As used in this article, the term "wild type" refers to the form of a gene or organism that is common or non-mutant in nature. In other words, it refers to the phenotype with the highest frequency observed in wild populations, or the system, organism, or gene that possesses this phenotype.
[0035] As used in this article, the term "mutation" refers to a change in the structure of a gene that results in a stable, heritable change in the genotype of a cell, virus, or microorganism.
[0036] II. Detailed Implementation Plan
[0037] On the other hand, this application provides a method for detecting contamination in a sample to be tested, wherein the method includes the following steps:
[0038] Step 1: Screening for mutation sites used to identify contamination in the sample to be tested;
[0039] Step 2: Based on the mutation sites of the test sample and the paired sample, construct indicators to determine the contamination of the test sample. The above indicators include any one or more of the following: correlation level, homozygous ratio, and average homozygous mutation abundance of the sample.
[0040] Step 3: Identify and determine the contamination of the sample to be tested based on at least one of the judgment indicators constructed in Step 2.
[0041] In some implementations, the aforementioned correlation level (homo.cor) is the Pearson correlation coefficient obtained by performing a Pearson correlation test on the mutation abundance of the test sample and the paired sample; when the aforementioned correlation level is below 90%, the test sample is determined to be contaminated.
[0042] In some implementations, the homozygosity ratio is the ratio of the number of mutation sites with mutation abundance greater than or equal to a first preset threshold in both the test sample and the paired sample to the number of mutation sites with mutation abundance greater than or equal to the first preset threshold in the paired sample; wherein, the formula for calculating the homozygosity ratio (homo.ratio) is:
[0043]
[0044] Wherein, homo.ratio represents the homozygous proportion of the sample, N1 represents the number of mutation sites in both the test sample and the paired sample whose mutation abundance is higher than or equal to the first preset threshold, and N2 represents the number of mutation sites in the paired sample whose mutation abundance is higher than or equal to the first preset threshold. The mutation sites in the paired sample whose mutation abundance is higher than or equal to the first preset threshold can include those in both the test sample and the paired sample whose mutation abundance is higher than or equal to the first preset threshold. That is, the mutation sites corresponding to the number of mutation sites represented by N1 are a subset of the mutation sites corresponding to the number of mutation sites represented by N2.
[0045] In some preferred embodiments, N2 ≥ 100.
[0046] In some implementations, when the homo.ratio is below 90%, the sample is determined to be contaminated, and when the homo.ratio is above or equal to 90%, the sample is determined to be uncontaminated.
[0047] In some implementations, the average homozygous mutation abundance (homoAF) of the above-mentioned samples is the average mutation abundance of mutations in the test sample whose mutation abundance in the paired samples is higher than or equal to a first preset threshold (e.g., 90%-98%, preferably 95%); when the average mutation abundance is set to be lower than 0.975, the test sample is determined to be contaminated.
[0048] In some implementations, the test sample is preferably derived from the subject's tumor tissue or its nucleic acid.
[0049] In some implementations, the paired samples are derived from normal tissues or normal cells of the same subject.
[0050] In some preferred embodiments, the aforementioned normal tissue includes adjacent tissue, leukocytes, etc.
[0051] In some implementations, the mutation sites screened in step one above are the sites corresponding to mutations detected in at least one of the test sample or the paired sample that have passed mutation quality control.
[0052] In some implementations, the above-mentioned mutation quality control is performed using mutation detection software.
[0053] In some preferred embodiments, the aforementioned variant detection software is selected from Vardict, Varscan, GATK (Genome Analysis Toolkit), or Mutect, etc.
[0054] In some implementations, the aforementioned mutation detection software is Vardict.
[0055] In some implementations, the mutation sites screened in step one above are the sites corresponding to any mutations in the test sample or paired sample whose mutation abundance is higher than or equal to the wild-type filtering threshold; preferably, the wild-type filtering threshold is 30%.
[0056] In some implementations, the mutation sites screened in step one above are the sites corresponding to mutations whose maximum population frequency in different population genomes is greater than or equal to 0.1%.
[0057] In some implementations, population frequencies in the population genome are queried from one or more population genome databases.
[0058] In some preferred embodiments, the aforementioned population genome database is selected from the 1000 Genomes Project, dbSNP, gnomAD (genome aggregation database), and ExAC (the Exome Aggregation Consortium), etc.
[0059] In some implementations, the first preset threshold is 90%-98%.
[0060] In some preferred embodiments, the first preset threshold is 95%.
[0061] In some implementations, the second preset threshold is 65%-90%.
[0062] In some preferred embodiments, the second preset threshold is 75%.
[0063] In one aspect, this application provides a method for identifying sample mismatches, wherein the method includes the following steps:
[0064] Step 1: Screening for mutation sites to identify sample mismatches;
[0065] Step 2: Based on the mutation sites of the test sample and the paired sample, construct an index to determine sample mismatch. The above-mentioned index includes any one or more of the following: correlation level, homozygous ratio, sample homozygous mutation abundance, and the average value of the paired homozygous variation index.
[0066] Step 3: Identify and determine sample mismatches based on at least one of the judgment indicators constructed in Step 2.
[0067] In some implementation schemes, the correlation level mentioned in step two is the Pearson correlation coefficient obtained by performing a Pearson correlation test on the mutation abundance of the test sample and the paired sample; when the correlation level is less than 50%, it is determined whether the sample is mismatched.
[0068] In some implementations, the homozygosity ratio is the ratio of the number of mutation sites with mutation abundance greater than or equal to a first preset threshold in both the test sample and the paired sample to the number of mutation sites with mutation abundance greater than or equal to the first preset threshold in the paired sample; wherein, the formula for calculating the homozygosity ratio (homo.ratio) is:
[0069]
[0070] Wherein, homo.ratio represents the homozygous proportion of the sample, N1 represents the number of mutation sites in both the test sample and the paired sample whose mutation abundance is higher than or equal to the first preset threshold, and N2 represents the number of mutation sites in the paired sample whose mutation abundance is higher than or equal to the first preset threshold. The mutation sites in the paired sample whose mutation abundance is higher than or equal to the first preset threshold can include those in both the test sample and the paired sample whose mutation abundance is higher than or equal to the first preset threshold. That is, the mutation sites corresponding to the number of mutation sites represented by N1 are a subset of the mutation sites corresponding to the number of mutation sites represented by N2.
[0071] In some preferred embodiments, N2 ≥ 100.
[0072] In some implementations, a determination is made as to whether the above samples are mismatched when the homo.ratio is below 75%.
[0073] In some implementations, the average homozygous mutation abundance (homoAF) is the average mutation abundance of mutations with a mutation abundance of 95% or higher in the paired samples in the test sample; when the average homozygous mutation abundance is lower than 0.9, it is determined whether the sample is mismatched.
[0074] In some embodiments, the determination of whether a sample is mismatched includes: determining whether the sample is mismatched based on the pair index. In some embodiments, the pair index includes a pair ratio and a homozygous pair ratio; wherein, the pair ratio is the ratio of the number of mutation sites in both the test sample and the paired sample whose mutation abundance is higher than or equal to a first preset threshold to the number of mutation sites in the test sample whose mutation abundance is higher than or equal to the first preset threshold; the homozygous pair ratio is the ratio of the number of mutation sites in both the test sample and the paired sample whose mutation abundance is higher than or equal to the first preset threshold to the number of sites in the test sample whose mutation abundance is higher than or equal to the first preset threshold and in the paired sample whose mutation abundance is higher than or equal to a second preset threshold.
[0075] In some implementations, a sample mismatch is determined when pair.ratio is below 85% and homo.pair.ratio is above or equal to 95%.
[0076] In some implementations, the formula for calculating the pair ratio is as follows:
[0077]
[0078] Wherein, pair.ratio is part of the paired homozygous variation index, N1 represents the number of mutation sites in both the test sample and the paired sample whose mutation abundance is higher than or equal to the first preset threshold, and N3 represents the number of mutation sites in the test sample whose mutation abundance is higher than the first preset threshold; where N3≥100.
[0079] In some implementations, the formula for calculating the homozygous pair ratio (homo.pair.ratio) is as follows:
[0080]
[0081] Here, homo.pair.ratio is part of the paired homozygous variation index, N1 represents the number of mutation sites in both the test sample and the paired sample where the mutation abundance is higher than or equal to the first preset threshold, and N4 represents the number of mutation sites in the test sample where the mutation abundance is higher than or equal to the first preset threshold and in the paired sample where the mutation abundance is higher than or equal to the second preset threshold; where N4≥100.
[0082] In some implementations, the determination of whether a sample is mismatched further includes: verifying the input information and / or identification number of the sample; verifying the pairing information between the sample to be tested and the corresponding paired sample; and reviewing the inspection records or results. The verification and review can be performed manually or automatically.
[0083] In some implementations, the test sample is preferably derived from the subject's tumor tissue or its nucleic acid.
[0084] In some implementations, the paired samples are derived from normal tissues or normal cells of the same subject.
[0085] In some preferred embodiments, the aforementioned normal tissue includes adjacent tissue, leukocytes, etc.
[0086] In some implementations, the mutation sites screened in step one above are the sites corresponding to mutations detected in at least one of the test sample or the paired sample that have passed mutation quality control.
[0087] In some implementations, the above-mentioned mutation quality control is performed using mutation detection software.
[0088] In some preferred embodiments, the aforementioned variant detection software is selected from Vardict, Varscan, GATK (Genome Analysis Toolkit), or Mutect, etc.
[0089] In some implementations, the aforementioned mutation detection software is Vardict.
[0090] In some implementations, the mutation sites screened in step one above are the sites corresponding to any mutations in the test sample or paired sample whose mutation abundance is higher than or equal to the wild-type filtering threshold; preferably, the wild-type filtering threshold is 30%.
[0091] In some implementations, the mutation sites screened in step one above are the sites corresponding to mutations whose maximum population frequency in different population genomes is greater than or equal to 0.1%.
[0092] In some implementations, population frequencies in the population genome are queried from one or more population genome databases.
[0093] In some preferred embodiments, the aforementioned population genome database is selected from the 1000 Genomes Project, dbSNP, gnomAD (genome aggregation database), and ExAC (the Exome Aggregation Consortium), etc.
[0094] In some implementations, the first preset threshold is 90%-98%.
[0095] In some preferred embodiments, the first preset threshold is 95%.
[0096] In some implementations, the second preset threshold is 65%-90%.
[0097] In some preferred embodiments, the second preset threshold is 75%.
[0098] On the other hand, this application provides a method for identifying sample mismatches, wherein the method includes:
[0099] Perform the method described in any of the foregoing aspects on the sample to be tested to detect whether the sample is contaminated; and
[0100] If contamination is detected in the sample to be tested, perform the method described in any of the above-mentioned aspects on the sample to be tested to identify whether there is a mismatch between the sample to be tested and the paired sample.
[0101] On the other hand, this application provides an apparatus for detecting contamination in a test sample and / or identifying sample mismatch, comprising:
[0102] The screening module is configured to screen for detecting contamination in the sample to be tested and / or identifying mutation sites of sample mismatch;
[0103] The building module is configured to construct indicators for determining contamination and / or identifying sample mismatches based on the mutation sites of the test sample and the paired sample. The determination indicators include any one or more of the following: correlation level, homozygous ratio, and average homozygous mutation abundance of the sample.
[0104] The determination module is configured to identify and determine contamination and / or mismatch of the test sample based on at least one determination index constructed in step two.
[0105] On the other hand, this application provides an apparatus for detecting contamination in a test sample and / or identifying sample mismatch, comprising:
[0106] One or more processors;
[0107] Storage device, on which one or more programs are stored,
[0108] When the above one or more programs are executed by the above one or more processors, the above one or more processors implement the method as described above.
[0109] On the other hand, this application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by one or more processors, implements the method as described in any of the above aspects.
[0110] For the purpose of clarity and concise description, the features are described herein as part of some identical or separate embodiments; however, it will be understood that the scope of this application may include some embodiments having a combination of all or some of the features described.
[0111] Example
[0112] Example 1: Generating indicators for judging sample mismatch and contamination of test samples
[0113] 1. Sample preparation and sequencing
[0114] DNA extraction from tumor tissue samples and paired samples was performed according to the instructions provided with the kit (QIAamp DNA FFPE Tissue Kit, manufactured by QIAGEN). The extracted DNA was fragmented into DNA fragments averaging 200 bp. A pre-library was then prepared using the classic sonication-induced double-strand ligation method. The procedure included end repair, 3' end A addition, adapter ligation, adapter ligation product purification, pre-library amplification, and amplified pre-library purification. The purified pre-library yield was above 500 ng (assessed using a Qubit HS analysis kit). Specific RNA was selected to capture Agilent probes targeting the target gene region of the kit. These probes hybridized with the pre-library to capture specific fragments, and non-specific fragments were eluted. Amplification was performed by Post-PCR, and the amplified products were purified. The fragment size and yield of the final purified library were evaluated. The peak DNA fragment length was approximately 350 bp, and the yield was between 10-300 ng. Finally, sequencing was performed according to the sequencer instructions using an Inmena sequencer.
[0115] After obtaining the sequencing results, the method of this application is performed on the offline data.
[0116] 2. Screening for mutation sites to identify sample mismatches and contamination in the test sample includes the following steps:
[0117] (1) Generate sequence alignment files
[0118] After quality control of the FASTQ files from high-throughput sequencing, the FASTQ files were aligned and back-aligned using the human reference genome (hg19 / b37) with the alignment software BWA-MEM (0.7.10) to generate SAM files. The SAM files were then converted to BAM files using Samtools (0.1.19) software. The detection method proposed in this application was then used for detection.
[0119] (2) Variants were detected at all sites in the target region using the vardict software. The mutation site results output by the software were then filtered to remove mutations with the output label “Failed”.
[0120] (3) Calculate the mutation abundance of the obtained mutation sites: mutation abundance is the ratio of the number of reads of the variant supported by the mutation site to the total number of reads covered by the site.
[0121] (4) Using the chromosomal location of the mutation site and the wild-type and variant genotypes as unique markers of mutation, the mutations detected in at least one of the test sample and the paired sample are counted. For example... Figure 1-4As shown, the mutation abundance distribution of samples S1, S2, S3, and S4 is displayed. The x-axis and y-axis represent the abundance of the mutation in the paired sample and the test sample, respectively. Sites with a mutation abundance of 30% or higher in at least one of the test sample and the paired sample were selected.
[0122] (5) Further select mutations with a maximum population frequency of 0.1% or higher in different population genomes, which are then used as mutation sites for sample mismatch and contamination identification. The population genomes used are from one or more population genome databases such as 1000GenomeProject and dbSNP database (e.g., 1000GenomeProject and ExAC database).
[0123] 3. Construct a discriminant index for identifying sample mismatch and contamination of test samples.
[0124] (1) Calculate homo.cor, which is the Pearson correlation coefficient of mutation abundance between the test sample and the paired sample. Figure 1-4 As shown, the homo.cor values for samples S1, S2, S3, and S4 are 0.971, 0.841, 0.485, and 0.307, respectively.
[0125] (2) Calculate homo.ratio.
[0126]
[0127] Where N1 is the number of mutations with a mutation abundance greater than or equal to 95% in both the test sample and the paired sample; N2 is the number of mutations with a mutation abundance greater than or equal to 95% in the paired sample; N2 ≥ 100. Figure 1-4 As shown, the homoratios of samples S1, S2, S3 and S4 are 0.994, 0.78, 0.645 and 0.597, respectively; N1 is 163, 156, 120 and 132, respectively; and N2 is 164, 200, 186 and 221, respectively.
[0128] (3) Calculate homoAF, which is the average mutation abundance in the test sample of mutations with a mutation abundance of 95% or higher in the paired samples. Figure 1-4 As shown, the homoAF values for samples S1, S2, S3 and S4 are 0.993, 0.971, 0.835 and 0.755, respectively.
[0129] (4) Calculate the pair.ratio and homo.pair.ratio contained in the pair index.
[0130]
[0131] Wherein, pair.ratio is part of the paired homozygous variation index, N1 represents the number of mutation sites with a mutation abundance of ≥95% in both the test sample and the paired sample, and N3 represents the number of mutation sites with a mutation abundance of ≥95% in the test sample; N3≥100. Figures 1-4 As shown, the pair.ratios of samples S1, S2, S3 and S4 are 0.994, 1, 0.93 and 0.695, respectively, and the N3 is 164, 156, 129 and 190, respectively.
[0132]
[0133] Here, homo.pair.ratio is part of the paired homozygous variation index, N1 represents the number of mutation sites with a mutation abundance of ≥95% in both the test sample and the paired sample, and N4 represents the number of mutation sites with a mutation abundance of ≥95% in the test sample and ≥75% in the paired sample; N4 ≥ 100. Figures 1-4 As shown, the homo.pair.ratios of samples S1, S2, S3 and S4 are 0.994, 1, 0.992 and 1, respectively, and the N4 is 164, 156, 121 and 132, respectively.
[0134] (5) Analysis of the results of samples S1, S2, S3 and S4
[0135] The homo.cor of sample S1 is 0.971 (≥90%), homo.ratio is 0.994 (≥90%), and homoAF is 0.993 (≥0.975). Based on one or more of these indicators, it can be determined that sample S1 is uncontaminated and properly paired.
[0136] Sample S2 has homo.cor = 0.841 (<90%), homo.ratio = 0.78 (<90%), and homoAF = 0.971 (<0.975). Based on one or more of these indicators, it can be directly determined that sample S2 is contaminated. In this embodiment, for sample S2 that has been detected as contaminated, further investigation can be conducted to determine whether there is a mismatch between sample S2 and its paired sample. Sample S2 has pair.ratio = 1 and homo.pair.ratio = 1. If a mismatch is not met, the pair... The pair index requirement means that sample S2 is contaminated but not mismatched. However, it should be understood that mismatch is similar to an extreme form of contamination. Therefore, from the perspective of index values, when the test sample and the paired sample are mismatched, the above indicators (homo.cor, homo.ratio, and / or homoAF) will show significantly worse performance than ordinary contamination. When the test sample clearly falls within the threshold range of contamination without mismatch (e.g., 50% ≤ homo.cor < 90%, 75% ≤ homo.ratio < 90%, and / or 0.9 ≤ homoAF < 0.975), the pair index of the sample does not need to be calculated. However, when the above indicators of the test sample indicate a high risk of mismatch (e.g., homo.cor < 50%, homo.ratio < 75%, and / or homoAF < 0.9), further investigation should be conducted to determine if the test sample is mismatched. The investigation methods include: in addition to the above-mentioned pair index calculations... In addition to the calculations performed by the index, the following steps may also be included, either manually or automatically: verifying the input information and / or identification number of the sample; verifying the pairing information between the sample to be tested and the corresponding paired sample; and reviewing the inspection records or results.
[0137] Sample S3 has homo.cor = 0.485 (<50%), homo.ratio = 0.645 (<75%), and homoAF = 0.835 (<0.9). While one or more of these indicators meet the criteria for contamination, they all fall within the threshold range requiring further investigation of mismatch possibilities. This threshold indicates a high risk of mismatch in sample S3. Therefore, the pair index of sample S3 needs to be calculated. Sample S3 has pair.ratio = 0.93 (>0.85%) and homo.pair.ratio = 0.992 (>95%), which does not meet the requirements for the pair index in cases of mismatch. Therefore, sample S3 is contaminated but not mismatched. It should be understood that when the contamination level of a sample is high, the homo.cor, homo.ratio, and / or homoAF provided in this application will indicate a higher risk of mismatch than ordinary contamination. In this case, it is necessary to further determine whether the sample is indeed mismatched. The determination method, in addition to the above-mentioned pair index... In addition to the calculations performed by the index, the following steps may also be included, either manually or automatically: verifying the input information and / or identification number of the sample; verifying the pairing information between the sample to be tested and the corresponding paired sample; and reviewing the inspection records or results.
[0138] Sample S4 has a homo.cor of 0.307 (<50%), a homo.ratio of 0.597 (<75%), and a homoAF of 0.755 (<0.9). All of these indicators fall within the threshold range that requires further investigation of the possibility of mismatch. Subsequently, the pair index of sample S4 was calculated. The pair.ratio of sample S4 is 0.695 (<85%) and the homo.pair.ratio is 1 (>95%), which meets the requirements for the pair index under mismatch conditions. Therefore, sample S4 is a mismatch.
[0139] Example 2: Performance Confirmation of Discriminant Indicators
[0140] 1. Performance evaluation of indicators under sample mismatch conditions
[0141] One hundred pairs of paired real sample DNA (sample i to be tested and paired sample i') were shuffled to form 100 mismatched sample pairs, where the control of sample i to be tested is paired sample j', and j ≠ i. NGS sequencing was performed on the samples, and analysis was conducted according to the steps described in Example 1. The homozygosity ratio (homo.ratio), correlation level (homo.cor), and average homozygosity mutation abundance (homoAF) of the 100 mismatched sample pairs were calculated. The range, mean, and variance of the three parameters for the 100 mismatched sample pairs are shown in Table 1. It can be seen that, under the simulated sample mismatch conditions, the homozygosity ratio (homo.ratio), correlation level (homo.cor), and pair index are all stable within the thresholds for determining sample mismatch or high contamination proposed in Example 1.
[0142] Table 1: Summary of Pollution Assessment Parameters for Mismatched Combinations
[0143]
[0144] 2. Performance evaluation of indicators under contaminated sample conditions
[0145] For 100 test samples, 40 tumor samples with contamination rates of 5%, 10%, and 20% were generated by simulating the mixing of sequencing data from other sources. The contamination index and the range, mean, and variance of the three types of parameters for these 40 simulated contamination test samples are shown in Table 2.
[0146] Table 2: Summary of Evaluation Parameters for Pollution Simulation Data
[0147]
[0148] 3. DNA mixing experiments with real samples to evaluate the performance of the sample contamination discrimination index.
[0149] Two pairs of real samples were selected, one containing tumor tissue and the other containing normal adjacent normal tissue as control DNA. The tissue DNA was mixed into the control DNA of another sample at ratios of 10% and 20%, respectively. The resulting contamination parameters are shown in Table 3.
[0150] Table 3: Summary of DNA mixing simulation parameters for real samples
[0151]
[0152]
Claims
1. A method for detecting contamination in a test sample, wherein, The method includes the following steps: Step 1: Screening for mutation sites used to identify contamination in the sample to be tested; Step 2: Based on the mutation sites of the test sample and the paired sample, construct an index to determine the contamination of the test sample. The index includes any one or more of the following: correlation level, homozygous ratio, and average homozygous mutation abundance of the sample. Step 3: Identify and determine the contamination of the sample to be tested based on at least one indicator constructed in Step 2; The correlation level (homo.cor) is the Pearson correlation coefficient obtained by performing a Pearson correlation test on the mutation abundance of the test sample and the paired sample. When the correlation level is below 90%, the sample to be tested is determined to be contaminated. The homozygosity ratio is the ratio of the number of mutation sites with mutation abundance higher than or equal to a first preset threshold in both the test sample and the paired sample to the number of mutation sites with mutation abundance higher than or equal to the first preset threshold in the paired sample. The formula for calculating the homozygous ratio (homo.ratio) is as follows: Where homo.ratio represents the homozygosity ratio of the sample, N1 represents the number of mutation sites in both the test sample and the paired sample whose mutation abundance is higher than or equal to the first preset threshold, and N2 represents the number of mutation sites in the paired sample whose mutation abundance is higher than or equal to the first preset threshold. When the homo.ratio is below 90%, the sample to be tested is determined to be contaminated; when the homo.ratio is above or equal to 90%, the sample to be tested is determined to be uncontaminated. The average homozygous mutation abundance of the sample (homoAF) is the average mutation abundance of mutations in the test sample whose mutation abundance is higher than or equal to the first preset threshold in the paired samples. If the average abundance of homozygous mutations in a sample is less than 0.975, the sample to be tested is determined to be contaminated.
2. The method according to claim 1, wherein, N2≥100。 3. A method for identifying sample mismatches, wherein, The method includes the following steps: Step 1: Screening for mutation sites to identify sample mismatches; Step 2: Based on the mutation sites of the test sample and the paired sample, construct an index to determine sample mismatch. The index includes any one or more of the following: correlation level, homozygous ratio, average abundance of homozygous mutations in the sample, and paired homozygous variation index. Step 3: Identify and determine sample mismatches based on at least one of the indicators constructed in Step 2; The correlation level (homo.cor) is the Pearson correlation coefficient obtained by performing a Pearson correlation test on the mutation abundance of the test sample and the paired sample. When the correlation level is below 50%, it is determined whether the sample is mismatched. The homozygosity ratio (homo.ratio) is the ratio of the number of mutation sites with mutation abundance higher than or equal to the first preset threshold in both the test sample and the paired sample to the number of mutation sites with mutation abundance higher than or equal to the first preset threshold in the paired sample. The formula for calculating the homozygous ratio (homo.ratio) is as follows: Where homo.ratio represents the homozygosity ratio of the sample, N1 represents the number of mutation sites in both the test sample and the paired sample whose mutation abundance is higher than or equal to the first preset threshold, and N2 represents the number of mutation sites in the paired sample whose mutation abundance is higher than or equal to the first preset threshold. When homo.ratio is below 75%, a determination is made as to whether the sample is mismatched. The average homozygous mutation abundance of the sample (homoAF) is the average mutation abundance of mutations in the test sample whose mutation abundance is higher than or equal to the first preset threshold in the paired samples. When the average abundance of homozygous mutations in the sample is less than 0.9, it is determined whether the sample is mismatched.
4. The method according to claim 3, wherein, N2≥100。 5. The method according to claim 3, wherein, The determination of whether the sample is mismatched includes: The pair index is used to determine whether the sample is mismatched. The pair index includes the pair ratio and the homozygous pair ratio. Wherein, the pair ratio is the ratio of the number of mutation sites in both the test sample and the paired sample whose mutation abundance is higher than or equal to the first preset threshold to the number of mutation sites in the test sample whose mutation abundance is higher than or equal to the first preset threshold. Wherein, the homozygous pair ratio (homo.pair.ratio) is the ratio of the number of mutation sites in both the test sample and the paired sample whose mutation abundance is higher than or equal to the first preset threshold to the number of sites in the test sample whose mutation abundance is higher than or equal to the first preset threshold and in the paired sample whose mutation abundance is higher than or equal to the second preset threshold. When pair.ratio is below 85% and homo.pair.ratio is above or equal to 95%, a mismatch is determined to exist in the sample.
6. The method according to claim 5, wherein, The formula for calculating the pair ratio (pair.ratio) is as follows: Wherein, pair.ratio is part of the paired homozygous mutation index, N1 represents the number of mutation sites in both the test sample and the paired sample where the mutation abundance is higher than or equal to the first preset threshold, and N3 represents the number of mutation sites in the test sample where the mutation abundance is higher than or equal to the first preset threshold. Where N3≥100.
7. The method according to claim 6, wherein, The formula for calculating the homozygous pairing ratio (homo.pair.ratio) is as follows: Wherein, homo.pair.ratio is part of the paired homozygous mutation index, N1 represents the number of mutation sites in both the test sample and the paired sample where the mutation abundance is higher than or equal to the first preset threshold, and N4 represents the number of mutation sites in the test sample where the mutation abundance is higher than or equal to the first preset threshold and in the paired sample where the mutation abundance is higher than or equal to the second preset threshold. Where N4≥100.
8. A method for identifying sample mismatches, wherein, The method includes: Perform the method as described in claim 1 on the test sample to detect whether the test sample is contaminated; and If contamination is detected in the sample to be tested, the method described in any one of claims 5-7 is performed on the sample to be tested to further identify whether there is a mismatch between the sample to be tested and the paired sample.
9. The method according to claim 8, wherein, The paired samples were derived from normal tissues or normal cells from the same subject.
10. The method according to claim 9, wherein, The test sample was derived from the subject's tumor tissue or its nucleic acid.
11. The method according to claim 9, wherein, The normal tissues include adjacent tissues and white blood cells.
12. The method according to claim 3, wherein, The mutation sites screened in step one are the sites corresponding to mutations detected in at least one of the test sample or the paired sample that have passed mutation quality control. The mutation quality control is performed using mutation detection software.
13. The method according to claim 12, wherein, The mutation detection software is selected from Vardict, Varscan, GATK (Genome Analysis Toolkit) or Mutect.
14. The method according to claim 13, wherein, The mutation detection software is Vardict.
15. The method according to claim 14, wherein, The mutation sites obtained in step one are the sites corresponding to any mutation in the test sample or paired sample whose mutation abundance is higher than or equal to the wild-type filtering threshold.
16. The method according to claim 15, wherein, The wild-type filtering threshold is 30%.
17. The method according to claim 16, wherein, The mutation sites screened in step one are the sites corresponding to mutations whose maximum population frequency in different population genomes is greater than or equal to 0.1%.
18. The method according to claim 17, wherein, Population frequencies in a population genome are queried from one or more population genome databases.
19. The method according to claim 18, wherein, The population genome database is selected from 1000 GenomesProject, dbSNP, gnomAD (genome aggregation database), and ExAC (the Exome Aggregation Consortium).
20. The method according to claim 7, wherein, The first preset threshold is 90%-98%.
21. The method according to claim 5 or 7, wherein, The second preset threshold is 65%-90%.
22. The method according to claim 20, wherein, The first preset threshold is 95%.
23. The method according to claim 21, wherein, The second preset threshold is 75%.
24. An apparatus for detecting contamination in a test sample and / or identifying sample mismatch, the apparatus comprising: The screening module is configured to screen for detecting contamination in the sample to be tested and / or identifying mutation sites of sample mismatch; The construction module is configured to construct indicators for determining contamination of the test sample and / or identifying sample mismatch based on the mutation sites of the test sample and the paired sample. The indicators include any one or more of the average values of correlation level, homozygous ratio and sample homozygous mutation abundance. The determination module is configured to identify and determine contamination and / or mismatch of the test sample based on at least one indicator constructed in step two. The homozygosity ratio is the ratio of the number of mutation sites with mutation abundance higher than or equal to a first preset threshold in both the test sample and the paired sample to the number of mutation sites with mutation abundance higher than or equal to the first preset threshold in the paired sample. The formula for calculating the homozygous ratio (homo.ratio) is as follows: Where homo.ratio represents the homozygosity ratio of the sample, N1 represents the number of mutation sites in both the test sample and the paired sample whose mutation abundance is higher than or equal to the first preset threshold, and N2 represents the number of mutation sites in the paired sample whose mutation abundance is higher than or equal to the first preset threshold. When the homo.ratio is below 90%, the sample to be tested is determined to be contaminated; when the homo.ratio is above or equal to 90%, the sample to be tested is determined to be uncontaminated. The average homozygous mutation abundance of the sample (homoAF) is the average mutation abundance of mutations in the test sample whose mutation abundance is higher than or equal to the first preset threshold in the paired samples. If the average abundance of homozygous mutations in a sample is less than 0.975, the sample to be tested is determined to be contaminated.
25. The apparatus according to claim 24, wherein, N2≥100。 26. An apparatus for detecting sample contamination and / or identifying sample mismatch, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-23.
27. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by one or more processors, it implements the method as described in any one of claims 1-23.