Method and device for judging sample mismatch and pollution based on SNP mutation frequency similarity and computer equipment
By calculating the similarity of SNP mutation frequencies in samples and using Pearson correlation coefficient and dual threshold judgment system, the problems of sample mismatch and large-scale contamination in tumor gene detection are solved, thus improving the accuracy and reliability of detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-04-10
AI Technical Summary
Current technologies lack effective means to identify sample mismatches and high levels of contamination in tumor gene testing, leading to inaccurate test results and impacting clinical decision-making.
By calculating the similarity of SNP mutation frequencies among samples in the same batch, the Pearson correlation coefficient is used to determine the matching relationship and contamination status between samples. A dual-threshold judgment system is set up to identify sample mismatch and contamination.
It enables accurate identification of sample mismatches and high-proportion contamination, improves the accuracy and reliability of tumor gene detection, and provides a reliable quality control method.
Smart Images

Figure CN121838869A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of gene detection technology, specifically involving methods, devices, and computer equipment for judging sample mismatch and contamination based on SNP mutation frequency similarity. Background Technology
[0002] Tumor gene testing is an important tool in modern medical diagnosis. By performing gene sequencing analysis on tumor tissue samples, it can provide crucial information for clinical diagnosis and treatment. Each batch of tumor gene testing involves multiple different samples from different patients. Some adjacent normal tissue or leukocyte blood samples (NC samples) are used to filter germline mutations; these are called paired samples, and they come from the same patient. In clinical testing, because blood samples are easy to collect and extract, the vast majority of paired samples are blood samples. In practice, contamination may be introduced during the embedding of tumor tissue in the paraffin block; human error may introduce contamination during sample extraction and library construction; errors in library construction and incorrect completion of the data entry form directly lead to sample mismatches. Sample contamination and sample mismatches have a significant impact on the accuracy of tumor gene testing results.
[0003] In current tumor gene testing, sample contamination and mismatch issues are primarily prevented through rigorous quality control processes. However, effective detection methods are lacking to identify existing contamination and mismatches. Especially in cases of high contamination rates (above 30%), traditional quality control methods struggle to accurately identify these issues, potentially leading to erroneous test results used in clinical decisions and posing potential risks to patients. Furthermore, effective technical means for the rapid and accurate identification of sample mismatches are also lacking.
[0004] Detection of sample mismatch and high-proportion contamination are significant technical challenges in the field of tumor gene testing. Sample mismatch leads to a mismatch between tumor tissue and normal control samples, affecting the accurate filtering of germline mutations; high-proportion contamination distorts test results, failing to accurately reflect the patient's gene mutation status.
[0005] Therefore, there is an urgent need to develop a technical method that can accurately identify sample mismatches and large-scale contamination to further improve the accuracy and reliability of tumor gene detection. Summary of the Invention
[0006] Based on this, one embodiment of this application provides a method, apparatus, and computer equipment for judging sample mismatch and contamination based on SNP mutation frequency similarity.
[0007] This application provides a method for determining sample mismatch and contamination based on SNP mutation frequency similarity, including the following steps:
[0008] Obtain mutation frequency data of SNP sites for all samples in the same batch of samples submitted for testing;
[0009] Based on the mutation frequency data, the mutation frequency similarity between any two samples in the batch is calculated, and the correspondence between any two sample numbers and mutation frequency similarity is obtained. Based on whether the sample numbers are the same and the relationship between the corresponding mutation frequency similarity and the preset similarity abnormality threshold range, it is determined whether there is sample mismatch or contamination in the two samples.
[0010] In some embodiments, all samples in the same batch of samples submitted for testing include:
[0011] Samples from multiple different submitted tests originating from the same individual; a sample from one submitted test and its corresponding paired sample; or samples from multiple different submitted tests and their paired sample corresponding to at least one of the submitted tests; or
[0012] Samples from different individuals for the same or different submitted items, optionally, wherein at least one sample from at least one submitted item of the same individual is matched with a corresponding paired sample;
[0013] Different individuals have different sample numbers, while the same individual has the same sample number for the same submitted item. Paired samples have the same sample number as their corresponding submitted samples. In some embodiments, calculating the mutation frequency similarity between pairs of samples within the submitted batch includes:
[0014] Mutation frequency similarity is calculated by measuring the Pearson correlation coefficient between the mutation frequencies of SNP sites in each pair of samples within the submitted batch.
[0015] In some embodiments, determining whether there is sample mismatch or contamination in a pair of samples based on whether the sample numbers are the same and the relationship between the corresponding mutation frequency similarity and a preset similarity anomaly threshold includes:
[0016] If two samples have the same sample number, and the corresponding mutation frequency similarity is greater than the upper limit of the similarity anomaly threshold range, it is judged as normal; if it is within the similarity anomaly threshold range, it is judged as sample contamination; if it is less than the lower limit of the similarity anomaly threshold range, it is judged as sample mismatch.
[0017] If two samples have different sample numbers, and the corresponding mutation frequency similarity is within the similarity anomaly threshold range and the two samples originate from different individuals, it is judged that there is sample contamination. If it is greater than the upper limit of the similarity anomaly threshold range, it is judged that there is sample mismatch. If it is less than the lower limit of the similarity anomaly threshold range, it is judged as normal.
[0018] If two samples have different sample numbers, and the corresponding mutation frequency similarity is greater than the upper limit of the similarity anomaly threshold range and the two samples originate from the same individual, they are judged to be normal. If the similarity anomaly threshold range is within the range, they are judged to be sample contamination. If the similarity anomaly threshold range is less than the lower limit of the range, they are judged to be sample mismatch.
[0019] In some embodiments, the method for obtaining the similarity anomaly threshold includes:
[0020] Simulate the process of a sample being contaminated by a sample from another different source at different proportions to generate a gradient contamination simulation sample;
[0021] Calculate the pairwise similarity between pollution source samples, contaminated original samples, gradient pollution simulation samples, and control samples;
[0022] The correlation between contamination ratio and similarity was analyzed, and the threshold for distinguishing normal samples from abnormal samples was determined based on the similarity distribution of large-scale retrospective data.
[0023] In some embodiments, the similarity anomaly threshold ranges from 0.58 to 0.90.
[0024] This application also provides an apparatus for determining sample mismatch and contamination based on SNP mutation frequency similarity, the apparatus comprising:
[0025] The data acquisition module is used to acquire mutation frequency data of SNP sites for all samples in the same batch of samples submitted for testing;
[0026] The similarity calculation module calculates the mutation frequency similarity between any two samples in the submitted batch based on the mutation frequency data, and obtains the correspondence between any two sample numbers and mutation frequency similarity.
[0027] The judgment module determines whether there is sample mismatch or contamination between pairs of samples based on whether the sample numbers are the same and the relationship between the corresponding mutation frequency similarity and the preset similarity abnormality threshold range.
[0028] The output module is used to output the judgment results and pollution source information.
[0029] In another aspect, this application provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method.
[0030] In another aspect, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described.
[0031] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method.
[0032] This application provides a method for determining sample mismatch and contamination based on SNP mutation frequency similarity. By utilizing the linear correlation characteristic of SNP mutation frequencies within the same patient, it can accurately identify matching relationships and contamination status between samples. When samples come from the same patient, their SNP mutation frequency distributions show a linear correlation, with a similarity close to 1, while the similarity between samples from different patients is much lower than 1. Furthermore, based on whether each pair of sample numbers is identical and the relationship between the corresponding mutation frequency similarity and a preset abnormal similarity threshold range, it determines whether sample mismatch or contamination exists between the pairs of samples, thus establishing a scientific threshold judgment system. This method can effectively distinguish between normal samples, mismatched samples, and contaminated samples, solving the problem that traditional quality control methods struggle to identify large-scale contamination, providing reliable quality assurance for clinical testing, and offering a reliable technical means for quality control in tumor gene testing. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of this application and to more completely understand this application and its beneficial effects, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 Linear regression fitting diagram of SNP site mutation frequency of paired samples provided in an embodiment of this application;
[0035] Figure 2 The similarity of SNP mutation frequencies between pairs of samples within a batch, calculated for an embodiment of this application;
[0036] Figure 3 This is a schematic diagram illustrating a data simulation method using three samples A, B, and BNC in one embodiment of this application.
[0037] Figure 4 A line graph showing the pollution ratio and similarity of one embodiment of this application is drawn;
[0038] Figure 5 In one embodiment of this application, the distribution statistics of sample similarity values are performed, and the similarity values of samples A and B are plotted in the distribution graph after sample B is contaminated by gradient A.
[0039] Figure 6This is a schematic diagram illustrating an embodiment of the present application regarding the determination of whether three samples match, whether there is a large proportion of contamination, and the source of contamination. Detailed Implementation
[0040] The present application will be further described in detail below with reference to the embodiments and examples. It should be understood that these embodiments and examples are for illustrative purposes only and are not intended to limit the scope of the present application. The purpose of providing these embodiments and examples is to enable a more thorough and comprehensive understanding of the disclosure of the present application. It should also be understood that the present application can be implemented in many different forms and is not limited to the embodiments and examples described herein. Those skilled in the art can make various modifications or alterations without departing from the spirit of the present application, and the equivalent forms obtained also fall within the protection scope of the present application. Furthermore, numerous specific details are set forth in the following description to provide a fuller understanding of the present application. It should be understood that the present application can be implemented without one or more of these details.
[0041] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0042] Unless otherwise stated or in case of contradiction, the terms or phrases used herein shall have the following meanings:
[0043] The terms "and / or," "or / and," and "and / or" as used herein include any one of two or more of the related listed items, as well as any and all combinations of the related listed items. These arbitrary and all combinations include any two related listed items, any more related listed items, or a combination of all related listed items. It should be noted that when at least three items are connected using at least two conjunctions selected from "and / or," "or / and," and "and / or," it should be understood that in this application, the technical solution undoubtedly includes technical solutions connected by "logical AND," and also undoubtedly includes technical solutions connected by "logical OR." For example, "A and / or B" includes three parallel solutions: A, B, and A+B. For example, the technical solution of "A, and / or, B, and / or, C, and / or, D" includes any one of A, B, C, and D (that is, a technical solution that is connected by "logical OR"), as well as any and all combinations of A, B, C, and D, that is, combinations of any two or three of A, B, C, and D, and also combinations of all four of A, B, C, and D (that is, a technical solution that is connected by "logical AND").
[0044] In this application, the terms "multiple", "various", "multiple times", "multi-dimensional", etc., unless otherwise specified, refer to a quantity greater than or equal to 2. For example, "one or more" means one or more than or equal to two.
[0045] In this application, terms such as "further," "even further," and "particularly" are used to describe purposes and indicate differences in content, but should not be construed as limiting the scope of protection of this application.
[0046] In this application, "optionally," "optionally," and "optional" mean that something is optional, that is, it means that it is selected from either "with" or "without." If there are multiple "optional" entries in a technical solution, unless otherwise specified, and there are no contradictions or mutual constraints, each "optional" entry shall be independent.
[0047] In this application, the technical features described in an open-ended manner include both closed technical solutions composed of the listed features and open technical solutions composed of the listed features.
[0048] All references to documents mentioned in this application are incorporated herein by reference as if each document were individually incorporated herein by reference. Unless they conflict with the inventive purpose and / or technical solution of this application, all cited documents are incorporated herein by reference in their entirety and for all purposes. When citing documents in this application, the definitions of relevant technical features, terms, nouns, phrases, etc., are also incorporated herein by reference. When citing documents in this application, examples and preferred embodiments of the cited technical features may also be incorporated herein by reference, but only to the extent that they enable the implementation of this application. It should be understood that when the cited content conflicts with the description in this application, this application shall prevail or modifications shall be made adaptably to the description in this application.
[0049] The term "SNP mutation frequency" refers to the frequency or proportion of mutations occurring at single nucleotide polymorphism (SNP) sites during gene testing. In the field of tumor gene testing, SNP mutation frequency reflects the genotypic status of a sample at that site. In diploid organisms, 0% or 100% typically indicates a homozygous genotype, while ~50% indicates a heterozygous genotype. Its calculation usually relies on data obtained from high-throughput sequencing technology, and the formula is: Mutation frequency = Number of sequencing reads of the mutated base / (Number of reads of the reference base + Number of reads of the mutated base). Specifically, SNP mutation frequency can include different types of mutation frequencies such as point mutation frequency, insertion / deletion mutation frequency, and copy number variation frequency. In clinical applications, SNP mutation frequency can be used for various aspects of tumor diagnosis, prognostic assessment, and treatment selection.
[0050] The term "mutation frequency similarity" refers to the degree of similarity in the distribution of SNP mutation frequencies between two samples. In the field of gene testing quality control, mutation frequency similarity is a key indicator for judging the consistency of sample sources, and it is calculated using statistical methods. Specifically, mutation frequency similarity can include various calculation methods such as Pearson correlation coefficient, Spearman correlation coefficient, cosine similarity, and Euclidean distance similarity. In practical applications, mutation frequency similarity can be used to identify quality problems such as sample mismatch, sample contamination, and sample mixing.
[0051] The term "similarity anomaly threshold" refers to a critical value used to determine whether sample similarity falls within the normal range. In the field of clinical testing quality control, the similarity anomaly threshold is a crucial standard for distinguishing between normal and abnormal samples, determined through statistical analysis and clinical validation. Specifically, similarity anomaly thresholds can include different types such as fixed thresholds, dynamic thresholds, and machine learning thresholds. In practical applications, the similarity anomaly threshold needs to be customized based on factors such as different testing platforms, sample types, and disease types.
[0052] The term "high-proportion contamination" in this field refers to a situation where a sample is contaminated by samples from other sources at a high proportion (typically higher than 30%). In tumor gene testing, high-proportion contamination is a serious quality issue that can lead to severely distorted test results. Specifically, high-proportion contamination can take many forms, including contamination of tumor tissue with normal tissue, cross-contamination between samples from different patients, and contamination of laboratory reagents. Identifying high-proportion contamination is crucial for ensuring the accuracy of test results because traditional quality control methods often struggle to effectively identify this type of contamination.
[0053] In this field, the term "sample mismatch" refers to a situation where a tumor tissue sample does not match its corresponding normal control sample. In tumor gene testing, sample mismatch is a serious operational error that can lead to incorrect germline mutation filtering, affecting the final test results. Specifically, sample mismatch can include various situations such as mismatch between tumor tissue and paired blood samples, sample confusion between different patients, and incorrect sample labeling. Identifying sample mismatch is crucial for ensuring the reliability of test results, as incorrect sample matching can lead to serious biases in clinical decision-making.
[0054] The term "paired sample" in this field refers to a normal tissue or blood sample from the same individual as the tumor tissue sample. In tumor gene testing, paired samples are primarily used to filter germline mutations and improve the accuracy of somatic mutation detection. Specifically, paired samples can include various types such as adjacent normal tissue samples, peripheral blood leukocyte samples, and oral mucosal cell samples. Quality control of paired samples is crucial for ensuring the accuracy of tumor gene testing, as the quality of the paired samples directly affects the effectiveness of germline mutation filtering.
[0055] The term "submission batch" in this field refers to a group of samples tested at the same time and under the same conditions. In clinical genetic testing, the submission batch is the basic unit of quality control and typically contains samples from multiple different patients. Specifically, a submission batch can include multiple samples for the same test, combinations of samples for different tests, sample groups containing paired samples, and so on. Quality control of the submission batch is crucial for ensuring the reliability of the entire testing process because there is a risk of cross-contamination among samples within a batch.
[0056] This application provides a method for determining sample mismatch and contamination based on SNP mutation frequency similarity, including the following steps:
[0057] Obtain the mutation frequency data of SNP sites for all samples in the same batch of samples submitted for testing; compare the high-throughput sequencing data with the reference genome using bioinformatics software, and calculate the mutation frequency of SNP sites. The calculation formula is: mutation frequency = number of reads of mutated bases / (number of reads of reference bases + number of reads of mutated bases).
[0058] Based on the mutation frequency data, the mutation frequency similarity between any two samples in the batch is calculated, and the correspondence between any two sample numbers and mutation frequency similarity is obtained. Based on whether the sample numbers are the same and the relationship between the corresponding mutation frequency similarity and the preset similarity abnormality threshold range, it is determined whether there is sample mismatch or contamination in the two samples.
[0059] In some embodiments, all samples in the same batch of samples submitted for testing include:
[0060] Samples from multiple different submitted tests originating from the same individual; a sample from one submitted test and its corresponding paired sample; or samples from multiple different submitted tests and their paired sample corresponding to at least one of the submitted tests; or
[0061] Samples from different individuals for the same or different submitted items, optionally, wherein at least one sample from at least one submitted item of the same individual is matched with a corresponding paired sample.
[0062] It is understood that the submitted sample can be a single sample or multiple samples. Different individuals have different sample numbers, while the same individual has the same sample number for the same submitted item. Paired samples share the same sample number as their corresponding submitted samples. This number refers to the numerical part preceding the test item identifier such as "T, P, or NC" or the paired sample identifier, excluding the following letters. For example: 454003086200P0. Name A and 454003086200NC. Name A have the same number; 453985161500T0. Name B and 454003086200NC. Name A have different numbers. In some embodiments, calculating the mutation frequency similarity between pairs of samples within the submitted batch includes:
[0063] Mutation frequency similarity is calculated by determining the Pearson correlation coefficient between the mutation frequencies of SNP sites in each pair of samples within the submitted batch. The Pearson correlation coefficient is a statistical indicator that measures the degree of linear correlation between two variables. In bioinformatics, the Pearson correlation coefficient is commonly used to compare the similarity of gene expression profiles or SNP mutation frequency profiles, and its value ranges from -1 to 1. Specifically, the Pearson correlation coefficient can be calculated based on raw data, standardized data, or logarithmically transformed data. In gene testing quality control, a Pearson correlation coefficient close to 1 indicates a high correlation between the two samples, close to 0 indicates no correlation, and close to -1 indicates a negative correlation.
[0064] Understandably, the Pearson correlation coefficient effectively measures the linear correlation between the mutation frequencies of SNP sites between two samples, making it a reliable statistical indicator for judging the consistency of sample origins. Specifically, the formula for calculating the Pearson correlation coefficient is r = cov(X,Y) / (σXσY), where X and Y represent the SNP mutation frequency vectors of the two samples, cov(X,Y) is the covariance, and σX and σY are the standard deviations.
[0065] In some embodiments, determining whether there is sample mismatch or contamination in a pair of samples based on whether the sample numbers are the same and the relationship between the corresponding mutation frequency similarity and a preset similarity anomaly threshold includes:
[0066] If two samples have the same sample number, and the corresponding mutation frequency similarity is greater than the upper limit of the similarity anomaly threshold range, it is judged as normal; if it is within the similarity anomaly threshold range, it is judged as sample contamination; if it is less than the lower limit of the similarity anomaly threshold range, it is judged as sample mismatch.
[0067] If two samples have different sample numbers, and the corresponding mutation frequency similarity is within the similarity anomaly threshold range and the two samples originate from different individuals, then it is judged that there is sample contamination. If the similarity frequency is greater than the upper limit of the similarity anomaly threshold range, it is judged that there is sample mismatch. If the similarity frequency is less than the lower limit of the similarity anomaly threshold range, it is judged as normal.
[0068] If two samples have different sample numbers, and the corresponding mutation frequency similarity is greater than the upper limit of the similarity anomaly threshold range and the two samples originate from the same individual, they are judged to be normal. If the similarity anomaly threshold range is within the range, they are judged to be sample contamination. If the similarity anomaly threshold range is less than the lower limit of the range, they are judged to be sample mismatch.
[0069] In some embodiments, the method for obtaining the similarity anomaly threshold includes:
[0070] Simulate the process of a sample being contaminated by a sample from another different source at different proportions to generate a gradient contamination simulation sample;
[0071] Calculate the pairwise similarity between pollution source samples, contaminated original samples, gradient pollution simulation samples, and control samples;
[0072] This study analyzes the correlation between contamination ratio and similarity, and determines a threshold for distinguishing normal and abnormal samples based on the similarity distribution of large-scale retrospective data. The method combines simulation experiments and big data analysis to ensure the scientific validity and reliability of the threshold.
[0073] In some embodiments, the similarity anomaly threshold ranges from 0.58 to 0.90.
[0074] This application utilizes SNP mutation frequency similarity analysis technology to accurately identify sample mismatches and high-proportion contamination, solving the problem that traditional quality control methods struggle to detect high-proportion contamination.
[0075] This application adopts a dual-threshold judgment system, which can set judgment criteria for the same patient sample and different patient samples respectively, thereby improving the specificity and sensitivity of the detection.
[0076] This application, through its multi-level judgment logic, can comprehensively cover various sample anomalies, including sample mismatch, sample contamination, and high-proportion contamination, providing a reliable quality assurance means for tumor gene detection.
[0077] Based on the same inventive concept, this application also provides an apparatus for determining sample mismatch and contamination based on SNP mutation frequency similarity, the apparatus comprising:
[0078] The data acquisition module is used to acquire mutation frequency data of SNP sites for all samples in the same batch of samples submitted for testing;
[0079] The similarity calculation module calculates the mutation frequency similarity between any two samples in the submitted batch based on the mutation frequency data, and obtains the correspondence between any two sample numbers and mutation frequency similarity.
[0080] The judgment module determines whether there is sample mismatch or contamination between pairs of samples based on whether the sample numbers are the same and the relationship between the corresponding mutation frequency similarity and the preset similarity abnormality threshold range.
[0081] The output module is used to output the judgment results and pollution source information.
[0082] The modules in the aforementioned device for determining sample mismatch and contamination based on SNP mutation frequency similarity can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0083] In an exemplary embodiment, a computer device is provided, which may be a server. The computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is connected to the system bus via the I / O interfaces. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores data for determining sample mismatches and contamination based on SNP mutation frequency similarity. The I / O interfaces of the computer device are used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a method for determining sample mismatches and contamination based on SNP mutation frequency similarity.
[0084] This application, in another aspect, provides a computer device including a memory and a processor. The memory stores a computer program, and the processor, when executing the computer program, implements the steps of the method described. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of this computer device provides computational and control capabilities. The memory of this computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and a computer program.
[0085] Another aspect of this application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, performs the following steps:
[0086] Obtain mutation frequency data of SNP sites for all samples in the same batch of samples submitted for testing;
[0087] Based on the mutation frequency data, the mutation frequency similarity between any two samples in the submitted batch is calculated, and the correspondence between any two sample numbers and mutation frequency similarity is obtained.
[0088] The presence of sample mismatch or contamination in each pair of samples is determined by whether the sample numbers are the same and the relationship between the corresponding mutation frequency similarity and the preset similarity anomaly threshold range.
[0089] Another aspect of this application provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0090] Obtain mutation frequency data of SNP sites for all samples in the same batch of samples submitted for testing;
[0091] Based on the mutation frequency data, the mutation frequency similarity between any two samples in the submitted batch is calculated, and the correspondence between any two sample numbers and mutation frequency similarity is obtained.
[0092] The presence of sample mismatch or contamination in each pair of samples is determined by whether the sample numbers are the same and the relationship between the corresponding mutation frequency similarity and the preset similarity anomaly threshold range.
[0093] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Furthermore, any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory.
[0094] The embodiments of this application will be described in detail below with reference to examples. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of this application. For experimental methods in the following embodiments where specific conditions are not specified, please refer to the guidelines given in this application, or follow experimental manuals or conventional conditions in the art, or follow the conditions recommended by the manufacturer, or refer to experimental methods known in the art.
[0095] In the specific embodiments described below, the measurement parameters involving raw material components may have slight deviations within the weighing accuracy range unless otherwise specified. For temperature and time parameters, acceptable deviations due to instrument testing accuracy or operational precision are permissible.
[0096] It should be understood that, in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0097] Example 1
[0098] I. Technical Principles
[0099] The SNP mutation frequencies of the same patient showed a linear correlation in distribution. The linear regression fit plot of a paired sample is shown below. Figure 1 As shown, the similarity between two linearly correlated data can be calculated using the Pearson correlation coefficient. If the samples are paired from the same patient, the calculated sample similarity will be close to 1.
[0100] This application calculated the similarity of SNP mutation frequencies between pairs of samples within a batch, as follows: Figure 2 As shown: each sample has a similarity of 1 with its own sample; paired samples from the same patient, i.e., tumor tissue and adjacent normal control samples, also have a similarity close to 1. However, the similarity of samples from different patients is much lower than 1.
[0101] To explore the correlation between sample similarity and sample contamination ratio, and to define the threshold for abnormal similarity, this application used three samples A, B, and BNC for data simulation.
[0102] The simulation method is illustrated below. Figure 3 As shown, samples A and B are from different patients, while samples BNC and B are from different samples of the same patient, with BNC serving as a normal control. Sample B was subjected to gradient contamination using samples A at different proportions. This application does not consider the sample similarity between samples with varying degrees of contamination; it only focuses on the relationship between contaminated and uncontaminated samples. The pairwise sample similarities are shown in Table 1 below:
[0103] Table 1
[0104]
[0105] Line graphs showing the contamination rate and similarity of the data in the table are shown below. Figure 4 As shown.
[0106] As can be seen from the above figure:
[0107] (1) For different patients, the higher the proportion of contamination in the contaminated sample, the higher the similarity to the sample from the source of contamination.
[0108] (2) For the same sample, the greater the proportion of contamination, the lower the similarity.
[0109] (3) For different samples from the same patient, the greater the proportion of contaminated tissue samples, the lower the similarity.
[0110] II. Obtain the preset similarity anomaly threshold range
[0111] This application retrospectively analyzed the sample similarity of 5332 samples and 157 samples from different batches, generating a total of 110511 similarity values. These similarity values were statistically analyzed for distribution, and the similarity values between samples A and B after sample B was contaminated by gradient A were plotted in the distribution graph below. Figure 5 As shown.
[0112] from Figure 5 As can be seen, when the pollution rate is between 30% and 70%, the corresponding similarity value falls between 0.58 and 0.90, which is in the extremely low frequency range of the similarity value distribution.
[0113] Therefore, this application can conclude that:
[0114] (1) When samples come from the same patient, the sample similarity should be >0.90;
[0115] (2) When the samples come from different patients, the sample similarity should be <0.58.
[0116] Therefore, the following similarity anomaly screening criteria are derived:
[0117] (1) Samples from the same patient with a similarity of <0.90;
[0118] (2) Samples from different patients with a similarity >0.58.
[0119] III. On-machine verification
[0120] Based on the above similarity anomaly criteria, the sample similarity of the 157 batches of samples was screened, and the screening results are shown in Table 2 below:
[0121] Table 2
[0122]
[0123] Example 2
[0124] This embodiment uses 211 SNP sites selected by the detection center as examples for analysis.
[0125] 1. The mutation frequency of SNP sites in each sample of a batch of samples was detected and analyzed, as shown in Table 3 below (mutation frequency is in %):
[0126] Table 3
[0127]
[0128] 2. Calculate the sample similarity between pairs of samples within a batch, as shown in Table 4 below:
[0129] Table 4
[0130]
[0131] 3. Determine if the samples match, and whether there is a large proportion of contamination and its source:
[0132] The sample similarity table was filtered according to the following conditions, and the results are shown in Table 5.
[0133] (1) Samples from the same patient with a similarity of <0.90;
[0134] (2) Samples from different patients with a similarity >0.58.
[0135] Table 5
[0136]
[0137] As can be seen from the table above:
[0138] 1. The sample similarity between 453668377800T4. Zhang*zhong and 511313808900T4. Dai*qiu is as high as 0.999, which means that the SNP site similarity is extremely high, and it can be determined that the samples come from the same patient.
[0139] 2. The sample similarity between 511313809000T4.Li*Huai and 511313809000NC.Li*Huai is 0.999, meaning that all paired samples of Li*Huai come from the same patient. As mentioned earlier in this article, the probability of contamination in NC samples is very low; therefore, it is inferred that 511313809000T4.Li*Huai was not contaminated.
[0140] 3. The sample similarity between 453668377800T4.Zhang*zhong and 511313809000T4.Li*huai is 0.863, while the sample similarity between 511313808900T4.Dai*qiu and 511313809000T4.Li*huai is 0.862. This means that there is a large proportion of contamination between 453668377800T4.Zhang*zhong and 511313808900T4.Dai*qiu and 511313809000T4.Li*huai, with the contamination rate ranging from 60% to 70%.
[0141] 4. From ② and ③, it can be deduced that 453668377800T4. Zhang*zhong and 511313808900T4. Dai*qiu were both heavily polluted by 511313809000T4. Li*huai.
[0142] The diagram is as follows: Figure 6 As shown, the conclusion is as follows: 453668377800T4. Zhang*zhong and 511313808900T4. Dai*qiu are samples from the same patient; this sample was highly contaminated by 511313809000T4. Li*huai, with a contamination rate between 60% and 70%.
[0143] The embodiments described above are merely illustrative of several implementation methods of this application, intended to facilitate a detailed understanding of the technical solutions of this application, but should not be construed as limiting the scope of protection of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the scope of protection of this application. Furthermore, it should be understood that after reading the above teachings of this application, those skilled in the art can make various alterations or modifications to this application, and the equivalent forms obtained also fall within the scope of protection of this application. It should also be understood that technical solutions obtained by those skilled in the art based on the technical solutions provided in this application through logical analysis, reasoning, or limited experimentation are all within the scope of protection of the appended claims. Therefore, the scope of protection of this patent application should be determined by the content of the appended claims, and the specification can be used to interpret the content of the claims.
Claims
1. A method for determining sample mismatch and contamination based on SNP mutation frequency similarity, characterized in that, The method comprises the following steps: obtaining mutation frequency data of SNP sites of all samples in the same submission batch; based on the mutation frequency data, calculating the mutation frequency similarity between each two samples in the submission batch, and obtaining the corresponding relationship between any two sample numbers and the mutation frequency similarity; determining whether sample mismatch or contamination exists in the two samples according to whether the two sample numbers are the same and the size relationship between the corresponding mutation frequency similarity and the preset similarity abnormal threshold range. 2.The method for determining sample mismatch and contamination based on SNP mutation frequency similarity according to claim 1, characterized in that, All samples in the same submission batch include: samples of different submission items from the same individual, samples of one submission item and corresponding paired samples, or samples of different submission items and corresponding paired samples of at least one of the submission items; and / or samples of the same or different submission items from different individuals, and optionally, at least one sample of at least one submission item of at least one individual is matched with the corresponding paired sample; The sample numbers of different individuals are different, the sample numbers of the same submission item of the same individual are the same, and the sample numbers of the paired samples and the corresponding submission samples are the same. 3.The method for determining sample mismatch and contamination based on SNP mutation frequency similarity according to claim 2, characterized in that, The calculation of the mutation frequency similarity between each two samples in the submission batch comprises: calculating the mutation frequency similarity by calculating the Pearson correlation coefficient of the SNP site mutation frequency between each two samples in the submission batch. 4.The method for determining sample mismatch and contamination based on SNP mutation frequency similarity according to claim 2, characterized in that, The determination of whether sample mismatch or contamination exists in the two samples according to whether the two sample numbers are the same and the size relationship between the corresponding mutation frequency similarity and the preset similarity abnormal threshold range comprises: if the sample numbers of the two samples are the same, and the corresponding mutation frequency similarity is greater than the upper limit value of the similarity abnormal threshold range, it is determined to be normal, if it is within the similarity abnormal threshold range, it is determined to exist sample contamination, and if it is less than the lower limit value of the similarity abnormal threshold range, it is determined to exist sample mismatch; if the sample numbers of the two samples are different, and the corresponding mutation frequency similarity is within the similarity abnormal threshold range and the two samples are from different individuals, it is determined to exist sample contamination, if it is greater than the upper limit value of the similarity abnormal threshold range, it is determined to exist sample mismatch, and if it is less than the lower limit value of the similarity abnormal threshold range, it is determined to be normal; if the sample numbers of the two samples are different, and the corresponding mutation frequency similarity is greater than the upper limit value of the similarity abnormal threshold range and the two samples are from the same individual, it is determined to be normal, if it is within the similarity abnormal threshold range, it is determined to exist sample contamination, and if it is less than the lower limit value of the similarity abnormal threshold range, it is determined to exist sample mismatch.
5. The method for determining sample mismatch and contamination based on SNP mutation frequency similarity according to any one of claims 1 to 4, characterized in that, The method for obtaining the similarity abnormal threshold range comprises: simulate a process in which one sample is contaminated by another sample from a different source at different proportions, and generate gradient contamination simulation samples; calculate the similarity between each two of the contaminated source sample, the original contaminated sample, the gradient contamination simulation sample, and the control sample; analyze the corresponding relationship between the contamination proportion and the similarity, and determine the threshold for distinguishing normal samples from abnormal samples according to the similarity distribution of large-scale retrospective data.
6. The method for determining sample mismatch and contamination based on SNP mutation frequency similarity according to claim 5, characterized in that, The similarity abnormal threshold range is 0.58-0.
90.
7. A device for determining sample mismatch and contamination based on SNP mutation frequency similarity, characterized in that, The device comprises: The data acquisition module is configured to acquire mutation frequency data of SNP sites of all samples in the same submission batch. The similarity calculation module is configured to calculate mutation frequency similarities between two samples in the submission batch based on the mutation frequency data, and acquire a corresponding relationship between any two sample numbers and the mutation frequency similarities. The judgment module is configured to determine whether sample mismatch or contamination exists in the two samples according to whether the two sample numbers are the same and a size relationship between the corresponding mutation frequency similarity and a preset similarity abnormal threshold range. The output module is configured to output a judgment result and contamination source information.
8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6.
Citation Information
Patent Citations
Molecular quality assurance methods for use in sequencing
CN108137642A
Cyclic RNA compositions and methods
CN116322788A
Genomic DNA mutation assays and uses thereof
US20180298430A1
Methods for fingerprinting of biological samples
US20210151126A1
Systems and methods for assessing similarity between samples using genotype signatures
WO2025090739A1