Sample contamination detection methods, apparatuses, systems, and related devices
By calculating the mutation genotype frequency value (BAF) of a sample at a preset site, the proportion of homozygous sites is determined, solving the problem of sample contamination judgment, providing a basis for the accuracy of nucleic acid test results, avoiding missed detection of low-contamination samples, and improving the accuracy of detection.
Patent Information
- Application Number
- CN202210390216.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-14
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2042-04-14
AI Technical Summary
The risk of sample contamination during sample preparation can lead to erroneous nucleic acid test results. Current technology makes it difficult to accurately determine whether a sample is contaminated, thus affecting the accuracy of the test results.
By calculating the mutation genotype frequency value (BAF) of the sample at a preset site, the proportion of homozygous sites is determined. If the proportion of homozygous sites is less than the first proportion, the sample is judged to be a contaminated sample, and the contamination value is further calculated to determine the degree of contamination.
It enables a simple and direct way to determine whether a sample is contaminated, provides a basis for the accuracy of nucleic acid test results, avoids missed detection of low-contamination samples, and improves the accuracy of test results.
Smart Images

Figure CN116959564B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of chip technology, specifically to a sample contamination detection method, apparatus, system and related equipment. Background Technology
[0002] Nucleic acid testing is a viral nucleic acid detection technology. It determines whether a sample is infected by a virus by identifying whether the nucleic acid of an invading virus is present in the sample. When the test result is "positive" for nucleic acid, it is considered that the virus is present in the sample.
[0003] However, samples are at risk of contamination during preparation, such as the introduction of DNA from other samples. If contaminated samples are not detected promptly, they can interfere with nucleic acid testing results and even cause errors, such as false positives.
[0004] Therefore, how to determine whether a sample is contaminated, and thus provide a basis for the accuracy of the nucleic acid test results, has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a sample contamination detection method, apparatus, system and related equipment, which can simply and directly determine whether a sample is contaminated, thereby providing a basis for the accuracy of nucleic acid detection results of the sample.
[0006] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions.
[0007] In a first aspect, embodiments of the present invention provide a sample contamination detection method, comprising:
[0008] Obtain nucleic acid information from the sample;
[0009] Calculate the mutation genotype frequency value (BAF) of the sample at a preset site; wherein, there are multiple preset sites, and in uncontaminated samples, the proportion of homozygous sites among the preset sites is greater than a first proportion;
[0010] Based on the BAF of the preset site, determine the homozygous sites among the preset sites;
[0011] Calculate the proportion of homozygous loci among preset loci. If the proportion of homozygous loci is less than the first proportion, then the sample is a contaminated sample.
[0012] Optionally, if the sample is a contaminated sample, the method further includes: calculating the contamination value of the contaminated sample.
[0013] Optionally, calculating the contamination value of the contaminated sample includes:
[0014] Based on the BAF of the preset sites, the preset sites are sorted to obtain a site sequence;
[0015] Remove homozygous sites from the site sequence to obtain the first site sequence;
[0016] Using the site in the preset position of the first site sequence as the marker point, the contamination value of the sample is calculated based on the BAF of the marker point;
[0017] The preset position is the position obtained by extending the sequence endpoint of the site by a first preset number of positions; the first preset number of positions is calculated based on the proportion of heterozygous genotype contamination homozygous sites in the preset site.
[0018] Optionally, determining the homozygous sites among the preset sites based on the BAF of the preset sites includes:
[0019] Obtain the first threshold;
[0020] When the BAF at the preset site is less than the first threshold, the preset site is considered a homozygous site.
[0021] Optionally, if the proportion of homozygous sites is greater than or equal to the first proportion, then the sample is a sample to be determined; the method further includes:
[0022] If the sample is a pending sample, determine whether the pending sample is a low-contamination sample.
[0023] Optionally, determining whether the sample to be determined is a low-contamination sample includes:
[0024] Based on the BAF of the preset sites, the preset sites are sorted to obtain a site sequence;
[0025] Remove sites in the site sequence whose BAF is greater than or equal to a preset value to obtain a second site sequence; wherein the preset value is greater than the first threshold and less than or equal to 5 times the first threshold;
[0026] Using the site in the second site sequence as the site to be processed, obtain the BAF of the neighboring bases of the site to be processed; wherein, the neighboring bases are a second preset number of bases adjacent to the site to be processed;
[0027] Calculate the significance of the difference between the sites to be treated and their neighboring bases. If the difference parameter between the sites to be treated and their neighboring bases is within a preset range, then the sample to be determined is a low-contamination sample.
[0028] Optionally, if the sample to be determined is a low-contamination sample, the method further includes: calculating the contamination value of the low-contamination sample;
[0029] The calculation of the contamination value of the low-contamination sample includes:
[0030] The third preset number of sites with the highest BAF in the second site sequence are taken as effective sites. The median value of the BAF of the effective sites is determined, and the median value is used as the contamination value.
[0031] The third preset number can be calculated based on the proportion of homozygous contamination sites to the total number of homozygous sites.
[0032] Optionally, in the step of calculating the mutation genotype frequency value BAF of the sample at the preset site, the BAF of the adjacent bases at the preset site is also calculated;
[0033] The step of obtaining the BAF of the neighboring bases of the site to be processed specifically involves obtaining the BAF of the neighboring bases of the site to be processed in the preset site.
[0034] Optionally, the preset site is selected based on one or more factors of the site; the factors for selecting the preset site include: heterozygosity of the site, BAF of the site, BAF fluctuation range of the site, noise level of the site, capture performance of the site and distance between adjacent sites.
[0035] Optionally, obtaining the nucleic acid information of the sample includes:
[0036] Obtain test data of the sample on the test equipment;
[0037] The test data was split and processed to obtain nucleic acid data;
[0038] The nucleic acid data is sorted and deduplicated to obtain the nucleic acid information.
[0039] Secondly, embodiments of the present invention provide a sample contamination detection device, comprising:
[0040] The information acquisition module is used to acquire the nucleic acid information of the sample;
[0041] The first calculation module is used to calculate the mutation genotype frequency value (BAF) of the sample at a preset site; wherein, there are multiple preset sites, and in the uncontaminated sample, the proportion of homozygous sites among the preset sites is greater than a first proportion;
[0042] A homozygous site determination module is used to determine homozygous sites among the preset sites based on the BAF of the preset sites;
[0043] The second calculation module is used to calculate the proportion of homozygous loci in the preset loci. If the proportion of homozygous loci is less than the first proportion, then the sample is a contaminated sample.
[0044] Thirdly, embodiments of the present invention provide a sample contamination detection system, which is configured to perform the sample contamination detection method provided in the embodiments of the present invention.
[0045] Fourthly, embodiments of the present invention provide a computer device, including at least one memory and at least one processor; the memory stores one or more computer-executable instructions, and the processor invokes the one or more computer-executable instructions to execute the sample contamination detection method provided in the embodiments of the present invention.
[0046] Fifthly, embodiments of the present invention provide a storage medium storing one or more computer-executable instructions, which are used to execute the sample contamination detection method provided in embodiments of the present invention.
[0047] The present invention provides a sample contamination detection method, apparatus, system, and related equipment. The method includes: acquiring nucleic acid information of a sample; calculating the mutation genotype frequency value (BAF) of the sample at a preset site; wherein there are multiple preset sites, and in uncontaminated samples, the proportion of homozygous sites among the preset sites is greater than a first proportion; determining homozygous sites among the preset sites based on the BAF of the preset sites; calculating the proportion of homozygous sites among the preset sites; if the proportion of homozygous sites is less than the first proportion, then the sample is a contaminated sample.
[0048] As can be seen, the embodiments of the present invention determine whether the sample is a contaminated sample by calculating the proportion of homozygous sites in the preset sites, thereby making a simple and direct judgment on whether the sample is contaminated, and thus providing a basis for the accuracy of the nucleic acid detection results of the sample.
[0049] In a preferred embodiment, the present invention further calculates the contamination value of the contaminated sample when the sample is a contaminated sample, thereby accurately determining the degree of contamination of the sample and providing a more accurate basis for the accuracy of the nucleic acid detection results of the sample.
[0050] Furthermore, in a further preferred embodiment, the present invention further performs detection and calculation on low-contamination samples when the sample is a low-contamination sample, thereby avoiding missed detection due to low sample contamination level, and can also accurately determine the contamination level of low-contamination samples, further providing a more accurate basis for the accuracy of nucleic acid detection results of the test samples. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0052] Figure 1 A statistical diagram illustrating the proportion of homozygous loci in uncontaminated and contaminated samples to the total number of loci.
[0053] Figure 2 This is an optional flowchart of a sample contamination detection method provided in an embodiment of the present invention;
[0054] Figure 3 An optional flowchart of step S10 provided in an embodiment of the present invention;
[0055] Figure 4 An optional flowchart for calculating the contamination value of a contaminated sample, provided for embodiments of the present invention;
[0056] Figure 5 A schematic diagram illustrating the statistical distribution of the proportion of homozygous sites contaminated by heterozygous genotypes, provided in an embodiment of the present invention.
[0057] Figure 6 An optional flowchart for calculating the contamination value of a sample to be determined is provided in an embodiment of the present invention;
[0058] Figure 7 A schematic diagram illustrating the statistical data of the proportion of homozygous contamination sites to the total number of homozygous sites provided in the embodiments of the present invention;
[0059] Figure 8 This is a schematic diagram of the pollution value results obtained by different algorithms provided in the embodiments of the present invention;
[0060] Figure 9 This is a schematic diagram of BAF statistical data on the relative positions of preset sites provided in an embodiment of the present invention;
[0061] Figure 10 This is an optional block diagram of the sample contamination detection device provided in an embodiment of the present invention;
[0062] Figure 11 Another optional block diagram of the sample contamination detection device provided in the embodiments of the present invention. Detailed Implementation
[0063] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0064] As described in the background section, test samples are at risk of contamination during the preparation process. For example, paraffin section samples are a type of sample with a very complex preparation process, and there is a risk of introducing contamination at any stage of the preparation process.
[0065] It is understandable that the nucleic acid information of a sample includes a large number of SNP (single nucleotide polymorphism) sites. These sites can be homozygous or heterozygous.
[0066] The inventors believe that, for a specific species, the ratio of homozygous to heterozygous sites is limited to a certain range, and when a sample is contaminated, the number of homozygous sites will decrease to some extent. (Refer to reference...) Figure 1 The diagram illustrates the statistical data of the proportion of homozygous loci in uncontaminated and contaminated samples relative to the total number of loci. The horizontal axis represents the proportion of homozygous loci in the total number of loci, and the vertical axis represents the probability density of the statistical samples. The right peak represents the proportion of homozygous loci in uncontaminated samples, and the left peak represents the proportion of homozygous loci in contaminated samples. A clear, flat region can be observed between the two peaks, effectively distinguishing between contaminated and uncontaminated samples. Therefore, the inventors believe that the proportion of homozygous loci can be used to determine whether a sample is contaminated.
[0067] Based on this, the present invention provides a sample contamination detection device, system, and related equipment. The method includes: acquiring nucleic acid information of a sample; calculating the mutation genotype frequency value (BAF) of the sample at a preset site; wherein, there are multiple preset sites, and in uncontaminated samples, the proportion of homozygous sites among the preset sites is greater than a first proportion; determining homozygous sites among the preset sites based on the BAF of the preset sites; calculating the proportion of homozygous sites among the preset sites; if the proportion of homozygous sites is less than the first proportion, then the sample is a contaminated sample.
[0068] As can be seen, the embodiments of the present invention determine whether the sample is a contaminated sample by calculating the proportion of homozygous sites in the preset sites, thereby making a simple and direct judgment on whether the sample is contaminated, and thus providing a basis for the accuracy of the nucleic acid detection results of the sample.
[0069] In a preferred embodiment, the present invention further calculates the contamination value of the contaminated sample when the sample is a contaminated sample, thereby accurately determining the degree of contamination of the sample and providing a more accurate basis for the accuracy of the nucleic acid detection results of the sample.
[0070] Furthermore, in a further preferred embodiment, the present invention further performs detection and calculation on low-contamination samples when the sample is a low-contamination sample, thereby avoiding missed detection due to low sample contamination level, and can also accurately determine the contamination level of low-contamination samples, further providing a more accurate basis for the accuracy of nucleic acid detection results of the test samples.
[0071] In one alternative implementation, Figure 2 An exemplary flowchart of an optional sample contamination detection method provided by an embodiment of the present invention is shown. Figure 2 As shown, the method includes:
[0072] Step S10: Obtain the nucleic acid information of the sample;
[0073] In this embodiment of the invention, the sample is a test sample to be tested. By determining whether the sample is contaminated, the accuracy of the nucleic acid test result of the sample can be determined.
[0074] The nucleic acid information of the sample can be test data describing the nucleic acid information of the sample obtained after the sample is tested by the testing device, or it can be processed data obtained after the test data has been processed.
[0075] Step S11: Calculate the mutation genotype frequency value (BAF) of the sample at a preset site;
[0076] The preset sites are pre-defined sites that meet the requirements for judging sample contamination. There can be multiple preset sites, set according to the actual patterns of the sites. For example, the preset sites can be SNP sites, and the number of preset sites can range from 100 to 500. In one optional example of this invention, the number of preset sites is 139. In other examples, the number of preset sites can also be 200, 400, etc. Based on the concept of this invention, the number of preset sites can be determined according to actual needs, which will not be elaborated further here.
[0077] The preset site can be selected based on one or more factors; specifically, the factors for selecting the preset site may include at least: the heterozygosity of the site, the BAF of the site, the BAF fluctuation range of the site, the noise level of the site, the capture performance of the site, and the distance between adjacent sites.
[0078] The mutant genotype frequency value (BAF) is used to characterize whether the preset locus is homozygous or heterozygous. By calculating the mutant genotype frequency value (BAF), it can be determined whether the preset locus is homozygous.
[0079] In a specific example, specialized computing software, such as Sambamba, can be used to perform deep calculations based on SNP locus information to calculate the BAF value of a preset locus.
[0080] In this invention, the proportion of homozygous sites at preset sites in an uncontaminated sample can be predetermined, thereby allowing the determination of whether a sample is contaminated based on this proportion range. In this embodiment, the proportion of homozygous sites at the preset sites in an uncontaminated sample is greater than a first proportion. This first proportion can be the minimum value of the proportion range; therefore, when the number of homozygous sites is less than this first proportion, the sample is considered contaminated.
[0081] Reference Figure 1 Taking the statistical data shown as an example, it can be seen that the proportion of homozygous sites in uncontaminated samples ranges from approximately 0.36 to 0.66. Accordingly, 0.36 can be used as the first proportion to distinguish between contaminated and uncontaminated samples.
[0082] In other examples, a first proportion can be determined based on the range of homozygous sites at preset sites in contaminated samples and the range of homozygous sites at preset sites in uncontaminated samples. Specifically, the first proportion can be any value between the range of homozygous sites in contaminated samples and the range of homozygous sites in uncontaminated samples. For more details, please refer to [link / reference]. Figure 1 The statistical data shown indicates that the proportion of homozygous loci in contaminated samples ranges from 0.03 to 0.28, including endpoints, while the proportion of homozygous loci in uncontaminated samples ranges from 0.36 to 0.66, including endpoints. Accordingly, the first proportion can be any value between 0.28 and 0.36 (including endpoints). It is understood that the closer the value of the first proportion is to the range of homozygous loci in uncontaminated samples, the more accurate the method.
[0083] It should be noted that, due to the unpredictability of samples, even if the preset sites are determined based on various factors, there may be situations where the BAF calculation requirements cannot be met during sample testing. For example, the depth of the preset sites may not meet the capture performance requirements. In this case, the preset sites can be further understood as sites that meet the BAF calculation requirements.
[0084] Continue to refer to Figure 2 Step S12: Determine the homozygous sites in the preset sites based on the BAF of the preset sites;
[0085] After determining the BAF of the preset site, it can be used to determine whether it is a homozygous site.
[0086] In an optional example, a fixed BAF (Body Accumulation Factor) can be used as a first threshold to distinguish between homozygous and heterozygous sites. This first threshold can be obtained, and if the BAF of a preset site is less than the first threshold, the preset site is considered homozygous. Specifically, if the BAF of a preset site is greater than or equal to the first threshold, the preset site is considered heterozygous; if the BAF is less than the first threshold, the preset site is considered homozygous. The first threshold can be determined based on statistical data from the corresponding samples. In a specific example, 2% can be used as the first threshold to distinguish between homozygous and heterozygous sites.
[0087] Continue to refer to Figure 2 Step S13: Calculate the proportion of homozygous sites in the preset sites. If the proportion of homozygous sites is less than the first proportion, then the sample is a contaminated sample.
[0088] The proportion of homozygous loci among the preset loci can be specifically the quotient of the number of homozygous loci to the number of preset loci. It is understood that the number of homozygous loci will decrease to some extent when the sample is contaminated. Therefore, after calculating the proportion of homozygous loci among the preset loci, it is determined whether this proportion is less than a first proportion. If so, the sample is a contaminated sample; otherwise, the sample is a pending sample.
[0089] As can be seen, the embodiments of the present invention determine whether the sample is a contaminated sample by calculating the proportion of homozygous sites in the preset sites, thereby simply and directly realizing the determination of whether the sample is contaminated, and thus providing a basis for the accuracy of the nucleic acid detection results of the sample.
[0090] In a further implementation of this invention, the nucleic acid information can be obtained after testing with a testing device. Specifically, refer to... Figure 3 The optional flowchart of step S10 shown may include:
[0091] Step S101: Obtain test data of the sample on the test equipment;
[0092] This involves obtaining test data from samples on testing equipment to extract the test data.
[0093] Step S102: The test data is split and reposted to obtain nucleic acid data;
[0094] By splitting and commenting on the test data, the nucleic acid information contained in the test data is extracted as nucleic acid data.
[0095] Step S103: Sort and remove duplicates from the nucleic acid data to obtain the nucleic acid information;
[0096] It is understandable that nucleic acid data may have issues with disorder and duplication. By sorting and deduplicating the nucleic acid data, concise nucleic acid information with a preset order can be obtained.
[0097] In an optional example, the test data of the sample can be test data from a nucleic acid sequencer (e.g., an Illumina sequencer). The process of splitting and re-encoding the data, and further sorting and deduplicating the data, can be as follows: using specialized software (e.g., bcl2fastq) to split the data from the nucleic acid sequencer to generate a test file (e.g., a fastq file) containing the test data; further using specialized software (e.g., bwa) to re-encode the test data in the test file to generate nucleic acid data (e.g., BAM file data); and further sorting and deduplicating the nucleic acid data to obtain the nucleic acid information of the sample.
[0098] In a further optional implementation of the present invention, the preset site can be selected based on one or more factors; wherein, the factors for selecting the preset site may include: heterozygosity of the site, BAF of the site, BAF fluctuation range of the site, noise level of the site, capture performance of the site, and spacing between adjacent sites, etc. The relevant factors can be constrained based on the actual condition of the sample, thereby selecting preset sites that match the sample.
[0099] In an optional example, based on the factors described above, the selection rules for the preset site may include one or more of the following rules:
[0100] The site heterozygosity rate is within a first preset range;
[0101] The BAF of heterozygous sites is within a second preset range;
[0102] The BAF fluctuation range of the site is within the third preset range;
[0103] The noise level at the site is within the fourth preset range;
[0104] The site capture performance meets the preset requirements;
[0105] The distance between adjacent sites is greater than the preset distance.
[0106] The heterozygosity rate describes the proportion of heterozygous loci among the selected loci, and the first preset range can be determined based on the actual situation of the sample. It is understood that the higher the heterozygosity rate of an SNP locus, the more likely it is to carry different genotypes in the population, and the better this locus can distinguish between two samples, which is more conducive to detecting sample contamination. In an optional example, the first preset range can be [0.3, 0.7].
[0107] The BAF of heterozygous sites is within a second preset range to constrain the degree of deviation of the BAF of heterozygous sites, thereby adapting to the preferred range of the testing equipment, such as the site's preference in probe capture, while avoiding excessive deviation of BAF from affecting the subsequent calculation of contamination values. In an optional example, the preferred range of BAF for heterozygous sites is [0.4, 0.55].
[0108] The BAF fluctuation range of a locus is within a third preset range to avoid CNV (Copy number variations), which could affect the nucleic acid detection results of the sample or the subsequent calculation of contamination values. In an optional example, the BAF fluctuation range of a locus can be constrained based on NIQR (Normalized Interquartile Range), specifically, the NIQR interval of BAF can be [0, 0.2].
[0109] The noise level at the site is within a fourth preset range, used to constrain the noise level of the site and prevent excessive noise from affecting the lower limit of contamination detection. In an optional example, the noise level of the site can be constrained based on the noise level in tissue and / or the noise level in blood. Specifically, the noise level in the tissue sample can be less than 2%, and / or the noise level in the blood sample can be less than 0.2%.
[0110] The site capture performance meets preset requirements to ensure easy site capture, thereby ensuring site noise stability based on sufficient site depth. In an optional example, the site capture performance can be constrained based on sample depth; specifically, the compliance rate for depths greater than 100X in tissue samples can be greater than 90%, and / or the compliance rate for depths greater than 1000X in blood samples can be greater than 90%.
[0111] The spacing between adjacent sites is greater than a preset distance to ensure the independence between sites and prevent site linkage. In an optional example, the spacing between adjacent sites can be greater than 1 million base pairs (1M).
[0112] In an optional example, the preset site can simultaneously meet the requirements of the above factors, thereby ensuring the detection accuracy of the embodiments of the present invention.
[0113] In a further implementation of the present invention, when the sample is a contaminated sample, the contamination value of the contaminated sample is further calculated, thereby providing a more accurate basis for the accuracy of the nucleic acid detection results of the sample.
[0114] Specifically, the contamination value of a sample can be calculated based on the BAF (Bipolar Anomaly Frame) of heterozygous genotype contamination homozygous sites. Based on this, embodiments of the present invention further provide an optional procedure for calculating the contamination value of the contaminated sample. This procedure can be used for highly contaminated samples, referring to... Figure 4 The illustrated optional flowchart for calculating the contamination value of the contaminated sample includes:
[0115] Step S20: Sort the preset sites according to their BAF values to obtain a site sequence;
[0116] The sorting of the preset sites can be done by arranging them in ascending or descending order based on their BAF (Balance of Faith). Sorting the preset sites is used to distinguish their types based on the range of BAF values. In this example, ascending order is used for illustration.
[0117] Step S21: Remove homozygous sites from the site sequence to obtain the first site sequence;
[0118] Wherein, if the BAF of a homozygous site is less than a first threshold, the preset site is considered a homozygous site. Accordingly, removing the homozygous site can be understood as removing several preset sites at the beginning of the site sequence until the BAF of the preset sites is greater than or equal to the first threshold. The site sequence composed of the remaining preset sites is the first site sequence.
[0119] Step S22: Using the site in the preset position of the first site sequence as the marker point, calculate the contamination value of the sample according to the BAF of the marker point;
[0120] The preset position is the position obtained by extending the first preset number of positions sequentially from the endpoint of the first point sequence. The first preset number is calculated based on the proportion of heterozygous genotype contamination homozygous sites in the preset site. The proportion of heterozygous genotype contamination homozygous sites in the preset site can be obtained based on statistical data. In an optional example, this proportion can be 13%.
[0121] Heterozygous genotype contamination sites, also known as homozygous sites, can be understood as sites that transform from homozygous sites to heterozygous sites after contamination. Therefore, the contamination value of a sample can be calculated based on the BAF value of these contaminant sites. It's important to note that even if a heterozygous genotype contamination site appears heterozygous, its BAF value will still be smaller than that of the original heterozygous site (i.e., the original heterozygous site in the uncontaminated sample).
[0122] The marker can be a heterozygous contamination site of a preset site, and the marker can be selected based on the order of the first site sequence, choosing a site with a large BAF.
[0123] The proportion of heterozygous genotypes contaminating homozygous sites can be obtained based on statistical data. In a specific example of this invention, refer to... Figure 5 The diagram illustrates the statistical distribution of the proportion of homozygous sites contaminated with heterozygous genotypes. The horizontal axis represents the proportion of homozygous sites contaminated with heterozygous genotypes out of the preset sites, and the vertical axis represents the probability density. It can be seen that the proportion of homozygous sites contaminated with heterozygous genotypes out of the preset sites is greater than or equal to 13% (i.e., the position marked by the dashed line in the diagram, with a value of approximately 0.13). This means that in the first site sequence after removing homozygous sites, at least the first 13% of the sites are homozygous sites contaminated with heterozygous genotypes.
[0124] Accordingly, the first preset number can be the product of the total number of preset sites and the proportion of heterozygous genotype contamination homozygous sites, and the preset position can be the position obtained by extending the first preset number of sites from the first site sequence starting point, thereby determining the marker position.
[0125] In one example, the contamination value of the sample can be twice the BAF of the marked location, i.e., BAF*2.
[0126] Furthermore, in embodiments of the present invention, when a sample is contaminated, the contamination value of the contaminated sample is calculated, thereby accurately determining the degree of contamination of the sample and providing a more accurate basis for the accuracy of nucleic acid detection results of the sample.
[0127] It is understood that the degree of contamination in contaminated samples varies, and at lower levels of contamination, false negatives may occur. Therefore, this invention further provides a sample contamination detection method for undetermined samples, which can reconfirm whether the undetermined sample is contaminated. Based on this, this invention further provides a method for calculating the contamination value of the undetermined sample, referring to… Figure 6 An optional flowchart for calculating the contamination value of the undetermined sample is shown. This process can be used for low-contamination undetermined samples, and the process includes:
[0128] Step S30: Sort the preset sites according to their BAF values to obtain a site sequence;
[0129] The sorting of the preset sites can be done by arranging them in ascending or descending order based on their BAF (Balance of Faith). Sorting the preset sites is used to distinguish their types based on the range of BAF values. In this example, ascending order is used for illustration.
[0130] Step S31: Remove sites in the site sequence where BAF is greater than or equal to a preset value to obtain a second site sequence;
[0131] The preset value is used to screen homozygous sites and homozygous contamination homozygous sites from the preset sites as sites to be treated, thereby determining whether the sample is contaminated based on the site to be treated.
[0132] The preset value cannot be too large or too small. A value that is too large may lead to false positives due to CNVs (Copynumber Variations), while a value that is too small may result in false negatives due to insufficient loci in the sequence. In one example, the preset value is greater than a first threshold and less than or equal to five times the first threshold, used to retain homozygous loci and at least some homozygous contaminated homozygous loci in the sequence. Homozygous contaminated homozygous loci can be understood as homozygous loci that have been contaminated but have not transformed into heterozygous contaminated homozygous loci. In a specific example, the preset value can be [0.03, 0.1], preferably 0.05.
[0133] Step S32: Using the remaining preset sites as sites to be processed, obtain the BAF of the adjacent bases of the sites to be processed;
[0134] The adjacent bases are a second preset number of bases adjacent to the site to be treated. This second preset number needs to meet the requirements for calculating the significance of differences. It is understood that the more bases in the second preset number, the more accurate the significance of differences will be; however, from the perspective of saving computational resources, the fewer bases in the second preset number, the better. In this embodiment, the second preset number can be 10. Furthermore, the adjacent bases should simultaneously cover the upstream and downstream of the site to be treated. The specific allocation of the upstream and downstream bases can be an equal distribution, i.e., the number of bases upstream and downstream of the site to be treated is the same, or it can be allocated according to the actual situation, which will not be elaborated upon here.
[0135] The BAF of the adjacent bases of the site to be treated can be calculated based on nucleic acid information, or the BAF of the adjacent bases of the preset site can be calculated simultaneously during the calculation of the mutation genotype frequency value BAF of the sample at the preset site in step S11. Accordingly, this step can obtain the BAF of the adjacent bases of the site to be treated in the preset site.
[0136] Step S33: Calculate the significance of the difference between the adjacent bases of the site to be treated and the calculated site. If the difference parameter between the site to be treated and the site to be treated is within a preset range, then the sample to be determined is a low-contamination sample.
[0137] Specifically, the preset range can be determined based on specific difference parameters. In this embodiment of the invention, the difference parameter can be a p-value. In an optional example, the preset range corresponding to the p-value can be less than 0.001, thereby confirming that the sample to be determined is a low-contamination sample.
[0138] Wherein, if the difference parameter between the site to be processed and the site to be processed exceeds a preset range, the sample to be determined is considered to be an uncontaminated sample.
[0139] In a specific example, there are n sites to be treated. Let v1 be a vector composed of the BAFs of the n sites to be treated, denoted as {b1, b2, ..., bn}, and v2 be a vector composed of the BAFs of the reference bases of each site to be treated, denoted as {s1, s2, ..., sn}. The BAF of the reference bases of the site to be treated is the median value of the BAFs of the multiple neighboring bases of the site to be treated. Then, the Wilcox rank-sum test is used to test the significance of v1 and v2. If the difference is significant (p-value < 0.001), the sample is considered contaminated; otherwise, it is considered uncontaminated.
[0140] By conducting further contamination detection on the samples to be tested, missed detections can be minimized, thereby improving detection accuracy.
[0141] In a preferred implementation of this invention, when the sample to be determined is a low-contamination sample, the contamination value of the low-contamination sample can be further calculated. Specifically, please refer to... Figure 6 Calculating the contamination value of the low-contamination sample may include:
[0142] Step S34: Take the third preset number of sites with the highest BAF in the second site sequence as the effective sites, determine the median value of the BAF of the effective sites, and use the median value as the contamination value.
[0143] The third preset number can be calculated based on the proportion of homozygous contamination homozygous sites to the total number of homozygous sites. The proportion of homozygous contamination homozygous sites in the total number of homozygous sites can be obtained based on statistical data. In an optional example, refer to... Figure 7 The diagram shows the statistical data of the proportion of homozygous contaminated homozygous sites to the total number of homozygous sites. The horizontal axis represents the proportion of homozygous contaminated homozygous sites to the total number of homozygous sites, and the vertical axis represents the probability density. It can be seen that this proportion can be 20%.
[0144] Accordingly, the effective site can be understood as the homozygous contamination site identified in the second site sequence. After determining the effective site, the median value of the BAF of the effective site can be selected as the contamination value of the low-contamination sample.
[0145] The median BAF value of the effective sites can be understood as the BAF of the site located in the middle of the sequence after being arranged in ascending or descending order based on the BAF value.
[0146] In this embodiment of the invention, when the sample to be determined is a low-contamination sample, the contamination value of the low-contamination sample is further calculated, thereby providing a more accurate basis for the accuracy of the nucleic acid detection results of the sample.
[0147] It should be noted that the example values given in the embodiments of the present invention are determined based on actual statistical data, thereby making the calculation method of the embodiments of the present invention more closely reflect reality.
[0148] In addition, the sample contamination detection method provided in this embodiment of the invention requires fewer sites, thereby reducing computational costs while achieving good contamination detection performance.
[0149] Furthermore, the contamination detection process for the undetermined sample provided in this embodiment of the invention has higher accuracy. In theory, only the background noise level of the sample can affect it, thus making the accuracy of this contamination detection process even higher.
[0150] Furthermore, the contamination detection method provided in this embodiment of the invention is not affected by CNV (Copy number variations), thus achieving higher accuracy.
[0151] In this embodiment of the invention, contaminated sample simulation verification was further performed based on the above method. Specifically, 2500 uncontaminated clinical samples were used to simulate contaminated samples, with contamination value gradients of 2%, 5%, 10%, 20%, 30%, 40%, and 50%, respectively. 2500 samples were mixed for each gradient to observe the detection performance and stability of the contamination values in this embodiment of the invention.
[0152] In a specific simulation verification, a reference method was provided for comparison with the method provided in the embodiments of the present invention. The reference method combines Naive Bayes and maximum likelihood method for contamination detection. The results compared with the embodiments of the present invention are as follows:
[0153] 1. Specificity calculations were performed on uncontaminated samples, and the specific results are as follows:
[0154] False positive sample number Total number of simulation Specificity Reference method 284 2500 88.6% Embodiment method of the present application 3 2500 99.9%
[0155] As can be seen, the method provided in this embodiment of the invention has a very low number of false positive samples for the detection of uncontaminated samples, with a specificity of 99.9%, which is much higher than the 88.6% of the reference method.
[0156] Furthermore, in the false positive samples calculated by the method of this invention, the contamination value is generally low (less than 5%), while in the false positive samples calculated by the reference method, the contamination value is generally high (greater than 10%). Obviously, the reference method is very prone to producing significant false positive results, and the detection accuracy is clearly insufficient.
[0157] 2. Based on different methods, the contamination values of contaminated samples were calculated using the SSE (sum of squared errors) statistic. The specific results are as follows:
[0158] Simulated pollution level Reference method SSE Embodiment SSE of the present application 0.02 1.69 0.02 0.05 1.74 0.10 0.1 2.03 0.19 0.2 3.25 0.68 0.3 6.11 1.54 0.4 16.9 3.47 0.5 57.7 13.8
[0159] It can be seen that the method provided in this embodiment of the invention, for polluted samples with different pollution levels, yields a significantly lower SSE (Self-Solution Scale) for the pollution values compared to the reference method. Clearly, the method provided in this embodiment of the invention demonstrates significant accuracy and stability in calculating pollution values.
[0160] 3. Performance analysis was conducted on the contamination values of the contaminated samples based on different methods. Specific results are as follows: Figure 8 The diagrams showing the pollution value results obtained by different algorithms demonstrate that the pollution detection method provided by the embodiments of the present invention is more accurate and more stable at different pollution value levels.
[0161] Furthermore, embodiments of the present invention perform further contamination detection on the samples to be tested, effectively detecting low-contamination samples. In further statistical data, referencing... Figure 9 The diagram shows BAF statistics for the relative positions of preset loci, where the vertical axis represents the BAF value and the horizontal axis represents the relative position of the selected SNP loci (i.e., the preset loci). In the area where the SNP is 0, the contamination value is approximately 0.2%, combined with... Figure 8 The statistical data shows that the embodiments of the present invention can detect samples with a contamination value of only 0.2%, which is obviously a level that other methods cannot achieve.
[0162] The sample contamination detection device provided in the embodiments of the present invention will be described below. The device described below can be considered as a computer device, which consists of functional modules required to implement the sample contamination detection method provided in the embodiments of the present invention. The device described below can be referred to in correspondence with the method described above.
[0163] Figure 10 An optional block diagram of the sample contamination detection device provided in an embodiment of the present invention is shown. For example... Figure 10 As shown, the device may include:
[0164] Information acquisition module 200 is used to acquire nucleic acid information of the sample;
[0165] The first calculation module 210 is used to calculate the mutation genotype frequency value (BAF) of the sample at a preset site; wherein, there are multiple preset sites, and in the uncontaminated sample, the proportion of homozygous sites among the preset sites is greater than a first proportion.
[0166] The homozygous site determination module 220 is used to determine the homozygous site in the preset site based on the BAF of the preset site;
[0167] The second calculation module 230 is used to calculate the proportion of homozygous sites in the preset sites. If the proportion of homozygous sites is less than the first proportion, then the sample is a contaminated sample.
[0168] Optional, see reference Figure 11 Another optional block diagram of the sample contamination detection device shown further includes: a third calculation module 240 for calculating the contamination value of the contaminated sample when the sample is a contaminated sample.
[0169] Optionally, the third calculation module 240 is used to calculate the contamination value of the contaminated sample, including:
[0170] Based on the BAF of the preset sites, the preset sites are sorted to obtain a site sequence;
[0171] Remove homozygous sites from the site sequence to obtain the first site sequence;
[0172] Using the site in the preset position of the first site sequence as the marker point, the contamination value of the sample is calculated based on the BAF of the marker point;
[0173] The preset position is the position obtained by extending the sequence endpoint of the site by a first preset number of positions; the first preset number of positions is calculated based on the proportion of heterozygous genotype contamination homozygous sites in the preset site.
[0174] Optionally, the homozygous site determination module 220 is used to determine homozygous sites among the preset sites based on the BAF of the preset sites, including:
[0175] Obtain the first threshold;
[0176] When the BAF at the preset site is less than the first threshold, the preset site is considered a homozygous site.
[0177] Optionally, if the proportion of homozygous sites is greater than or equal to the first proportion, then the sample is a pending sample; the device further includes: a pending sample determination module 250, used to determine whether the pending sample is a low-contamination sample when the sample is a pending sample.
[0178] Optionally, the undetermined sample determination module 250 is used to determine whether the undetermined sample is a low-contamination sample, including:
[0179] Based on the BAF of the preset sites, the preset sites are sorted to obtain a site sequence;
[0180] Remove sites in the site sequence whose BAF is greater than or equal to a preset value to obtain a second site sequence; wherein the preset value is greater than the first threshold and less than or equal to 5 times the first threshold;
[0181] Using the site in the second site sequence as the site to be processed, obtain the BAF of the neighboring bases of the site to be processed; wherein, the neighboring bases are a second preset number of bases adjacent to the site to be processed;
[0182] Calculate the significance of the difference between the sites to be treated and their neighboring bases. If the difference parameter between the sites to be treated and their neighboring bases is within a preset range, then the sample to be determined is a low-contamination sample.
[0183] Optionally, if the sample to be determined is a low-contamination sample, the device further includes: a fourth calculation module 260, used to calculate the contamination value of the low-contamination sample;
[0184] The fourth calculation module 260 is used to calculate the contamination value of the low-contamination sample, including:
[0185] The third preset number of sites with the highest BAF in the second site sequence are taken as effective sites. The median value of the BAF of the effective sites is determined, and the median value is used as the contamination value.
[0186] The third preset number can be calculated based on the proportion of homozygous contamination sites to the total number of homozygous sites.
[0187] Optionally, the first calculation module 210, in the step of calculating the mutation genotype frequency value BAF of the sample at the preset site, also calculates the BAF of the adjacent bases at the preset site.
[0188] The undetermined sample judgment module 250 is used to obtain the BAF of the neighboring bases of the site to be processed, specifically, to obtain the BAF of the neighboring bases of the site to be processed in the preset site.
[0189] Optionally, the preset site is selected based on one or more factors of the site; the factors for selecting the preset site include: heterozygosity of the site, BAF of the site, BAF fluctuation range of the site, noise level of the site, capture performance of the site and distance between adjacent sites.
[0190] Optionally, the information acquisition module 200 is used to acquire nucleic acid information of the sample, including:
[0191] Obtain test data of the sample on the test equipment;
[0192] The test data was split and processed to obtain nucleic acid data;
[0193] The nucleic acid data is sorted and deduplicated to obtain the nucleic acid information.
[0194] This invention also provides a sample contamination detection system, which can be configured to execute the sample contamination detection method provided in this invention.
[0195] This invention also provides a computer device, which may include at least one memory and at least one processor; the memory stores one or more computer-executable instructions, and the processor invokes the one or more computer-executable instructions to execute the sample contamination detection method provided in this invention.
[0196] This invention provides a storage medium that stores one or more computer-executable instructions, which are used to execute the above-described sample contamination detection method.
[0197] The foregoing describes multiple embodiments of the present invention. The optional methods described in each embodiment can be combined and cross-referenced without conflict, thereby extending to a variety of possible embodiments. These can all be considered as embodiments disclosed or made public by the present invention.
[0198] While the embodiments of the present invention have been disclosed above, the present invention is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of the present invention; therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.
Claims
1. A method of sample contamination detection, characterized by, The method comprises the following steps: obtaining nucleic acid information of a sample; calculating a mutation genotype frequency value BAF of the sample at preset sites; wherein the preset sites are multiple, and in a non-contaminated sample, the proportion of homozygous sites in the preset sites is greater than a first proportion; determining homozygous sites in the preset sites according to the BAF of the preset sites, which comprises: obtaining a first threshold value; when the BAF of the preset sites is less than the first threshold value, the preset sites are determined as homozygous sites; calculating the proportion of the homozygous sites in the preset sites, if the proportion of the homozygous sites is less than the first proportion, the sample is a contaminated sample; if the proportion of the homozygous sites is greater than or equal to the first proportion, the sample is a pending sample; if the sample is a pending sample, determining whether the pending sample is a low-contaminated sample, which comprises: sorting the preset sites according to the BAF of the preset sites to obtain a site sequence; removing sites with BAF greater than or equal to a preset value in the site sequence to obtain a second site sequence; wherein the preset value is greater than the first threshold value and less than or equal to 5 times the first threshold value; taking sites in the second site sequence as to-be-processed sites, obtaining the BAF of adjacent bases of the to-be-processed sites; wherein the adjacent bases are a second preset number of bases adjacent to the to-be-processed sites; calculating the difference significance of the to-be-processed sites and the adjacent bases of the to-be-processed sites, if the difference parameters of the to-be-processed sites and the to-be-processed sites are within a preset range, the pending sample is a low-contaminated sample.
2. The method of claim 1, wherein, If the sample is a contaminated sample, the method further comprises: calculating the contamination value of the contaminated sample.
3. The method of claim 2, wherein, The calculation of the contamination value of the contaminated sample comprises: sorting the preset sites according to the BAF of the preset sites to obtain a site sequence; removing homozygous sites in the site sequence to obtain a first site sequence; taking sites in a preset position in the first site sequence as a calibration site, calculating the contamination value of the sample according to the BAF of the calibration site; wherein the preset position is a position obtained by extending a first preset number from the end of the site sequence; the first preset number is calculated based on the proportion of heterozygous genotype contaminated homozygous sites in the preset sites in the preset sites.
4. The method of claim 1, wherein, If the pending sample is a low-contaminated sample, the method further comprises: calculating the contamination value of the low-contaminated sample; The calculation of the contamination value of the low-contaminated sample comprises: taking a third preset number of to-be-processed sites with the highest BAF in the second site sequence as effective sites, determining the median value of the BAF of the effective sites, and taking the median value as the contamination value; wherein the third preset number can be calculated based on the proportion of homozygous contaminated homozygous sites in total homozygous sites.
5. The method according to claim 1 or 4, characterized in that, In the step of calculating the mutation genotype frequency value BAF of the sample at the preset sites, the BAF of the adjacent bases of the preset sites is also calculated. The BAF of the adjacent base of the to-be-processed site is obtained, specifically, the BAF of the adjacent base of the to-be-processed site in the preset site is obtained.
6. The method of claim 1, wherein, The preset site is selected based on one or more factors of the site. The factors for selecting the preset site include: heterozygosity of the site, BAF of the site, BAF fluctuation range of the site, noise level of the site, capture performance of the site, and spacing of adjacent sites.
7. The method of claim 1, wherein, The nucleic acid information of the sample includes: Obtain the test data of the sample in the test device; Split and post the test data to obtain nucleic acid data; Sort and de-duplicate the nucleic acid data to obtain the nucleic acid information.
8. A sample contamination detection device, characterized by, It includes: An information acquisition module for acquiring nucleic acid information of a sample; A first calculation module for calculating the mutation genotype frequency value BAF of the sample at a preset site; wherein the preset site is multiple, and in a non-contaminated sample, the proportion of homozygous sites in the preset site is greater than a first proportion; A homozygous site determination module for determining homozygous sites in the preset site according to the BAF of the preset site, which includes: obtaining a first threshold; when the BAF of the preset site is less than the first threshold, the preset site is a homozygous site; A second calculation module for calculating the proportion of the homozygous site in the preset site, if the proportion of the homozygous site is less than the first proportion, the sample is a contaminated sample; if the proportion of the homozygous site is greater than or equal to the first proportion, the sample is a pending sample; If the sample is a pending sample, determine whether the pending sample is a low-contaminated sample, the determination of whether the pending sample is a low-contaminated sample includes: sorting the preset site according to the BAF of the preset site to obtain a site sequence; removing sites in the site sequence with BAF greater than or equal to a preset value to obtain a second site sequence; wherein the preset value is greater than the first threshold and less than or equal to 5 times the first threshold; taking the sites in the second site sequence as to-be-processed sites, obtaining the BAF of the adjacent base of the to-be-processed sites; wherein the adjacent base is the second preset number of bases adjacent to the to-be-processed site; calculate the difference significance of the to-be-processed site and the adjacent base of the to-be-processed site, if the difference parameters of the to-be-processed site and the to-be-processed site are in a preset range, the pending sample is a low-contaminated sample.
9. A sample contamination detection system, characterized by, The sample contamination detection system is configured to perform the sample contamination detection method of any one of claims 1-7.
10. A computer device, comprising: It includes: At least one memory and at least one processor; The memory stores one or more computer executable instructions, and the processor invokes the one or more computer executable instructions to execute the sample contamination detection method of any one of claims 1-7.
11. A storage medium, characterized by The storage medium stores one or more computer executable instructions for executing the sample contamination detection method of any one of claims 1-7.
Citation Information
Patent Citations
A method for screening SNP sites used for detecting sample contamination in high-throughput sequencing and a sample contamination detecting method
CN109022562A
Sample quality evaluation method and device
CN111477277A
Method for detecting sample cross contamination and method for predicting cross contamination source
CN112746097A