Genotype detection method, sample contamination detection method, device, equipment and medium
Through the combination of dynamic abundance threshold and Gaussian mixed model, the problem of inaccurate sample pollution identification in existing genotype detection is solved, and high accuracy and low cost genotype and pollution detection are achieved.
Patent Information
- Application Number
- CN202210749217.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-28
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-06-28
AI Technical Summary
Existing genotype detection methods rely on paired sequencing data, resulting in inaccurate identification of sample pollution and require additional control samples for detection, which increases cost and difficulty, making it difficult to identify multiple sources of pollution and low-level pollution.
By obtaining at least 10 SNP sites to be tested in the sample to be detected, the abundance fluctuation level is calculated, the dynamic abundance threshold model is constructed, and the Gaussian mixed model is used for optimization and solution to determine the genotype and pollution status.
Improves the accuracy of genotype detection, can identify multiple sources of pollution and low-level pollution without additional control samples, reducing detection costs and complexity.
Smart Images

Figure CN115035950B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and in particular to a genotype detection method, a sample contamination detection method, a device, equipment and a medium. Background Art
[0002] SNP site genotyping and batch SNP site genotyping are necessary steps in site genotype identification based on high-throughput sequencing and copy number variation detection based on high-throughput sequencing. SNP (Single Nucleotide Polymorphism) mainly refers to the DNA (Deoxyribo Nucleic Acid) sequence polymorphism caused by the variation of a single nucleotide at the genomic level. The polymorphism expressed by the SNP site only involves the variation of a single base, which can be caused by the conversion or transversion of a single base, or by the insertion or deletion of a base. Data show that specific enzyme SNP sites are related to drug metabolism. For example, breast cancer patients carrying the CYP2D6*10 homozygous mutation may have a higher risk of recurrence with adjuvant therapy with amoxifen than wild-type patients. In addition, the identification of chromosomal copy number variation based on the transformation of batch heterozygous SNP abundance first depends on the identification of SNP genotypes. Typically, SNP genotyping uses a static abundance threshold method (e.g., Allele-specific copy number analysis of tumors, published by Peter Van Loo et al., PNAS vol. 107, no. 30: 16910-16915), whereby mutations with abundance exceeding a preset threshold (e.g., 95%) are considered homozygous. This method is significantly affected by tumor proportion, chromosomal structural variation, sample quality, and contamination, resulting in limited stability and accuracy.
[0003] On the other hand, in high-throughput sequencing, heterologous DNA contamination caused by sample storage and preparation processes is a non-negligible problem. Sample contamination directly leads to the presence of a large number of low-abundance mutations from heterologous DNA in mutation detection, which are difficult to distinguish from the true systemic mutations in the sample being tested, thus leading to misinterpretation of mutation results. Therefore, the identification of sample contamination is a key step in high-throughput sequencing quality control, and accurate quantification of sample contamination helps to identify and extract reliable systemic mutations. Typically, the identification of sample contamination relies on control samples, which in practice poses problems such as sampling difficulties and additional costs. Quantification of sample contamination is even more challenging, including the possibility that the sample contains a large number of copy number variations, may be contaminated by multiple sources, and is insensitive to the identification and quantification of low-level contamination. Some studies have constructed heterozygous SNP abundance models to quantify contamination levels. However, the static threshold-based SNP genotyping method directly leads to the unreliability of heterozygous SNP datasets.
[0004] For example, Fiévet et al.'s ART-DeCo: Easy Tool for Detection and Characterization of Cross-Contamination of DNA Samples in Diagnostic Next-Generation Sequencing Analysis (European Journal of Human Genetics (2019) 27:792–800) describes a method for detecting contamination in sequencing samples, including screening for non-heterozygous SNPs and determining the proportion of contamination sources. The method uses static thresholds for screening non-heterozygous SNPs, with AR (allele ratio, similar to allele frequency) between [0-0.005] and [0.995-1] considered non-heterozygous. There are two methods for determining the proportion of contamination sources: one is a simple calculation of the proportion of contamination sources based on sequencing data of non-heterozygous SNPs in the test sample using the WCS formula (see page 794, right column of the paper); the other is a more refined calculation of the proportion of contamination sources based on the SNP genotype of a contaminating sample and sequencing data of non-heterozygous SNPs in the test sample.
[0005] However, due to the variability between samples, using this static threshold to determine the SNP genotypes of all samples is sometimes not accurate enough. Furthermore, to identify the source of contamination, an additional set of contaminated samples must be tested in addition to the sample to be tested. Contamination identification cannot be achieved through the test data of a single sample to be tested. Summary of the Invention
[0006] The present invention provides a genotype detection method, a sample contamination detection method, an apparatus, a device and a medium to solve the problem that existing genotype detection methods rely on paired sequencing data, and provide data guarantee for the accuracy of subsequent sample contamination detection.
[0007] According to one aspect of the present invention, a genotype detection method is provided, the method comprising:
[0008] Obtain at least 10 SNP sites to be tested in the sample to be tested and obtain at least 3 abundance thresholds to be tested;
[0009] For each abundance threshold to be tested, the abundance fluctuation level of the sample to be tested is calculated based on the abundance threshold to be tested and the minor allele frequency corresponding to each SNP site to be tested;
[0010] The abundance threshold to be measured corresponding to the abundance fluctuation level that meets the inflection point characteristics among at least three abundance fluctuation levels is used as the first abundance threshold;
[0011] Based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested, the genotype corresponding to each SNP site to be tested is determined.
[0012] According to another aspect of the present invention, a method for detecting sample contamination is provided, the method comprising:
[0013] Obtain at least 10 SNP sites to be detected in the sample to be tested, and determine the genotype corresponding to each SNP site to be detected;
[0014] Constructing a Gaussian mixture model based on the minor allele frequency corresponding to the SNP site to be tested that meets the genotype conditions; wherein the genotype conditions include the genotype being wild type and / or the genotype being homozygous;
[0015] An optimization solution operation is performed on the Gaussian mixture model to obtain contamination status data corresponding to the sample to be tested; wherein the contamination status data includes at least one of a contamination ratio, a minor allele frequency variance, and the number of contamination sources.
[0016] According to another aspect of the present invention, a genotype detection device is provided, comprising:
[0017] A SNP site acquisition module to be tested is used to obtain at least 10 SNP sites to be tested in the sample to be tested and obtain at least 3 abundance thresholds to be tested;
[0018] an abundance fluctuation level determination module, configured to calculate, for each abundance threshold to be measured, the abundance fluctuation level of the sample to be tested based on the abundance threshold to be measured and the minor allele frequency corresponding to each SNP site to be measured;
[0019] a first abundance threshold determination module, configured to use the abundance threshold to be measured corresponding to the abundance fluctuation level that meets the inflection point feature among the at least three abundance fluctuation levels as the first abundance threshold;
[0020] The genotype determination module is used to determine the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested.
[0021] According to another aspect of the present invention, there is provided a sample contamination detection device, the device comprising:
[0022] The genotype determination module is used to obtain at least 10 SNP sites to be detected in the sample to be detected and determine the genotype corresponding to each SNP site to be detected;
[0023] A Gaussian mixture model construction module is used to construct a Gaussian mixture model based on the minor allele frequency corresponding to the SNP site to be tested that meets the genotype conditions; wherein the genotype conditions include the genotype being wild type and / or the genotype being homozygous;
[0024] A contamination status data determination module is used to perform an optimization solution operation on the Gaussian mixture model to obtain contamination status data corresponding to the sample to be tested; wherein the contamination status data includes at least one of the contamination ratio, the variance of the minor allele frequency and the number of contamination sources.
[0025] According to another aspect of the present invention, an electronic device is provided, comprising:
[0026] at least one processor; and
[0027] a memory communicatively connected to the at least one processor; wherein,
[0028] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the genotype detection method or sample contamination detection method described in any embodiment of the present invention.
[0029] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the genotype detection method or sample contamination detection method described in any embodiment of the present invention when executed.
[0030] The technical solution of the embodiment of the present invention determines the abundance fluctuation level of the sample to be tested based on the minor allele frequency corresponding to the abundance threshold to be tested and each SNP site to be tested for each obtained abundance threshold to be tested, and uses the abundance threshold to be tested corresponding to the abundance fluctuation level that meets the inflection point characteristics among at least three abundance fluctuation levels as the first abundance threshold. The embodiment of the present invention determines the first abundance threshold through the constructed dynamic abundance threshold model, and determines the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested.
[0031] This dynamic threshold approach effectively avoids inaccurate genotype determination caused by deviations in homozygous, heterozygous, and wild-type SNP abundance from 0%, 50%, and 100% due to high tumor proportions, copy number variations, and sample contamination. Furthermore, the algorithm does not rely on control samples, addressing the difficulties of obtaining control samples and increased testing costs.
[0032] Furthermore, the present invention can have higher accuracy for distinguishing SNP heterozygous and non-heterozygous genotypes by relying on dynamic thresholds, i.e., lower false positives or false negatives. At the same time, the hybrid model method of the present invention can confirm the number of sample contamination sources and can be used for contamination identification and quantification of multiple contamination sources, identification and quantification of low-level contamination, and stable identification and quantification of contamination levels in samples with different chromosomal states, including stable chromosomes and samples with a large number of copy number variations. It also does not require sequencing information of other samples (e.g., reference samples, contaminated samples, etc.) in addition to the sample to be tested.
[0033] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0035] Figure 1 This is a flow chart of a genotype detection method provided in Example 1 of the present invention;
[0036] Figure 2 This is a flow chart of a sample contamination detection method provided according to the second embodiment of the present invention;
[0037] Figure 3 This is a flow chart of a sample contamination detection method provided in accordance with the third embodiment of the present invention;
[0038] Figure 4 This is a schematic structural diagram of a genotype detection device provided according to a fourth embodiment of the present invention;
[0039] Figure 5 This is a structural diagram of a sample contamination detection device provided according to a fifth embodiment of the present invention;
[0040] Figure 6 FIG. 1 is a structural diagram of an electronic device provided according to a sixth embodiment of the present invention. DETAILED DESCRIPTION
[0041] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0042] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0043] Example 1
[0044] Figure 1 This is a flow chart of a genotype detection method provided according to the first embodiment of the present invention. This embodiment is applicable to the case of performing genotype detection on germline SNP sites in a single tissue sample. The method can be performed by a genotype detection method device. The genotype detection device can be implemented in the form of hardware and / or software. The genotype detection device can be configured in a terminal device. Figure 1 As shown, the method includes:
[0045] S110, obtaining at least 10 SNP sites to be tested in the sample to be tested and obtaining at least 3 abundance thresholds to be tested.
[0046] In this example, the sample to be tested is a single tissue sample, such as a cancer tissue sample. In this embodiment, the type of SNP site to be tested is a germline SNP site. There are a large number of variant sites in the human genome, which can be divided into germline mutations and somatic mutations according to their sources. Germline mutations, also known as germ cell mutations, are mutations that originate from reproductive cells such as sperm or eggs. Therefore, usually all cells in the body carry this mutation, including cancer tissue samples. Somatic mutations, also known as acquired mutations, are mutations acquired during growth and development or under the influence of environmental factors. Usually, only some cells in the body carry this mutation.
[0047] Specifically, after quality control of the FASTQ file (off-machine data) containing the high-throughput sequencing sequence of the sample to be tested, the high-throughput sequencing sequence is aligned to the human reference genome using alignment software, and the SNP sites in the high-throughput sequencing sequence that are aligned to the human reference genome are used as the SNP sites to be tested, and a SAM file is generated. Exemplarily, the alignment software can be BWA-MEM (0.7.10), and the human reference genome can be hg19 or b37. Exemplarily, the SAM file is converted into a BAM file using samtools (0.1.19) software.
[0048] As an optional embodiment, obtaining at least 10 SNPs to be tested in the sample to be tested includes: selecting at least 10 SNPs in the sample to be tested whose population frequencies fall within a preset population frequency range as the SNPs to be tested; wherein the preset population frequency range includes 0.4-0.6. The population frequency represents the frequency of occurrence of the SNP in a human genome. Specifically, SNPs whose population frequencies, when mapped to a human reference genome, fall within a range of 0.4-0.6 are selected as the SNPs to be tested.
[0049] Specifically, the abundance thresholds to be tested are used to represent the screening threshold corresponding to the minor allele frequency. The threshold differences between the abundance thresholds to be tested can be the same or different. For example, a set of abundance thresholds to be tested, th, satisfies th∈{0%, 1%, 2%, ..., 10%, 30%}. The specific parameters of the abundance thresholds to be tested are not limited here; users can set each abundance threshold to be tested according to their actual needs.
[0050] S120. For each abundance threshold to be tested, calculate the abundance fluctuation level of the sample to be tested based on the abundance threshold to be tested and the minor allele frequency corresponding to each SNP site to be tested.
[0051] As an optional embodiment, the abundance fluctuation level of the sample to be tested is calculated based on the abundance threshold to be tested and the minor allele frequency corresponding to each SNP site to be tested, including: for each chromosome where each SNP site to be tested is located, based on the copy numbers corresponding to at least two SNP sites to be tested on the chromosome, a segmentation operation is performed on the chromosome to obtain one or more chromosome segments; wherein the copy number difference corresponding to any two SNP sites to be tested in each chromosome segment is less than a preset difference threshold; for each chromosome segment, the SNP site to be tested in the chromosome segment whose minor allele frequency is greater than the abundance threshold to be tested is used as the target SNP site; based on the minor allele frequencies corresponding to at least two target SNP sites, the abundance fluctuation level of the sample to be tested is calculated.
[0052] Specifically, the types of chromosomes include autosomes and / or sex chromosomes. Of the 23 pairs of human chromosomes, 22 are autosomes and one is a sex chromosome. The chromosomes in this embodiment include all chromosomes that may have heterozygous variations.
[0053] Specifically, copy number represents the number of times a gene or gene sequence occurs in a genome, or the number of times a gene segment is repeated in the genome. Normally, the copy number is 2. When copy number variation occurs, the genome rearranges, and the gene copy number changes from 2 to n, where n is a natural number not equal to 2.
[0054] Specifically, the coverage depth of the SNP site to be tested is calculated by GATK software, and the copy number of the SNP site to be tested is calculated based on the coverage depth. Wherein, the coverage depth refers to the ratio of the total number of bases (bp) obtained by sequencing to the genome size (Genome). In this embodiment, the coverage depth of the SNP site to be tested is greater than 200X.
[0055] Wherein, illustratively, a CBS (circular binary segmentation) algorithm is used to segment the chromosome into one or more chromosome segments based on the copy number. Wherein, the CBS algorithm can be implemented by the R software package PSCBS. In this embodiment, the copy number difference corresponding to any two SNP sites to be tested in each chromosome segment is less than a preset difference threshold, where, illustratively, the preset difference threshold can be 2. For example, the copy numbers of the four SNP sites to be tested on the chromosome are 2, 3, 5 and 6, respectively. Then, chromosome segment A contains two SNP sites to be tested with copy numbers of 2 and 3, respectively, and chromosome segment B contains two SNP sites to be tested with copy numbers of 5 and 6, respectively.
[0056] For each mutated SNP, the minor allele frequency (MAF) represents the ratio of the number of minor alleles to the total number of alleles. Alleles are genes located at the same position on a pair of homologous chromosomes that control different forms of the same trait. For example, suppose a SNP on a chromosome in 100 individuals has three alleles: A, C, and G. Genome sequencing reveals that the number of bases A, C, and G occurring is 100, 80, and 20, respectively. The MAF for this SNP is 80 / 200 = 0.4.
[0057] As an optional embodiment, the calculation of the abundance fluctuation level of the sample to be tested based on the minor allele frequencies corresponding to at least two target SNP sites includes: sorting each target SNP site based on the site position of each target SNP site in the chromosome; based on the sorting result, taking the difference between the minor allele frequency of the current target SNP site and the minor allele frequency of the next target SNP site as the abundance difference; performing averaging processing on at least one abundance difference to obtain the abundance fluctuation level of the sample to be tested.
[0058] For example, assuming that there are four target SNP sites on a chromosome, namely target SNP site A, target SNP site B, target SNP site C, and target SNP site D, with corresponding copy numbers of 2, 6, 5, and 3, respectively. Then, the target SNP sites contained in chromosome segment A and their order are target SNP site A and target SNP site D, and the target SNP sites contained in chromosome segment B and their order are target SNP site B and target SNP site C.
[0059] For example, assuming that chromosome segment A contains 100 target SNP sites, the number of calculated abundance differences is 99. The abundance differences are averaged to obtain the abundance fluctuation level corresponding to chromosome segment A.
[0060] S130 , using the abundance threshold to be measured corresponding to the abundance fluctuation level that meets the inflection point feature among the at least three abundance fluctuation levels as a first abundance threshold.
[0061] Specifically, an abundance fluctuation level curve is constructed with the abundance threshold to be measured as the horizontal axis and the abundance fluctuation level as the vertical axis. As the abundance threshold to be measured increases, the abundance fluctuation level first increases and then tends to stabilize.
[0062] S140 : Determine the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested.
[0063] Among them, mutation abundance, also known as mutant allele frequency, is used to characterize the ratio of the number of reads of the mutant allele to the total number of reads of the entire SNP site. Specifically, the samtools mpileup software is used to calculate the number of reads A of the mutant allele and the number of reads B of the wild-type allele at the SNP site to be tested. Mutation abundance = A / (A+B). For example, if the alleles at the SNP site are represented as AAA, BBB, ABB, and AAB, the mutation abundance corresponding to each allele representation is 100%, 0%, 25%, and 75%, respectively.
[0064] In this embodiment, genotype includes heterozygous, wild type and homozygous. Heterozygous is used to characterize individuals with genotypes containing different types of alleles on homologous chromosomes, such as AB. Wild type is used to characterize individuals with genotypes containing wild alleles on homologous chromosomes, such as AA, and homozygous is used to characterize individuals with genotypes containing homozygous alleles on homologous chromosomes, such as BB.
[0065] As an optional embodiment, the method of determining the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested includes: calculating a second abundance threshold based on the first abundance threshold; wherein the sum of the first abundance threshold and the second abundance threshold is one; for each SNP site to be tested, if the mutation abundance corresponding to the SNP site to be tested is less than or equal to the first abundance threshold, the genotype of the SNP site to be tested is wild type; if the mutation abundance corresponding to the SNP site to be tested is greater than or equal to the second abundance threshold, the genotype of the SNP site to be tested is homozygous; if the mutation abundance corresponding to the SNP site to be tested is greater than the first abundance threshold and less than the second abundance threshold, the genotype of the SNP site to be tested is heterozygous.
[0066] Specifically, the second abundance threshold = 1 - the first abundance threshold, and the genotype of the SNP site to be tested satisfies the following relationship:
[0067]
[0068] Where AF represents mutation abundance, tBest represents the first abundance threshold, and mBest represents the second abundance threshold.
[0069] The data in this example are based on high-throughput sequencing results from 100 real tissue samples. The tissue samples included samples with varying tumor percentages, varying degrees of chromosomal stability, and samples with both uncontaminated and varying degrees of contamination. Simultaneously, high-throughput sequencing was performed on control white blood cell samples from the same individuals as the tissue samples to determine the true genotypes of the SNPs, serving as a performance reference standard. Unlike cancerous tissue samples, the genomes of white blood cell samples are stable, with virtually no contamination from sample storage and preparation. To ensure the reliability of the reference standard, we employed a simple model based on the principle that close to 0% indicates wild type, close to 100% indicates homozygous, and close to 50% indicates heterozygous for direct genotyping. To account for data noise, a 5% abundance fluctuation was allowed, meaning that an abundance of 0%-5% indicates wild type, 95%-100% indicates homozygous, and 47.5%-52.5% indicates heterozygous. For tissue samples, we compared genotyping based on two static threshold combinations with the dynamic threshold algorithm of the present invention.
[0070] Among them, static threshold combination 1:
[0071]
[0072] Among them, static threshold combination 2:
[0073]
[0074] Furthermore, heterozygous SNP sites were set as positive sites and non-heterozygous sites as negative sites, and the consistency of the three strategies with the white blood cell reference standard was evaluated, as shown in Table 1. The dynamic threshold method showed good and stable performance in both sensitivity and specificity, and had the highest accuracy.
[0075] Table 1 is a consistency list of genotype identification provided in Example 1 of the present invention.
[0076]
[0077]
[0078] The technical solution of this embodiment determines the abundance fluctuation level of the sample to be tested based on the minor allele frequency corresponding to the abundance threshold to be tested and each SNP site to be tested for each obtained abundance threshold to be tested, and uses the abundance threshold to be tested corresponding to the abundance fluctuation level that meets the inflection point characteristics among at least three abundance fluctuation levels as the first abundance threshold. The embodiment of the present invention determines the first abundance threshold through the constructed dynamic abundance threshold model, and determines the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested, thereby solving the problem of existing genotype detection methods relying on paired sequencing data and improving the accuracy of genotype detection results.
[0079] Example 2
[0080] Figure 2 This is a flow chart of a sample contamination detection method provided according to the second embodiment of the present invention. This embodiment is applicable to the case of detecting and evaluating the contamination level of a single tissue sample. The method can be executed by a sample contamination detection method device. The genotype detection device can be implemented in the form of hardware and / or software. The sample contamination detection device can be configured in a terminal device. Figure 2 As shown, the method includes:
[0081] S210: Obtain at least 10 SNP sites to be detected in the sample to be detected, and determine the genotype corresponding to each SNP site to be detected.
[0082] In this example, the sample to be tested is a single tissue sample, such as a cancer tissue sample. In this embodiment, the type of SNP site to be tested is a germline SNP site. There are a large number of variant sites in the human genome, which can be divided into germline mutations and somatic mutations according to their sources. Germline mutations, also known as germ cell mutations, are mutations that originate from reproductive cells such as sperm or eggs. Therefore, usually all cells in the body carry this mutation, including cancer tissue samples. Somatic mutations, also known as acquired mutations, are mutations acquired during growth and development or under the influence of environmental factors. Usually, only some cells in the body carry this mutation.
[0083] Specifically, after quality control of the FASTQ file (off-machine data) containing the high-throughput sequencing sequence of the sample to be tested, the high-throughput sequencing sequence is aligned to the human reference genome using alignment software, and the SNP sites in the high-throughput sequencing sequence that are aligned to the human reference genome are used as the SNP sites to be tested, and a SAM file is generated. Exemplarily, the alignment software can be BWA-MEM (0.7.10), and the human reference genome can be hg19 or b37. Exemplarily, the SAM file is converted into a BAM file using samtools (0.1.19) software.
[0084] As an optional embodiment, obtaining at least 10 SNPs to be tested in the sample to be tested includes: selecting at least 10 SNPs in the sample to be tested whose population frequencies fall within a preset population frequency range as the SNPs to be tested; wherein the preset population frequency range includes 0.4-0.6. The population frequency represents the frequency of occurrence of the SNP in a human genome. Specifically, SNPs whose population frequencies, when mapped to a human reference genome, fall within a range of 0.4-0.6 are selected as the SNPs to be tested.
[0085] As an optional embodiment, determining the genotype corresponding to each SNP site to be tested includes: using a preset detection method to determine the genotype corresponding to each SNP site to be tested; wherein the preset detection method includes at least one of sequencing method, chip method and mass spectrometry method.
[0086] In another embodiment, optionally, determining the genotype corresponding to each SNP site to be tested includes: determining the genotype corresponding to each SNP site to be tested includes: obtaining a first abundance threshold, and determining the mutation abundance corresponding to each SNP site to be tested; based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested, determining the genotype corresponding to each SNP site to be tested.
[0087] As an optional embodiment, the first abundance threshold is preset by the user. For example, the first preset abundance threshold may be 0.01 or 0.03. The specific value of the first preset abundance threshold is not limited here, and the user may set it according to actual needs.
[0088] Among them, mutation abundance, also known as mutant allele frequency, is used to characterize the ratio of the number of reads of the mutant allele to the number of reads of the entire SNP site. Among them, alleles are used to describe genes that control different forms of the same trait at the same position on a pair of homologous chromosomes. Specifically, the samtools mpileup software is used to calculate the number of reads A of the mutant allele and the number of reads B of the wild-type allele at the SNP site to be tested. Mutation abundance = A / (A+B). For example, if the alleles at the SNP site are represented as AAA, BBB, ABB, and AAB, the mutation abundance corresponding to each allele representation is 100%, 0%, 25%, and 75%, respectively.
[0089] In this embodiment, genotype includes heterozygous, wild type and homozygous. Heterozygous is used to characterize individuals with genotypes containing different types of alleles on homologous chromosomes, such as AB. Wild type is used to characterize individuals with genotypes containing wild alleles on homologous chromosomes, such as AA, and homozygous is used to characterize individuals with genotypes containing homozygous alleles on homologous chromosomes, such as BB.
[0090] As an optional embodiment, the method of determining the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested includes: calculating a second abundance threshold based on the first abundance threshold; wherein the sum of the first abundance threshold and the second abundance threshold is one; for each SNP site to be tested, if the mutation abundance corresponding to the SNP site to be tested is less than or equal to the first abundance threshold, the genotype of the SNP site to be tested is wild type; if the mutation abundance corresponding to the SNP site to be tested is greater than or equal to the second abundance threshold, the genotype of the SNP site to be tested is homozygous; if the mutation abundance corresponding to the SNP site to be tested is greater than the first abundance threshold and less than the second abundance threshold, the genotype of the SNP site to be tested is heterozygous.
[0091] Specifically, the second abundance threshold = 1 - the first abundance threshold, and the genotype of the SNP site to be tested satisfies the following relationship:
[0092]
[0093] Where AF represents mutation abundance, tBest represents the first abundance threshold, and mBest represents the second abundance threshold.
[0094] S220. Construct a Gaussian mixture model based on the minor allele frequency corresponding to the SNP site to be tested that meets the genotype conditions.
[0095] In this embodiment, the genotype condition includes the genotype being a wild type and / or the genotype being a homozygous type.
[0096] Specifically, SNPs typically contain only two alleles. For a set of SNPs to be tested, assume that the alleles for the four tested SNPs are AAA, BBB, ABB, and AAB, respectively. When calculating the mutation abundance, we need to determine the proportion of A or B in the total number of alleles. If A is a mutant, the mutation abundances for the four tested SNPs are 100%, 0%, 25%, and 75%, respectively. For non-heterozygous SNPs, we only care about the degree of disparity between the number of A and B alleles, not the proportion of A or B in the total number of alleles. Therefore, this allele frequency well meets this requirement. The minor allele frequencies for the four tested SNPs are 0%, 0%, 25%, and 25%, respectively. From the above example, we can see that by mapping the mutation abundance around 50%, the mapped mutation abundance equals the minor allele frequency.
[0097] As an optional implementation, the Gaussian mixture model satisfies the formula:
[0098]
[0099] in,
[0100]
[0101] in,
[0102]
[0103] Where maf represents the minor allele frequency, δ 2 represents the variance of the minor allele frequency, n represents the number of pollution sources, α represents the pollution ratio, pbinom represents the Bernoulli probability distribution, P(C=i) represents the probability distribution corresponding to the i pollution sources, and N represents the probability distribution of the Gaussian mixture model.
[0104] S230 , performing an optimization solution operation on the Gaussian mixture model to obtain pollution status data corresponding to the sample to be detected.
[0105] Exemplary optimization methods include, but are not limited to, at least one of maximum likelihood estimation, expectation maximization, and Markov chain Monte Carlo algorithms. Specifically, the optim function in the R software package stats can be used to optimize the Gaussian mixture model using the L-BFGS-B algorithm to obtain the pollution state data.
[0106] In this embodiment, the contamination status data includes at least one of a contamination ratio, a minor allele frequency variance, and the number of contamination sources.
[0107] The technical solution of this embodiment obtains at least 10 SNP sites to be tested in the sample to be tested, determines the genotype corresponding to each SNP site to be tested, constructs a Gaussian mixture model based on the minor allele frequency corresponding to the non-heterozygous SNP site to be tested, and performs an optimization solution operation on the Gaussian mixture model to obtain contamination status data corresponding to the sample to be tested. This solves the problem that existing sample contamination measurement methods are greatly affected by copy number variation / loss of heterozygosity and have difficulty in detecting low-level contamination, thereby improving the accuracy of sample contamination measurement results.
[0108] Example 3
[0109] Figure 3 This is a flow chart of a sample contamination detection method provided according to the third embodiment of the present invention. This embodiment further refines the technical feature of "obtaining the first abundance threshold" in the above embodiment. Figure 3 As shown, the method includes:
[0110] S310: Obtain at least 10 SNP sites to be tested in the sample to be tested and obtain at least 3 abundance thresholds to be tested.
[0111] Specifically, the abundance thresholds to be tested are used to represent the screening threshold corresponding to the minor allele frequency. The threshold differences between the abundance thresholds to be tested can be the same or different. For example, a set of abundance thresholds to be tested, th, satisfies th∈{0%, 1%, 2%, ..., 10%, 30%}. The specific parameters of the abundance thresholds to be tested are not limited here; users can set each abundance threshold to be tested according to their actual needs.
[0112] S320. For each abundance threshold to be tested, calculate the abundance fluctuation level of the sample to be tested based on the abundance threshold to be tested and the minor allele frequency corresponding to each SNP site to be tested.
[0113] As an optional embodiment, the abundance fluctuation level of the sample to be tested is calculated based on the abundance threshold to be tested and the minor allele frequency corresponding to each SNP site to be tested, including: for each chromosome where each SNP site to be tested is located, based on the copy numbers corresponding to at least two SNP sites to be tested on the chromosome, a segmentation operation is performed on the chromosome to obtain one or more chromosome segments; wherein the copy number difference corresponding to any two SNP sites to be tested in each chromosome segment is less than a preset difference threshold; for each chromosome segment, the SNP site to be tested in the chromosome segment whose minor allele frequency is greater than the abundance threshold to be tested is used as the target SNP site; based on the minor allele frequencies corresponding to at least two target SNP sites, the abundance fluctuation level of the sample to be tested is calculated.
[0114] Specifically, the types of chromosomes include autosomes and / or sex chromosomes. Of the 23 pairs of human chromosomes, 22 are autosomes and one is a sex chromosome. The chromosomes in this embodiment include all chromosomes that may have heterozygous variations.
[0115] Specifically, copy number represents the number of times a gene or gene sequence occurs in a genome, or the number of times a gene segment is repeated in the genome. Normally, the copy number is 2. When copy number variation occurs, the genome rearranges, and the gene copy number changes from 2 to n, where n is a natural number not equal to 2.
[0116] Specifically, the coverage depth of the SNP site to be tested is calculated by GATK software, and the copy number of the SNP site to be tested is calculated based on the coverage depth. Wherein, the coverage depth refers to the ratio of the total number of bases (bp) obtained by sequencing to the genome size (Genome). In this embodiment, the coverage depth of the SNP site to be tested is greater than 200X.
[0117] Wherein, illustratively, a CBS (circular binary segmentation) algorithm is used to segment the chromosome into one or more chromosome segments based on the copy number. Wherein, the CBS algorithm can be implemented by the R software package PSCBS. In this embodiment, the copy number difference corresponding to any two SNP sites to be tested in each chromosome segment is less than a preset difference threshold, where, illustratively, the preset difference threshold can be 2. For example, the copy numbers of the four SNP sites to be tested on the chromosome are 2, 3, 5 and 6, respectively. Then, chromosome segment A contains two SNP sites to be tested with copy numbers of 2 and 3, respectively, and chromosome segment B contains two SNP sites to be tested with copy numbers of 5 and 6, respectively.
[0118] As an optional embodiment, the calculation of the abundance fluctuation level of the sample to be tested based on the minor allele frequencies corresponding to at least two target SNP sites includes: sorting each target SNP site based on the site position of each target SNP site in the chromosome; based on the sorting result, taking the difference between the minor allele frequency of the current target SNP site and the minor allele frequency of the next target SNP site as the abundance difference; performing averaging processing on at least one abundance difference to obtain the abundance fluctuation level of the sample to be tested.
[0119] For example, assuming that there are four target SNP sites on a chromosome, namely target SNP site A, target SNP site B, target SNP site C, and target SNP site D, with corresponding copy numbers of 2, 6, 5, and 3, respectively. Then, the target SNP sites contained in chromosome segment A and their order are target SNP site A and target SNP site D, and the target SNP sites contained in chromosome segment B and their order are target SNP site B and target SNP site C.
[0120] For example, assuming that chromosome segment A contains 100 target SNP sites, the number of calculated abundance differences is 99. The abundance differences are averaged to obtain the abundance fluctuation level corresponding to chromosome segment A.
[0121] S330: Using the abundance threshold to be measured corresponding to the abundance fluctuation level that meets the inflection point feature among the at least three abundance fluctuation levels as the first abundance threshold.
[0122] Specifically, an abundance fluctuation level curve is constructed with the abundance threshold to be measured as the horizontal axis and the abundance fluctuation level as the vertical axis. As the abundance threshold to be measured increases, the abundance fluctuation level first increases and then tends to stabilize.
[0123] S340: Determine the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested.
[0124] S350. Construct a Gaussian mixture model based on the minor allele frequency corresponding to the SNP site to be tested that meets the genotype conditions.
[0125] S360: performing an optimization solution operation on the Gaussian mixture model to obtain pollution status data corresponding to the sample to be detected.
[0126] The third embodiment of the present invention further provides measurement results of simulated data from 20 groups of human cancer tissue samples using the sample contamination detection method described in this embodiment.
[0127] Specifically, 20 groups of human cancer tissue samples were mixed at a preset contamination ratio to contaminate the sample data and obtain samples to be tested. Table 2 below is a sample contamination description table provided in Example 3 of the present invention.
[0128]
[0129] Table 3 below shows the contamination ratio data obtained by the sample contamination detection method provided in Example 3 of the present invention.
[0130] Pollution level Pollution detection rate Predicted pollution ratio 0.5% 100% 0.47% 1% 100% 0.81% 5% 100% 5.2% 10% 100% 9.6% 20% 100% 18%
[0131] It can be seen from Table 3 that the embodiment of the present invention is applicable to pollution situations with single sample and multiple sample sources, can still stably detect pollution at a pollution level of 0.5%, and can accurately predict the pollution level.
[0132] Table 4 below shows the pollution detection rates obtained by measuring different pollution levels using different detection methods provided in Example 3 of the present invention.
[0133]
[0134]
[0135] Table 5 below shows the pollution ratio data obtained by an existing conPair (v0.2) method provided in Example 3 of the present invention.
[0136]
[0137] Tables 4 and 5 show that the existing ART-DeCo (v1.1) method is unable to identify low-level contaminated samples. Although the existing conPair (v0.2) method can detect low-level contaminated samples, the contamination ratio obtained for contaminated samples containing a large number of copies is significantly higher than the actual contamination ratio.
[0138] Therefore, the sample contamination detection method provided by the embodiment of the present invention is higher than similar software in terms of detection sensitivity and accuracy of contamination level assessment.
[0139] The technical solution of this embodiment determines the abundance fluctuation level of the sample to be tested based on the minor allele frequency corresponding to the abundance threshold to be tested and each SNP site to be tested for each obtained abundance threshold to be tested, and uses the abundance threshold to be tested corresponding to the abundance fluctuation level that meets the inflection point characteristics among at least three abundance fluctuation levels as the first abundance threshold. The embodiment of the present invention determines the first abundance threshold through the constructed dynamic abundance threshold model, and determines the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested. This solves the problem of existing genotype detection methods relying on paired sequencing data, improves the accuracy of genotype detection results, and thus further improves the accuracy of sample contamination detection results.
[0140] Example 4
[0141] Figure 4 Schematic diagram of a genotype detection device according to the fourth embodiment of the present invention. Figure 4 As shown, the device includes: a SNP site acquisition module 410 to be detected, an abundance fluctuation level determination module 420, a first abundance threshold determination module 430 and a genotype determination module 440.
[0142] The SNP site acquisition module 410 is used to obtain at least 10 SNP sites to be tested in the sample to be tested and obtain at least 3 abundance thresholds to be tested;
[0143] The abundance fluctuation level determination module 420 is configured to calculate the abundance fluctuation level of the sample to be tested based on each abundance threshold to be tested and the minor allele frequency corresponding to each SNP site to be tested;
[0144] a first abundance threshold determination module 430, configured to use the abundance threshold to be measured corresponding to the abundance fluctuation level that meets the inflection point feature among the at least three abundance fluctuation levels as the first abundance threshold;
[0145] The genotype determination module 440 is configured to determine the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested.
[0146] The technical solution of this embodiment determines the abundance fluctuation level of the sample to be tested based on the minor allele frequency corresponding to the abundance threshold to be tested and each SNP site to be tested for each obtained abundance threshold to be tested, and uses the abundance threshold to be tested corresponding to the abundance fluctuation level that meets the inflection point characteristics among at least three abundance fluctuation levels as the first abundance threshold. The embodiment of the present invention determines the first abundance threshold through the constructed dynamic abundance threshold model, and determines the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested, thereby solving the problem of existing genotype detection methods relying on paired sequencing data and improving the accuracy of genotype detection results.
[0147] Based on the above embodiment, optionally, the abundance fluctuation level determination module 420 includes:
[0148] a chromosome segment determination unit configured to, for each chromosome where each SNP site to be tested is located, perform a segmentation operation on the chromosome based on the copy numbers corresponding to at least two SNP sites to be tested on the chromosome, thereby obtaining one or more chromosome segments; wherein the copy number difference between any two SNP sites to be tested in each chromosome segment is less than a preset difference threshold;
[0149] a target SNP site determination unit, configured to, for each chromosome segment, determine a SNP site to be tested in the chromosome segment whose minor allele frequency is greater than the abundance threshold to be tested as a target SNP site;
[0150] The abundance fluctuation level unit is used to calculate the abundance fluctuation level of the sample to be tested based on the minor allele frequencies corresponding to at least two target SNP sites.
[0151] Based on the above embodiment, optionally, the abundance fluctuation level unit is specifically used to:
[0152] Sort each target SNP site based on its location in the chromosome;
[0153] Based on the sorting results, the difference between the minor allele frequency of the current target SNP site and the minor allele frequency of the next target SNP site is used as the abundance difference;
[0154] Averaging processing is performed on at least one abundance difference to obtain an abundance fluctuation level of the sample to be detected.
[0155] Based on the above embodiment, optionally, the genotype determination module 440 is specifically configured to:
[0156] Calculating a second abundance threshold based on the first abundance threshold; wherein the sum of the first abundance threshold and the second abundance threshold is one;
[0157] For each SNP site to be tested, if the mutation abundance corresponding to the SNP site to be tested is less than or equal to the first abundance threshold, the genotype of the SNP site to be tested is wild type;
[0158] If the mutation abundance corresponding to the SNP site to be tested is greater than or equal to the second abundance threshold, the genotype of the SNP site to be tested is homozygous;
[0159] If the mutation abundance corresponding to the SNP site to be tested is greater than the first abundance threshold and less than the second abundance threshold, the genotype of the SNP site to be tested is heterozygous.
[0160] Based on the above embodiment, optionally, the SNP site acquisition module 410 to be tested is specifically configured to:
[0161] At least 10 SNP sites in the sample to be tested whose population frequencies meet the preset population frequency range are respectively used as SNP sites to be tested; wherein, the preset population frequency range includes 0.4-0.6.
[0162] The genotype detection device provided in the embodiment of the present invention can execute the genotype detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0163] Example 5
[0164] Figure 5 FIG. 1 is a schematic diagram of the structure of a sample contamination detection device according to the fifth embodiment of the present invention. Figure 5 As shown, the device includes: a genotype determination module 510, a Gaussian mixture model construction module 520 and a pollution status data determination module 530.
[0165] The genotype determination module 510 is used to obtain at least 10 SNP sites to be detected in the sample to be detected, and determine the genotype corresponding to each SNP site to be detected;
[0166] A Gaussian mixture model construction module 520 is configured to construct a Gaussian mixture model based on the minor allele frequencies corresponding to the SNP sites to be tested that meet genotype conditions; wherein the genotype conditions include a wild type genotype and / or a homozygous genotype;
[0167] The contamination status data determination module 530 is used to perform an optimization solution operation on the Gaussian mixture model to obtain the contamination status data corresponding to the sample to be tested; wherein the contamination status data includes at least one of the contamination ratio, the variance of the minor allele frequency and the number of contamination sources.
[0168] The technical solution of this embodiment obtains at least 10 SNP sites to be tested in the sample to be tested, determines the genotype corresponding to each SNP site to be tested, constructs a Gaussian mixture model based on the minor allele frequency corresponding to the non-heterozygous SNP site to be tested, and performs an optimization solution operation on the Gaussian mixture model to obtain contamination status data corresponding to the sample to be tested. This solves the problem that existing sample contamination measurement methods are greatly affected by copy number variation / loss of heterozygosity and have difficulty in detecting low-level contamination, thereby improving the accuracy of sample contamination measurement results.
[0169] Based on the above embodiment, optionally, the Gaussian mixture model satisfies the formula:
[0170]
[0171] in,
[0172]
[0173] in,
[0174]
[0175] Where maf represents the minor allele frequency, δ 2 represents the variance of the minor allele frequency, n represents the number of pollution sources, α represents the pollution ratio, pbinom represents the Bernoulli probability distribution, P(C=i) represents the probability distribution corresponding to the i pollution sources, and N represents the probability distribution of the Gaussian mixture model.
[0176] Based on the above embodiment, optionally, the genotype determination module 510 includes:
[0177] A first abundance threshold acquisition unit is used to obtain a first abundance threshold and determine the mutation abundance corresponding to each SNP site to be tested;
[0178] The genotype determination unit is used to determine the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested.
[0179] Based on the above embodiment, optionally, the genotype determination unit is specifically configured to:
[0180] Calculating a second abundance threshold based on the first abundance threshold; wherein the sum of the first abundance threshold and the second abundance threshold is one;
[0181] For each SNP site to be tested, if the mutation abundance corresponding to the SNP site to be tested is less than or equal to the first abundance threshold, the genotype of the SNP site to be tested is wild type;
[0182] If the mutation abundance corresponding to the SNP site to be tested is greater than or equal to the second abundance threshold, the genotype of the SNP site to be tested is homozygous;
[0183] If the mutation abundance corresponding to the SNP site to be tested is greater than the first abundance threshold and less than the second abundance threshold, the genotype of the SNP site to be tested is heterozygous.
[0184] Based on the above embodiment, optionally, the first abundance threshold obtaining unit includes:
[0185] The abundance threshold value acquisition subunit to be measured is used to obtain at least three abundance threshold values to be measured;
[0186] an abundance fluctuation level determination subunit, configured to calculate, for each abundance threshold to be measured, the abundance fluctuation level of the sample to be tested based on the abundance threshold to be measured and the minor allele frequency corresponding to each SNP site to be measured;
[0187] The first abundance threshold determination subunit is configured to use the abundance threshold to be measured corresponding to the abundance fluctuation level that meets the inflection point characteristics among the at least three abundance fluctuation levels as the first abundance threshold
[0188] Based on the above embodiment, optionally, the abundance fluctuation level determination subunit is specifically configured to:
[0189] For each chromosome where each SNP site to be tested is located, a segmentation operation is performed on the chromosome based on the copy numbers corresponding to at least two SNP sites to be tested on the chromosome to obtain one or more chromosome segments; wherein the copy number difference corresponding to any two SNP sites to be tested in each chromosome segment is less than a preset difference threshold;
[0190] For each chromosome segment, the SNP site to be tested in which the minor allele frequency in the chromosome segment is greater than the abundance threshold to be tested is taken as the target SNP site;
[0191] Based on the minor allele frequencies corresponding to at least two target SNP sites, the abundance fluctuation level of the sample to be tested is calculated.
[0192] Based on the above embodiment, optionally, the abundance fluctuation level determination subunit is specifically configured to:
[0193] Sort each target SNP site based on its location in the chromosome;
[0194] Based on the sorting results, the difference between the minor allele frequency of the current target SNP site and the minor allele frequency of the next target SNP site is used as the abundance difference;
[0195] Averaging processing is performed on at least one abundance difference to obtain an abundance fluctuation level of the sample to be detected.
[0196] Based on the above embodiment, optionally, the genotype determination module 510 includes:
[0197] The SNP site determination unit to be tested is used to select at least 10 SNP sites in the sample to be tested whose population frequencies meet the preset population frequency range as the SNP sites to be tested; wherein, the preset population frequency range includes 0.4-0.6.
[0198] The sample contamination detection device provided in the embodiment of the present invention can execute the sample contamination detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0199] Example 6
[0200] Figure 6 1 is a structural diagram of an electronic device provided according to embodiment six of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0201] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0202] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0203] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the genotype detection method or the sample contamination detection method.
[0204] In some embodiments, the genotype detection method or the sample contamination detection method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the genotype detection method or the sample contamination detection method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the genotype detection method or the sample contamination detection method in any other appropriate manner (e.g., by means of firmware).
[0205] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0206] Computer programs for implementing the genotype detection methods or sample contamination detection methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that, when executed by the processor, the computer programs implement the functions / operations specified in the flowcharts and / or block diagrams. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0207] Example 7
[0208] Embodiment 7 of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a processor to execute a genotype detection method, the method comprising:
[0209] Obtain at least 10 SNP sites to be tested in the sample to be tested and obtain at least 3 abundance thresholds to be tested;
[0210] For each abundance threshold to be tested, the abundance fluctuation level of the sample to be tested is calculated based on the abundance threshold to be tested and the minor allele frequency corresponding to each SNP site to be tested;
[0211] The abundance threshold to be measured corresponding to the abundance fluctuation level that meets the inflection point characteristics among at least three abundance fluctuation levels is used as the first abundance threshold;
[0212] Based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested, the genotype corresponding to each SNP site to be tested is determined.
[0213] Alternatively, the computer instructions are used to cause a processor to perform a sample contamination detection method, the method comprising:
[0214] Obtain at least 10 SNP sites to be detected in the sample to be tested, and determine the genotype corresponding to each SNP site to be detected;
[0215] Constructing a Gaussian mixture model based on the minor allele frequency corresponding to the SNP site to be tested that meets the genotype conditions; wherein the genotype conditions include the genotype being wild type and / or the genotype being homozygous;
[0216] An optimization solution operation is performed on the Gaussian mixture model to obtain contamination status data corresponding to the sample to be tested; wherein the contamination status data includes at least one of a contamination ratio, a minor allele frequency variance, and the number of contamination sources.
[0217] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0218] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0219] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0220] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0221] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0222] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A genotype detection method, characterized in that: include: Obtain at least 10 SNP sites to be tested in the sample to be tested and obtain at least 3 abundance thresholds to be tested; For each abundance threshold to be tested, the abundance fluctuation level of the sample to be tested is calculated based on the abundance threshold to be tested and the minor allele frequency corresponding to each SNP site to be tested; The abundance threshold to be measured corresponding to the abundance fluctuation level that meets the inflection point characteristics among at least three abundance fluctuation levels is used as the first abundance threshold; Determining the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested; The calculating the abundance fluctuation level of the sample to be tested based on the abundance threshold to be tested and the minor allele frequency corresponding to each SNP site to be tested comprises: For each chromosome where each SNP site to be tested is located, a segmentation operation is performed on the chromosome based on the copy numbers corresponding to at least two SNP sites to be tested on the chromosome to obtain one or more chromosome segments; wherein the copy number difference corresponding to any two SNP sites to be tested in each chromosome segment is less than a preset difference threshold; For each chromosome segment, the SNP site to be tested in which the minor allele frequency in the chromosome segment is greater than the abundance threshold to be tested is taken as the target SNP site; Based on the minor allele frequencies corresponding to at least two target SNP sites, the abundance fluctuation level of the sample to be tested is calculated.
2. The method according to claim 1, characterized in that The calculating the abundance fluctuation level of the sample to be tested based on the minor allele frequencies corresponding to the at least two target SNP sites respectively includes: Sort each target SNP site based on its location in the chromosome; Based on the sorting results, the difference between the minor allele frequency of the current target SNP site and the minor allele frequency of the next target SNP site is used as the abundance difference; Averaging processing is performed on at least one abundance difference to obtain an abundance fluctuation level of the sample to be detected.
3. The method according to claim 1, characterized in that The determining of the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested comprises: Calculating a second abundance threshold based on the first abundance threshold; wherein the sum of the first abundance threshold and the second abundance threshold is one; For each SNP site to be tested, if the mutation abundance corresponding to the SNP site to be tested is less than or equal to the first abundance threshold, the genotype of the SNP site to be tested is wild type; If the mutation abundance corresponding to the SNP site to be tested is greater than or equal to the second abundance threshold, the genotype of the SNP site to be tested is homozygous; If the mutation abundance corresponding to the SNP site to be tested is greater than the first abundance threshold and less than the second abundance threshold, the genotype of the SNP site to be tested is heterozygous.
4. The method according to claim 1, wherein The step of obtaining at least 10 SNP sites to be detected in the sample to be detected comprises: At least 10 SNP sites in the sample to be tested whose population frequencies meet the preset population frequency range are respectively used as SNP sites to be tested; wherein, the preset population frequency range includes 0.4-0.
6.
5. A method for detecting sample contamination, characterized in that: include: Obtain at least 10 SNP sites to be detected in the sample to be tested, and determine the genotype corresponding to each SNP site to be detected; Constructing a Gaussian mixture model based on the minor allele frequency corresponding to the SNP site to be tested that meets the genotype conditions; wherein the genotype conditions include the genotype being wild type and / or the genotype being homozygous; Performing an optimization solution operation on the Gaussian mixture model to obtain pollution status data corresponding to the sample to be detected; Wherein, the pollution status data includes at least one of pollution ratio, minor allele frequency variance and number of pollution sources; The Gaussian mixture model satisfies the formula: in, in, Where maf represents the minor allele frequency, δ 2 represents the variance of the minor allele frequency, n represents the number of pollution sources, α represents the pollution ratio, pbinom represents the Bernoulli probability distribution, P(C=i) represents the probability distribution corresponding to the i pollution sources, and N represents the probability distribution of the Gaussian mixture model; Determining the genotype corresponding to each SNP site to be tested includes: Obtaining a first abundance threshold, and determining the mutation abundance corresponding to each SNP site to be tested; Determining the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested; The obtaining of the first abundance threshold comprises: Obtain at least three abundance thresholds to be measured; For each abundance threshold to be tested, the abundance fluctuation level of the sample to be tested is calculated based on the abundance threshold to be tested and the minor allele frequency corresponding to each SNP site to be tested; The abundance threshold to be measured corresponding to the abundance fluctuation level that meets the inflection point characteristics among at least three abundance fluctuation levels is used as the first abundance threshold; The calculating the abundance fluctuation level of the sample to be tested based on the abundance threshold to be tested and the minor allele frequency corresponding to each SNP site to be tested comprises: For each chromosome where each SNP site to be tested is located, a segmentation operation is performed on the chromosome based on the copy numbers corresponding to at least two SNP sites to be tested on the chromosome to obtain one or more chromosome segments; wherein the copy number difference corresponding to any two SNP sites to be tested in each chromosome segment is less than a preset difference threshold; For each chromosome segment, the SNP site to be tested in which the minor allele frequency in the chromosome segment is greater than the abundance threshold to be tested is taken as the target SNP site; Based on the minor allele frequencies corresponding to at least two target SNP sites, the abundance fluctuation level of the sample to be tested is calculated.
6. The method according to claim 5, characterized in that The step of determining the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested comprises: Calculating a second abundance threshold based on the first abundance threshold; wherein the sum of the first abundance threshold and the second abundance threshold is one; For each SNP site to be tested, if the mutation abundance corresponding to the SNP site to be tested is less than or equal to the first abundance threshold, the genotype of the SNP site to be tested is wild type; If the mutation abundance corresponding to the SNP site to be tested is greater than or equal to the second abundance threshold, the genotype of the SNP site to be tested is homozygous; If the mutation abundance corresponding to the SNP site to be tested is greater than the first abundance threshold and less than the second abundance threshold, the genotype of the SNP site to be tested is heterozygous.
7. The method according to claim 5, characterized in that The calculating the abundance fluctuation level of the sample to be tested based on the minor allele frequencies corresponding to the at least two target SNP sites respectively includes: Sort each target SNP site based on its location in the chromosome; Based on the sorting results, the difference between the minor allele frequency of the current target SNP site and the minor allele frequency of the next target SNP site is used as the abundance difference; Averaging processing is performed on at least one abundance difference to obtain an abundance fluctuation level of the sample to be detected.
8. The method according to claim 5, characterized in that The step of obtaining at least 10 SNP sites to be detected in the sample to be detected comprises: At least 10 SNP sites in the sample to be tested whose population frequencies meet the preset population frequency range are respectively used as SNP sites to be tested; wherein, the preset population frequency range includes 0.4-0.
6.
9. A genotype detection device, characterized in that: include: A SNP site acquisition module to be tested is used to obtain at least 10 SNP sites to be tested in the sample to be tested and obtain at least 3 abundance thresholds to be tested; an abundance fluctuation level determination module, configured to calculate, for each abundance threshold to be measured, the abundance fluctuation level of the sample to be tested based on the abundance threshold to be measured and the minor allele frequency corresponding to each SNP site to be measured; a first abundance threshold determination module, configured to use the abundance threshold to be measured corresponding to the abundance fluctuation level that meets the inflection point feature among the at least three abundance fluctuation levels as the first abundance threshold; a genotype determination module, configured to determine the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested; The abundance fluctuation level determination module comprises: a chromosome segment determination unit configured to, for each chromosome where each SNP site to be tested is located, perform a segmentation operation on the chromosome based on the copy numbers corresponding to at least two SNP sites to be tested on the chromosome, thereby obtaining one or more chromosome segments; wherein the copy number difference between any two SNP sites to be tested in each chromosome segment is less than a preset difference threshold; a target SNP site determination unit, configured to, for each chromosome segment, determine a SNP site to be tested in the chromosome segment whose minor allele frequency is greater than the abundance threshold to be tested as a target SNP site; The abundance fluctuation level unit is used to calculate the abundance fluctuation level of the sample to be tested based on the minor allele frequencies corresponding to at least two target SNP sites.
10. A sample contamination detection device, characterized in that: include: The genotype determination module is used to obtain at least 10 SNP sites to be detected in the sample to be detected and determine the genotype corresponding to each SNP site to be detected; A Gaussian mixture model construction module is used to construct a Gaussian mixture model based on the minor allele frequency corresponding to the SNP site to be tested that meets the genotype conditions; wherein the genotype conditions include the genotype being wild type and / or the genotype being homozygous; A pollution status data determination module is used to perform an optimization solution operation on the Gaussian mixture model to obtain pollution status data corresponding to the sample to be detected; Wherein, the pollution status data includes pollution ratio, minor allele frequency variance and number of pollution sources; The Gaussian mixture model satisfies the formula: in, in, Where maf represents the minor allele frequency, δ 2 represents the variance of the minor allele frequency, n represents the number of pollution sources, α represents the pollution ratio, pbinom represents the Bernoulli probability distribution, P(C=i) represents the probability distribution corresponding to the i pollution sources, and N represents the probability distribution of the Gaussian mixture model; The genotype determination module comprises: A first abundance threshold acquisition unit is used to obtain a first abundance threshold and determine the mutation abundance corresponding to each SNP site to be tested; a genotype determination unit, configured to determine the genotype corresponding to each SNP site to be tested based on the first abundance threshold and the mutation abundance corresponding to each SNP site to be tested; The first abundance threshold acquisition unit includes: The abundance threshold value acquisition subunit to be measured is used to obtain at least three abundance threshold values to be measured; an abundance fluctuation level determination subunit, configured to calculate, for each abundance threshold to be measured, the abundance fluctuation level of the sample to be tested based on the abundance threshold to be measured and the minor allele frequency corresponding to each SNP site to be measured; a first abundance threshold determination subunit, configured to use the abundance threshold to be measured corresponding to the abundance fluctuation level that meets the inflection point feature among the at least three abundance fluctuation levels as the first abundance threshold; The abundance fluctuation level determination subunit is specifically used to: For each chromosome where each SNP site to be tested is located, a segmentation operation is performed on the chromosome based on the copy numbers corresponding to at least two SNP sites to be tested on the chromosome to obtain one or more chromosome segments; wherein the copy number difference corresponding to any two SNP sites to be tested in each chromosome segment is less than a preset difference threshold; For each chromosome segment, the SNP site to be tested in which the minor allele frequency in the chromosome segment is greater than the abundance threshold to be tested is taken as the target SNP site; Based on the minor allele frequencies corresponding to at least two target SNP sites, the abundance fluctuation level of the sample to be tested is calculated.
11. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can implement the genotype detection method described in any one of claims 1 to 4 or the sample contamination detection method described in any one of claims 5 to 8 when executed.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the genotype detection method described in any one of claims 1 to 4 or the sample contamination detection method described in any one of claims 5 to 8 when executed.
Citation Information
Patent Citations
Screening method of SNP (Single Nucleotide Polymorphism) sites for detecting pollution level of sample and detection method of pollution level of sample
CN114530198A
Individual single nucleotide polymorphisms locus genotyping method and apparatus
WO2016078067A1