Method and apparatus for determining sample contamination, electronic device, and storage device
By analyzing the number of crossings and likelihood ratios of autosomal disomes and trisomies, and combining Zscore values to determine sample contamination, the low cost and low complexity of sample contamination detection in high-throughput sequencing is solved, and accurate pollution detection is achieved at low sequencing depth.
Patent Information
- Application Number
- PCT/CN2025/073593
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-08
- Filing Date
- 2025-01-21
- Publication Date
- 2025-08-14
AI Technical Summary
In existing high-throughput sequencing technologies, sample contamination problems are difficult to effectively detect, especially at low cost and low complexity, which affects the accuracy and reliability of sequencing results.
By analyzing the number of crossings between autosomal disomes and trisomals in the sample, using autosomal trisomal likelihood ratio and monomer likelihood ratio, we can determine whether there is contamination in the sample. Combined with the standard data on the ploidy difference fragment crossing times of the reference sample, the Zscore value is calculated to determine the contamination.
It realizes low-cost and low-complex sample pollution detection, compatible with a variety of library construction methods and sequencing methods, and can accurately determine whether the sample has pollution at low sequencing depth, save sequencing costs and improve detection efficiency.
Smart Images

Figure CN2025073593_14082025_PF_FP_ABST
Abstract
Description
Method, device, electronic device and storage device for determining sample contamination Technical Field
[0001] The present invention belongs to the field of bioinformatics, and specifically relates to an analysis method, device, electronic device and computer-readable storage medium for judging sample contamination based on the number of autosomal disomy and trisomy crossovers. Background Art
[0002] In recent years, high-throughput sequencing (NGS) has become an essential tool for genomic research, widely recognized for its efficiency and accuracy. However, as the application of high-throughput sequencing technology expands, the problem of sample contamination has become increasingly prominent. Sample contamination can arise from cross-contamination with exogenous DNA or RNA molecules, seriously compromising the accuracy and reliability of sequencing results.
[0003] Short tandem repeat (STR) analysis is a commonly used method for contamination monitoring. STRs are composed of tandem repeats of 2-7 base pairs as core units, and the number of repeats varies between individuals. Therefore, by analyzing the differences in the number of repeats in these highly polymorphic STR regions among individuals, it can be used to determine contamination between samples. STR analysis is used in prenatal diagnosis to exclude maternal contamination by comparing loci in the test sample with those in the maternal blood. Other methods for detecting contamination include analyzing the ratio of homozygous loci between reference and analysis samples and analyzing the frequency of minor alleles within the sample.
[0004] Conventional STR methods require paired samples for comparison, thus requiring additional sampling. This invention establishes a new, low-cost, low-complexity method for detecting contamination based on the differences in likelihood ratio plots between autosomal disomy, trisomy, and contaminated samples. Summary of the Invention
[0005] The present invention provides an analytical method, device, electronic device and computer-readable storage medium for judging sample contamination based on the number of autosomal disomy and trisomy crossovers. It can realize low-cost and low-complexity contamination detection for samples, and is compatible with multiple library construction methods, multiple sequencing methods, and low-depth sequencing data.
[0006] In a first aspect, the present invention provides a method for detecting sample contamination, comprising: obtaining sequencing information obtained after sequencing a sample to be analyzed, the sequencing information including filtered reads of the sample to be analyzed; determining an autosomal trisomy likelihood ratio of the sample to be analyzed based on the filtered reads; determining the number of autosomal ploidy difference fragment crossovers of the sample to be analyzed based on the autosomal trisomy likelihood ratio; and determining whether the sample to be analyzed is contaminated based on the number of autosomal ploidy difference fragment crossovers of the sample to be analyzed.
[0007] In some optional embodiments, the above-mentioned determination of the autosomal ploidy of the sample to be analyzed based on the autosomal monosomal likelihood ratio includes: determining the mean of the autosomal monosomal likelihood ratio of each autosomal sample to be analyzed according to the positive and negative value relationship of the autosomal monosomal likelihood ratios between each adjacent window in the sample to be analyzed: in response to the sum of the monosomal likelihood ratios of each autosomal window divided by the number of windows, if the mean of the monosomal likelihood ratio of each autosomal sample to be analyzed is >0, it proves that the chromosome is a monomer; if the mean is <0, it proves that the chromosome is disomic; comprehensively considering the mean of all autosomal monosomal likelihood ratios, if the mean of all chromosome monosomal likelihood ratios is <0, it proves that the sample to be analyzed is non-haploid, and is diploid by default.
[0008] In some optional embodiments, the above-mentioned determination of the number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed based on the autosomal trisomy likelihood ratio includes: determining the number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed according to the positive and negative value relationship of the autosomal trisomy likelihood ratios between each adjacent window in the sample to be analyzed; in response to the opposite positive and negative values of the autosomal trisomy likelihood ratio in the front window and the autosomal trisomy likelihood ratio in the rear window, determining that there is one crossover in the sample to be analyzed.
[0009] In some optional embodiments, before determining whether the sample to be analyzed is contaminated based on the number of autosomal ploidy difference segment crossovers of the sample to be analyzed, the method further includes: determining the standard data of the number of autosomal ploidy difference segment crossovers corresponding to the reference sample, the reference sample includes a known diploid sample and a known triploid sample, and the standard data of the number of autosomal ploidy difference segment crossovers corresponding to the reference sample includes the mean and standard deviation of the number of autosomal ploidy difference segment crossovers corresponding to the reference sample.
[0010] In some optional embodiments, determining whether the sample to be analyzed is contaminated is based on the number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed, including: determining the Zscore value corresponding to the sample to be analyzed based on the mean and standard deviation of the number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed and the number of crossovers of the autosomal ploidy difference segments corresponding to the reference sample, wherein the Zscore value includes a Z2score value corresponding to the number of crossovers of the ploidy difference segments on the diploid autosomes and a Z3score value corresponding to the number of crossovers of the ploidy difference segments on the triploid autosomes; determining whether the sample to be analyzed is contaminated is based on the Zscore value corresponding to the sample to be analyzed and / or the number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed.
[0011] In some optional embodiments, the above-mentioned determination of whether the sample to be analyzed is contaminated based on the Zscore value corresponding to the sample to be analyzed and / or the number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed further includes: in response to (1) the Z2score value is greater than a first preset threshold and the Z3score value is less than a second preset threshold, or (2) the number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed is less than a third preset threshold and all the autosomes of the sample to be analyzed are three different chromosomes, determining that the sample to be analyzed is contaminated.
[0012] In some optional embodiments, before obtaining the sequencing information obtained after sequencing the sample to be analyzed, it also includes: obtaining a reference haplotype data set, wherein the reference haplotype data set includes chromosome numbers, position information of SNP sites and / or genotype information of SNP sites.
[0013] In some optional embodiments, the read lengths after filtering are used to determine the autosomal trisomy likelihood ratios of the samples to be analyzed, including: determining a window (bin) on the autosomal genome of the sample to be analyzed according to a preset base length; for each window, determining the read length containing the preset haplotype variation SNP in the window according to the information of the preset haplotype variation SNP in the reference haplotype data set; according to the differences in the allele probabilities of the read lengths containing the preset haplotype variation SNP under different ploidy assumptions, in response to the assumption that the read length pair comes from different numbers of chromosomes, calculating the corresponding allele frequency of the read length pair, wherein, if it is assumed that the read length pair comes from two different chromosomes, the corresponding allele frequency P2 of the read length pair is calculated, and if it is assumed that the read length comes from three different chromosomes, the corresponding allele frequency P3 of the read length pair is calculated; based on the allele frequency P2 and the allele frequency P3, determining the autosomal trisomy likelihood ratio LR3 of the window under the different ploidy assumptions.
[0014] In some optional embodiments, the above-mentioned determination of the number of crossovers of the autosomal ploidy difference fragments of the sample to be analyzed based on the autosomal trisomy likelihood ratio includes: merging the windows (bins) on the autosomal genome of the sample to be analyzed according to the characteristics of the autosomal ploidy difference fragments in the sample to be analyzed to obtain a large window after the sample to be analyzed is merged, and the large window includes a large window corresponding to an identical chromosome, a large window corresponding to two different chromosomes and / or a large window corresponding to three different chromosomes, wherein the characteristics include that the window is a monomeric fragment, the window is a disomic fragment or the window is a trisomic fragment; according to the large window after the sample to be analyzed is merged, the number of crossovers of the autosomal ploidy difference fragments of the sample to be analyzed is determined.
[0015] In some optional embodiments, before determining the autosomal trisomy likelihood ratio of the sample to be analyzed, the method further includes: determining the autosomal monosomy likelihood ratio of the sample to be analyzed based on the filtered read length, and screening to obtain non-monosomal samples to be analyzed.
[0016] In some optional embodiments, the above-mentioned determination of the autosomal monomer likelihood ratio of the sample to be analyzed and screening to obtain a non-monomeric sample to be analyzed includes: determining a window (bin) on the autosomal genome of the sample to be analyzed according to a preset base length; for each window, determining the read length containing the preset haplotype variation SNP in the window according to the information of the preset haplotype variation SNP in the reference haplotype data set; according to the difference in the allele probabilities of the read lengths of each read length containing the preset haplotype variation SNP under different ploidy assumptions, in response to the assumption that the read length or read length pair comes from different numbers of Chromosome, calculate the corresponding allele frequency of the read length or read length pair, wherein, if it is assumed that the read length pair comes from the same chromosome, calculate the corresponding allele frequency P1 of the read length pair, if it is assumed that the read length pair comes from two different chromosomes, calculate the corresponding allele frequency P2 of the read length pair; according to the allele frequency P1 and the allele frequency P2, determine the autosomal monomer likelihood ratio LR1 of the window under the different ploidy assumptions, if it is assumed that the read length pair comes from three different chromosomes, calculate the corresponding allele frequency P3 of the read length pair; according to the autosomal monomer likelihood ratio LR1, determine that the sample to be analyzed is non-monosomal.
[0017] In a second aspect, the present invention provides a device for detecting sample contamination, comprising: an acquisition module configured to acquire sequencing information obtained after sequencing the sample to be analyzed, the sequencing information including the filtered reads of the sample to be analyzed; a likelihood ratio determination module configured to determine the autosomal trisomy likelihood ratios of the sample to be analyzed based on the filtered reads; a crossover number determination module configured to determine the number of crossovers of the autosomal ploidy difference fragments of the sample to be analyzed based on the autosomal trisomy likelihood ratios; and a judgment module configured to determine whether the sample to be analyzed is contaminated based on the number of crossovers of the autosomal ploidy difference fragments of the sample to be analyzed.
[0018] In a third aspect, the present invention provides an electronic device comprising: one or more processors; a storage device storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement a method as described in any one of the above aspects.
[0019] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by one or more processors, implements the method as described in any of the above aspects.
[0020] The method of the present invention has the following beneficial effects:
[0021] 1. Efficient determination of sample contamination: This method uses the distribution of autosomal ploidy difference fragments to determine whether a sample is contaminated. Compared to traditional contamination detection methods, this method is a novel contamination detection method that accurately indicates contamination without relying on other samples or methods, thus offering significant economic and time advantages.
[0022] 2. Compatible with low-depth sequencing data: Data indicates that contamination detection can be achieved with data as low as 0.06X. This reduces sequencing costs. In specific scenarios, low-depth sequencing data can be used to determine whether a sample is contaminated, providing a basis for deciding whether to proceed with high-depth sequencing, enabling low-cost pre-judgment of sample contamination.
[0023] 3. Compatible with multiple library construction methods: This method does not require a special library construction method. It is compatible with common library construction methods such as whole genome library construction, degenerate genome library construction, and capture library construction, and can be easily applied in various experimental processes.
[0024] 4. Compatible with multiple gene sequencing technologies and highly versatile: such as chips, high-depth sequencing, and degenerate genome sequencing. This makes this method widely applicable to genetic research and clinical diagnosis in different fields.
[0025] 5. Improving Accuracy by Utilizing Linkage Disequilibrium: This method leverages the haplotype reference database and the principle of linkage disequilibrium to further improve the accuracy of low-depth sequencing. This allows for reliable results even at low sequencing depths. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the specification and, together with the description, serve to explain the principles of the specification.
[0027] FIG1 shows a schematic diagram of an implementation environment in which an embodiment of the present invention may be applied.
[0028] FIG. 2A shows a flow chart of an embodiment of a method for determining sample contamination according to the present invention.
[0029] FIG. 2B shows a partial flow chart of an embodiment of a method for determining sample contamination according to the present invention.
[0030] FIG3 shows a schematic structural diagram of an embodiment of a device for determining sample contamination according to the present invention.
[0031] FIG4 shows a schematic structural diagram of a computer system of an electronic device suitable for implementing an embodiment of the present invention.
[0032] FIG5 shows a comparison of the number of crossovers among ploidy difference segments on autosomes of diploid, triploid and contaminated samples.
[0033] FIG6 shows the results of detecting normal diploid using the method of the present invention.
[0034] FIG7 shows the results of detecting normal triploid using the method of the present invention.
[0035] FIG8 shows the result of detecting the contaminated sample 1 using the method of the present invention.
[0036] FIG. 9 shows the results of detecting the contaminated sample 2 using the method of the present invention. DETAILED DESCRIPTION
[0037] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the invention are shown in the accompanying drawings.
[0038] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0039] FIG1 is a schematic diagram showing an implementation environment of the analysis method, device, electronic device, and storage medium for determining sample contamination according to the present invention.
[0040] As shown in Figure 1, the implementation environment includes an electronic device 100. The analysis method for determining sample contamination in the embodiment of the present invention can be executed by terminal devices 101, 102, 103, and 104. For example, the electronic device 100 can include at least one of a terminal device and a server.
[0041] Terminal devices 101, 102, 103, 104 can be hardware or software. When terminal devices 101, 102, 103, 104 are hardware, they can be various electronic devices with a display screen and support information input (such as text input and / or voice input, etc.), including but not limited to smart phones, tablet computers, laptop computers and desktop computers, etc. When terminal devices 101, 102, 103, 104 are software, they can be installed in the terminal devices listed above. It can be implemented as multiple software or software modules (for example, for providing services for judging sample contamination), or it can be implemented as a single software or software module. It is not specifically limited here.
[0042] It should be understood that the number of terminal devices in FIG1 is merely illustrative and any number of terminal devices may be provided according to implementation requirements.
[0043] 2 , which shows a process 200 of an embodiment of an analysis method for determining sample contamination according to the present invention. The analysis method for determining sample contamination includes the following steps:
[0044] Step 201: Acquire sequencing information obtained after sequencing the sample to be analyzed.
[0045] In this embodiment, the execution subject of the analysis method (eg, the terminal device shown in FIG1 ) may first obtain sequencing information obtained after sequencing the sample to be analyzed.
[0046] Here, the above sequencing information includes the filtered read length (reads) of the sample to be analyzed. Read length refers to the sequence fragment of a specific length that can be measured by the sequencing reaction. The read length of the first-generation sequencing dideoxy chain termination method (Sanger method) is a short sequence fragment of about 1000bp, the second-generation sequencing is a short sequence fragment of 50bp-300bp, and the third generation can reach a long sequence fragment of more than 5000bp. The specific length of the read length is not limited in this embodiment. It should be understood that the method of the present invention can be applied to sequence fragments of various read lengths.
[0047] The method of the present invention can be applied to a variety of different sequencing methods, including but not limited to first-generation, second-generation, and third-generation sequencing technologies, as well as microarrays, high-depth sequencing, and degenerate genome sequencing. Therefore, the method of the present invention is highly versatile and compatible with a variety of gene sequencing technologies, and can be widely used in genetic research and clinical diagnosis in different fields. Similarly, the method of the present invention is compatible with low sequencing data. Data show that 0.06x data can be used to achieve contamination detection using this method. This means that sequencing costs can be saved.
[0048] In this embodiment, the filtered read length refers to the sequencing result data that has been subjected to post-sequencing quality control. In some optional embodiments, the method of the present invention may further include the following steps a-c of performing quality control on the sequencing read length after sequencing:
[0049] Step a: DNA extraction, library construction, and sequencing. The library construction method can be whole genome library construction, degenerate genome library construction, capture library construction, etc.
[0050] Step b: filter the raw sequencing data to remove low-quality bases and adapters.
[0051] Step c: Align to the human reference genome and remove repetitive sequences and low alignment quality sequences.
[0052] Step 202: Based on the filtered read lengths, determine the autosomal trisomy likelihood ratios of the samples to be analyzed.
[0053] In this embodiment, the execution subject of the analysis method has obtained the read length of the sample to be analyzed after sequencing and filtering in step 201, and calculated the probability of the sample to be analyzed being trisomic and disomic respectively. The allele probability of reads extracted from the same haploid type is different under different ploidy assumptions. If a pair of reads comes from the same chromosome, the frequency of their related alleles is observed to be P1=f(AB). If a pair of reads comes from two different chromosomes, the frequency of their related alleles is observed to be If a pair of reads comes from three different chromosomes, the probability of observing their related alleles is f(A) and f(B) represent the frequencies of the two alleles A and B in the reference haplotype data set, respectively.
[0054] In some optional embodiments, at least two reads are randomly sampled, and for each pair of reads, the allele frequencies P1, P2, and P3 observed under three different hypotheses are calculated and averaged. A statistical model is then established for two scenarios: the likelihood ratios LR1 and LR3 (i.e., the likelihood ratio for the entire window) for one chromosome segment and three different chromosome segments.
[0055] In some optional embodiments, the step of calculating the autosomal trisomy likelihood ratio LR3 further includes the following steps a to f:
[0056] Step a: Select filtered reads.
[0057] In step b, the genome is divided into windows of fixed preset length.
[0058] Here, the fixed preset length may be one window every 500 kb. According to the actual application scenario of the method, the window length may be adjusted to 100 kb to 5M.
[0059] Step c: Use 3SD and KS test algorithm to standardize the window position selection.
[0060] In step d, reads that overlap with SNPs that mark common haplotype variations in the reference haplotype dataset are selected within the determined window.
[0061] Exemplarily, the conditions for selecting SNPs are: the ref and alt base lengths are both 1, and the SNP is any one of the four ATCG bases and has a non-zero minor allele number.
[0062] In step e, the allele frequencies P2 from two different chromosomes and the allele frequencies P3 from three different chromosomes are calculated respectively.
[0063] Step f, calculate the likelihood ratio of the two results. The likelihood ratio of this window is three different chromosomes.
[0064] In this embodiment, the executor of the analysis method calculates the probability of the sample to be analyzed being triploid and diploid based on the read length of the sample to be analyzed after sequencing and filtering. In the subsequent steps, the above-mentioned autosomal trisomy likelihood ratio LR3 is used to further calculate the number of autosomal ploidy difference fragment crossovers of the sample to be analyzed.
[0065] Step 203: Determine the number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed based on the autosomal trisomy likelihood ratio.
[0066] Here, the segments formed by drawing the autosomal trisomy likelihood ratio LR3>0 and LR3<0 in the window are called ploidy difference segments, and the number of crossovers formed by the ploidy conversion and the horizontal axis is called the number of crossovers. Based on the autosomal trisomy likelihood ratio, for example, if the autosomal trisomy likelihood ratio LR3>0 in the first window and the autosomal trisomy likelihood ratio LR3<0 in the second window are recorded as the number of crossovers of the autosomal ploidy difference segment is 1.
[0067] Based on the autosomal trisomy likelihood ratio of the sample to be analyzed, the number of crossovers in the autosomal ploidy difference fragments is further determined to judge whether the sample is contaminated. This directly utilizes the difference in the crossover number maps between the sample to be analyzed and the non-contaminated sample, without relying on other reference samples, and has significant advantages in contamination detection efficiency and cost.
[0068] Step 204 , determining whether the sample to be analyzed is contaminated based on the number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed.
[0069] Here, determining whether the sample to be analyzed is contaminated based on the number of crossovers of the autosomal ploidy difference segments in the sample to be analyzed not only includes the case where the number of crossovers of the autosomal ploidy difference segments is used alone for limitation. In this embodiment, it also involves further calculating the Zscore value based on the number of crossovers of the autosomal ploidy difference segments, or combining it with whether the autosomes of the sample to be analyzed are all three different chromosomes for judgment.
[0070] In some optional embodiments, the calculation method of the number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed in the above-mentioned step 203 is specifically as follows: based on the positive and negative value relationship of the autosomal trisomy likelihood ratios between each adjacent window in the sample to be analyzed, the number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed is determined: in response to the opposite positive and negative values of the autosomal trisomy likelihood ratios in the front window and the autosomal trisomy likelihood ratios in the rear window, it is determined that there is one crossover in the sample to be analyzed.
[0071] In this embodiment, the above-mentioned judgment on whether there is crossover in the sample to be analyzed is based on the positive and negative value relationship of the autosomal trisomy likelihood ratio corresponding to the continuous windows in the genome. Exemplarily, when the autosomal trisomy likelihood ratio in the front window is positive (or negative) and the autosomal trisomy likelihood ratio in the rear window is negative (or positive), it is determined that there is one crossover in the sample to be analyzed. Alternatively, for example, if the autosomal trisomy likelihood ratio LR3 in the first window is greater than 0 and the autosomal trisomy likelihood ratio LR3 in the second window is less than 0, then the number of crossovers of the autosomal ploidy difference segment is recorded as 1.
[0072] In some optional implementations, before the above step 204, the process further includes step 203':
[0073] Step 203 ′: determine the standard data of the number of crossovers of the autosomal ploidy difference segments corresponding to the reference sample.
[0074] Here, the reference samples include known diploid samples and known triploid samples, and the standard data of the number of crossovers of the autosomal ploidy difference segments corresponding to the reference samples include the mean and standard deviation of the number of crossovers of the autosomal ploidy difference segments corresponding to the reference samples.
[0075] In this embodiment, the aforementioned known diploid samples and known triploid samples can be known diploid and triploid samples obtained from screened known ploidy samples using low-depth ploidy analysis software and microarray verification. It is understood that the known diploid and triploid samples in this embodiment are not limited to the aforementioned acquisition methods and can be directly obtained from other public databases or public channels.
[0076] In some optional implementations, the above step 203' further includes the following steps a to c:
[0077] Step a: Select a known diploid sample as a reference sample and obtain the mean (μ2) and standard deviation (σ2) of the number of crossovers of ploidy difference fragments on the autosomes.
[0078] Step b: select a known triploid sample as a reference sample and obtain the mean (μ3) and standard deviation (σ3) of the number of crossovers of ploidy difference fragments on the autosomes.
[0079] Step c: Use μ2 and μ3 as standard values and set the confidence interval.
[0080] Here, illustratively, the confidence interval may be 90%, 95%, and preferably, 99.7% is selected as the confidence interval.
[0081] In some optional embodiments, the above step 204 determines whether the sample to be analyzed is contaminated based on the number of autosomal ploidy difference fragment crossovers of the sample to be analyzed, and also includes determining whether the sample to be analyzed is contaminated based on the Zscore value corresponding to the sample to be analyzed.
[0082] In this embodiment, the Zscore value corresponding to the sample to be analyzed is determined based on the mean and standard deviation of the number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed and the number of crossovers of the autosomal ploidy difference segments corresponding to the reference sample, wherein the Zscore value includes a Z2score value corresponding to the number of crossovers of the ploidy difference segments on the diploid autosomes and a Z3score value corresponding to the number of crossovers of the ploidy difference segments on the triploid autosomes.
[0083] Here, the Z2score value and Z3score value are calculated according to the following formulas:
[0084] Among them, Cnum is the number of crossovers of ploidy difference fragments on the autosomes of the sample to be analyzed.
[0085] In some optional embodiments, in response to (1) the Z2score value is greater than a first preset threshold and the Z3score value is less than a second preset threshold, or (2) the number of crossovers of the autosomal ploidy difference fragments of the sample to be analyzed is less than a third preset threshold and all the autosomes of the sample to be analyzed are three different chromosomes, it is determined that the sample to be analyzed is contaminated.
[0086] In this embodiment, the selection of the first preset threshold and the second preset threshold is based on the acceptable range of differences between the observed data points and the mean value in the specific application environment. For example, under a normal distribution, a Zscore value of 3 or -3 indicates that the data point is 3 standard deviations away from the mean value, which means that approximately 99.73% of the data points fall within this range. If the Zscore of a data point is 3, then the difference between the data point and the mean value is relatively large, and it is located at the tail of the normal distribution with a very low probability. The above-mentioned third preset threshold is an empirical value and can be adjusted according to the needs of the specific application environment.
[0087] Here, illustratively, the first preset threshold is 3, the second preset threshold is -3, and the third preset threshold is 6.
[0088] In some optional implementations, step 201′ is further included before step 201:
[0089] Step 201 ′: obtaining a reference haplotype data set.
[0090] The reference haplotype data set includes chromosome number, position information of SNP sites and / or genotype information of SNP sites.
[0091] Here, reference database is a pre-built reference haplotype data set. This reference haplotype data set selects to capture the haplotype data of library construction and chip data construction, and the sequencing depth of these haplotypes is higher, and its genotype has been corrected through family haplotype, and the accuracy of the genotype and haplotype of SNP sites can be guaranteed. The family data quantity selected should be between 1,000 and 3,000. Further, own data is filtered, and the sample with more tripping sites is excluded, and threshold value is set to screen out the accurate sample of haplotype, and self-built haplotype database is constructed, and is used to calculate the autosomal ploidy difference fragment intersection number of times for the reference sample and the sample to be analyzed.
[0092] Here, in order to further improve the accuracy of this method, the reference haplotype data set may include but is not limited to self-built databases and public databases, and haplotype data of more ethnic groups (eg, Chinese) have been added.
[0093] In some optional implementations, the above step 202 further includes the following steps 2021-2024:
[0094] Step 2021 : Determine a bin on the autosomal genome of the sample to be analyzed based on a preset base length.
[0095] Here, the above-mentioned preset base length can be adjusted according to actual conditions. For example, in this embodiment, the preset base length is 500 kb, which can be adjusted to 100 kb to 5M.
[0096] Step 2022: For each window, based on the information of the preset haplotype variation SNP in the reference haplotype data set, determine the read length containing the preset haplotype variation SNP in the window.
[0097] Here, illustratively, the information of the preset haplotype variation SNP is a SNP marked as common in the reference haplotype data set, and after determination, reads overlapping with the SNP marked as common haplotype variation in the self-built haplotype database (or the database after the self-built haplotype database is fused with the public database) are selected in this window. In some embodiments, the conditions for selecting SNPs are: the ref and alt base lengths are both 1, and the SNP is any one of the four bases of ATCG and has a non-zero minor allele number.
[0098] Step 2023, based on the differences in allele probabilities of each read containing the preset haplotype variation SNP under different ploidy assumptions, in response to the assumption that the read pair comes from different numbers of chromosomes, calculate the corresponding allele frequency of the read pair.
[0099] Here, if it is assumed that the read pair comes from two different chromosomes, the corresponding allele frequency P2 of the read pair is calculated. If it is assumed that the read pair comes from three different chromosomes, the corresponding allele frequency P3 of the read pair is calculated.
[0100] Here, the allele frequency Allele frequency Where f(A) and f(B) represent the frequencies of the two alleles A and B in the reference haplotype data set, respectively.
[0101] Step 2024: Determine the autosomal trisomy likelihood ratio LR3 for the window under different ploidy assumptions based on the allele frequency P2 and the allele frequency P3.
[0102] Here, the likelihood ratio for autosomal trisomy is
[0103] In some optional implementations, the above step 2021 further includes step 2021':
[0104] Step 2021': Merge the windows (bins) on the autosomal genome of the sample to be analyzed according to the characteristics of the autosomal ploidy difference fragments in the sample to be analyzed.
[0105] Here, the execution subject of the present analysis method merges the windows to obtain a large window after merging the samples to be analyzed. The large window includes a large window corresponding to one identical chromosome, a large window corresponding to two different chromosomes, and / or a large window corresponding to three different chromosomes, wherein the characteristics include that the window is a monomeric fragment, the window is a trisomic fragment, or the window is a disomic fragment. Exemplarily, if the first window is a trisomic fragment, the second window is a trisomic fragment, the third, fourth, and fifth windows are disomic fragments, and the sixth and seventh windows are monomeric fragments, the first and second windows are merged into a large trisomic window, the third, fourth, and fifth windows are merged into a large disomic window, and the sixth and seventh windows are merged into a large monomeric window.
[0106] In some optional implementations, the above step 203 further includes step 202':
[0107] Step 202 ′: determine the autosomal monosomal likelihood ratio of the sample to be analyzed based on the filtered read length, and screen out non-monosomal samples to be analyzed.
[0108] In this embodiment, after obtaining the filtered read length, the execution subject of this method eliminates the possibility that the sample to be analyzed is a monomer. If the sample to be analyzed is determined to be a monomer, the execution subject no longer executes step 203 and subsequent steps.
[0109] In some optional implementations, the above step 202' further includes the following steps a to e:
[0110] Step a: determining a bin on the autosomal genome of the sample to be analyzed according to a preset base length.
[0111] Here, the above-mentioned preset base length can be adjusted according to actual conditions. For example, in this embodiment, the preset base length is 500 kb, which can be adjusted to 100 kb to 5M.
[0112] Step b: for each window, based on the information of the preset haplotype variation SNP in the reference haplotype data set, determine the read length containing the preset haplotype variation SNP in the window.
[0113] Here, illustratively, the information of the preset haplotype variation SNP is a SNP marked as common in the reference haplotype data set, and after determination, reads overlapping with the SNP marked as common haplotype variation in the self-built haplotype database (or the database after the self-built haplotype database is fused with the public database) are selected in this window. In some embodiments, the conditions for selecting SNPs are: the ref and alt base lengths are both 1, and the SNP is any one of the four bases of ATCG and has a non-zero minor allele number.
[0114] Step c, based on the differences in allele probabilities of each read containing the preset haplotype variation SNP under different ploidy assumptions, in response to the assumption that the read or read pair comes from a different number of chromosomes, calculate the corresponding allele frequency of the read or read pair.
[0115] Here, if it is assumed that the read pair comes from the same chromosome, the corresponding allele frequency P1 of the read pair is calculated. If it is assumed that the read pair comes from two different chromosomes, the corresponding allele frequency P2 of the read pair is calculated. If it is assumed that the read pair comes from three different chromosomes, the corresponding allele frequency P3 of the read pair is calculated.
[0116] Here, the allele frequency is P1=f(AB), and the allele frequency Where f(A) and f(B) represent the frequencies of the two alleles A and B in the reference haplotype data set, respectively.
[0117] Step d: Determine the autosomal monosomal likelihood ratio LR1 of the window under different ploidy assumptions based on the allele frequency P1 and the allele frequency P2.
[0118] Here, the autosomal monosomal likelihood ratio is
[0119] Step e: Determine that the sample to be analyzed is non-monosomal based on the autosomal monosomal likelihood ratio LR1.
[0120] Here, when the autosomal monosomy likelihood ratio LR1 of the sample to be analyzed meets the preset conditions, the sample to be analyzed can be determined to be non-monosomal. For example, if all window segments on the autosome have LR1 < 0, the possibility of haploidy can be ruled out, and the sample to be analyzed is defaulted to be non-haploid.
[0121] In some optional embodiments, the step of calculating the autosomal trisomy likelihood ratio LR3 and the autosomal monosomy likelihood ratio LR1 further includes the following steps a to f:
[0122] Step a: Select filtered reads.
[0123] In step b, the genome is divided into windows of fixed preset length.
[0124] Here, the fixed preset length may be one window every 500 kb. According to the actual application scenario of the method, the window length may be adjusted to 100 kb to 5M.
[0125] Step c: Use 3SD and KS test algorithm to standardize the window position selection.
[0126] In step d, reads that overlap with SNPs that mark common haplotype variations in the reference haplotype dataset are selected within the determined window.
[0127] Exemplarily, the conditions for selecting SNPs are: the ref and alt base lengths are both 1, and the SNP is any one of the four ATCG bases and has a non-zero minor allele number.
[0128] In step e, the allele frequencies P1 from one different chromosome, P2 from two different chromosomes, and P3 from three different chromosomes are calculated respectively.
[0129] Step f, calculate the likelihood ratio of the two results. The likelihood ratio of autosomal monomer is This window is the likelihood ratio of three different chromosomes
[0130] In this embodiment, the execution body of the analysis method calculates the probability that the sample to be analyzed is a monomer based on the read length of the sample to be analyzed after sequencing and filtering, and confirms that the sample to be analyzed is non-haploid (non-haploid samples are diploid by default).
[0131] In some optional embodiments, in subsequent steps, the autosomal trisomy likelihood ratio LR3 obtained above is used to further calculate the number of autosomal ploidy difference segment crossovers of the sample to be analyzed after excluding the possibility that the sample is a monosomy.
[0132] In some optional embodiments, the above-mentioned determination of the autosomal ploidy of the sample to be analyzed based on the autosomal monosomal likelihood ratio LR1 includes: determining the mean of the autosomal monosomal likelihood ratio of each autosomal sample to be analyzed according to the positive and negative value relationship of the autosomal monosomal likelihood ratios between each adjacent window in the sample to be analyzed: in response to the sum of the monomer likelihood ratios of each autosomal window divided by the number of windows, if the mean of the monosomal likelihood ratios of each autosomal sample to be analyzed is >0, it proves that the chromosome is a monomer; if the mean is <0, it proves that the chromosome is disomic; and comprehensively considering the mean of all autosomal monosomal likelihood ratios, if the mean of all chromosome monosomal likelihood ratios is <0, it proves that the sample to be analyzed is non-haploid, and is diploid by default.
[0133] In this embodiment, in order to further verify the accuracy of this method, three families were selected for experiments, with a total of 16 samples. Through sampling, the parental data were mixed with the offspring samples at a ratio of 10%, 20%, 30%, 40%, and 50%. The experimental results show that as long as the parental contamination ratio is greater than or equal to 20%, the method of this application can identify the presence of contamination in the sample. In addition, the experiment was also carried out on non-family samples, which were also mixed at a ratio of 10%, 20%, 30%, 40%, and 50%. The experimental results show that as long as the non-family contamination ratio is greater than or equal to 10%, the method of this embodiment can also identify the presence of contamination in the sample.
[0134] In some optional embodiments, this method can be further applied to the detection of triploids. For example, based on the mean and standard deviation of the crossover times of known reference triploids, if the Z3score of the crossover times of the autosomal ploidy difference segments of the sample to be analyzed is between -3 and 3, the sample to be analyzed is identified as a triploid.
[0135] With further reference to FIG3 , as an implementation of the methods shown in the above figures, the present invention provides an embodiment of a device for determining sample contamination. This device embodiment corresponds to the method embodiment shown in FIG2 , and the device can be specifically applied to various electronic devices.
[0136] As shown in Figure 3, the device 300 for determining sample contamination in this embodiment includes: an acquisition module 301, configured to obtain sequencing information obtained after sequencing the sample to be analyzed, and the sequencing information includes the filtered read length (reads) of the sample to be analyzed; a likelihood ratio determination module 302, configured to determine the autosomal trisomy likelihood ratio of the sample to be analyzed based on the filtered read length; a crossover number determination module 303, configured to determine the number of crossovers of the autosomal ploidy difference fragments of the sample to be analyzed based on the autosomal trisomy likelihood ratio; and a judgment module 304, configured to determine whether the sample to be analyzed is contaminated based on the number of crossovers of the autosomal ploidy difference fragments of the sample to be analyzed.
[0137] In this embodiment, the specific processing of the acquisition module 301, the likelihood ratio determination module 302, the crossover number determination module 303 and the judgment module 304 of the device 300 for judging sample contamination and the technical effects brought about by them can be referred to the relevant descriptions of step 201, step 202, step 203 and step 204 in the corresponding embodiment of Figure 2, and will not be repeated here.
[0138] It should be noted that the implementation details and technical effects of each module in the device for determining sample contamination provided by the embodiment of the present invention can be referred to the description of other embodiments of the present invention, and will not be repeated here.
[0139] As shown in Figure 4, it shows a schematic diagram of the structure of a computer system 400 suitable for implementing the electronic device of the present invention. The computer system 400 shown in Figure 4 is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.
[0140] As shown in FIG4 , a computer system 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. Various programs and data required for the operation of the computer system 400 are also stored in the RAM 403. The processing device 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0141] Typically, the following devices may be connected to I / O interface 405: input device 406 including, for example, a touch screen, touchpad, keyboard, mouse, camera, microphone, etc.; output device 407 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage device 408 including, for example, a magnetic tape, hard disk, etc.; and communication device 409. Communication device 409 may allow computer system 400 to communicate with other devices wirelessly or by wire to exchange data. Although FIG4 illustrates computer system 400 as an electronic device having various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0142] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 408, or installed from ROM 402. When the computer program is executed by processing device 401, the above-mentioned functions defined in the method of the embodiment of the present invention are performed.
[0143] It should be noted that the computer-readable medium described above in the present invention may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present invention, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0144] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0145] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device implements the method for determining sample contamination as shown in the embodiment and its optional implementation manner shown in FIG2 .
[0146] Computer program code for performing the operations of the present invention may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0147] The flow charts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0148] The units involved in the embodiments of the present invention may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.
[0149] The above description is merely a preferred embodiment of the present invention and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present invention is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the present invention.
[0150] Exemplarily, the following illustrates the process and effect embodiments of the method of the present invention in specific application scenarios.
[0151] Example
[0152] Example 1:
[0153] This embodiment uses the above method to detect normal diploids, including the following steps:
[0154] (1) Construction of genome sequencing library;
[0155] (2) sequencing to obtain whole genome sequencing data of normal diploid sample 1;
[0156] (3) Obtain analysis parameters: reads per window;
[0157] (5) Analyze the sample based on the sequencing read information according to the steps to determine whether the sample is contaminated.
[0158] The test results are shown in Figure 6. There is no contamination in the sample. The Z2score of the number of crossovers of the autosomal ploidy difference fragments in the test sample is 0.76, which is between (-3, 3), indicating that it is a normal diploid and there is no contamination.
[0159] Example 2:
[0160] This embodiment uses the above method to detect triploid sample 1, including the following steps:
[0161] (1) Construction of genome sequencing library;
[0162] (2) sequencing to obtain whole genome sequencing data of triploid sample 1;
[0163] (3) Obtain analysis parameters: reads per window;
[0164] (5) Analyze the sample based on the sequencing read information according to the steps to determine whether the sample is contaminated.
[0165] The test results are shown in Figure 7. There is no contamination in the sample. The Z3score of the number of crossovers of the autosomal ploidy difference fragments in the test sample is -1.32, which is between (-3, 3), indicating a normal triploid and no contamination.
[0166] Example 3:
[0167] This embodiment uses the above method to detect contaminated sample 1, which is a complete XY sample plus another XY 10% mixture, including the following steps:
[0168] (1) Construction of genome sequencing library;
[0169] (2) sequencing to obtain whole genome sequencing data of contaminated sample 1;
[0170] (3) Obtain analysis parameters: reads per window;
[0171] (5) Analyze the sample based on the sequencing read information according to the steps to determine whether the sample is contaminated.
[0172] The test results are shown in Figure 8. The sample is contaminated. The Z3score of the number of crossovers of the autosomal ploidy difference fragments in the test sample is -4.18, and the Z2score is 4.16. Z2score>3 and Z3score<-3, which means it is a contaminated sample.
[0173] Example 4:
[0174] This embodiment uses the above method to detect contaminated sample 2, which is a complete XY sample plus another XY 30% mixture, including the following steps:
[0175] (1) Construction of genome sequencing library;
[0176] (2) sequencing to obtain whole genome sequencing data of contaminated sample 2;
[0177] (3) Obtain analysis parameters: reads per window;
[0178] (5) Analyze the sample based on the sequencing read information according to the steps to determine whether the sample is contaminated.
[0179] The test results are shown in FIG9 . The sample is contaminated. The number of crossovers Cnum of the autosomal ploidy difference fragments in the test sample is 2 and <6, and all autosomes are three different chromosomes, which is a contaminated sample.
Claims
1. A method for detecting sample contamination, comprising: Acquiring sequencing information obtained after sequencing the sample to be analyzed, wherein the sequencing information includes filtered reads of the sample to be analyzed; Determining an autosomal trisomy likelihood ratio of the sample to be analyzed based on the filtered read length; Determining the number of autosomal ploidy difference segment crossovers of the sample to be analyzed based on the autosomal trisomy likelihood ratio; Determine whether the sample to be analyzed is contaminated based on the number of crossovers of the autosomal ploidy difference fragments of the sample to be analyzed.
2. The method according to claim 1, wherein Determining the number of autosomal ploidy difference segment crossovers of the sample to be analyzed based on the autosomal trisomy likelihood ratio includes: According to the positive and negative value relationship between the autosomal trisomy likelihood ratios of each adjacent window in the sample to be analyzed, the number of crossovers of the autosomal ploidy difference fragments of the sample to be analyzed is determined: in response to the opposite positive and negative values of the autosomal trisomy likelihood ratio in the front window and the autosomal trisomy likelihood ratio in the rear window, it is determined that there is one crossover in the sample to be analyzed.
3. The method according to claim 1, wherein Before determining whether the sample to be analyzed is contaminated according to the number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed, the method further includes: Determine the standard data of the number of crossovers of the autosomal ploidy difference segments corresponding to the reference samples, where the reference samples include known diploid samples and known triploid samples, and the standard data of the number of crossovers of the autosomal ploidy difference segments corresponding to the reference samples include the mean and standard deviation of the number of crossovers of the autosomal ploidy difference segments corresponding to the reference samples.
4. The method according to claim 3, wherein: The determining whether the sample to be analyzed is contaminated according to the number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed includes: Determine the Zscore value corresponding to the sample to be analyzed according to the mean and standard deviation of the number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed and the number of crossovers of the autosomal ploidy difference segments corresponding to the reference sample, wherein the Zscore value includes a Z2score value corresponding to the number of crossovers of the ploidy difference segments on the diploid autosomes and a Z3score value corresponding to the number of crossovers of the ploidy difference segments on the triploid autosomes; Determine whether the sample to be analyzed is contaminated based on the Zscore value corresponding to the sample to be analyzed and / or the number of autosomal ploidy difference segment crossovers of the sample to be analyzed.
5. The method according to claim 3, wherein The determining whether the sample to be analyzed is contaminated based on the Zscore value corresponding to the sample to be analyzed and / or the number of autosomal ploidy difference fragment crossovers of the sample to be analyzed further includes: In response to (1) the Z2score value is greater than the first preset threshold and the Z3score value is less than the second preset threshold, or (2) the number of crossovers of the autosomal ploidy difference fragments of the sample to be analyzed is less than the third preset threshold and all the autosomes of the sample to be analyzed are three different chromosomes, it is determined that the sample to be analyzed is contaminated.
6. The method according to claim 1, wherein Before obtaining sequencing information obtained after sequencing the sample to be analyzed, the method further includes: A reference haplotype data set is obtained, wherein the reference haplotype data set includes chromosome number, position information of SNP sites and / or genotype information of SNP sites.
7. The method according to claim 1, wherein Determining the autosomal trisomy likelihood ratio of the sample to be analyzed based on the filtered read length includes: Determining a bin on the autosomal genome of the sample to be analyzed according to a preset base length; For each window, determining the read length containing the preset haplotype variation SNP in the window based on the information of the preset haplotype variation SNP in the reference haplotype data set; According to the difference in allele probabilities of the reads containing the preset haplotype variation SNP under different ploidy assumptions, in response to assuming that the reads or read pairs come from different numbers of chromosomes, calculating the corresponding allele frequency of the reads or read pairs, wherein if it is assumed that the read pair comes from two different chromosomes, the corresponding allele frequency P2 of the read pair is calculated; if it is assumed that the read pair comes from three different chromosomes, the corresponding allele frequency P3 of the read pair is calculated; According to the allele frequency P2 and the allele frequency P3, the autosomal trisomy likelihood ratio LR3 of the window under the different ploidy assumptions is determined.
8. The method according to claim 1, wherein Determining the number of autosomal ploidy difference segment crossovers of the sample to be analyzed based on the autosomal trisomy likelihood ratio includes: Merging bins on the autosomal genome of the sample to be analyzed according to characteristics of the autosomal ploidy difference fragments in the sample to be analyzed to obtain a large window after merging the sample to be analyzed, wherein the large window includes a large window corresponding to one identical chromosome, a large window corresponding to two different chromosomes, and / or a large window corresponding to three different chromosomes, wherein the characteristics include that the window is a monomeric fragment, the window is a disomic fragment, or the window is a trisomic fragment; The number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed is determined based on the large window after the sample to be analyzed is merged.
9. The method according to claim 1, wherein Before determining the autosomal trisomy likelihood ratio of the sample to be analyzed, the method further comprises: Based on the filtered read length, the autosomal monosomal likelihood ratio of the sample to be analyzed is determined, and non-monosomal samples to be analyzed are screened.
10. The method according to claim 9, wherein: Determining the autosomal monosomal likelihood ratio of the sample to be analyzed and screening to obtain non-monosomal samples to be analyzed includes: Determining a bin on the autosomal genome of the sample to be analyzed according to a preset base length; For each window, determining the read length containing the preset haplotype variation SNP in the window based on the information of the preset haplotype variation SNP in the reference haplotype data set; According to the difference in allele probabilities of the reads containing the preset haplotype variation SNPs under different ploidy assumptions, in response to assuming that the reads or read pairs come from different numbers of chromosomes, calculating the corresponding allele frequencies of the reads or read pairs, wherein if it is assumed that the reads come from the same chromosome, the corresponding allele frequency P1 of the reads is calculated; if it is assumed that the read pairs come from two different chromosomes, the corresponding allele frequency P2 of the read pairs is calculated; According to the allele frequency P1 and the allele frequency P2, determining the autosomal monosomal likelihood ratio LR1 of the window under the different ploidy assumptions; According to the autosomal monosomal likelihood ratio LR1, the sample to be analyzed is determined to be non-monosomal.
11. A device for detecting sample contamination, comprising: an acquisition module configured to acquire sequencing information obtained after sequencing the sample to be analyzed, wherein the sequencing information includes filtered reads of the sample to be analyzed; A likelihood ratio determination module is configured to determine the autosomal trisomy likelihood ratios of the sample to be analyzed based on the filtered read lengths; a crossover number determination module, configured to determine the crossover number of the autosomal ploidy difference segment of the sample to be analyzed based on the autosomal trisomy likelihood ratio; The determination module is configured to determine whether the sample to be analyzed is contaminated based on the number of crossovers of the autosomal ploidy difference segments of the sample to be analyzed.
12. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method according to any one of claims 1 to 10.
13. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by one or more processors, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Method for detecting chromosome microdeletion and micro-duplication of human embryo
CN104745718A
Integrated identification method for embryo CNV, SV and SGD abnormity and abnormity source
CN115433777A
Detection system, device and method for embryo chromosome aneuploid and parent pollution analysis
CN117238375A
Method and device for judging sample pollution, electronic equipment and storage equipment
CN118098350A
Methods and related aspects for analyzing chromosome number status
US20230307130A1