An evaluation method for identifying sequencing data contamination

By using multiple parameters for comprehensive evaluation and pollution probability model calculation, the problem of difficulty in identifying species classification and cross-contamination in samples in the prior art is solved, and the accurate identification and evaluation of sequencing data pollution is achieved, and the accuracy of sequencing results is improved.

CN118675617BActive Publication Date: 2025-05-30ZHEJIANG LUOXI MEDICAL LAB CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410683447.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-29
Publication Date
2025-05-30
Estimated Expiration
2044-05-29

AI Technical Summary

Technical Problem

Existing sequencing data contamination identification methods are difficult to identify species classification and cross-contamination within samples at the same time, especially in low-biomass pathogenic microorganisms and common pathogenic microorganisms in the environment, making it difficult to accurately judge the pollution situation.

Method used

A comprehensive assessment of the authenticity of species annotation and cross-contamination within the sample was used to calculate the possibility of sample contamination using a contamination probability model.

Benefits of technology

Effectively identify and evaluate the contamination of sequencing data, thereby improving the identification accuracy of low-biomass pathogenic microorganisms and common pathogenic microorganisms in the environment, and reducing the impact of contamination on sequencing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118675617B_ABST
    Figure CN118675617B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of sequencing data processing, and discloses an evaluation method for identifying sequencing data contamination, comprising the following steps: Step 1: Obtain the sequencing data of samples in the same batch and negative control sequencing data and perform quality control to generate a microbial sequence data set; Step 2: Perform species alignment on the microbial sequence data set of each sample to obtain sequence alignment parameters; Step 3: Judge the authenticity of the sequences of each sample according to the sequence alignment parameters to obtain species alignment parameters; Step 4: Determine a list of suspected contaminant species, suspected pollution sources and contaminated lists; Step 5: According to the list of suspected pollution sources, use a pollution probability model to calculate the pollution possibility of the suspected contaminant species in the sample and determine the pollution probability. The present invention can comprehensively evaluate the pollution situation of the sample from two aspects of the authenticity of species annotation and cross-contamination within the sample, and use a pollution probability model to calculate the pollution possibility of the sample, providing data support for subsequent pathogen identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sequencing data processing, and in particular to an evaluation method for identifying sequencing data contamination. Background Art

[0002] With the continuous development of sequencing technology, metagenomic sequencing technology has gradually become one of the important means for clinical identification of pathogenic microorganisms. However, during the whole experiment process, metagenomic sequencing technology may be affected by various factors, resulting in the occurrence of microbial contamination. Among them, the sources of contamination include the sampling process, laboratory environment, kits, reagent contamination, etc. If the factors of contaminants are not considered during the experiment process and subsequent bioinformatics analysis, it may affect the identification results, especially the detection results of pathogenic microorganisms with low biomass.

[0003] Currently, the use of internal reference quantification method and negative control method are common methods for identifying contamination in metagenomic sequencing. Among them, the internal reference quantification method corrects the errors generated in the sequencing experiment by adding internal reference genes. However, the selection of internal reference, the concentration of added internal reference, etc. will all increase the complexity of sequencing, and thus increase the overall cost of sequencing; the negative control method judges contaminants by using negative control samples and removes potential contaminated microorganisms in the negative control samples. However, it is difficult to judge whether there is contamination for pathogenic microorganisms with low biomass and common pathogenic microorganisms in the environment.

[0004] In addition, there are also some software tools for identifying contaminants in sequencing data, including Decontam and Recentruge, etc. The Decontam software combines the frequency method and the popularity-based method. By using the frequency and popularity-based analysis, the contaminants in the final sample are obtained. However, it only uses the characteristics of sequence number, DNA concentration and pathogenic prevalence to judge contaminants, and cannot filter the problem of contaminants generated by misclassification; while Recentruge evaluates the confidence of each classification result in a score-oriented manner to achieve the purpose of identifying the contaminants generated in the sequencing data. However, this can only identify contamination for the classification of species and cannot well identify cross-contamination within species. Therefore, there is an urgent need to develop a method that can simultaneously identify sequencing contaminants from two aspects of species classification and cross-contamination within the sample. Summary of the Invention

[0005] Based on the above description, the present invention provides an evaluation method for identifying sequencing data contamination. This method comprehensively evaluates the authenticity of species annotation of the sample and the cross-contamination within the sample by using multiple parameters (kmer ratio, Unique ratio, relative abundance, genome coverage, standard sequence number and gene similarity), and uses a contamination probability model to obtain the sample contamination possibility, providing data support for subsequent sequencing result analysis.

[0006] The present invention provides an evaluation method for identifying sequencing data contamination, comprising the following steps:

[0007] Step 1: Obtain the sequencing data of samples in the same batch and negative control sequencing data, and perform quality control to generate a microbial sequence dataset for each sample;

[0008] Step 2: Perform species alignment on the microbial sequence dataset of each sample to obtain the sequence alignment parameters kmer ratio and Unique ratio;

[0009] Step 3: Judge the sequence authenticity of each sample according to the sequence alignment parameters of the species, determine the false positive alignment results, and obtain the true species sequence set of each sample and its species alignment parameters. The species alignment parameters include: genome coverage, relative abundance, number of sequences, and standard number of sequences;

[0010] Step 4: Determine the list of suspected contaminant species, suspected pollution sources, and contaminated lists for the true species sequence set;

[0011] Step 5: According to the list of suspected pollution sources, calculate the gene similarity of the suspected contaminant species in the list of suspected pollution sources and the list of contaminated sources, and use the pollution probability model to calculate the possibility of each sample in the list of suspected pollution sources being contaminated, and determine the pollution probability.

[0012] In the sequence alignment parameters described in Step 2, the kmer ratio represents the ratio of the number of kmers aligned to the same species in the sample to the total number of kmers aligned to that species, and the Unique ratio represents the ratio of the number of reads in which all kmers are aligned to the same species among the reads aligned to the same species to the total number of reads aligned to that species.

[0013] Preferably, the authenticity judgment in Step 3 is divided into false positive species judgment and false positive sequence judgment. The false positive species judgment rule is: if the kmer ratio or Unique ratio of the detected species is less than 50%, then the species is considered a false positive species; the false positive sequence judgment rule is: if the kmer ratio of a single sequence is less than 60%, then the sequence is considered a false positive sequence.

[0014] Preferably, the steps of determining the list of suspected contaminant species, suspected pollution sources, and contaminated lists in Step 4 are as follows:

[0015] Step 401: Calculate the species alignment parameters of each sample, including the number of sequences, standard number of sequences, relative abundance, and genome coverage;

[0016] Step 402: Count the detection frequency of species in each sample under the same batch;

[0017] Step 403: If a species is detected in both the same-batch samples and the negative control samples, and this species does not belong to the species that widely exist in the environment, then this species is determined to be a suspected contaminant species;

[0018] Step 404: If the number of detected species is greater than half of the number of samples in the same batch, then this species is determined to be a suspected contaminant species;

[0019] Step 405: According to the screened suspected contaminant species, select the sample with the highest relative abundance of this species as the suspected pollution source, and take the relative abundance of this species as the standard. If the ratio of the relative abundance of this species to the relative abundance of this species in other samples > 3, then include other samples in the contaminated list; if the ratio of the relative abundance of this species to the relative abundance of this species in other samples ≤ 3, then consider this sample as the suspected contaminated list.

[0020] Preferably, in step 5, if there are multiple samples in the suspected pollution source list, calculate the pollution probability of each sample and the suspected pollution source list respectively, and add the pollution probabilities of the same sample to obtain the final pollution probability of this sample.

[0021] Preferably, the pollution probability model described in step 5 is calculated using Bayes' formula, and the calculation formula is as follows:

[0022]

[0023] Among them, P(C i ,t) represents the probability of sample t being contaminated by sample C in the suspected pollution source list i , P(C i ) represents the probability of sample C i occurring pollution, and P(t) represents the probability of sample t being contaminated.

[0024] Preferably, the calculation formula of P(C i ,t) is as follows:

[0025]

[0026] I m represents the pollution index of sample t being contaminated by sample C under each index. The indexes include species relative abundance, gene similarity, genome coverage, and standard sequence number; i The pollution index of gene similarity is the genomic distance between sample C

[0027] and sample t; the calculation formulas of the pollution indexes of species relative abundance, genome coverage, and standard sequence number are as follows: i

[0028]

[0029] ​RA is the relative abundance of contaminant species in the sample, GC represents the genome coverage of the contaminant species in the sample, and SN represents the number of standard sequences of the contaminant species in the sample;

[0030] Preferably, the calculation formula of P(C i ) is as follows:

[0031] P(C i ) = P 0 (C i ) × A × B

[0032] Among them, P 0 (C i ) represents the prior probability of possible contamination caused by sample C i ; A represents the possibility of the contamination probability of sample C i in all samples, that is, the proportion of the number of standard sequences of the suspected contaminant species in this sample in the total number of standard sequences (20M) of this contaminant species in all samples; B represents the possibility of contamination of the suspected contaminant species in this sample, that is, the proportion of the number of sequences of the suspected contaminant species in this sample in the total number of microbial sequences in this sample;

[0033] Preferably, the calculation formula of P(t) is as follows:

[0034]

[0035] Among them, n is the total number of samples in the same batch and negative control samples.

[0036] Compared with the prior art, the beneficial effects of the present invention:

[0037] The present invention provides an evaluation method for identifying sequencing data contamination, and comprehensively evaluates the possibility of sample contamination from two aspects: the authenticity of species annotation and the probability of cross-contamination within the sample:

[0038] From the perspective of the authenticity of species annotation, the present invention uses the kmer ratio and Unique ratio parameters to judge the authenticity of species annotation from two aspects of species annotation and sequence authenticity respectively, which can fully consider the classification errors caused in the alignment process and effectively identify true positive sequence information;

[0039] From the aspect of cross-contamination within the sample, the present invention first uses relative abundance to preliminarily determine the source sample and the contaminated sample, and uses Bayes' formula and multiple parameters (relative abundance, genome coverage, number of standard sequences, and gene similarity) to calculate the possibility of the contaminated source sample being contaminated by the suspected source sample, providing data support for subsequent pathogen identification. Description of the Drawings

[0040] Figure 1This is the flowchart of the method for evaluating the contamination of sequencing data in the embodiments of the present invention. Detailed implementation manners

[0041] The present invention will be further described below in conjunction with embodiments.

[0042] Embodiment 1, calculating the contamination probability of Mycobacterium tuberculosis in the same batch of samples:

[0043] The present invention provides a method for evaluating the contamination of sequencing data. The specific technical route is as Figure 1 shown, and the specific steps include:

[0044] Step 1: Obtain the sequencing data and negative control sequencing data of the same batch and perform quality control to generate the microbial sequence dataset of each sample;

[0045] Step 2: Align the microbial sequence dataset of each sample with Mycobacterium tuberculosis to obtain the kmer ratio and Unique ratio of the sequence alignment parameters;

[0046] Step 3: Judge the sequence authenticity of each sample according to the sequence alignment parameters of Mycobacterium tuberculosis, determine the false positive alignment results, and obtain the set of real Mycobacterium tuberculosis sequences and their species alignment parameters in each sample. The species alignment parameters include: genome coverage, relative abundance, sequence number, and standard sequence number;

[0047] Step 4: Determine whether the Mycobacterium tuberculosis sequence set is a suspected contaminant species, a list of suspected pollution sources, and a list of contaminated samples;

[0048] Step 5: According to the list of suspected pollution sources, calculate the gene similarity of Mycobacterium tuberculosis in the list of suspected pollution sources and the list of contaminated sources, and use the contamination probability model to calculate the possibility of each sample being contaminated by the list of suspected pollution sources to determine the contamination probability.

[0049] Specifically, the quality control is performed using the software FastQC, Trimmomatic, and bowtie2, and low-quality (quality parameter Q30 < 85%) sequences, adapter sequences, and human source sequences (reference genomes of Homo sapiens GRCh37, Telomere-to-Telomere CHM13, and "YH1") are filtered. The original sequence numbers and microbial sequence numbers of the same batch of samples are shown in Table 1, where S1 to S7 represent the samples of the same batch, and NC represents the negative control samples of the same batch;

[0050] Table 1: Original sequence numbers and microbial sequence numbers of the same batch of samples:

[0051] Sample number Original sequence number Microbial sequence number NC 25681473 61846 S1 31167343 3752421 S2 33022285 10572 S3 34465981 196830 S4 42609034 4962 S5 33320418 387443 S6 38554751 110703 S7 30628562 46000

[0052] Specifically, in step 2, Kraken2 software is used for species annotation, where the length of the kmer is set to 35bp. The kmer ratio is the proportion of the number of kmers aligned to the same species in the sample to the total number of kmers aligned to that species, and the Unique ratio is the proportion of the number of reads in which all kmers are aligned to the same species among all reads aligned to that species.

[0053] In this embodiment, each read generates 16 kmer fragments, so the total number of kmers for that species is 16 * the number of reads, and the number of kmers for each read aligned to the same species is 16.

[0054] Specifically, in this embodiment, Mycobacterium tuberculosis detected in the sample is taken as an example for illustration. Table 2 shows the number of sequences of Mycobacterium tuberculosis detected, the Unique ratio, and the kmer ratio in samples S1 - S7 and the negative control sample NC.

[0055] Table 2: Detection of Mycobacterium tuberculosis in samples S1 - S7 and negative control sample NC:

[0056] Sample number Sequence number Unique proportion kmer proportion NC 35 71.875% 88.867% S1 3039056 82.623% 79.982% S2 16 57.143% 76.339% S3 20 40% 67.812% S4 24 65.217% 77.174% S5 29 66.667% 74.074% S6 33322 89.483% 82.140% S7 17 70.588% 83.824%

[0057] Specifically, the authenticity judgment in step 3 is divided into false - positive species and false - positive sequences. Among them, the filtering rule for false - positive species is: if the kmer ratio or Unique ratio of the detected species is less than 50%, then the species is considered a false - positive species; the filtering rule for false - positive sequences is: if the kmer ratio of a single sequence is less than 60%, then the sequence is considered a false - positive sequence.

[0058] According to the judgment criteria in step 3, in this embodiment, the Unique ratio of sample S3 is less than 50%, so the Mycobacterium tuberculosis in this sample may be a false - positive species. Then, false - positive sequence judgment is carried out on Mycobacterium tuberculosis in other samples, and finally, the genome coverage, relative abundance, number of sequences, and standard number of sequences of Mycobacterium tuberculosis in each sample are obtained.

[0059] Among them, the genome coverage refers to the proportion of the number of sequences aligned to the species obtained by sequencing to the total number of genomic sequences of that species; the number of sequences refers to the number of sequences of each species detected in the sample; the relative abundance refers to the proportion of the number of sequences annotated to that species to the total number of microbial sequences in the sample; the standard number of sequences refers to the number of sequences of the species detected in the sample at the 20M level. In this embodiment, the number of sequences, standard number of sequences, relative abundance, and genome coverage information of Mycobacterium tuberculosis in samples S1 - S7 and negative control sample NC are shown in Table 3.

[0060] Table 3: Species alignment parameter information of true - positive samples of Mycobacterium tuberculosis detected:

[0061] Sample number Sequence number Standard sequence number Relative abundance Genome coverage NC 32 430 0.115% 0.012% S1 3038562 11578139 81.221% 94.470% S2 14 372 0.176% 0.016% S4 23 638 0.98% 0.026% S5 27 323 0.007% 0.029% S6 33270 675730 30.842% 29.726% S7 17 279 0.071% 0.011%

[0062] Specifically, the steps for screening the list of suspected contaminant species, suspected pollution sources, and contaminated lists described in step 4 are as follows:

[0063] Step 401: Count the number of Mycobacterium tuberculosis sequences, standard sequence numbers, relative abundances, and genome coverages in each sample, as shown in Table 3;

[0064] Step 402: Count the detection frequency of Mycobacterium tuberculosis in each sample under the same batch;

[0065] Step 403: If Mycobacterium tuberculosis is detected in both the same-batch samples and the negative control samples, then Mycobacterium tuberculosis is determined to be a suspected contaminant species;

[0066] Step 404: If it is not detected in the negative control samples and the detection rate of Mycobacterium tuberculosis in the same-batch samples is greater than 50%, then Mycobacterium tuberculosis is determined to be a suspected contaminant species;

[0067] Step 405: According to the determined suspected contaminant species, select the sample with the highest relative abundance in this species as the suspected pollution source, and take the relative abundance of this species as the standard. If the ratio of the relative abundance of this species to the relative abundance of this species in other samples > 3, then include other samples in the contaminated list. If the ratio of the relative abundance of this species to the relative abundance of this species in other samples ≤ 3, then consider this sample as the suspected contamination list.

[0068] In this embodiment, according to the screening in step 4, Mycobacterium tuberculosis is detected in all samples, indicating that this species may cause contamination of the same-batch samples. Among them, sample S1 and sample S6 are included in the suspected pollution source list, and sample S2, S4, S5, S7, and the negative control sample NC are included in the polluted source list for subsequent calculation of the pollution possibility.

[0069] Specifically, step 5 uses a pollution probability model to calculate the pollution probability of each sample by the suspected pollution source list. The pollution probability model is calculated using Bayes' formula, and the calculation formula is as follows:

[0070]

[0071] P(C i ,t) represents the probability of sample t being contaminated by sample C in the suspected pollution source list, P(C i ) represents the probability of sample C i occurring contamination, P(t) represents the probability of sample t being contaminated. In this embodiment, sample C i i ​For sample S1 and sample S6, sample t is all samples of the same batch and the negative control sample;

[0072] P(C i ) is calculated as follows:

[0073] P(C i ) = P 0 (C i ) × A × B

[0074] P 0 (C i ) represents the prior probability of possible contamination caused by sample C i . In this embodiment, there are two samples in the candidate pollution source list. Therefore, for sample C i , P 0 (C i ) is 0.5; A represents the possibility of the pollution probability of sample C i among all samples, that is, the ratio of the number of standard sequences of the pollutant species of this sample C i to the total number of standard sequences of the pollutant species in all samples; B represents the pollution possibility of the pollutant species of sample C i , that is, the ratio of the number of sequences of the pollutant species of sample C i to the total number of microbial sequences in sample C i .

[0075] P(C i , t) is calculated as follows:

[0076]

[0077] I m represents the pollution index of sample t being contaminated by sample C i under each index. The indexes include species relative abundance, gene similarity, genome coverage, and the number of standard sequences;

[0078] Among them, the pollution index of gene similarity is the genome similarity between sample t and sample C i . The software Mash is used to evaluate the species genome distance between two samples, and the pollution index of sample genome similarity is calculated according to the parameter Mash - distance. When the parameter Mash - distance is smaller, it indicates that the genome similarity between the two samples is closer; the calculation formulas for the pollution indexes of species relative abundance, genome coverage, and the number of standard sequences are as follows:

[0079]

[0080] Among them, RA represents the relative abundance of contaminant species in the sample, GC represents the genomic coverage of contaminant species in the sample, and SN represents the number of standard sequences of contaminant species in the sample.

[0081] The calculation formula of P(t) is as follows:

[0082]

[0083] Using this pollution probability model to calculate the probability that each sample is contaminated by the list of suspected pollution sources. Considering that in addition to the possibility of being contaminated by the list of suspected pollution, all samples may also be contaminated by other samples. In this embodiment, n is the number of all samples and negative control samples in the same batch;

[0084] According to the pollution probability model in step 5, the pollution possibility that each sample is contaminated by sample S1 and sample S6 can be obtained; among them, Table 4 shows the pollution possibility that sample S1 contaminates other samples, and Table 5 shows the pollution possibility that sample S6 contaminates other samples.

[0085] Table 4: Pollution possibility that sample S1 contaminates other samples:

[0086]

[0087] Table 5: Pollution possibility that sample S6 contaminates other samples:

[0088]

[0089]

[0090] Since there are multiple samples (sample S1 and S6) in the list of suspected pollution sources, the pollution probability that each list of suspected pollution sources contaminates other samples is calculated separately, and the pollution probabilities of the same sample are added together to obtain the final pollution probability of the sample. Table 6 shows the final pollution probability of Mycobacterium tuberculosis in each sample and the PCR results of Mycobacterium tuberculosis.

[0091] Table 6: Final pollution probability and PCR results of samples S1 - S7 and negative control sample NC:

[0092] Sample number Final contamination probability CT value Whether tuberculosis is detected NC 99045163.93 / Not detected S1 0.002828392 22.46 Detected S2 44752448.92 / Not detected S4 7358604.312 37.24 Detected S5 997679940.6 / Not detected S6 55.27569329 26.74 Detected S7 216901962.3 / Not detected

[0093] Note: " / " indicates no CT value, indicating that tuberculosis is not detected in this sample. CT value ≤ 38 indicates positive tuberculosis, and CT value > 38 indicates negative tuberculosis.

[0094] According to the results in Table 6, the final contamination probabilities ranked from smallest to largest are: S1, S6, S4, S2, NC, S7, S3, and S5. According to the PCR results, S1, S6, and S4 were positive, and Mycobacterium tuberculosis was detected. This example shows that the method for calculating the contamination possibility of sequencing data proposed by the present invention can obtain the contamination probability of a sample based on multiple sequencing alignment parameters, providing data support for subsequent contamination judgment.

Claims

1. A method for evaluating sequencing data contamination, characterized in that: include: Step 1: Obtain sequencing data and negative control sequencing data of samples from the same batch and perform quality control to generate a microbial sequence dataset for each sample; Step 2: Perform species alignment on the microbial sequence dataset of each sample described in step 1 to obtain the sequence alignment parameters kmer ratio and unique ratio; Step 3: According to the sequence alignment parameters of the species described in step 2, the authenticity of the sequence of each sample is judged, the false positive alignment results are determined, and the true species sequence set and species alignment parameters of each sample are obtained. The species alignment parameters include: genome coverage, relative abundance, sequence number and standard sequence number; The authenticity judgment is divided into false positive species judgment and false positive sequence judgment, wherein the false positive species judgment rule is: if the kmer proportion or unique proportion of the detected species is less than 50%, the species is considered to be a false positive species; the false positive sequence judgment rule is: if the kmer proportion of a single sequence is less than 60%, the sequence is considered to be a false positive sequence; Step 4: Determine the suspected polluting species, suspected pollution source list and polluted list for the real species sequence set described in step 3; The steps for determining the suspected pollutant species, suspected pollution source list and polluted list are as follows: Step 401: Calculate species alignment parameters for each sample, including sequence number, standard sequence number, relative abundance, and genome coverage; Step 402: Count the detection frequency of species in each sample in the same batch; Step 403: If the species is detected in both the samples from the same batch and the negative control samples, and the species is not a species that is widely present in the environment, then the species is determined to be a suspected contaminant species; Step 404: If the number of species detected is greater than half of the number of samples in the same batch, the species is determined to be a suspected contaminant species; Step 405: Based on the screened suspected pollutant species, select the sample with the highest relative abundance of the species as the suspected pollution source, and use the relative abundance of the species as the standard. If the ratio of the relative abundance of the species to the relative abundance of the species in other samples is >3, then the other samples are included in the polluted list; if the ratio of the relative abundance of the species to the relative abundance of the species in other samples is ≤3, then the sample is considered to be in the suspected pollution list; Step 5: According to the list of suspected pollution sources described in step 4, calculate the gene similarity of the suspected polluting species in the list of suspected pollution sources and the list of contaminated sources, and use the pollution probability model to calculate the possibility of contamination of each sample in the list of suspected pollution sources to determine the pollution probability.

2. The method for evaluating sequencing data contamination according to claim 1, characterized in that: The contamination probability model described in step 5 is calculated using the Bayesian formula, and the calculation formula is as follows: P(C i ,t) represents sample C from the suspected pollution source list i The probability of causing sample t to be contaminated, P(C i ) represents sample C i The probability of contamination, P(t) represents the probability that sample t is contaminated; The P(C i ,t)、P(C i ) and P(t) are calculated as follows: (1) Among them, I m Indicates that sample t under each index is affected by sample C i The pollution index of pollution includes species relative abundance, gene similarity, genome coverage and number of standard sequences; The contamination index of the gene similarity is sample C i The calculation formulas for the genome similarity distance between sample t, species relative abundance pollution index, genome coverage pollution index and standard sequence number pollution index are as follows: Among them, RA is the relative abundance of the contaminant species in the sample, GC represents the genome coverage of the contaminant species in the sample, and SN represents the number of standard sequences of the contaminant species in the sample; (2)P(C i )=P0(C i )×A×B Among them, P0(C i ) indicates that the sample C i Prior probability of causing contamination; A represents sample C i The probability of contamination in all samples is the ratio of the number of standard sequences of the suspected contaminant species in the sample to the total number of standard sequences of the contaminant species in all samples; B represents the probability of contamination of the suspected contaminant species in the sample, that is, the ratio of the number of sequences of the suspected contaminant species in the sample to the total number of microbial sequences in the sample; (3) Where n is the total number of samples and negative control samples in the same batch.

Citation Information

Patent Citations

  • Method for judging background introduced microorganism sequence and application thereof

    CN113270145A

  • Method for identifying and analyzing pathogenic microorganisms based on metagenome sequencing data

    CN118038989A