Microbial sequencing sample pollution bacterium identification method based on single nucleotide polymorphism
By constructing a method for identifying contaminating bacteria in microbial sequencing samples based on single nucleotide polymorphisms and using BWA alignment and machine learning algorithms, the problem of distinguishing contaminating sequences from true colonization sequences in microbial sequencing samples was solved, and the accurate identification and abundance estimation of contaminating bacteria and colonizing bacteria were achieved, thereby improving the accuracy of microbial community analysis.
Patent Information
- Application Number
- CN202510758465.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies make it difficult to accurately distinguish between contaminant sequences and true colonization sequences in microbial sequencing samples, especially for microorganisms that may be both contaminants and colonizers. Commonly used methods may lead to false positive results and mislead scientific conclusions.
By constructing a method for identifying contaminating bacteria in microbial sequencing samples based on single nucleotide polymorphisms, using BWA alignment software and machine learning algorithms, SNP similarity and dissimilarity indices were calculated, and a RandomForest classifier model was established to identify and quantify the abundance of contaminating bacteria and colonizing bacteria.
It achieves accurate identification of contaminating bacteria and colonizing bacteria in microbial sequencing samples, improves the accuracy and robustness of identification, can accurately distinguish contaminating bacteria from colonizing bacteria in low biomass samples, and improves the accuracy of microbial community composition characterization.
Smart Images

Figure CN120690277A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of microbial sequencing technology, and in particular relates to a method for identifying contaminating bacteria in microbial sequencing samples based on single nucleotide polymorphisms, aiming to solve the problem of identifying background bacterial contamination in low-load metagenomic sequencing. Background Art
[0002] High-throughput DNA sequencing has become a key tool for studying the composition and function of microbial communities. Whole-genome shotgun metagenomic sequencing, in particular, allows for precise characterization of microbial community diversity and its functional potential by sequencing all DNA recovered from a sample. However, the accuracy of metagenomic sequencing is often compromised by contamination from various sources in practice, such as reagents, experimental equipment, and exogenous DNA introduced by engineered bacteria. This contamination is particularly pronounced in low-biomass environmental samples, where low levels of endogenous DNA make it more likely that contamination will mask true microbial signals, leading to controversial conclusions regarding the presence of microorganisms in low-biomass environments such as blood or body tissue. In high-biomass environments, although the impact of contamination is relatively small, contaminating bacteria may still represent low-frequency sequences in the data, interfering with the interpretation of low-frequency variants. This not only limits understanding of the true structure of the community but can also lead to false-positive results in subsequent microbial community analyses, misleading scientific conclusions and subsequent research. Currently, a common approach to removing contaminating bacteria involves using statistical methods to distinguish between contaminating and authentic microorganisms. Specifically, correlations between contaminating bacterial sequences and sample DNA concentrations can be calculated to identify and remove contaminating bacteria with significant negative correlations. However, this approach is not always effective, especially for microorganisms that may be both contaminants and true colonizers. In addition, the commonly used method of setting up negative blank controls, by removing strains with a high frequency in the negative blank control, also has certain limitations, as these methods may mistakenly remove colonizers that are actually present in the sample.
[0003] Therefore, current technology urgently needs a more accurate method to distinguish between contamination sequences and true colonization sequences in samples, and to effectively estimate the abundance of microorganisms that may be both contaminants and colonizers. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for identifying contaminating bacteria in microbial sequencing samples based on single nucleotide polymorphisms, to provide a more accurate method to distinguish contaminating sequences from true colonizing sequences in samples, and to effectively estimate the abundance of microorganisms that may be both contaminating bacteria and colonizing bacteria.
[0005] The technical solution adopted by the present invention is: a method for identifying contaminating bacteria in microbial sequencing samples based on single nucleotide polymorphisms comprises the following steps:
[0006] Step S1: Using BWA alignment software, the sequencing short reads of the microbial sequencing PC sample and the negative blank control NC sample are aligned with the microbial reference genome respectively to obtain the aligned BAM files, and the BAM files are sorted;
[0007] Step S2: Based on the BAM file, output the alignment information of each sequencing short read segment at each site on the reference genome, and extract the SNP and InDel information present in each short read segment sequence;
[0008] Step S3, comparing the SNP characteristics of strains with the same coverage area in the microbial sequencing PC sample and the negative blank control NC sample, and calculating the SNP similarity index and SNP dissimilarity index;
[0009] Step S4: Integrate the similarity and dissimilarity metrics of SNP features between microbial sequencing PC samples and negative blank control NC samples, use them as the core input features of the RandomForest classification algorithm in machine learning, and establish a three-classifier model to distinguish pure contaminating bacteria, contaminated colonizing mixed bacteria, and pure colonizing bacteria, and output the judgment results;
[0010] Step S5, quantitatively evaluating the richness of colonizing bacteria in the contaminated colonizing mixed bacteria;
[0011] Step S6: After the output of the model is obtained, a confidence score of the model determination result is performed to obtain an indication of the reliability of the model output result.
[0012] Preferably, in step S2, variant information detected at the beginning and end of the short reads is excluded, and variant information fully supported by multiple aligned short reads is not considered.
[0013] Preferably, in step S3, strains are considered whose common coverage region is longer than 100 bp and the number of polymorphic sites contained in the common coverage region is greater than 8.
[0014] Preferably, in step S3, the method for calculating the SNP similarity index is:
[0015]
[0016] Among them, All_Multi_Allele represents all polymorphic sites detected in the common coverage area. These sites show single nucleotide polymorphisms (SNPs) in at least one sample. Shared_Multi_Allele is the polymorphic sites in All_Multi_Allele that share the same coverage base between different samples. At least one of these sites can show consistent coverage bases with the negative blank control NC sample at this specific site. SNV Similarity is based on Shared_allele_ratio and further incorporates the sequencing depth information of each polymorphic site. It uses weighted calculation to more finely evaluate the similarity. PopANI is an overall similarity assessment of the common coverage area, which not only considers the polymorphic sites in the common coverage area, but also covers non-polymorphic sites, thereby comprehensively measuring the sequence similarity of the entire region.
[0017] The calculation method for the SNP dissimilarity index is:
[0018]
[0019] Among them, Dif_Multi_Allele represents the polymorphic sites in All_Multi_Allele that have different coverage bases between different samples, and contains at least one coverage base that is completely different from the negative blank control NC sample at this specific site. The read unsupported ratio is calculated by accumulating the number of reads at each polymorphic site that are different from the SNP of the negative blank control NC sample. PopDIF is an assessment of the overall dissimilarity of the common coverage area.
[0020] Preferably, in step S4, the three-classifier model outputs three independent scores, namely Contaminatorscore, Mixture score and Colonizer score, each score is between 0 and 1, representing the possibility of the strain being determined to be a corresponding category.
[0021] Preferably, in step S5, the quantitative assessment method of the richness of the colonizing bacteria in the contaminating colonizing bacteria mixture is estimated based on the read set covering the polymorphic site;
[0022] For strains with low sequencing depth, it is necessary to correct the coverage read amount of these polymorphic sites that do not have consistent coverage bases with the negative blank control NC sample, that is, to estimate the sequencing depth λ of these sites. The calculation method is:
[0023]
[0024] λ=1-log(1-Coverage),
[0025] The coverage of the sites in the sequencing process is estimated based on the Poisson distribution. Coverage is the estimated coverage of the contaminating bacteria from the negative blank control NC sample in the microbial sequencing PC sample.
[0026]
[0027] For strains with high sequencing depth, i.e., strains with an average sequencing depth greater than 10, only reads covering heterozygous polymorphic sites were considered for richness estimation.
[0028] Preferably, in step S6, the confidence score of the model judgment result is performed based on the length of the common coverage area of the strain in the two samples, the sequencing depth, the number of polymorphic sites detected, and the number of covered reads. The longer the common coverage area of the strain in the two samples, the higher the number of covered polymorphic sites, and the higher the number of covered reads, the higher the accuracy of the model judgment result. The confidence score calculation method is to use the sequencing depth and the number of covered reads as the input of the RandomForest model on the simulated training data set, and train with the correctness of the model judgment results as labels to evaluate the reliability of the model results under different conditions.
[0029] The advantages and positive effects of the present invention are:
[0030] The present invention provides a method for identifying contaminating bacteria in microbial sequencing samples based on single nucleotide polymorphisms, which has the following advantages:
[0031] 1) A multi-dimensional SNP assessment system was constructed when comparing the SNP signatures of two sample strains. By simultaneously considering both SNP similarity and SNP dissimilarity, this method not only identifies shared strain signatures (similarity) between samples but also sensitively captures unique variant information (dissimilarity). This enables more precise distinction between pure contaminating bacteria, pure colonizing bacteria, and mixed contaminating and colonizing bacteria in microbial sequencing samples, resolving the difficulty of accurately identifying complex contamination scenarios with traditional methods.
[0032] 2) By comparing a series of extracted SNP signatures, a machine learning model is established for prediction, enabling automated determination of strain colonization and contamination. Compared to traditional methods based on fixed threshold screening, machine learning models can automatically learn and optimize classification rules, uncovering deeper patterns from complex SNP signature data, thereby improving prediction accuracy and robustness.
[0033] 3) It has a wider scope of application. Even in biological samples with low biomass, that is, when the sequencing depth is limited, the present invention can still accurately identify contaminating bacteria and colonizing bacteria through fine SNP feature extraction and efficient machine learning algorithms. This feature is particularly important for clinical microbial samples with limited sequencing. In addition, the present invention can not only be used to assess whether the strain in the sample is a contaminating strain, but also to perform comparative analysis of SNP characteristics of different samples to determine whether the bacteria are strongly cross-contaminated between samples.
[0034] 4) A SNP-based abundance quantification method was proposed, which can restore the true abundance of colonizing bacteria in contaminated colonization mixed strains, thereby improving the accuracy of microbial community composition characterization. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a flowchart of the method of the present invention;
[0036] Figure 2 It is a schematic illustration of the polymorphic sites in the common coverage region detected in the microbial sequencing PC sample and the negative blank control NC sample;
[0037] Figure 3 is a schematic diagram of the simulated metagenomic sample generation process;
[0038] Figure 4 Schematic diagram comparing the richness of contaminating bacteria in the contaminated colonizing mixed bacteria estimated by the method of the present invention and the actual richness;
[0039] Figure 5 Schematic diagram of the correlation between the model-estimated richness of contaminating bacteria in the contaminating colonizing mixture and the actual richness. DETAILED DESCRIPTION
[0040] In order to further understand the content, features and effects of the present invention, the following embodiments are given to illustrate in detail.
[0041] The present invention provides a method for identifying contaminating bacteria in microbial sequencing samples based on single nucleotide polymorphism, which can be named ScanTrack.
[0042] The theoretical basis of ScanTrack is that strains from different sources usually have unique single nucleotide polymorphism (SNP) characteristics. Therefore, by comparing the SNP differences between strains in negative control samples and microbial sequencing samples, the colonization status of strains in microbial sequencing samples can be effectively determined. The present invention first relies on the short read data generated by microbial sequencing technology to accurately map it to the microbial reference genome, and then extracts discernible SNP features from it. Subsequently, these SNP features are deeply compared with the SNP features displayed by the same short read coverage area in the negative control sample to construct a series of quantitative indicators to measure SNP similarity and dissimilarity. Similarity indicators include but are not limited to the proportion of identical polymorphic sites and SNV similarity assessment based on sequencing depth weighting; while dissimilarity indicators focus on the proportion of different polymorphic sites and the proportion of reads that do not support the negative control sample SNP in the reads covering the polymorphic sites. Furthermore, the present invention uses the above-mentioned SNP comparison indicators as input features of the machine learning model to perform model training and optimization on the constructed simulated training data set. This process aims to build an efficient and accurate classification model capable of intelligently distinguishing between different strain states, including pure contaminants, mixed contaminants and colonizers, and pure colonizers. To quantitatively assess the richness of colonizers in mixed contaminant-colonizers, this paper proposes an innovative metric: based on reads mapped to polymorphic sites within the common coverage region, the proportion of reads that do not support the SNPs in the negative control sample is calculated to accurately estimate the abundance of colonizers.
[0043] See Figure 1 The method for identifying contaminating bacteria in microbial sequencing samples based on single nucleotide polymorphisms of the present invention comprises the following steps:
[0044] Step S1: Using BWA alignment software, the sequencing short reads of the microbial sequencing PC sample and the negative blank control NC sample are aligned with the microbial reference genome respectively to obtain the aligned BAM files, and the BAM files are sorted;
[0045] Step S2: Based on the BAM file, output the alignment information of each sequencing short read segment at each site on the reference genome, and extract the SNP and InDel information present in each short read segment sequence;
[0046] In this step, variant information detected at the beginning and end of the short reads is excluded, and variant information fully supported by multiple aligned short reads is not considered.
[0047] Step S3, comparing the SNP characteristics of strains with the same coverage area in the microbial sequencing PC sample and the negative blank control NC sample, and calculating the SNP similarity index and SNP dissimilarity index;
[0048] In this step, strains with a common coverage region length greater than 100 bp and a number of polymorphic sites greater than 8 were considered.
[0049] In this step, the calculation method for the SNP similarity index is:
[0050]
[0051] Among them, All_Multi_Allele represents all polymorphic sites detected in the common coverage area. These sites show single nucleotide polymorphisms (SNPs) in at least one sample. Shared_Multi_Allele is the polymorphic sites in All_Multi_Allele that share the same coverage base between different samples. At least one of these sites can show consistent coverage bases with the negative blank control NC sample at this specific site. SNV Similarity is based on Shared_allele_ratio and further incorporates the sequencing depth information of each polymorphic site. It uses weighted calculation to more finely evaluate the similarity. PopANI is an overall similarity assessment of the common coverage area, which not only considers the polymorphic sites in the common coverage area, but also covers non-polymorphic sites, thereby comprehensively measuring the sequence similarity of the entire region.
[0052] The calculation method for the SNP dissimilarity index is:
[0053]
[0054] Among them, Dif_Multi_Allele represents the polymorphic sites in All_Multi_Allele that have different coverage bases between different samples, and contains at least one coverage base that is completely different from the negative blank control NC sample at this specific site. The read unsupported ratio is calculated by accumulating the number of reads at each polymorphic site that are different from the SNP of the negative blank control NC sample. PopDIF is an assessment of the overall dissimilarity of the common coverage area.
[0055] See Figure 2, the above example is further explained. The number of polymorphic sites in the common coverage area detected in the PC sample and the NC sample of this strain is 4, of which the number of polymorphic sites shared_Multi_Allele sharing the same coverage base is 3 (site 2, site 3, site4), and the number of polymorphic sites Dif_Multi_Allele with at least one different coverage base is 3 (site 1, site 2, site4).
[0056] In addition, for strains with high sequencing depth (average sequencing depth greater than 10), SNPs with a read support ratio of less than 0.1 will be removed to eliminate false positive SNPs caused by sequencing or alignment errors. This filtering process is not performed for strains with low sequencing depth.
[0057] In step S4, the similarity and dissimilarity metrics of SNP features between microbial sequencing PC samples and negative blank control NC samples are integrated and used as the core input features of the RandomForest classification algorithm in machine learning. A three-classifier model is established to distinguish and output the results of pure contaminating bacteria, contaminated colonizing mixed bacteria, and pure colonizing bacteria, and to quantitatively evaluate the richness of colonizing bacteria in contaminated colonizing mixed bacteria. The RandomForest classification algorithm model is the algorithm model provided by the Python library.
[0058] In this step, the three-classifier model outputs three independent scores: Contaminatorscore, Mixture score, and Colonizerscore. Each score is between 0 and 1, representing the probability that the strain is classified into the corresponding category.
[0059] Step S5, quantitatively evaluating the richness of colonizing bacteria in the contaminated colonizing mixed bacteria;
[0060] In this step, the quantitative assessment method of the colonizing bacterial richness in the contaminating colonizing bacterial mixture is estimated based on the read set covering the polymorphic sites;
[0061] For strains with low sequencing depth, it is necessary to correct the coverage read amount of these polymorphic sites that do not have consistent coverage bases with the negative blank control NC sample, that is, to estimate the sequencing depth λ of these sites. The calculation method is:
[0062]
[0063] λ=1-log(1-Coverage),
[0064] The coverage of the sites in the sequencing process is estimated based on the Poisson distribution. Coverage is the estimated coverage of the contaminating bacteria from the negative blank control NC sample in the microbial sequencing PC sample.
[0065]
[0066] For strains with high sequencing depth, i.e., strains with an average sequencing depth greater than 10, only reads covering heterozygous polymorphic sites were considered for richness estimation.
[0067] Step S6: After the output of the model is obtained, a confidence score of the model determination result is performed to obtain an indication of the reliability of the model output result.
[0068] In this step, the confidence score of the model judgment result is calculated based on the length of the common coverage area of the strain in the two samples, the sequencing depth, the number of polymorphic sites detected, and the number of covered reads. The longer the common coverage area of the strain in the two samples, the higher the number of covered polymorphic sites, and the higher the number of covered reads, the higher the accuracy of the model judgment result. The confidence score is calculated by using the sequencing depth and the number of covered reads as the input of the RandomForest model on the simulated training data set, and using the correctness or incorrectness of the model judgment results as labels for training to evaluate the reliability of the model results under different conditions.
[0069] Example 1 ScanTrack is applied to a simulated metagenomic test sequencing dataset
[0070] Step S1, the simulated metagenomic sample generation process is as follows Figure 3As shown in the figure, sample generation is primarily based on the genomes of the 400 most common reagent bacteria in the negative blank control (NC) sample. These bacterial genomes were first randomly introduced with different SNP signatures into the negative blank control (NC) sample and the microbial sequencing (PC) sample, respectively, to serve as NC sample contaminants and PC sample colonizers. Subsequently, each NC sample was generated based on 200 randomly selected SNP-introduced NC sample contaminants, and NC metagenome samples were generated based on a lognormal (5, 0.5) abundance distribution. Paired PC sample metagenome samples were generated based on both the SNP-introduced NC sample contaminants and PC sample colonizers. A portion of randomly selected bacterial reads were generated based solely on NC sample contaminants as pure contaminants, while a portion of bacterial reads were generated based on both NC sample contaminants and PC sample colonizers as contaminant-colonizer mixtures. The remaining portion of bacterial reads were generated solely based on PC sample colonizers, serving as complete colonizers. All of these sequencing short reads were generated using the Illumina sequencing simulation software ART. Four datasets with different sequencing depths were simulated here, each of which contained 30 PC sample metagenomic samples and 30 paired NC sample metagenomic samples.
[0071] In step S2, the simulated metagenomic sequencing reads are aligned with the genomes of the 400 selected strains to obtain a BAM file.
[0072] In step S3, the BAM files obtained from NC and PC are used as input to ScanTrack to obtain the SNP characteristics of different strains in NC and PC, respectively.
[0073] Step S4: compare the SNP characteristics of the NC and PC strains, generate SNP similarity and SNP dissimilarity indicators for each strain, and output the prediction results of the strains.
[0074] In step S5, the prediction results of ScanTrack are compared with the known strain label results (Contaminator: completely contaminated bacteria; Mixture: contaminated and colonized mixed bacteria; Colonizer: completely colonized bacteria) to evaluate the performance and accuracy of ScanTrack. The AUC values of the model for discriminating Contaminator, Mixture, and Colonizer are 0.97, 0.96, and 0.90, respectively. Figure 4 As shown in .
[0075] Step S6: The accuracy of ScanTrack's richness estimation was evaluated by comparing the ScanTrack-estimated contaminant bacterial richness with the actual richness. The model-estimated richness of contaminant bacterial richness in the contaminant colonization mixture had a richness prediction error of 0.08, and the predicted richness had a significant Pearson correlation with the actual richness (Pearson's r = 0.66, p < 2.2e-16). Figure 5 As shown in .
[0076] Example 2: Testing for the presence of maltophilia colonization in pulmonary alveolar lavage fluid samples
[0077] In step S1, a clinical sample of bronchoalveolar lavage fluid (BALF) was collected and subjected to metagenomic sequencing. Sequencing results showed that Maltophilia was detected in both the negative control and the BALF sample, but it was unclear whether the bacteria colonized the BALF sample.
[0078] In step S2, the sequencing short reads of the BALF sequencing sample and the paired negative control sample are aligned to the reference genome of Bacillus maltophilia, and the resulting BAM file is used as input to ScanTrack SNV_profile. This yields the SNP signature of the strain. Subsequently, the SNP signature result files of the two samples are used as input to ScanTrack SNV_compare to compare the SNP signatures of Bacillus maltophilia to identify whether the bacterium is colonized in the BALF sample. The output result file for the bacterium is shown in the following table:
[0079] Table 1 ScanTrack output result file for Maltophilia in BALF samples
[0080]
[0081]
[0082] The results showed that Maltophilia was a contaminating mixed colonizing bacteria in the bronchoalveolar lavage fluid samples, and the colonizing bacteria accounted for a large proportion.
[0083] Example 3: Whether there is strong positive contamination of Pseudomonas aeruginosa in bronchoalveolar lavage fluid samples
[0084] In step S1, clinical samples of bronchoalveolar lavage fluid were collected and subjected to metagenomic sequencing. Sequencing results revealed a high abundance of Maltophilia in bronchoalveolar lavage fluid samples from the same batch, but this bacterium was not detected in the paired negative control samples from the same batch. It is unclear whether the high abundance of this bacterium in the samples was due to cross-contamination with a strong positive result.
[0085] In step S2, the sequencing short reads of the same batch of alveolar lavage fluid sequencing samples are aligned to the reference genome of Bacillus maltophilia, and the obtained BAM file is used as the input of ScanTrack SNV_profile. Thus, the SNP characteristics of the strain are obtained. The SNP feature result file of the obtained sample is then used as the input of ScanTrack SNV_compare to compare the SNP characteristics of Pseudomonas aeruginosa to identify whether there is strong positive cross-contamination and colonization of the bacteria in the alveolar lavage fluid sample. Here, the alveolar lavage fluid sample with high richness of Pseudomonas aeruginosa from the same batch is used as the PC sample, and the other alveolar lavage fluid sample is used as the NC sample. The output results of the bacteria are as follows:
[0086] Table 2 ScanTrack output result file for Pseudomonas aeruginosa in bronchoalveolar lavage fluid samples
[0087] Pseudomanasaeruginosa Covered_region 25790 Read_count_PC 34067 Read_count_NC 528 SNV_count_PC 1529 SNV_count_NC 95 All_Multi_Allele 1546 Shared_allele_ratio 0.86 Dif_allele_ratio 0.992 PopANI 0.990 PopDIF 0.059 SNV_similarity_PC 0.895 SNV_similarity_NC 0.847 SNV_unsupport_readnum_PC 7305 SNV_unsupport_ratio_PC 0.291 SNV_unsupport_readnum_NC 182 SNV_unsupport_ratio_NC 0.369 Cantaminator_abundance 0.70 Contaminatorscore 0.28 Mixturescore 0.69 Colonizerscore 0.03 ScanTrack_predict Mixture Confidence 0.76
[0088] The results showed that there was indeed strong cross-contamination of Pseudomonas aeruginosa in the bronchoalveolar lavage fluid sample, but in addition to the contaminating bacteria, about 20% were the true colonizing bacteria of the sample itself.
[0089] The above results for clinical samples show that the ScanTrack method developed in this application can accurately identify the colonization status of strains in clinical samples.
Claims
1. A method for identifying contaminating bacteria in microbial sequencing samples based on single nucleotide polymorphisms, characterized by: Including the following step, Step S1: Using BWA alignment software, the sequencing short reads of the microbial sequencing PC sample and the negative blank control NC sample are aligned with the microbial reference genome respectively to obtain the aligned BAM files, and the BAM files are sorted; Step S2: Based on the BAM file, output the alignment information of each sequencing short read segment at each site on the reference genome, and extract the SNP and InDel information present in each short read segment sequence; Step S3, comparing the SNP characteristics of strains with the same coverage area in the microbial sequencing PC sample and the negative blank control NC sample, and calculating the SNP similarity index and SNP dissimilarity index; Step S4: Integrate the similarity and dissimilarity metrics of SNP features between microbial sequencing PC samples and negative blank control NC samples, use them as the core input features of the RandomForest classification algorithm in machine learning, and establish a three-classifier model to distinguish pure contaminating bacteria, contaminated colonizing mixed bacteria, and pure colonizing bacteria, and output the judgment results; Step S5, quantitatively evaluating the richness of colonizing bacteria in the contaminated colonizing mixed bacteria; Step S6: After the output of the model is obtained, a confidence score of the model determination result is performed to obtain an indication of the reliability of the model output result.
2. The method for identifying contaminating bacteria in microbial sequencing samples based on single nucleotide polymorphisms according to claim 1, wherein: In step S2, variant information detected at the beginning and end of the short read segment is excluded, and variant information fully supported by multiple aligned short read segments is not considered.
3. The method for identifying contaminating bacteria in microbial sequencing samples based on single nucleotide polymorphisms according to claim 1, wherein: In step S3, strains with a common coverage region length greater than 100 bp and a number of polymorphic sites greater than 8 are considered.
4. The method for identifying contaminating bacteria in microbial sequencing samples based on single nucleotide polymorphisms according to claim 3, wherein: In step S3, the calculation method for the SNP similarity index is: Among them, All_Multi_Allele represents all polymorphic sites detected in the common coverage area. These sites show single nucleotide polymorphisms (SNPs) in at least one sample. Shared_Multi_Allele is the polymorphic sites in All_Multi_Allele that share the same coverage base between different samples. At least one of these sites can show consistent coverage bases with the negative blank control NC sample at this specific site. SNV Similarity is based on Shared_allele_ratio and further incorporates the sequencing depth information of each polymorphic site. It uses weighted calculation to more finely evaluate the similarity. PopANI is an overall similarity assessment of the common coverage area, which not only considers the polymorphic sites in the common coverage area, but also covers non-polymorphic sites, thereby comprehensively measuring the sequence similarity of the entire region. The calculation method for the SNP dissimilarity index is: Among them, Dif_Multi_Allele represents the polymorphic sites in All_Multi_Allele that have different coverage bases between different samples, and contains at least one coverage base that is completely different from the negative blank control NC sample at this specific site. The read unsupported ratio is calculated by accumulating the number of reads at each polymorphic site that are different from the SNP of the negative blank control NC sample. PopDIF is an assessment of the overall dissimilarity of the common coverage area.
5. The method for identifying contaminating bacteria in microbial sequencing samples based on single nucleotide polymorphisms according to claim 1, wherein: In step S4, the three-classifier model outputs three independent scores: Contaminator score, Mixture score, and Colonizer score. Each score is between 0 and 1, representing the probability that the strain is classified into the corresponding category.
6. The method for identifying contaminating bacteria in microbial sequencing samples based on single nucleotide polymorphisms according to claim 1, wherein: In step S5, the quantitative assessment method of the richness of the colonizing bacteria in the contaminating colonizing bacterial mixture is estimated based on the read set covering the polymorphic sites; For strains with low sequencing depth, it is necessary to correct the coverage read amount of these polymorphic sites that do not have consistent coverage bases with the negative blank control NC sample, that is, to estimate the sequencing depth λ of these sites. The calculation method is: λ=1-log(1-Coverage), The coverage of the sites in the sequencing process is estimated based on the Poisson distribution. Coverage is the estimated coverage of the contaminating bacteria from the negative blank control NC sample in the microbial sequencing PC sample. For strains with high sequencing depth, i.e., strains with an average sequencing depth greater than 10, only reads covering heterozygous polymorphic sites were considered for richness estimation.
7. The method for identifying contaminating bacteria in microbial sequencing samples based on single nucleotide polymorphisms according to claim 1, wherein: In step S6, the confidence score of the model judgment result is performed based on the length of the common coverage area of the strain in the two samples, the sequencing depth, the number of polymorphic sites detected, and the number of covered reads. The longer the common coverage area of the strain in the two samples, the higher the number of covered polymorphic sites, and the higher the number of covered reads, the higher the accuracy of the model judgment result. The confidence score calculation method is to use the sequencing depth and the number of covered reads as the input of the RandomForest model on the simulated training data set, and use the correctness of the model judgment results as labels for training to evaluate the reliability of the model results under different conditions.
Citation Information
Cited By
Tumor single sample pollution judgment method, judgment model construction method and electronic device thereof
CN121959292A