Pathogen identification method based on Bayesian model and application thereof
Through the Bayesian model-based pathogen identification method, combined with tNGS and multi-target joint technology, the problem of difficult to distinguish high-similar pathogens and subspecies in the prior art is solved, and efficient and accurate pathogen identification and mixed infection detection are achieved.
Patent Information
- Application Number
- CN202510230736.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-28
AI Technical Summary
The prior art has the problem that it is difficult to distinguish high-similar pathogens, subspecies and sequencing results cannot be typing due to mutations. Some methods have limited targets, making it difficult to achieve subspecies identification and mixed infection detection.
The Bayesian model-based pathogen identification method is used, combined with the direct homologous gene or sequence comparison method of tNGS methodology, Bayesian theorem and multi-target joint, and the precise identification of pathogens is achieved through decision analysis.
High sensitivity and specificity of pathogen identification is achieved, which can effectively identify mixed infections and provide reliable speculation and dynamic updates of prediction results when data is insufficient.
Smart Images

Figure CN120148644A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of high-throughput sequencing and gene detection, and specifically relates to a pathogen identification method based on a Bayesian model and its application. Background Art
[0002] Currently, there are mainly two common methods for determining whether two bacterial genomes belong to the same species: Average Nucleotide Identity (ANI) and 16S rRNA gene sequence similarity. Currently, it is generally believed that when the ANI value reaches more than 95%, or the 16S rRNA gene sequence similarity reaches 98.65%, the sequences of the two strains can be considered to be of the same species.
[0003] Currently, the gold standard for bacterial species identification is the direct homologous gene or sequence comparison method, that is, bacteria are identified to the species level by analyzing the compositional differences of homologous DNA sequences.
[0004] Whole Genome Sequencing (WGS) of pathogens is a relatively optimal method for pathogen typing. However, pathogen WGS requires a pathogen culture, and some strains have defects such as long culture time or difficult culture, making it difficult to be universal. Metagenomics Next Generation Sequencing (mNGS) can directly detect clinical samples for pathogen typing, but it has defects such as possible incomplete sequencing coverage and high cost, and is also difficult to be universal. Amplifying and sequencing the specific homologous DNA sequences of pathogens in clinical samples is a more economical and universal solution.
[0005] However, in actual pathogen typing and detection, the genomic similarities of the specific homologous DNA sequences of some pathogens are relatively high (exceeding 98%, even 99%), or the genomic similarities of the specific homologous DNA sequences of pathogen subspecies are extremely high (exceeding 99.3%). Moreover, pathogens are prone to mutation, such as various viruses like Mycobacterium, Nocardia, influenza A virus, novel coronavirus, adenovirus, etc. This leads to the situation that when analyzing homologous DNA sequences, due to mutations, the sequence may be aligned to pathogen A, pathogen B, pathogen C, etc. simultaneously, and thus the pathogen sequencing results cannot be typed. At the same time, some pathogens are difficult to distinguish through individual homologous DNA sequences.
[0006] Most of the current mycobacterium identification methods are based on multiplex PCR technology, including real-time fluorescence quantitative PCR technology, isothermal amplification technology, probe-reverse hybridization technology, probe-melting curve technology, etc., which are relatively convenient and fast. However, these reagents usually have limited targets and can cover a limited number of typing sites. Therefore, fewer pathogens can be identified, and it is difficult to identify subspecies. Some pathogens with close genetic relationships cannot be accurately identified.
[0007] tNGS (Targeted Next Generation Sequencing) is a new generation sequencing technology based on the combination of super-multiplex PCR amplification and high-throughput sequencing technology. It uses a large number of primers targeting specific gene sequences to perform super-multiplex PCR amplification on the nucleic acids in the sample to be tested, obtains a large number of target nucleic acid fragments, and then performs high-throughput sequencing, and then performs bioinformatics analysis on the obtained sequences, thereby achieving high-sensitivity and high-resolution identification. Therefore, the development of an efficient and accurate pathogen identification method combined with tNGS methodology has important application value. Summary of the invention
[0008] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a pathogen identification method based on a Bayesian model and its application. The present invention combines the direct homologous gene or sequence comparison method of tNGS methodology, Bayesian theorem, and multi-target combination to perform decision analysis on pathogens, thereby achieving accurate pathogen identification with high sensitivity and good specificity, and can effectively identify mixed infections.
[0009] In order to achieve the purpose of the invention, the present invention adopts the following technical solutions:
[0010] In a first aspect, the present invention provides a method for pathogen identification based on a Bayesian model, the method comprising:
[0011] (S1) Specific homologous DNA sequences of pathogens that can be used for typing of closely related pathogens or pathogen subspecies are collected from the database, and SNP sites with high consistency are screened from them, and the SNP sites are divided into different cgSNPs combinations based on different pathogen typing;
[0012] (S2) Calculate the proportion of the reference genome of a certain pathogen typing in all the reference genomes of typing as the prior probability P(A);
[0013] (S3) organizing the data of cgSNPs combination corresponding to each pathogen typing into a data set, wherein the data set includes different cgSNPs combination types and the number of occurrences of different cgSNPs combination types in different pathogen typing; and calculating the distribution ratio of each pathogen typing in each cgSNPs combination type in the data set as the conditional probability P(B|A);
[0014] (S4) Based on the values of P(B|A) and P(A), through Bayesian statistics, when a certain combination type of cgSNPs is detected in the sequencing data analysis results, calculate the occurrence probability of the detected pathogen typing.
[0015]
[0016] In the formula, P(A) is the prior probability, the judgment of event A before the occurrence of event B; P(A|B) is the conditional probability, that is, the probability of A after the occurrence of event B; P(B) is the marginal probability, that is, the probability of observing event B under all possible event A. is the likelihood function, adjusted to make the estimated probability closer to the true probability.
[0017] In the present invention, when the value of P(A|B) is greater than 0.85, it is considered that the confidence level of pathogen typing is relatively high, and when it is greater than 0.95, it is considered pathogen typing.
[0018] Preferably, in (S1), the specific homologous DNA sequence for near-source pathogen typing or pathogen subspecies typing is a set of genes carried in all typings or subspecies typings of the pathogen.
[0019] Preferably, in (S1), the selection criteria for the SNP sites with high consistency are: stably present in multiple strains of a certain bacterial species or subspecies, and not present in SNP sites of other genera, other species within the genus, or other subspecies.
[0020] Preferably, in (S1), the cgSNPs combination includes at least one SNP site.
[0021] In the present invention, the genes where the cgSNPs selected sites are located need to be widely covered and highly conserved among different strains to ensure the stability and reliability of typing; the cgSNPs combination needs to meet the resolution required for typing and be able to distinguish different typings.
[0022] Preferably, in (S2), the calculation formula for P(A) is: P(A) = the statistical number of the reference genome of a certain pathogen typing / the statistical number of the reference genomes of all typings.
[0023] In the present invention, P(A) is the prior probability, that is, the initial belief in the occurrence probability of a certain pathogen typing before observing other characteristics, and the probability of detecting a pathogen as a certain pathogen typing can be obtained through the above calculation formula according to the existing data set.
[0024] Preferably, in (S3), the calculation formula of P(B|A) is: P(B|A) = the statistical frequency of a certain cgSNPs combination in a certain pathogen typing / the statistical frequency of all cgSNPs combinations in a certain pathogen typing.
[0025] Preferably, in (S4), the calculation formula of P(B) is as follows:
[0026] P(B) = P(B|A)·P(A) + P(B|A′)·P(A′);
[0027] In the formula, P(B|A) is the occurrence frequency of the cgSNPs combination when the pathogen is of type A; P(A) is the proportion of the pathogen of type A in the data set; P(B|A') is the occurrence frequency of the cgSNPs combination when the pathogen is not of type A; P(A') is the proportion of the pathogen other than type A in the data set.
[0028] In the present invention, P(B) is the marginal probability, that is, the probability of observing a certain cgSNPs combination without considering the pathogen typing, and can be calculated by the above formula.
[0029] Preferably, the calculation formula of P(A') is: P(A') = 1 - P(A).
[0030] Preferably, P(B|A') is the sum of the probabilities of observing a certain cgSNPs combination for all pathogens that are not of type A.
[0031] Preferably, the calculation method of P(B|A') is: the statistical frequency of a certain cgSNPs combination in a certain pathogen that is not of type A / the statistical frequency of all cgSNPs combinations in a certain pathogen that is not of type A.
[0032] The present invention provides a method for pathogen identification based on a Bayesian model. The inference principle of the Bayesian model is: The Bayesian model is a class of statistical models based on Bayes' theorem, which can update and infer the probability distribution of unknown quantities by combining the observed data with prior knowledge.
[0033] By consulting literature, public databases, etc., specific homologous DNA sequences of pathogens that can be used for typing of closely related pathogens or pathogen subspecies are collected. According to the analysis, there are polymorphic SNP (single-nucleotide polymorphism) sites on different homologous DNA sequences of the same pathogen, but different homologous DNA sequences also have SNP sites with high consistency (core genome single-nucleotide polymorphism, cgSNP); usually, multiple cgSNPs can be combined to distinguish closely related pathogens or specific subtypes.
[0034] The decision analysis is performed using Bayesian theorem, mainly based on the known frequency of cgSNP and pathogen typing data set. When the sequencing data analysis results detect the site, the probability of occurrence of the detected pathogen typing is analyzed and derived through Bayesian statistics.
[0035] Conditional probability:
[0036]
[0037] Total probability formula:
[0038] P(B)=P(B|A)·P(A)+P(B|A′)·P(A′)
[0039]
[0040] Bayesian Inference:
[0041] Transform the conditional probability formula:
[0042]
[0043] P(A) is the prior probability, which is the judgment of event A before event B occurs; P(A|B) is the conditional probability, which is the probability of event A after event B occurs; P(B) is the marginal probability, which is the probability of observing event B under all possible events A. is the possibility function, and the adjustment makes the estimated probability closer to the true probability.
[0044] In the present invention, by consulting literature, public databases, etc., it was found that a single highly consistent SNP site cannot distinguish some pathogens with higher similarity, and cgSNPs in multiple target regions need to be combined. Therefore, cgSNPs and cgSNPs are combined in a multi-target manner to form a combination of cgSNPs (cgSNPs), including a single SNP site, multiple SNP sites in a certain region, and several regions.
[0045] That is, there are certain differences in the sequences of different strains of the same pathogen, and cgSNPs (single SNP sites or multiple SNP sites in a certain region or several regions) are almost present in different strains of the pathogen or pathogens closely related to it.
[0046] The tNGS method of ultra-multiplex PCR can effectively achieve the simultaneous detection of multiple targets. The sequencing method can directly read SNPs and can effectively detect cgSNPs. An example of the methodology for simultaneous detection of multiple targets is Figure 1 as shown.
[0047] In the case where the effective dataset is large enough, each cgSNPs can have extremely few "abnormal" genotypings. The finer the cgSNPs in the dataset can be divided, the finer the obtained probability will be.
[0048] In summary, the present invention combines the Bayesian model and multi-target simultaneous detection to achieve precise identification of pathogens.
[0049] In a second aspect, the present invention provides an application of the pathogen identification method based on the Bayesian model described in the first aspect in pathogen identification.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] (1) By combining prior knowledge with observed data through the Bayesian model, the present invention can provide more reliable inferences. When the observed data is insufficient, prior knowledge helps to make up for the lack of information and improve the accuracy of pathogen identification.
[0052] (2) When dealing with small sample data, the Bayesian model can still obtain meaningful inferences through prior knowledge and experimental functions. Especially in the case of rare pathogens or newly emerging genotypings, the data volume may be very limited. Using the Bayesian model, more reliable genotyping identification can be provided by combining prior knowledge and limited data.
[0053] (3) Through Bayes' rule, the present invention can dynamically update the prediction results of the model to adapt to new data. According to this feature, the technology can flexibly update the prediction parameters, improve the accuracy of diagnosis, and be used for the identification of newly emerging genotypings. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is an example of the methodology for simultaneous detection of multiple targets. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] The technical solutions of the present invention will be further described below through specific embodiments. Those skilled in the art should understand that the embodiments are only for helping to understand the present invention and should not be regarded as specific limitations to the present invention.
[0056] For those without specific technologies or conditions indicated in the examples, they shall be in accordance with the technologies or conditions described in the literature in this field or in accordance with the product specifications. For reagents or instruments without the manufacturer indicated, they are all conventional products that can be obtained through regular channels.
[0057] Example 1
[0058] Application of Mycobacterium identification
[0059] Genomes of 50 non-tuberculous mycobacteria, namely Mycobacterium avium, Mycobacterium intracellulare, Mycobacterium paraintracellulare, Mycobacterium chimaera, Mycobacterium lianjiangdongense, Mycobacterium colombiense, Mycobacterium massiliense, Mycobacterium timonense, Mycobacterium rhodesiae, Mycobacterium ulcerans, Mycobacterium abscessus subsp. abscessus, Mycobacterium abscessus subsp. massiliense, Mycobacterium abscessus subsp. bolletii, Mycobacterium chelonae, Mycobacterium gordonae, Mycobacterium kansasii, Mycobacterium persicum, Mycobacterium attenuatum, Mycobacterium gastri, Mycobacterium xenopi, Mycobacterium fortuitum, Mycobacterium suis, Mycobacterium nonchromogenicum, Mycobacterium marinum, Mycobacterium ulcerans, Mycobacterium szulgai, Mycobacterium malmoense, Mycobacterium scrofulaceum, Mycobacterium haemophilum, Mycobacterium smegmatis, Mycobacterium phlei, Mycobacterium margarethae, Mycobacterium cosmeticum, were obtained from the publicly available genomic databases. The prior probability of each species was taken as the proportion of all reference genomes of that species in the total.
[0060] The 16S RNA coding gene (16S rRNA), the intergenic region between 16S-23S rRNA (ITS), the B subunit of RNA polymerase (rpoB), and the heat shock protein 65 coding gene (hsp65) of the above non-tuberculous mycobacteria were analyzed to obtain the 16S rRNA, ITS, rpoB, and hsp65 genotype information of each non-tuberculous mycobacterium. The conditional probability P(B|A) of each different cgSNPs combination of each mycobacterium and the marginal probability P(B) of all cgSNPs combinations of each mycobacterium were calculated.
[0061] More than 6,000 mycobacterium genomes were collected in total. In the entire dataset, the proportions of different mycobacteria are shown in Table 1; a total of 51 groups of cgSNPs were analyzed and sorted out. The proportion distribution of each pathogen typing at each cgSNPs in this dataset is shown in Table 2.
[0062] Table 1
[0063]
[0064]
[0065] Table 2
[0066]
[0067]
[0068] In Table 2, taking Mycobacterium avium, Mycobacterium intracellulare, and Mycobacterium paraintracellulare as examples, the proportional distributions of each pathogen typing at each cgSNP were calculated respectively.
[0069] Taking the conditional probability P(cgSNPs1|Mycobacterium avium) of observing cgSNPs1 in the Mycobacterium avium typing as an example, its calculation method is as follows:
[0070] The statistical count of cgSNPs1 of Mycobacterium avium / The statistical count of all cgSNPs of Mycobacterium avium, that is, 1281 / 1306 = 0.98085758.
[0071] Derivation process:
[0072]
[0073] Mycobacterium avium purchased from the provincial preservation center was detected by tNGS. By comparing the sequencing data, the detected pathogen of the sample was Mycobacterium avium, and the cgSNP of the detected sample was cgSNP1. Then the probability calculation process of this sample being Mycobacterium avium is as follows:
[0074] Since,
[0075] P(Mycobacterium avium) = 0.23367329
[0076] P(cgSNPs a |Mycobacterium avium) = 0.98085758
[0077] P(cgSNPs a ) = P(cgSNPs a |Mycobacterium avium)·P(Mycobacterium avium) + P(cgSNPs a |not Mycobacterium avium)·P(not Mycobacterium avium)
[0078] = 0.98085758 * 0.23367329 + 0 * 0.76632671 = 0.22920022
[0079]
[0080] Based on the above analysis, the pathogen finally detected in this sample is Mycobacterium avium.
[0081] Example 2
[0082] Application of Influenza B Virus Subtyping
[0083] Genomes of two Influenza B viruses, namely Influenza B virus Victoria and Influenza B virus Yamagata, were obtained from the publicly available genomic databases. The prior probability of each species was taken as the proportion of all reference genomes of that species in the total.
[0084] The HA genes of the above Influenza B viruses were analyzed to obtain the HA genotype information of each Influenza B virus. And the conditional probability P(B|A) of each different cgSNPs combination of each Influenza B virus subspecies, as well as the marginal probability P(B) of all cgSNPs combinations of each Influenza B virus subspecies, were calculated.
[0085] More than 1800 Influenza B virus genomes were collected in total. In the entire dataset, the proportions of different Influenza B virus subspecies are shown in Table 3; a total of 5 groups of cgSNPs were analyzed and sorted out, and the proportion distribution of each pathogen typing at each cgSNPs in this dataset was statistically analyzed, as shown in Table 4.
[0086] Table 3
[0087]
[0088]
[0089] Table 4
[0090]
[0091] Influenza B virus Yamagata purchased from the Chinese Collection Center was detected by tNGS. By comparing the sequencing data, the detected pathogens in the sample were Influenza B virus Victoria and Influenza B virus Yamagata, and the cgSNPs of the detected sample were cgSNPs5.
[0092] Then the probability calculation process for this sample being Influenza B virus Victoria is as follows:
[0093] Since,
[0094] P(Influenza B virus Victoria) = 0.51964582
[0095] P(cgSNPs a |Influenza B virus Victoria) = 0.00209864
[0096] P(cgSNPsa ) = P(cgSNPs a | Influenza B virus Victoria) · P(Influenza B virus Victoria) + P(cgSNPs a | not Influenza B virus Victoria) · P(not Influenza B virus Victoria)
[0097] = 0.00209864 * 0.51964582 + 0.04617117 * 0.48035418
[0098] = 0.02326906
[0099]
[0100] The probability calculation process for the sample to be Influenza B virus Yamagata is as follows:
[0101] Since,
[0102] P(Influenza B virus Yamagata) = 0.48035418
[0103] P(cgSNPs a | Influenza B virus Yamagata) = 0.04617117
[0104] P(cgSNPs a ) = P(cgSNPs a | Influenza B virus Yamagata) · P(Influenza B virus Yamagata) + P(cgSNPs a | not Influenza B virus Yamagata) · P(not Influenza B virus Yamagata)
[0105] = 004617117 * 048035418 + 0.00209864 * 0.51964582
[0106] = 0.02326906
[0107]
[0108] Based on the above analysis, the pathogen finally detected in this sample is determined to be Influenza B virus Yamagata.
[0109] Example 3
[0110] Application of adenovirus typing
[0111] Obtain the genomes of 4 types of human adenoviruses, namely human adenovirus group B, human adenovirus group C, human adenovirus group D, and human adenovirus group E, from the publicly available genomic databases. The prior probability of each species is taken as the proportion of all reference genomes of that species in the overall population.
[0112] Analyze the E3 gene and L3 gene of the above adenoviruses to obtain the E3 and L3 genotype information of each adenovirus. Calculate the conditional probability P(B|A) of each different cgSNPs combination of each adenovirus subspecies, as well as the marginal probability P(B) of all cgSNPs combinations of each adenovirus subspecies.
[0113] A total of more than 100 adenovirus genomes were collected. In the entire dataset, the proportions of different adenovirus subspecies are shown in Table 5; a total of 4 groups of cgSNPs were analyzed and sorted out. The proportion distribution of each pathogen typing at each cgSNPs in this dataset is shown in Table 6.
[0114] Table 5
[0115]
[0116]
[0117] Table 6
[0118]
[0119] Detect the human adenovirus group B purchased from the Chinese Collection Center by tNGS. Compare the sequencing data to obtain that the detected pathogen of the sample is human adenovirus group B, and the cgSNPs of the detected sample is cgSNP1. The probability calculation process for this sample to be human adenovirus group B is as follows:
[0120] Since,
[0121] P(human adenovirus group B) = 0.46601942
[0122] P(cgSNPs1|human adenovirus group B) = 0.98085758
[0123] P(cgSNPs1) = P(cgSNPs1|human adenovirus group B)·P(human adenovirus group B) + P(cgSNPs1|not human adenovirus group B)·P(not human adenovirus group B)
[0124] = 0.95833333 * 0.46601942 + 0.09699697 * 0.53398058
[0125] = 0.49839644
[0126]
[0127] Based on the above analysis, the pathogen detected in the sample was finally determined to be human adenovirus group B.
[0128] Example 4
[0129] Application of rhinovirus genotyping
[0130] The genomes of three types of rhinoviruses, namely rhinovirus A, rhinovirus B, and rhinovirus C, were obtained from the publicly available genomic database. The prior probability of each species was taken as the proportion of all reference genomes of that species in the total.
[0131] The 5’UTR genes of the above rhinoviruses were analyzed to obtain the 5’UTR genotype information of each rhinovirus. The conditional probability P(B|A) of each different cgSNPs combination of each rhinovirus subspecies was calculated, as well as the marginal probability P(B) of all cgSNPs combinations of each rhinovirus subspecies.
[0132] More than 700 rhinovirus genomes were collected in total. In the entire dataset, the proportions of different rhinovirus subspecies are shown in Table 7; a total of 3 groups of cgSNPs were analyzed and sorted out, and the proportion distribution of each pathogen typing in each cgSNPs in this dataset was statistically analyzed, as shown in Table 8.
[0133] Table 7
[0134] Rhinovirus Proportion Rhinovirus A 0.55012853 Rhinovirus B 0.16323907 Rhinovirus C 0.28663239
[0135] Table 8
[0136]
[0137] The rhinovirus A purchased from the China National GeneBank was detected by tNGS. By comparing the sequencing data, the detected pathogen in the sample was rhinovirus A, and the cgSNPs of the detected sample was cgSNP1. The probability calculation process for this sample to be rhinovirus A is as follows:
[0138] Since,
[0139] P(rhinovirus A) = 0.55012853
[0140] P(cgSNPs1|rhinovirus A) = 0.98598131
[0141] P(cgSNPs1) = P(cgSNPs1|rhinovirus A)·P(rhinovirus A) + P(cgSNPs1|not rhinovirus A)·P(not rhinovirus A)
[0142] = 0.98598131 * 0.55012853 + 0.00787402 * 0.44987147
[0143] = 0.54595875
[0144]
[0145] In summary, the pathogen detected in this sample was finally determined to be rhinovirus A type.
[0146] Example 5
[0147] Application of enterovirus genotyping
[0148] Genomes of 8 enteroviruses, namely enterovirus group A, enterovirus A71, coxsackievirus group A, enterovirus group B, coxsackievirus group B, echovirus, enterovirus group C, and enterovirus group D, were obtained from publicly available genomic databases. The prior probability of each species was taken as the proportion of all reference genomes of that species in the total.
[0149] The 5'UTR genes of the above enteroviruses were analyzed to obtain the 5'UTR genotype information of each enterovirus. The conditional probability P(B|A) of each different cgSNPs combination of each enterovirus group and subspecies was calculated, as well as the marginal probability P(B) of all cgSNPs combinations of each enterovirus group and subspecies.
[0150] More than 2,000 enterovirus genomes were collected in total. In the entire dataset, the proportions of different enterovirus groups and subspecies are shown in Table 9; a total of 9 groups of cgSNPs were analyzed and sorted out. The proportion distribution of each pathogen typing in each cgSNPs in this dataset is shown in Tables 10 and 11.
[0151] Table 9
[0152]
[0153]
[0154] Table 10
[0155]
[0156] Table 11
[0157]
[0158]
[0159] The enterovirus A71 purchased from the China Collection Center was detected by tNGS. By comparing the sequencing data, the detected pathogens in the sample were enterovirus group A and enterovirus A71, and the cgSNPs of the detected sample were cgSNP1.
[0160] Then the probability calculation process for this sample to be enterovirus group A is as follows:
[0161] Since,
[0162] P(Enterovirus A) = 0.43082226
[0163] P(cgSNPs1|Enterovirus A) = 0.49828571
[0164] P(cgSNPs1) = P(cgSNPs1|Enterovirus A)·P(Enterovirus A)
[0165] = P(cgSNPs1|not Enterovirus A)·P(not Enterovirus A)
[0166] = 0.49828571 * 0.43082226 + 0 * 0.55012853 = 0.21467258
[0167]
[0168] The probability calculation process for the sample to be Enterovirus A71 is as follows:
[0169] Since,
[0170] P(Enterovirus A71) = 0.10388971
[0171] P(cgSNPs1|Enterovirus A71) = 0.99052133
[0172] P(cgSNPs1) = P(cgSNPs1|Enterovirus A71)·P(Enterovirus A71) + P(cgSNPs1|not Enterovirus A71)·P(not Enterovirus A71)
[0173] = 0.99052133 * 0.10388971 + 0.00153610 * 0.89611029
[0174] = 0.10428149
[0175]
[0176] Enterovirus A71 belongs to Enterovirus A. Based on the above analysis, it is finally determined that the detected pathogen is Enterovirus A71 and Enterovirus A.
[0177] Example 6
[0178] Application of co - infection identification
[0179] Mycobacterium abscessus, Mycobacterium avium, and Mycobacterium kansasii purchased from the provincial collection center were selected and mixed in a 1:1 ratio. tNGS library construction and sequencing were performed, and the constructed Bayesian model was used for bacterial species typing to evaluate its detection performance for mixed infections.
[0180] The results are shown in Table 12: The Bayesian model-based method can accurately identify mixed infection strains.
[0181] Table 12
[0182]
[0183] In summary, the present invention combines the tNGS method, Bayesian theorem and multi-target combination to perform decision analysis on pathogens and achieve accurate identification of pathogens.
[0184] The applicant declares that the above is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Those skilled in the art should understand that any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention are within the protection scope and disclosure scope of the present invention.
Claims
1. A pathogen identification method based on a Bayesian model, characterized in that: The method comprises: (S1) Specific homologous DNA sequences of pathogens that can be used for typing of closely related pathogens or pathogen subspecies are collected from the database, and SNP sites with high consistency are screened from them, and the SNP sites are divided into different cgSNPs combinations based on different pathogen typing; (S2) Calculate the proportion of the reference genome of a certain pathogen typing in all the reference genomes of typing as the prior probability P(A); (S3) organizing the data of cgSNPs combination corresponding to each pathogen typing into a data set, wherein the data set includes different cgSNPs combination types and the number of occurrences of different cgSNPs combination types in different pathogen typing; and calculating the distribution ratio of each pathogen typing in each cgSNPs combination type in the data set as the conditional probability P(B|A); (S4) Based on the P(B|A) and P(A) values, by Bayesian statistics, when a certain cgSNPs combination type is detected in the sequencing data analysis results, the probability of occurrence of the detected pathogen typing is calculated; In the formula, P(A) is the prior probability, which is the judgment of event A before event B occurs; P(A|B) is the conditional probability, which is the probability of event A after event B occurs; P(B) is the marginal probability, which is the probability of observing event B under all possible events A. is the possibility function, and the adjustment makes the estimated probability closer to the true probability.
2. The pathogen identification method based on the Bayesian model according to claim 1, characterized in that: In (S1), the specific homologous DNA sequence used for typing of closely related pathogens or typing of pathogen subspecies is a collection of genes carried in all typing or subspecies typing of the pathogen.
3. The pathogen identification method based on the Bayesian model according to claim 1 or 2, characterized in that: In (S1), the selection criteria for the SNP sites with higher consistency are: SNP sites that are stably present in multiple strains of a certain species or subspecies and do not exist in other genera, other species within a genus, or other subspecies.
4. The pathogen identification method based on the Bayesian model according to any one of claims 1 to 3, characterized in that: In (S1), the cgSNPs combination includes at least one SNP site.
5. The pathogen identification method based on the Bayesian model according to any one of claims 1 to 4, characterized in that: In (S2), the calculation formula of P(A) is: P(A)=the statistical number of reference genomes of a certain pathogen typing / the statistical number of reference genomes of all typings.
6. The pathogen identification method based on the Bayesian model according to any one of claims 1 to 5, characterized in that: In (S3), the calculation formula of P(B|A) is: P(B|A)=the statistical number of a certain cgSNPs combination in a certain pathogen typing / the statistical number of all cgSNPs combinations in a certain pathogen typing.
7. The pathogen identification method based on the Bayesian model according to any one of claims 1 to 6, characterized in that: In (S4), the calculation formula of P(B) is as follows: P(B)=P(B|A)·P(A)+·P(B|A′)·P(A′); Where, P(B|A) is the frequency of occurrence of the cgSNPs combination when the pathogen is type A; P(A) is the proportion of the pathogen in the data set when the pathogen is type A; P(B|A') is the frequency of occurrence of the cgSNPs combination when the pathogen is not type A; P(A') is the proportion of the pathogen in the data set when the pathogen is other than type A.
8. The pathogen identification method based on the Bayesian model according to claim 7, characterized in that: The calculation formula of P(A') is: P(A')=1-P(A).
9. The pathogen identification method based on the Bayesian model according to claim 7 or 8, characterized in that: The P(B|A') is the sum of the probabilities of observing a certain cgSNPs combination for all pathogens that are not typed as A; Preferably, the calculation method of P(B|A') is: the statistical number of a certain cgSNPs combination in a pathogen that is not typed A / the statistical number of all cgSNPs combinations in a pathogen that is not typed A.
10. Application of the pathogen identification method based on the Bayesian model according to any one of claims 1 to 9 in pathogen identification.
Citation Information
Patent Citations
Genome identification system
CN102007407A
Genotype correction device and method
CN109785899A
Method and device for improving metagenome species identification accuracy, storage medium and application thereof
CN117831614A
Direct identification and measurement of relative populations of microorganisms with direct DNA sequencing and probabilistic methods
US20120004111A1
Systems and methods for characterization of viability and infection risk of microbes in the environment
WO2017156431A1