A pathogen identification method based on a bayesian model and application thereof

By combining Bayesian models with tNGS and multi-target synergistic techniques, SNP sites of specific homologous DNA sequences are screened, solving the difficulty of identifying pathogens with high similarity and variation, and achieving accurate identification of pathogens and mixed infections.

CN120148644BActive Publication Date: 2025-11-18GUANGZHOU JINQIRUI BIOTECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510230736.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-11-18
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

Existing pathogen identification methods have difficulties in identifying pathogens with high similarity and variants, making it difficult to achieve high sensitivity and high resolution identification, especially in cases of mixed infections where accurate identification is challenging.

Method used

A pathogen identification method based on a Bayesian model, combining tNGS methodology and multi-target combination technology, is adopted. By screening SNP sites of specific homologous DNA sequences, Bayes' theorem is used for decision analysis to calculate prior probabilities and conditional probabilities, thereby achieving accurate identification of pathogens.

Benefits of technology

It improves the accuracy and sensitivity of pathogen identification, effectively identifies mixed infections, adapts to small sample data and emerging subtypes, and dynamically updates the model to improve diagnostic accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148644B_ABST
    Figure CN120148644B_ABST
Patent Text Reader

Abstract

The application provides a pathogen identification method based on a Bayesian model and application thereof, and the method comprises the following steps: collecting DNA information of a pathogen, screening SNP sites with higher consistency from the DNA information, and dividing the SNP sites into different cgSNPs combination types based on different pathogen typing; calculating the proportion of the reference genome of a certain pathogen typing in the reference genomes of all typing; calculating the proportion distribution of various pathogen typing in various cgSNPs combination types; and calculating the occurrence probability of the detected pathogen typing under the condition that a certain cgSNPs combination type is detected in the sequencing data analysis result through Bayesian statistics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of high-throughput sequencing and gene detection technology, specifically relating to a pathogen identification method based on a Bayesian model and its application. Background Technology

[0002] Currently, the two most common methods for determining whether two bacterial genomes belong to the same species are: average nucleotide identity (ANI) and 16S rRNA gene sequence similarity. It is generally believed that when the ANI value reaches 95% or higher, or the 16S rRNA gene sequence similarity reaches 98.65%, the sequences of the two strains can be considered to belong to the same species.

[0003] The current gold standard for bacterial species identification is the direct homologous gene or sequence comparison method, which identifies bacteria to the species level by analyzing differences in homologous DNA sequence composition.

[0004] Whole-genome sequencing (WGS) is a superior method for pathogen typing; however, WGS requires pathogen cultures, and some strains have drawbacks such as long culturing times or difficulty in culturing, limiting its universality. Metagenomics Next Generation Sequencing (mNGS) can directly detect clinical samples for pathogen typing, but it also suffers from drawbacks such as potentially incomplete sequencing coverage and high cost, similarly limiting its universality. Amplifying and sequencing specific homologous DNA sequences of pathogens in clinical samples is a more economical and universally applicable approach.

[0005] However, in actual pathogen typing and detection, some pathogens exhibit high genomic similarity (over 98%, even 99%) for specific homologous DNA sequences, or extremely high genomic similarity (over 99.3%) for specific homologous DNA sequences of pathogen subspecies. Furthermore, pathogens are prone to mutation, such as mycobacteria, nocardia, influenza A virus, SARS-CoV-2, and adenovirus. This can lead to situations where, during homologous DNA sequence analysis, mutations may cause sequences to simultaneously align with pathogens A, B, and C, resulting in unclassifiable pathogen sequencing results. Additionally, some pathogens are difficult to distinguish using only a few homologous DNA sequences.

[0006] Most current methods for identifying mycobacteria are based on multiplex PCR technology, including real-time quantitative PCR, isothermal amplification, probe-reverse hybridization, and probe-melting curve techniques. These methods are convenient and fast, but these reagents usually have limited targets and can only cover a limited number of typing sites. Therefore, they can identify fewer pathogens and it is difficult to identify subspecies. Some closely related pathogens cannot be accurately identified.

[0007] tNGS (Targeted Next Generation Sequencing) is a next-generation sequencing technology that combines ultra-multiplex PCR amplification with high-throughput sequencing. It uses a large number of primers targeting specific gene sequences to perform ultra-multiplex PCR amplification of nucleic acids in the sample, obtaining a large number of target nucleic acid fragments, followed by high-throughput sequencing. The resulting sequences are then analyzed using bioinformatics, achieving high-sensitivity and high-resolution identification. Therefore, developing an efficient and accurate pathogen identification method combining tNGS methodology has significant application value. Summary of the Invention

[0008] To address the shortcomings of existing technologies, the present invention aims to provide a pathogen identification method based on a Bayesian model and its application. This invention combines direct homologous gene or sequence comparison methods from tNGS methodology, Bayes' theorem, and multi-target synergy to perform decision analysis on pathogens, thereby achieving accurate pathogen identification with high sensitivity, good specificity, and effective identification of mixed infections.

[0009] To achieve this objective, the present invention adopts the following technical solution:

[0010] In a first aspect, the present invention provides a pathogen identification method based on a Bayesian model, the method comprising:

[0011] (S1) Collect specific homologous DNA sequences of pathogens from the database that can be used for close pathogen typing or pathogen subspecies typing, screen for SNP sites with high consistency, and divide the SNP sites into different cgSNPs combinations based on different pathogen typing.

[0012] (S2) Calculate the proportion of the reference genome of a certain pathogen type in the reference genome of all types, as the prior probability P(A);

[0013] (S3) Organize the data of cgSNPs combinations corresponding to each pathogen subtype into a dataset. The dataset includes different cgSNPs combination types and the number of times different cgSNPs combination types appear in different pathogen subtypes. Statistically analyze the proportion distribution of each pathogen subtype in each cgSNPs combination type in the dataset as the conditional probability P(B|A).

[0014] (S4) Based on P(B|A) and P(A) values, the probability of pathogen typing is calculated by Bayesian statistics when a certain combination of cgSNPs is detected in the sequencing data analysis results.

[0015]

[0016] In the formula, P(A) is the prior probability, the judgment of event A before event B occurs; P(A|B) is the conditional probability, that is, the probability of A after event B occurs; P(B) is the marginal probability, that is, the probability of observing event B under all possible events A. Let be the probability function, and adjust it to make the estimated probability closer to the true probability.

[0017] In this invention, when the P(A|B) value is greater than 0.85, the pathogen typing is considered to have high reliability, and when it is greater than 0.95, the pathogen typing is considered to be reliable.

[0018] Preferably, in (S1), the specific homologous DNA sequence used for close pathogen typing or pathogen subtyping is a set of genes carried in all pathogen typing or subtyping.

[0019] Preferably, in (S1), the selection criteria for the SNP sites with high consistency are: they are stably present in multiple strains of a certain species or subspecies, and are not present in SNP sites of other genera, other species within the same genus, or other subspecies.

[0020] Preferably, in (S1), the cgSNPs combination includes at least one SNP site.

[0021] In this invention, the genes containing the selected cgSNP sites need to be widely covered and highly conserved across different strains to ensure stable and reliable typing; the cgSNP combinations need to meet the resolution required for typing and be able to distinguish between different types.

[0022] Preferably, in (S2), the formula for calculating P(A) is: P(A) = the number of times the reference genome of a certain pathogen subtype is counted / the number of times the reference genome of all subtypes is counted.

[0023] In this invention, P(A) is the prior probability, which is the initial belief about the probability of a certain pathogen type occurring before other characteristics are observed. The probability of detecting a pathogen as a certain pathogen type can be obtained from existing datasets using the above calculation formula.

[0024] Preferably, in (S3), the formula for calculating P(B|A) is: P(B|A) = the number of times a certain cgSNP combination is counted in a certain pathogen subtype / the number of times all cgSNP combinations are counted in a certain pathogen subtype.

[0025] Preferably, in (S4), the calculation formula for P(B) is as follows:

[0026] P(B)=P(B|A)·P(A)+P(B|A′)·P(A′);

[0027] In the formula, P(B|A) is the frequency of cgSNPs combination when the pathogen is classified as A; P(A) is the proportion of the pathogen in the dataset when it is classified as A; P(B|A') is the frequency of cgSNPs combination when the pathogen is not classified as A; and P(A') is the proportion of the pathogen in the dataset when it is classified as something other than A.

[0028] In this invention, P(B) is the marginal probability, which is the probability of observing a certain combination of cgSNPs without considering pathogen typing. It can be calculated using the above formula.

[0029] Preferably, the formula for calculating P(A') is: P(A') = 1 - P(A).

[0030] Preferably, P(B|A') is the sum of probabilities of observing a certain combination of cgSNPs for all pathogens that are not of type A.

[0031] Preferably, the calculation method of P(B|A') is: the number of times a certain cgSNP combination in a pathogen that is not of type A / the number of times all cgSNP combinations in a pathogen that is not of type A.

[0032] This invention provides a method for pathogen identification based on a Bayesian model. The inference principle of the Bayesian model is as follows: the Bayesian model is a type of statistical model based on Bayes' theorem, which can update and infer the probability distribution of unknown quantities by combining observed data with prior knowledge.

[0033] By reviewing literature and public databases, specific homologous DNA sequences of pathogens that can be used for close-related pathogen typing or pathogen subspecies typing were collected. According to the analysis, polymorphic SNP (single-nucleotide polymorphism) sites exist on different homologous DNA sequences of the same pathogen, but different homologous DNA sequences also have SNP sites with high consistency (core genome single-nucleotide polymorphism, cgSNP). Usually, multiple cgSNPs can be combined to distinguish close-related pathogens or specific subtypes.

[0034] Decision analysis using Bayes' theorem mainly involves analyzing and deriving the probability of detecting a pathogen type when the sequencing data analysis results show that a cgSNP has a known occurrence frequency and pathogen typing, using Bayesian statistics.

[0035] Conditional probability:

[0036]

[0037] The law of total probability:

[0038] P(B)=P(B|A)·P(A)+P(B|A′)·P(A′)

[0039]

[0040] Bayesian inference:

[0041] Transform the conditional probability formula:

[0042]

[0043] P(A) is the prior probability, the judgment of event A before event B occurs; P(A|B) is the conditional probability, that is, the probability of A after event B occurs; P(B) is the marginal probability, that is, the probability of observing event B under all possible events A. Let be the probability function, and adjust it to make the estimated probability closer to the true probability.

[0044] In this invention, by consulting literature and public databases, it was found that a single highly consistent SNP site cannot distinguish some pathogens with high similarity. It is necessary to combine cgSNPs from multiple target regions. Therefore, by combining cgSNPs with multiple targets, cgSNPs are combined to form cgSNP combinations (cgSNPs), which include a single SNP site, multiple SNP sites in a certain region, and several regions.

[0045] That is, different strains of the same pathogen have certain differences in sequence, and cgSNPs (single SNP sites or multiple SNP sites or several regions in a certain region) are present in almost all different strains of the pathogen or closely related pathogens.

[0046] The tNGS method in ultramultiplex PCR can effectively achieve multi-target synergy, and sequencing methods can directly read SNPs, effectively detecting cgSNPs. Methodological examples of multi-target synergy are as follows... Figure 1 As shown.

[0047] When the effective dataset is large enough, each cgSNP can exhibit a very small number of "abnormal" subtypes. The more finely the cgSNPs in the dataset can be divided, the more precise the probabilities will be.

[0048] In summary, this invention combines Bayesian models and multi-target synergy to achieve accurate identification of pathogens.

[0049] Secondly, the present invention provides the application of the pathogen identification method based on the Bayesian model described in the first aspect in pathogen identification.

[0050] Compared with the prior art, the present invention has the following beneficial effects:

[0051] (1) This invention combines prior knowledge with observational data through a Bayesian model, which can provide more reliable inferences. When observational data is insufficient, prior knowledge helps to make up for the lack of information and improve the accuracy of pathogen identification.

[0052] (2) When dealing with small sample data, Bayesian models can still make meaningful inferences through prior knowledge and experimental functions. Especially in the case of rare pathogens or newly emerging subtypes, the amount of data may be very limited. Bayesian models can provide more reliable subtype identification through prior knowledge of tuberculosis and limited data.

[0053] (3) This invention can dynamically update the model’s prediction results through Bayes’ rule to adapt to new data. Based on this feature, the technology can flexibly update prediction parameters, improve the accuracy of diagnosis, and be used to identify newly emerging subtypes. Attached Figure Description

[0054] Figure 1 This is a methodological example of multi-target synergy. Detailed Implementation

[0055] The technical solution of the present invention will be further illustrated below through specific embodiments. Those skilled in the art should understand that the embodiments described are merely illustrative of the present invention and should not be construed as limiting the invention in any way.

[0056] Where specific techniques or conditions are not specified in the examples, they shall be performed in accordance with the techniques or conditions described in the literature in this field, or in accordance with the product instructions. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased through legitimate channels.

[0057] Example 1

[0058] Application of mycobacterial identification

[0059] The following mycobacteria were obtained from publicly available genome databases: Mycobacterium avium, Mycobacterium intracellulare, Mycobacterium paraintracellulare, Mycobacterium chimera, Mycobacterium licheniformis, Mycobacterium coronarium, Mycobacterium massa, Mycobacterium timonene, Mycobacterium rosenbergii, Mycobacterium woundum, Mycobacterium abscessum subsp. abscessum, Mycobacterium abscessum subsp. massa, Mycobacterium abscessum subsp. borax, Mycobacterium turcica, Mycobacterium Gordon, Mycobacterium Kansas, Mycobacterium Persianum, Mycobacterium attenuation, Mycobacterium gastrum, Mycobacterium bufotatum, Mycobacterium occulta, Mycobacterium suis, Mycobacterium exogenum, Mycobacterium septicemia, and Mycobacterium brisbaneum. The genomes of 50 nontuberculous mycobacteria were collected, including *Mycobacterium aureum*, *Mycobacterium neo-auspiciousum*, *Mycobacterium difficile*, *Mycobacterium bovis*, *Mycobacterium simianum*, *Mycobacterium stoloniferum*, *Mycobacterium cerevisiae*, *Mycobacterium aquilinum*, *Mycobacterium tumefaciens*, *Mycobacterium tumefaciens*, *Mycobacterium kusmotoum*, *Mycobacterium minorum*, *Mycobacterium nonchromogenicum*, *Mycobacterium marinum*, *Mycobacterium ulcerans*, *Mycobacterium sulgare*, *Mycobacterium margaritiferum*, *Mycobacterium scrofula*, *Mycobacterium hemoglobinum*, *Mycobacterium smegmatis*, *Mycobacterium styridis*, *Mycobacterium margaritiferum*, and *Mycobacterium cosmetum*. The prior probability of each species was calculated as the proportion of that species' entire reference genome in the total population.

[0060] The 16S RNA encoding gene (16S rRNA), the 16S-23S rRNA intergenic region (ITS), the RNA polymerase B subunit (rpoB), and the heat shock protein 65 encoding gene (hsp65) of the above nontuberculous mycobacteria were analyzed to obtain the genotype information of 16S rRNA, ITS, rpoB, and hsp65 for each nontuberculous mycobacterium. The conditional probability P(B|A) of each different cgSNP combination for each mycobacterium and the marginal probability P(B) of all cgSNP combinations for each mycobacterium were calculated.

[0061] A total of over 6,000 mycobacterial genomes were collected. The proportion of different mycobacteria in the entire dataset is shown in Table 1. A total of 51 cgSNPs were analyzed and compiled. The distribution of the proportion of each pathogen type in each cgSNP in the dataset is shown in Table 2.

[0062] Table 1

[0063]

[0064]

[0065] Table 2

[0066]

[0067]

[0068] Table 2 shows the proportion of each pathogen type in each cgSNP, taking Mycobacterium avium, Mycobacterium intracellulare, and Mycobacterium paraintracellulare as examples.

[0069] Taking the conditional probability P(cgSNPs1|Mycobacterium avium) of observing cgSNPs1 in Mycobacterium avium typing as an example, its calculation method is as follows:

[0070] The number of cgSNPs1 in Mycobacterium avium / the total number of cgSNPs in Mycobacterium avium, i.e., 1281 / 1306 = 0.98085758.

[0071] Derivation process:

[0072]

[0073] The avian mycobacteria purchased from the provincial collection center were tested using tNGS. The sequencing data showed that the pathogen detected in the sample was avian mycobacteria, and the cgSNPs detected in the sample were cgSNP1. The probability calculation process for this sample being avian mycobacteria is as follows:

[0074] because,

[0075] P(Mycobacterium avium) = 0.23367329

[0076] P(cgSNPs a |Mycobacterium avium) = 0.98085758

[0077] P(cgSNPs a )=P(cgSNPs a |Mycobacterium avium)·P(Mycobacterium avium)+P(cgSNPs a |Not Mycobacterium avium)·P(Not Mycobacterium avium)

[0078] =0.98085758*0.23367329+0*0.76632671=0.22920022

[0079]

[0080] Based on the above analysis, the pathogen detected in this sample was ultimately determined to be Mycobacterium avium.

[0081] Example 2

[0082] Application of influenza B virus typing

[0083] The genomes of two influenza B viruses, Victoria and Yamagata, were obtained from publicly available genome databases. The prior probability of each species was calculated as the proportion of all reference genomes of that species in the total genome.

[0084] The HA gene of the above influenza B viruses was analyzed to obtain the HA genotype information of each influenza B virus. The conditional probability P(B|A) of each different cgSNP combination for each influenza B virus subspecies, and the marginal probability P(B) of all cgSNP combinations for each influenza B virus subspecies were calculated.

[0085] A total of over 1,800 influenza B virus genomes were collected. The proportion of different influenza B virus subspecies in the entire dataset is shown in Table 3. Five groups of cgSNPs were analyzed and compiled. The distribution of the proportion of each pathogen type in each cgSNP in the dataset is shown in Table 4.

[0086] Table 3

[0087]

[0088]

[0089] Table 4

[0090]

[0091] The influenza B virus Yamagata purchased from the China National Archives was tested using tNGS. The sequencing data showed that the pathogens detected in the sample were influenza B virus Victoria and influenza B virus Yamagata, and the cgSNPs detected in the sample were cgSNPs5.

[0092] The probability calculation process for this sample being influenza B virus Victoria is as follows:

[0093] because,

[0094] P(Influenza B virus Victoria) = 0.51964582

[0095] P(cgSNPs a |Influenza B virus (Victoria) = 0.00209864

[0096] P(cgSNPsa )=P(cgSNPs a |Influenza B virus Victoria)·P(Influenza B virus Victoria)+P(cgSNPs a |Not a type B influenza virus (Victoria)·P (Not a type B influenza virus (Victoria))

[0097] =0.00209864*0.51964582+0.04617117*0.48035418

[0098] =0.02326906

[0099]

[0100] The probability calculation process for this sample being the influenza B virus Yamagata is as follows:

[0101] because,

[0102] P(Influenza B virus Yamagata) = 0.48035418

[0103] P(cgSNPs a |Influenza B virus (Yamagata) = 0.04617117

[0104] P(cgSNPs a )=P(cgSNPs a |Influenza B virus Yamagata)·P(Influenza B virus Yamagata)+P(cgSNPs) a (Not a type B influenza virus, Yamagata)·P (Not a type B influenza virus, Yamagata)

[0105] =004617117*048035418+0.00209864*0.51964582

[0106] =0.02326906

[0107]

[0108] Based on the above analysis, the pathogen detected in this sample was ultimately determined to be influenza B virus Yamagata.

[0109] Example 3

[0110] Application of adenovirus typing

[0111] The genomes of four human adenovirus groups (groups B, C, D, and E) were obtained from publicly available genome databases. The prior probability of each species was calculated as the proportion of all reference genomes of that species in the total genome.

[0112] The E3 and L3 genes of the above adenoviruses were analyzed to obtain the E3 and L3 genotype information of each adenovirus. The conditional probability P(B|A) of each different cgSNP combination for each adenovirus subspecies, and the marginal probability P(B) of all cgSNP combinations for each adenovirus subspecies were calculated.

[0113] More than 100 adenovirus genomes were collected. The proportion of different adenovirus subspecies in the entire dataset is shown in Table 5. Four groups of cgSNPs were analyzed and compiled. The distribution of the proportion of each pathogen type in each cgSNP in the dataset is shown in Table 6.

[0114] Table 5

[0115]

[0116]

[0117] Table 6

[0118]

[0119] The human adenovirus group B virus purchased from the China National Archives was detected using tNGS. The sequencing data showed that the pathogen detected in the sample was human adenovirus group B, and the cgSNPs detected in the sample were cgSNP1. The probability calculation process for this sample being human adenovirus group B is as follows:

[0120] because,

[0121] P(human adenovirus group B) = 0.46601942

[0122] P(cgSNPs1|human adenovirus group B)=0.98085758

[0123] P(cgSNPs1) = P(cgSNPs1|human adenovirus group B)·P(human adenovirus group B) + P(cgSNPs1|non-human adenovirus group B)·P(non-human adenovirus group B)

[0124] =0.95833333*0.46601942+0.09699697*0.53398058

[0125] =0.49839644

[0126]

[0127] Based on the above analysis, the pathogen detected in this sample was ultimately determined to be human adenovirus group B.

[0128] Example 4

[0129] Application of rhinovirus typing

[0130] The genomes of three rhinoviruses—rhinovirus A, rhinovirus B, and rhinovirus C—were obtained from publicly available genome databases. The prior probability for each species was calculated as the proportion of all reference genomes for that species in the total genome.

[0131] The 5'UTR genes of the above rhinoviruses were analyzed to obtain the 5'UTR genotype information of each rhinovirus. The conditional probability P(B|A) of each different cgSNP combination in each rhinovirus subspecies, and the marginal probability P(B) of all cgSNP combinations in each rhinovirus subspecies were calculated.

[0132] More than 700 rhinovirus genomes were collected. The proportion of different rhinovirus subspecies in the entire dataset is shown in Table 7. Three groups of cgSNPs were analyzed and compiled. The distribution of the proportion of each pathogen type in each cgSNP in the dataset is shown in Table 8.

[0133] Table 7

[0134] rhinovirus Proportion Rhinovirus type A 0.55012853 Rhinovirus type B 0.16323907 Rhinovirus type C 0.28663239

[0135] Table 8

[0136]

[0137] The rhinovirus type A sample purchased from the China National Archives was tested using tNGS. The sequencing data showed that the pathogen detected in the sample was rhinovirus type A, and the cgSNPs detected in the sample were cgSNP1. The probability calculation process for this sample being rhinovirus type A is as follows:

[0138] because,

[0139] P(rhinovirus A) = 0.55012853

[0140] P(cgSNPs1|rhinovirus type A)=0.98598131

[0141] P(cgSNPs1) = P(cgSNPs1|rhinovirus A)·P(rhinovirus A) + P(cgSNPs1|not rhinovirus A)·P(not rhinovirus A)

[0142] =0.98598131*0.55012853+0.00787402*0.44987147

[0143] =0.54595875

[0144]

[0145] Based on the above analysis, the pathogen detected in this sample was ultimately identified as rhinovirus type A.

[0146] Example 5

[0147] Application of enterovirus typing

[0148] The genomes of eight enteroviruses—Enterovirus A, Enterovirus A71, Coxsackievirus A, Enterovirus B, Coxsackievirus B, Echovirus, Enterovirus C, and Enterovirus D—were obtained from publicly available genome databases. The prior probability of each species was calculated as the proportion of all reference genomes of that species in the total population.

[0149] The 5'UTR genes of the above enteroviruses were analyzed to obtain the 5'UTR genotype information of each enterovirus. The conditional probability P(B|A) of each different cgSNP combination for each enterovirus group and subspecies, and the marginal probability P(B) of all cgSNP combinations for each enterovirus group and subspecies were calculated.

[0150] More than 2,000 enterovirus genomes were collected. The proportion of different enterovirus groups and subspecies in the entire dataset is shown in Table 9. A total of 9 cgSNPs were analyzed and compiled. The distribution of the proportion of each pathogen type in each cgSNP in the dataset is shown in Tables 10 and 11.

[0151] Table 9

[0152]

[0153]

[0154] Table 10

[0155]

[0156] Table 11

[0157]

[0158]

[0159] Enterovirus A71 purchased from the China National Archives was tested using tNGS. The sequencing data showed that the pathogens detected in the samples were enterovirus group A and enterovirus A71, and the cgSNPs detected in the samples were cgSNP1.

[0160] The probability calculation process for this sample being enterovirus group A is as follows:

[0161] because,

[0162] P(Enterovirus A group) = 0.43082226

[0163] P(cgSNPs1|Enterovirus A group)=0.49828571

[0164] P(cgSNPs1) = P(cgSNPs1|Enterovirus A group)·P(Enterovirus A group)

[0165] = P(cgSNPs1|not a group A enterovirus)·P(not a group A enterovirus)

[0166] =0.49828571*0.43082226+0*0.55012853=0.21467258

[0167]

[0168] The probability calculation process for this sample being enterovirus A71 is as follows:

[0169] because,

[0170] P (Enterovirus A71) = 0.10388971

[0171] P(cgSNPs1|Enterovirus A71) = 0.99052133

[0172] P(cgSNPs1) = P(cgSNPs1|Enterovirus A71)·P(Enterovirus A71) + P(cgSNPs1|Not Enterovirus A71)·P(Not Enterovirus A71)

[0173] =0.99052133*0.10388971+0.00153610*0.89611029

[0174] =0.10428149

[0175]

[0176] Enterovirus A71 belongs to group A of enteroviruses. Based on the above analysis, the pathogen was finally identified as enterovirus A71, group A of enteroviruses.

[0177] Example 6

[0178] Application of mixed infection identification

[0179] Mycobacterium abscessus, Mycobacterium avium, and Mycobacterium kansasum purchased from the provincial collection center were mixed in a 1:1 ratio, and tNGS library construction and sequencing were performed. The constructed Bayesian model was used for bacterial typing to evaluate its detection performance for mixed infections.

[0180] The results are shown in Table 12: the Bayesian model-based method can accurately identify mixed-infection strains.

[0181] Table 12

[0182]

[0183] In summary, this invention combines the tNGS method, Bayes' theorem, and multi-target syntactic analysis to perform decision analysis on pathogens, thereby achieving accurate identification of pathogens.

[0184] The applicant declares that the above description is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Those skilled in the art should understand that any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention fall within the protection and disclosure scope of the present invention.

Claims

1. A pathogen identification method based on a Bayesian model, characterized in that, The method includes: (S1) Collect specific homologous DNA sequences of pathogens that can be used for close pathogen typing or pathogen subspecies typing from the database, screen for SNP sites with high consistency, and divide the SNP sites into different cgSNPs combinations based on different pathogen typing. (S2) Calculate the proportion of the reference genome of a certain pathogen type in the reference genome of all types, as the prior probability P(A); (S3) Organize the data of cgSNPs combinations corresponding to each pathogen subtype into a dataset. The dataset includes different cgSNPs combination types and the number of times different cgSNPs combination types appear in different pathogen subtypes. Statistically analyze the proportion distribution of each pathogen subtype in each cgSNPs combination type in the dataset as the conditional probability P(B|A). (S4) Based on P(B|A) and P(A) values, the probability of pathogen typing is calculated by Bayesian statistics when a certain combination of cgSNPs is detected in the sequencing data analysis results. ; In the formula, P(A) is the prior probability, the judgment of event A before event B occurs; P(A|B) is the conditional probability, that is, the probability of A after event B occurs; P(B) is the marginal probability, that is, the probability of observing event B under all possible events A. The probability function is adjusted to make the estimated probability closer to the true probability. In (S1), the specific homologous DNA sequence used for the typing of closely related pathogens or the typing of pathogen subspecies is a set of genes carried in all types or subtypes of pathogens; In (S1), the selection criteria for the SNP sites with high consistency are: they are stably present in multiple strains of a certain species or subspecies, and are not present in SNP sites of other genera, other species within the same genus, or other subspecies. In (S2), the formula for calculating P(A) is: P(A) = Number of reference genome statistics for a certain pathogen subtype / Number of reference genome statistics for all subtypes; In (S3), the calculation formula for P(B|A) is: P(B|A) = the number of times a certain cgSNP combination is counted in a certain pathogen type / the number of times all cgSNP combinations are counted in a certain pathogen type; In (S4), the formula for calculating P(B) is as follows: ; In the formula, P(B|A) is the frequency of cgSNPs combination when the pathogen is classified as A; P(A) is the proportion of the pathogen in the dataset when it is classified as A; P(B|A') is the frequency of cgSNPs combination when the pathogen is not classified as A; and P(A') is the proportion of the pathogen in the dataset when it is classified as something other than A. The formula for calculating P(A') is: P(A') = 1 - P(A); The P(B|A') is the sum of probabilities of observing a certain combination of cgSNPs for all pathogens that are not of type A. The calculation method for P(B|A') is: the number of times a certain cgSNP combination is found in a pathogen that is not classified as A / the number of times all cgSNP combinations are found in a pathogen that is not classified as A.

2. The pathogen identification method based on a Bayesian model according to claim 1, characterized in that, In (S1), the cgSNPs combination includes at least one SNP site.

3. The application of the Bayesian model-based pathogen identification method according to any one of claims 1-2 for non-diagnostic purposes in pathogen identification.

Citation Information

Patent Citations

  • Genotype correction device and method

    CN109785899A

  • Method and device for improving metagenome species identification accuracy, storage medium and application thereof

    CN117831614A

  • Direct identification and measurement of relative populations of microorganisms with direct DNA sequencing and probabilistic methods

    US20120004111A1