SNP (Single Nucleotide Polymorphism) site kit for identifying mycobacterium tuberculosis lineage branches and application of SNP site kit
By providing a test kit for 37 SNP sites, the problem of identifying 6 new subbranches of Mycobacterium tuberculosis lineage 3.1.1 was solved, and high-accuracy and high-specificity identification was achieved, supporting the precise prevention and control of tuberculosis.
Patent Information
- Application Number
- CN202510872246.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-06-26
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies make it difficult to efficiently identify the six new subclades of Mycobacterium tuberculosis lineage 3.1.1, resulting in insufficient tuberculosis prevention and control capabilities.
A kit including 37 SNP sites is provided for identifying 6 new subclades of Mycobacterium tuberculosis lineage 3.1.1, which is verified by whole genome sequence alignment and genetic evolutionary tree to ensure high accuracy and specificity.
It has achieved precise prevention and control of Mycobacterium tuberculosis lineage 3.1.1, improved the accuracy of identification of 6 new branches and the accuracy of evolutionary analysis, and supported the precise prevention and control of tuberculosis.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] This patent relates to the field of SNP molecular markers, specifically, to a 37 SNP site kit and its application for identifying six genetic evolutionary branches of Mycobacterium tuberculosis lineage 3.1.1 except the two subbranches L3.1.1.i1 and L3.1.1.i2. Background Art
[0002] Tuberculosis (TB) is one of the diseases with the longest history of mankind. It is also the disease with the highest number of deaths caused by a single pathogen in the world. It has become a major global public health problem. Mycobacterium tuberculosis (M. tuberculosis), commonly known as tuberculosis bacillus, is the pathogen that causes tuberculosis. As MTBC accompanied human migration out of Africa, 9 human adaptive lineages (L1 to L9) have been differentiated. Different lineages show extremely strong biogeographical correlations, and there are large differences in the proportion of transmission and prevalence. Compared with lineage 2 and lineage 4, which are widely prevalent in the world, lineage 3 is less in number. [2] , limited to globally endemic areas [3] Shuaib et al. [4] Based on a dataset of 2,682 lineage 3 (L3) strains from 38 countries across five continents, MIRU-VNTR genotyping analysis revealed that L3 originated in South Asia and further spread to Northeast Africa and East Africa. L3-endemic areas have the highest tuberculosis burden in the world. Therefore, studying the epidemic patterns of L3 is of great significance to the global end of tuberculosis. Our team previously completed the invention patent (201910805565.X) for "Application of 1,073 SNP sites in Mycobacterium tuberculosis lineage 3" based on the whole genome data of 30 strains of Mycobacterium tuberculosis lineage 3, which facilitated the epidemiological investigation and tracing of L3-infected patients.
[0003] SNP typing technology is based on the polymorphism of genomic DNA sequences caused by single nucleotide variations at the genomic level. It is obtained from whole genome sequence analysis. Its advantage is that it has higher resolution and overcomes the disadvantage of homology and heterogeneity of VNTR typing methods. At present, foreign researchers have screened SNP sites as typing targets. In the early stage of our team [5], we achieved the identification of its six different branches based on the specific SNP combination of M.TBlineage 2.3 whole genome data, and were approved for "A method for identifying Mycobacterium tuberculosis lineage 2.3 subtype strains and its application" national invention patent (2023103402392).
[0004] Some genetic experts hope to define M.TB into infinitely small branches based on genomic data. However, taking the Beijing lineage as an example, the high virulence / transmission capacity shown is not universally possessed by all members of the Beijing family, but is limited to a few branches [5]. Based on the whole genome sequencing data of Mycobacterium tuberculosis, Luca et al. [1] found that L3 can be divided into L3.1.1, L3.1.2, L3.2.1 and L3.2.2. At the same time, they further divided L3.1.1 into two sub-branches, L3.1.1.i1 and L3.1.1.i2. Given that L3.1.1 accounts for a very high proportion and that the strains of L3.1.1 that have not been typed in the evolutionary tree show significant heterogeneity in evolutionary structure, it is of great significance to expand the classification research of L3.1.1 on the basis of the original two sub-branches of L3.1.1.i1 and L3.1.1.i2, and to study its transmissibility / virulence and other aspects.
[0005] It is generally believed that strains from the origin of the disease have the highest genetic diversity. To more comprehensively genotype L3 strains, our team initially used a series of genetic evolutionary methods (see the "Specific Implementation Methods" section) based on genomic data from 1754 L3 strains from India, the origin of L3, to define six new L3.1.1 clades, in addition to L3.1.1.i1 and L3.1.1.i2: L3.1.1.1, L3.1.1.2, L3.1.1.3, L3.1.1.4, L3.1.1.5, and L3.1.1.6. They also integrated SNP combinations to identify these six new clades, totaling 37 SNPs. Using genomic data from 5126 L3.1.1 strains from various countries and regions around the world, these 37 SNPs were validated for typing M.TB L3.1.1, achieving high specificity.
[0006] In view of this, a patent application is filed for a set of 37 SNP sites that define 3.1.16 new branches of the Mycobacterium tuberculosis lineage.
[0007] References
[0008] [1]Luca Freschi, Roger Vargas Jr., Ashaque Hussain, SM Mostofa Kamal, Alena Skrahina, Sabira Tahseen, Nazir Ismail, Anna Barbova, Stefan Niemann, Daniel Maria Cirillo, Anna S Dean, Matteo Zignol, Maha Reda Farhat tuberculosis.NatCommun.2021 Oct 2012(1):6099.doi:10.1038 / s41467-021-26248-1.
[0009] [2]Abigail L.Manson,Keira A.Cohen,Thomas Abeel,ChristopherA.Desjardins,Derek T.Armstrong,Clifton E.Barry,III,Jeannette Brand,Sinéad B.Chapman,Sang-Nae Cho,Andrei Gabrielian,James Gomez,Andrea M.Jodals,MosesJoloba,Pontus Jureen,Jong Kathryn Winglee, Aksana Zalutskaya, Laura E. Via, Gail H. Cassell, Susan E. Dorman, Jerrold Ellner, Parisa Farnia, James E. Galagan, Alex Rosenthal, Valeriu Crudu, Daniel Homorodean, Po-Ren Hsueh, Sujatha Narayanan, Alexander S. Pym, Alena Skrahina, Soumya Swaminathan, Martie Van der Walt, David Alland, William R. Bishai, Ted Cohen, Sven Hoffner, Bruce W. Birren, and AshleeM.Earl.Genomic analysis of globally diverse Mycobacterium tuberculosisstrains provides insights into the emergence and spread of multidrug resistance[J].Nature Genetics,2017,49(3):395-402.DOI:10.1038 / ng.3767.
[0010] [3]Abigail L.Manson,Thomas Abel,James E.Galagan,Jagadish ChandraboseSundaramurthi,Alex Salazar,Thies Gehrmann,Siva Kumar Shanmugam,KannanPalaniyandi,Sujatha Narayanan,Soumya Swaminathan,and AshleeM.Earl.Mycobacterium tuberculosis Whole Genome Sequences from Southern IndiaSuggest Novel Resistance Mechanisms and the Need for Region-SpecificDiagnostics.Clinical Infectious Diseases.2017:64(11):1494–501.
[0011] [4]Shuaib YA,Khalil EAG,Wieler LH,Schaible UE,Bakheit MA,Mohamed-NoorSE,Abdalla MA,Kerubo G,Andrew S,Hillemann D,Richter E,Kranzer K,Niemann S,Merker M.Mycobacterium tuberculosis Complex Lineage 3 as a Causative Agent ofPulmonary Tuberculosis,Eastern Sudan.Emerg Infect Dec.2020 Mar26(3):427-4
[0012] [5]Chendi Zhu,Tingting Yang,Jinfeng Yin,Hui Jiang,Howard E.Takiff,Qian Gao,Qingyun Liu,Weimin Li.The Global Success of MycobacteriumTuberculosis Modern Beijing Family Is Driven by a Few Recently EmergedStrains.Microbiology Spectrum.2023 Aug 17;11(4):e0333922. Summary of the Invention
[0013] A first objective of the present invention is to provide a kit for identifying 37 SNPs from six new subclades of Mycobacterium tuberculosis lineage 3.1.1, in addition to the two subclades L3.1.1.i1 and L3.1.1.i2. The kit includes reagents or components for identifying the 37 SNPs from the six new subclades of Mycobacterium tuberculosis lineage 3.1.1. The kit is capable of identifying the 37 SNPs from the six new subclades of Mycobacterium tuberculosis lineage 3.1.1, thereby discovering the epidemiological characteristics of different subclades of Mycobacterium tuberculosis lineage 3.1.1 and improving the ability to accurately prevent and control Mycobacterium tuberculosis lineage 3.1.1.
[0014] A second object of the present invention is to provide a kit for use in identifying six new clades within the Mycobacterium tuberculosis lineage 3.1.1, excluding subclades L3.1.1.i1 and L3.1.1.i2, or in phylogenetic analysis. This kit can improve the accuracy of identification or phylogenetic analysis of the six new clades within the Mycobacterium tuberculosis lineage 3.1.1.
[0015] The third object of the present invention is to provide an application of the above-mentioned kit in the preparation of diagnostic products for judging tuberculosis infectious pathogen sublineages, L3.1.1.1, L3.1.1.2, L3.1.1.3, L3.1.1.4, L3.1.1.5 and L3.1.1.6.
[0016] The fourth object of the present invention is to provide a method for identifying or evolutionarily analyzing six new subbranches of Mycobacterium tuberculosis lineage 3.1.1 in addition to the two subbranches L3.1.1.i1 and L3.1.1.i2. The method uses the above-mentioned kit to detect 37 SNP sets of the six new subbranches of Mycobacterium tuberculosis lineage 3.1.1.
[0017] In order to achieve the above-mentioned purpose of the present invention, the following technical solutions are adopted:
[0018] The present invention provides a kit for identifying 37 SNP sites in 6 new branches of Mycobacterium tuberculosis lineage 3.1.1, excluding the two subbranches L3.1.1.i1 and L3.1.1.i2. The kit includes reagents and / or components for identifying the 37 SNP sets in 6 new branches of Mycobacterium tuberculosis lineage 3.1.1. The SNP set includes the following 37 SNP sites, the physical locations of which are determined based on the whole genome sequence alignment of Mycobacterium tuberculosis H37Rv, the whole genome sequence of Mycobacterium tuberculosis H37Rv having the Accession Number NC_000962.3 (Table 1):
[0019] Table 1. 37 SNP screening sites
[0020]
[0021] The above-mentioned kit of the present invention can detect a SNP set selected from 37 SNP sites.
[0022] The 37 SNP sites described in the present invention were obtained by comparison based on the whole genome data of 1,754 Mycobacterium tuberculosis lineage 3 from India, and verified based on the genetic evolutionary tree constructed based on the genome data of 5,126 Mycobacterium tuberculosis 3.1.1 strains from various countries and regions around the world, thereby representing the SNP sites of each of the 6 new evolutionary branches on the Mycobacterium tuberculosis lineage 3.1.1 evolutionary tree except the two sub-branches L3.1.1.i1 and L3.1.1.i2. The above-mentioned SNP sites excellently characterize the genetic evolutionary information of the 3.1.16 new branches of the Mycobacterium tuberculosis lineage and the mutation information of key genes. The kit described in the present invention can identify the above-mentioned SNP sites of the 3.1.16 new genetic evolutionary branches of the Mycobacterium tuberculosis lineage and provide the above-mentioned information, which is of great significance for the genetic classification of the 3.1.16 new branches of the Mycobacterium tuberculosis lineage and the prevention and control of tuberculosis.
[0023] In some embodiments, the above-mentioned SNP set includes: the 1 SNP site defining branch 1, or the 2 SNP sites defining branch 2, or the 4 SNP sites defining branch 3, or the 9 SNP sites defining branch 4, or the 15 SNP sites defining branch 5, or the 6 SNP sites defining branch 6.
[0024] The above embodiment of the present invention optimizes the number of SNP sites included in the SNP set. The optimized SNP site set can better identify and analyze the evolution of the six new branches of Mycobacterium tuberculosis lineage 3.1.1.
[0025] In some embodiments, the above-mentioned SNP set includes 37 of the above-mentioned SNP sites.
[0026] The data sample test set used by the present invention to screen SNP sites comes from the whole genome sequences of 1754 Mycobacterium tuberculosis lineage 3.1.1 in India, and the training set covers 5126 Mycobacterium tuberculosis lineage 3.1.1 strains that have been discovered and sequenced in the world at this stage. The SNP set including 37 SNP sites described in the present invention can more comprehensively reflect the overall situation of the identification and evolution of 6 new branches of Mycobacterium tuberculosis lineage 3.1.1 in addition to the two sub-branches L3.1.1.i1 and L3.1.1.i2. Secondly, the above-mentioned SNP set of the present invention not only ensures high resolution, but also reduces the number of SNP sites that do not need to be detected, achieving a balance between the two. Moreover, the above-mentioned 37 SNP sites of the present invention are used to identify the global Mycobacterium tuberculosis lineage 3.1.1 in addition to the two sub-branches L3.1.1.i1 and L3.1.1.i2, with high accuracy and strong specificity. This will provide support for the identification and precise prevention and control of sub-branches of the East African-Indian lineage Mycobacterium tuberculosis with strong pathogenicity and transmissibility.
[0027] In some embodiments, the reagents and / or components for detecting a SNP set include one or more of PCR primers, molecular probes, biosensors, and chips.
[0028] The present invention also relates to an application of the above kit in the identification and evolutionary analysis of 6 new branches of Mycobacterium tuberculosis lineage 3.1.1.
[0029] The present invention also relates to a use of the above-mentioned kit in preparing a diagnostic product for determining the 2.3 sublineage of Mycobacterium tuberculosis lineage, a pathogen of tuberculosis infection.
[0030] The present invention also relates to a method for identifying or evolutionarily analyzing six new branches of Mycobacterium tuberculosis lineage 3.1.1 except the two subbranches L3.1.1.i1 and L3.1.1.i2. The method uses the above-mentioned kit to detect 37 SNP sites of the six new branches of Mycobacterium tuberculosis lineage 3.1.1.
[0031] In some embodiments, the method for detecting the 37 SNP sites includes one or more of the following: SNP detection method based on gel electrophoresis, DNA sequencing method, DNA chip method, denaturing high performance liquid chromatography method or mass spectrometry detection method.
[0032] In some embodiments, the gel electrophoresis-based SNP detection method includes one or more of single-strand conformational polymorphism detection method, denaturing gradient gel electrophoresis detection method, enzyme-amplified polymorphic sequence detection method, and allele-specific PCR detection method.
[0033] In some embodiments, the mass spectrometry detection method comprises matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF) detection method.
[0034] Beneficial effects
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] 1). The 37 SNP sites involved in the kit, application and method of the present invention are derived from the whole genome sequence of 1754 Mycobacterium tuberculosis lineage 3 from India, the country of origin, and are subjected to genetic differentiation analysis. At the same time, they are SNP sites obtained after verification of 5126 L3.1.1 Mycobacterium tuberculosis genome sequences from various countries and regions around the world. The above 37 specific SNP sites were functionally enriched at the same time, indicating that they are all key SNP sites that characterize the 6 new genetic evolutionary branches of the Mycobacterium tuberculosis lineage 3.1.1 and are closely related to the gene function of Mycobacterium tuberculosis. Therefore, the detection of the above SNP sites is of great significance for the identification and evolutionary analysis of the 6 new branches of the Mycobacterium tuberculosis lineage 3.1.1 in addition to the two sub-branches L3.1.1.i1 and L3.1.1.i2.
[0037] 2) The 37 SNP sites involved in the kit, application and method of the present invention have a high accuracy rate when identifying a single Mycobacterium tuberculosis, and the specificity of some branches reaches 100%. It can accurately identify 6 new genetic strains of Mycobacterium tuberculosis lineage 3.1.1 except the two sub-branches L3.1.1.i1 and L3.1.1.i2 at the strain level. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0039] Figure 1 Phylogenetic tree constructed based on the SNP set of L3 genomes of 1754 Mycobacterium tuberculosis strains from India.
[0040] Figure 2 A verified evolutionary tree constructed based on 5126 L3.1.1 worldwide combined with the above-mentioned specific sites.
[0041] Figure 3 Bayesian skylines of each Clade.
[0042] Figure 4 Distribution of terminal branch lengths of each clade.
[0043] Figure 5 Clustering rate of each Clade strain. DETAILED DESCRIPTION
[0044] The following is a further description of specific embodiments of the present invention. It should be noted that the description of these embodiments is intended to facilitate understanding of the present invention and does not constitute a limitation of the present invention. In addition, the technical features involved in the embodiments described below may be combined with each other as long as they do not conflict with each other.
[0045] The experimental methods in the following examples are conventional methods unless otherwise specified, and the experimental materials used in the following examples are commercially available unless otherwise specified.
[0046] Example 1 Screening of SNP sites across the entire genome of Mycobacterium tuberculosis
[0047] 1. Download strain data
[0048] The whole genome data of 1754 Mycobacterium tuberculosis strains from India were downloaded from the NCBI database.
[0049] 2. Quality assessment of strain sequencing data
[0050] FastQC software was used to evaluate the GC content of the strain sequencing fastq files: (1) sequencing data with a GC content of less than 60% or greater than 70% were removed; (2) sequencing data with a SNP miss call rate greater than 15% were removed. Finally, all 1754 Mycobacterium tuberculosis lineage 3.1.1 sequencing data were included in this study.
[0051] 3. Statistical comparison rate
[0052] Bowtie 2 software was used to align the sequencing files of the strains to the reference genome of Mycobacterium tuberculosis H37Rv (NC_000962.3), and the strain alignment rate was counted.
[0053] 4. SNP Calling
[0054] A previously validated analysis pipeline was used to map short sequencing reads to the reference genome. Briefly, the Sickle tool was used to trim the WGS data. Sequencing reads with a Phred base quality higher than 20 and a length greater than 30 were retained for analysis. The whole genome sequence of Mycobacterium tuberculosis H37Rv strain (NC_000962.2) was used as a reference template for read mapping. Sequencing reads were mapped to the reference genome analysis using Bowtie 2. SAMtools was used for SNP calling with a map quality greater than 30. Fixed mutations (frequency ≥ 75%) were included in results supported by at least 10 reads using VarScan, and the chain skew filter option was enabled. All SNPs located in genomic repetitive regions that are difficult to characterize with short gene sequencing technology were excluded (e.g., PPE / PE-PGRS family genes, phage sequences, insertions or mobile genetic elements). Small insertions or deletions identified by VarScan were also excluded.
[0055] Finally, 106,438 SNP sites were obtained from 1,754 sequenced strains.
[0056] Example 2 Construction of a phylogenetic tree of Mycobacterium tuberculosis
[0057] 1754 L3 strains were used to calculate SNPs specific to the six subclades. Briefly, ancestral sequences for the different L3 subclades were reconstructed using Baseml software to generate the most recent common ancestor (MRCA) sequence, which was then compared with the most recent L3 MRCA sequence. This yielded a minimal set of identifiable L3.1.1 SNPs. As described in Table 1, the following steps were then followed to construct a phylogenetic tree of Mycobacterium tuberculosis:
[0058] Mixed infection isolates were excluded from phylogenetic reconstruction by investigating SNP genotypic heterozygosity. For all phylogenetic reconstructions, SNPs from MTBC isolates were merged into a consensus and non-redundant list, and nucleotide positions with gaps exceeding 5% of the taxa were excluded (possibly due to insertions or deletions, low coverage, etc.). Alignments of polymorphic positions from all strains were used for phylogenetic reconstruction using MEGAX. When the number of taxa was large, the neighbor-joining method was used for initial inference of phylogenetic structure. However, for the final estimation of the phylogeny, maximum likelihood bootstrapping with at least 100 replicates was applied under a general time-reverse model with confidence levels.
[0059] The phylogenetic tree was visualized in FigTree (http: / / tree.bio.ed.ac.uk / software / figtree / ) or iTOL. The final phylogenetic tree differentiation nodes of phylogenetic 3.1.1 were obtained (see Figure 1Specifically, the above SNP sites were used to subdivide the Mycobacterium tuberculosis lineage 3.1.1 into the following 6 new branches: L3.1.1.1, L3.1.1.2, L3.1.1.3, L3.1.1.4, L3.1.1.5 and L3.1.1.6.
[0060] Example 3 Evaluation of SNP panel strain-level identification results
[0061] For specific test methods, refer to the contents in Examples 1 and 2.
[0062] 5127 L3.1.1 strains of Mycobacterium tuberculosis with known genome sequences from 51 countries were downloaded from the NCBI database ( Figure 2 ), these strains were typed according to the final 37 SNP sets (see Table 1), and the typing results of the 37 SNP sets were completely consistent with the typing results of the whole genome sequence data.
[0063] Example 4 L3.1.1 Analysis of Epidemic Characteristics of Each Sub-branch
[0064] Through multi-method joint analysis, the newly defined six sub-clades (L3.1.1.1-L3.1.1.6, Clade1-6) showed significant heterogeneity at the level of transmission dynamics (Table 2).
[0065] Table 2. The six newly defined clades (Clades 1-6) show significant heterogeneity in terms of transmission dynamics.
[0066]
[0067] Specifically, the effective population growth curve generated by Bayesian skyline analysis shows that the population sizes of Clade 1, 3, 5, and 6 have shown a clear expansion trend in history ( Figure 3 ); The length of the terminal branch is used as an indicator of the bacterial population's expansion capacity. The smaller the terminal branch length, the faster the expansion rate. The distribution of terminal branch lengths shows that the recent spread of Clade 1 and 2 is stronger than that of other sublineages (see Appendix). Figure 4 , Table 3; Clustering rate mainly reflects the local transmission activity of pathogens. Analysis shows that Clade 1, 3, and 4 have more serious local transmission ( Figure 5 ).
[0068] Table 3. Statistical differences in terminal branch length distribution among groups
[0069]
[0070] The analysis method proposed in this application divides the six evolutionary branches (Clade1-6) according to the expansion trend, recent transmission intensity and the severity of local transmission, further clarifies the characteristics between different subgroups within the L3.1.1 lineage, and can accurately identify the six genetic subbranches of the Mycobacterium tuberculosis lineage L3.1.1.1 at the strain level.
[0071] The new typing marker system proposed in this application can effectively distinguish sublineages of Mycobacterium tuberculosis with different transmission characteristics, and provide a reliable technical solution for accurate traceability and molecular epidemiological research. In the actual prevention and control of tuberculosis, if a group of patients are found to be infected with Mycobacterium tuberculosis lineage 3 strains, the new typing scheme proposed in this patent can be used to efficiently identify the specific subgroup to which the strain infected by the patient belongs. Accordingly, more prevention and control resources can be preferentially allocated to patient groups infected with high-transmission risk strains (such as L3.1.1.1-L3.1.1.4), thereby providing a scientific basis for the precise prevention and control of tuberculosis.
[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention is described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for identifying six new subclade SNPs (single nucleotide polymorphisms, single nucleotide polymorphisms) of Mycobacterium tuberculosis lineage 3.1.1 A kit for detecting nucleotide polymorphism) sites, characterized in that The kit includes reagents and / or components for detecting a SNP set, wherein the SNP set includes the following SNP sites; the physical locations of the SNP sites are determined based on a full genome sequence alignment of Mycobacterium tuberculosis H37Rv, wherein the Accession Number of the full genome sequence of Mycobacterium tuberculosis H37Rv is NC_000962.3: Table 1: 37 screening sites SNP site combination 1: located at position 3571834 on the chromosome of Mycobacterium tuberculosis, with a nucleotide sequence of C / A; SNP site combination 2: located at position 911261 on the chromosome of Mycobacterium tuberculosis, with a nucleotide sequence of C / T; Located at position 3874726 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is C / A; SNP site combination 3: located at position 633562 on the chromosome of Mycobacterium tuberculosis, with nucleotides G / A; Located at position 2221584 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is G / C; Located at position 2941179 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is C / T; Located at position 4059186 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is G / T; SNP site combination 4: located at position 358473 on the chromosome of Mycobacterium tuberculosis, with nucleotide sequence G / C; Located at position 635139 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is C / G; Located at position 1111852 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is G / T; Located at position 2186236 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is G / A; Located at position 3247874 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is G / A; Located at position 3385218 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is G / A; Located at position 3678094 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is C / T; Located at position 3964930 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is G / A; Located at position 4060100 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is C / T; SNP site combination 5: located at position 208320 on the Mycobacterium tuberculosis chromosome, with nucleotides G / A; Located at position 301341 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is G / T; Located at position 686123 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is C / A; Located at position 957306 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is C / T; Located at position 996263 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is C / T; Located at position 1728161 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is C / A; Located at position 2128372 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is G / A; Located at position 2289047 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is C / T; Located at position 2504177 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is G / A; Located at position 3112700 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is G / A; Located at position 3459081 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is G / C; Located at position 3714639 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is G / A; Located at position 4007272 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is G / A; Located at position 4097569 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is C / A; Located at position 4313156 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is G / A; SNP site combination 6: located at position 224338 on the Mycobacterium tuberculosis chromosome, with nucleotide sequence C / T; Located at position 1836417 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is G / T; Located at position 1924959 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is C / T; Located at position 2631641 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is C / A; Located at position 3219500 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is A / T; Located at position 4289953 on the chromosome of Mycobacterium tuberculosis, its nucleotide sequence is C / T.
2. The kit according to claim 1, wherein: The SNP set includes the 37 SNP sites that define branches 1-6; or the 1 SNP site that defines branch 1, or the 2 SNP sites that define branch 2, or the 4 SNP sites that define branch 3, or the 9 SNP sites that define branch 4, or the 15 SNP sites that define branch 5, or the 6 SNP sites that define branch 6.
3. The kit according to claim 2, wherein The SNP set includes the 37 SNP sites mentioned above.
4. The kit according to claim 1, wherein The reagents and / or components for detecting the 37 SNP set include one or more of PCR primers, molecular probes, biosensors and chips.
5. The kit according to any one of claims 1 to 4, used in the identification or evolutionary analysis of six branches of the Mycobacterium tuberculosis lineage 3.1.1 except the two subbranches L3.1.1.i1 and L3.1.1.i2.
6. The kit according to any one of claims 1 to 4, used in the preparation of a diagnostic product for determining the six branches of the tuberculosis pathogen, Mycobacterium tuberculosis lineage 3.1.1, excluding the two subbranches L3.1.1.i1 and L3.1.1.i2.
7. A method for identifying or analyzing the evolution of Mycobacterium tuberculosis lineage 3, characterized in that: The method uses the kit according to any one of claims 1 to 4 to detect 37 SNP sites defined by 6 branches of the Mycobacterium tuberculosis lineage 3.1.
1.
8. The method according to claim 7, wherein Methods for detecting the 37 SNP sites include one or more of the following: SNP detection methods based on gel electrophoresis, DNA sequencing, DNA chip, denaturing high performance liquid chromatography and / or mass spectrometry.
9. The method according to claim 8, wherein The gel electrophoresis-based SNP detection method includes one or more of a single-strand conformational polymorphism detection method, a denaturing gradient gel electrophoresis detection method, an enzyme-amplified polymorphic sequence detection method, and / or an allele-specific PCR detection method.
10. The method according to claim 8, wherein The mass spectrometry detection method includes matrix-assisted laser desorption ionization time-of-flight mass spectrometry detection method.
Citation Information
Patent Citations
Application of 1073 SNP loci in Mycobacterium tuberculosis lineage 3
CN110453001B