Characteristic target and primer pair for Klebsiella pneumoniae level rapid identification and serotype typing and application of characteristic target and primer pair

By constructing a phylogenetic tree using the characteristic target yhaJ_3 and primer pairs AACTGGAGGAGGAACTGGAC, TCGGCAATATCCTTCTCCAC, as well as the wzc and cps genes, the problems of accuracy and efficiency in Klebsiella pneumoniae species identification and serotyping were solved, and a rapid and highly specific detection method was achieved.

CN120989273APending Publication Date: 2025-11-21JINAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511401305.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing methods for identifying Klebsiella pneumoniae species have insufficient resolution and are easily confused with closely related species. Traditional serotyping methods are cumbersome to operate and have poor specificity, making it difficult to meet the needs of rapid detection.

Method used

Species identification was performed using the characteristic target yhaJ_3. PCR amplification was performed using primer pairs AACTGGAGGAGGAACTGGAC and TCGGCAATATCCTTCTCCAC. Serotyping was performed using wzc and cps genes. A phylogenetic tree was constructed through pan-genome analysis and multi-site sequence alignment.

Benefits of technology

It enables rapid, accurate, and specific identification of Klebsiella pneumoniae species and serotypes, and is suitable for high-throughput screening and POCT platforms, reducing testing costs and improving result stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120989273A_ABST
    Figure CN120989273A_ABST
Patent Text Reader

Abstract

The invention discloses a characteristic target and a primer pair for Klebsiella pneumoniae level rapid identification and serotype typing and application of the characteristic target and the primer pair. The nucleotide sequence of a characteristic target yhaJ3 for rapid identification of the species level is shown as SEQ ID NO.1, the characteristic targets for serotype typing are wzc and cps, the nucleotide sequence of wzc is shown as SEQ ID NO.2, and the nucleotide sequence of cps is shown as SEQ ID NO.3. The invention further discloses a kit for identifying the species level of the serotype. According to the invention, the specific gene yhaJ3 is screened for species level recognition, wzc and cps genes are screened for serotype typing, and the specific primer is designed for PCR amplification, so that the problems of low recognition accuracy and poor detection efficiency in the prior art are solved, and rapid, accurate and specific recognition of klebsiella pneumoniae and main serotypes thereof is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of bioinformatics and microbial gene detection services, in particular to a characteristic target for rapid identification and serotyping of Klebsiella pneumoniae species, a primer pair and application thereof BACKGROUND

[0002] Klebsiella pneumoniae is an important opportunistic pathogen, widely exists in natural environment, human intestine, hospital and community environment. The bacteria can cause various clinical infections, including pneumonia, urinary tract infection, bacteremia, liver abscess, wound infection and meningitis, especially in the elderly, infants and patients with low immune function. Due to its diverse transmission routes and strong environmental adaptability, it has become one of the main pathogens of hospital-acquired infections, which has caused a continuous challenge to public health.

[0003] Traditional species identification methods (such as 16S rRNA sequence alignment) have limited resolution in distinguishing K. pneumoniae from its closely related species members (such as K. variicola and K. quasipneumoniae), which can easily lead to misjudgment or confusion, thereby affecting the accuracy of epidemiological investigation and the development of clinical treatment plan.

[0004] In addition, the capsule of K. pneumoniae is a major virulence factor, and the K antigen in the capsular polysaccharide determines the serotype of the strain. Different serotypes are closely related to their virulence, pathogenicity and prevalence. At present, common high virulence serotypes include K1, K2, K5, K20, K54 and K57. Traditional slide agglutination serotyping method has problems such as complicated operation, long time-consuming, poor specificity, etc., which is difficult to meet the needs of high-throughput detection and rapid clinical diagnosis.

[0005] Therefore, it is urgent to develop an efficient, accurate and high-throughput K. pneumoniae identification and typing method to support the development and application of related kits and POCT systems. SUMMARY

[0006] The existing species identification method for K. pneumoniae has insufficient resolution, which is easy to be confused with closely related species, affecting the accuracy of identification. Traditional serotyping method has problems such as complicated operation, poor specificity and low efficiency, which is difficult to meet the rapid detection requirement. The present application aims to provide a K. pneumoniae species level rapid identification and serotyping method based on characteristic target gene.

[0007] The first object of the present application is to provide the application of characteristic target yhaJ_3 in identifying K. pneumoniae species level for non-disease diagnosis and / or treatment purposes, and the nucleotide sequence of yhaJ_3 is shown in SEQ ID NO. 1.

[0008] A second object of the present application provides application of the primer pair comprising AACTGGAGGAGGAACTGGAC and TCGGCAATATCCTTCTCCAC in identification of the Klebsiella pneumoniae species level for non-disease diagnosis and / or treatment purposes.

[0009] A third object of the present application provides application of the kit comprising the primers AACTGGAGGAGGAACTGGAC and TCGGCAATATCCTTCTCCAC in identification of the Klebsiella pneumoniae species level for non-disease diagnosis and / or treatment purposes.

[0010] A fourth object of the present application provides a method for identification of the Klebsiella pneumoniae species level for non-disease diagnosis and / or treatment purposes, which comprises the following steps: extracting genomic DNA of a sample to be tested, performing PCR amplification with the genomic DNA as a template and the primers AACTGGAGGAGGAACTGGAC and TCGGCAATATCCTTCTCCAC, and if a specific 634bp amplification band appears, the sample to be tested is Klebsiella pneumoniae, otherwise, it indicates that the sample to be tested is not Klebsiella pneumoniae.

[0011] Preferably, the annealing temperature of the PCR amplification is 60-65°C.

[0012] Preferably, the annealing temperature of the PCR amplification is 63°C.

[0013] Preferably, the PCR amplification system is 2xHieffUltra-Rapid II HotStart PCR Master Mix 10μL, template DNA 1μL, 10μM upper and lower primers 1.2μL each, and sterile water to make up the volume to 20μL; the PCR amplification procedure is: 95℃ pre-denaturation for 5min, 1 cycle; 95℃ denaturation for 30s, 63℃ annealing for 30s, 72℃ extension for 10s, a total of 35 cycles; 72℃ extension for 5min, 1 cycle.

[0014] A fifth object of the present application provides application of the characteristic targets wzc and cps in Klebsiella pneumoniae serotype typing, wherein the nucleotide sequence of the wzc is shown as SEQ ID NO. 2, and the nucleotide sequence of the cps is shown as SEQ ID NO. 3.

[0015] The sixth object of the present application is to provide a K. pneumoniae serotyping method based on characteristic target wzc and cps, which comprises the following steps: extracting DNA of a sample of a strain to be tested, obtaining whole genome and whole genome information of all strain samples; extracting serotype characteristic gene sequences wzc and cps of each serotype K. pneumoniae of the strain sample and identity confirmation, and performing sequence alignment and alignment, then removing redundant nucleotide sequences, and finally constructing a maximum likelihood method phylogenetic tree in series of the aligned sequences to identify the serotype of the strain to be tested.

[0016] Preferably, the maximum likelihood method phylogenetic tree is set to 5000 times Ultrafast.

[0017] The present application has the following outstanding technical advantages:

[0018] (1) Based on bioinformatics platform, K. pneumoniae species characteristic genes are screened through pan-genome analysis, which improves the accuracy and specificity of K. pneumoniae identification.

[0019] (2) Pan-genome analysis and MLST are integrated to establish a rapid, high-throughput and low-cost K. pneumoniae serotyping method.

[0020] (3) K. pneumoniae species-specific genes are systematically screened to improve the accuracy and specificity of species identification.

[0021] (4) The PCR amplification system is optimized to improve the primer recognition ability and reduce false positive interference.

[0022] (5) The established K. pneumoniae species and serotype identification method is suitable for large-scale sample screening and can be expanded to POCT platform and commercial molecular diagnostic kit development.

[0023] The present application solves the problems of low recognition accuracy and poor detection efficiency in the prior art by screening specific gene yhaJ_3 for species level identification, screening wzc and cps genes for serotyping, and designing specific primers for PCR amplification, and realizes rapid, accurate and specific identification of K. pneumoniae and its main serotypes.

[0024] The application establishes a molecular detection method for rapid identification and serotyping of Klebsiella pneumoniae by screening the species level characteristic gene yhaJ_3 and the serotype classification characteristic gene wzc and cps, which has the advantages of high specificity, accurate recognition, simple operation, fast detection speed and wide application range. The experimental results show that yhaJ_3 has good amplification effect in target bacteria and no amplification in non-target bacteria, and the specific recognition accuracy is 100%; the phylogenetic tree constructed by wzc and cps can accurately distinguish six common high virulence serotypes. Compared with the traditional method, the detection cost of the application is lower, the result is more stable, it is suitable for high-throughput screening and POCT platform development, and has significant technical popularization and application value.

[0025] The application selects the highly specific yhaJ_3 gene from 1833 strains of Klebsiella pneumoniae and 25 strains of non-target bacteria through pan-genome analysis, as a characteristic target for species level recognition, effectively overcoming the problem of insufficient recognition accuracy of 16S rRNA and MLST method. At the same time, wzc and cps genes are selected as serotype classification targets because they are involved in the synthesis of capsular polysaccharide and have a clear serotype determining function, which can accurately distinguish common high virulence serotypes and make up for the low specificity of traditional agglutination method. In addition, the annealing temperature of the PCR primer is optimized to ensure stable amplification in target bacteria and no amplification in non-target bacteria, and a simple and efficient detection system is constructed.

[0026] The application has a sound theoretical basis, and through pan-genome analysis, the whole genome is systematically compared and characteristic extraction is performed, and the core gene stably existing in the target bacteria or specific serotype is selected, which significantly improves the specificity and consistency of recognition. At the same time, combined with multi-locus sequence alignment (MLST) and phylogenetic tree construction, the genetic differences between different serotypes are accurately reflected, so that the serotype classification has a clear system evolution basis, further enhancing the scientificity and accuracy of the classification.

[0027] This technology is fully verified by experimental data: first, PCR verification shows that the yhaJ_3 gene has good amplification effect and high specificity in many strains of Klebsiella pneumoniae, the target bacteria have high amplification rate and the non-target bacteria have no amplification, the accuracy is 100%, and the false positive rate is 0%, which is better than other candidate genes; secondly, through annealing temperature gradient optimization, 63℃ is determined as the optimal temperature, realizing single clear amplification band, improving the stability and applicability of the reaction system; finally, the serotype phylogenetic tree constructed based on wzc and cps genes is highly consistent with the traditional identification results, further proving the accuracy and systematicness of the method in serotype classification. BRIEF DESCRIPTION OF DRAWINGS

[0028] Figure 1Figure 1 is a K. pneumoniae analysis strain clustering result map based on ANI analysis.

[0029] Figure 2 Figure 3 is a K. pneumoniae species characteristic gene PCR verification result map in target bacteria. Wherein A is the preliminary verification of the candidate gntR_1, yhaJ_3, rapz, pgK, nuoH, rpoH, glyA, ducB, rne, cra_2 target in 3 K. pneumoniae strains; B is the verification of the gntR_1, yhaJ_3, rapz, pgK, nuoH, rpoH, rne, cra_2 target in 22 K. pneumoniae strains after preliminary screening.

[0030] Figure 3 Figure 4 is a K. pneumoniae species characteristic gene PCR verification result map in non-target bacteria. Wherein A is the yhaJ_3 amplification result; B is the nuoH amplification result; C is the rapz amplification result; D is the rne amplification result; E is the gntR_1 amplification result.

[0031] Figure 4 Figure 5 is a K. pneumoniae species characteristic gene PCR annealing temperature optimization map. Wherein A is gntR_1 annealing temperature optimization; B is rapz annealing condition optimization; C is yhaJ_3 annealing condition optimization.

[0032] Figure 5 Figure 6 is a K. pneumoniae six serotype phylogenetic tree constructed based on new characteristic genes wzc and cps.

[0033] Figure 6 Figure 7 is a K. pneumoniae phylogenetic tree analysis map constructed based on traditional six serotype characteristic genes. DETAILED DESCRIPTION

[0034] The present application provides characteristic targets, primer pairs and their applications for rapid identification and serotype typing of K. pneumoniae species level. Through the integration of pan-genome analysis, core gene screening, multi-site sequence alignment, phylogenetic construction and molecular experiment verification, accurate identification of K. pneumoniae species and serotypes is realized, which provides theoretical and technical support for infection monitoring and molecular diagnosis.

[0035] The technical scheme of the present application specifically includes the following steps:

[0036] (1) Genomic data acquisition and screening

[0037] Representative K. pneumoniae complete genome data with Assembly Level of "Complete" were downloaded from NCBI database, and species confirmation was performed using fastANI tool.

[0038] (2) Gene function annotation

[0039] Genomes were annotated using Prokka, and GFF, FFN, FAA, etc. format files were generated for subsequent feature gene mining.

[0040] (3) Pan-genome analysis to mine K. pneumoniae species characteristic genes

[0041] By comparing the pan-genome data of K. pneumoniae and non-K. pneumoniae strains, the characteristic genes only present in K. pneumoniae were screened from the presence / absence matrix.

[0042] (4) K. pneumoniae species characteristic gene verification

[0043] For the screened species characteristic genes, specific primers were designed using Oligo 7.0 software, and PCR verification was performed in actual standard strains.

[0044] (5) K. pneumoniae serotype characteristic gene mining

[0045] First, according to the known related genes (such as wzc, cps, etc.) of K1, K2, K5, K20, K54, K57 six serotypes, the serotype of K. pneumoniae strains was confirmed. Subsequently, pan-genome analysis was performed on all K. pneumoniae genomes using Roary, with a BLASTP homology threshold of 95%, and a presence / absence matrix was generated. On this basis, genes only present in a certain serotype and absent in other serotypes were screened as the characteristic genes of the serotype.

[0046] (6) Construction of phylogenetic tree based on K. pneumoniae serotype characteristic genes

[0047] First, MAFFT was used for multiple sequence alignment of serotype characteristic genes, and trimAI was used to remove redundant regions in the alignment. Subsequently, the Concatenate function was used to splice multiple gene sequences, and a phylogenetic tree was constructed on the PhyloSuite software.

[0048] For those skilled in the art to better understand the present application, the present application is further described below in conjunction with specific examples, but the protection scope of the present application is not limited thereto.

[0049] Example 1: Whole genome data download and screening

[0050] Enter "Klebsiella pneumoniae" in the NCBI Genome database for retrieval. Select "Annotated genomes" in Filters and preliminarily screen according to "ASSEMBLY LEVEL". Then click "Download-Download Package", select "Genome sequence (FASTA)" and download in bulk through the browser. If the number of genomes of this species is large (e.g., as of March 6, 2025, there are 83,394 K. pneumoniae genome information), only download the representative strain genomes with "ASSEMBLY LEVEL" as "complete".

[0051] Then remove duplicate redundant sequences, and as of September 30, 2024, a total of 1833 K. pneumoniae strain genome data were obtained. The fastANI tool was used for species identification of the downloaded genomes, using an ANI standard threshold of 0.95, and the results showed that the similarity between these genomes was above 97%, confirming that they were K. pneumoniae. Figure 1 In addition, 25 strains of E. coli were downloaded on October 1, 2024 as non-target control strains.

[0052] The Genbank numbers corresponding to the 25 strains of E. coli are: CP056249.1, CP137935.1, CP115333.1, CP157958.1, NC_002695.2, NZ_CP026939.2, NZ_CP068041.1, NZ_CP081192.1, NZ_CP060974.1, NZ_CP095155.1, NZ_CP117049.1, NZ_CP117053.1, NZ_CP074680.1, NZ_CP073962.1, NZ_CP097213.1, NZ_CP103745.1, NZ_CP109601.1, NZ_CP116480.1, NZ_CP126297.1, NZ_AP027978.1, NZ_CP122499.1, NZ_CP115361.1, NZ_CP146654.1, NZ_CP150989.1, NZ_CP158503.1.

[0053] The genome was annotated using Prokka to generate GFF, FFN, FAA, and other format files for subsequent feature gene mining.

[0054] Example 2: K. pneumoniae species characteristic gene primer pair design and verification

[0055] According to the identity-confirmed K. pneumoniae strains (1833 strains) and non-K. pneumoniae (25 strains) strain information, the pan-genome analysis was performed again, with the change of the BLASTP homology threshold to 95%. The gene_presence_absence.csv output by the pan-genome analysis contains all the genes downloaded to the local to generate the pan-genome results. The local matrix conversion script was run in the Windows local CMD command operator to convert the absence and presence of the above genes into a matrix of 0 and 1, i.e. 0 represents the absence of a gene and 1 represents the presence of a gene. Using the matrix, the specific genes carried only by K. pneumoniae and not by non-K. pneumoniae strains were selected by screening the carrying status, thereby completing the mining and screening of K. pneumoniae characteristic gene targets.

[0056] By comparing the pan-genome data of K. pneumoniae and non-K. pneumoniae strains, the characteristic genes present only in K. pneumoniae were screened from the presence / absence matrix.

[0057] For the 10 candidate characteristic genes (gntR_1, yhaJ_3, rapz, pgK, nuoH, rpoH, glyA, ducB, rne, cra_2) obtained by pan-genome analysis, Oligo 7.0 was used to design forward and reverse primers, and the specific sequences are shown in Table 1. First, 3 strains of K. pneumoniae were selected to extract genomic DNA as a template for preliminary PCR amplification verification (the PCR reaction system was 20 μL, mixed in a 100 μL PCR tube, and the system included: 2x HieffUltra-Rapid II HotStart PCR Master Mix 10 μL (Yixing Biotechnology (Shanghai) Co., Ltd.; Catalog No.: 10157ES03), template DNA 1 μL (negative control was 1 μL of sterile water), 10 μM upper and lower primers 1.2 μL each, sterile water to make up the volume to 20 μL; the PCR amplification program was: 95°C pre-denaturation for 5 min, 1 cycle; 95°C denaturation for 30 s, 58°C annealing for 30 s, 72°C extension for 10 s, a total of 35 cycles; 72°C extension for 5 min, 1 cycle), as shown in Table 1. Figure 2A. Among them, glyA and ducB have non-specific amplification or no amplification product, and the remaining 8 genes have good amplification effect. Further, the same PCR amplification method was used to verify the remaining 8 genes (gntR_1, yhaJ_3, rapz, pgK, nuoH, rpoH, rne, cra_2) of 22 strains of K. pneumoniae strains, as shown in Figure 2 B. The results showed that rpoH, pgK, and cra_2 had different degrees of non-specific amplification, so they were excluded. Subsequently, yhaJ_3, nuoH, rapz, rne, and gntR_1 were selected as potential molecular targets, and the same PCR amplification method was used to verify the specificity in 1 strain of K. pneumoniae and 18 strains of non-K. pneumoniae (Table 2), as shown in Figure 3 Although most genes showed certain specificity, nuoH and rne still appeared similar amplification products in non-target bacteria as in target bacteria, so they were also excluded. Finally, the remaining 3 genes (gntR_1, rapz, yhaJ_3) corresponding to the primer pairs were subjected to PCR amplification program optimization using the same PCR amplification method. The annealing temperature was set to 60°C, 61°C, 63°C, and 65°C, respectively, to improve the specificity of amplification, as shown in Figure 4 After optimization, the yhaJ_3 primer pair had good amplification specificity in K. pneumoniae and did not produce false positive results in non-target bacteria. The nucleotide sequence of yhaJ_3 is shown in SEQ ID NO. 1.

[0058] Table 1. K. pneumoniae species characteristic genes and primer sequences

[0059]

[0060] Table 2 Figure 3 and Figure 4 Test strains corresponding to each lane

[0061] Lane Strain Lane Strain 1( Figure 3 ) Klebsiella pneumoniae 11 Acinetobacter baumannii 2 Escherichia coli 12 Acinetobacter baumannii 3 Salmonella enteritidis 13 Pseudomonas aeruginosa 4 Bacillus thuringiensis 14 Campylobacter jejuni 5 Bacillus cereus 15 Enterobacter cloacae 6 Cronobacter sakazakii 16 Shigella sonnei 7 Pseudomonas aeruginosa 17 Shigella flexneri 8 Listeria monocytogenes 18 Vibrio parahaemolyticus 9 Staphylococcus aureus 19 Proteus vulgaris 10 Yersinia enterocolitica 1( Figure 4 ) Klebsiella pneumoniae

[0062] Example 3: Phylogenetic analysis based on K. pneumoniae serotype characteristic genes

[0063] The downloaded K. pneumoniae whole genome (FASTA file) was automatically annotated using Prokka, and the command line was generated in batches through a local Perl script. Taking the ATCC 46117 strain as an example, Prokka was run in the background using the nohup command (for example, for the K. pneumoniae ATCC 46117 strain, the command line is: nohup prokka --outdir --cpu 120 --prefix KP_46117 --addgenes --locustag KB_46117 --strain KB_46117 --kingdom Bacteria --genus Klebsiella --species pneumoniae --gcode 11. / GCA_030845375.1_ASM3084537v1_genomic.fasta), and 120 CPU cores were specified to accelerate the annotation of the target genome sequence (GCA_030845375.1_ASM3084537v1_genomic.fasta). The output results are stored in the specified directory, and the files are prefixed with "KP_46117". The gene name, local tag, strain information, biological classification, and genetic code table are set in the parameters, and finally the result file containing complete annotation information is generated.

[0064] In order to obtain K. pneumoniae with serotypes K1, K2, K5, K20, K54, and K57, the prokka annotation of the obtained K. pneumoniae whole genome sequence was used to perform BLAST comparison with the serotype characteristic gene sequence (SEQ ID NO. 4-9) using TBtools 2.0, and the identity-confirmed K. pneumoniae of each serotype was obtained, wherein K129, K225, K517, K2021, K5411, and K5722 were obtained.

[0065] A rapid pan-genome analysis was performed on the filtered K. pneumoniae whole genomes of the six serotypes using Roary v3.11.286 with a core genome threshold of 99% (Roary default) and a BLASTP homology threshold of 95%; the input was the GFF3 file generated by Prokka annotation. An example command: nohup roary -p 120 -e -n -i 95 -f fpangenome_95result / *.gff > history_95.log &. This command calls 120 CPUs in the background to process all GFF files in the current directory with a homology clustering threshold of 95%. The homology threshold (95%) can be adjusted according to the research needs. Specifically, a rapid large-scale pan-genome analysis was performed on the filtered K. pneumoniae whole genomes of the six serotypes using Roary v3.11.286. A core genome threshold of 99% was used, and the BLASTP homology threshold was 95%. Only the GFF file (GFFV3 format) containing the sequence and annotation was required as the input file, which was one of the output files of Prokka annotation. The command line example used was: nohup roary -p 120 -e -n -i 95 -f fpangenome_95result / *.gff > history_95.log &. The nohup command was used to ensure that roary ran continuously in the background, and the program would continue to execute even if the terminal session ended. The program took all files with the suffix.gff in the current path as input and performed pan-genome analysis based on a 95% BLASTP similarity threshold.

[0066] Before performing multilocus sequence typing (MLST), the gene presence absence.csv file generated by Roary was used to screen the unique feature gene sequences of each of the 6 serotypes of K. pneumoniae, and the number of serotype-specific feature genes in the corresponding.ffn file of each K. pneumoniae was screened according to the functional annotation information of the screened genes, and the strain name in the gene presence absence.csv table header, the first column of the feature gene name and the gene number of the feature gene in the bacteria were retained. These information was used to extract the nucleotide sequences of the required feature genes from the.ffn format annotation file of each K. pneumoniae strain through a specific script in the Windows command line interface. This process ensures that only relevant feature gene sequences are retained, providing the necessary data for MLST typing. Through this method, non-target sequences and redundant sequences can be effectively removed, ensuring the accuracy and efficiency of the analysis, and finally the sequence data of the feature genes are arranged in separate files, and each feature gene file contains the corresponding feature gene sequence of each K. pneumoniae strain. Specifically, based on the gene presence absence.csv file generated by Roary, the feature genes of the 6 serotypes of K. pneumoniae were screened, and the nucleotide sequences were extracted from the corresponding.ffn file. Then, PhyloSuite was used for multiple sequence alignment (MAFFT), redundant sequence removal (trimAI) and concatenation of feature genes to generate a joint sequence file. The optimal partition scheme and evolutionary model were determined by PartitionFinder2 (AICc criterion, MrBayes inference), and then IQ-TREE was used to construct a maximum likelihood phylogenetic tree under Edge-linked partitioning, with 5000 times of Ultrafast bootstrap method to evaluate branch support. The final results are consistent with the reference genome clustering, thereby confirming the true identity of K. pneumoniae in the NCBI database.

[0067] Finally, the serotype-specific feature genes wzc and cps were mined, and based on this, the K. pneumoniae of each serotype (K129 strain, K225 strain, K517 strain, K2021 strain, K5411 strain, K5722 strain) whose identity was confirmed was used to construct a K. pneumoniae serotype phylogenetic tree Figure 5). The method is as follows: using PhyloSuite to align and concatenate the screened serotype characteristic gene sequences wzc and cps, requiring the sequence names of different strains in each characteristic gene file to be consistent, and using regular expressions in Notepad++ to change the sequence numbers of different strains to different strain names. In the "Alignment" function of PhyloSuite, select the "MAFFT" plug-in for sequence alignment and alignment, then remove redundant nucleotide sequences in the "trimAI" plug-in in the "Alignment" function, and finally concatenate the aligned sequences in the "Concatenate Sequence" in the "Alignment" function. If there are some gene data missing during the running process, it will be ignored. In addition, other parameters are selected as default parameters. In the "Phylogeny" function of PhyloSuite, by selecting the "PartitionFinder2" plug-in, the concatenated sequence file concatenation.phy output by the version is automatically imported, and the optimal partition model and evolution scheme of the used gene concatenated sequence are determined. Parameter settings include selecting Nucleotide as Configuration, setting branchlengths=linked to assume that the branch lengths of all genes or data partitions are related to each other, using models=mrbays to specify MrBayes for Bayesian inference, and applying model_selection=AICc90 (The corrected Akaike information criterion) for model selection to consider the goodness of fit and complexity of the model, and avoid overfitting. The DataBlocks checkbox will automatically identify different characteristic gene partitions in concatenation.phy, if it cannot be automatically identified, copy the contents in partitionfinder style in the previous step result file partition.txt. In the "Phylogeny" function of PhyloSuite, by selecting the "IQ-TREE" plug-in, the output file concatenation.fas and IQ_partition.txt are automatically imported into Alignment File and Partition Mode. In the parameter setting, the Partition Style is selected as Edge-linked, which represents that all partitions (such as genes or coding regions) share the same set of branch lengths in phylogenetic tree inference, similar to the -q option in RAxML. The outgroup is selected as K. pneumoniae, which will root the phylogenetic tree according to this reference species, which is conducive to clearly showing the relationship and evolution direction between species.Bootstrap was set to Ultrafast bootstrap with 5000 resampling to quickly and unbiasedly calculate the branch support to assess the reliability of each branch in the phylogenetic tree. Other parameters were kept default, i.e. using the default setting of IQ-TREE. In the above analysis, the maximum likelihood phylogenetic tree was inferred by 5000 Ultrafast bootstrap under the edge-connected partition model. Meanwhile, the approximate Bayesian test and Shimodiara-Hasegawa-like approximate likelihood ratio test were used to infer the maximum likelihood tree of phylogeny. According to the strains that were consistent with the clustering results of the reference genome, the true identity of K. pneumoniae bacteria in NCBI Datasets taxonomy was confirmed. As can be seen from the figure, there is obvious clustering differentiation between different serotypes, which is consistent with the differentiation results based on the traditional six serotype characteristic genes. Figure 6 Table 3 gives the specific sequence information and annotation function of the characteristic genes wzc and cps.

[0068] Table 3. Sequence information and annotation function of K. pneumoniae serotype characteristic genes wzc and cps

[0069] SEQ ID NO. 1 (nucleotide sequence of yhaJ_3)

[0070] ATGGCCAAAGAGAGAGCATTAACGCTGGAAGCGTTACGCGTGATGGATGCGATCGACCGGCGCGGAAGCTTCGCTGCCGCAGCTGATGAGCTGGGACGCGTGCCTTCCGCCCTGAGCTACACCATGCAAA AACTGGAGGAGGAA CTGGACGTTGACCGCTCAGGTCATCGCACGAAGTTCACTAACGTTGGAAGAATGCTGCTGGAGCGCGGGCGCGTGCTGCTTGAAGCGGCGGATAAACTCACCACCGATGCTGAAGCCCTGTCGCGCGGTTGGGAAACCCACCTGACGATCGTTACCGAAGCGCTGGTGCCCACTCCGGATCTGTTTCCGCTGATTGAGAAACTGGCGACTAAATCAAACACCCAGCTGTCGATTATCACTGAGGTGCTGGCCGGCGCCTGGGAGCGGCTGGAGCAGGGGCGGGCGGACATCGTTGTCGCACCGGATATGCATTTTCGTTCCTCTTCGGAGATTAACTCACGCAAACTCTATTCGGTGCTCAGCGTGTACGTCGCGGCGCCGGATCACCCGATCCATCAGGAGCCGGAGCCGCTCTCTGAGGTTACCCGCGTGAAATACCGCGGCGTCGCGGTGGCGGATACCGCGCGCGAGCGGCCCGTTCTCACCGTTCAACTGCTGGATAAACAGCCGCGTTTAACGGTGAGTACGATAGAAGATAAACGTCAGGCGCTACTGGCCGGACTGGGCGTGGCGACTATGCCCTATCCGCTG GTGGAGAA GGATATTGCCGA AGGTCGCCTGCGGGTCGTCAGCCCGGAATACACCAATGAAATTGACATTATCATGGCCTGGCGACGGGACAGTATGGGCGAAGCGAAGTCCTGGTGTCTGCGTGAGATCCCCAAGCTGTTTGCCGGCAGATAGSEQ ID NO. 2 (nucleotide sequence of wzc)

[0071] ATGACTGATAAACTTAATTCAATGTCTACAGGAAGCCAAGAAAGTGACGGGATCGA

[0072] TCTAGGTCGTTTAATTGGTGAAGTTATTGATCATAAAGGATTGATAATTTCTGTAACTTTC

[0073] TTTTTTATGGGGCTAGGTTTTTTATATGCATTTTTAGCTCCACCCGTATATGAAGCAGATTC

[0074] ATTAGTGCAGGTTGAACAAAATGCAGGAAATAATTTTTTATCGAGCCTTTCTGACGTTCT

[0075] TCCAACAACACCGCCTCAATCAGCAGCTGAAATTGAACTAATTAAATCTAGGATGGTGCT

[0076] TGGTAAAACTATCAATGATTTAAATCTGAGCGTAGTAATTGAAGAGAAAACTACGCCTAT

[0077] TATAGGTCAGTTTTTAAAGAAAATACGTGGTGATGACGGCAGTAAAATAAATGTAAAATA

[0078] TTTTAATGTCCCTCATGATGCGTTAGACACGAAGTTTACAATCAAAATAACTGGTAAAGA

[0079] CAGCTATACCATTAATCTTGATGATGCCGGCGAATTGAAAGGTCAAGTTAATGAGCCCGT

[0080] ATCGAAAAATGGATTCGAAATATTGCTGACGCAAATTGAGTCTCCGCCAGGTACTGAATT

[0081] TAGTATAAAACGCAAAGATACTCTTCAAGTACTCAGTGATCTTAATGACGCATTCACTGT

[0082] CGCAGATACTGGAAAAGATACAGGTGTATTGTCGCTATCATTAACAGGAAATGATCCTGA

[0083] AAAAATAAAGACGATACTTCAAAGTATTACTGATAATTATTTGTTACAGAATATTGAGAGA

[0084] AAATCGGAAGAGGCAGCAAAAAGTTTAAACTTTTTGGATCGAAAAATCCCTGATGTAAA

[0085] AAATGAGCTTAATGCTGCAGAAAATAAGCTAAATTATTACAGACAGCAAAATAGTTCTGT

[0086] TGATTTAACTATGGAGGCAAAATCTCTTCTGGATACAATGGTTCAATTAGATGCGCAGATT

[0087] AATCAACTAACATTTTCTGAGGCCGAAGTATCCAAACTATATACAAAAGAGCATCCAACA

[0088] TATAGAGCTCTTTTAGAAAAAAGAAAAACTCTGGAAGAAGAAAAAAATAATCTTAAGAA

[0089] AAAGATTAATAATCTTCCAGAAACACAACAAGAAGTATTACGTTTAACTCGTGATGTCCA

[0090] GGTTGGGCAAGATGTCTATTTACAACTATTAAACAAAGAACAAGAATTGAGCATTACTAA

[0091] AGCAAGTACAGTGGGTAATGTTCGAATAATTGATAATGCCGTAACACAGCCCGATCCAAT

[0092] CAAACCTAAAAAGCTGCTTGTAATAGTCGTTCTAACGGTGTTTGGAACAATGCTGTCCTT

[0093] GATGTATGTCATAGTTAAGGTTGCGTTCCATAAAGGTGTTCAAAGTGCCGAACAACTTGA

[0094] AGAAAATGGTTTGAATGTGTACGCGAGTATCCCTTTATCTGACTGGCAGTTGAAAAATGT

[0095] AAACAACAGCAGGAAAAAGGATAAGAAAAAATTTGATGTATTAATTAAGGAGAATGGTG

[0096] CAGATTTAGCTATTGAAGCTATTAGAGGGTTACGGACGCGTTTATATTTCGCAATGTTAGA

[0097] AGCGAAAAATAATATTTTAATGATATCAGGTCCAAGTCCTGAAATAGGCAAAACCTTCGT

[0098] CAGTACTAACCTTGCAGGGGTAGTTGCTCAGGCAGGACAAAAAGTTTTATTAATTGATGC

[0099] AGATATGCGACGAGGATATATGCATCACTACTTTGGCGGTTTGCCCAACAATGGCTTGTC

[0100] AGAGATACTCACAGGTAGGGTAGATTACGAGAAGGCTGTTGTTCATACTGATATAGCTGG

[0101] CTTAGATTACATTGGTCGTGGTGAAATACCACCAAATCCAGCGGAGCTATTAATGGGGAG

[0102] TCGCATTGAAAAGTTCCTTGAGTGGGCAAGTGGAAAGTATGATCTTGTATTAGTAGACAC

[0103] TCCGCCAATTTTAGCAGTGACCGATGCTGCAATAATTGGTCGCCATGTAGGAACCACATT

[0104] ATTAGTTGCTCGCTTTGAAAAAAACACTGTCAAAGAAATCGATGTTGCTAAAAATAGATT

[0105] GGAACATAGCGGTGTTATAGTTAAAGGCGTAATATTAAATGCTGTGACCAGAAAAGCTAG

[0106] TAATAAATATGGTGAATATGCGTATTATGAATACGAGTATAAATCGAAAGAATAG SEQ ID NO. 3 (nucleotide sequence of cps)

[0107] ATGGCATCGGTAACGAATAACAAACAAACACCTACCGAATCGGATGATATCGATTTA

[0108] GGTAAAATTGTCGGTGAGTTAATAGACCATCGCAAGCTAATTATTGCTATCACAACAGCT

[0109] TTTACGGTTATAGCAGTGCTCTATGCATTACTAGCGACCCCAATATATCAATCAACAGCAC

[0110] TAATTCAGGTTGAGCAAAAACAAGGTAACGCCATTTTAGATAGTCTGAGTCAGATGCTTC

[0111] CTGATAGCCAACCACAGTCAGCGCCTGAAATAGCATTAATTCAGTCTCGAATGATATTAG

[0112] GGAAAACTGTAGATGATCTAAATTTACAAGCAGAAATCGAGCCAAAGTATTTCCCGATAT

[0113] TTGGACGAGGTTTGGCAAGATTGCTGGGTAAAGAGCCCGGGACTATATCAGTCCCAAGA

[0114] TTCTATTTAGATACTGGAAGTAATGATGTACCCTCTGAGGTTACTTTAACTATCTTGGGTG

[0115] AGAATAATTTTGAAATTGAGGGGGAAGGGTTTTCTTTAAAAGGTAAAAAGGGTGTTCTG

[0116] TTGGAAGATAAGGGAGTTTCAATTCTCGTAGACTCTATCGATGCCCAACCTGGAAGTCAA

[0117] TTCAAAATAACCTATATAAGTCGTTTGAAAGCTATTAGCAACTTGTTAGAGTCATTAAATG

[0118] TAGCAGACCAAGGGAAAGATACTGGGATGTTAAATTTGACTTTTACAGGCGATAATCCG

[0119] ACTTTAATATCACAAGTTTTAAGTAGCATTACTCAGAACTATCTTGCGCAAAATGTTGCA

[0120] AGACAAGCAGCTCAAGATGCAAAAAGTCTTGAATTTCTGAATGAACAATTACCAAAGGT

[0121] AAGAACTGATTTAGATGCAGCTGAGGATAAGTTAAATAGCTATCGTAAGCAAAAAGACT

[0122] CTGTTGATTTAACAATGGAGGCTAAATCTGTTTTGGATCAGATAGTAAACGTTGATAATCA

[0123] ACTTAATGAACTTACTTTCAGAGAAGCCGAGATTTCTCAGCTATATACGAAAGAACATCC

[0124] AACTTATAAAGCGTTAATGGAGAAAAGACAAACATTACAAACGGAGAGGAATAAATTAA

[0125] ACAAGAAAGTTAGCTCAATGCCCTCGACTCAACAAGAGGTTTTGAGATTAAGCAGAGAT

[0126] GTTGAGTCCGGTAGGGCTGTGTATTTGCAATTGCTGAGCCGACAGCAAGAGCTAAATATT

[0127] GCCAAATCTAGTGCCATAGGTAACGTACGTATTATTGATAATGCAATTACTGAACCGAAA

[0128] CCGGTAAAACCTAAAAAAATTCTTGTGATTGCTTTAGGTATTATCATCGGGTTATTCTTTAT

[0129] CTGTAGGTTTTGTTTTCGTCAGAGTATTCTTGCGGAGAGGTATCGAGTCCCCTGAGCAGC

[0130] TAGAAGAAATGGGAATCAATGTATACGCGAGCATCCCTGTATCAGAATGGCTTACTAAAA

[0131] ACACTAATAAAAAATAAAGACAAAAAAATGAATCTGATACATTGTTAGCTGTTGAAAAC

[0132] CCAGCAGATTTGGCTGTTGAAGCTATCAGAAGTTTAAAGAACTAGTCTTCATTTTGCAATG

[0133] ATGGAGTCGAAAAAATACATATTAATGATTTCTGGAGCTAGTCCAAATGCGGGCAAAACA

[0134] TTCGTAAGTACGAATTTAGCCGCTACAATCGCGATGACAGGAAAGAAAGTTCTGTTTATC

[0135] GACTCTGATCTCCGAAAGGATACGTTCACAAAATGTTGGGTTCGGAAAATGTCAAAGG

[0136] TTTATCTGATATTTATCTGGCCAAGCGAAAGTTGAAAGTATTATCAAAAGAGTCAGTGG

[0137] GGGGGGATTTGATTATATTGGTCGTGGACAAACACCACCAAATCCTGCAGAGTTGCTGAT

[0138] GCATCCTCGATTCAAGGAACTATTATCTTGGGCATCGCAGAACTATGAATTAGTAATTGTC

[0139] GATACGCCTCCAATTTTGGCTGTTACCGATGCGGCAATAATCGGGCAATATGCTGGAACA

[0140] ACTCTACTTGTAGCTCGTTTTGAAGCAAATACAGCTAAAGAAATTGCTGTAAGTATTAAA

[0141] CGCTTCGAACAAACAGGCGTAGTTATTAAAGGGTGTATCTTGAATGGTGTAATGAAAAA

[0142] AGCGAGTAGCTATTATAGCTATGGCTATAGCCAATATGGCTACTCATATACAGATAATAAAT

[0143] CTAAATAASEQ ID NO.4 (K1 serotype characteristic gene sequence)

[0144]

[0145]

[0146]

[0147]

[0148]

Claims

1. Use of the signature target yhaJ_3 in the identification of Klebsiella pneumoniae species level for non-disease diagnostic and / or therapeutic purposes, characterized in that, The nucleotide sequence of yhaJ_3 is shown as SEQ ID NO.

1.

2. Use of a primer pair for the identification of the species level of Klebsiella pneumoniae for diagnostic and / or therapeutic purposes other than disease, characterized in that, The primer pair comprises AACTGGAGGAGGAACTGGAC and TCGGCAATATCCTTCTCCAC.

3. Use of a kit for the identification of the species level of Klebsiella pneumoniae for non-disease diagnostic and / or therapeutic purposes, characterized in that, The kit comprises the primers AACTGGAGGAGGAACTGGAC and TCGGCAATATCCTTCTCCAC.

4. A method for identifying the species level of Klebsiella pneumoniae for non-disease diagnostic and / or therapeutic purposes, characterized in that, The method comprises the following steps: Genomic DNA of the sample to be tested is extracted, and PCR amplification is performed by taking the genomic DNA as a template and AACTGGAGGAGGAACTGGAC and TCGGCAATATCCTTCTCCAC as amplification primers; if a specific 634bp amplification band appears in the amplification product, the sample to be tested is Klebsiella pneumoniae, otherwise, it is indicated that the sample to be tested is not Klebsiella pneumoniae.

5. The method of claim 4, wherein, The annealing temperature of the PCR amplification is 60-65°C.

6. The method of claim 5, wherein, The annealing temperature of the PCR amplification is 63°C.

7. The method of claim 4, wherein, The PCR amplification system is 2xHieffUltra-RapidII HotStart PCR Master Mix 10μL, template DNA 1μL, 10μM upper and lower primers 1.2μL each, and sterile water is added to make up the volume to 20μL; the PCR amplification procedure is: 95°C pre-denaturation for 5min, 1 cycle; 95°C denaturation for 30s, 63°C annealing for 30s, 72°C extension for 10s, a total of 35 cycles; 72°C extension for 5min, 1 cycle.

8. Application of characteristic targets wzc and cps in serotype typing of Klebsiella pneumoniae.

9. A method for serotyping of Klebsiella pneumoniae based on the characteristic targets wzc and cps, characterized in that, The method comprises the following steps: DNA of the sample to be tested is extracted, and all strain samples are obtained to obtain whole genome and whole genome information; the strain sample and the serotype characteristic gene sequence wzc and cps of each serotype K. pneumoniae with identity certification are extracted and subjected to sequence alignment and alignment, and then redundant nucleotide sequences are removed, and finally the aligned sequences are connected in series to construct a maximum likelihood method phylogenetic tree, so as to identify the serotype of the strain to be tested.

10. The method of claim 9, wherein, The maximum likelihood method phylogenetic tree is set to 5000 times Ultrafast.