Construction method and application of pathogenic microorganism genome database

Through screening and simulation amplification operations, a pathogenic microbial genome database integrating multiple strain sequences was constructed, solving the problem of difficult to balance database coverage and scale in the prior art, achieving higher detection accuracy and faster analysis speed.

CN119993288APending Publication Date: 2025-05-13GUANGZHOU JINQIRUI BIOTECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510075497.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, when constructing a database of pathogenic microbial genomes, it is difficult to balance the coverage and scale of the database, resulting in limited accuracy and efficiency of the detection results.

Method used

Through screening, filtering and simulation amplification, a genomic database integrating all reliable strain deduplication sequences of target species and reliable strain deduplication sequences of similar species is constructed to reduce false positives and false negatives and improve detection accuracy.

Benefits of technology

It has achieved the reduction of the risk of errors in identification of pathogenic species, improved the detection accuracy of microbial classification, and reduced the database scale and computing resource requirements, and improved the analysis speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993288A_ABST
    Figure CN119993288A_ABST
Patent Text Reader

Abstract

The invention discloses a construction method and application of a pathogenic microorganism genome database. The method comprises the steps of genome acquisition, genome screening, target region grabbing, strain verification, target region sequence clustering, redundancy elimination and database merging. The construction method of the pathogenic microorganism genome database is suitable for targeted high-throughput sequencing, and the corresponding amplification relationship between the primer and the corresponding microorganism genome sequence is obtained through the operations of screening, filtering, simulation amplification and the like on pathogenic microorganism genome data; the risk of wrong identification of pathogenic species can be reduced by constraining the corresponding relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of high-throughput sequencing and relates to a method for constructing a pathogenic microorganism genome database and its application. Background Art

[0002] Targeted Next Generation Sequencing (tNGS) is a precise molecular diagnostic technology that designs specific primers to amplify the DNA fragments of target pathogens and then uses a high-throughput sequencing platform for sequencing analysis.

[0003] In the field of pathogen detection, tNGS technology is favored for its advantages of high sensitivity, high specificity and low cost. Compared with traditional culture methods and non-targeted high-throughput sequencing, tNGS technology is not only more economical, but also provides results faster, and is especially suitable for the detection of low-abundance pathogens. At present, tNGS technology has been widely used in clinical infection diagnosis, epidemic monitoring, public health and safety, and other fields, and is of great significance for guiding clinical precision treatment and epidemic prevention and control.

[0004] Since tNGS technology selects specific target areas for pathogenic microorganisms to amplify the disease, the sequencing area is limited, which puts higher requirements on the establishment of a reference genome database. How to establish a pathogenic microorganism reference genome database that is accurately classified, fully covered, and as concise as possible is a difficult point. At present, most pathogenic microorganism genome databases are collected in public databases such as NCBI and constructed after screening and processing. In public databases, each species may have multiple different strain genomes. How to select these strain genomes to represent the species and apply them to subsequent tNGS comparisons is also a problem that needs to be solved.

[0005] At present, the mutation rate of genomes between most strains is as high as more than 3%. In clinical practice, if only one strain is selected as the representative genome of the species, it is difficult to cover some strains with higher mutation rates of the species, which often results in missed detection and false negative test results. The use of a database strategy containing the genomes of all strains of the same species can significantly reduce missed detection, but this method also has significant defects. First, this method will make the database large, thereby prolonging the analysis time, which is impractical for clinical applications that require rapid response. At the same time, it also greatly increases the demand for computing resources, resulting in a surge in analysis costs. Secondly, the quality of strain genome sequences obtained from public databases varies, and some sequences may contain contaminated or misclassified strains. If they are not screened and filtered, they may cause false positives, causing trouble for clinical diagnosis and treatment.

[0006] Therefore, there is an urgent need to provide a method for constructing a pathogenic microorganism genome database suitable for targeted high-throughput sequencing. Summary of the invention

[0007] In view of the deficiencies in the prior art and actual needs, the present invention provides a method for constructing a pathogenic microorganism genome database and its application. The genome database constructed using this method not only integrates the deduplicated sequences of all reliable strains of the target species and retains rich species and strain information, but also incorporates the deduplicated sequences of reliable strains of similar species, thereby reducing false positives and false negatives and improving the detection accuracy of microbial classification. It also has the advantages of small hard disk space occupation, convenient external installation, fast comparison speed, and saving computing resources.

[0008] In order to achieve the purpose of the invention, the present invention adopts the following technical solutions:

[0009] In a first aspect, the present invention provides a method for constructing a pathogenic microorganism genome database, the method comprising the following steps:

[0010] (1) Genome acquisition: obtaining specific pathogenic microorganism genome data from public databases;

[0011] (2) Genome screening: Selecting the genome of species and strains according to predetermined screening rules;

[0012] (3) Target region capture: simulate the amplification of the strain genome selected in step (2) according to the primers to obtain the target region sequence;

[0013] (4) Strain verification: Filter out incorrectly classified strains and analyze the causes of strains that do not produce products in simulated amplification;

[0014] (5) Target region sequence clustering: cluster analysis of target region amplified sequences;

[0015] (6) Redundancy removal: De-redundancy is performed on sequences with a specific similarity based on the clustering results, and representative reference sequences are retained;

[0016] (7) Database merging: Replace the primers in step (3), repeat steps (3) to (6), obtain multiple representative genome sequences, and integrate them to generate a representative database of the pathogenic microorganism genome.

[0017] In the present invention, the method for constructing a pathogenic microorganism genome database is suitable for targeted high-throughput sequencing. By screening, filtering and simulating amplification of pathogenic microorganism genome data, the corresponding amplification relationship between primers and corresponding microbial genome sequences is obtained. By constraining the corresponding relationship, the risk of incorrect identification of pathogenic species can be reduced.

[0018] The database of the present invention is established as follows:Figure 1 shown.

[0019] Preferably, the genome acquisition in step (1) includes: obtaining genome sequences of all attribution levels of specific pathogenic microorganisms through a public database.

[0020] In the present invention, the specific pathogenic microorganisms may be any hierarchical system in biological taxonomy, namely, kingdom, phylum, class, order, family, genus, species, or specific strain typing.

[0021] Preferably, the genome sequences of all attribution levels include the genome of the upper level or the genome of the upper level in biological taxonomy.

[0022] Preferably, when the specific pathogenic microorganism is at the species level or a specific typing strain, the genome sequence at the genus level is obtained.

[0023] Preferably, when the specific pathogenic microorganism is at the genus level, the genome sequence at the family level is obtained.

[0024] Preferably, the public database includes any one of the PATRIC database, the RefSeq database, the Genbank database or the Nucleotide database, or a combination of at least two thereof.

[0025] Preferably, the step (2) of selecting the species strain genome according to a predetermined screening rule comprises the following steps:

[0026] (a) Remove species with unclear genus and species names, remove bacteriophages from viruses, remove species with one of the following keywords in the Latin name: Uncultured, Unclassified, Unidentified, Cloning, clone, Synthetic, phage, plasmad, Recombinant, Predicted, environmental, and species without genus names in the Latin name;

[0027] (b) removing genome sequences with a length of less than 300 bp or a N base ratio higher than 50%;

[0028] (c) Removal of genome similarity ANI does not match the species level genome.

[0029] Preferably, the step (c) of removing genomes whose genome similarity ANI does not meet the species level includes: removing genomes whose genome similarity ANI is less than 95% from species-level microbial genomes, and removing genomes whose genome similarity ANI is less than 90% from genus-level genomes.

[0030] The ANI mentioned in the present invention refers to Average Nucleotide Identity (ANI), which is an indicator used to evaluate the similarity of genomes between microbial species. The higher the ANI value, the greater the similarity between the two genomes.

[0031] Preferably, the simulation amplification software in step (3) includes any one of BLAST, NCBI Primer-BLAST, SnapGene, e-PCR or UCSC In-Silico PCR, or a combination of at least two thereof.

[0032] Preferably, the filtering of misclassified strains in step (4) includes aligning the strain sequences with the misclassified strains into the database through the BLAST module of NCBI, checking the closest microbial classification, and verifying according to the alignment results; the analysis of the cause of the strains with no products in the simulated amplification includes aligning and locating the sequences 100-1000bp upstream and downstream of the primer sequence (for example, 100bp, 200bp, 300bp, 500bp, 800bp, 1000bp).

[0033] Preferably, the cause analysis of the strain with no product in the simulated amplification includes comparing and locating the sequences 150-250 bp upstream and downstream of the primer sequence (eg, 150 bp, 200 bp, 250 bp).

[0034] Preferably, the cluster analysis software in step (5) includes any one of cd-hit, SeqKit, USEARCH, VSEARCH, QIIME, Mothur or DADA2.

[0035] Preferably, the specific similarity in step (6) is 80%-100% (eg, 80%, 81%, 85%, 90%, 100%).

[0036] Preferably, the specific similarity Identity is set to 100%, which can ensure that all product sequences generate representative sequences in the reference genome.

[0037] In a second aspect, the present invention provides a pathogenic microorganism genome database, which is constructed by the construction method described in the first aspect.

[0038] In a third aspect, the present invention provides an application of the pathogenic microorganism genome database described in the second aspect in targeted high-throughput sequencing.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] (1) The method for constructing a pathogenic microorganism genome database of the present invention is suitable for targeted high-throughput sequencing. By screening, filtering and simulating amplification of pathogenic microorganism genome data, the corresponding amplification relationship between primers and corresponding microorganism genome sequences is obtained. By constraining the corresponding relationship, the risk of incorrect identification of pathogenic species can be reduced;

[0041] (2) The genome database obtained by the present invention not only integrates the deduplicated sequences of all reliable strains of the target species, retains rich species strain information, but also incorporates the deduplicated sequences of reliable strains of similar species, thereby reducing false positives and false negatives and improving detection accuracy; at the same time, since the repeated target region sequences within the species are removed, the amount of data is significantly reduced, which not only occupies less hard disk space, but also speeds up analysis time and saves computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 Establishing a flow chart for the database of the present invention;

[0043] Figure 2 This is the result diagram of strain verification comparison in Example 1;

[0044] Figure 3 A database analysis time box plot is established for different methods in Example 2. DETAILED DESCRIPTION

[0045] To further illustrate the technical means and effects of the present invention, the present invention is further described below in conjunction with the embodiments and drawings. It should be understood that the specific implementation methods described herein are only used to explain the present invention, rather than to limit the present invention.

[0046] If no specific techniques or conditions are specified in the examples, the techniques or conditions described in the literature in the field or the product instructions are used. If no manufacturer is specified for the reagents or instruments used, they are all conventional products that can be purchased through regular channels.

[0047] Example 1

[0048] The present invention provides a method for establishing a pathogenic microorganism genome database suitable for targeted high-throughput sequencing, comprising the following steps:

[0049] In order to better illustrate the process of building a database, the primers primer01 and primer02 of Neisseria gonorrhoeae are taken as an example (Table 1).

[0050] Table 1

[0051]

[0052]

[0053] (1) Genome acquisition: Neisseria gonorrhoeae is at the species level. Therefore, when acquiring the genome, it is necessary to obtain the genome of the Neisseria genus. The specific method is to download the genome data of the Neisseria genus and its sub-classifications from the RefSeq and Genbank databases of NCBI (ftp: / / ftp.ncbi.nlm.nih.gov / genomes / ), which contains a total of 4290 strains of genome information, divided into 45 species.

[0054] (2) Genome screening: For the downloaded genomes, species with one of the following keywords in the Latin name: Uncultured, Unclassified, Unidentified, Cloning, clone, Synthetic, phage, plasmid, Recombinant, Predicted, environmental, and species without genus names in the Latin name were removed. Then, genomes with a genome sequence length of less than 300 bp or a N base ratio of more than 50% were removed, leaving 4177 strains, which were divided into 43 species.

[0055] Furthermore, for the remaining strain sequences, the fastANI software was used to calculate the ANI similarity between each pair to obtain the ANI matrix. Strains with ANI similarity within the species less than 95% and ANI similarity between species less than 90% were deleted, leaving 4100 genome sequences.

[0056] (3) Target region capture: For the screened strain sequences, use electronic e-PCR software to perform simulated amplification on primers prmer01 and prmer02 respectively, obtain the amplification product sequence of the primers, and append it to the back of each strain. The simulated amplification results are shown in Table 2 below.

[0057] Table 2

[0058]

[0059]

[0060] (4) Strain verification: For both primer01 and primer02, there was a strain of Neisseria gonorrhoeae without a simulated amplification product sequence, with the assembly number GCF_900087905.2. Through prmer01, in other genomes with product sequences, taking GCF_900087635.2 as an example, its product sequence region was NZ_LT591897.1:937922-938066, and its downstream 200bp sequence was aligned to the GCF_900087905.2 genome sequence for positioning. The results are shown in Figure 2 ,From the results, it can be seen that the primer cannot amplify the NZ_LT592159.1 sequence number in the GCF_900087905.2 strain, but its product sequence can be successfully obtained through the downstream sequence alignment position. At the same time, it is suspected that the strain number may have a species classification error;

[0061] Furthermore, the species of the strain numbered GCF_900087905.2 was confirmed by placing its subordinate genome sequence numbered NZ_LT592159.1 into the BLAST module of NCBI for comparison, selecting core_nt for the genome library, and obtaining the first comparison result with its own number, but the subsequent best comparison results were all Neisseriameningitidis, i.e. Neisseria meningitidis, indicating that the classification of the strain may be incorrect and needs to be corrected or deleted.

[0062] (5) Target region sequence clustering: The target region amplified sequences were clustered using cd-hit software. Specifically, for the Neisseria Gonorrhoeae species of the prmer01 primer, a total of 182 strains were clustered, and then the product sequences of the two strains of Neisseria Meningitidis were clustered. The prmer02 primer was clustered in the same manner as above.

[0063] (6) De-redundancy: De-redundancy was performed on sequences with a similarity of 100% according to the clustering results, and representative reference sequences were retained. Specifically, the prmer01 primer retained only two strain numbers after de-redundancy of the product of Neisseria Gonorrhoeae, and also retained two strain numbers after de-redundancy of the product of Neisseria Meningitidis. The primer02 primer was also de-redundanted according to the above operation, and the products of Neisseria Gonorrhoeae and Neisseria Meningitidis were both 2 strains.

[0064] (7) Database merging: Through steps (3) to (6), representative genome sequences corresponding to primers primer01 and primer02 are obtained respectively, and they are integrated to generate a representative reference database of the genome of Neisseria gonorrhoeae.

[0065] Through the above method, a microbial genome database of amplification primers corresponding to Neisseria gonorrhoeae was obtained, which not only included the target species Neisseria gonorrhoeae, but also the similar species Neisseria meningitidis.

[0066] Comparative Example

[0067] Compared with conventional library construction methods.

[0068] In order to evaluate the effect of the reference genome database of Neisseria gonorrhoeae constructed in the above Example 1, all untreated genomes of Neisseria gonorrhoeae in Example 1 were identified as ALL_genome, one strain was selected as a representative genome and identified as One_genome, and the genome constructed using the method of the present invention in Example 1 was identified as This_genome. The three databases were compared in terms of database size, accuracy, and analysis time.

[0069] (1) Database size comparison

[0070] Table 3

[0071] database size One_genome 2.2M ALL_genome 467M This_genome 13M

[0072] The database size comparison results are shown in Table 3. The data of ALL_genome is 36 times that of This_genome by the method of the present invention, indicating that the database constructed by the method of the present invention can significantly reduce storage and computing resources.

[0073] (2) Comparison of accuracy evaluation

[0074] Three strains were extracted from the genomes of Neisseria gonorrhoeae and Neisseria meningitidis, respectively, and the product sequences of the primer sequence of primer01 on their genomes were extracted. Their sequences were simulated into simulated data sets with a sequencing length of 101bp and a depth of 5000×. The mem module of bwa ​​software was used to compare these simulated samples with the above three databases, and the statistical results are shown in Table 4 below.

[0075] Table 4

[0076]

[0077]

[0078]

[0079] From the results in Table 4, it can be seen that the species classification accuracy of One_genome is 50%, among which the best matching rate of two simulated samples of Neisseria gonorrhoeae fails to reach 100%, and although the best matching rate of ALL_genome for Neisseria gonorrhoeae is 100%, it fails to correctly classify the products of Neisseria meningitidis, and the species classification accuracy is still 50%, while the database This_genome established by the method of the present invention can correctly distinguish the product sequences of Neisseria meningitidis and Neisseria gonorrhoeae, and the species classification accuracy and the best matching rate are both 100%. It shows that the accuracy of the present invention is higher than that of the conventional method.

[0080] (3) Comparison of analysis time

[0081] In the above accuracy evaluation comparison, the comparison time fluctuations of the six samples in the three databases are shown in Figure 3 Box plot. Figure 3 It can be seen that the average analysis time of the One_genome database is 1.2s, the average analysis time of the ALL_genome database is 22.1s, and the average analysis time of the This_genome database of the method of the present invention is 2.0s. It can be seen that the database constructed by the method of the present invention has the advantages of low analysis resource requirements and short analysis time, and is superior to the database established by the conventional method.

[0082] In summary, the method for constructing a pathogenic microorganism genome database of the present invention is suitable for targeted high-throughput sequencing. By screening, filtering and simulating amplification of pathogenic microorganism genome data, the corresponding amplification relationship between primers and corresponding microbial genome sequences is obtained. By constraining the corresponding relationship, the risk of incorrect identification of pathogenic species can be reduced.

[0083] The applicant declares that the present invention illustrates the detailed method of the present invention through the above-mentioned embodiments, but the present invention is not limited to the above-mentioned detailed method, that is, it does not mean that the present invention must rely on the above-mentioned detailed method to be implemented. Those skilled in the art should understand that any improvement of the present invention, equivalent replacement of various raw materials of the product of the present invention, addition of auxiliary components, selection of specific methods, etc., all fall within the protection scope and disclosure scope of the present invention.

Claims

1. A method for constructing a pathogenic microorganism genome database, characterized in that: The method comprises the following steps: (1) Genome acquisition: obtaining specific pathogenic microorganism genome data from public databases; (2) Genome screening: Selecting the genome of species and strains according to predetermined screening rules; (3) Target region capture: simulate the amplification of the strain genome selected in step (2) according to the primers to obtain the target region sequence; (4) Strain verification: Filter out incorrectly classified strains and analyze the causes of strains that do not produce products in simulated amplification; (5) Target region sequence clustering: cluster analysis of target region amplified sequences; (6) Redundancy removal: De-redundancy is performed on sequences with a specific similarity based on the clustering results, and representative reference sequences are retained; (7) Database merging: Replace the primers in step (3), repeat steps (3) to (6), obtain multiple representative genome sequences, and integrate them to generate a representative reference database of the pathogenic microorganism genome.

2. The method for constructing a pathogenic microorganism genome database according to claim 1, characterized in that: The genome acquisition in step (1) includes: obtaining genome sequences of all attribution levels of specific pathogenic microorganisms through a public database; Preferably, the genome sequences of all levels of attribution include the genome of the upper level or the genome of the upper level in biological taxonomy; Preferably, when the specific pathogenic microorganism is at the species level or a specific typing strain, the genome sequence at the genus level is obtained; Preferably, when the specific pathogenic microorganism is at the genus level, the genome sequence at the family level is obtained; Preferably, the public database includes any one of the PATRIC database, the RefSeq database, the Genbank database or the Nucleotide database, or a combination of at least two thereof.

3. The method for constructing a pathogenic microorganism genome database according to claim 1 or 2, characterized in that: The step (2) of selecting the species strain genome according to the predetermined screening rules comprises the following steps: (a) Remove species with unclear genus and species names, remove bacteriophages from viruses, remove species with one of the following keywords in the Latin name: Uncultured, Unclassified, Unidentified, Cloning, clone, Synthetic, phage, plasmad, Recombinant, Predicted, environmental, and species without genus names in the Latin name; (b) removing genome sequences with a length of less than 300 bp or a N base ratio higher than 50%; (c) Removal of genome similarity ANI does not match the species level genome.

4. The method for constructing a pathogenic microorganism genome database according to claim 3, characterized in that: The step (c) of removing genomes whose genome similarity ANI does not meet the species level includes: removing genomes whose genome similarity ANI is less than 95% from the species-level microbial genomes, and removing genomes whose genome similarity ANI is less than 90% from the genus-level genomes.

5. The method for constructing a pathogenic microorganism genome database according to any one of claims 1 to 4, characterized in that: The simulation amplification software in step (3) includes any one of BLAST, NCBI Primer-BLAST, SnapGene, e-PCR or UCSC In-Silico PCR, or a combination of at least two thereof.

6. The method for constructing a pathogenic microorganism genome database according to any one of claims 1 to 5, characterized in that: The filtering of misclassified strains in step (4) includes comparing the strain sequences with misclassified strains to the database through the BLAST module of NCBI, checking the closest microbial classification, and verifying according to the comparison results; the analysis of the cause of the strains with no products in the simulated amplification includes comparing and locating the sequences 100-1000bp upstream and downstream of the primer sequences.

7. The method for constructing a pathogenic microorganism genome database according to any one of claims 1 to 6, characterized in that: The cluster analysis software in step (5) includes any one of cd-hit, SeqKit, USEARCH, VSEARCH, QIIME, Mothur or DADA2.

8. The method for constructing a pathogenic microorganism genome database according to any one of claims 1 to 7, characterized in that: The specific similarity in step (6) is 80%-100%.

9. A pathogenic microorganism genome database, characterized in that: The pathogenic microorganism genome database is constructed by the construction method described in any one of claims 1-8.

10. Use of the pathogenic microorganism genome database according to claim 9 in targeted high-throughput sequencing.

Citation Information

Cited By

  • Measurement method of sexual reproduction animal germline mutation rate and application

    CN121171342A

  • Nucleic acid sequence database construction method, device and equipment and readable storage medium

    CN121350002A

  • Nucleic acid sequence database construction method, device, equipment and readable storage medium

    CN121350002B