Screening methods and models for target pathogen-specific core genome sequences

By aligning with the reference genome, aligning and filtering, the core genomic region sequences of the target pathogen were screened, and the problem of insufficient screening efficiency and accuracy of multi-target combinations in the prior art was solved, and efficient and specific pathogen-specific sequence screening was achieved.

CN119339801BActive Publication Date: 2025-05-23BEIJING GOLDEN KEY MEDICAL LAB CO LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410768242.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-14
Publication Date
2025-05-23
Estimated Expiration
2044-06-14

AI Technical Summary

Technical Problem

When facing multiple target combinations, existing pathogen-specific target sequence screening methods are difficult to take into account the interference problems between primers of each target sequence, and the parameters and processing methods during the filtration process are inconsistent, resulting in insufficient efficiency and accuracy.

Method used

By directly comparing with the selected reference genome sequence, the genome coverage depth and number of covered genomes were counted statistically, and the sequence variation of the coverage area was analyzed to screen out the core genome region sequences of the target pathogen. Then, these sequences were aligned and filtered with the human reference genome, the reference genome of other species in the same family, and the NT library to obtain the final target pathogen-specific sequence.

Benefits of technology

This method can efficiently and simply screen core region sequences within the genome level, taking into account the interference problems in the design of multiple PCR primer probes, and improve the screening efficiency and accuracy of target pathogen-specific sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119339801B_ABST
    Figure CN119339801B_ABST
Patent Text Reader

Abstract

The present application relates to the field of bioinformatics technology, and specifically discloses a method for screening target pathogen-specific core genome sequences and a model thereof. The present application targets different target pathogens, based on direct comparison of the genome sequence with a selected reference genome sequence, and statistically analyzes the coverage depth of different regions of the genome and the number of covered genomes, etc., to screen the target pathogen core genome region sequence; then the core genome region sequence is compared with the human reference genome sequence, the reference genome sequence of all other species of the same family as the target pathogen, and the NT library, and the final target pathogen-specific sequence is filtered out. The present application can screen out specific core sequences with high coverage and high copy number, and has the advantages of high efficiency, simplicity, and short time consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of bioinformatics, and in particular to a screening method, model and application of a target pathogen-specific core genome sequence. Technical Background

[0002] Compared with traditional culture and other detection methods, mNGS (metagenomic next generation sequencing) sequencing detection technology has the advantages of high sensitivity, specificity and wide coverage. It has been widely recognized in the clinic and is widely used in the detection and diagnosis of clinical infectious pathogens. However, since the current detection cost of mNGS is still relatively high, in comparison, detection methods based on multiple amplification, such as tNGS detection (multiple targeted sequencing technology) and ddPCR (multiple digital PCR), have higher detection sensitivity, cost-effectiveness and / or detection timeliness, and are therefore more easily accepted by patients and favored by clinicians.

[0003] Affected by and limited by the number of multiple target combinations, the development of pathogen detection methods based on multiple amplification (tNGS or ddPCR) first requires the selection of appropriate target pathogen-specific target sequences. In addition to collecting target genes by consulting existing literature or materials, the existing methods for screening pathogen-specific target sequences are to directly screen from scratch based on the whole genome level. There are also some literature or patented technologies for de novo screening of pathogen-specific sequences. Overall, the general idea is to find the core sequence through comparative genomic analysis of all genomic sequences of the target species, and then filter out non-specific alignments by comparing the gene sequences of closely related non-target species. However, these published methods use different processing methods or parameters in the processes of core sequence search and non-specific sequence filtering. They can achieve the search for a single pathogen-specific sequence, but when faced with the need to combine multiple targets in the future, it is still unknown whether the impact of interference between primers of each target sequence is taken into account.

[0004] In view of this, it is necessary to develop a more effective screening and evaluation process for specific sequences of different target pathogens. Summary of the invention

[0005] In order to solve the above technical problems, this application proposes a method for screening target pathogen-specific core genome sequences. This method targets different target pathogens, based on the genome sequence, by directly comparing it with the selected reference genome sequence, statistically calculating the coverage depth of different regions of the genome and the number of covered genomes, and analyzing the sequence variation of the covered region, to screen the target pathogen core genome region sequence; then compare the core genome region sequence with the human reference genome sequence, the reference genome sequence of all other species in the same family of the target pathogen, and the NT library, and obtain the final target pathogen-specific sequence after filtering. Specifically, this application proposes the following technical solutions:

[0006] The present application first provides a method for screening target pathogen-specific core genome sequences, comprising the following steps:

[0007] 1) Obtain the target pathogen genome sequence;

[0008] 2) Align with the reference genome sequence and statistically calculate the genome coverage depth, genome coverage number and / or variation information;

[0009] 3) Select the target pathogen core genome sequence;

[0010] 4) Obtaining the target pathogen-specific core genome sequence: The target pathogen core genome sequence is compared, annotated, and filtered with the human reference genome, the constructed reference genomes of all species in the same family as the target pathogen, and the NCBI NT database to obtain the target pathogen-specific core genome sequence.

[0011] The specific core genome sequence described herein refers to a sequence fragment that is carried by the genome of each strain in the target pathogen and is unique to the target pathogen relative to other bacteria;

[0012] In some aspects, in step 1), the downloading is downloading the target pathogen genome sequence from the NCBI genome database;

[0013] Preferably, the step 1) specifically comprises: for the target pathogen, downloading assembly_summary_genbank.txt and assembly_summary_refseq.txt files from the NCBI genome database; more preferably, if the number of target pathogen genomes recorded in assembly_summary_refseq.txt is greater than or equal to 1000, downloading the corresponding genome sequence according to the assembly_summary_refseq.txt file, otherwise downloading the corresponding genome sequence according to the assembly_summary_genbank.txt file.

[0014] In some aspects, the step 2) is specifically as follows: selecting a reference genome of the target pathogen in the NCBI genome database as a reference genome sequence, aligning all downloaded genomes with the reference genome sequence, and outputting the SNP and InDel information of the reference sequence in each alignment; filtering each alignment result of the query sequence, and retaining only the best alignment as the final result; statistically calculating the number of genome coverages, genome coverage depth, and / or SNP / InDel variation information and its frequency at each site of the reference genome sequence;

[0015] Preferably, if the target pathogen is at the complex group level or genus level, the genome coverage number, genome coverage depth and / or SNP / InDel variation information and its frequency of each species in the complex group or genus are statistically calculated simultaneously.

[0016] In some aspects, in step 3), the selection is: based on the genome coverage number information at each site of the reference genome sequence obtained in step 2, calculate the ratio of the genome coverage number at each site of the reference genome to the total number of strain genomes, filter sites with a ratio lower than 95%, and then connect the bases of adjacent sites to obtain the core genome sequence of the target pathogen;

[0017] Preferably, if the target pathogen is at the complex group level or the genus level, the sites where the ratio of the number of genomes covered by each species within the complex group or genus to the total number of genomes of the strain is less than 95% are filtered;

[0018] More preferably, the core genome sequence is not less than 70 bp in length.

[0019] In some aspects, the step 4) further includes: further screening out sequences with a genome coverage depth that is more than 1.5 times the genome coverage number as multi-copy core genome sequences.

[0020] In some aspects, in step 4), the comparison, annotation and filtering specifically include:

[0021] The target pathogen core genome sequence was aligned with the human reference genome. For hits that were aligned with the human reference genome and had a homology of >= 90%, the corresponding query core sequence was first filtered out; then it was aligned with the constructed reference genome sequences of all species in the same family as the target pathogen, and different target sequences of each query sequence were aligned, and only the best alignment result was retained; based on the filtered query sequence alignment results, the ratio of target pathogen-specific alignment and non-specific alignment at each site was calculated in combination with the frequency statistics of each target sequence; if the proportion of non-specific alignment was > 2%, it was considered to be a non-specific alignment region, and only the specific alignment region was retained for each query sequence in turn; the specific alignment region was then aligned with the NCBI NT library, and the ratio of specific alignment to non-specific alignment at each site of the query sequence was counted, and only the region with a non-specific alignment ratio of < 2% was retained, and finally the target pathogen-specific core genome sequence was obtained.

[0022] In some aspects, in step 4), the method for constructing the reference genome of all species in the same family as the target pathogen comprises the following steps:

[0023] 1) Obtain the complete taxonomy information of each strain genome according to the taxonomy ID recorded in the assembly_summary_refseq.txt and assembly_summary_genbank.txt files of the NCBI genome database;

[0024] 2) Find the family name of the target pathogen, and then find all genome information of the same family as the target pathogen; for each species, only one strain genome is selected and retained, and the reference genome is preferentially selected;

[0025] 3) Download and select all representative strain genome sequences at the species level of the same family as the target pathogen, and use makeblastdb to build a library; based on the assembly_summary_genbank.txt file, count the number of strain genomes at different species levels as various occurrence frequency information.

[0026] The present application also provides a screening model for target pathogen-specific core genome sequences, including the following modules:

[0027] Module 1: used to download the target pathogen genome sequence;

[0028] Module 2: used to compare with the reference genome, statistically calculate the genome coverage depth, genome coverage number and variation information;

[0029] Module 3: used to select the core genome sequence of the target pathogen;

[0030] Module 4: Used to obtain the target pathogen-specific core genome sequence: The target pathogen core genome sequence is compared, annotated and filtered with the human reference genome, the constructed reference genomes of all species in the same family as the target pathogen, and the NCBI NT database to obtain the target pathogen-specific core genome sequence.

[0031] Furthermore, each module 1)-4) is specifically used to implement the specific details corresponding to the aforementioned steps 1)-4).

[0032] The present application also provides an electronic device, comprising: a processor and a memory; the processor and the memory are connected, wherein the memory is used to store a computer program, and the processor is used to call the computer program to execute any of the methods described above.

[0033] The present application also provides a computer storage medium, wherein the computer storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, any of the methods described above is executed.

[0034] Beneficial technical effects of this application:

[0035] 1. When screening the core genomic regions of the target pathogen, the present application does not require gene prediction and the corresponding core-pan gene analysis process, but directly obtains the core region sequence (including coding and non-coding regions) at the whole genome level based on the comparison and calculation of each genomic sequence with the reference genome; at the same time, the calculation process is efficient, simple and time-saving.

[0036] 2. When screening the core regions of the genome, this application also calculated the number of strain genomes and the depth of genome coverage covered by each region of the reference genome, and mainly evaluated and selected based on the number of strain genomes covered. In particular, for the search of core genome sequences of pathogenic complexes, the number of genome coverages within the complex must meet the threshold requirements to ensure high target coverage of complex-specific sequences; when targeting super-multiplex PCR scenarios, this application further compares the differences in coverage depth and number of covered genomes and considers mutation information, thereby screening out specific regional sequences with high coverage, high copy number and low variation.

[0037] 3. When filtering non-specific core genome sequences, in addition to using the NCBI NT library to filter non-specific sequences, this application also adds a filtering method of comparing with the reference genome sequences of all other pathogenic species in the same family as the target pathogen, so as to further ensure the high specificity of the core genome sequences finally screened.

[0038] 4. The screening method of the present application can take into account the interference problem between primers of each target sequence during the design process of multiplex PCR primer probes. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 , technical flow chart of this application.

[0040] Figure 2 , Klebsiella pneumoniae reference genome coverage depth, number of covered genomes, and variation distribution.

[0041] Figure 3 , genome coverage and variation distribution of the screened Klebsiella pneumoniae-specific core sequences. DETAILED DESCRIPTION

[0042] The embodiments of the present application will be described in detail below in conjunction with the examples, but it will be appreciated by those skilled in the art that the following examples are only used to illustrate the present application and should not be considered as limiting the scope of the present application. In the examples, if no specific conditions are specified, the conditions are carried out according to normal conditions or manufacturer recommendations. If the manufacturer is not specified for the reagents or instruments used, they are all conventional products that can be purchased from the market.

[0043] Definition of some terms

[0044] Unless otherwise defined below, the meaning of all technical terms and scientific terms used in the specific embodiments of the present application is intended to be the same as those generally understood by those skilled in the art. Although it is believed that the following terms are well understood by those skilled in the art, the following definitions are still set forth to better explain the present application.

[0045] As used in this application, an indefinite or definite article used when referring to a singular noun eg "a" or "an", "the" includes a plural of that noun.

[0046] As used in this application, the terms "comprises", "comprising", "having", "containing" or "involving" are inclusive or open-ended and do not exclude other unrecited elements or method steps. The term "consisting of" is considered a preferred embodiment of the term "comprising". If a group is defined below as comprising at least a certain number of embodiments, this should also be understood to disclose a group that preferably consists of only these embodiments.

[0047] The term "approximately" in this application indicates an accuracy range that can be understood by those skilled in the art and still ensures the technical effect of the characteristic in question. The term generally indicates ±10%, preferably ±5%, of the indicated value.

[0048] In addition, the terms first, second, third, (a), (b), (c), and the like in the specification and claims are used to distinguish similar elements and are not necessarily required to describe a sequential or chronological order. It should be understood that the terms so used are interchangeable under appropriate circumstances, and the embodiments described in this application can be implemented in other sequences than those described or illustrated in this application.

[0049] The above terms or definitions are provided only to help understand the present application. These definitions should not be construed as having a scope less than that understood by those skilled in the art.

[0050] The method for screening target pathogen-specific core genome sequences of the present application comprises the following steps:

[0051] 1) Obtain the target pathogen genome sequence;

[0052] 2) Align with the reference genome sequence and statistically calculate the genome coverage depth, genome coverage number and / or variation information;

[0053] 3) Select the target pathogen core genome sequence;

[0054] 4) Obtaining the target pathogen-specific core genome sequence: The target pathogen core genome sequence is compared, annotated, and filtered with the human reference genome, the constructed reference genomes of all species in the same family as the target pathogen, and the NCBI NT database to obtain the target pathogen-specific core genome sequence.

[0055] The specific core genome sequence described herein refers to a sequence fragment that is carried by the genome of each strain in the target pathogen and is unique to the target pathogen relative to other bacteria;

[0056] In some embodiments, in step 1), the downloading is downloading the target pathogen genome sequence from the NCBI genome database;

[0057] In some preferred embodiments, the step 1) specifically includes: for the target pathogen, downloading assembly_summary_genbank.txt and assembly_summary_refseq.txt files from the NCBI genome database; more preferably, if the number of target pathogen genomes recorded in assembly_summary_refseq.txt is greater than or equal to 1000, then downloading the corresponding genome sequence according to the assembly_summary_refseq.txt file, otherwise downloading the corresponding genome sequence according to the assembly_summary_genbank.txt file.

[0058] In some embodiments, the specific steps of step 2) are: selecting the reference genome of the target pathogen in the NCBI genome database as the reference genome sequence, aligning all downloaded genomes with the reference genome sequence, and outputting the SNP and InDel information of the reference sequence in each alignment; filtering each alignment result of the query sequence, and retaining only the best alignment as the final result; statistically calculating the number of genome coverage, genome coverage depth and / or SNP / InDel variation information and its frequency at each site of the reference genome sequence;

[0059] In some preferred embodiments, if the target pathogen is at the complex group level or genus level, the genome coverage number, genome coverage depth and / or SNP / InDel variation information and its frequency of each species in the complex group or genus are statistically calculated simultaneously.

[0060] In some embodiments, in step 3), the selection is: based on the genome coverage number information at each site of the reference genome sequence obtained in step 2, calculate the ratio of the genome coverage number at each site of the reference genome to the total number of strain genomes, filter sites with a ratio lower than 95%, and then connect the bases of adjacent sites to obtain the core genome sequence of the target pathogen;

[0061] In some preferred embodiments, if the target pathogen is at the complex group level or the genus level, the sites where the ratio of the number of genomes covered by each species within the complex group or genus to the total number of genomes of the strain is less than 95% are filtered;

[0062] In some more preferred embodiments, the core genome sequence is no less than 70 bp in length.

[0063] In some aspects, the step 4) further includes: further screening out sequences with a genome coverage depth that is more than 1.5 times the genome coverage number as multi-copy core genome sequences.

[0064] In some embodiments, in step 4), the comparison, annotation and filtering specifically include:

[0065] The target pathogen core genome sequence was aligned with the human reference genome. For hits that were aligned with the human reference genome and had a homology of >= 90%, the corresponding query core sequence was first filtered out; then it was aligned with the constructed reference genome sequences of all species in the same family as the target pathogen, and different target sequences of each query sequence were aligned, and only the best alignment result was retained; based on the filtered query sequence alignment results, the ratio of target pathogen-specific alignment and non-specific alignment at each site was calculated in combination with the frequency statistics of each target sequence; if the proportion of non-specific alignment was > 2%, it was considered to be a non-specific alignment region, and only the specific alignment region was retained for each query sequence in turn; the specific alignment region was then aligned with the NCBI NT library, and the ratio of specific alignment to non-specific alignment at each site of the query sequence was counted, and only the region with a non-specific alignment ratio of < 2% was retained, and finally the target pathogen-specific core genome sequence was obtained.

[0066] In some embodiments, in step 4), the method for constructing the reference genome of all species of the same family as the target pathogen comprises the following steps:

[0067] 1) Obtain the complete taxonomy information of each strain genome according to the taxonomy ID recorded in the assembly_summary_refseq.txt and assembly_summary_genbank.txt files of the NCBI genome database;

[0068] 2) Find the family name of the target pathogen, and then find all genome information of the same family as the target pathogen; for each species, only one strain genome is selected and retained, and the reference genome is preferentially selected;

[0069] 3) Download and select all representative strain genome sequences at the species level of the same family as the target pathogen, and use makeblastdb to build a library; based on the assembly_summary_genbank.txt file, count the number of strain genomes at different species levels as various occurrence frequency information.

[0070] The screening model of the target pathogen-specific core genome sequence of the present application includes the following modules:

[0071] Module 1: used to download the target pathogen genome sequence;

[0072] Module 2: used to compare with the reference genome, statistically calculate the genome coverage depth, genome coverage number and variation information;

[0073] Module 3: used to select the core genome sequence of the target pathogen;

[0074] Module 4: Used to obtain the target pathogen-specific core genome sequence: The target pathogen core genome sequence is compared, annotated and filtered with the human reference genome, the constructed reference genomes of all species in the same family as the target pathogen, and the NCBI NT database to obtain the target pathogen-specific core genome sequence.

[0075] It can be understood that each module 1)-4) described in the present application is specifically used to implement the specific details corresponding to the aforementioned steps 1)-4).

[0076] The electronic device of the present application includes: a processor and a memory; the processor and the memory are connected, wherein the memory is used to store a computer program, and the processor is used to call the computer program to execute any of the methods described above.

[0077] The computer storage medium of the present application stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, any of the methods described above is executed.

[0078] The present application is described below in conjunction with specific embodiments.

[0079] Example 1 Screening of Klebsiella pneumoniae-specific core genome sequences

[0080] This example uses Klebsiella pneumoniae as the target pathogen for interpretation. The specific screening method is based on the aforementioned screening steps of the target pathogen-specific core genome sequence (see Figure 1 The specific contents are as follows.

[0081] Step 1. Search and download the Klebsiella pneumoniae genome sequence from the NCBI genome database. Download the assembly_summary_refseq.txt file from the NCBI genome database (download path: https: / / ftp.ncbi.nlm.nih.gov / genomes / ASSEMBLY_REPORTS / assembly_summary_refseq.tx t), select the record entry of Klebsiella pneumoniae from it, and complete the download of the genome of all strains of Klebsiella pneumoniae according to the download path in the record entry.

[0082] Step 2, based on the direct comparison of genome sequence, the coverage depth, the number of covered genomes, the variation and its frequency information of different sites of Klebsiella pneumoniae reference genome are statistically calculated. With the Klebsiella pneumoniae reference genome (GCF_000240185.1) in the NCBI genome database as a reference, all Klebsiella pneumoniae strain genomes downloaded are compared with it (ncbi-blast-2.9.0+ / blastn-task blastn-outfmt 0-evalue 1e-5-num_alignments10000), the output format is m0, and then the m0 format file is parsed into m6 format, and SNP and InDel variation information are output simultaneously. Then the comparison result is filtered, and for the comparison area of ​​each query (query sequence) and subject (target sequence), the best comparison is selected as the final comparison result. Finally, the comparison results of all strains are counted, and the coverage depth and the number of covered genomes (including the number of covered genomes and the proportion of each species or subspecies) at each site on the reference genome are calculated. At the same time, considering the subsequent design of primers and probe sequences, it is also necessary to count the variations and their frequencies at different sites in the genome (such as Figure 2 and Table 1).

[0083] Table 1. Coverage depth and number of genomes covered at each site of the Klebsiella pneumoniae reference genome

[0084]

[0085]

[0086] Step 3, obtain the core genome region sequence of Klebsiella pneumoniae. First, calculate the ratio of the number of genomes covered on each site of the reference genome to the total number of genomes of the strain and the ratio of the coverage depth to the total number of genomes of the strain, select and filter the core sites according to the proportion of the number of genomes covered, and retain the sites with a higher proportion (the default threshold is 95%) as the core sites. Then the bases of the core sites in the continuous positions are connected to obtain the core genome sequence fragments. Further, the coverage depth is significantly higher than the genome number fragments covered (the former is more than 1.5 times of the latter), which is the core genome sequence of multiple copies. Finally, the fragments with sequence length less than 70bp are filtered out to obtain the final core genome sequence.

[0087] Step 4: Select the specific core genome sequence of Klebsiella pneumoniae.

[0088] Step 4.1 builds a reference sequence database. Similarly, according to the assembly_summary_refseq.txt file information, the human reference genome sequence (GRCh38) and all species reference genome sequences of the same family (Enterobacteriaceae) as Klebsiella pneumoniae are downloaded from the NCBI genome database, and a localized database (ncbi-blast-2.9.0+ / makeblastdb-dbtype nucl-parse_seqids) is constructed. Specifically, when constructing the reference genome sequences of all species of the same family (Enterobacteriaceae) as Klebsiella pneumoniae, first, according to the taxonomy ID recorded in the assembly_summary_refseq.txt and assembly_summary_genbank.txt files of the NCBI genome database, the kingdom, phylum, class, order, family, genus, species and complete classification information of each strain genome are obtained using TaxonKit software (taxonkit lineage). Then, the family level name to which the target pathogen belongs is found, and all genome information of the same family as the target pathogen is found accordingly. For each species, only one strain genome is selected and retained, and the reference genome is preferentially selected. Finally, the genome sequences of all representative strains of the same family as the target pathogen were downloaded and the library was built using makeblastdb (makeblastdb-dbtype nucl-parse_seqids).

[0089] The assembly_summary_genbank.txt file (https: / / ftp.ncbi.nlm.nih.gov / genomes / ASSEMBLY_REPORTS / assembly_summary_genbank.txt) was used to count the number of strain genomes at different species levels as various frequency information. The NCBI NT database was directly downloaded (ascp-i asperaweb_id_dsa.openssh-l 100M-k 1-T anonftp@ftp-trace.ncbi.nlm.nih.gov: / blast / db / nt.*.tar.gz), and it can be used as a localized library after decompression.

[0090] Step 4.2 Select Klebsiella pneumoniae-specific sequences based on sequence alignment annotation. a) First, align the Klebsiella pneumoniae core genome sequence obtained in step 3 with the human reference genome (GRCh38) sequence (ncbi-blast-2.9.0+ / blastn-task blastn-outfmt 6-evalue 1e-5-num_alignments 10000), retain the alignment results with identity above 90%, and remove the corresponding core genome region alignment areas. b) Then, the core genome sequence after removing the human source is aligned with the reference genomes of all species in the same family as Klebsiella pneumoniae (ncbi-blast-2.9.0+ / blastn-task blastn-outfmt 6-evalue 1e-5-num_alignments10000, identity threshold set to 90%) and annotated, and for each query sequence, only retain the best alignment results for different subject alignments. Then, the ratio of Klebsiella pneumoniae-specific alignment and non-specific alignment at each site on the query sequence is calculated based on the frequency statistics of the species corresponding to each subject. If the proportion of non-specific alignment is high (>2%), it is considered to be a non-specific alignment region, and only the region with specific alignment is retained for each query sequence in turn. c) In the next step, the specific sequence obtained above is compared with the NCBI NT library. Similarly, the ratio of Klebsiella pneumoniae-specific alignment and non-specific alignment at each site of the query sequence is counted, and only the region with a low proportion of non-specific alignment (<2%) is retained. That is, the specific core genome sequence of Klebsiella pneumoniae can be finally obtained (see Tables 2 and Figure 3 ).

[0091] Table 2. Evaluation and selection of Klebsiella pneumoniae nuclear genome sequence specificity

[0092]

[0093]

[0094] Example 2 Screening of specific core genome sequences of Enterobacter cloacae complex

[0095] This example uses the Enterobacter cloacae complex as the target pathogen for interpretation, and adopts the screening steps of the target pathogen-specific core genome sequence described above in this application, which are as follows.

[0096] Step 1. Search and download the genome sequence of the Enterobacter cloacae complex from the NCBI genome database. Download the assembly_summary_refseq.txt file from the NCBI genome database (download path: https: / / ftp.ncbi.nlm.nih.gov / genomes / ASSEMBLY_REPORTS / assembly_summary_refseq.tx t), and screen out Enterobacter hormaechei, Enterobacter cloacae, (Enterobacter asburiae), Enterobacter kobei, Enterobacter ludwigii, Enterobacter roggenkampii, Enterobacter sichuanensis, Enterobacter cancerogenus, Enterobacter chengduensis, and Enterobacter sichuanensis. chuandaensis) and other strains belonging to the Enterobacter cloacae complex, and complete the download of the genome of all strains of the Enterobacter cloacae complex according to the download path in the record entry.

[0097] Step 2, based on direct alignment of genome sequence, statistically calculate the coverage depth, number of covered genomes, variation and frequency information of different sites of reference genome of Enterobacter cloacae complex group. With the reference genome of Enterobacter cloacae in NCBI genome database (GCF_023702375.1) as reference, all the genomes of Enterobacter cloacae strains downloaded are aligned with it (ncbi-blast-2.9.0+ / blastn-task blastn-outfmt 0-evalue 1e-5-num_alignments10000), the output format is m0, then parse the m0 format file into m6 format, and output SNP and InDel variation information at the same time. Then filter the alignment result, for each alignment area of ​​query and subject, select the best alignment as the final alignment result. Finally, the alignment results of all strains are counted, and the coverage depth, number of covered genomes (including the number of covered genomes and proportion of each species or subspecies) and variation and frequency (as shown in Table 3) at each site on the reference genome are calculated.

[0098] Table 3. Coverage depth and number of genomes covered at each locus of the reference genome of Enterobacter cloacae complex

[0099]

[0100]

[0101]

[0102]

[0103]

[0104]

[0105]

[0106]

[0107] Step 3: Obtain the core genome region sequence of the Enterobacter cloacae complex.

[0108] First, calculate the ratio of the number of genomes covered at each site of the reference genome to the total number of genomes of the strain and the ratio of the coverage depth to the total number of genomes of the strain, select and filter the core sites based on the proportion of the number of genomes covered, and retain the sites with a higher proportion (the default threshold is 95%) as core sites. Here, it is necessary to require that the number of genomes covered by each species in the Enterobacter cloacae complex (Enterobacter hallii, Enterobacter cloacae, Enterobacter kobe, Enterobacter lewyii, etc.) has a high proportion (>=95%), otherwise it is not considered to be a core site and is filtered out. Then connect the bases of the core sites at consecutive positions to obtain the core genome sequence fragments. Finally, filter out the fragments with a sequence length of less than 70bp to obtain the final core genome sequence.

[0109] Step 4: Select the specific core genome sequence of the Enterobacter cloacae complex.

[0110] Step 4.1 builds a reference sequence database. Similarly, according to the assembly_summary_refseq.txt file information, the human reference genome sequence (GRCh38) and all species reference genome sequences of the same family (Enterobacteriaceae) as Klebsiella pneumoniae are downloaded from the NCBI genome database, and a localized database (ncbi-blast-2.9.0+ / makeblastdb-dbtype nucl-parse_seqids) is constructed. Specifically, when constructing the reference genome sequences of all species of the same family (Enterobacteriaceae) as Klebsiella pneumoniae, first, according to the taxonomy ID recorded in the assembly_summary_refseq.txt and assembly_summary_genbank.txt files of the NCBI genome database, the kingdom, phylum, class, order, family, genus, species and complete classification information of each strain genome are obtained using TaxonKit software (taxonkit lineage). Then, the family level name to which the target pathogen belongs is found, and all genome information of the same family as the target pathogen is found accordingly. For each species, only one strain genome is selected and retained, and the reference genome is preferentially selected. Finally, download the genome sequences of all representative strains of the same family as the target pathogen, and use makeblastdb to build a library (makeblastdb-dbtype nucl-parse_seqids). At the same time, based on the assembly_summary_genbank.txt file (https: / / ftp.ncbi.nlm.nih.gov / genomes / ASSEMBLY_REPORTS / assembly_summary_genbank.txt), count the number of strain genomes at different species levels as various occurrence frequency information. Download the NCBI NT database directly (ascp-i asperaweb_id_dsa.openssh-l 100M-k 1-T anonftp@ftp-trace.ncbi.nlm.nih.gov: / blast / db / nt.*.tar.gz), and use it as a localized library after decompression.

[0111] Step 4.2: Based on the sequence alignment annotation, select the E. cloacae complex-specific sequences.

[0112] a) First, the core genome sequence of the Enterobacter cloacae complex obtained in step 3 was aligned with the human reference genome (GRCh38) sequence (ncbi-blast-2.9.0+ / blastn-task blastn-outfmt 6-evalue 1e-5-num_alignments 10000), and the alignment results with identity above 90% were retained, and the corresponding core genome region alignment regions were removed. b) Then, the core genome sequence after removing the human source was aligned with the constructed reference genomes of all species of the Enterobacteriaceae (ncbi-blast-2.9.0+ / blastn-task blastn-outfmt 6-evalue1e-5-num_alignments 10000, and the identity threshold was set to 90%) and annotated, and only the best alignment results were retained for the different subjects of each query sequence. Then, the ratio of specific alignment and non-specific alignment of the Enterobacter cloacae complex at each site on the query sequence was calculated in combination with the frequency statistics of the reference species corresponding to each subject. If the proportion of non-specific alignment is high (>2%), it is considered to be a non-specific alignment region, and only the specific alignment region is retained for each query sequence. c) In the next step, the specific sequence obtained above is compared with the NCBI NT library. Similarly, the ratio of Klebsiella pneumoniae specific alignment to non-specific alignment at each site of the query sequence is counted, and only the region with a low proportion of non-specific alignment (<2%) is retained. That is, the specific core genome sequence of the Enterobacter cloacae complex can be obtained in the end.

[0113] Table 2-2 Evaluation and selection of core genome sequence specificity of Enterobacter cloacae complex

[0114]

[0115]

[0116] In summary, the method of the present application can effectively screen out specific core sequences with high coverage and high copy number, whether for target pathogens at the species level or at the complex group or genus level; and compared with previous methods, it has the advantages of being more efficient and simpler.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for screening target pathogen-specific core genome sequences, characterized in that: The steps include: 1) Obtain the target pathogen genome sequence; 2) Align with the reference genome sequence and statistically calculate the genome coverage depth, genome coverage number and / or variation information; 3) Select the target pathogen core genome sequence; 4) Obtaining the target pathogen-specific core genome sequence: The target pathogen core genome sequence is sequentially compared, annotated, and filtered with the human reference genome, the constructed reference genomes of all species in the same family as the target pathogen, and the NCBI NT database to obtain the target pathogen-specific core genome sequence; In the step 1), the acquisition is to download the target pathogen genome sequence from the NCBI genome database; In the step 2), the comparison with the reference genome sequence and the statistical calculation of the genome coverage depth, the number of genome coverage and / or the variation information are as follows: the reference genome of the target pathogen in the NCBI genome database is selected as the reference genome sequence, all the downloaded genomes are compared with the reference genome sequence, and the SNP and InDel information of each reference genome sequence in the comparison is output; Each alignment result of the query sequence is filtered, and only the best alignment is retained as the final result; The number of genome coverage, genome coverage depth and / or SNP / InDel variation information and its frequency at each site of the reference genome sequence are statistically calculated; if the target pathogen is at the complex group level or genus level, the number of genome coverage, genome coverage depth and / or SNP / InDel variation information and its frequency of each species in the complex group or genus are statistically calculated at the same time; In the step 3), the selection is as follows: based on the genome coverage number information at each site of the reference genome sequence obtained in step 2), the ratio of the genome coverage number at each site of the reference genome to the total number of strain genomes is calculated, sites with a ratio lower than 95% are filtered out, and then the bases of adjacent sites are connected to obtain the core genome sequence of the target pathogen.

2. The screening method according to claim 1, characterized in that: In the step 1), the acquisition is specifically as follows: for the target pathogen, downloading assembly_summary_genbank.txt and assembly_summary_refseq.txt files from the NCBI genome database; if the number of target pathogen genomes recorded in assembly_summary_refseq.txt is greater than or equal to 1000, then downloading the corresponding genome sequence according to the assembly_summary_refseq.txt file, otherwise downloading the corresponding genome sequence according to the assembly_summary_genbank.txt file.

3. The screening method according to any one of claims 1-2, characterized in that: In the step 3), if the target pathogen is at the complex group level or genus level, the sites where the ratio of the number of genomes covered by various species within the complex group or genus to the total number of strain genomes is less than 95% are filtered; the core genome sequence length is not less than 70bp.

4. The screening method according to claim 3, characterized in that: The step 4) also includes: further screening out sequences with a genome coverage depth that is more than 1.5 times the genome coverage number as multi-copy core genome sequences.

5. The screening method according to claim 4, characterized in that: In the step 4), the comparison, annotation and filtering specifically include: performing sequence comparison on the core genome sequence of the target pathogen and the human reference genome, and filtering out the corresponding query core sequence for hits that are compared with the human reference genome and have a homology of ≥ 90%; then comparing with the constructed reference genome sequences of all species of the same family as the target pathogen, and comparing different target sequences of each query sequence, and retaining only the best comparison result; based on the comparison results of the filtered query sequences, the ratio of the target pathogen-specific comparison and the non-specific comparison of each site is calculated in combination with the frequency statistics of each target sequence; if the proportion of the non-specific comparison is > 2%, it is considered to be a non-specific comparison area, and only the specific comparison area is retained for each query sequence in turn; the specific comparison area is then compared with the NCBI NT library, and the ratio of the specific comparison and the non-specific comparison of each site of the query sequence is counted, and only the area with the non-specific comparison ratio < 2% is retained, and finally the target pathogen-specific core genome sequence is obtained.

6. The screening method according to claim 5, characterized in that: In step 4), the method for constructing the reference genome of all species of the same family as the target pathogen comprises the following steps: 1) Obtain the complete taxonomy information of each strain genome according to the taxonomy ID recorded in the assembly_summary_refseq.txt and assembly_summary_genbank.txt files of the NCBI genome database; 2) Find the family name of the target pathogen, and then find all genome information of the same family as the target pathogen; for each species, only one strain genome is selected and retained, and the reference genome is preferentially selected; 3) Download and select all representative strain genome sequences at the species level of the same family as the target pathogen, and use makeblastdb to build a library; based on the assembly_summary_genbank.txt file, count the number of strain genomes at different species levels as various occurrence frequency information.

7. An electronic device, characterized in that: include: Processor and memory; The processor is connected to a memory, wherein the memory is used to store a computer program, and the processor is used to call the computer program to execute the method according to any one of claims 1 to 6.

8. A computer storage medium, characterized in that The computer storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the method according to any one of claims 1 to 6 is executed.

Citation Information

Patent Citations

  • Pathogen specific nucleic acid gene of staphylococcus epidermidis and detection method

    CN115852004A