Construction method, application, device and electronic equipment of full-length ribosomal RNA database

By constructing a full-length ribosomal RNA database, the problem that microbial species classification in the prior art is difficult to reach the species level or strain level, and higher species classification resolution and more accurate clinical pathogen detection are achieved.

CN115148293BActive Publication Date: 2025-06-13ZHEJIANG DIGENA DIAGNOSTIC TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210617403.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-01
Publication Date
2025-06-13
Estimated Expiration
2042-06-01

AI Technical Summary

Technical Problem

Existing ribosomal RNA databases are difficult to achieve species-level or strain-level resolution in microbial species classification, limiting the accuracy of clinical pathogen detection.

Method used

By constructing a full-length ribosomal RNA database, the coordinate information, species classification information and species names of all microorganisms are obtained, and ribosomal RNA genes are screened and merged based on preset conditions to generate a fasta sequence format file, and finally the sequence information of all ribosomal RNA genes is merged to construct a full-length ribosomal RNA database.

Benefits of technology

It improves the resolution of microbial species classification, can achieve more accurately classification at the species level or strain level, and enhances the accuracy of clinical pathogen detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115148293B_ABST
    Figure CN115148293B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, application, device and electronic device for constructing a full-length ribosomal RNA database. The method includes: obtaining coordinate information, species classification information and species names of ribosomal RNA genes; screening the genomes where ribosomal RNAs are located according to species information; arranging the screened ribosomal RNA genes according to the sizes of the starting positions of the ribosomal RNAs. If two adjacent ribosomal RNA genes are respectively the small subunit and the large subunit of the same ribosome, and the distance between the two adjacent ribosomal RNA genes is less than a threshold, then merge the coordinate information of the two adjacent ribosomal RNA genes to obtain the coordinate information of the full-length ribosomal RNA; obtain sequence information (fasta format file) according to the coordinate information of the full-length ribosomal RNA; merge the fasta sequence format files of all ribosomal RNA genes to obtain a full-length ribosomal RNA database. Further classifying species according to the full-length ribosomal RNA database can improve the resolution of species classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics, and in particular to a method for constructing a full-length ribosomal RNA database, an application, a device and an electronic device. Background Art

[0002] Microorganisms are the most numerous organisms on Earth. Microorganisms include a large group of organisms such as bacteria, viruses, fungi, and some small protists, microalgae, etc. They are tiny in size and closely related to humans. Complex microbial communities form dynamic and diverse natural environments, with a distribution range covering from the mammalian gut to the soil. Understanding the diversity of the microbiota in the environment, that is, the composition of the microbial community, can help us understand their relevance to humans and the environment. For example, it can measure the impact of dietary or drug changes on the gut microbiota. Isolation and culture is a traditional technique for studying the types and composition of environmental microorganisms. However, most microorganisms cannot be isolated and cultured under laboratory conditions, which brings difficulties to microbial research.

[0003] With the emergence of next-generation high-throughput sequencing, sequencing of the 16S ribosomal RNA (rRNA), 23S ribosomal RNA (rRNA) gene, and internal transcribed spacer (ITS) region has become a cheap, routine, and direct method for diversity studies of complex samples (such as feces or soil) without the need for culturing. For example, the 16S rRNA gene contains 9 hypervariable regions (V1-V9). The short hypervariable region V1 of the 16S rRNA gene can provide species classification resolution at the family level and can distinguish common Staphylococcus and Streptococcus pathogens. However, the resolution of the existing ribosomal RNA databases using variable regions is still insufficient. In many cases, the resolution can only reach the genus level at most, and it is very difficult to reach the species level. Clinical pathogen detection often requires species resolution to reach the species or even strain level. Summary of the Invention

[0004] The present invention provides a method for constructing a full-length ribosomal RNA database, an application, a device and an electronic device to at least solve the above technical problems existing in the prior art.

[0005] On the one hand, the present invention provides a method for constructing a full-length ribosomal RNA database, the method comprising:

[0006] Obtaining coordinate information, species classification information, and species names of ribosomal RNA genes of all microorganisms, wherein the coordinate information includes the genome number where the ribosomal RNA gene is located, the start position of the ribosomal RNA, and the end position of the ribosomal RNA;

[0007] Screen the ribosomal RNA genes of all microorganisms based on preset conditions and the species classification information;

[0008] Arrange the screened ribosomal RNA genes according to the sizes of the ribosomal RNA start positions. If two adjacent ribosomal RNA genes are the small subunit and the large subunit of the same ribosome respectively, and the distance between the two adjacent ribosomal RNA genes is less than the threshold, then merge the coordinate information of the two adjacent ribosomal RNA genes to obtain the coordinate information of the full-length ribosomal RNA; if the merging condition is not met, then retain the coordinate information of the original ribosomal RNA gene;

[0009] According to the coordinate information of the full-length ribosomal RNA after screening and merging and the coordinate information of the ribosomal RNA genes that do not meet the merging condition, obtain the sequence information of the ribosomal RNA genes corresponding to the coordinate information, and generate a fasta sequence format file for the sequence information;

[0010] Merge the fasta sequence format files of all ribosomal RNA genes to obtain a full-length ribosomal RNA database.

[0011] In an implementable embodiment, the screening of the genome where the ribosomal RNA gene is located based on the preset conditions and the species classification information includes:

[0012] Determine the species classification level according to the species classification information;

[0013] If the species classification level has species-level information, then the species classification information meets the preset conditions, and retain the genome where the ribosomal RNA gene corresponding to the species classification information is located;

[0014] If the species classification level does not have species-level information, then the species classification information does not meet the preset conditions, and delete the genome where the ribosomal RNA gene corresponding to the species classification information is located.

[0015] In an implementable embodiment, the determination of the coordinate information of the full-length ribosomal RNA includes:

[0016] Determine the start position of the smallest ribosomal RNA in the small subunit and the large subunit as the start position of the full-length ribosomal RNA;

[0017] Determine the termination position of the largest ribosomal RNA in the small subunit and the large subunit as the termination position of the full-length ribosomal RNA.

[0018] In one implementable manner, the sequence information includes one or more of a sequence position name, a sequence length, a brief description of the sequence, a unique sequence number, sequence keywords, a species name, genomic region information, a gene name, and sequence base information.

[0019] In one implementable manner, if the species name in the sequence information is inconsistent with the obtained species name, then revise the obtained species name to the species name in the sequence information.

[0020] In one implementable manner, after obtaining the ribosomal RNA database, the method further includes:

[0021] Compare the sequencing result sequence of the sample with the sequence information of each ribosomal RNA gene in the full-length ribosomal RNA database to obtain comparison information;

[0022] Determine the species classification information of the sample according to the comparison information.

[0023] In one implementable manner, the comparison information includes one or more of the sequence length of the sample, the starting position of the sequence, the ending position of the sequence, the number of bases matched in the comparison, the comparison length, and the dynamic programming alignment score.

[0024] In one implementable manner, the determining the species classification information of the sample according to the comparison information includes:

[0025] Determine the comparison similarity between the sequencing result sequence of the sample and the sequence information of each ribosomal RNA gene in the full-length ribosomal RNA database according to the number of bases matched in the comparison and the comparison length, and filter the sequence information of ribosomal RNA genes with a comparison similarity less than the comparison similarity threshold according to the comparison similarity threshold;

[0026] Determine the comparison coverage of the sequencing result sequence of the sample and the sequence information of each ribosomal RNA gene in the filtered full-length ribosomal RNA database according to the sequence length, starting position, and ending position of the sample, and filter the sequence information of ribosomal RNA genes with a comparison coverage less than the comparison coverage threshold according to the comparison coverage threshold;

[0027] Determine at least one reference sequence from the filtered ribosomal RNA genes according to the dynamic programming alignment score and the dynamic programming alignment score threshold;

[0028] Calculate the classification information ratio of the reference sequences according to the classification hierarchy from low to high;

[0029] Determine the species classification information of the sample according to the classification information ratio and the reference sequences.

[0030] On the other hand, the present invention provides an apparatus for constructing a full-length ribosomal RNA database, which comprises:

[0031] A first acquisition module, configured to acquire the coordinate information, species classification information, and species name of ribosomal RNA genes of all microorganisms, wherein the coordinate information includes the genome number where the ribosomal RNA gene is located, the start position of the ribosomal RNA, and the end position of the ribosomal RNA;

[0032] A screening module, configured to screen the genomes where the ribosomal RNA genes of all microorganisms are located based on preset conditions and the species classification information;

[0033] A merging module, configured to arrange the screened ribosomal RNA genes in ascending order of the start position of the ribosomal RNA. If two adjacent ribosomal RNA genes are respectively the small subunit and the large subunit of the same ribosome, and the distance between the two adjacent ribosomal RNA genes is less than a threshold, then merge the coordinate information of the two adjacent ribosomal RNA genes to obtain the coordinate information of the full-length ribosomal RNA; if the merging condition is not satisfied, then retain the coordinate information of the original ribosomal RNA gene;

[0034] A second acquisition module, configured to acquire the sequence information of the ribosomal RNA gene corresponding to the coordinate information according to the coordinate information of the full-length ribosomal RNA after screening and merging and the coordinate information of the ribosomal RNA gene that does not satisfy the merging condition, and generate a fasta sequence format file for the sequence information;

[0035] A construction module, configured to merge the fasta sequence format files of all ribosomal RNA genes to obtain a full-length ribosomal RNA database.

[0036] On yet another aspect, the present invention provides a computer-readable storage medium storing a computer program for executing the method for constructing a full-length ribosomal RNA database according to the present invention.

[0037] On still another aspect, the present invention provides an electronic device, comprising:

[0038] A processor;

[0039] A memory for storing executable instructions of the processor;

[0040] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method for constructing a full-length ribosomal RNA database according to the present invention.

[0041] In the above solution of the present invention, by obtaining the coordinate information, species classification information, and species names of the ribosomal RNA genes of all microorganisms, in order to make the constructed full-length ribosomal RNA database sufficiently concise, preset conditions are set to screen the genomes where the ribosomal RNA genes are located, and further the coordinate information of the ribosomal RNA genes that meet the conditions is merged to obtain the coordinate information of the full-length ribosomal RNA. According to the coordinate information, the sequence information of the ribosomal RNA genes is obtained. Finally, after merging the sequence information of all ribosomal RNA genes, a full-length ribosomal RNA database is obtained. Since the longer the sequence of the ribosomal RNA gene, the more accurate the species classification information, therefore, using the full-length ribosomal RNA database of this solution for species classification can improve the resolution of species classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 FIG. shows a schematic flowchart of a method for constructing a full-length ribosomal RNA database provided by an embodiment of the present invention;

[0043] Figure 2 FIG. shows an example diagram of a GenBank format file;

[0044] Figure 3 FIG. shows an example diagram of a fasta format file;

[0045] Figure 4 FIG. shows a schematic flowchart of an application method of a full-length ribosomal RNA database provided by an embodiment of the present invention;

[0046] Figure 5 FIG. shows a schematic diagram of a sample sequencing result sequence;

[0047] Figure 6 FIG. shows a schematic diagram of alignment information;

[0048] Figure 7 FIG. shows a schematic diagram of annotation information;

[0049] Figure 8 FIG. shows a schematic structural diagram of a device for constructing a full-length ribosomal RNA database provided by an embodiment of the present invention;

[0050] Figure 9 FIG. shows a schematic structural diagram of an application device of a full-length ribosomal RNA database provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0051] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0052] like Figure 1 A method for constructing a full-length ribosomal RNA database provided by one embodiment of the present invention is shown, and the method comprises:

[0053] Step S101, obtaining coordinate information, species classification information and species name of ribosomal RNA genes of all microorganisms, wherein the coordinate information includes the genome number where the ribosomal RNA gene is located, the ribosomal RNA starting position and the ribosomal RNA ending position.

[0054] The coordinate information, species classification information and species name of the ribosomal RNA genes of all microorganisms of different sources and types are obtained from the bioinformatics database. The ribosomal RNA genes include small subunit ribosomal RNA sequences (including 16S / 18S) and large subunit ribosomal RNA sequences (including 23S / 28S). The bioinformatics database is the SILVA database (SILVAribosomal RNA database), which provides a comprehensive data set of small subunit (abbreviated as SSU, including 16S / 18S) and large subunit (abbreviated as LSU, including 23S / 28S) ribosomal RNA (rRNA) sequence data in three life domains (bacteria, archaea and eukaryotes). It is the most comprehensive database of ribosomal RNA genes currently collected. In addition to the SILVA database, other databases that collect ribosomal RNA gene data can also be selected, and the present invention does not limit this.

[0055] The coordinate information includes the genome number, the starting position of the ribosomal RNA, and the ending position of the ribosomal RNA. A genome region can be uniquely determined based on the coordinate information, and the base sequence of the region can be obtained through this coordinate information.

[0056] Species classification information is used to characterize the classification to which the species belongs. The species classification levels, from large to small, include domain, kingdom, phylum, class, order, family, genus, and species.

[0057] Download the coordinate information, species classification information, and species names of ribosomal RNA genes of all microorganisms from the SILVA database. The downloaded coordinate information, species classification information, and species names are stored in the taxmap file. For example, the file name of one ribosomal RNA gene is: taxmap_ncbi_ssu_ref_nr99_138.1.txt.gz. According to "ssu" in the file name, it can be determined that this ribosomal RNA gene is the small subunit. The file format is shown in Table 1 below:

[0058] Table 1 is an example of the taxmap file

[0059]

[0060]

[0061] According to Table 1, it can be determined that the coordinate information of this ribosomal RNA gene is AB003380: 1-2573; the species classification information is: Root; Cellular organisms; Bacteria <Prokaryotes>; FCB group; Fibrobacteres <phylum>; Fibrobacter; Fibrobacterales; Fibrobacteraceae; Fibrobacter; Fibrobacter succinogenes; the species name is: Fibrobacter succinogenes.

[0062] Step S102: Screen the genomes where the ribosomal RNA genes of all microorganisms are located based on preset conditions and the species classification information.

[0063] The preset condition is whether the species level in the species classification information is clear. When the species level in the species classification information is not clear, "unclassified", "uncultured", or "environmental samples" often appear in the species classification information. If "unclassified", "uncultured", or "environmental samples" appear in the species classification information of a certain ribosomal RNA gene, then delete this ribosomal RNA gene.

[0064] Step S103: Arrange the screened ribosomal RNA genes according to the size of the starting position of the ribosomal RNA. If two adjacent ribosomal RNA genes are respectively the small subunit and the large subunit of the same ribosome, and the distance between the two adjacent ribosomal RNA genes is less than the threshold, then merge the coordinate information of the two adjacent ribosomal RNA genes to obtain the coordinate information of the full-length ribosomal RNA; if the merging condition is not met, then retain the coordinate information of the original ribosomal RNA gene;

[0065] Arrange the filtered ribosomal RNA genes in ascending order according to the size of the ribosomal RNA start position. If two adjacent ribosomal RNA genes are the small subunit and the large subunit of the same ribosome respectively, and the distance between the two adjacent ribosomal RNA genes is less than the threshold value. Assume the threshold value is 2000bp (base pair, the number of base pairs), then merge the coordinate information of the two adjacent ribosomal RNA genes to obtain the coordinate information of the full-length ribosomal RNA. Assume the genome number is AAHN02000002, the start position of the ribosomal RNA large subunit is 439396, and the end position is 442264, that is, the ribosomal RNA large subunit interval is AAHN02000002:439396 - 442264. The genome number is AAHN02000002, the start position of the ribosomal RNA small subunit is 442892, and the end position is 444402, that is, the ribosomal RNA small subunit interval is AAHN02000002:442892 - 444402. The distance between the two genomic intervals is 442892 - 442264 = 628bp. Since 628bp does not exceed the threshold value of 2000bp, it meets the merging condition.

[0066] If the merging condition is not met, then retain the coordinate information of the original ribosomal RNA gene, that is, the coordinate information of the large subunit or the small subunit of the original ribosomal RNA;

[0067] In one example, determining the coordinate information of the full-length ribosomal RNA includes:

[0068] Determine the start position of the full-length ribosomal RNA as the smallest ribosomal RNA start position among the small subunit and the large subunit;

[0069] Determine the end position of the full-length ribosomal RNA as the largest ribosomal RNA end position among the small subunit and the large subunit.

[0070] The coordinate information of the full-length ribosomal RNA after merging is AAHN02000002:439396 - 444402.

[0071] Step S104, according to the coordinate information of the full-length ribosomal RNA and the coordinate information of the original ribosomal RNA gene that does not meet the merging condition, obtain the sequence information of the ribosomal RNA gene corresponding to the coordinate information, and generate the sequence information into a fasta sequence format file.

[0072] Download the sequence information of ribosomal RNA genes from the biological information database according to the coordinate information, where the coordinate information includes the coordinate information of the full-length ribosomal RNA after merging and the coordinate information of the original ribosomal RNA genes that do not meet the merging conditions. The biological information database is the NCBI (National Center for Biotechnology Information) database. By calling the E-utilities interface of the NCBI database, download the GenBank format file (suffix:.gbk) of ribosomal RNA according to the coordinate information of ribosomal RNA genes. The GenBank format file contains detailed information about this ribosomal RNA sequence. The sequence information includes one or more of the sequence position name, sequence length, sequence brief description, sequence unique number, sequence keywords, species name, genome region information, gene name, and sequence base information.

[0073] For example Figure 2 As shown in the example diagram of the GenBank format file, the format description of the GenBank format file is shown in Table 2 below.

[0074] Table 2 is the format description of the GenBank format file

[0075] Label Meaning LOCUS Information such as sequence position name, length, etc. DEFINITION Brief description of the sequence ACCESSION Unique sequence number KEYWORDS Sequence keywords SOURCE Species information REFERENCE References FEATURES Genomic region information, gene name, etc. ORIGIN Sequence base information

[0076] Extract the sequence information from the GenBank format file to generate a fasta format file. The fasta format is a text-based format used to represent nucleic acid sequences or polypeptide sequences, where nucleic acids or amino acids are represented by single letters, and it is allowed to add a sequence name and annotation before the sequence, which is beneficial to improving the readability of the full-length ribosomal RNA database.

[0077] The fasta format file is as Figure 3 shown, and the information in this file is as follows:

[0078] (1) AAHN02000002.1:439396 - 444402 is the "genome number: start position - end position" of the full-length ribosomal RNA;

[0079] (2) Burkholderia mallei ATCC 10399ctg_1047277000579, whole genome is the genome name;

[0080] (3) type = ssu + lsu represents that the type of ribosomal RNA is the small subunit (ssu) plus the large subunit (lsu). The value of type may also be only ssu or lsu because the NCBI database may only contain ssu sequences or lsu sequences;

[0081] (4) organism_name = Burkholderia mallei ATCC 10399, and the species name is Burkholderia mallei ATCC 10399;

[0082] (5) taxid = 412021, and the ncbi species number is 412021.

[0083] Step S105: Merge the fasta sequence format files of all ribosomal RNA genes to obtain a full-length ribosomal RNA database.

[0084] In the above solution of the present invention, by obtaining the coordinate information, species classification information, and species name of ribosomal RNA genes of all microorganisms, in order to make the constructed full-length ribosomal RNA database sufficiently concise, preset conditions are set to screen the ribosomal RNA genes, and further the ribosomal RNA genes that meet the conditions are merged to obtain full-length ribosomal RNA. Finally, after merging the sequence information of all ribosomal RNA genes, a full-length ribosomal RNA database is obtained. Since the longer the sequence of ribosomal RNA genes, the more accurate the species classification information, using the full-length ribosomal RNA database of this solution for species classification can improve the resolution of species classification.

[0085] In one example, screening the genomes where the multiple ribosomal RNA genes are located based on the preset conditions and the species classification information includes:

[0086] Determine the species classification level according to the species classification information;

[0087] If the species classification level has species-level information, the species classification information meets the preset conditions, and retain the genome where the ribosomal RNA gene corresponding to the species classification information is located;

[0088] If the species classification level does not have species-level information, the species classification information does not meet the preset conditions, and delete the genome where the ribosomal RNA gene corresponding to the species classification information is located.

[0089] Parse the species classification information in the taxdump file obtained in step S101 according to the species classification system of the NCBI database to obtain the classification levels of the species. If the classification levels include "unclassified", "uncultured", or "environmental samples", then delete the genome where the ribosomal RNA gene is located. Further, if the classification information meets any of the following conditions, it is retained:

[0090] Classified as bacteria, that is, the domain level is Bacteria;

[0091] Classified as fungi, that is, the kingdom level is Fungi;

[0092] Classified as viruses, that is, the domain level is Viruses;

[0093] Is a multicellular animal (including all parasites), that is, the kingdom level is Metazoa.

[0094] In one example, if the species name in the sequence information is inconsistent with the obtained species name, then determine the obtained species name as the species name in the sequence information.

[0095] If the species name in the sequence information downloaded through the NCBI database in step S104 is inconsistent with the species name downloaded through the SILVA database in step S101, then take the species name downloaded through the SILVA database as the standard, and use the species name downloaded through the SILVA database as the species name of the sequence information. Further, parse the NCBI species number (taxid) from the species name according to the species classification system file of the NCBI database.

[0096] As Figure 4 Shown is a flowchart of an application method of a full-length ribosomal RNA database provided by an embodiment of the present invention. The method includes:

[0097] Step S201: Compare the sequencing result sequence of the sample with the sequence information of each ribosomal RNA gene in the full-length ribosomal RNA database to obtain comparison information;

[0098] Step S202: Determine the species classification information of the sample according to the comparison information.

[0099] The constructed full-length ribosomal RNA database is indexed by indexing software, such as by minimap2 software. After gene sequencing of a sample, the sequencing result sequence of the sample is obtained. The sample sequencing result sequence is in the same format as the full-length ribosomal RNA database, both being fasta format files. For example Figure 5 is a schematic diagram of the sequencing result sequence of a certain sample. In the fasta file, the line starting with ">" is the sequence name line, and the subsequent lines are the base sequence lines.

[0100] The sequencing result sequence of the sample is input into the full-length ribosomal RNA database for alignment to obtain alignment information. The alignment information includes one or more of the sequence length, sequence start position, sequence end position, number of aligned and matched bases, alignment length, and dynamic programming alignment score of the sample. It may also include the database sequence name, database sequence length, database sequence start position, and database sequence end position.

[0101] The alignment information file is in paf format, as shown in the appendix Figure 6 The following is a schematic diagram of the alignment information. The paf format description is shown in Table 3 below:

[0102] Table 3 is the paf format description

[0103]

[0104] In addition to the 12 columns shown in Table 3, the alignment information also includes Figure 5 the information with the label "AS" in the last column shown in. The "AS" data represents the dynamic programming alignment score, and this value can reflect the overall quality value of the alignment. The higher the alignment score, the better and more accurate the alignment quality.

[0105] In one example, determining the species classification information of the sample according to the alignment information includes:

[0106] Determine the alignment similarity between the sequencing result sequence of the sample and the sequence information of each ribosomal RNA gene in the full-length ribosomal RNA database according to the number of aligned and matched bases and the alignment length, and filter the sequence information of ribosomal RNA genes with an alignment similarity less than the alignment similarity threshold according to the alignment similarity threshold;

[0107] Determine the alignment coverage of the sequencing result sequence of the sample and the sequence information of each ribosomal RNA gene in the filtered full-length ribosomal RNA database according to the sequence length, sequence start position, and sequence end position of the sample, and filter the sequence information of ribosomal RNA genes with an alignment coverage less than the alignment coverage threshold according to the alignment coverage threshold;

[0108] Determine at least one reference sequence from the filtered full-length ribosomal RNA database according to the dynamic programming alignment score and the dynamic programming alignment score threshold;

[0109] Calculate the classification information ratio of the reference sequence according to the classification hierarchy from low to high;

[0110] Determine the species classification information of the sample according to the classification information ratio and the reference sequence.

[0111] identity = match_length / mapping_length

[0112] where identity is the alignment similarity;

[0113] match_length: the number of bases that match in the alignment, that is, the data in the 10th column of the paf format file;

[0114] mapping_length: the alignment length (including the number of gaps), that is, the data in the 11th column of the paf format file.

[0115] The alignment similarity threshold is set to 70%. If the alignment similarity between a certain ribosomal RNA gene sequence and the sequencing result sequence of the sample is less than 70%, then filter out this ribosomal RNA gene sequence that does not meet the similarity threshold requirement. The qualified ribosomal RNA gene sequences after filtering are used for subsequent species classification and identification.

[0116] query_cov = (query_end - query_start + 1) / query_length

[0117] where:

[0118] query_cov is the alignment coverage;

[0119] query_end is the sequence alignment termination position, that is, the data in the 4th column of the paf format file;

[0120] query_start is the sequence alignment start position, that is, the data in the 3rd column of the paf format file;

[0121] query_length is the sequence length, that is, the data in the 2nd column of the paf format file.

[0122] Assume that the alignment coverage threshold is 50%. If the alignment coverage between a certain ribosomal RNA gene sequence and the sequencing result sequence of the sample is less than 50%, then filter out this ribosomal RNA gene sequence that does not meet the coverage threshold requirement.

[0123] Determine at least one reference sequence from the filtered full-length ribosomal RNA database according to the highest dynamic programming alignment score and a threshold value. The filtered full-length ribosomal RNA genes are sorted in descending order of the dynamic programming alignment score to obtain the highest dynamic programming alignment score. Calculate the percentage difference between the dynamic programming alignment score of each reference sequence and the highest dynamic programming alignment score, and filter out the reference sequences with a percentage difference higher than a certain threshold value. Assume the threshold value is 10%. If the percentage difference between the dynamic programming alignment score of a reference sequence and the highest dynamic programming alignment score is greater than 10%, then filter out the ribosomal RNA gene sequence that does not meet the percentage difference threshold. The finally retained ribosomal RNA gene sequences are used for species classification of the sequencing result sequences of the samples.

[0124] Suppose a certain sample sequencing result sequence is aligned with the full-length ribosomal RNA database. After filtering, there are 6 reference sequences, and the coordinate information and species classification information of each reference sequence are shown in Table 4 below.

[0125] Table 4 is the coordinate information and species classification information of the reference sequences

[0126]

[0127]

[0128] Calculate the classification information ratio of the reference sequences according to the classification levels from low to high, that is, in the order of species, genus, family, order, class, phylum, kingdom, and domain, and calculate the classification information ratio. The classification information ratio is determined according to the voting algorithm, and the classification information ratio is the voting ratio determined according to the voting algorithm. If the voting ratio is greater than or equal to 50% at the species level, then there is no need to calculate the voting ratio of the classification information at the genus level; if the voting ratio is less than 50% at the species level, then continue to calculate the voting ratio of the genus level; if the voting ratio is less than 50% at the genus level, then continue to calculate the voting ratio of the family level until the voting ratio is greater than or equal to 50% at a certain classification level.

[0129] Suppose the classification information ratios of the above 6 reference sequences at the species level are shown in Table 5 below.

[0130] Table 5 is the classification information ratios of the reference sequences at the species level

[0131]

[0132]

[0133] According to the voting ratio in Table 5, it can be determined that there is no unique taxonomic information for the reference sequence at the species level.

[0134] Table 6 shows the proportion of each taxonomic information at the genus level

[0135]

[0136] Therefore, the classification of Pseudomonas at the genus level is the final taxonomic information, which meets the voting ratio and belongs to the unique taxonomic information. Thus, the species taxonomic information of the sample is determined as:

[0137] d__Bacteria; p__Proteobacteria; c__Gammaproteobacteria; o__Pseudomonadales; f__Pseudomonadaceae; g__Pseudomonas.

[0138] In addition, it should be supplemented that: if there are two reference sequences and the voting ratios of the two reference sequences at a certain taxonomic level are both 50%, in this case, it means that the taxonomic information cannot be determined at this level and it is necessary to further search for the unique taxonomic information at a higher taxonomic level until it is found.

[0139] In addition to the species taxonomic information, the annotation information of the sample also includes sequence number, species number, species name, the lowest taxonomic level of the species, the taxonomic information of each taxonomic level of the species, the species number of each taxonomic level of the species, and the lowest alignment score of each comparison reference sequence of the species. The higher the score, the better the alignment quality. The annotation information file is as Figure 7 shown, and the meanings of the data in each column of the annotation information file are shown in Table 7 below.

[0140] Table 7 is the description of the annotation information file

[0141]

[0142]

[0143] As Figure 8 shown, an embodiment of the present invention provides a device for constructing a full-length ribosomal RNA database, and the device includes:

[0144] A first acquisition module 301, configured to acquire the coordinate information, species taxonomic information, and species name of the ribosomal RNA genes of all microorganisms, where the coordinate information includes the genome number where the ribosomal RNA gene is located, the starting position of the ribosomal RNA, and the ending position of the ribosomal RNA;

[0145] A screening module 302, configured to screen the genomes where the ribosomal RNA genes of all microorganisms are located based on preset conditions and the species classification information;

[0146] A merging module 303, configured to arrange the screened ribosomal RNA genes according to the sizes of the ribosomal RNA start positions. If two adjacent ribosomal RNA genes are respectively the small subunit and the large subunit of the same ribosome, and the distance between the two adjacent ribosomal RNA genes is less than a threshold, then merge the coordinate information of the two adjacent ribosomal RNA genes to obtain the coordinate information of the full-length ribosomal RNA; if the merging condition is not satisfied, then retain the coordinate information of the original ribosomal RNA genes;

[0147] A second obtaining module 304, configured to obtain the sequence information of the ribosomal RNA genes corresponding to the coordinate information according to the coordinate information of the full-length ribosomal RNA and the coordinate information of the original ribosomal RNA genes that do not meet the merging conditions, and generate a fasta sequence format file for the sequence information;

[0148] A construction module 305, configured to merge the fasta sequence format files of all ribosomal RNA genes to obtain a full-length ribosomal RNA database.

[0149] In one example, the screening module 302 is specifically configured to:

[0150] Determine the species classification level according to the species classification information;

[0151] If the species classification level has species-level information, then the species classification information meets the preset conditions, and retain the genome where the ribosomal RNA gene corresponding to the species classification information is located;

[0152] If the species classification level does not have species-level information, then the species classification information does not meet the preset conditions, and delete the genome where the ribosomal RNA gene corresponding to the species classification information is located.

[0153] In one example, the determining the coordinate information of the full-length ribosomal RNA includes:

[0154] Determine the start position of the full-length ribosomal RNA as the smallest ribosomal RNA start position among the small subunit and the large subunit;

[0155] Determine the end position of the full-length ribosomal RNA as the largest ribosomal RNA end position among the small subunit and the large subunit.

[0156] In one example, the sequence information includes one or more of a sequence position name, a sequence length, a brief description of the sequence, a unique sequence number, sequence keywords, a species name, genomic region information, a gene name, and sequence base information.

[0157] In one example, if the species name in the sequence information is inconsistent with the obtained species name, the obtained species name is revised to the species name in the sequence information.

[0158] In one example, the present invention further provides a schematic structural diagram of an application device for a full-length ribosomal RNA database. The device includes:

[0159] An alignment module 401, configured to align the sequencing result sequence of a sample with the sequence information of each ribosomal RNA gene in the full-length ribosomal RNA database to obtain alignment information;

[0160] A determination module 402, configured to determine the species classification information of the sample according to the alignment information.

[0161] In an implementable manner, the alignment information includes one or more of the sequence length of the sample, the starting position of the sequence, the ending position of the sequence, the number of bases aligned and matched, the alignment length, and the dynamic programming alignment score.

[0162] In one example, the determination module 407 is specifically configured to:

[0163] Determine the alignment similarity between the sequencing result sequence of the sample and the sequence information of each ribosomal RNA gene in the full-length ribosomal RNA database according to the number of bases aligned and matched and the alignment length, and filter the sequence information of ribosomal RNA genes with an alignment similarity less than the alignment similarity threshold according to the alignment similarity threshold;

[0164] Determine the alignment coverage of the sequencing result sequence of the sample and the sequence information of each ribosomal RNA gene in the filtered full-length ribosomal RNA database according to the sequence length, the starting position of the sequence, and the ending position of the sample, and filter the sequence information of ribosomal RNA genes with an alignment coverage less than the alignment coverage threshold according to the alignment coverage threshold;

[0165] Determine at least one reference sequence from the filtered ribosomal RNA genes according to the dynamic programming alignment score and the dynamic programming alignment score threshold;

[0166] Calculate the classification information ratio of the reference sequences according to the classification hierarchy from low to high;

[0167] Determine the species classification information of the sample according to the classification information ratio and the reference sequence.

[0168] In another aspect, the present invention provides a computer-readable storage medium storing a computer program for executing the method for constructing the full-length ribosomal RNA database of the present invention.

[0169] In yet another aspect, the present invention provides an electronic device, comprising:

[0170] a processor;

[0171] a memory for storing executable instructions of the processor;

[0172] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method for constructing the full-length ribosomal RNA database of the present invention.

[0173] In addition to the above methods and devices, embodiments of the present application may also be computer program products, which include computer program instructions that, when run on a processor, cause the processor to execute the steps in the methods according to various embodiments of the present application described in the "Exemplary Method" section above of this specification.

[0174] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0175] Furthermore, embodiments of the present application may also be computer-readable storage media storing computer program instructions that, when run on a processor, cause the processor to execute the steps in the methods according to various embodiments of the present application described in the "Exemplary Method" section above of this specification.

[0176] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0177] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present application are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present application. In addition, the above-disclosed specific details are only for the purposes of illustration and easy understanding, rather than limitations. The above details do not limit the present application to necessarily adopt the above specific details for implementation.

[0178] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present application are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended words, meaning "including but not limited to", and can be used interchangeably with each other. The words "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.

[0179] It should also be noted that in the devices, equipment, and methods of the present application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present application.

[0180] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0181] The foregoing description has been presented for purposes of illustration and description. In addition, this description is not intended to limit embodiments of the present application to the forms disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.

Claims

1. A method for constructing a full-length ribosomal RNA database, characterized in that, the method comprises: obtaining the coordinate information, species classification information, and species name of the ribosomal RNA genes of all microorganisms, where the coordinate information includes the genomic number where the ribosomal RNA is located, the start position of the ribosomal RNA, and the end position of the ribosomal RNA; screening the genomes where the ribosomal RNA genes of all microorganisms are located based on preset conditions and the species classification information; arranging the screened ribosomal RNA genes according to the size of the start position of the ribosomal RNA. If two adjacent ribosomal RNA genes are respectively the small subunit and the large subunit of the same ribosome, and the distance between the two adjacent ribosomal RNA genes is less than a threshold, then merge the coordinate information of the two adjacent ribosomal RNA genes to obtain the coordinate information of the full-length ribosomal RNA; if the merging condition is not met, then retain the coordinate information of the original ribosomal RNA gene; obtaining the sequence information of the ribosomal RNA genes corresponding to the coordinate information of the full-length ribosomal RNA after screening and merging and the coordinate information of the original ribosomal RNA that does not meet the merging condition, and generating a fasta sequence format file for the sequence information; merging the fasta sequence format files of all ribosomal RNA genes to obtain a full-length ribosomal RNA database.

2. The method according to claim 1, characterized in that, the screening of the genomes where the ribosomal RNA genes of all microorganisms are located based on preset conditions and the species classification information includes: determining the species classification level according to the species classification information; if the species classification level has species-level information, then the species classification information meets the preset conditions, and retain the genome where the ribosomal RNA gene corresponding to the species classification information is located; if the species classification level does not have species-level information, then the species classification information does not meet the preset conditions, and delete the genome where the ribosomal RNA gene corresponding to the species classification information is located.

3. The method according to claim 1, characterized in that, the obtaining of the coordinate information of the full-length ribosomal RNA includes: determining the start position of the full-length ribosomal RNA as the smallest start position of the small subunit and the large subunit; determining the end position of the full-length ribosomal RNA as the largest end position of the small subunit and the large subunit.

4. The method according to claim 1, characterized in that, the sequence information includes one or more of sequence position name, sequence length, sequence brief description, sequence unique number, sequence keyword, species name, genomic region information, gene name, and sequence base information.

5. The method according to claim 4, characterized in that, if the species name in the sequence information is inconsistent with the obtained species name, then determine the obtained species name as the species name in the sequence information.

6. An application method of a full-length ribosomal RNA database constructed by the method according to any one of claims 1-5, characterized in that, the application method includes: Align the sequencing result sequence of the sample with the sequence information of each ribosomal RNA gene in the full-length ribosomal RNA database to obtain alignment information; Determine the species classification information of the sample according to the alignment information.

7. According to the application method described in claim 6, it is characterized in that, the alignment information includes one or more of the sequence length, sequence start position, sequence end position, number of aligned bases, alignment length, and dynamic programming alignment score of the sample.

8. According to the application method described in claim 7, it is characterized in that, the determining the species classification information of the sample according to the alignment information includes: Determine the alignment similarity between the sequencing result sequence of the sample and the sequence information of each ribosomal RNA gene in the full-length ribosomal RNA database according to the number of aligned bases and the alignment length, and filter the sequence information of ribosomal RNA genes with an alignment similarity less than the alignment similarity threshold according to the alignment similarity threshold; Determine the alignment coverage of the sequencing result sequence of the sample and the sequence information of each ribosomal RNA gene in the filtered full-length ribosomal RNA database according to the sequence length, sequence start position, and sequence end position of the sample, and filter the sequence information of ribosomal RNA genes with an alignment coverage less than the alignment coverage threshold according to the alignment coverage threshold; Determine at least one reference sequence from the filtered full-length ribosomal RNA database according to the dynamic programming alignment score and the dynamic programming alignment score threshold; Calculate the classification information ratio of the reference sequences according to the classification hierarchy from low to high; Determine the species classification information of the sample according to the classification information ratio and the reference sequence.

9. A device for constructing a full-length ribosomal RNA database, it is characterized in that, the device includes: A first acquisition module for acquiring the coordinate information, species classification information, and species name of the ribosomal RNA genes of all microorganisms, where the coordinate information includes the genome number where the ribosomal RNA gene is located, the ribosomal RNA start position, and the ribosomal RNA end position; A screening module for screening the genomes where the ribosomal RNA genes of all microorganisms are located based on preset conditions and the species classification information; A merging module for arranging the screened ribosomal RNA genes in ascending order of the ribosomal RNA start position. If two adjacent ribosomal RNA genes are the small subunit and large subunit of the same ribosome respectively, and the distance between the two adjacent ribosomal RNA genes is less than the threshold, then merge the coordinate information of the two adjacent ribosomal RNA genes to obtain the coordinate information of the full-length ribosomal RNA; if the merging condition is not met, then retain the coordinate information of the original ribosomal RNA gene; A second acquisition module for acquiring the sequence information of the ribosomal RNA genes corresponding to the coordinate information of the full-length ribosomal RNA after screening and merging and the coordinate information of the ribosomal RNA that does not meet the merging condition, and generating a fasta sequence format file for the sequence information; A building block for merging fasta sequence format files of all ribosomal RNA genes to obtain a full-length ribosomal RNA database.

10. A computer-readable storage medium storing a computer program for executing the method for constructing the full-length ribosomal RNA database according to any one of claims 1-5 above.

11. An electronic device comprising: a processor; a memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method for constructing the full-length ribosomal RNA database according to any one of claims 1-5 above.

Citation Information

Patent Citations

  • System and method for splicing gene sequence fragments

    CN102867134A

  • Method for performing full genome sequence hole filling by means of long sequencing read segment

    CN109411020A