A method, device and related equipment for identifying microorganisms
By constructing a specific segment library of representative strains or representative strains, the problem of difficulty in identifying strain levels in the prior art is solved, high-precision strain recognition and new species recognition are achieved, and the accuracy and efficiency of functional strain screening are improved.
Patent Information
- Application Number
- CN202210784883.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-29
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-06-29
AI Technical Summary
The prior art is difficult to identify strain-level differences in metagenomic analysis, especially for strains with high similarity, and for new species to be identified.
A specific segment library is constructed that represents strains or species is obtained, and the specific segments are identified through specific alignment, combined with sequence alignment or biological probes, and a specific segment library is used to improve the recognition resolution and accuracy.
High-precision identification and distinction of strains with high similarity are achieved, the functional strain screening cycle is shortened, the accuracy and effectiveness of identification are improved, the strains of the same strain and species can be identified, and new species can be identified.
Smart Images

Figure CN115148288B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of microbial strain identification, and in particular to a microbial identification method, an identification device and related equipment. Background Art
[0002] Different strains of the same microorganism species may have functional differences even though their genomes are highly similar. Therefore, it is necessary to improve the resolution of microbial analysis to the strain level. Currently, there are two common methods for calculating the relative abundance of metagenomic microorganisms: one is to calculate using the metagenomic species composition analysis tool Metaphlan: (1) first find the marker genes of the reference genome, and (2) then perform species identification and relative abundance calculation based on the results of the alignment with the marker genes; however, the resolution of microbial identification based on marker genes is limited. Different strains of the same species can be expected to have highly similar genomes, so the strains are likely to have the same or highly similar marker genes. Therefore, different strains of the same species cannot be effectively identified, and it is difficult to distinguish strains with high genome similarity within the same species.
[0003] Another type of solution is the classification method based on the kmer-based Kraken and Centrifuge analysis tools: (1) first annotate the reference genome to the corresponding NCBI species classification database, (2) then annotate the sequencing fragments (reads) in the metagenome based on the kmer alignment, (3) based on the alignment results, determine which node on the evolutionary tree it belongs to and accurately locate it to the genus and species level; however, it is impossible to locate the strain; and new species do not have taxids and cannot be better annotated, so this KMER method cannot be used to identify new species and strains.
[0004] In addition, when building a custom database containing genomic data obtained by sequencing of self-screened strains, it is difficult to annotate the new species to the corresponding NCBI species classification database because there are no more reports on the new species before and there are no relevant records in the NCBI database.
[0005] In view of this, the present invention is proposed. Summary of the Invention
[0006] The purpose of the present invention is to provide a method, an identification device and related equipment for identifying microorganisms to solve the above technical problems.
[0007] Existing literature reports analyze metagenomic cohorts, but existing analysis methods can only reach the species level. However, in the real-world screening of live bacterial drugs or microbial pesticide candidates, strain-level analysis is required. Therefore, a method that can accurately identify strains is urgently needed, especially when analyzing metagenomic data, capable of identifying strains at the strain level.
[0008] When multiple samples in a metagenome are sequenced simultaneously, DNA is extracted and assembled into a genome, there will be problems such as insufficient coverage and different sequencing contigs. Incomplete assembly leads to limited accuracy of ANI recognition.
[0009] In summary, there is a need to provide a method with high resolution that can distinguish strains with high similarity, distinguish marker genes with high precision, and effectively and efficiently identify screened strains or new species. This method is easy to distinguish microorganisms and convenient to customize the database.
[0010] The present invention is achieved in that:
[0011] The present invention provides a method for identifying microorganisms, which comprises the following steps:
[0012] Obtain microbial genome sequences and build a microbial genome sequence database;
[0013] Perform cluster analysis on microbial genomes according to the set threshold;
[0014] After clustering, the representative genomes of each representative strain or the representative genomes of the representative bacterial species are selected to form a representative genome library;
[0015] Sequencing fragments (reads) of the genome library representing the strain or species are specifically aligned to obtain specific segments, and the specific segments of each strain or species are used to construct a specific segment library representing the strain or species;
[0016] Based on the specific segment library of the representative strain or representative bacterial species, the target microorganism is identified through sequence alignment or as a biological probe.
[0017] The above-mentioned method for identifying microorganisms can effectively identify target strains or strains through the specific segment library proposed in the present invention. It has high resolution for identifying microbial strains and can accurately identify and distinguish strains with high similarity, thereby achieving accurate identification of strains at the same strain level during the target strain screening process to obtain strains with the same target function. In addition, the above-mentioned method can also accurately identify strains of the same species to obtain strains of the same species. The provision of the method of the present invention helps to shorten the development cycle of microbial strains in the process of screening functional strains in medicines, foods, animal or plant microbiomes, environmental treatment, and agriculture, and improve the accuracy and effectiveness of functional strain screening. In addition, the method of the present invention can also effectively identify new species at the same time.
[0018] The above-mentioned step of constructing a specific segment library of representative strains or representative bacterial species includes: the sequencing fragments on the specific alignment are the position segments where the sequencing fragments (reads) of each representative strain are specific to other representative strains.
[0019] In an optional embodiment, the step of constructing a specific segment library of a representative strain or representative bacterial species includes: obtaining sequencing fragments (reads) of the representative strain or representative bacterial species that are specific to the sequencing fragments of other representative strains or species, determining the position segment of the sequencing fragments obtained based on the specific alignment of the representative strain or species, and merging the specific segments of each strain or species to obtain a specific segment library of the representative strain or representative bacterial species.
[0020] In an optional embodiment, the step of constructing a specific segment library representing a strain or a representative bacterial species includes: obtaining sequencing fragments (reads) of the genome sequence of the representative strain or representative bacterial species, using the sequencing fragments of any single strain representing the strain or representative bacterial species as input, and comparing the sequencing fragments (reads) of the genome of all representative strains or representative bacterial species in the representative genome library one by one; selecting the sequencing fragments (reads) of the genome of the above-mentioned representative strain that are specifically aligned and / or the positional segments where the sequencing fragments (reads) are located; using the positional segments of the sequencing fragments (reads) on the specific alignment or using the sequencing fragments on the specific alignment and the corresponding positional segments where the sequencing fragments (reads) are located to construct a specific segment library representing a strain or representative bacterial species.
[0021] In an optional embodiment, the steps for constructing a specific segment library of a representative strain or representative bacterial species are as follows: obtaining sequencing fragments (reads) of the genome sequence of the representative strain or representative bacterial species, using the sequencing fragments of a single strain of any representative strain or representative bacterial species as input, and comparing the genomes of all representative strains or representative bacterial species in the representative genome library one by one; selecting the sequencing fragments (reads) of the genome of the representative strain on the specific alignment, and recording the positions of the sequencing fragments (reads) on the representative genome, as set 1, selecting the sequencing fragments (reads) of the corresponding genome on the specific alignment and simultaneously aligning them to the sequencing fragments (reads) of the genomes of other representative strains or representative bacterial species with the same similarity, and recording their positions on the representative genome, as set 2; removing the intersection of set 1 and set 2 from set 1 to obtain set 3, which is the specific segment of the representative strain or representative bacterial species; using the same method to obtain the specific segments of all representative strains or representative bacterial species to form a specific segment library.
[0022] In an optional embodiment, the above-mentioned sequencing fragments (reads) may also be sequencing fragments that traverse the sequencing genome structure.
[0023] In the above method, constructing a specific segment library of representative strains or representative strains is a key technical means for improving recognition resolution, accurately distinguishing strains, and screening out strains with the same biological function. The inventor uses the sequencing fragments (reads) of a single strain of any representative strain or representative strain as input, compares the genomes of all representative strains or representative strains of the representative genome library one by one, selects the sequencing fragments (reads) of the representative strain genome on the specific alignment, obtains the position of each sequencing fragment (reads) on the representative genome, as set 1 (for example, recorded as locate-1); then compares the sequencing fragments (reads) with the representative genomes of all other representative strains, selects the sequencing fragments (reads) of the corresponding genome on the specific alignment, and simultaneously compares them to the sequencing fragments (reads) of the genomes of other representative strains or representative strains with the same similarity, and records their positions on the representative genome (for example, recorded as locate-n), as set 2.
[0024] Because there are some base mismatches (such as 1-2 bases) in set 2, the sequencing fragments (reads) are aligned with other representative genomes at the same high similarity. The inventors found that since locate-1 in set 1 is the most accurate and the longest. The other locate-n in set 2 has some positions overlapping with locate-1, and the length of the alignment must be less than the length of locate-1. Based on this, the position of the representative genome that is not aligned with other representative strains can be obtained by the difference between locate-1 and locate-n (for example, denoted as locate-x). That is, set 3 is obtained by removing the intersection with set 2 from set 1, and set 3 is the specific segment of the representative strain or representative strain; the same method is used to obtain the specific segments of all representative strains or representative strains to form a specific segment library.
[0025] Since this position (locate-x) can only exist in the representative genome of the representative strain with this sequencing fragment (reads), and does not exist in the representative genomes of other representative strains, it can more accurately identify single strains with higher resolution, and eliminate mismatched alignment results to a great extent, thereby obtaining more accurate alignment information and making the alignment results more accurate, thereby achieving specific identification of single strains.
[0026] By constructing set 1 and corresponding set 2 by sequencing the single-strain sequencing fragments (reads) of all representative strains, more accurate comparison results can be obtained for all representative strains, thereby facilitating more effective and efficient identification of screened strains or new species when screening for subsequent strains or strains required for microbial drugs, microbial agriculture, food microorganisms and other fields.
[0027] The present invention adopts this method to obtain comprehensive comparison information by comparing all representative strain sequences, so that the comparison results obtained are more accurate.
[0028] In a preferred embodiment of the present invention, the sequencing fragments (reads) of the genomic library representing the strain or the bacterial species are specifically aligned to obtain specific segments, and the specific segments of each strain or bacterial species are obtained to construct a specific segment library representing the strain or species, which also includes filtering all alignment results.
[0029] In an optional embodiment, the filtering condition is to allow a maximum of 4 base mismatches; preferably, the form of the base mismatch includes any one or more of the following: base mutation, insertion or deletion.
[0030] In an optional embodiment, all alignment results are filtered. After filtering, if there are still multiple alignment results for the same sequencing fragment (reads), the best alignment result of the sequencing fragment (reads) must meet the following conditions at the same time to obtain the best alignment result of the sequencing fragment (reads) and include it in set 1. The conditions are: (1) The best alignment result of the sequencing fragment (reads) allows at most 1 base mismatch; (2) The alignment result with the second highest alignment score of the sequencing fragment (reads) has at least 2 mismatches, preferably at least 3 mismatches, for example, 3-4. In an optional embodiment, the above score is calculated according to the form of base mismatch, and the form of base mismatch includes any one or more of the following: base mutation, insertion or deletion. Optionally, a base mutation is deducted 5 points, a base deletion or insertion is 15 points, and two base mutations are deducted 10 points.
[0031] The inventors found that the above-mentioned filtering strategy helps to further improve the resolution, can distinguish strains with high similarity, and can distinguish marker genes with high precision, so as to effectively and efficiently identify the screened strains or new species.
[0032] In a preferred embodiment of the present invention, the microbial genomes are clustered according to a set threshold to obtain the same strain cluster; the representative genomes of each representative strain are determined by any of the following methods:
[0033] When the same strains are obtained through clustering, the gene sequence with the longest length in the same strain cluster is selected as the representative genome of the same strain cluster;
[0034] Alternatively, when clustering yields the same strain, select each type of identical representative strain to calculate the average ANI, sort and select the strain gene sequence with the largest ANI as the representative genome of the representative strain;
[0035] Alternatively, when the same strain is obtained by clustering, the integrity and contamination are used as quality score scoring indicators, and the genome of the strain with the highest quality score value is calculated as the representative genome of the representative strain.
[0036] This helps to obtain more complete sequence information and avoid situations where the sequence cannot be obtained because it is located at both ends.
[0037] In the above strategy, the gene sequence with the longest length is selected as the representative genome of the strain, which helps to obtain more complete strain information.
[0038] In an optional embodiment, the representative genome of the representative bacterial species is determined by any of the following methods:
[0039] When strains of the same species are clustered, the gene sequence of the model strain is selected as the strain genome representing the species;
[0040] Alternatively, when strains of the same species are clustered, the strain gene sequence with the longest gene sequence length within the species is selected as the strain genome representing the species;
[0041] Alternatively, when strains of the same species are clustered, strains of the same species are selected to calculate the average ANI, and the strain gene sequence with the largest ANI is selected as the strain genome representing the species.
[0042] The genome of the model strain is more representative; therefore, the gene sequence of the model strain can be selected as the strain genome representing the species.
[0043] In a preferred embodiment of the present invention, the steps of identifying target microorganisms by sequence alignment or as biological probes based on a library of specific segments representing representative strains or species include:
[0044] Comparing the microbial strain or species to be identified with a library of specific segments representing the strain or species to identify the target strain or species;
[0045] Alternatively, the sequence information of a specific segment representing a strain or a bacterial species is used as a biological probe to detect a target strain or a target bacterial species;
[0046] Alternatively, based on the comparison information between the metagenomic sequencing data and the specific segment library of the representative strain or representative bacterial species; combining the length of the specific segment of the representative strain or representative bacterial species, the relative abundance of each strain is calculated; based on the relative abundance of each strain, the target strain or bacterial species is screened;
[0047] Alternatively, the relative abundance of each strain is calculated based on the comparison information between the metagenomic sequencing data and the specific segment library of the representative strain or representative bacterial species; and the target strain or bacterial species is screened based on the relative abundance of each strain and combined with biomarkers;
[0048] Alternatively, metagenome sequencing reads are used as input and compared to a representative genome library using a sequence alignment tool. In the alignment results, reads that can be specifically aligned to a specific segment library and / or positional segments of sequence fragments that can be specifically aligned to a specific segment library of representative strains are retained. Combined with the lengths of the specific segments of each representative strain, the relative abundance of each strain is calculated to screen for target strains or species. The sequence alignment tool is preferably Bowtie2. Compared to the following method, this method adds the step of aligning reads generated by metagenomic sequencing to a representative strain library using Bowtie2. This step helps improve alignment accuracy and strain resolution.
[0049] Another way is:
[0050] According to the comparison information between the metagenomic sequencing data and the specific segment library representing the strain or representative bacterial species; retaining the sequencing fragments (reads) that can be specifically aligned to the specific segment library and / or retaining the position segment that can be specifically aligned to the sequencing fragments of the specific segment library, and combining the length of the specific segment representing the genome, the step of calculating the relative abundance of each strain preferably includes: using the metagenomic sequencing fragments (reads) as input and directly aligning them with the specific segment library. This method is to directly align the reads obtained by metagenomic sequencing to the specific segment library representing the strain or representative bacterial species. This eliminates the step of aligning the reads obtained by metagenomic sequencing to the representative strain library through Bowtie2, simplifies the screening process, improves the alignment efficiency, and can also meet the resolution of a single strain.
[0051] In an optional embodiment, the sample source of the metagenome is a non-natural environment sample or a natural environment sample;
[0052] In an optional embodiment, the non-natural environment sample is a microbial population from an animal, a microbial population from a plant, a microbial population from a drug, a microbial population from a fertilizer, or a microbial population from food; the natural environment sample is a sample from soil, water, or air;
[0053] In an alternative embodiment, the microbial population from an animal body is a microbial population from the human intestine, the human stomach, the nasal cavity, the (inner and / or outer) ear canal, the eye, the skin, the human oral cavity, or the human reproductive tract.
[0054] The human intestine includes, but is not limited to, the small intestine, large intestine, and rectum.
[0055] The human reproductive tract includes but is not limited to: male internal genitalia, male external genitalia, female internal genitalia, and female external genitalia.
[0056] The male internal reproductive organs include, for example, the testicles, epididymis, vas deferens, ejaculatory ducts, prostate, seminal vesicles, and bulbourethral glands.
[0057] Female internal reproductive organs such as the ovaries, fallopian tubes, uterus, and vagina. Female external reproductive organs such as the labia majora, clitoris, and vestibule.
[0058] In an optional embodiment, the natural environment sample is from soil after application of bacterial fertilizer, soil after application of pesticide, domestic sewage or industrial sewage.
[0059] In a preferred embodiment of the present invention, clustering analysis of microbial genomes is performed with an ANI threshold of 95% or 99%.
[0060] In an optional embodiment, strains with ANI ≥ 99% are clustered as the same strain cluster.
[0061] In an alternative embodiment, strains with an ANI ≥ 95% are clustered as strains of the same species.
[0062] In a preferred embodiment of the present invention, obtaining a microbial genome sequence and constructing a microbial gene sequence database includes: obtaining a microbial genome sequence based on at least one of the following databases:
[0063] Human intestinal microbial genome sequence database, agricultural microbial sequence database, genome sequences obtained by sequencing microorganisms collected from the microbial resource platform, agricultural fertilizer microbial sequence database, fungal medicine microbial sequence database, sewage treatment microbial sequence database and food microbial field data.
[0064] The agricultural microbial genome sequence database can include gene sequences of bacteria, fungi, and other microorganisms. In other embodiments, the genome sequences obtained by sequencing microorganisms collected from the microbial resource platform can also be combined with existing database data to construct a database of agricultural fertilizer and fungal medicine microbial sequences.
[0065] The present invention also provides a device for microorganism identification, which includes: a microorganism gene sequence database construction unit, a microorganism genome clustering unit, a representative strain or representative bacterial species selection unit, a representative strain or representative bacterial species specific segment library construction unit, and a bacterial species or strain identification unit;
[0066] The microbial gene sequence database construction unit obtains the microbial genome sequence and constructs the microbial genome sequence database;
[0067] The microbial genome clustering unit performs cluster analysis on the microbial genome according to the set threshold;
[0068] The representative strain or representative bacterial species selection unit selects the representative genomes of each type of representative strain after clustering or selects the representative genomes of the representative bacterial species to form a representative genome library;
[0069] The unit for constructing a specific segment library representing a strain or a species of bacteria obtains specific segments by specific alignment of sequenced fragments (reads) of a genome library representing a strain or a species of bacteria, and constructs a specific segment library representing a strain or a species of bacteria from the obtained specific segments of each strain or species of bacteria;
[0070] The bacterial species or strain identification unit is based on the comparison information between the metagenomic sequencing data and the specific segment library representing the strain or representative bacterial species; the sequencing fragments (reads) that are specifically aligned to the specific segment library and / or the position segments that can be specifically aligned to the sequencing fragments of the specific segment library are retained, and the relative abundance of each strain is calculated in combination with the length of the specific segment representing the strain or representative bacterial species; the target strain or bacterial species is screened out based on the relative abundance at the strain level.
[0071] The bacterial species or strain identification unit is based on a library of specific segments representing the bacterial strain or bacterial species, and identifies the target bacterial strain or bacterial species through sequence comparison or as a biological probe.
[0072] In an optional embodiment, the microbial gene sequence database construction unit constructs the above-mentioned microbial gene sequence database.
[0073] In an optional embodiment, the microbial genome clustering unit is clustered with an ANI threshold of 95% or 99%; preferably, strains with ANI ≥ 99% are clustered as the same strain cluster; strains with ANI ≥ 95% are clustered as strains of the same species.
[0074] In an optional embodiment, representative strains or representative bacterial species selection units are used to construct the above-mentioned representative genome library.
[0075] In an optional embodiment, the construction unit of the specific segment library representing the strain or the bacterial species is further to construct the specific segment library representing the above-mentioned strain or the bacterial species.
[0076] In an optional embodiment, the bacterial species or strain identification unit screens out the target bacterial strain or bacterial species according to the above method.
[0077] The microorganism identification device provided by the present invention can be applied to the situation of determining microbial medicines, agricultural microorganisms, and microbial strains or species for environmental protection.
[0078] The present invention also provides a computer-readable storage medium comprising instructions, which, when executed on a computer, enable the computer to execute the above-mentioned method for identifying microorganisms.
[0079] The present invention has the following beneficial effects:
[0080] The method for identifying microorganisms provided by the present invention has high resolution for identifying microbial strains, can identify and differentiate strains with high similarity with high precision, and can accurately identify strains of the same strain level during the target strain screening process to obtain strains with the same target function. In addition, the above method can also accurately identify strains of the same species to obtain strains of the same species. The provision of the method of the present invention helps to shorten the development cycle and improve the accuracy and effectiveness of functional strain screening in the process of screening microbial strains in medicines, foods, animal or plant microbiomes, environmental treatment, and agriculture. In addition, the method of the present invention can also effectively identify new species at the same time.
[0081] The inventors constructed a specific segment library representing strains or representative bacterial species to improve the recognition resolution, effectively expand the recognition range, accurately identify and distinguish strains at the strain level, and provide assistance for subsequent strain screening to achieve the same biological function.
[0082] This can lead to the development of corresponding microbial identification devices, equipment and computer-readable storage media, thereby accurately identifying microorganisms in the fields of microbial medicine, agricultural microorganisms, food microorganisms, environmental treatment microorganisms, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0084] Figure 1 A flow chart of a microorganism identification method provided in Example 1;
[0085] Figure 2 A flow chart of a microorganism identification method provided in Example 2;
[0086] Figure 3 A schematic diagram of the structure of a device for identifying a microbial strain;
[0087] Figure 4 Schematic diagram of the hardware structure of the device;
[0088] Figure 5 Schematic diagram of the microorganism identification method of the present invention;
[0089] Figure numbers: 110 - microbial gene sequence database construction unit; 120 - microbial genome clustering unit; 130 - representative strain or representative bacterial species selection unit; 140 - representative strain or representative bacterial species specific segment library construction unit; 150 - bacterial species or strain identification unit; 210 - processor; 220 - memory; 230 - input device; 240 - output device; 250 - bus. DETAILED DESCRIPTION
[0090] Reference will now be made in detail to embodiments of the present invention, one or more examples of which are described below. Each example is provided to illustrate, not to limit, the present invention. Indeed, it will be apparent to those skilled in the art that various modifications and variations may be made to the present invention without departing from the scope or spirit of the invention. For example, features illustrated or described as part of one embodiment may be used in another embodiment to produce further embodiments.
[0091] The practice of the present invention will employ, unless otherwise indicated, conventional techniques of cell biology, molecular biology (including recombinant techniques), microbiology, biochemistry, and immunology, which are within the capabilities of a person skilled in the art. The technique is fully explained in the literature, for example, in Molecular Cloning: A Laboratory Manual, 2nd ed. (Sambrook et al., 1989); Oligonucleotide Synthesis (MJ Gait, ed., 1984); Animal Cell Culture (RI Freshney, ed., 1987); Methods in Enzymology (Academic Press, Inc.); Handbook of Experimental Immunology (DM Weir and CC Blackwell, eds.); Gene Transfer Vectors for Mammalian Cells (JM Miller and MP Calos, eds., 1987); Current Protocols in Molecular Biology (FM Ausubel et al., eds., 1987); and PCR: The Polymerase Chain Reaction. Reaction" (Mullis et al., eds., 1994); and Current Protocols in Immunology (JE Coligan et al., eds., 1991), each of which is expressly incorporated herein by reference.
[0092] To make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely below. Where specific conditions are not specified in the embodiments, conventional conditions or conditions recommended by the manufacturer are used. Where the manufacturer of the reagents or instruments is not specified, all are conventional products that can be purchased commercially.
[0093] The features and performance of the present invention are further described in detail below with reference to the embodiments.
[0094] Example 1
[0095] Reference Figure 1 As shown, this embodiment provides a method for identifying microorganisms.
[0096] It includes the following steps:
[0097] Step S1: Obtain microbial genome sequences and construct a microbial gene sequence database (such as a human intestinal microbial sequence database, an agricultural microbial sequence database);
[0098] One implementation method is: it can be obtained based on an existing human intestinal microbial genome sequence database, for example, a human intestinal microbial sequence database can be obtained through the UHGG database, or it can be based on microorganisms collected from a microbial resource platform and sequenced to obtain genome sequences and existing UHGG database data to jointly construct a human intestinal microbial sequence database; the present invention preferably uses microorganisms collected from the Muen microbial resource platform, and sequences 150,000 strains of genome sequences, and combines the UHGG database resources to construct a human intestinal microbial genome sequence database; when used for human intestinal microbial drug screening, the present invention also constructs a human intestinal microbial genome sequence database.
[0099] In other embodiments, a database of agricultural microbial genome sequences can be obtained, including gene sequences of bacteria, fungi, and other microorganisms. Similarly, a database of agricultural microbial fertilizer and fungal medicine sequences can be constructed by combining the genome sequences of microorganisms collected and sequenced from a microbial resource platform with existing database data. The present invention preferentially utilizes microorganisms collected and sequenced from the Muen Microbial Resource Platform, obtaining the genome sequences of 150,000 strains and combining them with known database resources to construct a database of agricultural microbial fertilizer and fungal medicine genome sequences. In other embodiments, the genome sequences of these strains can also be adaptively expanded as needed.
[0100] In other embodiments, the method provided by the present invention can also construct an environmental treatment, such as sewage treatment, microbial sequence database; in one embodiment, it can also construct a food microbial field database.
[0101] Step S2: Clustering the microbial genomes of the above S1.
[0102] All microbial genomes of interest are used as input, and the ANI between each pair is calculated using fastANI, and clustering is performed based on a set threshold. As an implementation method, the ANI threshold can be selected as 99% for clustering, and strains with ANI ≥ 99% will be clustered together as the same strain.
[0103] As another embodiment, the ANI threshold is selected as 95% for clustering, and strains with ANI ≥ 95% are clustered together as strains of the same species.
[0104] Step S3: Select representative strains of each type of clustered strains.
[0105] For strains in the same strain cluster, as an implementation method, the gene sequence with the longest gene sequence length is selected as the representative strain genome sequence.
[0106] When the same strains are obtained by clustering, as another embodiment, the same strains are selected for average ANI calculation, the calculated average ANIs are ranked, and the strain gene sequence with the largest ranking is selected as the representative strain genome;
[0107] When the same strain is obtained by clustering, the optional method also includes: the representative strain genome is the strain with the highest quality score value calculated according to the quality score;
[0108] When strains of the same species are obtained through clustering, as an embodiment, the gene sequence of the model strain is selected as the strain genome representing the species;
[0109] When strains of the same species are obtained through clustering, as an implementation method, the strain gene sequence with the longest gene sequence length within the strain is selected as the strain genome representing the strain species;
[0110] When strains of the same species are clustered, as an implementation method, strains of the same species are selected to calculate the average ANI, the calculated average ANIs are ranked, and the strain with the largest ranking is selected as the strain genome representing the species.
[0111] The average ANI calculation method and the method for selecting the largest ranking are shown in Table 1 below:
[0112] Table 1 ANI statistics of different strains.
[0113]
[0114] Taking strains B1, B2, B3, and B4 as examples, the average ANI of strain B1 relative to other strains was 99.4, the average ANI of strain B2 relative to other strains was 97.75, the average ANI of strain B3 relative to other strains was 98, and the average ANI of strain B4 relative to other strains was 95; the strains were sorted according to the average ANI values and strain 1 with the maximum average ANI value of 99.4% was selected as the representative strain.
[0115] Step S4: Construct a specific segment library representing the strain.
[0116] The steps of constructing a specific segment library of a representative strain or representative bacterial species include: the sequencing fragments on the specific alignment are the position segments where the sequencing fragments (reads) of each representative strain are specific to other representative strains; the sequencing fragments are sequencing fragments obtained by sequencing a single strain, or constructed sequencing fragments obtained by traversing the sequencing genome;
[0117] The positional segment of the sequenced fragments obtained by specific alignment of the representative strains or species is determined, and the specific segments of each strain or species are merged to obtain a specific segment library of the representative strain or species.
[0118] Specifically, Bowtie2 can be used to build a representative genome sequence library. Then, a single bacterial sequencing read of a representative strain or reads traversing its genome structure are used as input to align the representative genome library: reads that correspond to the representative strain genome in the specific alignment are selected, and their positions on the representative strain genome are recorded as set 1.
[0119] Select the genome reads corresponding to the specific alignment and align them simultaneously with reads from other representative strains or species at the same similarity. Note their positions on the representative genomes as Set 2. Remove the intersection of Set 1 and Set 2 from Set 1. The remaining position set is the specific segment for the representative strain. Similarly, obtain the specific segments of all representative strains as a strain-level specific segment library.
[0120] The inventors realized that since reads are often only part of the genome, a considerable portion of reads have mismatches in one or two bases. Therefore, in order to obtain all the alignment results of the reads, the present invention further implements the following method: all alignment results are filtered, and the filtering condition is to allow a maximum of 4 base mismatches. After filtering, if the same read still has multiple alignment results, the best alignment result of the read must meet the following conditions at the same time before the best alignment result of the read is included in Set1. The conditions are: (1) The best alignment result of the read is allowed to have a maximum of 1 mismatch; (2) The alignment result with the second highest score for the read has at least 2 mismatches, for example, at least 3 base mismatches.
[0121] Therefore, the present invention adopts this method to obtain comprehensive comparison information by comparing all representative strain sequences, so that the comparison results obtained are more accurate.
[0122] Step S5: Identify target microorganisms.
[0123] Based on the specific segment library of representative strains or representative bacterial species constructed above, target strains or bacterial species (target microorganisms) are identified by sequence alignment or as biological probes;
[0124] In one embodiment, the microbial strain or species to be identified can be compared with a library of specific segments representing the strain or species to identify the target strain or species;
[0125] Another embodiment is: using the sequence information of the specific segment representing the strain or species as a biological probe to detect the target strain or species;
[0126] Other possible implementations include: comparing metagenomic sequencing data with a library of specific segments representing strains or species; calculating the relative abundance of each strain based on the length of the specific segments representing the strains or species; and screening for target strains or species based on the relative abundance of each strain in combination with biomarkers.
[0127] The embodiments that can be adopted by the present invention also include: using metagenomic sequencing reads as input, and Bowtie2 aligning the representative strain genome library. In the alignment results, only those reads that are specifically aligned to the strain-level specific segment library are retained, and the relative abundance of each strain is calculated in combination with the length of the representative genome specific segment. Based on the strain abundance and combined with the sample grouping information, we can analyze the differences in the entire intestinal flora for the human intestinal flora, and can also be used for biomarker screening. At the same time, it can also be used to evaluate the health status of the human intestine. If there are other omics data, the association analysis between the metagenomics and other omics can also be analyzed; similarly, other microbial flora such as ear canal microorganisms, environmental microorganisms, oral microorganisms, and reproductive tract microorganisms are also applicable. According to the relative abundance at the strain level, biomarkers are used to screen the strains or strains required for microbial medicine, microbial agriculture, and food microorganisms.
[0128] According to the method of the present invention, the strain identification device of the present invention has the advantage of high resolution, can distinguish strains with high similarity, and can distinguish marker genes with high precision, and can effectively and efficiently identify the screened strains or new species; the specific segment library enables the method of the present invention to effectively expand the recognition range, and can accurately identify and distinguish strains at the strain level, providing assistance for subsequent strain screening to achieve the same biological function.
[0129] Example 2
[0130] This embodiment provides a method for identifying microorganisms. Figure 2 As shown, the principle refers to Figure 5 As shown, it includes the following steps:
[0131] Step S1: Obtain the microbial genome sequence and construct a microbial gene sequence database (such as a human intestinal microbial sequence database, an agricultural microbial sequence database).
[0132] In one embodiment, a human intestinal microbial genome sequence database can be obtained, for example, from the UHGG database. Alternatively, a human intestinal microbial sequence database can be constructed based on genome sequences obtained by sequencing microorganisms collected from a microbial resource platform and the existing UHGG database data. The present invention preferably uses microorganisms collected from the Muen microbial resource platform, and sequences 150,000 strains to construct a human intestinal microbial genome sequence database in combination with the UHGG database resources. When used for human intestinal microbial drug screening, the present invention constructs a human intestinal microbial genome sequence database.
[0133] Another embodiment of the present invention can include obtaining an agricultural microbial genome sequence database, which can include gene sequences of bacteria, fungi, and other microorganisms. Similarly, a database of agricultural microbial fertilizer and fungal medicine microbial sequences can be constructed based on the genome sequences of microorganisms collected from a microbial resource platform and sequenced, combined with existing database data. The present invention preferentially uses microorganisms collected from the Muen Microbial Resource Platform, sequencing the genome sequences of 150,000 strains, and combining these with known database resources to construct a database of agricultural microbial fertilizer microbial genome sequences.
[0134] In addition, in other embodiments, the present invention can also construct an environmental treatment, such as sewage treatment, microbial sequence database; one embodiment can also construct a food microbial field database.
[0135] Step S2: Clustering microbial genomes.
[0136] All microbial genomes of interest are used as input, and the ANI between each pair is calculated using fastANI, and clustering is performed based on a set threshold. As an implementation method, the ANI threshold can be selected as 99% for clustering, and strains with ANI ≥ 99% will be clustered together as the same strain.
[0137] As another embodiment, the ANI threshold of 95% is selected for clustering, and strains with ANI ≥ 95% are clustered together as strains of the same species;
[0138] Step S3: Select representative strains of each type of clustered strains;
[0139] The microbial genomes are clustered according to the set threshold to obtain the same strain cluster; the representative genomes of each representative strain are determined by any of the following methods:
[0140] When the same strain is obtained through clustering, for the strains in the same strain cluster, as an implementation method, the gene sequence with the longest gene sequence length is selected as the representative strain genome sequence;
[0141] When the same strains are obtained by clustering, as another embodiment, the same strains of various types are selected to calculate the average ANI, and the strain gene sequence with the largest ANI is selected as the representative strain genome;
[0142] When the same strain is obtained by clustering, an optional method also includes: using integrity and contamination as quality score scoring indicators, and calculating the genome of the strain with the highest quality score as the representative genome of the representative strain;
[0143] The representative genome of the representative bacterial species is determined by any of the following methods:
[0144] When strains of the same species are clustered, the gene sequence of the model strain is selected as the strain genome representing the species;
[0145] Alternatively, when strains of the same species are clustered, the strain gene sequence with the longest gene sequence length within the species is selected as the strain genome representing the species;
[0146] Alternatively, when strains of the same species are clustered, strains of the same species are selected to calculate the average ANI, and the strain gene sequence with the largest ANI is selected as the strain genome representing the species.
[0147] The calculation method and ranking of average ANI can be as shown in Table 1 of Example 1.
[0148] Step S4: Constructing a specific segment library representing the strain
[0149] The step of constructing a specific segment library of a representative strain or representative bacterial species includes: the sequencing fragments on the specific alignment are the position segments where the sequencing fragments (reads) of each representative strain are specific to other representative strains; the sequencing fragments are sequencing fragments obtained by sequencing a single strain, or constructed sequencing fragments obtained by traversing the sequencing genome;
[0150] Determine the position segment of the sequenced fragments obtained by specific alignment of the representative strain or species, and merge the specific segments of each strain or species to obtain a specific segment library of the representative strain or species;
[0151] Specifically, Bowtie2 can be used to build a representative genome sequence library. Then, a single bacterial sequencing read of a representative strain or reads traversing its genome structure are used as input to align the representative genome library: reads that correspond to the representative strain genome in the specific alignment are selected, and their positions on the representative strain genome are recorded as set 1.
[0152] Select the genome reads corresponding to the specific alignment and align them simultaneously with reads from other representative strains or species at the same similarity. Note their positions on the representative genomes as Set 2. Remove the intersection of Set 1 and Set 2 from Set 1. The remaining position set is the specific segment for the representative strain. Similarly, obtain the specific segments of all representative strains as a strain-level specific segment library.
[0153] In the present invention, the inventors realized that since reads are often only part of the genome, a considerable portion of reads have mismatches at one or two bases. Therefore, in order to obtain all alignment results of the reads, the present invention further implements the following method: all alignment results are filtered, and the filtering condition is to allow a maximum of 4 base mismatches. After filtering, if the same read still has multiple alignment results, the best alignment result of the read must meet the following conditions at the same time before the best alignment result of the read is included in Set1, the conditions are: (1) the best alignment result of the read is allowed to have a maximum of 1 mismatch; (2) the alignment result with the second highest score for the read has at least 2 mismatches, for example, 3 bases are preferred.
[0154] The inventors used Set 1 to align the sequencing reads of each representative strain with the representative genome of that strain, obtaining the position of each read in the representative genome (e.g., denoted as locate-1). They then compared the reads with the representative genomes of all other representative strains. If they could be aligned (or aligned with the same high similarity), the position of the read in the representative genome of the other representative strain that could be aligned was recorded (e.g., denoted as locate-n) as Set 2. Because Set 2 contained some base mismatches (e.g., 1-2 bases), which caused the reads to align with the other representative genomes with the same high similarity, the inventors found that locate-1 in Set 1 was the most accurate and longest. The other locate-n in set 2 has some positions overlapping with locate-1, and the length of the alignment must be less than the length of locate-1. Based on this, the position of the representative genome that is not aligned with other representative strains can be obtained through the difference between locate-1 and locate-n (for example, recorded as locate-x). Since this position can only exist in the representative genome of the representative strain with this read, and does not exist in the representative genomes of other representative strains, it can more accurately identify single strains with higher resolution, and eliminate mismatched alignment results to a great extent, thereby obtaining more accurate alignment information, making the alignment results more accurate, and thus achieving specific identification of single strains.
[0155] By constructing Set 1 and the corresponding Set 2 for the single-strain sequencing reads of all representative strains, more accurate comparison results can be obtained for all representative strains, thereby facilitating more effective and efficient identification of screened strains or new species when screening for strains or strains required for subsequent microbial pharmaceuticals, microbial agriculture, food microbiology and other fields.
[0156] Therefore, the present invention adopts this method to obtain comprehensive comparison information by comparing all representative strain sequences, so that the comparison results obtained are more accurate.
[0157] S5: Identification of target microorganisms:
[0158] Step S5: Identify target microorganisms.
[0159] Based on the specific segment library of representative strains or representative bacterial species constructed above, target strains or bacterial species (target microorganisms) are identified by sequence alignment or as biological probes;
[0160] In one embodiment, the microbial strain or species to be identified can be compared with the specific segment library of the representative strain or species to identify the target strain or species;
[0161] Another embodiment is: using the sequence information of the specific segment representing the strain or species as a biological probe to detect the target strain or species;
[0162] Other possible implementations include: comparing metagenomic sequencing data with a library of specific segments representing strains or species; calculating the relative abundance of each strain based on the length of the specific segments representing the strains or species; and screening for target strains or species based on the relative abundance of each strain in combination with biomarkers.
[0163] Specifically, metagenomic sequencing reads are used as input and mapped to a strain-specific segment library. Only reads that specifically map to the strain-specific segment library are retained. The relative abundance of each strain is calculated based on the length of the representative genome-specific segment. Based on the relative abundance at the strain level, biomarkers are used to screen for strains or bacterial species that are promising for applications in microbial medicine, microbial agriculture, and food microbiology.
[0164] According to the method of the present invention, the strain identification device of the present invention has the advantage of high resolution, can distinguish strains with high similarity, and can distinguish marker genes with high precision, and can effectively and efficiently identify the screened strains or new species; the specific segment library enables the method of the present invention to effectively expand the recognition range, and can accurately identify and distinguish strains at the strain level, providing assistance for subsequent strain screening to achieve the same biological function.
[0165] Example 3
[0166] This embodiment provides a device for identifying microbial strains, referring to Figure 3 This embodiment is applicable to the determination of microbial drugs, agricultural microorganisms, and microbial strains or species for environmental protection.
[0167] The device for identifying microbial strains includes: a microbial gene sequence database construction unit 110, a microbial genome clustering unit 120, a representative strain or representative bacterial species selection unit 130, a representative strain or representative bacterial species specific segment library construction unit 140, and a bacterial species or strain identification unit 150.
[0168] (1) The microbial gene sequence database construction unit 110 is used to obtain microbial genome sequences and construct a microbial gene sequence database (human intestinal microbial sequence database, agricultural fertilizer microbial sequence database, food microbial sequence database, and environmental protection microbial sequence database).
[0169] The above-mentioned human intestinal microbial sequence database can be obtained based on an existing human intestinal microbial genome sequence database, for example, a human intestinal microbial sequence database can be obtained through the UHGG database, or a human intestinal microbial sequence database can be constructed based on the genome sequences obtained by sequencing microorganisms collected from a microbial resource platform and the existing UHGG database data. The present invention preferably uses microorganisms collected from the Muen microbial resource platform, and sequences 150,000 strains to construct a human intestinal microbial genome sequence database in combination with the UHGG database resources. When used for human intestinal microbial drug screening, the present invention constructs a human intestinal microbial genome sequence database.
[0170] Another embodiment of the present invention can include obtaining an agricultural microbial genome sequence database, which can include gene sequences of bacteria, fungi, and other microorganisms. Similarly, a database of agricultural microbial fertilizer and fungal medicine microbial sequences can be constructed based on the genome sequences of microorganisms collected from a microbial resource platform and sequenced, combined with existing database data. The present invention preferentially uses microorganisms collected from the Muen Microbial Resource Platform, sequencing the genome sequences of 150,000 strains, and combining these with known database resources to construct a database of agricultural microbial fertilizer microbial genome sequences.
[0171] In addition, the present invention can also construct an environmental treatment, such as sewage treatment, microbial sequence database; one embodiment can also construct a food microbial field database.
[0172] (2) Microbial genome clustering unit 120: used to cluster microbial genomes.
[0173] All microbial genomes obtained from the genome database in step (1) are used as input, and the ANI between each pair is calculated by fastANI, and clustering is performed with a set threshold. As an embodiment, the ANI threshold can be selected as 99% for clustering, and strains with ANI ≥ 99% will be clustered together as the same strains;
[0174] As another embodiment, the ANI threshold of 95% is selected for clustering, and strains with ANI ≥ 95% are clustered together as strains of the same species;
[0175] (3) Representative strain or representative bacterial species selection unit 130: used to select representative strains or representative bacterial species of each type of clustered strains.
[0176] The microbial genomes are clustered according to the set threshold to obtain the same strain cluster; the representative genomes of each representative strain are determined by any of the following methods:
[0177] When the same strain is obtained through clustering, for the strains in the same strain cluster, as an implementation method, the gene sequence with the longest gene sequence length is selected as the representative strain genome sequence;
[0178] When the same strains are obtained by clustering, as another embodiment, the same strains of various types are selected to calculate the average ANI, and the strain gene sequence with the largest ANI is selected as the representative strain genome;
[0179] When the same strain is obtained by clustering, an optional method also includes: using integrity and contamination as quality score scoring indicators, and calculating the genome of the strain with the highest quality score as the representative genome of the representative strain;
[0180] The representative genome of the representative bacterial species is determined by any of the following methods:
[0181] When strains of the same species are clustered, the gene sequence of the model strain is selected as the strain genome representing the species;
[0182] Alternatively, when strains of the same species are clustered, the strain gene sequence with the longest gene sequence length within the species is selected as the strain genome representing the species;
[0183] Alternatively, when strains of the same species are clustered, strains of the same species are selected to calculate the average ANI, and the strain gene sequence with the largest ANI is selected as the strain genome representing the species.
[0184] The average ANI calculation method and the maximum selection sorting result can be shown in Table 1 of Example 1.
[0185] (4) A specific segment library construction unit 140 representing a strain or a species.
[0186] The step of constructing a specific segment library of a representative strain or representative bacterial species includes: the sequencing fragments on the specific alignment are the position segments where the sequencing fragments (reads) of each representative strain are specific to other representative strains; the sequencing fragments are sequencing fragments obtained by sequencing a single strain, or constructed sequencing fragments obtained by traversing the sequencing genome; the position segments where the sequencing fragments obtained by the specific alignment of the representative strain or bacterial species are located are determined, and the specific segments of each strain or bacterial species are merged to obtain the specific segment library of the representative strain or representative bacterial species;
[0187] First, use Bowtie2 to build a representative genome sequence library. Then, use the single bacterial sequencing reads of a representative strain or the reads of its genome structure as input to align the representative genome library: select the reads that correspond to the representative strain genome in the specific alignment, and record their positions on the representative strain genome as Set1;
[0188] Select the genome reads corresponding to the specific alignment and align them simultaneously with reads from other representative strains or species at the same similarity. Note their positions on the representative genomes as Set 2. Remove the intersection of Set 1 and Set 2 from Set 1. The remaining position set is the specific segment for the representative strain. Similarly, obtain the specific segments of all representative strains as a strain-level specific segment library.
[0189] The inventors realized that since reads are often only part of a genome, a considerable portion of reads have mismatches at one or two bases. Therefore, in order to obtain all alignment results for the reads, the present invention further implements the following method: all alignment results are filtered, and the filtering condition is to allow a maximum of 4 base mismatches. After filtering, if the same read still has multiple alignment results, the best alignment result of the read must meet the following conditions at the same time before the best alignment result of the read is included in Set1: (1) The best alignment result of the read is allowed to have a maximum of 1 mismatch; (2) The alignment result with the second highest score for the read has at least 3 mismatches, preferably 3 bases.
[0190] Therefore, the present invention adopts this method to obtain comprehensive comparison information by comparing all representative strain sequences, so that the comparison results obtained are more accurate.
[0191] (5) Bacteria species or strain identification unit 150.
[0192] The bacterial species or strain identification unit 150 is used to identify target microorganisms.
[0193] Based on the specific segment library of representative strains or representative bacterial species constructed above, target strains or bacterial species (target microorganisms) are identified by sequence alignment or as biological probes;
[0194] In one embodiment, the microbial strain or species to be identified can be compared with the specific segment library of the representative strain or species to identify the target strain or species;
[0195] Another embodiment is: using the sequence information of the specific segment representing the strain or species as a biological probe to detect the target strain or species;
[0196] Other possible implementations include: comparing metagenomic sequencing data with a library of specific segments representing strains or species; calculating the relative abundance of each strain based on the length of the specific segments representing the strains or species; and screening for target strains or species based on the relative abundance of each strain in combination with biomarkers.
[0197] The strain identification method of the present invention uses metagenomic sequencing reads as input and uses Bowtie2 to align the reads to a representative strain genome library. Only reads that specifically align to the strain-specific segment library are retained from the alignment results. The relative abundance of each strain is calculated based on the length of the representative genome segment.
[0198] Another possible implementation of the strain identification method of the present invention is to use metagenomic sequencing reads as input and align the reads obtained by metagenomic sequencing to a strain-specific segment library. From the alignment results, only reads that specifically align to the strain-specific segment library are retained. The relative abundance of each strain is calculated based on the length of the representative genome-specific segment.
[0199] The strain abundances obtained by the strain identification unit of the present invention, combined with sample grouping information, can be used to analyze differences in the entire human intestinal flora. This can also be used for biomarker screening and assessment of human intestinal health. If other omics data are available, the metagenome can also be analyzed for associations with other omics. Similarly, other microbial communities, such as environmental microbes, oral microbes, and genital microbes, are also applicable. Based on the relative abundance of strains, biomarkers can be used to screen for strains or strains needed for microbial medicine, microbial agriculture, and food microbiology.
[0200] The strain identification device according to the present invention has the advantage of high resolution, can distinguish strains with high similarity, and can distinguish marker genes with high precision, and can effectively and efficiently identify the screened strains or new species; the specific segment library enables the method of the present invention to effectively expand the recognition range, can accurately identify and distinguish strains at the strain level, and provide assistance for subsequent strain screening to achieve the same biological function
[0201] Example 4
[0202] This embodiment provides a device, which is a device for identifying microbial strains provided in Example 3 of the present invention. Specifically, the device includes: Figure 4As shown: one or more processors 210, memory 220, input device 230 and output device 240. The processor 210, memory 220, input device 230 and output device 240 in the device can be connected through a bus or other means.
[0203] Figure 4 Taking a processor 210 as an example, the processor 210 , the memory 220 , the input device 230 and the output device 240 are connected via a bus 250 .
[0204] The memory 220 is a non-transitory computer-readable storage medium that can be used to store software programs, computer executable programs, and modules, such as the program instructions / modules corresponding to the microorganism (strain, species) identification method in the first and second embodiments of the present invention (for example, Figure 3 The microbial gene sequence database construction unit 110, microbial genome clustering unit 120, representative strain or representative bacterial species selection unit 130, representative strain or representative bacterial species specific segment library construction unit 140, and bacterial species or strain identification unit 150 are shown. The processor 210 executes the software programs, instructions, and modules stored in the memory 220 to perform various functional applications and data processing of the device, that is, to implement a microbial identification method of the above method embodiment.
[0205] The required applications; the storage data area can store data created based on the use of the device, etc. In addition, the memory 220 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 220 may optionally include a memory remotely located relative to the processor 210, and these remote memories may be connected to the terminal device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0206] The input device 230 may be used to receive input digital or character information and generate key signal input related to user settings and function control of the device. The output device 240 may include a display device such as a display screen.
[0207] Example 5
[0208] An embodiment of the present invention also provides a computer-readable storage medium containing computer-executable instructions. When the computer-executable instructions are executed by a computer processor, they are used to execute the method for microbial identification in Example 1 or Example 2. Optionally, when the computer-executable instructions are executed by a computer processor, they can also be used to execute the technical solution of a method for microbial screening (identification) provided in any embodiment of the present invention.
[0209] Through the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented with the help of software and necessary general-purpose hardware, and of course it can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the method for identifying microorganisms in various embodiments of the present invention.
[0210] Example 6
[0211] Human intestinal microbial genome data were downloaded from the UHGG database, combined with the strain sequence data tested by Muen, and a specific segment library was obtained according to the method of the present invention, as shown in Table 2.
[0212] Table 2 Partial contents of the specific segment library.
[0213]
[0214]
[0215] The PRJNA541981 (melanoma patients treated with PD-1) project was selected from the NCBI SAR database, and the baseline samples in this project (27 stool samples before PD-1 treatment) were selected as the input metagenomic data; the method of the present invention and metaphlan were used for analysis at the same time; the running results of seven samples SRR9033749 to SRR9033754 and SRR9033760 were partially displayed, as shown in Table 3.
[0216] Table 3 Statistics of running results of different samples.
[0217]
[0218]
[0219] The results of the method of the present application were compared with those of Metaphlan, as shown in Table 4. Metaphlan's identification method can only identify bacteria at the species level. As shown in Table 4, the Metaphlan method identified 253 species and 122 genera, but only 270 strains (due to limited marker genes). However, the method of the present invention was able to identify 1732 strains, 1045 species, and 374 genera.
[0220] Table 4 Statistics of species, genus and strain identification results using different methods.
[0221]
[0222]
[0223]
[0224]
[0225] Furthermore, the genus Alistipes was randomly selected to further view the results and compare them. After the metagenomic data was input, the running results of seven samples SRR9033749 to SRR9033754 and SRR9033760 were partially displayed. Metaphlan and the method of the present invention were used to analyze the abundance distribution of each species under the genus Alistipes (the results are shown in Table 5). It can be seen from Table 5 that the method of the present invention can identify more species than metaphlan.
[0226] Table 5: Identification results of various species under the genus Alistipes using the methods of metaphlan and the present invention respectively.
[0227]
[0228] The identification results of Alistipes putredinis were further analyzed, and it was found that metaphlan could identify two strains of the same species under Alistipes putredinis (as shown in Table 6); while the present invention identified five results under Alistipes putredinis (as shown in Table 7).
[0229] Table 6: Metaphlan strain level identification
[0230]
[0231] Table 7: Strain level identification results using the method of the present invention (abundance values are displayed using normalized data)
[0232]
[0233] From the above results, it can be seen that the strain identification method of the present invention has a higher resolution than the prior art through the construction of the specific segment library proposed by the present invention, can distinguish strains with high similarity, and can distinguish marker genes with high precision, and can effectively and efficiently identify the screened strains or new species; it can accurately identify to the strain level.
[0234] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for identifying microorganisms, characterized in that: It includes the following steps: Obtain microbial genome sequences and build a microbial genome sequence database; Perform cluster analysis on microbial genomes according to the set threshold; After clustering, the representative genomes of each representative strain or the representative genomes of the representative bacterial species are selected to form a representative genome library; The sequencing fragments (reads) of the genomic library representing the strain or the bacterial species are specifically aligned to obtain specific segments, and the specific segments of each strain or bacterial species are used to construct a specific segment library representing the strain or the bacterial species; The step of constructing a specific segment library of representative strains or representative bacterial species includes: the sequencing fragments on the specific alignment are the position segments where the sequencing fragments (reads) of each representative strain are specific to other representative strains; The step of constructing a specific segment library of a representative strain or a representative bacterial species comprises: obtaining sequencing fragments (reads) of the representative strain or bacterial species that are specific to the sequencing fragments of other representative strains or bacterial species, determining the positional segments of the sequencing fragments on the representative genome based on the sequencing fragments obtained by specific alignment of the representative strain or bacterial species, and combining the specific segments of each strain or bacterial species to obtain a specific segment library of the representative strain or representative bacterial species; Based on the specific segment library of the representative strain or representative bacterial species, the target microorganism is identified by comparing the sequencing fragments on the specific alignment in the specific segment library and the position segment information of the sequencing fragments on the specific alignment on the representative genome or using them as biological probes.
2. The method according to claim 1, wherein: The sequencing fragments are sequencing fragments obtained by sequencing a single strain, or constructed sequencing fragments obtained by traversing a sequencing genome.
3. The method according to claim 1, wherein: The step of constructing a specific segment library representing a representative strain or a representative bacterial species includes: obtaining sequencing fragments (reads) of the genome sequence of the representative strain or representative bacterial species, using the sequencing fragments of a single strain of any representative strain or representative bacterial species as input, and comparing the sequencing fragments of the genomes of all representative strains or representative bacterial species in the representative genome library one by one; selecting the sequencing fragments of the representative strain genome on the differential alignment and / or the positional segments where the sequencing fragments are located; and constructing a specific segment library representing the representative strain or representative bacterial species using the positional segments of the sequencing fragments on the differential alignment or using the sequencing fragments on the differential alignment and the corresponding positional segments where the sequencing fragments are located.
4. The method according to claim 1, wherein: The steps for constructing a specific segment library of a representative strain or representative bacterial species are as follows: obtaining sequencing fragments (reads) of the genome sequence of the representative strain or representative bacterial species, using the single strain sequencing fragments of any representative strain or representative bacterial species as input, and comparing the genomes of all representative strains or representative bacterial species in the representative genome library one by one; selecting sequencing fragments (reads) that are aligned with the genome of the representative strain, and recording the positions of the sequencing fragments (reads) on the representative genome, as set 1, selecting sequencing fragments (reads) that are aligned with the corresponding genome and simultaneously aligning them with sequencing fragments (reads) of the genomes of other representative strains or representative bacterial species with the same similarity, and recording their positions on the representative genome, as set 2; removing the intersection of set 1 and set 2 from set 1 to obtain set 3, and set 3 is the position segment of the representative strain or representative bacterial species; using the same method to obtain the position segments of all representative strains or representative bacterial species to form a specific segment library.
5. The method according to claim 1, wherein The sequencing fragments (reads) of the genomic library representing the strain or the bacterial species are differentially aligned to obtain positional segments, and the obtained positional segments of each strain or bacterial species are used to construct a specific segment library representing the strain or bacterial species, which also includes filtering all alignment results.
6. The method according to claim 5, characterized in that The filtering condition allowed a maximum of 4 base mismatches.
7. The method according to claim 6, characterized in that The form of the base mismatch includes any one or more of the following: base mutation, insertion or deletion.
8. The method according to claim 6, characterized in that All alignment results are filtered. After filtering, if there are still multiple alignment results for the same sequencing fragment (reads), the best alignment result of the sequencing fragment (reads) must meet the following conditions at the same time to obtain the best alignment result of the sequencing fragment (reads) and include it in set 1. The conditions are: (1) the best alignment result of the sequencing fragment (reads) allows at most 1 base mismatch; (2) the alignment result with the second highest score of the sequencing fragment (reads) has at least 2 mismatches.
9. The method according to claim 8, characterized in that The number of mismatches in the alignment result with the second highest score among the sequencing reads is at least 3.
10. The method according to claim 8, characterized in that The score is calculated according to the form of base mismatch, and the form of base mismatch includes any one or more of the following: base mutation, insertion or deletion.
11. The method according to any one of claims 1 to 10, characterized in that: The microbial genomes are clustered according to the set threshold to obtain the same strain cluster; the representative genomes of each representative strain are determined by any of the following methods: When the same strains are obtained through clustering, the gene sequence with the longest length in the same strain cluster is selected as the representative genome of the same strain cluster; Alternatively, when clustering yields the same strain, select each type of identical representative strain to calculate the average ANI, sort and select the strain gene sequence with the largest ANI as the representative genome of the representative strain; Alternatively, when the same strain is obtained by clustering, the integrity and contamination are used as quality score scoring indicators, and the genome of the strain with the highest quality score value is calculated as the representative genome of the representative strain.
12. The method according to claim 11, characterized in that The representative genome of the representative bacterial species is determined by any of the following methods: When strains of the same species are clustered, the gene sequence of the model strain is selected as the strain genome representing the species; Alternatively, when strains of the same species are clustered, the strain gene sequence with the longest gene sequence length within the species is selected as the strain genome representing the species; Alternatively, when strains of the same species are clustered, strains of the same species are selected to calculate the average ANI, and the strain gene sequence with the largest ANI is selected as the strain genome representing the species.
13. The method according to any one of claims 1 to 10, characterized in that: The steps of identifying target microorganisms by sequence alignment or as biological probes based on the specific segment library of the representative strains or representative bacterial species include: Comparing the sequenced fragments of the microbial strain or species to be identified with the specific segment library of the representative strain or species to identify the target strain or species; Alternatively, the sequence information of the positional segment representing the strain or species is used as a biological probe to detect the target strain or species; Alternatively, based on the comparison information between the metagenomic sequencing data and the specific segment library of the representative strain or representative bacterial species; combined with the length of the positional segment of the representative strain or representative bacterial species, the relative abundance of each strain is calculated; based on the relative abundance of each strain, the target strain or bacterial species is screened; Alternatively, the metagenomic sequencing fragments (reads) are used as input, and the input sequences are aligned with the representative genome library using a sequence alignment tool; in the alignment results, the sequencing fragments (reads) that can be differentially aligned to the specific segment library and / or the position segments of the sequencing fragments that can be differentially aligned to the specific segment library of the representative strain are retained, and the relative abundance of each strain is calculated based on the length of the position segment of each representative strain to screen out the target strain or species; wherein the sequence alignment tool is Bowtie2; Alternatively, based on the comparison information between the metagenomic sequencing data and the specific segment library representing the strain or representative bacterial species; retain the sequencing fragments (reads) that can be differentially aligned to the specific segment library and / or retain the positional segments that can be differentially aligned to the sequencing fragments of the specific segment library, and calculate the relative abundance of each strain in combination with the length of the representative genomic positional segment.
14. The method according to claim 13, characterized in that The step of calculating the relative abundance of each strain includes: taking metagenomic sequencing fragments (reads) as input and directly comparing them with the specific segment library.
15. The method according to claim 13, characterized in that The sample source of the metagenome is a non-natural environment sample or a natural environment sample; The non-natural environmental sample is a microbial population from an animal, a microbial population from a plant, a microbial population from a drug, a microbial population from a fertilizer, or a microbial population from a food; The natural environment samples are samples from soil, water or air; The microbial community from the animal body is a microbial community from the human intestine, the human stomach, the nasal cavity, the ear canal, the eye, the skin, the human oral cavity or the human reproductive tract; The natural environment sample is from soil after application of bacterial fertilizer, soil after application of pesticide, domestic sewage or industrial sewage.
16. The method according to claim 13, characterized in that The clustering analysis of the microbial genome is performed with an ANI threshold of 95% or 99%.
17. The method according to claim 16, characterized in that Strains with ANI ≥ 99% were clustered as the same strain cluster.
18. The method according to claim 16, characterized in that Strains with ANI ≥ 95% were clustered as strains of the same species.
19. The method according to any one of claims 1 to 10 and claims 16 to 18, characterized in that: The obtaining of microbial genome sequences and constructing of a microbial gene sequence database comprises: obtaining microbial genome sequences according to at least one of the following databases: Human intestinal microbial genome sequence database, agricultural microbial sequence database, genome sequences obtained by sequencing microorganisms collected from the microbial resource platform, agricultural fertilizer microbial sequence database, fungal medicine microbial sequence database, sewage treatment microbial sequence database and food microbial field data.
20. A device for identifying microorganisms, characterized in that: It includes: Microbial gene sequence database construction unit, microbial genome clustering unit, representative strain or representative bacterial species selection unit, representative strain or representative bacterial species specific segment library construction unit, bacterial species or strain identification unit; The microbial gene sequence database construction unit obtains a microbial genome sequence and constructs a microbial genome sequence database; The microbial genome clustering unit performs cluster analysis on the microbial genome according to a set threshold; The representative strain or representative bacterial species selection unit selects representative genomes of each type of representative strain after clustering or selects representative genomes of representative bacterial species to form a representative genome library; The representative strain or representative bacterial species specific segment library construction unit obtains specific segments by specific alignment of sequencing fragments (reads) of the representative strain or representative bacterial species genomic library, and constructs the representative strain or representative bacterial species specific segment library with the obtained specific segments of each strain or species; The step of constructing a specific segment library of representative strains or representative bacterial species includes: the sequencing fragments on the specific alignment are the position segments where the sequencing fragments (reads) of each representative strain are specific to other representative strains; The step of constructing a specific segment library of a representative strain or a representative bacterial species comprises: obtaining sequencing fragments (reads) of the representative strain or bacterial species that are specific to the sequencing fragments of other representative strains or bacterial species, determining the position segment of the sequencing fragments obtained by specific alignment of the representative strain or bacterial species, and combining the specific segments of each strain or bacterial species to obtain a specific segment library of the representative strain or representative bacterial species; The bacterial species or strain identification unit is based on the specific segment library of the representative strain or representative bacterial species, and identifies the target microorganism by comparing the sequencing fragments on the specific segment library and the position segment information of the sequencing fragments on the specific segment library or using it as a biological probe.
21. The device for identifying microorganisms according to claim 20, characterized in that: The microbial gene sequence database construction unit constructs the microbial gene sequence database as claimed in claim 19.
22. The device for identifying microorganisms according to claim 20, characterized in that: The microbial genome clustering units are clustered with an ANI threshold of 95% or 99%.
23. The device for identifying microorganisms according to claim 20, characterized in that: Strains with an ANI ≥ 99% were clustered as the same strain cluster; strains with an ANI ≥ 95% were clustered as strains of the same species.
24. The device for identifying microorganisms according to claim 20, characterized in that: The representative strain or representative bacterial species selection unit constructs the representative genome library described in claim 1; The unit for constructing a library of specific segments of representative strains or representative bacterial species is to construct a library of specific segments of representative strains or representative bacterial species as claimed in claim 1; The bacterial species or strain identification unit screens out the target bacterial strain or bacterial species according to the method described in any one of claims 13-15.
25. A microbial strain identification device, characterized in that: The device comprises: one or more processors; A memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method for microorganism identification according to any one of claims 1 to 18.
26. The microorganism strain identification device according to claim 25, characterized in that: The apparatus further comprises a communication device for performing data communication; The communication device includes an input device and an output device.
27. A computer-readable storage medium comprising instructions, which, when executed on a computer, causes the computer to execute the method for microorganism identification according to any one of claims 1 to 18.
Citation Information
Patent Citations
Method and device for obtaining species-specific consensus sequences of microorganisms and application
CN111477276A
Optimized method for analyzing microbial communities through metagenome binning
CN111933218A