Method for rapidly detecting microorganisms based on next-generation sequencing
By constructing a dual-level reference database and using an iterative fitting algorithm, the problems of high computational cost and inaccurate quantification in metagenomic sequencing technology for microbial community detection are solved, achieving high-precision and rapid detection of microbial species, which is suitable for clinical applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG HONG KONG MACAO GREATER BAY AREA PRECISION MEDICINE RESEARCH INSTITUTE (GUANGZHOU)
- Filing Date
- 2025-12-22
- Publication Date
- 2026-05-01
AI Technical Summary
Existing metagenomic sequencing technologies face challenges such as high computational costs, numerous fuzzy mappings, and inaccurate quantification when dealing with microbial communities, especially at the level of high diversity at the subspecies or strain level, making it difficult to meet the needs of rapid clinical testing.
A dual-level reference database was constructed, including a pan-genome database and a strain-level reference library. Gene clusters were generated through simulated sequencing and clustering. The reference matrix and iterative fitting algorithm were used for alignment and abundance estimation, reducing computational load and improving quantitative accuracy.
It achieves high-precision and rapid detection of microbial species, meets real-time requirements, has a stable and reliable process, and has significant clinical application value.
Smart Images

Figure FT_1 
Figure FT_2 
Figure SMS_2
Abstract
Description
Rapid detection methods for microorganisms based on next-generation sequencing Technical Field
[0001] This invention relates to the field of bioinformatics, specifically to a method for rapid detection of microorganisms based on next-generation sequencing, and more specifically to a method, system, electronic device, and computer-readable storage medium for rapid detection of microbial species. Background Technology
[0002] Metagenomic sequencing technology, which can directly sequence the genetic material of all microorganisms in a sample without culturing, has become a powerful tool for pathogen detection, microbiome research, and infectious disease diagnosis. In the pathogen detection process, the core step is to accurately align the sequenced short reads to a reference genome database and calculate the species-level functional abundance, thereby achieving species identification.
[0003] Traditional mNGS analysis workflows typically rely on dynamic species identification (e.g., using tools like MetaPhlAn) and direct alignment of reads to a large reference genome. However, this approach faces significant challenges when dealing with microbial communities, especially the high diversity at the subspecies or strain level. First, the sheer size of microbial genomes makes direct alignment computationally expensive, hindering the needs of rapid clinical testing. Second, the presence of numerous homologous genes with high sequence similarity (including orthologs and paralogs) in microbial genomes leads to "multi-mapping" during alignment, where a single read may match multiple genes or species. Existing mainstream tools (such as HUMAnN) improve gene quantification by constructing species pan-genome databases, but a significant proportion of reads remain unmapped, and biases are easily introduced when allocating fuzzy reads, leading to errors in species identification and inaccurate functional abundance quantification, thus limiting their application in precision medicine.
[0004] Therefore, for large-scale sequencing data, there is still a need to develop a comprehensive, accurate, and rapid analysis method to meet the application needs of rapid detection and real-time monitoring of clinical pathogens. Summary of the Invention
[0005] This invention is based on the inventors' discovery and understanding of the following facts and problems: By constructing a two-level reference database for a fixed list of species (such as known pathogenic microorganisms) and pre-calculating gene cluster structures and design matrices reflecting gene similarity based on this database, it was found that when analyzing actual samples, only efficient read mapping and fast abundance estimation algorithms need to be executed, thereby achieving real-time and high-precision pathogen detection.
[0006] Therefore, in a first aspect, the present invention provides a method for rapid detection of microbial species. According to an embodiment of the present invention, the method includes: constructing a pan-genome database and a strain-level reference library based on the genomic sequences of multiple known microbial species, wherein the pan-genome database consists of the pan-genome genes of at least one microbial species, and the strain-level reference library consists of the coding gene sequences of at least one microbial strain, the coding genes having known functional annotations; generating a first simulated sequencing read based on the pan-genome database through a first simulated sequencing; performing a first alignment of the simulated sequencing read with the pan-genome database, and clustering the pan-genome genes to generate a pan-genome gene cluster of at least one species; generating a strain-level gene cluster corresponding to a given species based on the pan-genome gene cluster of the given species and the strain-level reference library of the given species; and generating a second simulated sequencing read based on the strain-level gene cluster through a second simulated sequencing, and aligning the second simulated sequencing read with the strain-level reference library. A second alignment is performed in the library to generate multiple first equivalence classes for a given strain based on the results of the second alignment, and to determine the contribution ratio of each gene to the reads in each first equivalence class; a reference matrix X is constructed based on the multiple equivalence classes and the contribution ratio of each gene to the reads in each equivalence class; next-generation sequencing is performed on the sample to be tested to obtain sequencing data consisting of multiple sequencing reads; the sequencing data is compared with the pan-genome gene cluster to obtain sequencing reads that match the pan-genome gene cluster; the sequencing reads that match the pan-genome gene cluster are compared with the strain-level gene cluster to generate multiple second equivalence classes for a given strain based on the results of the fourth alignment, and the number of sequencing reads falling into each second equivalence class is counted to construct a counting matrix Y; the gene abundance in the sample to be tested is determined by iterative fitting based on the counting matrix Y and the reference matrix X; and the microbial species in the sample to be tested are determined based on the gene abundance. The method proposed in this invention has high accuracy and fast analysis speed for detecting microorganisms, which can meet the real-time requirements. Moreover, the process is stable and reliable, and has important clinical application value and market prospects.
[0007] According to an embodiment of the present invention, the microbial species is a pathogen having a known gene sequence.
[0008] According to embodiments of the present invention, the genome sequence of the known microbial species is derived from at least one of the following databases: ChocoPhlAn, NCBI RefSeq, GenBank, NT database, PATRIC, and the National Pathogenic Microorganism Resource Center.
[0009] According to an embodiment of the present invention, the number of genes in the pan-genome gene cluster does not exceed 50% of the number of genes in the pan-genome database.
[0010] According to an embodiment of the present invention, both the first simulated sequencing and the second simulated sequencing adopt the following simulated sequencing conditions: coverage depth fold depth = 1~5, preferably depth = 1~3; minimum number of reads min_n_read = 200~1200, preferably min_n_read = 200~1000; read length read_len = 100~300, preferably read_len = 100~250.
[0011] According to an embodiment of the present invention, the iterative fitting is performed using the alternating EM algorithm.
[0012] According to an embodiment of the present invention, the method further includes: determining the functional abundance of microorganisms in the sample to be tested based on the gene abundance.
[0013] In a second aspect, the present invention provides a system for rapid detection of microbial species. According to an embodiment of the invention, the system comprises: a database construction unit for constructing a pan-genome database and a strain-level reference library based on the genomic sequences of multiple known microbial species, wherein the pan-genome database consists of the pan-genome genes of at least one microbial species, and the strain-level reference library consists of the coding gene sequences of at least one microbial strain, the coding genes having known functional annotations; a first simulated sequencing unit for generating a first simulated sequencing read based on the pan-genome database through a first simulated sequencing; a pan-genome gene cluster construction unit for performing a first alignment of the simulated sequencing read with the pan-genome database and clustering the pan-genome genes through clustering to generate a pan-genome gene cluster of at least one species; a strain-level gene cluster construction unit for generating a strain-level gene cluster corresponding to a given species based on the pan-genome gene cluster of the given species and the strain-level reference library of the given species; a second simulated sequencing unit for generating a second simulated sequencing read based on the strain-level gene cluster through a second simulated sequencing; and a reference matrix X construction unit for constructing a reference matrix X. A second alignment is performed between simulated sequencing reads and a strain-level reference library. Based on the results of this second alignment, multiple first equivalence classes for a given strain are generated, and the contribution ratio of each gene to each first equivalence class is determined. A reference matrix X is constructed based on the multiple equivalence classes and the contribution ratio of each gene to each equivalence class. A sequencing unit is used to perform next-generation sequencing on the sample to be tested to obtain sequencing data consisting of multiple sequencing reads. A alignment unit is used to perform a third alignment between the sequencing data and the pan-genome gene cluster to obtain the correlation between the sequencing data and the pan-genome gene cluster. The system comprises: a matching sequencing read unit; a counting matrix Y construction unit, used to perform a fourth alignment between the sequencing reads matched by the pan-genome gene cluster and the strain-level gene cluster, to generate multiple second equivalence classes for a given strain based on the results of the fourth alignment, and to count the number of sequencing reads falling into each second equivalence class to construct the counting matrix Y; an iterative fitting unit, used to determine the gene abundance in the test sample through iterative fitting based on the counting matrix Y and the reference matrix X; and a species determination unit, used to determine the microbial species in the test sample based on the gene abundance. The system proposed in this invention has high detection accuracy and fast analysis speed for microorganisms, can meet real-time requirements, and has a stable and reliable process, possessing significant clinical application value and market prospects.
[0014] In a third aspect, the present invention provides an electronic device. According to an embodiment of the invention, the electronic device includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program to implement the system described above. The electronic device proposed by the present invention is capable of efficiently detecting and identifying microorganisms.
[0015] In a fourth aspect, the present invention provides a computer-readable storage medium. According to an embodiment of the invention, the storage medium stores computer program instructions that, when executed on a processor, cause the processor to perform the system as described above. The storage medium proposed by the present invention enables automated and efficient detection and identification of microorganisms, and the detection results are reliable.
[0016] The beneficial effects of this invention are at least as follows: This invention proposes a method for rapid detection of microbial species. The method has fast analysis speed, high accuracy to meet real-time requirements, and stable and reliable process, which has important clinical application value and market prospects. Attached Figure Description
[0017] Figure 1 is a reference data example diagram according to an embodiment of the present invention.
[0018] Figure 2 is an example diagram of the output results according to an embodiment of the present invention. Detailed Implementation
[0019] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0020] The endpoints and any values of the ranges disclosed herein are not limited to the precise ranges or values, and these ranges or values should be understood to include values close to these ranges or values. For numerical ranges, the endpoint values of the various ranges, the endpoint values of the various ranges and individual point values, and individual point values can be combined with each other to obtain one or more new numerical ranges, which should be considered as specifically disclosed herein.
[0021] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0022] To facilitate understanding of this invention, certain technical and scientific terms are specifically defined below. Unless explicitly defined elsewhere in this document, all other technical and scientific terms used herein have the meanings commonly understood by one of ordinary skill in the art to which this invention pertains. In the description of this invention, the terms used herein have been explained and described; these explanations and descriptions are merely for the purpose of facilitating understanding and should not be construed as limiting the scope of protection of this invention.
[0023] In this paper, the term "pan-genome" refers to the sum of genomic information of all individuals within a species, including the core genome and the accessory genome. The core genome is the set of genes present in all tested strains / individuals and typically performs housekeeping functions (such as replication, transcription, translation, and basal metabolism). It is relatively conserved in number and represents the "essential genetic backbone" of the species. The accessory genome is the set of genes present only in some strains, including plasmids, phage fragments, transposons, resistance islands, virulence islands, and secondary metabolic synthesis clusters. Its content rapidly increases or decreases with changes in ecological niche or selective pressure and is the main genetic source of phenotypic differences (pathogenicity, drug resistance, environmental adaptation) among strains.
[0024] In this article, the terms "metagenomic sequencing" or "mNGS" refer to the use of high-throughput sequencing technology to obtain the sequence data of all microbial genomes in a sample, and then using bioinformatics analysis methods to perform sequence alignment to identify the types and abundance of microorganisms.
[0025] In this article, the term "equivalence class" or "eqclass" refers to a group of highly similar genes matched by simulated reads.
[0026] In this article, the term "gene cluster" refers to a set of genes that are physically adjacent in the genome and are functionally related.
[0027] In this article, the term "mapping" refers to the process of locating reads generated by sequencing onto a reference genome using sequence alignment algorithms, and determining their most likely origin.
[0028] In this article, the term "paralogous gene" refers to a homologous gene that is produced within the genome of the same species through a gene duplication event. They share a common ancestral gene, but their sequences and functions may diverge due to independent evolution within the genome after duplication.
[0029] In this article, the terms "next-generation sequencing" or "short-read sequencing" refer to large-scale sequencing technologies with short read lengths, high throughput, and high precision.
[0030] The technical solution of this application is described in detail below: A method for rapid detection of microbial species is proposed in some embodiments of this invention. The method proposed in this invention has high detection accuracy and fast analysis speed for microorganisms, meeting real-time requirements, and the process is stable and reliable, possessing significant clinical application value and market prospects.
[0031] According to an embodiment of the present invention, the method includes: constructing a pan-genome database and a strain-level reference library based on the genomic sequences of multiple known microbial species, wherein the pan-genome database consists of pan-genome genes of at least one microbial species, and the strain-level reference library consists of coding gene sequences of at least one microbial strain, wherein the coding genes have known functional annotations. The present invention, by constructing a dual-level reference database (pan-genome + strain-level genome) containing strain-level information, can more comprehensively capture microbial genetic diversity and improve the detection sensitivity and resolution for specific pathogen strains. Based on the pan-genome database, a first simulated sequencing is performed to generate a first simulated sequencing read; the simulated sequencing read is first aligned with the pan-genome database, and the pan-genome genes are clustered to generate pan-genome gene clusters of at least one species. By clustering the pan-genome into gene clusters, the number of genes can be significantly reduced, thereby reducing the computational load of subsequent analysis. For a given species, based on the pan-genome gene clusters of the given species and a strain-level reference library of the given species, a strain-level gene cluster corresponding to the given species is generated. Based on the strain-level gene clusters, a second simulated sequencing read is generated through a second simulated sequencing process. The second simulated sequencing read is then compared with the strain-level reference library in a second alignment. Based on the results of the second alignment, multiple first equivalence classes of the given strain are generated, and the contribution ratio of each gene to each first equivalence class is determined. Based on the multiple equivalence classes and the contribution ratio of each gene to each equivalence class, a reference matrix X is constructed. The reference matrix X is pre-calculated and is used to describe the co-abundance relationship between gene families, and can be used for all subsequent sample analyses. Next-generation sequencing is performed on the sample to be tested to obtain sequencing data consisting of multiple sequencing reads. The sequencing data is aligned a third time with the pan-genome gene cluster to obtain sequencing reads that match the pan-genome gene cluster. The sequencing reads matching the pan-genome gene cluster are then aligned a fourth time with the strain-level gene cluster. Based on the results of the fourth alignment, multiple second equivalence classes for a given strain are generated, and the number of sequencing reads falling into each second equivalence class is counted to construct a counting matrix Y. Based on the counting matrix Y and the reference matrix X, the gene abundance in the test sample is determined through iterative fitting. And based on the gene abundance, the microbial species in the test sample are determined. This invention employs a probabilistic allocation mechanism (through designing matrix X and iterative fitting) to handle the fuzzy mapping problem caused by homologous genes, avoiding the bias introduced by forced single allocation. The association between reads and multiple homologous genes is accurately quantified through equivalence classes and normalized weights, significantly improving quantitative accuracy.
[0032] According to embodiments of the present invention, to process paralogous genes, after obtaining the final gene abundance, a clustering algorithm (such as k-means) is used to identify gene sets with highly correlated abundance, and their abundances are summed as the overall abundance of a "paralogous group" for reporting. Finally, the species-level functional abundance is output, and microbial species identification is completed simultaneously. By introducing a mechanism for merging and reporting paralogous genes, the abundance estimation bias caused by gene duplication events can be further reduced.
[0033] According to embodiments of the present invention, the microbial species is a pathogen with a known gene sequence. The pathogens include intestinal microorganisms, oral microorganisms, environmental microorganisms, rhizosphere microorganisms, etc. It should be noted that this invention uses pathogens as an example and should not be construed as limiting the invention; the method proposed in this invention is also applicable to the detection of other microbial groups.
[0034] According to embodiments of the present invention, the genome sequence of the known microbial species is derived from at least one of the following databases: ChocoPhlAn, NCBI RefSeq, GenBank, NT database, PATRIC, and the National Pathogenic Microorganism Resource Center.
[0035] According to an embodiment of the present invention, the number of genes in the pan-genome gene cluster does not exceed 50% of the number of genes in the pan-genome database. By clustering the pan-genome into functionally similar gene clusters using a clustering algorithm, the number of genes can be significantly reduced, thereby reducing the computational load of subsequent data analysis.
[0036] According to an embodiment of the present invention, the simulated reads are randomly generated using the simx tool. The random generation includes: random fragment extraction, random orientation, and no functional bias. The random fragment extraction involves randomly selecting a starting position from each reference sequence (gene) and extracting fragments of normally distributed length. The random orientation ensures that each fragment has a 50% probability of reverse complementation (simulating the randomness of positive and negative strands in sequencing). The lack of functional bias means that all genes are generated with reads according to a uniform coverage standard, without considering differences in gene function or expression levels.
[0037] According to an embodiment of the present invention, minimap2 is used as the alignment tool when constructing the reference matrix X. For controlled simulated short reads, minimap2 is fast and flexible in short read mode, making it suitable for quickly completing self-alignment and generating equivalence classes. Bowtie2 is used as the alignment tool when constructing the counting matrix Y. Real samples need to be captured with high sensitivity for multiple alignments. The full output of bowtie2's `--very-sensitive -a` matches the equivalence class statistical process of bam2eqv, avoiding the underestimation of equivalence classes due to the absence of suboptimal alignments.
[0038] According to an embodiment of the present invention, both the first simulated sequencing and the second simulated sequencing adopt the following simulated sequencing conditions: coverage depth fold depth = 1~5, preferably depth = 1~3; minimum number of reads min_n_read = 200~1200, preferably min_n_read = 200~1000; read length read_len = 100~300, preferably read_len = 100~250.
[0039] According to embodiments of the present invention, the coverage depth multiple is depth=1~5. Exemplarily, the depth=1, 2, 3, 4 or 5, or any range between any two of the above values. According to some preferred embodiments of the present invention, the depth=1~3. The depth>5 usually does not provide significant benefits, but instead increases the computational burden.
[0040] According to embodiments of the present invention, the minimum number of read segments min_n_read = 200~1200. For example, the minimum number of read segments min_n_read = 200, 400, 600, 800, 1000 or 1200, or any range between any two of the above values. According to some preferred embodiments of the present invention, min_n_read = 200~1000; this ensures that even short sequences have a sufficient number of read segments to avoid insufficient coverage.
[0041] According to embodiments of the present invention, the read length read_len is 100~300. For example, the read length read_len is 100, 150, 200, 250 or 300, or any range between any two of the above values. According to some preferred embodiments of the present invention, the read_len is 100~250. The simulated sequencing read length can affect the alignment specificity.
[0042] According to embodiments of the present invention, the simulated sequencing conditions further include frag_len_mean=250 and frag_len_std=25; frag_len_mean is the mean fragment length, reflecting the expected value of the library fragment size; frag_len_std is the standard deviation of fragment length, reflecting the variability of the library fragment length. The generation of simulated reads is based on a dynamic length generation strategy and a minimum read threshold setting strategy, which dually ensures the sufficiency of the simulated read coverage. According to some specific embodiments of the present invention, the dynamic length generation strategy includes generating depth×length reads for long sequences (default 1x coverage), and forcing the generation of at least min_n_read=200 reads for short sequences, so that short sequences obtain higher coverage and ensure that small genes can also be fully represented.
[0043] According to an embodiment of the present invention, the core algorithm logic of the simulated read segment is as follows: 1 void simulate_read(int&count_read, const auto&ref_seq, const std::string&ref_name, 2 std::mt19937&norm_gen, std::default_random_engine&uniform_gen, 3 auto&out, 4 size_t read_len = 100, int depth = 1, int min_n_read = 200, 5 double frag_len_mean = 250, double frag_len_std = 25) According to an embodiment of the present invention, the algorithm steps of the simulated read segment are as follows: (1) Calculate the target number of read segments: n_seq = depth × seq_len, if it is less than the minimum value, then take min_n_read; (2) Generate read segments in a loop until the target number is reached; (3) Each iteration: from the normal distribution $N(\mu=250, Extract the segment length from $\sigma=25); randomly select the starting position of the segment (uniform distribution); randomly decide whether to reverse complement (Bernoulli distribution, p=0.5); extract a segment of length read_len and output it.
[0044] According to an embodiment of the present invention, a maximum of 5 mismatches are allowed when generating the equivalence class to control the comparison tolerance.
[0045] According to an embodiment of the present invention, the filtering parameters when constructing the reference matrix X are as follows: H_thres=0.025, targetsize=5000; H_thres is a high threshold used to filter low-count connections between transcripts and remove noise; targetsize is the target matrix size used to control the CRP clustering scale.
[0046] H_thres_min = 0.01~0.05. For example, H_thres_min = 0.01, 0.015, 0.02, 0.025, 0.03, 0.035, 0.04, 0.045 or 0.05, or any range between any two of the above values. According to some preferred embodiments of the present invention, H_thres_min = 0.02~0.05; H_thres_min is a minimum threshold used to exclude equivalence classes with excessively low read contribution (<1%). Too high a threshold will result in the loss of low-abundance genes, while too low a threshold will introduce noise.
[0047] According to an embodiment of the present invention, the filtering algorithm logic constructed by the reference matrix X is specifically as follows: 1# Calculate the contribution ratio of each read segment to the transcript 2rawmat$prop = rawmat$Weight / rawmat$txTrueCount 3rawmat$propEq = rawmat$Count / rawmat$txTrueCount 5# Exclude the equivalence classes with contribution degree <H_thres_min 6pick = which(minProp>H_thres_min | minPropEq>H_thres_min) According to an embodiment of the present invention, on the first equivalence class side, gentc (fast aggregation for low-noise simulated BAM) is used, and on the second equivalence class side, bam2eqv (supporting filtering such as nm_nonsharing to enhance strain-level specificity) is used.
[0048] According to a specific embodiment of the present invention, the iterative fitting is performed using the alternating EM algorithm. Based on the model Y Xβ, where Y is the counting matrix Y as described above, X is the pre-computed reference matrix X as described above, and β is the gene / IGC abundance matrix to be estimated. The alternating expectation maximization (AEM) algorithm is used to estimate β: E-step: Fix the currently estimated abundance β and calculate the posterior probability that each read segment belongs to each gene: <000,0104>
[0049] 0Where, is the matching weight between read segment r and gene g.
[0050] M-step: Fix the read segment attribution probability distribution and update the abundance β to maximize the likelihood function (assuming Y follows a Poisson distribution):
[0051] Iteratively execute the E-step and M-step until convergence (for example, the change in β is less than 1e-6).
[0052] According to an embodiment of the present invention, the method further includes: determining the microbial functional abundance in the test sample based on the gene abundance.
[0053] In some embodiments of the present invention, a system for rapid detection of microbial species is proposed. According to embodiments of the present invention, the system includes: a database construction unit, a first simulated sequencing unit, a pan-genome gene cluster construction unit, a strain-level gene cluster construction unit, a second simulated sequencing unit, a reference matrix X construction unit, a sequencing unit, an alignment unit, a counting matrix Y construction unit, an iterative fitting unit, and a species determination unit. The system proposed in this invention offers high accuracy and fast analysis speed for microbial detection, meets real-time requirements, and has a stable and reliable process, possessing significant clinical application value and market potential.
[0054] According to embodiments of the present invention, the database construction unit is used to construct a pan-genome database and a strain-level reference library based on the genomic sequences of multiple known microbial species. The pan-genome database consists of the pan-genome genes of at least one microbial species, and the strain-level reference library consists of the coding gene sequences of at least one microbial strain, wherein the coding genes have known functional annotations. The database contains a dual-level reference database (pan-genome + strain-level genome) containing strain-level information, which can more comprehensively capture microbial genetic diversity and improve the detection sensitivity and resolution of specific pathogen strains.
[0055] According to an embodiment of the present invention, the first simulated sequencing unit is used to generate a first simulated sequencing read based on the pan-genome database through a first simulated sequencing.
[0056] According to an embodiment of the present invention, the pan-genome gene cluster construction unit is used to perform a first alignment of the simulated sequencing read with the pan-genome database, and to cluster the pan-genome genes to generate pan-genome gene clusters of at least one species, so as to reduce the number of genes and thus reduce the amount of subsequent analysis and computation.
[0057] According to an embodiment of the present invention, the strain-level gene cluster construction unit is used to generate a strain-level gene cluster corresponding to a given species, based on the pan-genome gene cluster of the given species and the strain-level reference library of the given species.
[0058] According to an embodiment of the present invention, the second simulated sequencing unit is used to generate a second simulated sequencing read based on the strain-level gene cluster through a second simulated sequencing.
[0059] According to an embodiment of the present invention, the reference matrix X construction unit is used to perform a second alignment between the second simulated sequencing read and a strain-level reference library, so as to generate multiple first equivalence classes for a given strain based on the results of the second alignment, determine the read contribution ratio of each gene to each first equivalence class, and construct a reference matrix X based on the multiple equivalence classes and the read contribution ratio of each gene to each equivalence class. The reference matrix X is pre-calculated and can be used for all subsequent sample analyses.
[0060] According to an embodiment of the present invention, the sequencing unit is used to perform next-generation sequencing on the sample to be tested to obtain sequencing data consisting of multiple sequencing reads.
[0061] According to an embodiment of the present invention, the alignment unit is used to perform a third alignment of the sequencing data with the pan-genome gene cluster to obtain sequencing reads that match the pan-genome gene cluster.
[0062] According to an embodiment of the present invention, the counting matrix Y construction unit is used to perform a fourth alignment between the sequencing reads matched by the pan-genome gene clusters and the strain-level gene clusters, so as to generate multiple second equivalence classes for a given strain based on the results of the fourth alignment, and count the number of sequencing reads falling into each second equivalence class to construct the counting matrix Y.
[0063] According to an embodiment of the present invention, the iterative fitting unit is used to determine the gene abundance in the test sample by iterative fitting based on the counting matrix Y and the reference matrix X; and according to an embodiment of the present invention, the species determination unit is used to determine the microbial species in the test sample based on the gene abundance. The present invention employs a probabilistic allocation mechanism (through design matrix X and iterative fitting) to handle the fuzzy mapping problem caused by homologous genes, avoiding the bias introduced by forced single allocation. The association between reads and multiple homologous genes is accurately quantified through equivalence classes and normalized weights, significantly improving quantitative accuracy.
[0064] According to embodiments of the present invention, the present invention, based on a fixed reference database and analysis process, ensures consistency and repeatability between different analyses; through Docker containerization, deployment is simplified and the environment is unified, avoiding runtime errors caused by system differences, realizing one-click automated analysis, and lowering the user threshold.
[0065] In some embodiments of the present invention, an electronic device is proposed. According to embodiments of the present invention, it includes a memory and a processor; the memory is used to store a computer program; the memory, used to store the computer program, is any device capable of retaining data, such as a hard disk drive, a high-density disk read-only memory (CD-ROM), a DVD, or a solid-state storage device. The memory retains instructions and data used by the processor. The pointing device may be a mouse, a trackball, or other type of pointing device, and is used in conjunction with a keyboard to input data into the computer system. A graphics adapter displays images and other information on the display. A network adapter couples the computer system to a local area network (LAN) or a wide area network (WAN). The processor executes the computer program to implement the system described above. The electronic device proposed in this invention can efficiently detect and identify microorganisms.
[0066] In some embodiments of the present invention, a computer-readable storage medium is proposed. According to embodiments of the present invention, the storage medium stores computer program instructions that, when executed on a processor, cause the processor to perform the system as described above. When executed by the one or more processors, the instructions cause the electronic device to perform operations. These operations include providing sequencing data of a microbial sample to be detected, using the sequencing data as input, and outputting gene abundance and species identification of the microorganism. The storage medium proposed in this invention enables automated and efficient detection and identification of microorganisms, and the detection results are reliable.
[0067] Embodiments of the present invention will now be described in more detail, examples of which are illustrated in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the invention, and should not be construed as limiting the invention.
[0068] Example 1: Construction of Reference Matrix X. Input of reference data: Data was extracted from the NCBI and Chocophlan databases respectively, and the reference data was created using Megafun as shown in Table 1.
[0069] Table 1:
[0070] Wherein, taxid is the classification identifier; file_name is the file name, which is a compressed file storing the gene sequence data of the species, the core data carrier; gene_num is the number of genes; species_infos is the species information string; k~s are kingdom, phylum, class, order, family, genus and species respectively.
[0071] The reference data storage format is shown in Figure 1, and the specific input is as follows:
[0072] Merging pan-genome sequences:
[0073] Create a comparison index:
[0074] Simulated read generation (simx tool): 1# Generate simulated reads for each pangenome sequence 2. ` / bin / simx REF_DATA / SP1 / pangenome -o SP1_X_matrix / temp -t 8` 3# Parameters: depth=1, min_n_read=200, read_len=100, frag_len~N(250,25)` Alignment to generate BAM:
[0075] Generate equivalence classes (gentc tool):
[0076] Construct the CRP matrix (buildCRP_para.R):
[0077] Log output example:
[0078] The X matrix is the core design matrix of the MegaDock pipeline, used to describe the co-abundance relationship (CRP) between gene families in a pan-genome. The final output files are final_Y.RData (pan-genome abundance matrix), final_singleton.RData (strain-level gene abundance matrix), and final_paralog.RData (homologous / parallel gene abundance matrix).
[0079] The final_Y.RData contains Y: a list of length 11112 (corresponding to the number of CRPs); f: auxiliary variables (file paths, etc.). This matrix can be used for cluster-level differential analysis, summing by column (samples) to obtain the total abundance of clusters, and screening high-weight equivalence classes to identify major contribution patterns.
[0080] The structure of the Y object:
[0081] Output example:
[0082] To check the abundance of a single CRP:
[0083] Output example:
[0084] Explanation: List elements: Each element corresponds to a CRP (gene cluster); Rows: Equivalence class encoding (binary mode, representing the gene combination matched by reads); Columns: Sample names (e.g., test_1ex, test_2ex, or sample1, sample2); Values: Normalized abundance (proportion) after AEM correction, reflecting the relative abundance of the equivalence class in the sample; Source flow: create_count_matrix.R: Summarizes the raw counts from the sample's eqClass.txt; AEM_update_X_beta.R: Performs AEM (Abundance Estimation Method) correction using the X matrix, and deconvolves to obtain the final abundance.
[0085] The final_singleton.Rdata file contains final_singleton: a matrix (71630 × 2); f: an auxiliary variable. It can be used as a baseline control (the abundance of single-copy genes is generally more stable); for differential expression analysis; and for comparison with paralog genes to verify deconvolution effects.
[0086] View the structure:
[0087] Output example:
[0088] Meaning explanation: Row name: UniRef ID of single-copy gene (species | UniRef90 number); Column name: Sample name (test_1ex, test_2ex, etc.); Value: Abundance estimate of the gene in each sample (count or normalized value); Source: Summary of single-copy portions of AEM and create_count_matrix. These genes have only a unique mapping in the pan-genome and do not require complex deconvolution.
[0089] The final_paralog.RData contains beta.all: a matrix (34747 × 2); f: an auxiliary variable that can be used to obtain accurate quantification of homologous genes (to solve the problem of multiple alignment), to obtain the genome-wide abundance spectrum after merging with singleton, and to consider collinearity of homologous genes during differential analysis.
[0090] View the structure:
[0091] Output example:
[0092] Meaning explanation: Row name: UniRef ID of the homologous gene (paralog); Column: Sample (consistent with the column name of final_singleton); Value: Abundance coefficient (beta value) after AEM deconvolution, representing the true abundance estimate of the gene after considering homology.
[0093] Source: AEM_isolate.R or AEM_update_X_beta.R; For multicopy / homologous gene families, the beta coefficient is solved using the relationship between the X matrix and the Y matrix (Y = X·β).
[0094] After successfully constructing the reference matrix X, you only need to directly call the execution command shown below:
[0095] Example 2: Construction of a method for rapid detection of microbial species. Reference data preparation is the same as in Example 1.
[0096] The test comprehensively analyzed 16 million 75bp sequence data points. Utilizing 10 CPU cores, the test validated the system's efficiency and accuracy for four gold standard species: Bacillus Subtilis, Enterococcus faecalis, Limosilactobacillus Fermentum, and Listeria Monocytogenes.
[0097] The specific process is as follows:
[0098] Construction of reference matrix X:
[0099] Construct the Y matrix and complete the quantification:
[0100] View results:
[0101] The software outputs a ".md" report file under the ". / OUTPUT" path and a ".log" log file under the ". / LOG" path. The detection report file is saved in the user-defined path ". / OUTPUT". This markdown file contains the identification results from the Paralog and Singleton databases. An example of the results is shown in Figure 2, with four columns of information for each column: "Species": Species category name; "test_1_part1": Number of genes detected in the first FastQ file; "test_1_part2": Number of genes detected in the second FastQ file; "Total": Total number of genes detected (test_1_part1 + test_1_part2). The results show that the overall runtime is 35 minutes, demonstrating the high efficiency of the proposed method in real-time analysis of large-scale sequencing data and providing a reliable solution for applications such as pathogen detection.
[0102] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0103] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for rapid detection of microbial species, characterized in that, include: A pan-genome database and a strain-level reference library are constructed based on the genome sequences of multiple known microbial species. The pan-genome database consists of the pan-genome genes of at least one microbial species, and the strain-level reference library consists of the coding gene sequences of at least one microbial strain, wherein the coding genes have known functional annotations. Based on the pan-genome database, a first simulated sequencing is performed to generate a first simulated sequencing read. The simulated sequencing read is first aligned with the pan-genome database, and the pan-genome genes are clustered to generate a pan-genome gene cluster of at least one species. For a given species, a strain-level gene cluster corresponding to the given species is generated based on the pan-genome gene cluster of the given species and the strain-level reference library of the given species. Based on the strain-level gene clusters, a second simulated sequencing read is generated through a second simulated sequencing process. The second simulated sequencing read is then compared with a strain-level reference library. Based on the results of the second comparison, multiple first equivalence classes for a given strain are generated, and the contribution ratio of each gene to the reads in each first equivalence class is determined. Based on the multiple equivalence classes and the contribution ratio of each gene to the reads in each equivalence class, a reference matrix X is constructed. Next-generation sequencing is performed on the sample to be tested to obtain sequencing data consisting of multiple sequencing reads; the sequencing data is then aligned with the pan-genome gene cluster to obtain sequencing reads that match the pan-genome gene cluster; the sequencing reads that match the pan-genome gene cluster are then aligned with the strain-level gene cluster to generate multiple second equivalence classes for a given strain based on the results of the fourth alignment, and the number of sequencing reads falling into each second equivalence class is counted to construct a counting matrix Y; based on the counting matrix Y and the reference matrix X, the gene abundance in the sample to be tested is determined by iterative fitting. Based on the gene abundance, the microbial species in the sample to be tested are determined.
2. The method according to claim 1, characterized in that, The microbial species are pathogens with known gene sequences.
3. The method according to claim 1, characterized in that, The genome sequences of the known microbial species are derived from at least one of the following databases: ChocoPhlAn, NCBI RefSeq, GenBank, NT database, PATRIC, and the National Pathogenic Microorganism Resource Center.
4. The method according to claim 1, characterized in that, The number of genes in the pan-genome gene cluster does not exceed 50% of the number of genes in the pan-genome database.
5. The method according to claim 1, characterized in that, Both the first and second simulated sequencing use the following simulated sequencing conditions: coverage depth factor depth = 1~5, preferably depth = 1~3; minimum number of reads min_n_read = 200~1200, preferably min_n_read = 200~1000; read length read_len = 100~300, preferably read_len = 100~250.
6. The method according to claim 1, characterized in that, The iterative fitting was performed using the alternating EM algorithm.
7. The method according to claim 1, characterized in that, The method further includes: determining the functional abundance of microorganisms in the sample to be tested based on the gene abundance.
8. A system for rapid detection of microbial species, characterized in that, include: A database construction unit is used to construct a pan-genome database and a strain-level reference library based on the genome sequences of multiple known microbial species. The pan-genome database consists of the pan-genome genes of at least one microbial species, and the strain-level reference library consists of the coding gene sequences of at least one microbial strain, wherein the coding genes have known functional annotations. A first simulated sequencing unit is used to generate a first simulated sequencing read based on the pan-genome database through a first simulated sequencing. A pan-genome gene cluster construction unit is used to perform a first alignment of the simulated sequencing read with the pan-genome database and to cluster the pan-genome genes through clustering to generate a pan-genome gene cluster of at least one species. A strain-level gene cluster construction unit is used to generate a strain-level gene cluster corresponding to a given species, based on the pan-genome gene cluster of the given species and the strain-level reference library of the given species. The second simulated sequencing unit is used to generate a second simulated sequencing read based on the strain-level gene clusters through second simulated sequencing. The reference matrix X construction unit is used to perform a second alignment between the second simulated sequencing read and the strain-level reference library, so as to generate multiple first equivalence classes for a given strain based on the results of the second alignment, determine the read contribution ratio of each gene to each first equivalence class, and construct the reference matrix X based on the multiple equivalence classes and the read contribution ratio of each gene to each equivalence class; the sequencing unit is used to perform second-generation sequencing on the sample to be tested to obtain sequencing data consisting of multiple sequencing reads; The alignment unit is used to perform a third alignment between the sequencing data and the pan-genome gene cluster to obtain sequencing reads that match the pan-genome gene cluster; the counting matrix Y construction unit is used to perform a fourth alignment between the sequencing reads that match the pan-genome gene cluster and the strain-level gene cluster, so as to generate multiple second equivalence classes for a given strain based on the results of the fourth alignment, and count the number of sequencing reads that fall into each second equivalence class to construct the counting matrix Y; An iterative fitting unit is used to determine the gene abundance in the test sample by iterative fitting based on the counting matrix Y and the reference matrix X. A species determination unit is used to determine the microbial species in the sample to be tested based on the gene abundance.
9. An electronic device, characterized in that, It includes a storage unit and a processor; the storage unit is used to store a computer program; the processor is used to execute the computer program to implement the system of claim 8.
10. A computer-readable storage medium, characterized in that, The storage medium stores computer program instructions that, when executed on a processor, cause the processor to perform the system as described in claim 8.