A method for bulk detection of viruses in single-cell microbial genomes
By constructing a database with comprehensive and independent reference indexes, and combining a two-level detection strategy with quadruple table analysis, the problems of low flexibility and throughput in single-cell microbial genome detection were solved, achieving efficient virus/plasmid detection and correlation analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MOBIDROP (ZHEJIANG) CO LTD
- Filing Date
- 2026-03-31
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies are not flexible enough for single-cell microbial genome detection, have low throughput, and can only detect the correlation with one bacteriophage per run.
A database containing a comprehensive reference index and independent reference indexes was constructed. A two-level detection strategy was adopted: a first threshold was used to screen for SAGs, and a second threshold was used to verify the integrity of the virus/plasmid sequence. The significant association between the target strain and the virus/plasmid was determined by combining a quadruple table and statistical tests.
It achieves high-throughput, automated virus/plasmid detection, enabling simultaneous detection of all viruses/plasmids in a sample and performing correlation analysis, thus improving detection efficiency.
Smart Images

Figure CN121938472B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics and relates to a method for processing single-cell sequencing data, and more specifically to a method for batch detection of viruses in single-cell microbial genomes. Background Technology
[0002] The relationship between microorganisms and humans is complex and close. Microorganisms can be pathogens, causing diseases, such as bacteria and viruses. However, most microorganisms are beneficial to humans, playing important roles in food production (such as fermentation), medicine (such as antibiotic production), and environmental protection (such as wastewater treatment). Different microorganisms exert different biological functions through their different gene sequences. Sequencing techniques now allow for large-scale detection of microbial genes, leading to a deeper understanding of their mechanisms of action and significance.
[0003] Single-cell microbial genome sequencing (microbe-seq) is a technique that uses single-cell sequencing to detect the complete gene sequence of a single microorganism in high throughput. The technical article was published in Science in 2022 (PMID: 35653470). This technology can efficiently distinguish similar species, differentiate strains through mutations in a single microorganism, detect microbial-viral interactions, and discover horizontal gene transfer events within microbial communities. Compared to previous microbial genome sequencing technologies, single-cell microbial genome sequencing offers advantages such as high throughput, high detection resolution, and the ability to distinguish between strains.
[0004] The aforementioned article describes a method for detecting bacteriophages (viruses) and calculating their association with species or strains based on single-cell microbial sequencing data. The specific steps include: constructing a reference index for the bacteriophage using bowtie2; aligning all reads of the SAG (single-amplified genome, a collection of sequencing sequences from a single microorganism within a droplet) to this reference index; if the alignment rate is greater than 5%, the SAG is considered to contain the bacteriophage; and combining the SAG's presence with the target species / strain, using Fisher's exact test to calculate whether the target species / strain is significantly associated with the bacteriophage. However, this method still has the following problems: 1) the type of virus contained in the sample must be known; 2) the virus must have a known reference sequence; and 3) only one correlation with a single bacteriophage can be detected per run. Summary of the Invention
[0005] To address the issues of low flexibility and low throughput in existing technologies, this invention provides a method for high-throughput and automated detection of the presence of viruses / plasmids that are significantly related to a target species / strain in single-cell microbial genome data.
[0006] The technical solution adopted in this invention is: a method for batch detection of viruses in the genome of single-celled microorganisms, comprising the following steps:
[0007] S1. Obtain the genome sequences of viruses / plasmids from the database and standardize the sequences to construct a reference database that can be used for sequence alignment; the database includes known databases or assembled databases; the reference database includes a comprehensive reference index containing all virus / plasmid genome sequences in the database and an independent reference index for each virus / plasmid genome sequence in the database;
[0008] S2. Align the reads of each SAG to the comprehensive reference index obtained in step S1 and calculate the overall alignment rate; based on a preset first threshold, select SAGs with an overall alignment rate greater than the first threshold as primary screening SAGs containing viruses / plasmids; calculate the number of reads in each viral / plasmid genome contig aligned in each primary screening SAG, determine the plasmid / virus with the most read alignments as the primary virus / plasmid, then align the reads of the primary screening SAG to the independent reference index of the primary virus / plasmid and calculate the specific alignment rate; based on a preset second threshold, select primary screening SAGs with a specific alignment rate to overall alignment rate ratio greater than the second threshold as secondary screening SAGs containing the primary virus / plasmid; generate a corresponding Bin containing the secondary screening SAG for each primary plasmid / virus;
[0009] S3. For each Bin, construct a quadrant table based on the strain classification information of SAG and the major virus / plasmid determined in step S2, and perform statistical tests to determine whether there is a significant association between the target strain and the target virus / plasmid, and annotate it.
[0010] This invention provides a high-throughput and automated method for detecting viruses in single-celled microbial genomes. The method first constructs a reference database containing a comprehensive reference index and independent reference indices. The database can be sourced from large databases (e.g., Refseq) or known databases from other omics experiments, or from assembled databases containing viral / plasmid gene sequences (e.g., metagenomics or the single-celled microbial genome itself). This solves the problem in existing technologies that require knowledge of the viruses contained in the sample and the existence of known reference sequences for each virus, resulting in a unique virus / plasmid database. A two-stage detection strategy is then employed: first, each SAG read is aligned with the comprehensive reference index for high-throughput initial screening based on a first threshold; then, the virus / plasmid most frequently aligned in the initially positive SAGs is calculated as the primary virus / plasmid, and a second alignment is performed based on the independent reference index of this primary virus / plasmid. A second threshold is used to verify the integrity of the virus / plasmid sequence, effectively eliminating fragment interference. This achieves high-throughput and automated detection of viruses / plasmids and can detect unknown virus or plasmid sequences from the assembly results for subsequent association analysis. Finally, by constructing a quadruple table between the target strain and the target virus / plasmid, statistical tests are performed to output significant association results and annotation information of the associated target virus / plasmid.
[0011] Preferably, step S1 includes: obtaining the genome sequence of the virus / plasmid from a known database; the known database includes a separate FASTA file, a basic annotation file, and other annotation files for each virus / plasmid genome sequence to be constructed; constructing an independent reference index for each virus / plasmid based on the separate FASTA file; merging all virus / plasmid genome sequences to obtain a merged FASTA file, and constructing a comprehensive reference index for all viruses / plasmids based on the merged FASTA file.
[0012] Preferably, the individual FASTA files are named with the ID of the plasmid / virus, the basic annotation file contains the ID and name of the plasmid / virus, and the other annotation files contain the ID of the plasmid / virus; the name of each contig in the merged FASTA file contains the ID and name of its corresponding plasmid / virus; the basic annotation file and other annotation files are merged according to the ID of the plasmid / virus.
[0013] Preferably, step S1 includes: obtaining the genome sequence of the virus / plasmid from the assembly database; the assembly database contains assembly results in FASTA format and assembly result annotations; screening candidate contigs from the assembly results; annotating the candidate contigs; evaluating the integrity and contamination of the candidate contigs and giving a quality rating based on a preset quality threshold, the quality rating including high quality, medium quality, and low quality; for candidate contigs with the same annotation results, only the candidate contig with the highest integrity is retained as a reference contig; constructing an independent reference index for each virus / plasmid based on the reference contigs with a quality rating that is not low quality; merging all reference contigs with a quality rating that is not low quality to obtain a merged FASTA file, and constructing a comprehensive reference index for all viruses / plasmids based on the merged FASTA file.
[0014] Preferably, the annotation file for the candidate contig includes the ID and name of the plasmid / virus to which the candidate contig belongs; the name of each contig in the merged FASTA file contains the ID and name of its corresponding plasmid / virus; and the annotation files are merged according to the ID of the plasmid / virus.
[0015] Preferably, in step S2, the first threshold is 1% to 50%, and the second threshold is 50% to 90%.
[0016] Preferably, the quadruple table in step S3 includes the following values: the number of SAGs that contain both the target strain and the target virus / plasmid, the number of SAGs that contain the target virus / plasmid but not the target strain, the number of SAGs that contain the target strain but not the target virus / plasmid, and the number of SAGs that contain neither the target strain nor the target virus / plasmid.
[0017] Preferably, the statistical test method in step S3 includes: using Fisher's exact test to calculate the correlation between the target strain and the target virus / plasmid, obtaining the p-value and odds value, and thereby screening out and annotating the significant association results between the target strain and the target virus / plasmid.
[0018] Preferably, the method further includes step S4: outputting association analysis results, the results including at least the target strain ID, target virus / plasmid ID and its name, p-value and odd value obtained from statistical tests, number of SAGs containing the target strain, number of SAGs containing the target virus / plasmid, and number of SAGs containing both the target strain and the target virus / plasmid.
[0019] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any one of the methods described herein.
[0020] The beneficial effects of this invention are:
[0021] This invention provides a method for detecting the presence of viruses / plasmids significantly associated with a target bacterial species / strain in single-celled microbial genome data. This method can construct a unique virus / plasmid database based on known or assembled databases, thereby enabling the discovery and detection of unknown virus or plasmid sequences for subsequent association analysis. Compared to existing technologies that can only detect correlations with one bacteriophage per run, this invention achieves high-throughput, batch, and automated analysis by constructing a comprehensive reference index, independent reference indexes, and combining a two-level detection strategy. A single run can simultaneously detect the presence of all viruses / plasmids in the database within a sample and automatically complete association analysis with all strains, significantly improving detection efficiency. Attached Figure Description
[0022] Figure 1 This is a simplified flowchart of a method for batch detection of viruses in the genome of single-celled microorganisms provided in an embodiment of the present invention.
[0023] Figure 2 In step S1 of this embodiment of the invention, a reference database including a comprehensive reference index and an independent reference index is constructed from the data sourced from PhageScope (PMID: 37904614). Detailed Implementation
[0024] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.
[0025] refer to Figure 1 This invention provides a method for batch detection of viruses in single-celled microbial genomes. The method mainly includes three steps: constructing a reference database, analyzing the virus / plasmid information contained in the SAG (Self-Augmented Array), and performing Bin association analysis and annotation. The specific steps are as follows.
[0026] Step S1 first obtains the genome sequences of the virus / plasmid from a known database or an assembled database, and then standardizes the sequences to construct a reference database that can be used for sequence alignment. This reference database includes a comprehensive reference index containing all virus / plasmid genome sequences in the database, and an independent reference index for each virus / plasmid genome sequence in the database.
[0027] In one specific embodiment, the genome sequences of the virus / plasmid are obtained from a known database. The known database can be from a large database (e.g., Refseq) or other omics experiments, but it should include a separate FASTA file for each virus / plasmid genome sequence to be constructed, a basic annotation file, and other annotation files. Each plasmid / virus to be constructed is a separate FASTA file named ID.fasta. The ID is a unique identifier (ID number) for that plasmid / virus. The basic annotation file is a tab-delimited text file that should contain the ID column of the plasmid / virus (corresponding to the ID in the individual FASTA file) and its name (the name of the sequence, usually the common name of the plasmid or virus). The basic annotation file may also include other optional annotation columns. The other annotation files should contain the ID of the plasmid / virus (corresponding to the ID column in the basic annotation file). Subsequently, all virus / plasmid genome sequences are merged to obtain a combined FASTA file named combined.fasta. Within the combined FASTA file, the name of each contig is its original ID plus a "|" and the name of the plasmid or virus it belongs to. For example, if the virus name is Ecoil_phage_001 and the contig ID is V001, then the new contig ID is Ecoil_phage_001|V001. Using bowtie2, a bowtie2 index is constructed for all viruses / plasmids based on the merged FASTA file combined.fasta; this is called the combined reference index. A separate bowtie2 index is constructed for each plasmid / virus based on its individual FASTA file; this is called the single reference index. Finally, the basic annotation file and other annotation files are merged according to the plasmid / virus IDs to construct the reference database.
[0028] In another specific embodiment, the genome sequence of the virus / plasmid is obtained from an assembly database. Some genomic data (e.g., metagenomics or single-celled microbial genomes themselves) may contain gene sequences from viruses or plasmids. Furthermore, these sequences found in the assembly database may also include virus or plasmid sequences not included in the database, thus enabling the discovery and detection of unknown virus or plasmid sequences. The procedure for obtaining the genome sequence of the virus / plasmid from the assembly database can be found in Zhang Y. Environ Sci Ecotechnol. 2024 Sep12 (PMID: 39430728). Specifically, the genome sequence of the virus / plasmid is first obtained from the assembly database; the assembly results in the assembly database should be nucleic acid sequence files in FASTA format. The assembly database may also include assembly result annotations. Subsequently, using Virsorter2, candidate conitgs from the virus / plasmid are initially screened from the assembly results. Genomad is used to annotate the candidate conitgs, indicating their taxonomic names; the candidate conitg uses its original name as its ID and its taxonomic name as its name. After evaluating the integrity and contamination of candidate contigs using checkv, a quality rating is given based on a preset quality threshold. This quality rating includes high quality, medium quality, and low quality. Based on the genomad annotation results, for candidate contigs with the same annotation results, only the candidate contig with the highest integrity is retained as the reference contig. For all reference contigs with a quality rating of high or medium quality, a bowtie2 index is constructed using bowtie2, which is a single reference index for each virus / plasmid. All reference contigs with a quality rating of high or medium quality are merged to obtain a merged FASTA file named combined.fasta. Within the merged FASTA file, the name of each contig is its original ID plus a "|" followed by the name of its associated plasmid or virus. For example, if the virus name is Ecoil_phage_001 and the contig ID is V001, then the new contig ID is Ecoil_phage_001|V001. Using bowtie2, a bowtie2 index for all viruses / plasmids is constructed based on the merged FASTA file combined.fasta, called the combined_reference. Finally, annotation files based on genomad are generated according to the plasmid / virus IDs.
[0029] Step S2 analyzes the virus / plasmid information contained in the SAGs. First, bowtie2 is used to align the reads of each SAG to the comprehensive reference index obtained in Step S1, and the overall alignment rate psc is calculated. Based on a preset first threshold (P1), SAGs with an overall alignment rate psc ≤ P1 (1% < P1 < 50%) are removed, considering that these SAGs do not contain plasmids or viruses. SAGs with an overall alignment rate psc greater than P1 are selected as the initially screened SAGs containing viruses / plasmids. For the initially screened SAGs, samtools is used to calculate the number of reads aligned to each virus / plasmid genomic contig in each initially screened SAG. Based on which plasmid or virus the contig belongs to, the plasmid / virus with the most read alignments is determined as the main virus / plasmid. Then, the reads of this initially screened SAG are aligned to the independent reference index of this main virus / plasmid and the specific alignment rate psv is calculated. Based on a preset second threshold (P2), initially screened SAGs with (psv / psc) ≤ P2 (50% < Ps < 90%) are removed, considering that they contain virus or plasmid fragments and thus have a high contamination level. Initially screened SAGs with (psv / psc) > P2 are selected, considering that they contain the plasmid or virus with the most read alignments, as the finely screened SAGs. For each main plasmid / virus, a corresponding set Bin (L v ) is generated:
[0030] .
[0031] Step S3 performs association analysis and annotation on the Bin containing the finely screened SAGs obtained in Step S2. For each Bin, first, a four-way table Fvb as shown in Table 1 is constructed based on the strain classification information of the SAG and the main virus / plasmid determined in Step S2.
[0032] Table 1. Four-way table
[0033]
[0034]
[0035] Among them, L A is the全集 of SAGs, L b is the set of SAGs of the target strain, and L v is the set of finely screened SAGs containing the target virus / plasmid. That is, the four-way table includes the following values: the number of SAGs that contain both the target strain and the target virus / plasmid (N p,p,v,b ), the number of SAGs that contain the target virus / plasmid but do not contain the target strain (N p,n,v,b ), the number of SAGs that contain the target strain but do not contain the target virus / plasmid (N n,p,v,bThe number of SAGs (N) that contain neither the target strain nor the target virus / plasmid. n,n,v,b ).
[0036] Subsequently, the `fisher_exact` function of scipy is used to call Fisher's exact test to calculate the correlation between the target strain and the target virus / plasmid, obtaining the p-value and odds value:
[0037] .
[0038] Finally, based on the p-value and odds value, significant associations between the target strain and the target virus / plasmid were screened and their plasmid / virus information was annotated.
[0039] Step S4 is used to output the association analysis results. The results include at least the target strain ID, target virus / plasmid ID and name, p-value and odds value obtained from statistical tests, number of SAGs containing the target strain, number of SAGs containing the target virus / plasmid, and number of SAGs containing both the target strain and the target virus / plasmid.
[0040] In one specific embodiment, an apparatus for batch detection of viruses in single-celled microbial genomes is also provided, including a processor and a memory; the processor and the memory are connected via a communication bus; wherein, the processor is used to call and execute a program stored in the memory; the memory is used to store the program, which is at least used to execute the method for batch detection of viruses in single-celled microbial genomes.
[0041] In one specific embodiment, custom data is constructed from the data sourced from PhageScope (PMID: 37904614) according to step S1 of the above method. Specifically, the genome sequences of viruses / plasmids are obtained from known databases and processed to obtain a reference database including a comprehensive reference index and independent reference indexes. The results are shown in [the table below]. Figure 2 Among them, combined_bin.x.bt2l and combined_bin_rev.x.bt2l are the combined reference index files after merging all sequences; combined_bin.fasta is the FATA file after merging all sequences; meta_data.tsv contains the annotation information for each individual virus (Table 2); the fasta folder contains the FATA files for each individual virus sequence; and the bowtie2_refs folder contains the individual index reference files for each individual virus.
[0042] Table 2. Examples of meta_data.tsv
[0043]
[0044] Where Phage_ID is the virus ID number; Length is the length of the virus conig; GC_content is the virus GC ratio; Taxonomy is the virus classification information; Name is the virus name; Completeness is the assembly integrity; and Host is the virus host.
[0045] After constructing the reference database using the above process, the technical article (PMID: 35653470) was processed according to steps S2-S4, and the results shown in Table 3 were output. Using this method, six additional viruses and strains with significant correlations were detected in the article data, indicating that this method has the ability to discover and detect unknown viruses or plasmid sequences.
[0046] Table 3. Processing Results of Technical Article (PMID: 35653470)
[0047]
[0048] Wherein, bin_ID is the strain genome number; phage_ID is the virus / plasmid ID; p_value is the p-value obtained from Fisher's exact test; odd is the odd value obtained from Fisher's exact test. If the odd value is inf (infinity), it means that the count of the corresponding cell in the quadruple table is zero, indicating that the target virus / plasmid and the target strain show a very strong specific association in this dataset; SAG_with_strain is the number of SAGs containing the target strain; SAG_with_phage is the number of SAGs containing the target virus / plasmid; SAG_with_strain&phage is the number of SAGs containing both the target strain and the target virus / plasmid; and phage_name is the name of the virus or plasmid.
[0049] In one specific embodiment, the single-celled microbial genome sequencing results were processed according to the above method to obtain the genome sequences of viruses / plasmids from refseq and processed to obtain a reference database including a comprehensive reference index and independent reference indexes. The processing results are shown in Table 4.
[0050] Table 4. Results of single-celled microbial genome processing
[0051]
[0052] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Various modifications and improvements made by those skilled in the art to the technical solutions of the present invention without departing from the spirit of the present invention should fall within the protection scope of the present invention.
Claims
1. A method for batch detection of viruses in the genome of single-celled microorganisms, characterized in that, Includes the following steps: S1. Obtain the genome sequences of viruses / plasmids from the database and standardize the sequences to construct a reference database that can be used for sequence alignment; the database includes known databases or assembled databases; the reference database includes a comprehensive reference index containing all virus / plasmid genome sequences in the database and an independent reference index for each virus / plasmid genome sequence in the database; S2. Align the reads of each SAG to the comprehensive reference index obtained in step S1 and calculate the overall alignment rate; based on a preset first threshold, select SAGs with an overall alignment rate greater than the first threshold as primary screening SAGs containing viruses / plasmids; calculate the number of reads in each viral / plasmid genome contig aligned in each primary screening SAG, determine the plasmid / virus with the most read alignments as the primary virus / plasmid, then align the reads of the primary screening SAG to the independent reference index of the primary virus / plasmid and calculate the specific alignment rate; based on a preset second threshold, select primary screening SAGs with a specific alignment rate to overall alignment rate ratio greater than the second threshold as secondary screening SAGs containing the primary virus / plasmid; generate a corresponding Bin containing the secondary screening SAG for each primary plasmid / virus; S3. For each Bin, construct a quadrant table based on the strain classification information of SAG and the major virus / plasmid determined in step S2, and perform statistical tests to determine whether there is a significant association between the target strain and the target virus / plasmid, and annotate it.
2. The method as described in claim 1, characterized in that, Step S1 includes: Obtain the genome sequences of the virus / plasmid from a known database; the known database includes a separate FASTA file, a basic annotation file, and other annotation files for each virus / plasmid genome sequence to be constructed; An independent reference index is built for each virus / plasmid based on individual FASTA files; All viral / plasmid genome sequences are merged to obtain a merged FASTA file, and a comprehensive reference index of all viruses / plasmids is constructed based on the merged FASTA file.
3. The method as described in claim 2, characterized in that, The individual FASTA file is named with the ID of the plasmid / virus, the basic annotation file contains the ID of the plasmid / virus and its name, and the other annotation files contain the ID of the plasmid / virus. The name of each contig in the merged FASTA file contains the ID and name of its corresponding plasmid / virus. Merge basic annotation files and other annotation files according to plasmid / virus IDs.
4. The method as described in claim 1, characterized in that, Step S1 includes: The genome sequences of the virus / plasmid are obtained from the assembly database; the assembly database contains assembly results in FASTA format and assembly result annotations. Candidate conitgs from viruses / plasmids are screened from the assembly results; Candidate contigs are annotated; after evaluating the completeness and contamination of candidate contigs, a quality rating is given based on a preset quality threshold, which includes high quality, medium quality, and low quality; for candidate contigs with the same annotation results, only the candidate contig with the highest completeness is retained as the reference contig. An independent reference index is constructed for each virus / plasmid based on reference contigs whose quality rating is not low. Merge all reference contigs that are not of low quality to obtain a merged FASTA file, and build a comprehensive reference index for all viruses / plasmids based on the merged FASTA file.
5. The method as described in claim 4, characterized in that, The annotation file for the candidate contig includes the ID of the plasmid / virus to which the candidate contig belongs and its name; Each contig in the merged FASTA file contains its corresponding plasmid / virus ID and name; Merge annotation files according to plasmid / virus IDs.
6. The method as described in claim 1, characterized in that, In step S2, the first threshold is 1% to 50%, and the second threshold is 50% to 90%.
7. The method as described in claim 1, characterized in that, The quadruple table mentioned in step S3 includes the following values: the number of SAGs that contain both the target strain and the target virus / plasmid, the number of SAGs that contain the target virus / plasmid but not the target strain, the number of SAGs that contain the target strain but not the target virus / plasmid, and the number of SAGs that contain neither the target strain nor the target virus / plasmid.
8. The method as described in claim 1, characterized in that, The statistical testing method described in step S3 includes: using Fisher's exact test to calculate the correlation between the target strain and the target virus / plasmid, obtaining the p-value and odds value, and then screening out and annotating the significant association results between the target strain and the target virus / plasmid.
9. The method as described in claim 1, characterized in that, The method further includes step S4: outputting association analysis results, the results of which include at least the target strain ID, target virus / plasmid ID and its name, p-value and odd value obtained from statistical tests, number of SAGs containing the target strain, number of SAGs containing the target virus / plasmid, and number of SAGs containing both the target strain and the target virus / plasmid.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 9.