A method for omics analysis of environmental rna viruses
By combining metatranscriptome sequencing with the RdRp homology iteration method and the geNomad method, environmental RNA viruses can be identified and analyzed. This solves the problems of high cost and long cycle of traditional methods, and realizes efficient and economical RNA virus detection and analysis, providing detailed classification and host information.
Patent Information
- Application Number
- CN202510118509.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-01-24
AI Technical Summary
Existing technologies are insufficient for the efficient and economical detection and analysis of RNA viruses in the environment. In particular, due to the limited number of RNA virus sequences, traditional culture methods are costly, time-consuming, and difficult to identify unknown RdRp sequences.
Metatranscriptome sequencing combined with RdRp homology iteration and the geNomad method of machine learning was used to identify RNA virus operational taxonomic units (vOTUs) by assembling contigs, removing redundancy and clustering. Taxonomic analysis was performed using CAT and geNomad methods, and life history was predicted using VIBRANT software and host information was determined using the RNAvirHOST method.
It enables efficient detection and analysis of environmental RNA viruses, improves the ability to identify RNA virus sequences, provides rich taxonomic analysis results, and identifies potential high-risk viruses and host information.
Smart Images

Figure CN119920323B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of bioinformatics analysis, and in particular relates to an omics analysis method for environmental RNA viruses. Background Technology
[0002] Viruses are the most numerous organisms on Earth, with an estimated 10 million viruses worldwide. 31 Viral particles, including RNA viruses, use RNA as their genetic material and can infect a variety of organisms, posing a significant threat to humans and livestock, such as the well-known COVID-19 and avian influenza viruses. For over a century, cell or animal model-based culture methods have been the gold standard for virus detection. However, due to their long cycles, high costs, and labor intensity, and the fact that many viruses are difficult to culture, it is difficult to efficiently detect RNA virus species present in the environment. Metatranscriptomics sequencing has become a new method for discovering novel RNA viruses because it eliminates the need for virus isolation and culture in a laboratory environment. RNA-dependent RNA polymerase (RdRp) is a key protease essential for RNA virus replication, controlling the replication and transcription processes of RNA viruses. This provides a theoretical basis for identifying RNA viruses based on the homology of RdRp domains in metatranscriptomics. However, due to the limited number of known RNA virus sequences, homology analysis alone cannot fully identify unknown RdRp sequences. Therefore, combining RdRp homology analysis with emerging machine learning, k-mer analysis, and other principles to identify RNA virus sequences in the metatranscriptome can provide methods and technologies for the detection and analysis of RNA viruses in the environment, which has important practical significance. Summary of the Invention
[0003] To address the aforementioned problems, the present invention aims to provide an omics analysis method for environmental RNA viruses. First, the sample undergoes metatranscriptome sequencing, and the raw metatranscriptome sequencing data is subjected to quality control analysis to obtain effective reads. Then, contigs are assembled from these reads, and RNA virus contigs are further identified using various methods. Redundancy removal and clustering are then performed to obtain RNA virus operational taxonomic units (vOTUs) and their sequences. Next, quality control is conducted to obtain characteristics such as the length, contamination level, and integrity of RNA virus sequences in the sample. Then, abundance analysis is used to determine the relative abundance of different RNA viruses in the sample. Next, species classification analysis is performed, creatively coupling two viral bioinformatics classification methods, CAT and geNomad, to obtain classification results. Then, the life history of RNA viruses is predicted using VIBRANT software to obtain functional annotation results. Finally, the RNAvirHOST method is used to clarify the host information of RNA virus sequences, distinguishing bacteriophages, animal viruses, plant viruses, etc., with a focus on human and animal health and economic crops, to identify the potential risks of RNA viruses.
[0004] To achieve the above objectives, the present invention adopts the following technical solution: a method for omics analysis of RNA viruses, comprising the following steps, wherein steps 3-7 can be performed simultaneously:
[0005] 1) Perform metagenomic sequencing on the samples, and then decompress and perform quality control analysis on the raw data after the metagenomic sequencing to obtain valid reads;
[0006] 2) The effective reads are assembled into contigs, and the RNA virus contig sequences are identified by the RdRp-based homology iteration method and the geNomad method. The identified RNA virus contig sequences are then deredundant and clustered to obtain RNA virus operational taxonomic units (vOTUs) and their sequences.
[0007] 3) Perform checkV quality control on the sequences of RNA virus operational taxonomic units (vOTUs) obtained in step 2), and reveal the characteristics of RNA virus sequences in the sample, such as length, contamination level, and integrity, based on the quality control results;
[0008] 4) Use the coverM tool to perform abundance analysis on the sequences of RNA virus operational taxonomic units (vOTUs) obtained in step 2) to determine the relative abundance of different RNA viruses in the sample;
[0009] 5) Use geNomad and CAT software to perform taxonomic analysis on the sequences of RNA virus operational taxonomic units (vOTUs) obtained in step 2). Prioritize the classification results of geNomad, and then use CAT results as a supplement. By combining the geNomad method or the CAT method, obtain the taxonomic analysis results of RNA virus operational taxonomic units (vOTUs) at the phylum and family levels.
[0010] 6) Analyze the sequences of RNA virus operational taxonomic units (vOTUs) obtained in step 2) using VIBRANT software, predict the life cycle of RNA viruses, and obtain the functional annotation results of RNA virus operational taxonomic units (vOTUs).
[0011] 7) Using species classification information at the family level, the RNA virus operational taxonomic unit (vOTU) sequence obtained in step 2) is analyzed using the RNAvirHOST virus-host association method to determine the viral host information of the RNA virus operational taxonomic unit (vOTU).
[0012] In step 1) above, the sample undergoes metagenomic sequencing. The raw data from the metagenomic sequencing is then decompressed and subjected to quality control analysis to obtain valid reads. This includes the following steps:
[0013] Metatranscriptome sequencing was performed on the samples, and the raw data was decompressed. The read_qc module of MetaWrap software (v 1.3.2) was used to perform quality control on the decompressed data, and finally, the quality control filtered, adapter-free, high-quality reads, i.e., effective reads, were obtained. The echo tool was used to count the effective reads (clean reads).
[0014] In step 2) above, valid reads are assembled into contigs, and RNA virus contig sequences are identified using the RdRp-based homology iteration method and the geNomad method based on a machine learning model. The identified RNA virus contig sequences are then deredundant and clustered to obtain RNA virus operational taxonomic units (vOTUs) and their sequences. This includes the following steps, where steps 2.3-2.4 can be performed simultaneously:
[0015] 2.1) Use the assembly module of MetaWrap software (v 1.3.2) to assemble the valid reads of each sample into contigs.
[0016] 2.2) Screen out and merge contig sequences longer than 1.0 kb from the samples to form a preliminary RNA virus dataset, and predict the sequences in this dataset as a viral protein sequence dataset.
[0017] 2.3) By using the homology iteration method (i.e., using the hidden Markov model, iterating a total of 5 times), the RdRp protein sequences in the viral protein sequence database are compared to select viral sequences with RdRp protein sequences.
[0018] 2.4) RNA virus sequences were identified by classifying and screening the preliminary RNA virus dataset using the machine learning-based geNomad (v 1.7.0) software.
[0019] 2.5) The RNA virus sequences obtained by the homology iteration method based on RdRp protein sequences and the machine learning-based geNomad (v 1.7.0) software were merged, and the sequences with 100% similarity were clustered based on CD-HIT-EST to obtain the RNA virus contig sequence dataset.
[0020] 2.6) Cluster the RNA virus sequence dataset based on sequence similarity ≥90% and coverage ≥80% to obtain non-redundant RNA virus operational taxonomic units (vOTUs) and their sequences.
[0021] In step 3) above, the sequences of RNA virus operational taxonomic units (vOTUs) obtained in step 2) are subjected to checkV quality control, and based on the quality control results, the characteristics of the RNA virus sequences in the sample, such as length, contamination level, and integrity, are revealed, including the following steps:
[0022] The checkV tool was used to perform quality control on RNA virus operational taxonomic unit (vOTU) sequences, thereby obtaining characteristics such as the length, contamination level, and integrity of RNA virus sequences in the sample.
[0023] In step 4) above, the abundance of RNA virus operational taxonomic units (vOTUs) obtained in step 2) is analyzed using the coverM tool to determine the relative abundance of different RNA viruses in the sample. This includes the following steps:
[0024] 4.1) Use the coverM tool to align the sequences of RNA virus operational taxonomic units (vOTUs) with valid reads to generate the number of covered bases (base) and RPKM abundance information for each RNA virus operational taxonomic unit (vOTU) in different samples.
[0025] 4.2) Divide the number of bases covered by each RNA virus operational taxonomic unit (vOTU) in the sample by its length (the result obtained in step 3 above). If it is ≥70% (i.e., coverage ≥70%), it means that the RNA virus operational taxonomic unit (vOTU) exists in the sample.
[0026] 4.3) The relative abundance of the RPKM of the real RNA virus operational taxonomic units (vOTUs) in the sample will be calculated, that is, the abundance of all RNA viruses in a single sample will be taken as 100%, so as to obtain the relative abundance of RNA viruses in different samples.
[0027] In step 5) above, the sequences of RNA virus operational taxonomic units (vOTUs) obtained in step 2) are subjected to taxonomic analysis using geNomad and CAT software. The classification results from geNomad are given priority, supplemented by CAT results. By combining either the geNomad or CAT methods, the taxonomic analysis results of the RNA virus operational taxonomic units (vOTUs) at the phylum and family levels are obtained. This includes the following steps, where steps 5.1-5.5 and 5.6-5.8 are performed simultaneously:
[0028] 5.1) Use the geNomad (v 1.7.0) software to annotate the RNA virus operational taxonomic unit (vOTU) sequences obtained in step 2).
[0029] 5.2) Use CAT (v5.2.3) software to annotate the RNA virus operational taxonomic unit (vOTU) sequences obtained in step 2).
[0030] 5.3) Prioritize the classification results of geNomad (v 1.7.0), and then supplement with the results of CAT (v5.2.3) to finally obtain the classification results of RNA virus operational taxonomic units (vOTU).
[0031] 5.4) The abundance of sequences of non-redundant RNA virus operational taxonomic units (vOTUs) was analyzed using the coverM (v0.7.0) tool to determine different RNA virus populations and their relative abundance in the samples.
[0032] In step 6) above, the sequences of the RNA virus operational taxonomic units (vOTUs) obtained in step 2) are analyzed using VIBRANT software to predict the life cycle of RNA viruses. The functional annotation results of the RNA virus operational taxonomic units (vOTUs) include the following steps:
[0033] 6.1) Enable VIBRANT software (v1.2.1) to predict the life cycle of RNA viruses;
[0034] 6.2) The gene sequences of non-redundant RNA virus operational taxonomic units (vOTU) were predicted using Prodigal (v2.6.3), and compared with the SARG and CARD databases based on the blastx method;
[0035] 6.3) The coverM (v0.7.0) tool was used to perform abundance analysis on non-redundant RNA virus operational taxonomic unit (vOTU) sequences to determine the relative abundance of different RNA virus populations in the samples;
[0036] 6.4) Based on the comparison results of the SARG and CARD databases, the prediction results of the VIBRANT software, and the abundance results of coverM, the life history of RNA viruses, the characteristics of carrying antibiotic resistance genes (ARGs), and their relative abundance were determined.
[0037] In step 7) above, the RNA virus operational taxonomic unit (vOTU) sequence obtained in step 2) is analyzed using species classification information at the family level and the method of associating RNAvirHOST viruses with their hosts to determine the viral host information of the RNA virus operational taxonomic unit (vOTU), including the following steps:
[0038] 7.1) The RNAvirHOST software (v1.0.5) was used to perform host infectivity analysis on the sequences of non-redundant RNA virus operational taxonomic units (vOTUs);
[0039] 7.2) After classifying the non-redundant RNA virus operational taxonomic units (vOTUs), the family-level classification results are entered into the ICTV (https: / / ictv.global / ) website, which provides the classification of infectable hosts for that family, thereby obtaining the ICTV host characteristics;
[0040] 7.3) Combine the results of virus-host relationships obtained in steps 7.1)-7.2) to obtain the host characteristics of non-redundant RNA virus operational taxonomic units (vOTUs).
[0041] the term:
[0042] Hidden Markov models are known statistical models used to describe a Markov process with hidden unknown parameters. In this invention, a homology iteration method is used with a hidden Markov model. Following the standard operation of the model, the sequence is compared with known sequences (such as RdRp). The aligned sequence is then included in the known sequence for the next round of homology comparison, and this process is repeated five times to uncover as many viral sequences as possible.
[0043] geNomad is a tool for identifying viral and plasmid genomes from nucleotide sequences. It can quickly find mobile genetic elements from genomes and metagenomics. For instructions on how to use it, please see https: / / github.com / apcamargo / genomad.
[0044] Specifically, the present invention relates to the following technical solutions:
[0045] 1. A method for analyzing RNA viruses in samples, comprising the following steps:
[0046] 1) Metatranscriptomics data were obtained by performing metatranscriptomic sequencing on the samples. Then, the metatranscriptomics data were analyzed by the homology iteration method based on RdRp and the geNomad method based on machine learning model to screen RNA virus sequences in the samples and obtain RNA virus operational taxonomic units (vOTUs).
[0047] 2) Perform abundance analysis, classification identification, and host analysis on the RNA virus operational taxonomic units obtained in step 1);
[0048] The classification and identification process includes analyzing the RNA virus operational taxonomic unit sequences using two methods, geNomad and CAT, to obtain the classification results of the RNA virus operational taxonomic units in the sample.
[0049] 2. According to the method described in Project 1, wherein the host analysis includes analyzing the RNA virus operational taxonomic unit (vOTU) sequence obtained in step 1) using the RNAvirHOST method to obtain the host information of the RNA virus in the sample.
[0050] 3. The method according to Project 1 or 2, wherein the abundance analysis includes performing abundance analysis on the RNA virus operational taxonomic unit sequences obtained in step 1) using the coverM tool to calculate the relative abundance of RNA viruses in the sample.
[0051] 4. The method according to Project 1 or 2, wherein step 2) further includes performing checkV quality control on the RNA virus operational taxonomic unit sequence obtained in step 1) to reveal the characteristics of the RNA virus sequence in the sample, such as length, contamination level, and integrity.
[0052] 5. The method according to Project 1 or 2, wherein step 2) further includes analyzing the RNA virus operational taxonomic unit sequence obtained in step 1) using VIBRANT software to predict the life history of the RNA virus in the sample.
[0053] 6. According to the method described in Project 1, wherein, in step 1), the sample is subjected to metatranscriptome sequencing to obtain raw metatranscriptome data; the raw data is subjected to quality control analysis to obtain valid reads; the valid reads are assembled into contigs; then the contigs are screened by homology iteration and geNomad methods to obtain RNA virus contigs in the sample; the RNA virus contigs obtained by homology iteration and geNomad methods are merged and clustered to obtain RNA virus operational taxonomic units (vOTUs) and their sequences.
[0054] 7. The method according to Project 1, wherein the sample includes soil, water, atmospheric particulate matter, composting environment, sludge, river or seabed sediment, or anaerobic digestion environment.
[0055] 8. The method according to Project 7, wherein the anaerobic digestion environment includes livestock and poultry manure and food waste.
[0056] The present invention has the following advantages due to the adoption of the above technical solutions:
[0057] 1. This invention innovatively combines the analysis methods of the homology iteration method based on RdRp and the geNomad method based on machine learning to identify RNA virus sequences in samples, which can greatly improve the ability to detect RNA viruses in omics data.
[0058] 2. In classifying RNA viruses, this invention creatively couples two viral bioinformatics classification methods, the CAT method and the geNomad method, resulting in richer and more representative taxonomic analysis results.
[0059] 3. This invention innovatively combines species classification at the scientific level to determine the host, virus and host information, and uses humans, economic animals and plants as key hosts to determine potentially high-risk RNA viruses vOTU.
[0060] Therefore, this invention can be widely applied to the field of omics analysis of environmental RNA viruses. Attached Figure Description
[0061] Figure 1 This is a schematic flowchart of the omics analysis method for environmental RNA viruses of the present invention;
[0062] Figure 2 This is a distribution diagram of environmental RNA virus quality inspection results in Example 3 of the present invention;
[0063] Figure 3 This is a frequency distribution diagram of environmental RNA viruses in samples according to Example 4 of the present invention;
[0064] Figure 4This is the taxonomic result of environmental RNA viruses at the phylum and family level in Example 5 of the present invention; wherein, Figure 4 A shows the species composition at the RNA virus phylum level. Figure 4 B shows the species composition at the RNAviridae level.
[0065] Figure 5 This is a distribution diagram of the RNA life cycle in Example 6 of the present invention;
[0066] Figure 6 This is a diagram showing the host composition and abundance distribution of RNA viruses in Example 7 of the present invention; wherein, Figure 6 A shows the number of vOTUs that the RNA virus can infect in the host. Figure 6 B shows the relative abundance of RNA viruses that can infect the host. Detailed Implementation
[0067] The present invention is further illustrated by the following embodiments, but no embodiment or combination thereof should be construed as limiting the scope or embodiments of the invention. The scope of the invention is limited by the appended claims. Based on this specification and general knowledge in the art, those skilled in the art can clearly understand the scope of the claims. Without departing from the spirit and scope of the invention, those skilled in the art can make any modifications and alterations to the technical solutions of the invention, and such modifications and alterations are also included within the scope of the invention.
[0068] Example 1: Data Acquisition and Quality Control
[0069] Raw data from anaerobic digestion of feces from different livestock (chicken, pig, and cattle) worldwide was collected and downloaded from public databases of NCBI (https: / / www.ncbi.nlm.nih.gov) and GSA (https: / / ngdc.cncb.ac.cn), totaling 36 samples. Alternatively, raw data samples can be obtained by performing metagenomic sequencing on environmental samples.
[0070] According to the software's instructions, the read_qc module of MetaWrap software (v 1.3.2) was used to perform quality control on the raw sample data. This module uses FastQC (v0.11.8) software to check the overall quality of the raw readings and Trimmomactic (v0.39) software to remove paired-end connector sequences and low-quality sequences, ultimately obtaining quality-controlled filtered connector-free high-quality reads, i.e., clean reads, totaling 640 GB.
[0071] The specific steps for using the MetaWrap software in this example are as follows:
[0072] for F in 01RAW_READS / *_1.fastq; do
[0073] R=${F%_*}_2.fastq
[0074] BASE=${F##* / }
[0075] SAMPLE=${BASE%_*}
[0076] metawrap read_qc -1 $F -2 $R -t 20 --skip-bmtagger -o 01READ_QC / $SAMPLE
[0077] pigz $F $R
[0078] done
[0079] Example 2: Identification of RNA Virus Sequences and Analysis of Viral Operational Taxonomic Units (vOTUs)
[0080] According to the software's instructions, the assembly module of the MetaWrap software (v 1.3.2) is used to assemble valid reads to obtain contig sequences. Specifically, this module calls MEGAHIT (v1.1.3) to assemble contig sequences. Contig sequences longer than 1.0 kb in the sample are screened out and merged into a preliminary RNA virus dataset RNA_contigs1000bp.fasta.
[0081] The specific operation of the MetaWrap software is as follows:
[0082] for F in 01Reads_fq / *_1.fastq; do
[0083] R=${F%_*}_2.fastq
[0084] BASE=${F##* / }
[0085] SAMPLE=${BASE%_*}
[0086] metawrap assembly -1 $F -2 $R -m 500 -t 50 -o 05ASSEMBLY / $SAMPLE
[0087] rm -rf 05ASSEMBLY / $SAMPLE / megahit
[0088] done
[0089] for F in 01Reads_fq / *_1.fastq; do
[0090] R=${F%_*}_2.fastq
[0091] BASE=${F##* / }
[0092] SAMPLE=${BASE%_*}
[0093] sed "s / k141 / $SAMPLE / g" 05ASSEMBLY / $SAMPLE / final_assembly.fasta |seqtk seq -L 1000 >06VCs / 00Contigs / $SAMPLE.fa
[0094] done
[0095] cat 06VCs / 00Contigs / *.fa >06VCs / 00Contigs / RNA_contigs1000bp.fasta
[0096] The initial RNA virus dataset RNA_contigs1000bp.fasta obtained above was used to predict the corresponding protein sequences using Prodigal (v2.6.3) (see https: / / github.com / hyattpd / Prodigal for usage instructions). Then, using the homology iteration method (i.e., using a hidden Markov model, iterated a total of 5 times), the RdRp protein sequences in the contig protein sequences were compared, and viral sequences containing RdRp protein sequences were selected, resulting in a total of 1696 RNA virus contig sequences.
[0097] The specific method is as follows:
[0098] . / rdrp-search -f "test / *.faa" -t 50 -g -o test / out
[0099] samtools faidx test / out / final-hits.fa;
[0100] In addition, the initial RNA virus dataset RNA_contigs1000bp.fasta was filtered using geNomad (v1.7.0) to obtain a total of 1811 RNA virus contig sequences (see https: / / github.com / apcamargo / genomad for usage instructions).
[0101] Then, by merging 1696 RNA virus contig sequences obtained by the homology iteration method with 1811 RNA virus contig sequences obtained by the geNomad method, and clustering the sequences with 100% similarity based on CD-HIT-EST (v 4.7), 1994 RNA virus contig sequences were obtained, which is 17.57% and 10.1% higher than the homology iteration method (1696) and the geNomad (1811) method, respectively. Finally, these sequences were clustered according to sequence similarity ≥90% and coverage ≥80% to obtain 729 RNA virus operational taxonomic units (vOTUs) and their sequences.
[0102] The specific operation of this embodiment is as follows:
[0103] 1) Decompress the raw data after metatranscriptome sequencing, and use the read_qc module of MetaWrap software (version 1.3.2) to perform quality control on the decompressed data, and use the echo tool to count the effective reads (cleanreads).
[0104] 2) Use the metawrap assembly module to assemble the valid reads of each sample into a contiguous group.
[0105] 3) Screen out and merge contiguous sequences longer than 1.0 kb from the samples to form a preliminary RNA virus dataset, and predict the sequences in this dataset to become a viral protein sequence dataset.
[0106] 4) By using the homology iteration method (5 iterations in total), compare the RdRp protein sequences in the viral protein sequence database and select viral sequences that have RdRp protein sequences.
[0107] 5) RNA virus sequences were identified by classifying and screening the preliminary RNA virus dataset using the machine learning-based geNomad (v 1.7.0) software.
[0108] 6) The RNA virus sequences obtained by the RdRp protein sequence iteration method and the geNomad software based on machine learning were merged, and the sequences with 100% similarity were clustered based on CD-HIT-EST to obtain the RNA virus contiguous group sequence dataset.
[0109] 7) Cluster the RNA virus sequence dataset based on sequence similarity ≥90% and coverage ≥80% to obtain non-redundant RNA virus operational taxonomic units (vOTUs) and their sequences.
[0110] Example 3: RNA Virus Quality Control Analysis
[0111] After quality checking the 729 RNA virus operational taxonomic unit (vOTU) sequences obtained using checkV (v1.0.1) (see https: / / bitbucket.org / berkeleylab / checkv / src / master / for usage instructions), the viral sequence lengths were found to be between 1000 and 10786 bp. According to the software's built-in quality assessment criteria (e.g., length, contamination level, and integrity), 81, 136, and 283 viral sequences were identified as high-quality, medium-quality, and low-quality, respectively. The quality of the remaining 229 viral sequences could not be determined. Figure 2 ).
[0112] Example 4 Abundance Analysis of RNA Viruses
[0113] Abundance analysis of the RNA virus operational taxonomic units (vOTUs) obtained in Example 2 was performed using the coverM (v0.7.0) tool (see: https: / / github.com / wwood / CoverM for usage instructions). Based on coverage (≥70%) (a commonly reported standard in the literature), RNA viruses were found in only 29 out of 36 samples. Furthermore, most viruses (475 in total) had a very low actual frequency (≤3) in these 29 samples, and the most frequent virus appeared in only 15 samples. Figure 3 ).
[0114] The specific operation of this embodiment is as follows:
[0115] 1) The RNA virus operational taxonomic unit (vOTU) sequence obtained in Example 2 was compared with the effective reads using the coverM (v0.7.0) tool to generate the number of covered bases (base) and RPKM abundance information of each RNA virus operational taxonomic unit (vOTU) in different samples.
[0116] 2) Divide the number of bases covered by each RNA virus operational taxonomic unit (vOTU) sequence in the sample by its length (the result obtained in Example 3). If ≥70% (i.e., coverage ≥70%), it indicates that the RNA virus operational taxonomic unit (vOTU) is present in the sample.
[0117] 3) The relative abundance of RPKMs of real RNA virus operational taxonomic units (vOTUs) in the samples will be calculated, that is, the abundance of all RNA viruses in a single sample will be taken as 100%, so as to obtain the relative abundance of RNA viruses in different samples.
[0118] Example 5: Taxonomic Analysis of RNA Viruses
[0119] The sequences of the 729 RNA virus operational taxonomic units (vOTUs) obtained in Example 2 were analyzed for taxonomic purposes using two methods: geNomad (v1.7.0, usage instructions refer to https: / / github.com / apcamargo / genomad) and CAT (v5.2.3, usage instructions refer to https: / / github.com / MGXlab / CAT_pack). geNomad and CAT yielded classification results for 647 and 670 RNA virus operational taxonomic units (vOTUs), respectively. A combination of methods was used, prioritizing the geNomad results and supplementing them with CAT results. Ultimately, a classification of 727 RNA virus operational taxonomic units (vOTUs) was obtained.
[0120] Figure 4 This paper presents the final taxonomic analysis results of RNA viruses at the phylum and family levels in this embodiment. During the anaerobic digestion of livestock and poultry manure, RNA viruses mainly consist of phyla such as Duplopiviricetes (28.71%), Pisoniviricetes (17.75%), Leviviricetes (11.52%), and Howeltoviricetes (9.12%). Figure 4 A); At the family level, a total of 36 RNA virus families were identified. Among these families, five RNA virus families—Picobirnaviridae, Partitiviridae, Mitoviridae, Leviviridae, and Picornaviridae—were found to be the main core families in the anaerobic digestion of livestock and poultry manure (present in more than 80% of samples). Figure 4 B).
[0121] The specific operation of this embodiment is as follows:
[0122] 1) Species annotation was performed on the non-redundant RNA virus operational taxonomic unit (vOTU) sequences obtained in Example 2 using the geNomad (v 1.7.0) software;
[0123] 2) Species annotation was performed on the non-redundant RNA virus operational taxonomic unit (vOTU) sequences obtained in Example 2 using CAT (v5.2.3) software;
[0124] 3) Prioritize the classification results of geNomad (v 1.7.0), and then supplement with the results of CAT (v5.2.3) to finally obtain the classification results of RNA virus operational taxonomic units (vOTUs);
[0125] 4) The abundance of non-redundant RNA virus operational taxonomic units (vOTUs) was analyzed using the coverM (v0.7.0) tool to determine the relative abundance of different RNA virus populations in the sample.
[0126] Example 6: Life history and functional analysis of RNA viruses
[0127] After predicting the life history of RNA viruses using VIBRANT software (v1.2.1), only 28 RNA virus operational taxonomic units (vOTUs) could be predicted as lytic bacteriophages, accounting for an average relative abundance of 1.72% in the 29 samples. Figure 5 ).
[0128] In addition, by comparing with antibiotic resistance gene (ARG) databases such as CARD and SARG, a vOTU471 identified as Partitiviridae was found to carry the aadS resistance gene, with an average relative abundance of 0.04% in 29 samples.
[0129] The specific operation of this embodiment is as follows:
[0130] 1) Enable VIBRANT software (v1.2.1) to predict the life cycle of RNA viruses;
[0131] 2) The gene sequences of the non-redundant RNA virus operational taxonomic unit (vOTU) sequences obtained in Example 2 were predicted using Prodigal (v2.6.3), and compared with the SARG and CARD databases based on the blastx method;
[0132] 3) The abundance of non-redundant RNA virus operational taxonomic unit (vOTU) sequences was analyzed using the coverM (v0.7.0) tool to determine the relative abundance of different RNA virus populations in the sample;
[0133] 4) Based on the comparison results of the SARG and CARD databases, the prediction results of the VIBRANT software, and the abundance results of coverM, the life history of RNA viruses, the characteristics of carrying antibiotic resistance genes (ARGs), and their relative abundance were determined.
[0134] Example 7: Analysis of RNA Infection Hosts and Transmission Risk
[0135] Based on species classification information at the family level, the virus-host association method was used with RNAvirHOST software (v1.0.5). Specifically, the virus family classification names were searched on the ICTV website (https: / / ictv.global / ) to find the hosts infecting them, thus identifying the viral hosts of 474 vOTUs. The RNAvirHOST software (v1.0.5) was used to obtain the viral hosts of 370 vOTUs, resulting in a final viral host information for 585 vOTUs. These RNA viruses were mainly vertebrate viruses (227), bacteriophage viruses (149), and fungal viruses (97), while invertebrate viruses (57) and plant viruses were fewer in number (55). Figure 6 A). Abundance analysis revealed that RNA viruses in livestock and poultry manure were mainly vertebrate viruses with an average relative abundance of 27.62% ± 24.95%, followed by invertebrate viruses (25.36% ± 23.55%). High-risk viruses capable of infecting primates, mammals, and plants accounted for an average relative abundance of 21.03% ± 25.25%. Figure 6 B).
[0136] The specific operation of this embodiment is as follows:
[0137] 1) The RNAvirHOST software (v1.0.5) was used to perform host infectability analysis on non-redundant RNA virus operational taxonomic units (vOTUs);
[0138] 2) After classifying the non-redundant RNA virus operational taxonomic units (vOTUs), the family-level classification results are entered into the ICTV (https: / / ictv.global / ) website, which provides the classification of infectable hosts for that family, thereby obtaining the ICTV host characteristics;
[0139] 3) The obtained host results are aggregated to finally obtain the host characteristics of non-redundant RNA virus operational taxonomic units (vOTUs).
Claims
1. A method for analyzing RNA viruses in samples, comprising the following steps: 1) Metatranscriptomics sequencing was performed on the samples to obtain raw metatranscriptomics data. The raw data underwent quality control analysis to obtain valid reads. The valid reads were assembled into contigs. Then, the metatranscriptomics data were analyzed using the RdRp-based homology iteration method and the geNomad method based on a machine learning model to screen for RNA virus sequences in the samples and obtain RNA virus contigs in the samples. The RNA virus contigs obtained by the homology iteration method and the geNomad method were merged and clustered to obtain RNA virus operational taxonomic units and their sequences. 2) Perform abundance analysis, classification identification, and host analysis on the RNA virus operational taxonomic units obtained in step 1); The classification and identification process includes analyzing the RNA virus operational taxonomic unit sequences using two methods, geNomad and CAT, to obtain the classification results of the RNA virus operational taxonomic units in the sample.
2. The method according to claim 1, wherein, The host analysis includes analyzing the RNA virus operational taxonomic unit sequence obtained in step 1) using the RNAvirHOST method to obtain the host information of the RNA virus in the sample.
3. The method according to claim 1 or 2, wherein, The abundance analysis includes performing abundance analysis on the RNA virus operational taxonomic unit sequences obtained in step 1) using the coverM tool to calculate the relative abundance of RNA viruses in the sample.
4. The method according to claim 1 or 2, wherein, Step 2) also includes performing checkV quality control on the RNA virus operational taxonomic unit sequence obtained in step 1) to reveal the length, contamination level, and integrity characteristics of the RNA virus sequence in the sample.
5. The method according to claim 1 or 2, wherein, Step 2) also includes analyzing the RNA virus operational taxonomic unit sequence obtained in step 1) using VIBRANT software to predict the life history of the RNA virus in the sample.
6. The method according to claim 1, wherein, The samples included soil, water, atmospheric particulate matter, compost, and sludge.
Citation Information
Patent Citations
Multi-omics sequencing and analyzing method based on Nanopore sequencing technology
CN111455031A
Virus genome identification and splicing method and application
CN116072222A