A method and system for identifying foreign viruses based on metagenomic sequencing
By employing a multi-dimensional identification system based on metagenomic sequencing and parallel processing technology, the problems of long detection cycles, low sensitivity, and insufficient throughput in traditional exogenous virus detection have been solved, enabling rapid and accurate detection of exogenous viruses, which is suitable for large-scale biological product testing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CEFETY BIOSCIENCE
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-24
AI Technical Summary
Traditional methods for detecting exogenous viruses have long detection cycles, low efficiency, and limited sensitivity, making it difficult to meet the needs of modern biological products for rapid, broad-spectrum, and high-sensitivity detection. They cannot effectively detect low-copy, low-titer, latent infections, and unknown viruses. They are also cumbersome to operate, have highly subjective results, low throughput, and high costs, making them unsuitable for large-scale testing scenarios.
Using metagenomic sequencing, a multi-dimensional identification system was constructed through data quality control, host removal, species identification, parameter-free assembly, and contig identification and verification. Multi-process parallel processing was introduced to generate standardized identification reports. Techniques such as Kraken2, RVDB/nt_core database alignment, and Virsorter2/CheckV verification were integrated to achieve high sensitivity and broad-spectrum detection.
It enables rapid and accurate detection of exogenous viruses, shortens the detection cycle, improves detection efficiency, is compatible with high-throughput detection of large-scale cell banks and biological products, identifies unknown viruses, reduces costs, and has scientific research and practical value.
Smart Images

Figure SMS_1 
Figure SMS_2
Abstract
Description
Technical Field
[0001] This invention relates to the fields of bioinformatics and pathogen detection, specifically to a method and system for identifying exogenous viruses based on metagenomic sequencing. Background Technology
[0002] With the rapid development of biomedical technology, cell therapy, gene therapy, and recombinant protein drugs have become core areas of biopharmaceutical research and development. The production of these biopharmaceuticals involves a large number of live cell operations, such as Chinese hamster ovary cells (CHO), human embryonic kidney cells (HEK293), and various engineered cell lines. Exogenous viral contamination is a major hidden danger to the safety of biopharmaceuticals, which may not only affect the quality and efficacy of the products, but also pose serious health risks to patients. Therefore, establishing a highly sensitive and broad-spectrum method for detecting exogenous viruses is a key link in ensuring the safety of biopharmaceuticals. Currently, the detection of exogenous viruses in cell banks mainly follows the International Council for Harmonisation of Technical Requirements for Human Rights (ICH) Q5A guidelines and relevant pharmacopoeia standards. Traditional detection methods mainly include in vivo methods, i.e., animal inoculation experiments, and in vitro methods, i.e., cell culture experiments. Among them, the in vitro culture method involves lysing the sample to be tested and inoculating it into various indicator cells, such as Vero, MRC-5, and CHO, and then continuously passaged them for 28 days to observe the cytopathic effect (CPE). The endpoint methods such as hemagglutination, hemagglutination, and fluorescent antibody detection are combined to determine viral contamination. This method can directly prove the presence of the virus and its infectivity. However, traditional detection methods still have significant limitations and can no longer fully meet the needs of modern biopharmaceuticals for rapid, broad-spectrum, and highly sensitive detection. Traditional methods have long detection cycles and low efficiency. In vitro cell culture usually requires 28 days or even longer of continuous passage and observation, which seriously slows down the research and development and production progress and makes it impossible to achieve rapid screening and release. Their sensitivity is limited, and false negatives are prone to occur for viruses with low copy numbers, low titers, latent infections, and those inhibited by the sample matrix, making it difficult to meet the high safety requirements of biopharmaceuticals. The detection range is highly dependent on the virus's proliferation ability and cytopathic effect on indicator cells. It is difficult to effectively detect viruses that cannot be cultured, grow slowly, are unknown, or are non-cytopathic viruses, posing a risk of unknown pathogen contamination. At the same time, the operation is cumbersome, the result interpretation is highly subjective, and the repeatability is poor, making it difficult to achieve standardization and automation. Moreover, it requires the use of multiple indicator cells and multiple batches of repeated experiments, resulting in low throughput and high cost. It cannot be adapted to the high-throughput detection scenarios of large-scale cell banks, intermediate products, and finished products, and there is room for improvement. Summary of the Invention
[0003] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a method and system for identifying exogenous viruses based on metagenomic sequencing, thereby resolving the aforementioned problems.
[0004] (II) Technical Solution To achieve the above objectives, at least one embodiment of this disclosure provides a method for identifying exogenous viruses based on metagenomic sequencing, comprising the following steps: S1. Data quality control step: The raw metagenomic sequencing data of the sample is filtered for quality to obtain clean sequencing data; S2, Host Removal Step: The clean sequencing data is compared with the host reference genome and ribosomal RNA database, and the host sequences that are matched are removed to obtain the enriched microbial sequences; S3. Species identification step: Use a virus database to identify the species of the microbial sequences and preliminarily screen out virus sequences. S4. Parameterless assembly step: The initially screened viral sequences are assembled without parameters to generate viral contig sequences. S5. Contiguous group identification and verification step: The viral contiguous group sequence is identified and verified in multiple dimensions, including: The viral contig sequences were compared with viral nucleic acid databases and viral protein databases, respectively, to screen for positive contigs; The positive contigs were compared with a general nucleic acid database to determine the species attribution; The positive contigs were graded in quality based on their consistency and coverage to obtain virus identification results with different confidence levels. S6. Results Integration Step: Integrate the analysis results from the previous steps to generate an exogenous virus identification report.
[0005] For example, in a method for identifying exogenous viruses based on metagenomic sequencing provided in at least one embodiment of this disclosure, the steps of S5, contig identification and verification, further include: using virus screening tools and virus integrity assessment tools to assess the viral attributes and integrity of the positive contigs.
[0006] For example, in a method for identifying exogenous viruses based on metagenomic sequencing provided in at least one embodiment of this disclosure, in step S5, contig identification and verification, the quality grading criteria are as follows: contigs with alignment consistency ≥90% and coverage ≥90% are judged as high confidence; contigs with alignment consistency ≥50% and coverage ≥50% are judged as medium confidence; and contigs that do not meet the medium confidence standard are judged as low confidence.
[0007] For example, in a method for identifying exogenous viruses based on metagenomic sequencing provided in at least one embodiment of this disclosure, the S5 contig identification and verification step further includes potential new virus mining: screening out viral contig sequences other than those of high-confidence species, excluding sequences homologous to known high-confidence viral species, and using the remaining sequences as potential new virus sequences.
[0008] For example, in a method for identifying exogenous viruses based on metagenomic sequencing provided in at least one embodiment of this disclosure, the S5 contig identification and verification step further includes supplementary verification: for viruses detected in the species identification step but not in the contig identification and verification step, the sequencing sequence of the virus is extracted for comparison and verification, and the number of sequences per million sequencing sequences of the virus sequence is calculated, and virus sequences with a sequence number ≥ 1 per million sequencing sequences are screened.
[0009] For example, in at least one embodiment of the present disclosure, a method for identifying exogenous viruses based on metagenomic sequencing is provided, wherein the S1 data quality control step uses the FASTP tool, and the filtering parameters include: enabling front-end trimming and back-end trimming, an average front-end trimming quality threshold of 20, a maximum error rate of 20, a maximum number of unknown bases of 30, and a sequence length threshold of 50 bp; the S2 host removal step uses the Kneaddata tool to simultaneously align the host reference genome and the Silva database; the S3 species identification step uses the Kraken2 tool with a preset virus database as a reference; and the S4 parameter-free assembly step uses the Megahit tool.
[0010] For example, in a method for identifying exogenous viruses based on metagenomic sequencing provided in at least one embodiment of this disclosure, in step S5, contig identification and verification: the blastn tool is used to perform alignment with the RVDB viral nucleic acid database as a reference; the diamond tool is used to perform blastx alignment with the RVDB viral protein database as a reference; the blastn tool is used to verify species attribution with the nt_core universal nucleic acid database; the virsorter2 tool is used for virus screening; and the checkV tool is used for virus integrity and contamination assessment.
[0011] For example, in a method for identifying exogenous viruses based on metagenomic sequencing provided in at least one embodiment of this disclosure, the method further includes parallel processing: multi-process parallel processing of sequencing data of multiple samples, while simultaneously performing the data quality control step to the result integration step.
[0012] According to another aspect of the present invention, a system for identifying exogenous viruses based on metagenomic sequencing is also provided, comprising: The data quality control module is configured to perform quality filtering on the raw sequencing data to obtain clean data. The host removal module is configured to compare the clean data with the host reference genome and ribosomal RNA database, remove the host sequence, and obtain the microbial sequence. The species identification module is configured to use a virus database to identify microbial sequences and perform preliminary screening of virus sequences. The parameterless assembly module is configured to perform parameterless assembly on the initially screened viral sequences to generate viral contigs. The contig identification and verification module is configured to perform multi-dimensional identification and verification of viral contigs, including: screening positive contigs by comparing with viral nucleic acid databases and viral protein databases respectively, determining species attribution by comparing with a general nucleic acid database, and classifying quality based on comparison consistency and coverage. The results integration module is configured to integrate the results from various modules and generate an external virus identification report.
[0013] For example, in a metagenomic sequencing-based exogenous virus identification system provided in at least one embodiment of this disclosure, the contig identification and verification module is further configured to: evaluate the positive contigs using virus screening tools and virus integrity assessment tools; and perform potential new virus mining and supplementary verification, wherein the supplementary verification includes calculating the number of sequences per million sequencing sequences of the virus sequence and screening sequences with a value ≥1, and the system further includes a parallel processing module configured to perform multi-process parallel processing on the sequencing data of multiple samples.
[0014] (III) Beneficial Effects Compared with existing technologies, this invention provides a method and system for identifying exogenous viruses based on metagenomic sequencing, which has the following advantages: This invention integrates multiple steps, including Kraken2 species identification, RVDB / nt_core database alignment, and Virsorter2 / CheckV verification, to form a multi-dimensional identification system that improves the accuracy and reliability of virus identification. It introduces a multi-process parallel processing mechanism to support simultaneous analysis of batch samples, significantly improving analytical efficiency and adapting to the high-throughput detection needs of large-scale cell banks, intermediate products of biological products, and finished products. This solves the problem of low throughput in traditional methods and existing mNGS methods. Furthermore, it integrates key indicators from each analytical step to generate standardized identification reports, while simultaneously identifying potential new viruses and providing clues for subsequent research, thus balancing practicality and scientific research value. Detailed Implementation
[0015] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0016] It should be understood that in the various embodiments of the present invention, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0017] It should be understood that in this invention, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or device.
[0018] The applicant's research revealed that traditional detection methods still have significant limitations and can no longer fully meet the needs of modern biopharmaceuticals for rapid, broad-spectrum, and highly sensitive detection. Traditional methods suffer from long detection cycles and low efficiency; in vitro cell culture typically requires 28 days or even longer of continuous passage and observation, severely slowing down R&D and production progress and preventing rapid screening and release. Their sensitivity is limited, easily producing false negatives for low-copy, low-titer, latent, and sample matrix-inhibited viruses, failing to meet the high safety requirements of biopharmaceuticals. The detection range is highly dependent on the virus's proliferation ability and cytopathic effect on indicator cells, making it difficult to effectively detect unculturable, slow-growing, unknown, and non-cytopathic viruses, posing a risk of unknown pathogen contamination. Furthermore, the methods are cumbersome, result interpretation is highly subjective, and reproducibility is poor, making standardization and automation difficult. They also require multiple indicator cells and multiple batches of repeated experiments, resulting in low throughput and high cost, making them unsuitable for high-throughput detection scenarios involving large-scale cell banks, intermediate products, and finished products, indicating room for improvement. Addressing the core pain points of traditional detection methods, such as long cycle time, low sensitivity, narrow detection range, strong subjectivity, and insufficient throughput, the applicant, relying on the development trend of high-throughput sequencing technology, has carried out targeted technology research and development and solution design. The increasingly mature metagenomic next-generation sequencing (mNGS) technology has been introduced into the field of exogenous virus detection in biological products. The aim is to break through the technical bottlenecks of traditional methods and build a more efficient, accurate, and comprehensive exogenous virus analysis system. Compared with traditional detection methods, mNGS technology has significant technical advantages. Its core feature is that it does not require pre-assumptions about the type and type of pathogen, nor does it rely on virus culture or known virus characteristics. It can perform unbiased, high-throughput sequencing of all nucleic acids in the sample, and then compare, analyze, and annotate the obtained sequencing data with known microbial databases. Theoretically, it can screen all known viruses in the sample at once, while effectively identifying unknown viruses, greatly expanding the detection range. It perfectly solves the shortcomings of traditional methods in detecting unknown viruses and non-cytopathic viruses, demonstrating its great potential to replace traditional culture methods and achieve efficient screening of exogenous viruses. Therefore, the applicant has specifically optimized the bioinformatics analysis process of mNGS technology to develop a highly sensitive and accurate method for analyzing exogenous viruses that fully utilizes the inherent advantages of mNGS technology, such as high throughput, unbiasedness, and broad-spectrum screening. This method optimizes the bioinformatics analysis process to achieve precise removal of host nucleic acid interference, efficient mining of low-abundance viral signals, and accurate annotation of viral information. Its development can not only effectively make up for many shortcomings of traditional detection methods, but also fill the application gap of existing mNGS technology in the bioinformatics analysis layer of exogenous virus detection in biological products.
[0019] It should be noted that the method of the present invention is applicable to a variety of biological product matrices, including but not limited to animal cell banks such as CHO cells, HEK293 cells, and Vero cells and their derivatives.
[0020] The following specific embodiments illustrate a method and system for identifying exogenous viruses based on metagenomic sequencing provided in this disclosure.
[0021] Example 1 This embodiment uses a metagenomic sequencing-based method and system for identifying exogenous viruses according to the present invention to detect exogenous viruses in six CHO cell samples. Three of the samples, PC-1, PC-2, and PC-3, contain low concentrations of five model viruses: parvovirus MVM, reovirus Reo3, respiratory syncytial virus HRSV, feline leukemia virus FeLV, and Epstein-Barr virus EBV. The other three samples, NC-1, NC-2, and NC-3, are CHO cell samples without any viruses and serve as negative control samples.
[0022] Total nucleic acid was extracted from 6 samples, and sequencing libraries were constructed. Then, paired-end sequencing was performed using a metagenomic next-generation sequencing platform. The average data volume of each sample was 30 Gb. After sequencing, the raw sequencing data of each sample were obtained in FASTQ format.
[0023] According to the method and system for identifying exogenous viruses based on metagenomic sequencing of the present invention, the data quality control step is initiated, and the raw sequencing data of each sample is filtered for quality using the FASTP tool. The specific parameter settings are as follows: Enable front-end trimming (--cut_front) and back-end trimming (--cut_tail). Set the average quality threshold for front-end trimming to 20, the maximum error rate (-e) to 20, the maximum number of unknown bases (-n) to 30, and the sequence length threshold (--length_required) to 50 bp, i.e., retain sequences with a length ≥ 50 bp.
[0024] After quality filtering is completed, the FASP tool automatically generates a quality report file in JSON format. The system reads the JSON file, extracts key quality indicators, and constructs a quality statistics table. The results are shown in Table 1. The clean data Q20 ratio of all 6 samples is ≥98.58%, the Q30 ratio is ≥95.12%, and the average sequence length is between 135-139 bp. The quality control is qualified and meets the requirements for subsequent analysis.
[0025] Table 1. Sample data quality control and host removal results According to the present invention, a method and system for identifying exogenous viruses based on metagenomic sequencing is used to remove host sequences from clean sequencing data. This tool simultaneously compares the CHO cell host reference genome and the Silva ribosomal RNA database, removing sequences that match the host genome or ribosomal RNA and retaining unmatched sequences, i.e., potential viral and other microbial sequences.
[0026] After the host sequence removal was completed, the paired clean sequence files were extracted from the output directory of the kneaddata tool and used as input data for subsequent species identification and virus assembly. As shown in Table 1, the CHO host sequence removal rate of the 6 samples was ≥99.25%, which effectively eliminated the interference of host nucleic acid during the production process and ensured the effective enrichment of virus sequences.
[0027] According to the present invention, a method and system for identifying exogenous viruses based on metagenomic sequencing is used. The Kraken2 tool is employed, with a pre-defined virus database as a reference, to identify species from sequencing data after host sequence removal. The running time of Kraken2 species identification is recorded, log information is generated, and the identification results are initially screened, retaining species classified as "Viruses" to provide a target direction for subsequent virus sequence assembly and validation.
[0028] According to the present invention, a method and system for identifying exogenous viruses based on metagenomic sequencing, firstly, the NCBI classification number of the virus is used as the target, and a list of all relevant classification IDs corresponding to the classification number is extracted from the Kraken2 species identification report; then, based on the list of classification IDs, the corresponding sequencing sequence IDs are extracted from the Kraken2 alignment result output file, and the corresponding paired-end sequencing sequences are further screened to generate a virus sequence file.
[0029] The extracted viral sequences were assembled without parameters using the megahit tool, splicing short read sequencing sequences into longer contig sequences. After assembly, the final contig sequence file was generated, along with a megahit.log file, which recorded key operations, running parameters, and assembly results during the assembly process.
[0030] According to the present invention, a method and system for identifying exogenous viruses based on metagenomic sequencing is used to perform multi-dimensional identification and verification of Contig sequences generated by parameter-free assembly.
[0031] First, the blastn tool was used to perform alignment analysis of the Contig sequences against the RVDB viral nucleic acid database. Simultaneously, the diamond tool was used to perform blastx alignment against the RVDB viral protein database to supplement protein identification information. Contig sequences that aligned positively against the RVDB database were screened out and then aligned against the nt_core general nucleic acid database using the blastn tool to further verify the species attribution of the Contig sequences. At the same time, the NCBI classification number and species name were extracted from the alignment results.
[0032] Subsequently, the virsorter2 tool was used to screen the Contig sequences for viruses, extracting information such as virus scores and hallmarks. The checkV tool was used to assess the integrity and contamination of the Contig sequences, extracting information such as the number of viral genes, the number of host genes, and the degree of contamination.
[0033] The alignment results from the RVDB database, the nt_core database, the virsorter2 validation results, and the checkV validation results were merged with the Contig basic information table to integrate key indicators such as the length of the Contig sequence, alignment consistency, coverage, number of viral genes, number of host genes, and species information.
[0034] Contig sequences are graded for quality based on alignment consistency and coverage: Overlapping groups with an alignment consistency of ≥90% and a coverage of ≥90% are classified as high confidence (HQ). Overlapping groups with an alignment consistency of ≥50% and a coverage of ≥50% are classified as medium confidence (MQ). The judgment that does not meet the medium confidence level standard is low confidence (LQ).
[0035] By grading the quality, high-confidence viral contig sequences are preferentially retained.
[0036] By combining the alignment results from the RVDB and nt_core databases, the final species attribution of the Contig sequence was determined. For sequences whose alignment results from RVDB and nt_core are inconsistent, the RVDB alignment result was given priority.
[0037] In terms of identifying potential new viruses, this embodiment screens out Contig sequences other than those of high-confidence species, excludes sequences homologous to known high-confidence virus species, and uses the remaining sequences as potential new virus sequences. In the positive control samples of this embodiment, all detected viruses are known model viruses, and no potential new viruses were found.
[0038] For supplementary validation, if the viral sequence identified by Kraken2 is not detected in Contig, the sequencing sequence of the virus is further extracted for BLASTN alignment validation. At the same time, the RPM value of the viral sequence is calculated, and the number of viral sequences per million sequencing sequences is used to screen viral sequences with RPM≥1 to improve the detection reliability of low-abundance viruses.
[0039] The analysis results from each of the above steps are integrated to form a complete exogenous virus analysis report, linking key indicators of each step, such as the amount of clean data, host sequence removal rate, number of virus sequences, and contig quality level, to form a complete analysis chain.
[0040] Test results: As shown in Table 2, all five low-concentration model viruses—MVM, Reo3, HRSV, FeLV, and EBV—were successfully detected in the three samples PC-1, PC-2, and PC-3 incorporating model viruses, with no missed detections. The RPM values of each virus sequence were clearly distinguishable, and the confidence levels were all high, indicating that the method of this invention has high detection sensitivity and can meet the requirements for low-concentration virus detection.
[0041] Table 2 Virus detection results of positive control samples The three negative control samples NC-1, NC-2, and NC-3, which were not contaminated with any virus, did not detect any exogenous virus and had no false positives. Combined with the production scenario characteristics of unprocessed CHO cell harvest liquid, the method of this invention can effectively eliminate interference factors such as carriers and bacteria in the production process.
[0042] This embodiment employs a parallel processing module to process the sequencing data of six samples in a multi-process parallel manner, simultaneously executing data quality control steps to result integration steps. The system creates six process pools, with each module running collaboratively. The total running time is approximately 540 minutes, and the average running time for a single sample is approximately 90 minutes. Compared to traditional in vitro culture methods, the detection cycle is significantly shortened, demonstrating the significant advantages of parallel processing, which can meet the rapid detection needs of large-scale production samples.
[0043] During operation of this embodiment, the database management module can quickly call up the dedicated filtering database to effectively remove production-related interference sequences; the contiguous group identification and verification module can accurately screen viral sequences and eliminate false positives; and the parallel processing module can reasonably allocate resources to avoid interference between samples and ensure detection efficiency and accuracy.
[0044] It should be noted that this embodiment also verifies the exogenous virus identification system of the present invention. The system includes a data quality control module for performing data quality control steps; a host removal module for performing host removal steps; a species identification module for performing species identification steps; a parameter-free assembly module for performing parameter-free assembly steps; a contiguous group identification and verification module for performing contiguous group identification and verification steps; a result integration module for performing result integration steps; and a parallel processing module for scheduling system resources to achieve multi-process parallel processing. Through the collaborative work of the above modules, efficient detection of 6 CHO cell samples was achieved, proving that the system can realize the exogenous virus identification method of the present invention.
[0045] In summary, the present invention provides a method and system for identifying exogenous viruses based on metagenomic sequencing. By constructing a complete analysis chain, particularly through multi-dimensional identification and verification mechanisms such as multi-database comparison, quality grading, potential new virus mining, and supplementary verification, it achieves high sensitivity, high accuracy, and broad-spectrum detection of exogenous viruses in bioproducts. Examples demonstrate that the method and system of the present invention can be effectively adapted to sample detection in bioproduct production matrices such as CHO cells, significantly shortening the detection cycle to several hours. It can be stably applied to actual production scenarios, accurately detecting potential viral contamination and providing reliable support for quality control in the bioproduct production process. It solves the pain points of traditional detection methods, such as long cycles, complex operations, and narrow detection spectra, and has broad practical application value.
[0046] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for identifying exogenous viruses based on metagenomic sequencing, characterized in that, Includes the following steps: S1. Data quality control step: The raw metagenomic sequencing data of the sample is filtered for quality to obtain clean sequencing data; S2, Host Removal Step: The clean sequencing data is compared with the host reference genome and ribosomal RNA database, and the host sequences that are matched are removed to obtain the enriched microbial sequences; S3. Species identification step: Use a virus database to identify the species of the microbial sequences and preliminarily screen out virus sequences. S4. Parameterless assembly step: The initially screened viral sequences are assembled without parameters to generate viral contig sequences. S5. Contiguous group identification and verification step: The viral contiguous group sequence is identified and verified in multiple dimensions, including: The viral contig sequences were compared with viral nucleic acid databases and viral protein databases, respectively, to screen for positive contigs; The positive contigs were compared with a general nucleic acid database to determine the species attribution; The positive contigs were graded in quality based on their consistency and coverage to obtain virus identification results with different confidence levels. S6. Results Integration Step: Integrate the analysis results from the previous steps to generate an exogenous virus identification report.
2. The method for identifying exogenous viruses based on metagenomic sequencing according to claim 1, characterized in that, The S5 step of contig identification and verification also includes: using virus screening tools and virus integrity assessment tools to assess the virus attributes and integrity of the positive contigs.
3. The method for identifying exogenous viruses based on metagenomic sequencing according to claim 1, characterized in that, In step S5, the contiguous group identification and verification, the quality grading standard is as follows: contiguous groups with alignment consistency ≥ 90% and coverage ≥ 90% are judged as high confidence; contiguous groups with alignment consistency ≥ 50% and coverage ≥ 50% are judged as medium confidence; and contiguous groups that do not meet the medium confidence standard are judged as low confidence.
4. The method for identifying exogenous viruses based on metagenomic sequencing according to claim 3, characterized in that, The S5 contig identification and verification steps also include potential new virus mining: screening out viral contig sequences other than those of high-confidence species, excluding sequences homologous to known high-confidence viral species, and using the remaining sequences as potential new virus sequences.
5. The method for identifying exogenous viruses based on metagenomic sequencing according to claim 1, characterized in that, The S5 contig identification and verification step also includes supplementary verification: for viruses detected in the species identification step but not in the contig identification and verification step, the sequencing sequence of the virus is extracted for comparison and verification, and the number of sequences per million sequencing sequences of the virus is calculated, and virus sequences with a sequence number ≥ 1 per million sequencing sequences are screened.
6. The method for identifying exogenous viruses based on metagenomic sequencing according to claim 1, characterized in that, The S1 data quality control step uses the FASTP tool, with filtering parameters including: enabling front-end and back-end trimming, an average front-end trimming quality threshold of 20, a maximum error rate of 20%, a maximum number of unknown bases of 30, and a sequence length threshold of 50 bp. The S2 host removal step uses the Kneaddata tool to simultaneously align the host reference genome with the Silva database. The S3 species identification step uses the Kraken2 tool with a preset virus database as a reference. The S4 parameter-free assembly step uses the Megahit tool.
7. The method for identifying exogenous viruses based on metagenomic sequencing according to claim 1, characterized in that, In step S5, the contig identification and verification: the blastn tool is used to perform alignment with the RVDB viral nucleic acid database as a reference; the diamond tool is used to perform blastx alignment with the RVDB viral protein database as a reference; the blastn tool is used to verify species attribution with the nt_core universal nucleic acid database; the virsorter2 tool is used for virus screening; and the checkV tool is used for virus integrity and contamination assessment.
8. The method for identifying exogenous viruses based on metagenomic sequencing according to claim 1, characterized in that, The method also includes parallel processing: sequencing data of multiple samples are processed in parallel using multiple processes, while the data quality control step and the result integration step are performed simultaneously.
9. A system for identifying exogenous viruses based on metagenomic sequencing, characterized in that, include: The data quality control module is configured to perform quality filtering on the raw sequencing data to obtain clean data. The host removal module is configured to compare the clean data with the host reference genome and ribosomal RNA database, remove the host sequence, and obtain the microbial sequence. The species identification module is configured to use a virus database to identify microbial sequences and perform preliminary screening of virus sequences. The parameterless assembly module is configured to perform parameterless assembly on the initially screened viral sequences to generate viral contigs. The contig identification and verification module is configured to perform multi-dimensional identification and verification of viral contigs, including: screening positive contigs by comparing with viral nucleic acid databases and viral protein databases respectively, determining species attribution by comparing with a general nucleic acid database, and classifying quality based on comparison consistency and coverage. The results integration module is configured to integrate the results from various modules and generate an external virus identification report.
10. The exogenous virus identification system based on metagenomic sequencing according to claim 9, characterized in that, The contig identification and verification module is further configured to: evaluate the positive contigs using virus screening tools and virus integrity assessment tools; and to perform potential new virus mining and supplementary verification, wherein the supplementary verification includes calculating the number of sequences per million sequencing sequences of the virus sequence and screening sequences with a value ≥1. The system also includes a parallel processing module configured to perform multi-process parallel processing on sequencing data of multiple samples.