Method for identifying dna viruses in an organism based on high-throughput sequencing data and applications thereof

By performing quality control and comparative analysis on high-throughput sequencing data, the problems of inaccuracy and high cost in DNA virus identification have been solved, enabling accurate identification and efficient utilization of sequencing data, and improving the accuracy and efficiency of detection.

CN118398077BActive Publication Date: 2026-05-12SHENZHEN HAPLOX BIOTECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN HAPLOX BIOTECH
Filing Date
2024-04-01
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies, the interpretation of high-throughput sequencing data lacks a unified standard, resulting in strong subjectivity, high false positives, and waste of resources in DNA virus identification, and second-generation sequencing data is not fully utilized.

Method used

By performing quality control on high-throughput sequencing data, comparing with the NCBI database, removing overlapping parts of host and viral sequences, using BLAST software for high-quality comparison, setting sequence thresholds to determine the presence of DNA viruses, and constructing an identification device for identification.

Benefits of technology

It enables accurate identification of DNA viruses, reduces false positives, improves the utilization rate of sample sequencing data, and shortens the detection time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118398077B_ABST
    Figure CN118398077B_ABST
Patent Text Reader

Abstract

The application discloses a method for identifying DNA viruses in organisms based on high-throughput sequencing data, and the method comprises the following steps: taking a biological sample derived from the organism as a to-be-tested sample, performing high-throughput sequencing on the to-be-tested sample, and obtaining sequencing data in a FastQ file format; performing analysis and identification by using the sequencing data; and then identifying the DNA viruses in the to-be-tested sample. The method does not need to rely on metagenomic sequencing, can accurately identify the DNA viruses by using sequencing data of second-generation sequencing (high-throughput sequencing), avoids the generation of false positives, and greatly improves the utilization rate of sample sequencing data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedical technology, specifically to a method for identifying DNA viruses in organisms based on high-throughput sequencing data and its application. Background Technology

[0002] DNA viruses are a class of biological viruses containing deoxyribonucleic acid (DNA) genetic material, widely distributed in humans, vertebrates, insects, plants, and microorganisms. According to the latest ICTV international virus classification standard, DNA viruses can be classified by host type into vertebrate DNA viruses, plant DNA viruses, invertebrate DNA viruses, prokaryotic microorganism (bacteria and archaea) DNA viruses, and eukaryotic microorganism (algae, fungi, and protozoa) DNA viruses. Based on the type of DNA nucleic acid, DNA viruses are further divided into double-stranded DNA viruses (dsDNA) and single-stranded DNA viruses (ssDNA).

[0003] Some DNA viruses are pathogenic and tumorigenic to humans, and are closely related to human health. Furthermore, some DNA viruses can be transmitted from mother to child, thus identifying and detecting the specific types and subtypes of DNA viruses carried by humans can effectively prevent diseases caused by DNA viruses.

[0004] As a type of microorganism, DNA viruses are often detected and identified using metagenomic next-generation sequencing (mNGS) when the specific virus type contained in a sample is unknown. mNGS does not rely on traditional microbial culture and can extract nucleic acid information from all microorganisms in a sample for high-throughput sequencing without bias. Using the microbial community genome in the sample as the research object, combined with bioinformatics analysis, host information is removed before comparison with pathogen databases to obtain information on the type of pathogenic microorganism. However, the interpretation of mNGS reports is often subjective, and the standards for interpreting mNGS are limited. There is a lack of unified clinical standards for sequence thresholds, sensitivity assessment, and specificity evaluation. Furthermore, metagenomic sequencing is expensive and time-consuming.

[0005] Second-generation sequencing (high-throughput sequencing) is usually used for deep sequencing of the genome of a single organism and is widely used for somatic cell variation detection and differential gene expression detection. The data from second-generation sequencing is not used for gene variation analysis, which results in a waste of sequencing data.

[0006] Therefore, there is an urgent need for a method that can use sequencing data from second-generation sequencing (high-throughput sequencing) for DNA virus identification. Summary of the Invention

[0007] The purpose of this invention is to overcome the above-mentioned shortcomings of the prior art and provide a method for identifying DNA viruses in organisms based on high-throughput sequencing data and its application.

[0008] The first objective of this invention is to provide a method for identifying DNA viruses in organisms based on high-throughput sequencing data.

[0009] A second objective of this invention is to provide a device for identifying DNA viruses in organisms based on high-throughput sequencing data.

[0010] A third objective of this invention is to provide the application of the above-described method and / or apparatus in the identification of DNA viruses in organisms.

[0011] To achieve the above objectives, the present invention is implemented through the following solution:

[0012] A method for identifying DNA viruses in an organism based on high-throughput sequencing data, using a biological sample derived from the organism as the test sample, includes the following steps:

[0013] S1. Perform high-throughput sequencing on the sample to be tested, collect sequencing data in FastQ format from the sequencing data, perform quality control on the sequencing data, and obtain clean fq data;

[0014] S2. Obtain the sequence information of the reference genome of the organism from the NCBI database to form an organism database;

[0015] The genomic sequence information of DNA viruses whose host organism is described is obtained from the NCBI database and used as a DNA virus database.

[0016] S3. Align the clean fq data obtained in step S1 with the organism database shown in step S2 to obtain sequences that cannot be aligned with the organism database. Align the clean fq data obtained in step S1 with the DNA virus database shown in step S2 to obtain sequences aligned with the DNA virus database. Remove sequences that coexist with sequences that cannot be aligned with the organism database and sequences that coexist with sequences aligned with the DNA virus database. The sequences remaining after removing coexisting sequences from sequences that cannot be aligned with the organism database are recorded as unmap data. The sequences remaining after removing coexisting sequences from sequences aligned with the DNA virus database are recorded as virus_map data.

[0017] S4. Remove sequences with a clip ratio greater than 10% from the virus_map data obtained in step S3 to obtain virus_map_rc data;

[0018] S5. Select 9,000 to 11,000 sequences from the unmap data obtained in step S3 and the virus_map_rc data obtained in step S4, respectively, and denot them as unmap_some data and virus_map_rc_some data;

[0019] If the number of sequences in the unmap data obtained in step S3 or the virus_map_rc data obtained in step S4 is less than 10,000, then all sequences are selected;

[0020] S6. Summarize the whole genome sequence information of DNA viruses in the NCBI database as a comparison database, compare the unmap_some data obtained in step S5 with the comparison database, and obtain the number of high-quality aligned sequences of each DNA virus in the unmap_some data.

[0021] The virus_map_rc_some data obtained in step S5 is compared with the alignment database to obtain the number of high-quality aligned sequences of each DNA virus in the virus_map_rc_some data;

[0022] The high-quality alignment is defined as follows: the sequence in the unmap_some data and / or virus_map_rc_some data aligns to a sequence fragment in the alignment database that is at least 90% of the sequence length, and the sequence identity of the sequence fragment aligned to the alignment database is ≥90%.

[0023] S7. If the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S6 and / or the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data meet the judgment criteria, the sample to be tested contains the DNA virus.

[0024] If the number of high-quality aligned sequences of the DNA virus in the unmap_some data and / or the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data obtained in step S6 does not meet the judgment criteria, the sample to be tested does not contain the DNA virus.

[0025] The judgment criteria are: the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S6 is ≥1 and / or the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data is ≥1.

[0026] Preferably, the number of DNA virus reads in the test sample accounts for no less than 2% of the total number of reads in the test sample.

[0027] More preferably, the sample to be tested is a biological sample derived from a cancer patient.

[0028] More preferably, the tumor patient is an HPV patient, an HBV patient, an EBV patient, a Betapolyomavirushominis BK virus patient, and / or an HHV-8 virus patient.

[0029] More preferably, the biological sample is a digestive fluid, tissue fluid, and / or tissue.

[0030] More preferably, the biological sample is peripheral blood, exfoliated epithelial cells, and / or biological tissue.

[0031] Preferably, the high-throughput sequencing in step S1 is high-throughput sequencing with a sequencing depth > 30×.

[0032] Preferably, the quality control of sequencing data in step S1 specifically involves using Fastp software to perform quality control based on default parameters.

[0033] More preferably, the version number of the Fastp software is 0.19.7.

[0034] Preferably, the step S2 of obtaining the genome sequence information of a DNA virus whose host is the organism specifically involves: obtaining the genome sequence information of a DNA virus with complete nucleotide information, whose host is the organism, and whose genome molecular type is unknown DNA from the NCBI database.

[0035] Preferably, the comparison in step S3 is performed using the bwa mem software based on default parameters.

[0036] More preferably, the version number of the bwa mem software is 0.7.17.

[0037] Specifically, step S4 involves removing sequences with a clip percentage greater than 10% from the virus_map data obtained in step S3, which means removing sequences in the virus_map data where the cigar value is "S" or "H" and the percentage is greater than 10%.

[0038] Preferably, in step S5, 10,000 sequences are selected from the unmap data obtained in step S3 and the virus_map_rc data obtained in step S4, respectively.

[0039] Preferably, the comparison in step S6 is performed using BLAST software based on default parameters.

[0040] More preferably, the version number of the blast software is 2.12.0+.

[0041] The judgment condition in step S7 is: the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S6 is ≥100 and the proportion of the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S5 is ≥5% and / or the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data is ≥100 and the proportion of the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data obtained in step S5 is ≥5%;

[0042] Alternatively, the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S6 is ≥10 and the proportion of the high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S5 is ≥0.5%, and / or the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data is ≥10 and the proportion of the high-quality aligned sequences of the DNA virus in the virus_map_rc_some data obtained in step S5 is ≥0.5%.

[0043] The present invention also claims protection for an apparatus for identifying DNA viruses in organisms based on high-throughput sequencing data, the apparatus comprising a data acquisition component, an identification component, and a result output component;

[0044] The data acquisition component is used to acquire sequencing data in FastQ format from the high-throughput sequencing data of the sample to be tested, wherein the sample to be tested is a biological sample derived from the organism.

[0045] The identification component uses the sequencing data acquired by the data acquisition component as input data and executes the method as described in any one of claims 1 to 6 to obtain the DNA virus identification results in the organism.

[0046] The result output component is used to output the DNA virus identification results in the organism obtained by the identification component.

[0047] Preferably, the sample to be tested is a biological sample derived from a cancer patient.

[0048] More preferably, the tumor patient is an HPV patient, an HBV patient, an EBV patient, a Betapolyomavirushominis BK virus patient, and / or an HHV-8 virus patient.

[0049] The present invention also claims protection for the use of any of the methods and / or devices described above in the identification of DNA viruses in organisms.

[0050] Preferably, the DNA virus is HPV virus, HBV virus, EBV virus, Betapolyomavirus hominisBK virus and / or HHV-8 virus.

[0051] More preferably, the DNA virus is HPV virus, HBV virus and / or EBV virus.

[0052] Compared with the prior art, the present invention has the following beneficial effects:

[0053] This invention provides a method for identifying DNA viruses in organisms based on high-throughput sequencing data. The method uses biological samples derived from the organism as test samples, performs high-throughput sequencing on the test samples to obtain sequencing data in FastQ format, and analyzes and identifies the DNA viruses in the test samples. This method does not rely on metagenomic sequencing; it uses sequencing data from second-generation sequencing (high-throughput sequencing) to accurately identify DNA viruses, avoids false positives, and greatly improves the utilization rate of sample sequencing data. Attached Figure Description

[0054] Figure 1 This is a flowchart of the method shown in Example 1. Detailed Implementation

[0055] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. These embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Unless otherwise specified, the experimental methods used in the following embodiments are conventional methods; the materials and reagents used, unless otherwise specified, are commercially available.

[0056] Example 1: A method for identifying DNA viruses in organisms based on high-throughput sequencing data

[0057] 1. Experimental Methods

[0058] When using a human biological sample (digestive fluid, tissue fluid, or tissue) as the test sample, and performing high-throughput sequencing at a sequencing depth of 150×, the number of DNA viral reads contained in the test sample accounts for >2% of the total number of DNA reads in the test sample. The flowchart of the method is as follows. Figure 1 As shown, the specific steps are as follows:

[0059] S1. Perform high-throughput sequencing on the sample to be tested with a sequencing depth of 150×, collect sequencing data in FastQ format from the sequencing data, and use Fastp software (version 0.19.7) to remove adapters and filter the sequencing data based on preset parameters to obtain clean fq data;

[0060] S2. Obtain the sequence information of the human hg19 reference genome from the NCBI database to serve as an organism database;

[0061] DNA virus genome sequence information was selected from the NCBI database to construct a DNA virus database.

[0062] The criteria for selecting DNA virus genome sequence information are: complete nucleotide information, human host, and unknown genome molecule type; in NCBI data, this is represented as "Nucleotide Completeness: complete; Host: human; genome molecule type: dna, unknown".

[0063] S3. The clean fq data obtained in step S1 is compared with the organism database shown in step S2 using bwa mem (version 0.7.17) based on preset parameters. Sequences whose flag does not contain 5 in the comparison results are considered as unaligned sequences, resulting in sequences that cannot be aligned with the organism database shown. The clean fq data obtained in step S1 is compared with the DNA virus database shown in step S2 using bwa mem (version 0.7.17) based on preset parameters. Sequences whose flag contains 4 in the comparison results are considered as aligned sequences, resulting in sequences aligned with the DNA virus database shown.

[0064] Sequences that cannot be matched with the shown organism database and sequences that are matched with the shown DNA virus database are removed from each other. The sequences that are matched with the shown organism database after removing the sequences that are matched from each other are denoted as unmap data; sequences that are matched with the shown DNA virus database after removing the sequences that are matched from each other are denoted as virus_map data.

[0065] S4. Remove sequences with a clip ratio greater than 10% from the virus_map data obtained in step S3 to obtain virus_map_rc data;

[0066] S5. Randomly select 10,000 sequences from the unmap data obtained in step S3 and the virus_map_rc data obtained in step S4, and denot them as unmap_some data and virus_map_rc_some data, respectively.

[0067] If the number of sequences in the unmap data obtained in step S3 and the virus_map_rc data obtained in step S4 is less than 10,000, then all sequences in the unmap data obtained in step S3 and the virus_map_rc data obtained in step S4 are selected.

[0068] S6. Summarize the whole genome sequence information of DNA viruses in the NCBI database as a comparison database. Use Blast (version 2.12.0+) to compare the unmap_some data obtained in step S5 with the comparison database based on preset parameters to obtain the number of high-quality aligned sequences of each DNA virus in the unmap_some data.

[0069] The virus_map_rc_some data obtained in step 5 is compared with the alignment database using Blast (version 2.12.0+) based on preset parameters to obtain the number of high-quality aligned sequences of each DNA virus in the virus_map_rc_some data;

[0070] High-quality alignments are defined as follows: sequences in unmap_some or virus_map_rc_some data that align to the alignment database have a length of at least 90% of the sequence length and a sequence identity ≥ 90% for the sequence fragments aligned to the alignment database.

[0071] S7. If the number of high-quality aligned sequences of the DNA virus in the unmap_some data or the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data obtained in step S6 meets the judgment condition, the sample to be tested contains the DNA virus.

[0072] The number of high-quality aligned sequences of the DNA virus in the unmap_some data and the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data obtained in step S6 do not meet the judgment criteria, and the sample to be tested does not contain the DNA virus.

[0073] When the DNA virus is HBV, HPV, or EBV, the judgment condition is: the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S6 is ≥1 or the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data is ≥1.

[0074] When the DNA virus is a DNA virus other than HBV, HPV, and EBV, the judgment condition is: the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S6 is ≥100 and the proportion of the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S5 is ≥5%; or the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data is ≥100 and the proportion of the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data obtained in step S5 is ≥5%.

[0075] Alternatively, the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S6 is ≥10, and the proportion of the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S5 is ≥0.5%, or the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data is ≥10, and the proportion of the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data obtained in step S5 is ≥0.5%.

[0076] Example 2: Validation of a method for identifying DNA viruses in organisms based on high-throughput sequencing data

[0077] 1. Experimental Methods

[0078] (1) Data Acquisition

[0079] Validation set data 1: Tumor tissue sections from 150 liver cancer patients (HBV antigen detection and clinical phenotype confirmed that all 150 liver cancer patients were infected with HBV virus) were subjected to whole exome sequencing at a depth of 150×. The sequencing data were collected as validation set data 1, which contains 150 HBV positive samples.

[0080] Validation set data 2: contains sequencing data of 235 samples, including gDNA sequencing data of 78 cervical cancer patients (HBV negative samples), gDNA sequencing data of 55 nasopharyngeal carcinoma patients (HPV negative samples), tumor tissue section sequencing data of 58 cervical cancer patients (HPV positive samples), and tumor tissue section sequencing data of 44 nasopharyngeal carcinoma patients (EBV positive samples); among them, the tissue section samples were sequenced at a sequencing depth of 500× whole exome sequencing, and the remaining samples were sequenced at a sequencing depth of 150× whole exome sequencing.

[0081] The tumor patients involved in validation set data 1 and validation set data 2 all came from Shenzhen Haplos Medical Laboratory, and patient authorization was obtained before using patient gDNA or tumor tissue sequencing.

[0082] (2) Virus identification

[0083] Following the method shown in Example 1, DNA virus identification was performed on each sample in validation set data 1 and validation set data 2 respectively, and the DNA virus identification results of each sample in validation set data 1 and validation set data 2 were obtained. The results were then compared with the actual virus results of each sample, and the accuracy of the method shown in Example 1 was calculated.

[0084] When performing DNA identification on validation set data 1, the purpose was to identify whether each sample in validation set data 1 contained HBV virus. When performing DNA identification on validation set data 2, the gDNA of 78 cervical cancer patients was identified as containing HBV virus, the gDNA of 55 nasopharyngeal carcinoma patients was identified as containing HPV virus, the tumor tissue sections of 58 cervical cancer patients were identified as containing HPV virus, and the tumor tissue sections of 44 nasopharyngeal carcinoma patients were identified as containing EBV virus.

[0085] 2. Experimental Results

[0086] Following the method shown in Example 1, each sample in validation set data 1 was identified. Among 150 samples, 6 false negative samples and 144 true positive samples were identified, with an accuracy rate of 96%.

[0087] Following the method described in Example 1, each sample in the validation set data 2 was identified. All 133 negative samples were identified as true negatives, and 95 of the 102 positive samples were identified as true positives and 7 as false negatives, with an accuracy rate of 97%.

[0088] Example 3: A device for identifying DNA viruses in organisms based on high-throughput sequencing data

[0089] 1. Composition

[0090] A device for identifying DNA viruses in organisms based on high-throughput sequencing data includes a data acquisition component, an identification component, and a result output component.

[0091] The data acquisition component is used to acquire sequencing data in FastQ format from the high-throughput sequencing (sequencing depth > 30×) data of the sample to be tested.

[0092] The identification component uses the sequencing data acquired by the data acquisition component as input data and executes the method shown in Example 1 to obtain the identification results of DNA viruses in the organism.

[0093] The results output component is used to output the DNA virus identification results obtained by the identification component.

[0094] 2. Instructions for use

[0095] Using human biological samples (digestive fluid, tissue fluid, or tissue) as test samples, high-throughput sequencing with a sequencing depth >30× is performed on the test samples. The sequencing data in FastQ file format from the sequencing data is obtained using the data acquisition component in the above-mentioned device and DNA virus identification is performed to obtain DNA virus identification results.

[0096] Example 4: Consistency Test of a Method for Identifying DNA Viruses in Organisms Based on High-Throughput Sequencing Data

[0097] 1. Experimental Methods

[0098] Validation set data 3: Download datasets PRJNA508862, PRJDB6776, PRJNA737147, and PRJEB31886 from the NCBI bioproject database as validation set data 3.

[0099] Each dataset in validation set 3 was subjected to DNA virus identification according to the method shown in Example 1, and the results were compared with the validation results of each dataset; among them, PRJNA508862 and PRJDB6776 datasets were used to identify HPV virus, and PRJNA737147 and PRJEB31886 datasets were used to identify HBV virus.

[0100] The PRJNA508862 dataset was verified using a bioinformatics verification method, as shown in PMC6365579.

[0101] The PRJDB6776 dataset was validated using the PCR probe method, as shown in PMC5974501.

[0102] The PRJNA737147 dataset was validated using antigen validation, and the antigen validation method is shown in the prior art (DOI:10.13140 / RG.2.2.33619.71204).

[0103] The PRJEB31886 dataset was validated using antibody verification, and the antibody verification method is shown in PMC6506499.

[0104] 2. Experimental Results

[0105] The DNA virus identification and validation results of validation set data 3 are shown in Table 1.

[0106] Table 1. DNA virus identification and validation results in validation set data 3.

[0107] Dataset Virus type Master clone consistency Master clone inconsistency accuracy PRJNA508862 HPV 18 3 85.7% PRJDB6776 HPV 205 0 100% PRJNA737147 HBV 15 0 100% PRJEB31886 HBV 12 0 100% total 250 3 99%

[0108] When DNA virus identification was performed on each dataset in validation set 3 using the method shown in Example 1 of this invention, the results showed that the accuracy of DNA virus identification using the method shown in Example 1 reached 98% compared with the validation results of each dataset, and the master clone consistency was high. This indicates that the method shown in Example 1 can not only accurately identify DNA viruses, but also accurately detect different subtypes of DNA viruses. Comparative Example 1: A method for identifying DNA viruses in organisms.

[0109] 1. Experimental Methods

[0110] Experimental group 1 was set up as shown in Example 1.

[0111] Control group 1 was set up as follows: clean fq data and DNA virus database were obtained according to steps S1 and S2 of Example 1. The clean fq data and DNA virus database were compared with the DNA virus database using bwa mem (version 0.7.17) based on the pre-screening parameters. The sequence with flag 4 in the comparison result was taken as the sequence that was aligned, and the sequence of the DNA virus database shown above was obtained.

[0112] 1000 sequences were randomly selected from the DNA virus database shown above to obtain virus_rc data. The reference genome containing the whole genome sequence information of each DNA virus was used as the alignment database. The virus_rc data and the alignment database were aligned using Blast (version 2.12.0+) based on preset parameters to obtain the number of high-quality aligned sequences of each DNA virus in the virus_rc data.

[0113] The result interpretation method for control group 1 is as follows: when the DNA virus is HBV, HPV or EBV, if the number of high-quality aligned sequences of the DNA virus in the virus_rc data is ≥1, then the sample to be tested contains the DNA virus.

[0114] If the DNA virus is a DNA virus other than HBV, HPV, and EBV, and the number of high-quality aligned sequences of that DNA virus in the virus_rc data is ≥10, then the sample to be tested contains that DNA virus.

[0115] The difference between control group 2 and control group 1 is that before randomly selecting 1000 sequences, sequences with a clip ratio greater than 10% in the DNA virus database shown above are removed; the other steps are the same.

[0116] The difference between control group 3 and experimental group 1 is that step S5 of Example 1 is not performed, and the judgment method of step S7 is: when the DNA virus is HBV virus, HPV virus or EBV virus, the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S6 is ≥1 or the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data is ≥1, then the sample to be tested contains the DNA virus.

[0117] If the DNA virus is a DNA virus other than HBV, HPV, and EBV, then the sample to be tested contains the DNA virus if the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S6 is ≥10 or the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data is ≥10.

[0118] The methods described in Experimental Group 1 and Control Groups 1 to 3 were used to test samples of known virus types (a total of 685 samples of known virus types, including 150 HBV positive samples, 58 HPV positive samples, 44 EBV positive samples and 433 negative samples), and the sensitivity, specificity and testing time of each group were calculated.

[0119] 2. Experimental Results

[0120] When samples of known virus types were tested using the methods described in experimental group 1 and control groups 1 to 3, the test results are shown in Table 2.

[0121] Table 2 Detection Results

[0122]

[0123] The results showed that only by using the method shown in Example 1 to detect DNA viruses in samples could the detection time be shortened while still achieving excellent detection results (high sensitivity and high specificity).

[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. For those skilled in the art, other variations or modifications can be made based on the above description and ideas, and it is neither necessary nor possible to exhaustively describe all implementation methods here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.

Claims

1. A method for identifying DNA viruses in organisms based on high-throughput sequencing data, characterized in that, Using biological samples derived from the organism as test samples, the method includes the following steps: S1. Perform high-throughput sequencing on the sample to be tested, collect sequencing data in FastQ format from the sequencing data, perform quality control on the sequencing data, and obtain clean fq data; S2. Obtain the sequence information of the reference genome of the organism from the NCBI database to form an organism database; The genomic sequence information of DNA viruses whose host organism is described is obtained from the NCBI database and used as a DNA virus database. S3. Align the clean fq data obtained in step S1 with the organism database shown in step S2 to obtain sequences that cannot be aligned with the organism database. Align the clean fq data obtained in step S1 with the DNA virus database shown in step S2 to obtain sequences aligned with the DNA virus database. Remove sequences that coexist with sequences that cannot be aligned with the organism database and sequences that coexist with sequences aligned with the DNA virus database. The sequences remaining after removing coexisting sequences from sequences that cannot be aligned with the organism database are recorded as unmap data. The sequences remaining after removing coexisting sequences from sequences aligned with the DNA virus database are recorded as virus_map data. S4. Remove sequences with a clip ratio greater than 10% from the virus_map data obtained in step S3 to obtain virus_map_rc data; S5. Select 9,000 to 11,000 sequences from the unmap data obtained in step S3 and the virus_map_rc data obtained in step S4, respectively, and denot them as unmap_some data and virus_map_rc_some data; If the number of sequences in the unmap data obtained in step S3 or the virus_map_rc data obtained in step S4 is less than 10,000, then all sequences are selected; S6. Summarize the whole genome sequence information of DNA viruses in the NCBI database as a comparison database, compare the unmap_some data obtained in step S5 with the comparison database, and obtain the number of high-quality DNA virus sequences in the unmap_some data. The virus_map_rc_some data obtained in step S5 is compared with the alignment database to obtain the number of high-quality DNA virus sequences in the virus_map_rc_some data; The high-quality alignment is defined as follows: the sequence in the unmap_some data and / or virus_map_rc_some data aligns to a sequence fragment in the alignment database that is at least 90% of the sequence length, and the sequence identity of the sequence fragment aligned to the alignment database is ≥90%. S7. If the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S6 and / or the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data meet the judgment criteria, the sample to be tested contains the DNA virus. If the number of high-quality aligned sequences of the DNA virus in the unmap_some data and / or the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data obtained in step S6 does not meet the judgment criteria, the sample to be tested does not contain the DNA virus. The judgment criteria are: the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S6 is ≥1 and / or the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data is ≥1.

2. The method according to claim 1, characterized in that, The proportion of DNA virus reads in the test sample to the total number of reads in the test sample shall not be less than 2%.

3. The method according to claim 1, characterized in that, The high-throughput sequencing mentioned in step S1 refers to high-throughput sequencing with a sequencing depth > 30×.

4. The method according to claim 1, characterized in that, The step S2, which involves obtaining the genome sequence information of a DNA virus whose host is the organism, specifically involves obtaining the genome sequence information of a DNA virus with complete nucleotide information, whose host is the organism, and whose genome molecular type is unknown DNA from the NCBI database.

5. The method according to claim 1, characterized in that, The comparison in step S3 is as follows: the comparison is performed using the bwa mem software based on the default parameters.

6. The method according to claim 1, characterized in that, The judgment condition in step S7 is: the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S6 is ≥100 and the proportion of the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S5 is ≥5% and / or the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data is ≥100 and the proportion of the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data obtained in step S5 is ≥5%; Alternatively, the number of high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S6 is ≥10 and the proportion of the high-quality aligned sequences of the DNA virus in the unmap_some data obtained in step S5 is ≥0.5%, and / or the number of high-quality aligned sequences of the DNA virus in the virus_map_rc_some data is ≥10 and the proportion of the high-quality aligned sequences of the DNA virus in the virus_map_rc_some data obtained in step S5 is ≥0.5%.

7. A device for identifying DNA viruses in organisms based on high-throughput sequencing data, characterized in that, The device includes a data acquisition component, an identification component, and a result output component; The data acquisition component is used to acquire sequencing data in FastQ format from the high-throughput sequencing data of the sample to be tested, wherein the sample to be tested is a biological sample derived from the organism. The identification component uses the sequencing data acquired by the data acquisition component as input data and executes the method as described in any one of claims 1 to 6 to obtain the DNA virus identification results in the organism. The result output component is used to output the DNA virus identification results in the organism obtained by the identification component.

8. The apparatus according to claim 7, characterized in that, The samples to be tested are biological samples derived from cancer patients.

9. The use of the method according to any one of claims 1 to 6 and / or the apparatus according to any one of claims 7 to 8 in the identification of DNA viruses in organisms.

10. The application according to claim 9, characterized in that, The DNA virus is HPV, HBV, EBV, Betapolyomavirus hominis BK, and / or HHV-8.