Iridescent virus database construction method and device, equipment and storage medium
By constructing the iris virus database, the problem of data fragmentation is solved, the unified integration and standardized analysis of iris viruses are realized, the virus classification and new species discovery is supported, and a standardized application knowledge base is provided, which promotes the in-depth research of iris viruses.
Patent Information
- Application Number
- CN202510474337.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-29
AI Technical Summary
In the prior art, the fragmentation of data related to iris viruses is lacking unified integration standards and interactive operation capabilities, which hinders the interaction between viral genomic data and spatiotemporal distribution data, and lacks unified analysis standards for iris viruses.
Build an iridescence virus database, including page integration standardized data sets, geographic distribution systems, non-redundant protein databases, core genes, evolutionary trees, genome collinear query system, host range and visual detection technology, providing standardized analysis processes and interactive operation capabilities.
It has realized the unified integration and standardized analysis of iris viral data, supports virus classification and new species discovery, provides a standardized application knowledge base for host range, vaccine reagents and visual detection technology, and promotes the in-depth research on iris virals.
Smart Images

Figure CN120388623A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a method, apparatus, device, and storage medium for constructing an iridovirus database. Background Art
[0002] Iridoviruses are a collective term for viruses in the family Iridoviridae, order Pimascovirales, class Megaviricetes. They are a type of large icosahedral virus containing a double-stranded DNA genome. Since they were first reported in the 1950s, they have been found to infect more than 150 species of poikilothermic vertebrates and invertebrates, including reptiles, amphibians, fish, crustaceans, and insects. The good environmental adaptability and unique host-switching mechanism of iridoviruses have led to their widespread prevalence worldwide. Their high pathogenicity has brought huge economic losses to the aquaculture industry and also led to a reduction in the population of wild amphibians, including many endangered species.
[0003] For virological research, a dedicated database comprehensively containing multi-dimensional relevant content and functions is very necessary. It can not only provide relatively comprehensive relevant research content and research tools for researchers in related fields, but also provide refined and integrated latest prevention, control, and detection technologies for practitioners in industries related to aquaculture and other industries. Currently, for some viruses with relatively in-depth research, such as influenza virus, coronavirus, rhinovirus, and herpes simplex virus, there are relevant databases. However, the relevant content of iridoviruses currently exists in a fragmented form in repositories such as RefSeq and GenBank and related research literature, lacking a unified integration standard. At the same time, there are still two key systematic deficiencies in the current research on iridoviruses: fragmented raw data hinders the interoperability between viral genome data and viral spatio-temporal distribution data; there is a lack of a unified analysis standard tailored for iridoviruses. Summary of the Invention
[0004] In view of this, embodiments of this application provide a method, apparatus, device, and storage medium for constructing an iridovirus database to establish a comprehensive iridovirus database including interoperable standardized data and standardized analysis processes.
[0005] One aspect of the embodiments of this application provides a method for constructing an iridovirus database, the method comprising the following steps:
[0006] Create each page of the iridovirus database; wherein, the pages include the home page, virus dataset, detailed description of virus isolates, virus distribution, non-redundant protein database, core genes, phylogenetic tree, genomic collinearity, host range, vaccine reagents, visualization detection technology, download, and pages corresponding to relevant information;
[0007] Integrate the standardized dataset of iridovirus isolates on the page of the virus dataset, and set the jump link to the page of the detailed description of each virus isolate;
[0008] Integrate the annotation results obtained through the standardized genome annotation process for iridoviruses on the page of the detailed description of each virus isolate;
[0009] Integrate the geographical distribution system of iridovirus isolates on the page of the virus distribution;
[0010] Integrate the non-redundant proteins of iridoviruses and the core genes of iridoviruses obtained through the core gene extraction process for iridoviruses on the page of the non-redundant protein database and the page of the core genes respectively;
[0011] Integrate the virus phylogenetic tree calculated through the core genes of iridoviruses on the page of the phylogenetic tree;
[0012] Integrate the nucleotide identity alignment results of the genomes of iridoviruses and the genomic collinearity query system for iridoviruses on the page of the genomic collinearity;
[0013] Integrate the content of the host range, vaccine reagents, and visualization detection technology in the standardized application knowledge base of iridoviruses on the pages of the host range, vaccine reagents, and visualization detection technology respectively;
[0014] Provide downloads of the genomes of iridovirus isolates, open reading frame nucleic acid sequences, and protein sequences on the download page;
[0015] Provide copyright and citation information of relevant information on the relevant information page.
[0016] In some embodiments, the steps of constructing the standardized dataset of iridovirus isolates include the following steps:
[0017] Retrieve the original information of iridovirus isolates from the RefSeq database and GenBank database and collect and organize it into a unified format. The original information includes the name, taxonomic status, host or sample collection species, collection time, collection location, and reference of each virus isolate;
[0018] Search for virus isolates lacking a direct link to the reference in a search engine and establish a connection between the virus isolates and the corresponding reference.
[0019] Correct inconsistent information or blanks in the original information by checking the content in the corresponding reference. When the corresponding reference is missing or the relevant information in the reference is lacking, the records in the RefSeq database and the GenBank database shall prevail.
[0020] Standardize all data, including splitting the name of the virus isolate into the virus name and the isolate name, splitting the taxonomic status of the virus isolate into the subfamily, genus, species, and genotype of the virus isolate, standardizing the collection location of the virus isolate into a unified address form, and obtaining the endangered level of the host or sample collection species from the IUCN Red List of Threatened Species and NatureServe.
[0021] Attach a representative picture with a caption and source to the virus isolate by retrieving in the original reference reporting, collecting, or isolating the virus isolate and the original reference for genome sequencing of the virus isolate.
[0022] Visualize the annotations obtained according to the standardized genome annotation process for iridoviruses and the original annotations from the GenBank database through the NCBI Sequence Viewer.
[0023] In some embodiments, the steps of constructing the geographical distribution system of the iridovirus isolates include the following steps:
[0024] Use AMap.CircleMarker to draw dot markers on the base map.
[0025] For virus isolates with specific collection location coordinates, the coordinates of the dot marker on the base map are the specific collection location coordinates; for virus isolates without specific collection location coordinates, the coordinates of the dot marker on the base map are random coordinates within the collection location range of the virus isolate.
[0026] Use createPopupContent to create an information display box.
[0027] Use select2 to create a dropdown filter box.
[0028] In some embodiments, the steps of extracting the core genes of the iridovirus include the following steps:
[0029] Translate the open reading frames obtained by classifying all iridovirus isolates into protein sequences, and use CD-HIT to integrate proteins with a similarity greater than or equal to 50% under the parameter "-c 1 -n 5" to construct the non-redundant protein database of iridoviruses;
[0030] Use the blastp command in BLAST+ to compare the proteins in the non-redundant protein database of iridoviruses with the proteins translated from the open reading frames in each iridovirus isolate genome under the parameters "-evalue 1e-05 -max_target_seqs 1";
[0031] For Iridoviridae and Alphairidovirinae, which include more than 200 virus isolates, the non-redundant proteins that can be aligned in more than 75% of the virus strains are determined as the core genes at the corresponding taxonomic positions;
[0032] For Betairidovirinae, Megalocytivirus, and Ranavirus, which have less than 200 virus isolates, the non-redundant proteins that can be aligned in 100% of the virus strains are determined as the core genes at the corresponding taxonomic positions.
[0033] In some embodiments, the steps of constructing the phylogenetic tree of the iridovirus core genes include the following steps:
[0034] Use the blastp command in BLAST+ to compare each of the core genes of iridoviruses with the protein sequences translated from the open reading frames obtained by classifying each iridovirus isolate under the parameters "-evalue 1e-05 -max_target_seqs 1";
[0035] Combine the comparison results of each of the core genes of iridoviruses into a fasta file, and use MAFFT to perform multiple alignment on the fasta file;
[0036] Input multiple multiple alignment results into IQ-TREE, and use the partition model, model selection, and ultrafast bootstrap approximation to construct the phylogenetic tree of the iridovirus core genes.
[0037] In some embodiments, the method for obtaining the genomic nucleotide identity alignment results of iridoviruses includes the following steps:
[0038] Retrieve the complete genome sequences of iridovirus isolates from the RefSeq database and the GenBank database;
[0039] By performing reverse complementation on the nucleic acid sequence of the viral genome and adjusting the position of the terminal redundant sequence, the hypothetical major capsid protein gene is made the first gene in the nucleic acid sequence of the viral genome;
[0040] Use the PairwiseAligner function of the Bio.Align package in Biopython to perform global alignment on any two of the adjusted viral genome nucleic acid sequences to obtain the corresponding genomic nucleotide identity alignment results.
[0041] In some embodiments, the steps of constructing the iridovirus standardized application knowledge base include the following steps:
[0042] Collect and organize relevant literature through a search engine, extract the information in the relevant literature and standardize it;
[0043] For the data on the host range, include the data with molecular biological evidence, include the species with clear scientific names and do not include hybrids, and obtain the endangered level of the host species from the Red List of Endangered Species of the International Union for Conservation of Nature and NatureServe;
[0044] For the data on the vaccine reagents, include the vaccine reagents that have been tested on live animals and proven effective;
[0045] For the visualization detection technology, include the detection technologies whose results can be observed with the naked eye without the aid of additional instruments.
[0046] Another aspect of the embodiments of the present application also provides an iridovirus database construction device, and the device includes:
[0047] A page creation unit for creating each page of the iridovirus database; wherein, the page includes the home page, virus data set, detailed description of virus isolates, virus distribution, non-redundant protein database, core genes, phylogenetic tree, genomic collinearity, host range, vaccine reagents, visualization detection technology, download, and pages corresponding to relevant information;
[0048] A page configuration unit for:
[0049] Integrate the standardized data set of iridovirus isolates on the page of the virus data set and set the jump link to the page of the detailed description of each virus isolate;
[0050] Integrate the annotation results obtained through the standardized genome annotation process for iridoviruses on the page of the detailed description of each virus isolate;
[0051] Integrate the geographical distribution system of iridovirus isolates on the page of the virus distribution;
[0052] Integrate the non-redundant proteins of iridovirus and the core genes of iridovirus obtained through the core gene extraction process for iridovirus on the page of the non-redundant protein database and the page of the core genes respectively;
[0053] Integrate the viral phylogenetic tree calculated through the core genes of iridovirus on the page of the phylogenetic tree;
[0054] Integrate the nucleotide identity alignment results of the iridovirus genome and the genome collinearity query system for iridovirus on the page of genome collinearity;
[0055] Integrate the host range content, vaccine reagent content, and visualization detection technology content in the standardized application knowledge base of iridovirus on the pages of the host range, the vaccine reagent, and the visualization detection technology respectively;
[0056] Provide downloads of the iridovirus isolate genome, open reading frame nucleic acid sequences, and protein sequences on the download page;
[0057] Provide copyright and citation information of relevant information on the relevant information page.
[0058] Another aspect of the embodiments of the present application also provides an electronic device, including a processor and a memory;
[0059] The memory is used to store programs;
[0060] The processor executes the program to implement the method described in any one of the above.
[0061] Another aspect of the embodiments of the present application also provides a computer-readable storage medium, where the storage medium stores a program, and the program is executed by a processor to implement the method described in any one of the above.
[0062] The present application has at least the following beneficial effects:
[0063] The present application collects the original data of known iridovirus isolates and standardizes it, establishes a geographical distribution system for iridovirus isolates, proposes a standardized genome annotation process for iridovirus, obtains evolutionarily conserved core genes that can be used for virus classification, and proposes a series of evolutionary tools for iridovirus (including phylogenetic trees, virus genome nucleotide identity comparison results, and genome collinearity query systems) for virus classification and discovery of new species. At the same time, it also provides a standardized application knowledge base including host range, vaccine reagents, and visualization on-site detection technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0065] Figure 1 It is a schematic flowchart of a method for constructing an iridovirus database provided by an embodiment of the present application;
[0066] Figure 2 It is a structural block diagram of a device for constructing an iridovirus database provided by an embodiment of the present application. Specific embodiments
[0067] In order to make the purpose, technical solutions and advantages of the present application more clear, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0068] Referring to Figure 1 , an embodiment of the present application provides a method for constructing an iridovirus database, specifically including the following steps S100 to S190:
[0069] S100: Create each page of the iridovirus database; wherein, the page includes pages corresponding to the home page, virus dataset, detailed description of virus isolates, virus distribution, non-redundant protein database, core genes, phylogenetic tree, genomic collinearity, host range, vaccine reagents, visualization detection technology, download, and relevant information;
[0070] S110: Integrate the standardized dataset of iridovirus isolates on the page of the virus dataset, and set the jump link for each page of the detailed description of the virus isolates;
[0071] S120: Integrate the annotation results obtained through the standardized genome annotation process for iridovirus on each page of the detailed description of the virus isolates;
[0072] S130: Integrate the geographical distribution system of iridovirus isolates on the page of the virus distribution;
[0073] S140: Integrate the iridovirus non-redundant protein and iridovirus core genes obtained through the core gene extraction process for iridovirus on the page of the non-redundant protein database and the page of the core genes respectively;
[0074] S150: Integrate the virus phylogenetic tree calculated through the iridovirus core genes on the page of the phylogenetic tree;
[0075] S160: Integrate the genomic nucleotide identity alignment results of iridoviruses and the genomic collinearity query system for iridoviruses on the page of the genomic collinearity;
[0076] S170: Integrate the host range content, vaccine reagent content, and visualization detection technology content in the standardized application knowledge base of iridoviruses on the pages of the host range, the vaccine reagent, and the visualization detection technology respectively;
[0077] S180: Provide downloads of the genomic sequences of iridovirus isolates, open reading frame nucleic acid sequences, and protein sequences on the download page;
[0078] S190: Provide copyright and citation information for the relevant information on the relevant information page.
[0079] Optionally, the steps of constructing the standardized data set of iridovirus isolates include the following steps:
[0080] Retrieve the original information of iridovirus isolates from the RefSeq database and the GenBank database and collect and organize it into a unified format. The original information includes the name, taxonomic status, host or sample collection species, collection time, collection location, and reference of each virus isolate;
[0081] Search for virus isolates lacking a direct link to the reference in a search engine and establish the connection between the virus isolate and the corresponding reference;
[0082] Correct inconsistent information or blanks in the original information by checking the content in the corresponding reference. When the corresponding reference is missing or the relevant information in the reference is missing, the records in the RefSeq database and the GenBank database shall prevail;
[0083] Standardize all data, including splitting the name of the virus isolate into the virus name and the isolate name, splitting the taxonomic status of the virus isolate into the subfamily, genus, species, and genotype of the virus isolate, standardizing the collection location of the virus isolate into a unified address form, and obtaining the endangered level of the host or sample collection species from the IUCN Red List of Threatened Species and NatureServe;
[0084] Attach representative pictures with captions and sources to the virus isolates by retrieving in the original references reporting, collecting, or isolating the virus isolates and the original references for genome sequencing of the virus isolates;
[0085] Visualize the annotations obtained according to the standardized genomic annotation process for iridoviruses and the original annotations from the GenBank database through the NCBI Sequence Viewer.
[0086] Optionally, the steps of constructing the geographical distribution system of the iridovirus isolates include the following steps:
[0087] Use AMap.CircleMarker to draw dot markers on the base map;
[0088] For virus isolates with specific collection location coordinates, the coordinates of the dot markers on the base map are the specific collection location coordinates; for virus isolates without specific collection location coordinates, the coordinates of the dot markers on the base map are random coordinates within the collection location range of the virus isolates;
[0089] Use createPopupContent to create an information display box;
[0090] Use select2 to create a dropdown filter box.
[0091] Optionally, the steps of extracting the core genes of iridoviruses include the following steps:
[0092] Translate the open reading frames obtained by dividing all iridovirus isolates into protein sequences, and use CD-HIT to integrate proteins with a similarity greater than or equal to 50% under the parameter "-c 1 -n 5" to construct the non-redundant protein database of iridoviruses;
[0093] Use the blastp command in BLAST+ to compare the proteins in the non-redundant protein database of iridoviruses with the proteins translated from the open reading frames in each iridovirus isolate genome under the parameters "-evalue 1e-05 -max_target_seqs 1";
[0094] For Iridoviridae and Alphairidovirinae including more than 200 virus isolates, the non-redundant proteins that can be aligned in more than 75% of the virus strains are determined as the core genes at the corresponding taxonomic status;
[0095] For Betairidovirinae, Megalocytivirus and Ranavirus with less than 200 virus isolates, the non-redundant proteins that can be aligned in 100% of the virus strains are determined as the core genes at the corresponding taxonomic status.
[0096] Optionally, the steps of constructing the phylogenetic tree of the iridovirus core genes include the following steps:
[0097] Use the blastp command in BLAST+ to compare each of the core genes of the iridovirus with the protein sequences translated from the open reading frames obtained by partitioning each iridovirus isolate under the parameters "-evalue 1e-05 -max_target_seqs 1";
[0098] Merge the comparison results of each of the core genes of the iridovirus into a fasta file, and use MAFFT to perform multiple alignments on the fasta file;
[0099] Input multiple multiple alignment results into IQ-TREE, and use the partition model, model selection, and ultrafast bootstrap approximation to construct the phylogenetic tree of the iridovirus core genes.
[0100] Optionally, the method for obtaining the genomic nucleotide identity alignment results of the iridovirus includes the following steps:
[0101] Retrieve the complete genomic sequences of the iridovirus isolates from the RefSeq database and the GenBank database;
[0102] By performing reverse complementation on the viral genomic nucleic acid sequence and adjusting the position of the terminal redundant sequences, make the putative major capsid protein gene the first gene in the viral genomic nucleic acid sequence;
[0103] Use the PairwiseAligner function of the Bio.Align package in Biopython to perform global alignment on any two of the adjusted viral genomic nucleic acid sequences to obtain the corresponding genomic nucleotide identity alignment results.
[0104] Optionally, the steps of constructing the standardized application knowledge base of the iridovirus include the following steps:
[0105] Collect relevant literature through a search engine and organize it, extract the information in the relevant literature and standardize it;
[0106] For the data on the host range, include the data with molecular biological evidence, include the species with clear scientific names and do not include hybrids, and at the same time obtain the endangered levels of the host species from the IUCN Red List of Threatened Species and NatureServe;
[0107] For the data on the vaccine reagents, include the vaccine reagents that have been tested in live animals and proven effective;
[0108] For the visualization detection technology, detection technologies that can be observed with the naked eye without the aid of additional instruments are included.
[0109] Next, the solutions of the embodiments of the present application will be introduced and described in detail in combination with specific application examples.
[0110] Specifically, this embodiment provides a method for constructing an iridovirus database, which may specifically include the following technical solutions:
[0111] 1) Create independent html pages for the home page, virus dataset, detailed description of virus isolates, virus distribution, non-redundant protein database, core genes, phylogenetic tree, genomic collinearity, host range, vaccine reagents, visualization detection technology, download, and related information respectively;
[0112] 2) Integrate the standardized dataset of iridovirus isolates on the virus dataset page and provide links that can jump to the detailed description pages of each virus isolate;
[0113] 3) Integrate the annotation results obtained through the standardized genomic annotation process for iridoviruses on the detailed description page of each virus isolate;
[0114] 4) Integrate the geographical distribution system of iridovirus isolates on the virus distribution page;
[0115] 5) Integrate the non-redundant proteins of iridoviruses and the core genes of iridoviruses obtained through the core gene extraction process for iridoviruses on the non-redundant protein database page and the core gene page;
[0116] 6) Integrate the phylogenetic tree calculated through the core genes of iridoviruses on the phylogenetic tree page;
[0117] 7) Integrate the genomic nucleotide identity alignment results of iridoviruses and the genomic collinearity query system for iridoviruses on the genomic collinearity page;
[0118] 8) Integrate part of the content of the host range in the standardized application knowledge base of iridoviruses on the host range page;
[0119] 9) Integrate part of the content of the vaccine reagents in the standardized application knowledge base of iridoviruses on the vaccine reagents page;
[0120] 10) Integrate part of the content of the visualization detection technology in the standardized application knowledge base of iridoviruses on the visualization detection technology page;
[0121] 11) Provide downloads of the genomes of iridovirus isolates, open reading frame nucleic acid sequences, and protein sequences on the download page;
[0122] 12) Provide relevant information such as copyright and citation of the relevant information on the relevant information page.
[0123] As an alternative implementation, this embodiment provides a standardized data set of iridovirus isolates, which includes a summary statistics table and a detailed description of each iridovirus isolate.
[0124] Among them, the summary statistics table includes the following: virus isolate name, the numbers of the virus isolate nucleic acid sequence in the GenBank and RefSeq databases, the subfamily of the virus isolate, the genus of the virus isolate, the species of the virus isolate, the genotype of the virus isolate, the host of the virus isolate or the species from which the sample was collected, the collection time of the virus isolate, the collection location of the virus isolate, remarks, and a link that can jump to the detailed description page of each virus isolate.
[0125] The detailed description of each iridovirus isolate is divided into the following three parts: The first part includes the following standardized content: virus name, isolate name, the numbers of the virus isolate nucleic acid sequence in the GenBank and RefSeq databases, the subfamily of the virus isolate, the genus of the virus isolate, the species of the virus isolate, the genotype of the virus isolate, the host of the virus isolate or the species from which the sample was collected, the collection time of the virus isolate, the collection location of the virus isolate, the detailed collection location coordinates of the virus isolate, the original references reporting or collecting or isolating the virus isolate, and the original references for genome sequencing of the virus isolate; The second part is pictures related to the virus isolate with captions and sources, such as electron microscope pictures of virus particles or tissue section pictures of infected animals, etc.; The third part is the annotations obtained according to the standardized genome annotation process for iridoviruses and the original annotations from the GenBank database visualized through NCBI Sequence Viewer.
[0126] As an alternative implementation, this embodiment provides a method for constructing the above-mentioned standardized data set of iridovirus isolates:
[0127] 1) Retrieve the original information of iridovirus isolates from the RefSeq database and the GenBank database and collect and organize it into a unified format. The original information includes the name, taxonomic status, host or sample collection species, collection time, collection location, and reference of each virus isolate.
[0128] 2) Search for virus isolates lacking a direct link to the reference in a search engine to establish the connection between the virus isolate and the corresponding reference.
[0129] 3) Correct inconsistent information or blanks in the original information by checking the content in the references. When the corresponding reference is missing or the relevant information in the reference is lacking, the records in the RefSeq database and GenBank database shall prevail;
[0130] 4) Standardize all data, including splitting the name of the virus isolate into the virus name and the isolate name, splitting the taxonomic status of the virus isolate into the subfamily, genus, species, and genotype of the virus isolate, standardizing the collection location of the virus isolate into a five-level address in the form of "specific location - city (county) - province (state) - country - continent", and obtaining the endangered level of the host or the species from which the sample was collected from the IUCN Red List of Threatened Species and NatureServe;
[0131] 5) Attach a representative picture with a caption and source for the virus isolate by retrieving in the original references reporting, collecting, or isolating the virus isolate and the original references for genome sequencing of the virus isolate;
[0132] 6) Visualize the annotations obtained according to the standardized genome annotation process for iridoviruses and the original annotations from the GenBank database through the NCBI Sequence Viewer.
[0133] As an alternative implementation manner, this embodiment provides a geographical distribution system for iridovirus isolates. The geographical distribution system shows the collection locations of virus isolates in the form of dot markers on the world map, and uses different colors to represent different genera. At the same time, the system provides a screening function for users to screen and display virus isolates belonging to different subfamilies, genera, and different continents, and also provides a simple information display box for users to view information about the virus isolate when clicking on the dot marker representing the virus isolate.
[0134] As an alternative implementation manner, this embodiment provides a construction method for the geographical distribution system of iridovirus isolates, including the following steps:
[0135] 1) Use the map in the English version of the map software as the base map, and draw dot markers on the base map using AMap.CircleMarker;
[0136] 2) For virus isolates with specific collection location coordinates, the coordinates of the dot marker on the base map are the accurate coordinates. For virus isolates without specific collection location coordinates, the coordinates of the dot marker on the base map are random coordinates within the collection location range of the virus isolate, while ensuring that the random coordinate markers do not overlap excessively;
[0137] 3) Use createPopupContent to create an information display box;
[0138] 4) Use select2 to create a drop-down filtering box.
[0139] As an alternative implementation, this embodiment provides a standardized genomic annotation method for iridovirus, including the following steps:
[0140] 1) Retrieve the complete genomic sequences of iridovirus isolates from the RefSeq database and the GenBank database;
[0141] 2) Divide the complete iridovirus genomic sequences into open reading frames. First, an open reading frame should start with the nucleotide ATG and end with the nucleotides TAA, TAG, TGA, TRA, or TAR. Second, ignore nested open reading frames, including forward and reverse nesting. Finally, when forward open reading frames overlap, preferentially retain the longer open reading frame. When the lengths of the overlapping open reading frames are the same, retain both open reading frames.
[0142] As an alternative implementation, this embodiment provides a method for extracting core genes of iridovirus, including the following steps:
[0143] 1) Translate the open reading frames obtained by dividing all iridovirus isolates into protein sequences, and use CD-HIT to integrate proteins with a similarity of 50% or more under the parameter "-c 1 -n 5" to construct a non-redundant protein database of iridovirus;
[0144] 2) Use the blastp command in BLAST+ to compare the proteins in the non-redundant protein database of iridovirus with the proteins translated from the open reading frames in each iridovirus isolate genome under the parameters "-evalue 1e-05 -max_target_seqs 1";
[0145] 3) For Iridoviridae and Alphairidovirinae including more than 200 virus isolates, the non-redundant proteins that can be aligned in more than 75% of the virus isolates are determined as the core genes of the corresponding taxonomic status. For Betairidovirinae, Megalocytivirus, and Ranavirus with less than 200 virus isolates, the non-redundant proteins that can be aligned in 100% of the virus isolates are determined as the core genes of the corresponding taxonomic status.
[0146] As an alternative implementation, this embodiment provides a method for constructing an evolutionary tree using the core genes of iridovirus, including the following steps:
[0147] 1) Use the blastp command in BLAST+ to compare each iridovirus core gene with the protein sequences translated from the open reading frames obtained by partitioning each iridovirus isolate under the parameters "-evalue 1e-05 -max_target_seqs 1".
[0148] 2) Combine the comparison results of each iridovirus core gene into a fasta file and perform multiple alignment using MAFFT.
[0149] 3) Input the multiple alignment results into IQ-TREE and construct a phylogenetic tree using the partition model, model selection, and ultrafast bootstrap approximation.
[0150] As an alternative implementation, this example provides a method for aligning the genomic nucleotide identities of iridoviruses, including the following steps:
[0151] 1) Retrieve the complete genomic sequences of iridovirus isolates from the RefSeq database and the GenBank database.
[0152] 2) By performing reverse complementation on the viral genomic nucleic acid sequences and adjusting the positions of the terminal redundant sequences, make the putative major capsid protein gene the first gene in the viral genomic nucleic acid sequence.
[0153] 3) Use the PairwiseAligner function of the Bio.Align package in Biopython to perform global alignment on any two adjusted viral genomic nucleic acid sequences to obtain the alignment results of the corresponding genomic nucleotide sequence identities.
[0154] As an alternative implementation, this example provides a genomic collinearity query system for iridoviruses, which can query the genomic collinearity between two genomes by inputting or selecting two complete iridovirus genomes.
[0155] As an alternative implementation, this example provides a construction method for a genomic collinearity query system for iridoviruses, including the following steps:
[0156] 1) Retrieve the complete genomic sequences of iridovirus isolates from the RefSeq database and the GenBank database.
[0157] 2) Use nucmer in MUMmer to perform collinearity analysis on the selected genomic sequences of two viral isolates under the parameter "-l 10".
[0158] 3) Visualize the collinearity analysis results using mummerplot in MUMmer.
[0159] 4) Create a drop-down filtering box using select2 and provide an input filtering function.
[0160] As an alternative implementation, this embodiment provides an iridovirus standardization application knowledge base, which consists of three parts, including host range data, vaccine reagents, and visualization detection techniques. Each part includes a card-style user-friendly interface and a standardized data table.
[0161] As an alternative implementation, this embodiment provides a method for constructing an iridovirus standardization application knowledge base, including the following steps:
[0162] 1) Collect relevant literature through a search engine and organize it, refine and standardize the information therein;
[0163] 2) For host range data, only include data with molecular biological evidence, only include species with clear scientific names and do not include hybrids. At the same time, obtain the endangered level of the host species from the IUCN Red List of Threatened Species and NatureServe. The standardized data includes the following: host scientific name, host species category, host living environment, class to which the host species belongs, order to which the host species belongs, family to which the host species belongs, endangered level of the host in the IUCN Red List of Threatened Species, endangered level of the host in NatureServe, virus subfamily, virus genus, virus species, virus genotype, reference;
[0164] 3) For vaccine reagent data, only include vaccine reagents that have been tested in live animals and proven effective. The standardized data includes the following: virus name, virus subfamily, virus genus, virus species, virus genotype, vaccine name, vaccine type, virus gene concerned by the vaccine, host vaccinated, vaccination method, survival rate after vaccination, protection rate after vaccination, remarks, reference;
[0165] 4) For visualization detection techniques, only include detection techniques whose results can be observed with the naked eye without the need for additional instruments. The standardized data includes the following: virus name, virus subfamily, virus genus, virus species, virus genotype, detection method, virus gene concerned by the detection technique, experimentally verified detection species, type of sample detected, detection limit, reference.
[0166] The beneficial effects of this embodiment include:
[0167] In this embodiment, by collecting the original data of known iridovirus isolates and standardizing it, a geographical distribution system of iridovirus isolates is established, and a standardized genome annotation process for iridoviruses is proposed to obtain evolutionarily conserved core genes that can be used for virus classification. A series of evolutionary tools for iridoviruses (including phylogenetic trees, nucleotide identity comparison results of virus genomes, and genome collinearity query systems) are proposed for virus classification and discovery of new species. At the same time, a standardized application knowledge base including host range, vaccine reagents, and visualization on-site detection techniques is also provided.
[0168] In addition, through a large amount of creative work, the inventors of this application have constructed the iridovirus database of this embodiment, which provides an important foundation for iridovirus genome-related work, can promote further progress in different biological aspects of the iridovirus research field, provide a basis for better preventing and controlling diseases and protecting animals, and at the same time can provide valuable information for experimental teams researching related content to help narrow the experimental scope.
[0169] Refer to Figure 2 , this embodiment of the application provides an iridovirus database construction device, including:
[0170] A page creation unit for creating each page of the iridovirus database; wherein, the pages include pages corresponding to the home page, virus dataset, detailed description of virus isolates, virus distribution, non-redundant protein database, core genes, phylogenetic tree, genome collinearity, host range, vaccine reagents, visualization detection technology, download, and related information;
[0171] A page configuration unit for:
[0172] Integrating the standardized dataset of iridovirus isolates on the page of the virus dataset and setting the jump link for the page of each detailed description of virus isolates;
[0173] Integrating the annotation results obtained through the standardized genome annotation process for iridoviruses on the page of each detailed description of virus isolates;
[0174] Integrating the geographical distribution system of iridovirus isolates on the page of virus distribution;
[0175] Integrating the iridovirus non-redundant proteins and iridovirus core genes obtained through the core gene extraction process for iridoviruses on the page of the non-redundant protein database and the page of core genes respectively;
[0176] Integrating the virus phylogenetic tree calculated through iridovirus core genes on the page of the phylogenetic tree;
[0177] Integrate the genomic nucleotide identity alignment results of iridovirus and the genomic collinearity query system for iridovirus on the page of the genomic collinearity;
[0178] Integrate the host range content, vaccine reagent content, and visualization detection technology content in the iridovirus standardized application knowledge base on the pages of the host range, the vaccine reagent, and the visualization detection technology respectively;
[0179] Provide downloads of the genomic sequences of iridovirus isolates, open reading frame nucleic acid sequences, and protein sequences on the download page;
[0180] Provide copyright and citation information of the relevant information on the relevant information page.
[0181] It can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented in the device embodiments of the present application are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0182] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present application are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated where the order of various operations is changed and where sub-operations described as part of a larger operation are executed independently.
[0183] In addition, although the present application is described in the context of functional modules, it should be understood that unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It can also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present application. Rather, considering the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Therefore, those skilled in the art can implement the present application as set forth in the claims without undue experimentation using ordinary skills. It can also be understood that the specific concepts disclosed are illustrative only and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.
[0184] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0185] The logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in combination with an instruction execution system, apparatus, or device.
[0186] More specific examples (non-exhaustive list) of computer-readable media include the following: electrical connection parts with one or more wirings (electronic devices), portable computer disk cartridges (magnetic devices), random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memories), optical fiber devices, and portable compact disc read-only memories (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or otherwise processing it as necessary, and then storing it in a computer memory.
[0187] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0188] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0189] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the claims and their equivalents.
[0190] The above is a specific description of the preferred embodiments of the present application, but the present application is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.
Claims
1. A method for constructing an iridovirus database, characterized in that, The method includes the following steps: Create each page of the iridovirus database; wherein, the pages include the home page, virus dataset, detailed description of virus isolates, virus distribution, non-redundant protein database, core genes, phylogenetic tree, genomic collinearity, host range, vaccine reagents, visualization detection technology, download, and pages corresponding to relevant information; Integrate the standardized dataset of iridovirus isolates on the page of the virus dataset, and set the jump link for the page of the detailed description of each virus isolate; Integrate the annotation results obtained through the standardized genomic annotation process for iridoviruses on the page of the detailed description of each virus isolate; Integrate the geographical distribution system of iridovirus isolates on the page of the virus distribution; Integrate the non-redundant proteins of iridoviruses and the core genes of iridoviruses obtained through the core gene extraction process for iridoviruses on the page of the non-redundant protein database and the page of the core genes respectively; Integrate the virus phylogenetic tree calculated through the core genes of iridoviruses on the page of the phylogenetic tree; Integrate the genomic nucleotide identity alignment results of iridoviruses and the genomic collinearity query system for iridoviruses on the page of the genomic collinearity; Integrate the host range content, vaccine reagent content, and visualization detection technology content in the standardized application knowledge base of iridoviruses on the pages of the host range, the vaccine reagents, and the visualization detection technology respectively; Provide downloads of the genomes, open reading frame nucleic acid sequences, and protein sequences of iridovirus isolates on the page of the download; Provide copyright and citation information of relevant information on the page of the relevant information.
2. The method for constructing an iridovirus database according to claim 1, wherein The steps for constructing the standardized dataset of iridovirus isolates include the following steps: Retrieve the original information of iridovirus isolates from the RefSeq database and the GenBank database and collect and organize it into a unified format. The original information includes the name, taxonomic status, host or sample collection species, collection time, collection location, and reference of each virus isolate; Search for virus isolates lacking a direct link to the reference in the search engine to establish the connection between the virus isolate and the corresponding reference; Correct inconsistent information or blanks in the original information by checking the content in the corresponding reference. When the corresponding reference is missing or the relevant information in the reference is missing, the records in the RefSeq database and the GenBank database shall prevail; Standardize all data, including splitting the name of the virus isolate into the virus name and the isolate name, splitting the taxonomic status of the virus isolate into the subfamily, genus, species, and genotype of the virus isolate, standardizing the collection location of the virus isolate into a unified address form, and obtaining the endangered level of the host or sample collection species from the IUCN Red List of Threatened Species and NatureServe; Retrieve from the original references reporting, collecting, or isolating the virus isolates and the original references for genome sequencing of the virus isolates, and attach representative pictures with captions and sources to the virus isolates; Visualize the annotations obtained according to the standardized genome annotation process for iridoviruses and the original annotations from the GenBank database using the NCBI Sequence Viewer.
3. The method for constructing an iridovirus database according to claim 1, wherein Steps for constructing the geographical distribution system of the iridovirus isolates, including the following steps: Use AMap.CircleMarker to draw dot markers on the base map; For virus isolates with coordinates of specific collection locations, the coordinates of the dot markers on the base map are the coordinates of the specific collection locations; for virus isolates without coordinates of specific collection locations, the coordinates of the dot markers on the base map are random coordinates within the range of the virus isolate collection locations; Use createPopupContent to create an information display box; Use select2 to create a dropdown filtering box.
4. A method for constructing an iridovirus database according to claim 1, characterized in that, Steps for extracting the core genes of iridoviruses, including the following steps: Translate the open reading frames obtained by dividing all iridovirus isolates into protein sequences, and use CD-HIT to integrate proteins with a similarity greater than or equal to 50% under the parameter "-c 1 -n 5" to construct the non-redundant protein database of iridoviruses; Use the blastp command in BLAST+ under the parameters "-evalue 1e-05 -max_target_seqs 1" to compare the proteins in the non-redundant protein database of iridoviruses with the proteins translated from the open reading frames in each iridovirus isolate genome; For Iridoviridae and Alphairidovirinae including more than 200 virus isolates, the non-redundant proteins that can be aligned in more than 75% of the virus strains are determined as the core genes at the corresponding taxonomic positions; For Betairidovirinae, Megalocytivirus, and Ranavirus with less than 200 virus isolates, the non-redundant proteins that can be aligned in 100% of the virus strains are determined as the core genes at the corresponding taxonomic positions.
5. A method for constructing an iridovirus database according to claim 1, characterized in that Steps for constructing the phylogenetic tree of the core genes of iridoviruses, including the following steps: Use the blastp command in BLAST+ under the parameters "-evalue 1e-05 -max_target_seqs 1" to compare each core gene of iridoviruses with the protein sequences translated from the open reading frames obtained by dividing each iridovirus isolate; Merge the comparison results of each core gene of iridoviruses into a fasta file, and use MAFFT to perform multiple alignments on the fasta file; Input multiple multiple alignment results into IQ-TREE, and use the partition model, model selection, and ultrafast bootstrap approximation to construct the phylogenetic tree of the core genes of iridoviruses.
6. The method for constructing an iridovirus database according to claim 1, wherein A method for obtaining the genomic nucleotide identity alignment results of iridovirus, comprising the following steps: Retrieve the complete genomic sequences of iridovirus isolates from the RefSeq database and GenBank database; By performing reverse complementation on the viral genomic nucleic acid sequence and adjusting the position of the terminal redundant sequence, make the putative major capsid protein gene the first gene in the viral genomic nucleic acid sequence; Use the PairwiseAligner function of the Bio.Align package in Biopython to perform global alignment on any two of the adjusted viral genomic nucleic acid sequences, and obtain the corresponding genomic nucleotide identity alignment results.
7. A method for constructing an iridovirus database according to any one of claims 1 to 6, characterized in that, The steps for constructing the standardized application knowledge base of iridovirus include the following steps: Collect relevant literature through a search engine and organize it, extract the information in the relevant literature and standardize it; For the data on the host range, include the data with molecular biological evidence, include the species with clear scientific names and do not include hybrids, and obtain the endangered levels of the host species from the IUCN Red List of Threatened Species and NatureServe; For the data on the vaccine reagents, include the vaccine reagents that have been tested in live animals and proven effective; For the visualization detection technology, include the detection technology whose results can be observed with the naked eye without the aid of additional instruments.
8. An iridovirus database construction device, characterized in that, The device includes: A page creation unit for creating each page of the iridovirus database; wherein, the page includes the home page, virus dataset, detailed description of virus isolates, virus distribution, non-redundant protein database, core genes, phylogenetic tree, genomic collinearity, host range, vaccine reagents, visualization detection technology, download, and the page corresponding to relevant information; A page configuration unit for: Integrate the standardized dataset of iridovirus isolates on the page of the virus dataset, and set the jump link to the page of the detailed description of each virus isolate; Integrate the annotation results obtained through the standardized genomic annotation process for iridovirus on the page of the detailed description of each virus isolate; Integrate the geographical distribution system of iridovirus isolates on the page of the virus distribution; Integrate the iridovirus non-redundant protein and iridovirus core genes obtained through the core gene extraction process for iridovirus on the page of the non-redundant protein database and the page of the core genes respectively; Integrate the viral phylogenetic tree calculated through the core genes of iridovirus on the page of the phylogenetic tree; Integrate the genomic nucleotide identity alignment results of iridovirus and the genomic collinearity query system for iridovirus on the page of the genomic collinearity; Integrate the host range content, vaccine reagent content, and visualization detection technology content in the standardized application knowledge base of iridovirus on the pages of the host range, the vaccine reagents, and the visualization detection technology respectively; Provide downloads of the genomic sequences, open reading frame nucleic acid sequences, and protein sequences of iridovirus isolates on the page of the download; Provide copyright and citation information of relevant information on the page of the relevant information.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory; The memory is used for storing programs; The processor executes the program to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a program, and the program is executed by the processor to implement the method according to any one of claims 1 to 7.