Method and system for viral sequence genome structure and functional annotation
By aligning and functionally annotating viral genome sequences, identifying and processing the special structure of the viral genome, the problems of low efficiency and lack of flexibility of existing tools in viral genome annotation are solved, and efficient and flexible viral genome annotation and accurate virus classification are achieved.
Patent Information
- Application Number
- CN202411821975.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-12-11
AI Technical Summary
Existing viral genome annotation tools have problems such as low computational efficiency, insufficient flexibility and limited adaptability when dealing with terminal repeat sequences, overlapping open reading frames and bidirectional translation of ambiguous viruses, and are unable to accurately identify and process the special structure of viral genomes.
By aligning the viral genome sequence to identify terminal repeat sequences and open reading frames, the getORF software is used to predict the open reading frames, and the blastp tool is combined for functional annotation. Specific sequences and frames are deleted or retained according to user-set parameters. Open reading frames with viral protein functions are preferentially retained, the direction of the viral genome sequence is adjusted, ambiguous viruses are identified, and overlapping open reading frames are processed.
It improves the accuracy and efficiency of viral genome annotation, realizes efficient and flexible annotation of viral genomes, generates file formats that meet international standards, and facilitates data sharing and virus classification.
Smart Images

Figure CN119920306B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gene annotation, and in particular to a method and system for annotating the structure and function of viral sequence genomes. Background Art
[0002] With the rapid development of high-throughput sequencing technology, massive amounts of viral genomic data are continuously being generated. The annotation and functional prediction of viral genomes have become core tasks in fields such as virology research, viral vaccine development, and drug target identification. Sequence structure and functional annotation of viral genomes not only help reveal the genetic characteristics of viruses but also assist researchers in understanding the mechanisms of virus-host interactions, clarifying virus classification, and predicting their potential pathogenicity.
[0003] Traditional gene annotation tools such as GeneMarkS, PROKKA, and RAST, while performing well in annotating genes in bacteria and eukaryotes, have certain limitations when it comes to annotating viral genomes. Existing annotation tools are particularly unable to accurately identify and process the unique structures of viral genomes, such as terminal repeats, overlapping open reading frames (ORFs), and bidirectional translation in ambiguous viruses. Furthermore, the accuracy and efficiency of viral genome annotation are crucial for downstream functional research. However, existing tools often exhibit low computational efficiency, insufficient flexibility, and an inability to adapt to the classification of all viruses when processing these unique structures.
[0004] Therefore, the current viral genome annotation technology has many limitations and cannot meet the growing needs of the field of viral research. Especially when it comes to the special structure and function annotation of viral sequences, more efficient, flexible and adaptable tools are urgently needed. Summary of the Invention
[0005] The present invention provides a method and system for annotating the structure and function of viral sequence genomes, which are used to overcome the limitations of the existing viral genome annotation technology and achieve more efficient, flexible and adaptable viral genome annotation.
[0006] The present invention provides a method for annotating the structure and function of viral sequence genomes, comprising:
[0007] Identifying terminal repeat sequences in the viral genome sequence by aligning the viral genome sequence itself, and deleting or retaining the terminal repeat sequences according to a first parameter set by the user;
[0008] Predicting the open reading frame in the viral genome sequence using getORF software, and deleting or retaining the open reading frame according to a second parameter set by the user;
[0009] Annotating the viral genome sequence using a blastp tool, identifying open reading frames with viral protein functions based on functional annotations of the open reading frames in the annotations, determining virus classification information based on the annotations of the open reading frames with viral protein functions, and adjusting the orientation of the viral genome sequence based on the annotations;
[0010] Ambiguous viruses are identified according to the virus classification information, and overlapping open reading frames in different translation frames in the viral genome sequence of the ambiguous virus are retained according to the tolerance of open reading frame overlap set by the user, and open reading frames with viral protein functions in the same direction are preferentially retained.
[0011] According to a method for annotating the structure and function of a viral sequence genome provided by the present invention, the terminal repeat sequence includes a tandem repeat sequence and an inverted complementary repeat sequence.
[0012] According to a method for annotating viral sequence genome structure and function provided by the present invention, the first parameter includes whether to remove terminal repeats, the shortest terminal tandem repeat sequence retention length, and the shortest terminal reverse complementary repeat sequence retention length. Deleting or retaining the terminal repeat sequence according to the first parameter set by the user includes:
[0013] When the value of whether to remove terminal repeat sequences is TRUE, the tandem repeat sequences and reverse complementary repeat sequences are deleted;
[0014] If the value of whether to remove terminal repeats is FALSE, determining whether the length of the tandem repeats is less than the shortest retained length of the terminal tandem repeats, and whether the length of the reverse complementary repeats is less than the shortest retained length of the terminal reverse complementary repeats;
[0015] If the length of the tandem repeat sequence is less than the retained length of the shortest terminal tandem repeat sequence, the tandem repeat sequence is deleted; otherwise, the tandem repeat sequence is retained;
[0016] If the length of the reverse complementary repeat sequence is less than the retained length of the shortest terminal reverse complementary repeat sequence, the reverse complementary repeat sequence is deleted; otherwise, the reverse complementary repeat sequence is retained.
[0017] According to a method for annotating the genome structure and function of a viral sequence provided by the present invention, after aligning the viral genome sequence itself and identifying the terminal repeat sequence in the viral genome sequence, the method further includes:
[0018] Determining the position and content of the terminal repeat sequence in the viral genome sequence;
[0019] The position and content of the terminal repeat sequence are visualized.
[0020] The viral sequence genome structure and function annotation method provided by the application includes an open reading frame without a start codon and / or a stop codon.
[0021] The viral sequence genome structure and function annotation method provided by the application, the second parameter includes whether to retain all discovered open reading frames and the minimum open reading frame length, and the open reading frame is deleted or retained according to the second parameter set by the user, including:
[0022] In the case that the all discovered open reading frames are retained, all discovered open reading frames are retained.
[0023] In the case that the all discovered open reading frames are not retained, the longest open reading frame is retained in the same translation frame, and it is determined whether the length of the longest open reading frame is less than the minimum open reading frame length.
[0024] In the case that the length of the longest open reading frame is less than the minimum open reading frame length, the longest open reading frame is deleted, otherwise the longest open reading frame is retained.
[0025] The viral sequence genome structure and function annotation method provided by the application further includes:
[0026] According to the annotation of the viral genome sequence, an annotation file in GFF format, a tbl file for submission to NCBI and NGDC, and a GenBank file for local viewing and editing are generated.
[0027] The application further provides a viral sequence genome structure and function annotation system, including:
[0028] A terminal repeat sequence detection module is configured to identify terminal repeat sequences in the viral genome sequence by aligning the viral genome sequence itself, and delete or retain the terminal repeat sequences according to a first parameter set by a user.
[0029] An open reading frame prediction module is configured to predict open reading frames in the viral genome sequence by using getORF software, and delete or retain the open reading frames according to a second parameter set by a user.
[0030] a functional annotation module, configured to annotate the viral genome sequence using a blastp tool, identify open reading frames having viral protein functions based on the functional annotations of the open reading frames in the annotations, determine virus classification information based on the annotations of the open reading frames having viral protein functions, and adjust the orientation of the viral genome sequence based on the annotations;
[0031] The ambisense virus identification module is used to identify ambisense viruses according to the virus classification information, retain overlapping open reading frames in different translation frames in the viral genome sequence of the ambisense virus according to the tolerance of open reading frame overlap set by the user, and give priority to retaining open reading frames with viral protein functions in the same direction.
[0032] The present invention also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for annotating the viral sequence genome structure and function as described above is implemented.
[0033] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for annotating the viral sequence genome structure and function.
[0034] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-mentioned methods for annotating the viral sequence genome structure and function.
[0035] The method and system for annotating the genome structure and function of viral sequences provided by the present invention take into account the special structure of the viral genome, including terminal repeat sequences, overlapping characteristics of open reading frames, and bidirectional translation of ambiguous viruses, when annotating the viral genome sequence. This allows for automatic identification and processing, thereby improving the accuracy of viral genome annotation while demonstrating greater efficiency and flexibility. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0037] Figure 1 This is one of the flow charts of the method for structural and functional annotation based on viral genome sequences provided by the present invention;
[0038] Figure 2Schematic diagram of the open reading frame prediction results in the structure and function annotation method based on the viral genome sequence provided by the present invention;
[0039] Figure 3 It is a visualization diagram of the terminal repeat sequence of the viral genome in the structure and function annotation method based on the viral genome sequence provided by the present invention;
[0040] Figure 4 This is the second flow chart of the method for structural and functional annotation based on viral genome sequences provided by the present invention;
[0041] Figure 5 Schematic diagram of the structure and function annotation system based on viral genome sequences provided by the present invention;
[0042] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0043] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0044] In the field of viral gene annotation, existing technologies such as GLIMMER and VAPiD have certain uses in annotating specific viral genomes. However, these tools have shortcomings in the following aspects:
[0045] Identification and processing of terminal repeats: Viral genomes often contain terminal repeats (such as tandem repeats and reverse complementary repeats), which play an important role in the evolution and function of viral genomes. Existing tools are often unable to automatically identify and process these repeats, resulting in inaccurate annotation results.
[0046] Open reading frame prediction and annotation: The common phenomenon of open reading frame overlap in viral genomes is not well handled by existing tools. In particular, for ambiguous viruses, existing tools are unable to accurately distinguish and retain functional open reading frames in the same and different orientations.
[0047] Classification and direction adjustment: Sequence direction adjustment and classification determination after virus annotation depend on the annotation results, but existing tools have limited support for virus classification, cannot automatically adjust sequence direction according to annotation results, and have limited applicability to different viruses.
[0048] Efficiency and flexibility: Existing technologies often take a long time to process viral genomes, especially when dealing with large amounts of data or complex genome structures, and lack flexibility in customizing the retention and discarding of reading frames.
[0049] The present invention provides a method and system for structural and functional annotation based on viral genome sequences, which effectively improves the deficiencies in the prior art.
[0050] The following combination Figure 1 A method for annotating the structure and function of viral sequence genomes of the present invention is described, comprising:
[0051] Step 101, identifying terminal repeat sequences in the viral genome sequence by aligning the viral genome sequence itself, and deleting or retaining the terminal repeat sequences according to a first parameter set by the user;
[0052] This embodiment realizes the identification of terminal repeat sequences, automatically detects whether tandem repeat sequences and reverse complementary repeat sequences exist at the end of the viral genome sequence through sequence alignment, and flexibly selects to delete or retain these repeat sequences according to the first parameter set by the user.
[0053] Step 102, predicting the open reading frame in the viral genome sequence using getORF software, and deleting or retaining the open reading frame according to a second parameter set by the user;
[0054] This embodiment achieves flexible retention of open reading frames by predicting all open reading frames in the viral genome sequence through getORF, especially being able to identify open reading frames without start codons or stop codons and process these open reading frames. Figure 2 shown.
[0055] The open reading frame can be flexibly selected to be deleted or retained according to the second parameter set by the user, thereby improving the flexibility of reading frame selection.
[0056] Step 103, annotating the viral genome sequence using a blastp tool, identifying open reading frames with viral protein functions based on functional annotations of the open reading frames in the annotations, determining virus classification information based on the annotations of the open reading frames with viral protein functions, and adjusting the orientation of the viral genome sequence based on the annotations;
[0057] The embodiment realizes virus function annotation and classification. The blastp tool is used to perform function annotation on the reserved open reading frame, and identify the open reading frame with virus protein function. The value Evalue of the virus protein function annotation of the blastp tool is identified according to the user setting. For example, the value Evalu of the virus protein function annotation is set to 1e-5, and the default is 1e-5.
[0058] According to the annotation result of the open reading frame with virus protein function, the minimum classification unit of the virus is determined, and the virus classification information is output, so as to improve the accuracy and efficiency of virus classification.
[0059] Uniport protein database ProtDB can be established by using Diamond, and a correspondence table of protein ID and virus classification ID in the protein database is established. The protein ID is obtained from the annotation of the open reading frame with virus protein function, the virus classification ID is obtained by searching the correspondence table, and thus the virus classification information is obtained.
[0060] In addition, the sequence direction can be automatically adjusted according to whether the direction of the virus genome sequence determined according to the function annotation result is consistent with the direction of the target sequence, so as to ensure the correctness of the virus genome sequence.
[0061] In step 104, according to the virus classification information, the ambisense virus is identified, according to the user-set tolerance degree of the overlap of the open reading frame, the overlapping open reading frame in different translation frames in the virus genome sequence of the ambisense virus is reserved, and the open reading frame with virus protein function in the same direction is preferentially reserved.
[0062] The embodiment realizes the identification and direction optimization of the ambisense virus. The ambisense virus is automatically distinguished according to the virus classification information obtained from the annotation result, and the virus protein open reading frame along the 5'->3' or 3'->5' direction is preferentially reserved according to the annotation result, so as to ensure the accuracy of the annotation.
[0063] According to the total length or total number (the total length is used by default) of the open reading frame (ORF) of the virus protein annotated in the same direction, the sequence is sorted. If the total length of the 5'->3' direction is greater than the total length of the 3'->5' direction or the total number of the 5'->3' direction is greater than the total number of the 3'->5' direction, the sequence direction does not need to be adjusted; otherwise, the sequence is subjected to reverse complement processing.
[0064] According to the user-selected tolerance degree of the overlap of the open reading frame (default 70%), the overlapping open reading frame in different translation frames is reserved. The open reading frame in different translation frames is allowed to exist a specified degree of overlap, which solves the defect that the existing annotation tool cannot handle the overlap of the open reading frame.
[0065] The intersection of the newly discovered open reading frame and the gap between the known functional protein open reading frame can also be set, such as ORFoverlap=0.1, which means that the default gap must overlap the new open reading frame by at least 10%.
[0066] When annotating viral genome sequences, this embodiment takes into account the special structure of the viral genome, including terminal repeats, overlapping characteristics of open reading frames, and bidirectional translation of ambiguous viruses, and can automatically identify and process them, thereby improving the accuracy of viral genome annotation while demonstrating higher efficiency and flexibility.
[0067] Based on the above embodiment, the terminal repeat sequence in this embodiment includes a tandem repeat sequence and an inverted complementary repeat sequence.
[0068] Based on the above embodiment, in this embodiment, the first parameter includes whether to remove terminal repeats, the shortest terminal tandem repeat length to be retained, and the shortest terminal reverse complementary repeat length to be retained. Deleting or retaining the terminal repeat according to the first parameter set by the user includes:
[0069] When the value of whether to remove terminal repeat sequences is TRUE, the tandem repeat sequences and reverse complementary repeat sequences are deleted;
[0070] If the value of whether to remove terminal repeats is FALSE, determining whether the length of the tandem repeats is less than the shortest retained length of the terminal tandem repeats, and whether the length of the reverse complementary repeats is less than the shortest retained length of the terminal reverse complementary repeats;
[0071] If the length of the tandem repeat sequence is less than the retained length of the shortest terminal tandem repeat sequence, the tandem repeat sequence is deleted; otherwise, the tandem repeat sequence is retained;
[0072] If the length of the reverse complementary repeat sequence is less than the retained length of the shortest terminal reverse complementary repeat sequence, the reverse complementary repeat sequence is deleted; otherwise, the reverse complementary repeat sequence is retained.
[0073] Whether to remove terminal repeat sequences removeTR includes whether to remove tandem repeat sequences and reverse complementary repeat sequences. You can set removeTR=TRUE or removeTR=FALSE. The default is not to delete.
[0074] The minimum terminal tandem repeat length can be set to MIN_LENGTH_DTR = 10, which is 20 bp by default. The minimum terminal inverse complementary repeat length can be set to MIN_LENGTH_ITR = 10, which is 20 bp by default. Repeat sequences shorter than 20 bp will be deleted.
[0075] Based on the above embodiment, this embodiment further includes:
[0076] Determining the position and content of the terminal repeat sequence in the viral genome sequence;
[0077] The position and content of the terminal repeat sequence are visually displayed.
[0078] A visual diagram of the terminal repeat sequence of the viral genome is shown in Figure 3 This embodiment supports visual display of the location and content of repeated sequences, which helps researchers to intuitively understand the structural characteristics of the genome.
[0079] Based on the above embodiment, the open reading frame in this embodiment includes an open reading frame without a start codon and / or a stop codon.
[0080] Based on the above embodiment, in this embodiment, the second parameter includes whether to retain all found open reading frames and the shortest open reading frame length. Deleting or retaining the open reading frame according to the second parameter set by the user includes:
[0081] When the value of "whether to retain all found open reading frames" is TRUE, all found open reading frames are retained;
[0082] When the question of whether to retain all found open reading frames is FALSE, for the open reading frames in the same translation frame, retain the longest open reading frame, and determine whether the length of the longest open reading frame is less than the length of the shortest open reading frame;
[0083] If the length of the longest open reading frame is less than the length of the shortest open reading frame, the longest open reading frame is deleted; otherwise, the longest open reading frame is retained.
[0084] You can set whether to keep all found open reading frames (KeepAllOrff = TRUE or KeepAllOrff = FALSE). The default is not to keep them. For open reading frames in the same translation frame, the longest open reading frame is kept.
[0085] Allow users to customize the minimum open reading frame length MinLenORF=200, which defaults to 200bp. Open reading frames shorter than 200bp will be deleted, and open reading frames can be selectively retained or deleted to achieve flexible screening of open reading frames.
[0086] Based on the above embodiment, this embodiment further includes:
[0087] Based on the annotation of the viral genome sequence, an annotation file in GFF format that meets the submission standards of NCBI and NGDC, a tbl file for submission to NCBI and NGDC, and a GenBank file for local viewing and editing are automatically generated.
[0088] This embodiment generates standardized file formats, can automatically generate gene annotation files in GFF format, feature files in tbl format, and GenBank files for local viewing and editing, supports direct submission to NCBIgenbank and NGDC-CNCBgenbase, and shares genomic data, greatly simplifying the submission and sharing process of genomic data.
[0089] Figure 4 This is a flow chart of the structural and functional annotation method based on viral genome sequences provided by the present invention. First, a FASTA file of the viral genome sequence is input. The system detects the presence of tandem repeats at both ends of the viral sequence and removes them according to user-defined parameters. Next, the system predicts all open reading frames in the genome and identifies multiple open reading frames with viral protein functions through blastp functional annotation. If the system identifies an ambiguous virus, it prioritizes retaining viral protein ORFs in the 5'->3' direction, successfully resolving the ORF overlap issue. The generated GFF and tbl files are submitted directly to NCBI via the system. The entire annotation process takes only a few minutes, significantly reducing manual adjustment time compared to traditional annotation tools.
[0090] Compared to commonly used command-line tools like PROKKA, GLIMMER, and VAPiD, this system demonstrates greater efficiency and flexibility in processing viral genomes. While traditional tools often require manual intervention when handling ambiguous viruses and ORF overlap, this system automatically identifies and processes these features, significantly improving the accuracy of viral genome annotation. Furthermore, the generated files conform to international standards, facilitating the sharing of genomic data.
[0091] Compared to existing tools, this method significantly reduces runtime when processing large-scale viral genome data. Furthermore, it is applicable to all viral sequences, not just specific viral taxa, enhancing the tool's versatility and flexibility. Operation is simple, with all functional modules highly integrated, requiring only a single command.
[0092] The viral sequence genome structure and function annotation system provided by the present invention is described below. The viral sequence genome structure and function annotation system described below and the viral sequence genome structure and function annotation method described above can be referenced to each other.
[0093] like Figure 5As shown, the system includes a terminal repeat sequence detection module 501, an open reading frame prediction module 502, a function annotation module 503 and an ambiguous virus identification module 504, wherein:
[0094] The terminal repeat sequence detection module 501 is used to identify the terminal repeat sequence in the viral genome sequence by aligning the viral genome sequence itself, and delete or retain the terminal repeat sequence according to a first parameter set by the user;
[0095] The open reading frame prediction module 502 is used to predict the open reading frame in the viral genome sequence using the getORF software, and delete or retain the open reading frame according to a second parameter set by the user;
[0096] The function annotation module 503 is used to annotate the viral genome sequence using the blastp tool, identify open reading frames with viral protein functions based on the functional annotations of the open reading frames in the annotations, determine virus classification information based on the annotations of the open reading frames with viral protein functions, and adjust the orientation of the viral genome sequence based on the annotations;
[0097] The ambisense virus identification module 504 is used to identify ambisense viruses based on the virus classification information, retain overlapping open reading frames in different translation frames in the viral genome sequence of the ambisense virus according to the tolerance of open reading frame overlap set by the user, and give priority to retaining open reading frames with viral protein functions in the same direction.
[0098] This embodiment takes into account the special structure of the viral genome, including terminal repeats, overlapping characteristics of open reading frames, and bidirectional translation of ambiguous viruses, when annotating the viral genome sequence. It can automatically identify and process the data, thereby improving the accuracy of viral genome annotation and demonstrating higher efficiency and flexibility.
[0099] Figure 6 An example of a physical structure diagram of an electronic device is shown below. Figure 6As shown, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communications bus 640, wherein the processor 610, the communications interface 620, and the memory 630 communicate with each other via the communications bus 640. The processor 610 may call logic instructions in the memory 630 to execute a method for annotating the structure and function of a viral sequence genome, the method comprising: identifying terminal repeats in a viral genome sequence, deleting or retaining the terminal repeats according to a first parameter set by a user; predicting open reading frames in a viral genome sequence, deleting or retaining the open reading frames according to a second parameter set by a user; annotating the viral genome sequence using a blastp tool, identifying open reading frames with viral protein functions based on the functional annotations of the open reading frames, determining virus classification information, and adjusting the orientation of the viral genome sequence based on the annotations; identifying ambiguous viruses based on the virus classification information, retaining overlapping open reading frames in different translation frames of the ambiguous viruses based on a tolerance for overlap of open reading frames set by the user, and prioritizing open reading frames in the same orientation.
[0100] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0101] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the viral sequence genome structure and function annotation method provided by the above methods, the method including: identifying terminal repeat sequences in the viral genome sequence, deleting or retaining the terminal repeat sequences according to a first parameter set by the user; predicting open reading frames in the viral genome sequence, deleting or retaining the open reading frames according to a second parameter set by the user; annotating the viral genome sequence by a blastp tool, identifying open reading frames with viral protein functions according to the functional annotation of the open reading frames, determining virus classification information, and adjusting the direction of the viral genome sequence according to the annotation; identifying ambiguous viruses according to the virus classification information, retaining overlapping open reading frames in different translation frames of ambiguous viruses according to the tolerance degree of overlap of open reading frames set by the user, and giving priority to retaining open reading frames in the same direction.
[0102] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the viral sequence genome structure and function annotation method provided by the above-mentioned methods, the method comprising: identifying terminal repeat sequences in the viral genome sequence, deleting or retaining the terminal repeat sequences according to a first parameter set by a user; predicting open reading frames in the viral genome sequence, deleting or retaining the open reading frames according to a second parameter set by a user; annotating the viral genome sequence by a blastp tool, identifying open reading frames with viral protein functions according to the functional annotations of the open reading frames, determining the virus classification information, and adjusting the direction of the viral genome sequence according to the annotations; identifying ambiguous viruses according to the virus classification information, retaining overlapping open reading frames in different translation frames of ambiguous viruses according to the tolerance degree of overlap of the open reading frames set by the user, and giving priority to retaining open reading frames in the same direction.
[0103] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0104] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for annotating the structure and function of viral sequence genomes, characterized in that: include: Identifying terminal repeat sequences in the viral genome sequence by aligning the viral genome sequence itself, and deleting or retaining the terminal repeat sequences according to a first parameter set by the user; Predicting the open reading frame in the viral genome sequence using getORF software, and deleting or retaining the open reading frame according to a second parameter set by the user; Annotating the viral genome sequence using a blastp tool, identifying open reading frames with viral protein functions based on functional annotations of the open reading frames in the annotations, determining virus classification information based on the annotations of the open reading frames with viral protein functions, and adjusting the orientation of the viral genome sequence based on the annotations; Ambiguous viruses are identified according to the virus classification information, and overlapping open reading frames in different translation frames in the viral genome sequence of the ambiguous virus are retained according to the tolerance of open reading frame overlap set by the user, and open reading frames with viral protein functions in the same direction are preferentially retained.
2. The method for annotating viral sequence genome structure and function according to claim 1, characterized in that: The terminal repeat sequence includes a tandem repeat sequence and an inverted complementary repeat sequence.
3. The method for annotating viral sequence genome structure and function according to claim 2, characterized in that: The first parameter includes whether to remove the terminal repeat sequence, the shortest terminal tandem repeat sequence retention length, and the shortest terminal reverse complementary repeat sequence retention length. Deleting or retaining the terminal repeat sequence according to the first parameter set by the user includes: When the value of whether to remove terminal repeat sequences is TRUE, the tandem repeat sequences and reverse complementary repeat sequences are deleted; If the value of whether to remove terminal repeats is FALSE, determining whether the length of the tandem repeats is less than the shortest retained length of the terminal tandem repeats, and whether the length of the reverse complementary repeats is less than the shortest retained length of the terminal reverse complementary repeats; If the length of the tandem repeat sequence is less than the retained length of the shortest terminal tandem repeat sequence, the tandem repeat sequence is deleted; otherwise, the tandem repeat sequence is retained; If the length of the reverse complementary repeat sequence is less than the retained length of the shortest terminal reverse complementary repeat sequence, the reverse complementary repeat sequence is deleted; otherwise, the reverse complementary repeat sequence is retained.
4. The method for annotating viral sequence genome structure and function according to claim 1, characterized in that: After identifying the terminal repeat sequence in the viral genome sequence by aligning the viral genome sequence itself, the method further includes: Determining the position and content of the terminal repeat sequence in the viral genome sequence; The position and content of the terminal repeat sequence are visually displayed.
5. The method for annotating viral sequence genome structure and function according to claim 1, wherein: The open reading frame includes an open reading frame without a start codon and / or a stop codon.
6. The method for annotating viral sequence genome structure and function according to claim 1, characterized in that: The second parameter includes whether to retain all found open reading frames and the shortest open reading frame length. Deleting or retaining the open reading frame according to the second parameter set by the user includes: When the value of "whether to retain all found open reading frames" is TRUE, all found open reading frames are retained; When the question of whether to retain all found open reading frames is FALSE, for the open reading frames in the same translation frame, retain the longest open reading frame, and determine whether the length of the longest open reading frame is less than the length of the shortest open reading frame; If the length of the longest open reading frame is less than the length of the shortest open reading frame, the longest open reading frame is deleted; otherwise, the longest open reading frame is retained.
7. The method for annotating the viral sequence genome structure and function according to any one of claims 1 to 6, characterized in that: Also includes: Based on the annotation of the viral genome sequence, an annotation file in GFF format, a tbl file for submission to NCBI and NGDC, and a GenBank file for local viewing and editing were generated.
8. A viral sequence genome structure and function annotation system, characterized in that: include: a terminal repeat sequence detection module, configured to identify terminal repeat sequences in the viral genome sequence by aligning the viral genome sequence itself, and to delete or retain the terminal repeat sequences according to a first parameter set by a user; an open reading frame prediction module, configured to predict the open reading frame in the viral genome sequence using the getORF software, and to delete or retain the open reading frame according to a second parameter set by the user; a functional annotation module, configured to annotate the viral genome sequence using a blastp tool, identify open reading frames having viral protein functions based on the functional annotations of the open reading frames in the annotations, determine virus classification information based on the annotations of the open reading frames having viral protein functions, and adjust the orientation of the viral genome sequence based on the annotations; The ambisense virus identification module is used to identify ambisense viruses according to the virus classification information, retain overlapping open reading frames in different translation frames in the viral genome sequence of the ambisense virus according to the tolerance of open reading frame overlap set by the user, and give priority to retaining open reading frames with viral protein functions in the same direction.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for annotating the viral sequence genome structure and function as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for annotating the viral sequence genome structure and function as claimed in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method, system, storage medium and device for annotating macro virus group original sequencing data short-read sequence
CN114093416A
Method for splicing and annotating mitochondrial genomes of earthbees
CN116030891A