Virus genome splicing method and system combining reference sequence and de novo splicing
By combining reference sequences and de novo splicing methods, the problems of low accuracy of viral genome splicing, incomplete removal of host sequences and difficulty in assembly of non-coding regions in the prior art are solved, and high accuracy and high efficiency of viral genome assembly are achieved.
Patent Information
- Application Number
- CN202411821976.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-11
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-11
AI Technical Summary
In the prior art, viral genome assembly tools based on de novo splicing have problems such as low splicing accuracy, incomplete host sequence removal and difficulty in assembly of non-coding regions.
Using a method of combining reference sequences and de novo splicing, the host genome sequence is removed through quality control and k-mers strategy to generate preliminary assembly sequences of the viral genome, and then the reference sequence is selected according to the virus classification annotation of the blastn tool, and the final viral genome sequence is adjusted through sequence cluster division and consensus sequence generation.
It significantly improves the accuracy of viral genome splicing, reduces base differences in de novo splicing, and can accurately determine the length of the 5' and 3' non-coding regions of the virus, improving assembly consistency and efficiency.
Smart Images

Figure CN119943150A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of viral genome splicing, and in particular to a viral genome splicing method and system combining reference sequence and de novo splicing. Background Art
[0002] The sequencing and assembly of viral genomes are core tasks in virology research. With the rapid development of next-generation sequencing (NGS), researchers can quickly generate large amounts of viral genome data, promoting research in areas such as viral evolution, pathogenic mechanisms, and epidemiology. The precise assembly of viral genomes is the basis for understanding viral genetic information, predicting its host range, and developing antiviral drugs and vaccines. For example, in recent years, through genome sequencing technology, scientists have been able to quickly identify new viruses and reveal their transmission pathways and potential pathogenic mechanisms through the comparison and annotation of gene sequences.
[0003] However, although high-throughput sequencing technology has greatly reduced sequencing costs and can generate massive amounts of data in a short period of time, these data are usually short fragments of reads that need to be assembled to restore the complete viral genome sequence. Assembly of viral genomes usually faces the following challenges:
[0004] Short read length and complex assembly: The viral genome is usually small, but its genome structure is complex and diverse, including highly repetitive regions, non-coding regions (such as 5' and 3' non-coding regions), and overlapping regions of open reading frames, which increase the difficulty of assembly. Short read sequencing data (usually 100 to 300bp) is difficult to cover all regions of the viral genome, especially when the sequencing depth is insufficient or the diversity of viral sequences is high, splicing errors are prone to occur.
[0005] Host genome contamination: Virus sequencing data usually comes from host samples (such as blood or tissue samples), so a large number of host genome sequences are mixed in the sequencing data. How to efficiently remove these host sequences and retain viral sequences is a key step in high-quality assembly. Host sequence contamination not only increases the complexity of data processing, but also interferes with the assembly and annotation of viral genomes.
[0006] Diversity of viral genomes and assembly of non-coding regions: The viral genome is highly diverse, and the genome sequences of different virus species and strains vary greatly. Even different regions within the same virus have different mutation rates. Especially in the 5' and 3' non-coding regions, the length and sequence of these regions are highly uncertain. Existing de novo assembly tools are difficult to accurately determine the length and sequence of these regions, which leads to errors in the splicing of viral genomes. Summary of the invention
[0007] The present invention provides a method and system for assembling viral genomes by combining reference sequences and de novo splicing, which is used to solve the limitations of viral genome assembly tools based on de novo splicing in the prior art and the defects of low accuracy of viral genome splicing, and realizes an extended splicing strategy that can effectively combine reference sequences to reduce errors in de novo splicing, and can automatically select appropriate reference sequences to ensure assembly accuracy. At the same time, it has an efficient host sequence removal function and the ability to process low-quality reference sequences.
[0008] The present invention provides a method for assembling viral genomes by combining reference sequences and de novo assembly, comprising:
[0009] After quality control of the sequencing data of the viral genome sequence, the host genome sequence in the sequencing data is removed based on the k-mers strategy;
[0010] De novo assembly of the sequencing data after removal of the host genome sequence to generate a preliminary assembly sequence of the viral genome;
[0011] Performing virus classification annotation on the preliminary assembled sequence of the viral genome based on the blastn tool, selecting a viral genome sequence from a virus library as a reference sequence according to the virus classification annotation, wherein the similarity between the reference sequence and the preliminary assembled sequence of the viral genome is greater than a first preset threshold;
[0012] Generate a data set based on the reference sequence and the preliminary assembled sequence of the viral genome, divide the data set into multiple sequence clusters based on the similarity between the sequences in the data set, and each sequence cluster contains at least one preliminary assembled sequence of the viral genome, and generate a consensus sequence based on the base with the largest proportion at the same position of the sequence in each sequence cluster;
[0013] The reads in the sequencing data after removing the host genome sequence are aligned with the consensus sequence, and the consensus sequence is adjusted according to the alignment result to obtain the final viral genome sequence.
[0014] According to a method for assembling viral genomes combining reference sequences and de novo assembly provided by the present invention, before selecting a viral genome sequence as a reference sequence from a virus library according to the virus classification annotation, the method further comprises:
[0015] The virus-related sequences in the preliminary assembled sequence of the viral genome are retained, and the sequences unrelated to the virus are removed.
[0016] According to a method for assembling a viral genome by combining a reference sequence and de novo assembly provided by the present invention, sequences related to the virus in the preliminary assembly sequence of the viral genome are retained, and sequences unrelated to the virus are removed, comprising:
[0017] Retaining sequences in the preliminary assembled sequence of the viral genome whose nucleotide similarity is greater than a second preset threshold and whose nucleotide coverage is greater than a third preset threshold as sequences related to the virus;
[0018] The sequences other than the virus-related sequences in the preliminary assembled sequence of the viral genome are removed as sequences unrelated to the virus.
[0019] According to a method for assembling a viral genome combining a reference sequence and de novo assembly provided by the present invention, a viral genome sequence is selected from a virus library as a reference sequence according to the virus classification annotation, comprising:
[0020] Selecting, from the virus library according to the virus classification annotation, a virus genome sequence whose similarity to the preliminary assembled virus genome sequence is greater than the first preset threshold;
[0021] After replacing the non-ATGC bases in the selected viral genome sequence with N bases, filtering out viral genome sequences in which the proportion of N bases exceeds a fourth preset threshold or the number of N bases is greater than a fifth preset threshold;
[0022] The filtered viral genome sequence is used as the reference sequence.
[0023] According to a viral genome assembly method combining a reference sequence and de novo assembly provided by the present invention, quality control of sequencing data of a viral genome sequence is performed, comprising:
[0024] The adapter sequences, reads with quality values less than a sixth preset threshold, and sequences with consecutive repeated bases in the sequencing data are removed.
[0025] According to a viral genome splicing method combining reference sequence and de novo splicing provided by the present invention, the sequencing data is double-end sequencing data or single-end sequencing data.
[0026] The present invention also provides a viral genome splicing system combining reference sequence and de novo splicing, comprising:
[0027] A quality control module, used to perform quality control on sequencing data of viral genome sequences;
[0028] A host removal module, used to remove host genome sequences in the sequencing data based on a k-mers strategy;
[0029] A de novo assembly module, used to assemble the sequencing data from scratch after removing the host genome sequence to generate a preliminary assembly sequence of the viral genome;
[0030] A classification annotation module, used for performing virus classification annotation on the preliminary assembled sequence of the virus genome based on the blastn tool;
[0031] A reference sequence pull-down module, used to select a viral genome sequence from a virus library as a reference sequence according to the viral classification annotation, wherein the similarity between the reference sequence and the preliminary assembled sequence of the viral genome is greater than a first preset threshold;
[0032] A consensus sequence generation module, used to generate a data set based on the reference sequence and the preliminary assembly sequence of the viral genome, divide the data set into multiple sequence clusters according to the similarity between the sequences in the data set, and each sequence cluster contains at least one preliminary assembly sequence of the viral genome, and generate a consensus sequence according to the base with the largest proportion at the same position of the sequence in each sequence cluster;
[0033] A reads comparison module, used to compare the reads in the sequencing data after removing the host genome sequence with the consensus sequence;
[0034] The sequence adjustment module is used to adjust the consensus sequence according to the comparison result to obtain the final viral genome sequence.
[0035] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for assembling viral genomes combining reference sequences and de novo assembly as described above is implemented.
[0036] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for assembling viral genomes combining reference sequences and de novo assembly.
[0037] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements any of the above-described methods for assembling viral genomes combining reference sequences and de novo assembly.
[0038] The present invention provides a method and system for splicing viral genomes in combination with reference sequences and de novo splicing. The method and system perform extended splicing in combination with reference sequences, firstly obtain a preliminary assembly sequence of the viral genome by de novo splicing, then automatically select a reference sequence with high similarity to the preliminary assembly sequence of the viral genome, and integrate the preliminary assembly sequence of the viral genome with the reference sequence, thereby reducing base differences in the de novo splicing, ensuring that the spliced sequence is closer to the real viral genome, and significantly improving the accuracy of the splicing. In addition, the lengths of the 5' and 3' non-coding regions of the virus can be accurately determined through the extended splicing of the reference sequence, which is of great significance to the functional research of the viral genome. The method for dividing sequence clusters based on similarity ensures that each cluster contains at least one preliminary assembly sequence of the viral genome spliced from scratch, and generates a consensus sequence through multiple sequence alignment. It not only improves the accuracy of assembly, but also generates more consistent splicing results, which is suitable for the splicing of diverse virus strains; by aligning the host-free reads to the generated consensus sequence and correcting the consensus sequence according to the alignment results, the base errors in splicing are further reduced; compared with the completely de novo splicing method, by combining the extended splicing of the reference sequence, the splicing time and computing resource requirements are significantly reduced, and high-quality splicing results can still be quickly generated when processing data with low sequencing depth or poor sequencing quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0040] Figure 1 It is one of the flow diagrams of the viral genome splicing method combining reference sequence and de novo splicing provided by the present invention;
[0041] Figure 2 This is the second schematic diagram of the process of the viral genome splicing method combining reference sequence and de novo splicing provided by the present invention;
[0042] Figure 3 It is a schematic diagram of the splicing method of the viral genome combining reference sequence and de novo splicing provided by the present invention;
[0043] Figure 4 It is a schematic diagram of the comparison between the final viral genome sequence and Sanger first generation sequencing in the viral genome splicing method combining the reference sequence and de novo splicing provided by the present invention;
[0044] Figure 5It is a schematic diagram of the structure of the viral genome splicing system combining reference sequence and de novo splicing provided by the present invention;
[0045] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0047] Currently, viral genome assembly tools based on de novo assembly, such as SPAdes, Velvet, IDBA-UD, etc., have been widely used in the assembly of viral genomes. These tools attempt to reconstruct complete genomes through complex graph algorithms (such as Hamiltonian paths or Eulerian paths). However, these tools have the following limitations when processing viral genomes:
[0048] Poor splicing accuracy: De novo splicing tools are prone to introduce erroneous splicing results when the sequencing depth is insufficient or the sequencing data quality is poor, especially in the highly repetitive regions or non-coding regions of the viral genome, where the splicing results differ greatly from the true sequence.
[0049] The length of the non-coding region is difficult to determine: The 5' and 3' non-coding regions of the viral genome are crucial for viral replication and regulatory functions, but the length and sequence of these regions are often difficult to accurately assemble by de novo splicing. In addition, due to the lack of obvious reading frames or protein coding signals in the non-coding regions, existing splicing tools have difficulty accurately identifying and assembling these regions.
[0050] Incomplete host sequence removal: Although existing tools provide the function of host sequence removal, the effect is limited, especially for incompletely annotated host genomes, the removal effect is not ideal. This results in viral genome assembly data often being mixed with a large number of host background sequences, interfering with the accurate identification of viral sequences.
[0051] Reference sequence selection problem: In the reference sequence-based assembly method, how to select a suitable reference sequence is an important factor affecting the accuracy of assembly. The difference between the reference sequence and the sequence to be assembled may lead to incorrect sequence assembly, especially in the case of large diversity of viral genomes. The wrong reference sequence will introduce incorrect base information, affecting the accuracy of the final assembly result.
[0052] In order to improve the accuracy of assembly when the sequencing depth is insufficient, many studies have begun to combine reference sequences for assembly. For example, the QUAST software can evaluate the quality of the assembly results and identify errors in the assembly by comparing with the reference sequence. However, the existing reference sequence-based splicing tools still face several problems: first, how to select a reference sequence similar to the target sequence is still a difficulty; second, the difference between the reference sequence and the sequence to be assembled will introduce splicing bias, especially when dealing with virus strains with large variations, which can easily lead to incorrect splicing.
[0053] In the assembly of viral genomes, removing host genome sequences is a crucial step. Usually, the method of removing host genomes relies on alignment tools, such as BWA, Bowtie2, etc., to remove host sequences by aligning sequencing data to the host genome reference sequence. However, this method has shortcomings in the following situations:
[0054] Incomplete annotation of host genomes: For some incompletely annotated host genomes, especially novel or rare host species, existing host removal methods are ineffective, resulting in a large number of host sequences not being effectively removed, affecting the splicing and annotation of the viral genome.
[0055] Limitations of the alignment strategy: Some low-complexity regions or sequences with high similarity are prone to alignment errors, and host sequences cannot be accurately eliminated and mixed into the virus splicing results.
[0056] To solve these problems, researchers have developed a variety of host removal strategies, such as the host removal method based on the k-mers strategy, which can improve the effect of host removal to a certain extent through more efficient alignment and filtering algorithms. This method can not only process known host genomes, but also effectively remove some unannotated genomes. However, the existing technology still needs to be further optimized to improve its ability to handle complex host backgrounds.
[0057] De novo assembly is a common method for assembling the entire viral genome, especially when there is no suitable reference sequence or an unknown viral sequence. Common de novo assembly tools such as SPAdes and Velvet rely on the coverage depth and sequencing quality of short reads. However, low sequencing depth, short read length, and highly repetitive regions can lead to assembly errors, especially in the non-coding and low-complexity regions of the viral genome, where the assembly results differ greatly from the true sequence.
[0058] In addition, de novo assembly is particularly difficult for the 5' and 3' non-coding regions, which are usually regulatory regions of the viral genome and have large length and sequence variations. Traditional assembly tools have difficulty accurately determining the length and sequence of these regions.
[0059] In order to make up for the shortcomings of de novo splicing, many studies have proposed a splicing method based on reference sequences, that is, by aligning the sequencing data with the known reference sequence, the splicing is extended. This method has important application value in the assembly of viral genomes, especially when the sequencing depth is not enough, the reference sequence can be used to correct the splicing results. However, how to choose a suitable reference sequence is the key to successful splicing.
[0060] Regarding the selection of reference sequences, the sequences of different virus strains vary greatly. If the selected reference sequence differs too much from the virus sequence to be assembled, it is easy to introduce erroneous base information, resulting in splicing deviation. Therefore, the splicing method based on the reference sequence needs to be able to automatically select a reference sequence with high similarity to the sequence to be spliced, and perform quality control on the reference sequence itself.
[0061] For the quality control of the reference sequence, there may be low-quality regions in the known reference sequence, such as regions with more non-ATGC bases or N bases, which will have a negative impact on the splicing results. Therefore, the reference sequence must be strictly quality controlled and filtered to ensure that the reference sequence used for splicing has a high accuracy.
[0062] Combine the following Figure 1 A method for assembling viral genomes combining reference sequences and de novo assembly is described, comprising:
[0063] Step 101, after quality control of the sequencing data of the viral genome sequence, the host genome sequence in the sequencing data is removed based on the k-mers strategy;
[0064] First, the sequencing data of the viral genome sequence is quality controlled to ensure the quality of the data for subsequent analysis.
[0065] Then, based on the k-mers strategy, the host genome sequences in the sequencing data were removed to ensure that the retained sequences were mainly derived from viruses.
[0066] The host removal method based on the k-mers strategy uses the host genome reference sequence for comparison to remove host contamination sequences and ensure the integrity of the viral sequence. The host genome reference sequence HOSTdb is a specified host sequence, which can include all host sequences at the order, family, and genus levels.
[0067] The host sequence removal technology based on the k-mers strategy can effectively remove the host genome sequence in the alignment step, especially for incompletely annotated host species. This technology can handle complex host backgrounds, ensure that only viral sequences are retained, and reduce the interference of host contamination on subsequent splicing.
[0068] Step 102, de novo splicing of the sequencing data after removing the host genome sequence to generate a preliminary assembly sequence of the viral genome;
[0069] After removing the host data, the remaining viral data are assembled de novo to generate a preliminary assembly sequence of the viral genome.
[0070] Step 103, performing virus classification annotation on the preliminary assembled viral genome sequence based on the blastn tool, selecting a viral genome sequence from a virus library as a reference sequence according to the virus classification annotation, and the similarity between the reference sequence and the preliminary assembled viral genome sequence is greater than a first preset threshold;
[0071] The de novo assembled preliminary viral genome was classified and annotated using the blastn tool to ensure that the assembly results were consistent with the viral genome.
[0072] According to the similarity score between the preliminary assembled viral genome sequence and the viral sequence in the virus library NTdb in the classification annotation results, the known viral sequences with high similarity to the preliminary assembled viral genome sequence, such as those with a similarity greater than 99%, are automatically pulled down as reference sequences. The pulled-down reference sequences can be further quality controlled to ensure the high quality of the reference sequences.
[0073] Step 104, generating a data set according to the reference sequence and the preliminary assembled sequence of the viral genome, dividing the data set into a plurality of sequence clusters according to the similarity between the sequences in the data set, and each sequence cluster contains at least one preliminary assembled sequence of the viral genome, and generating a consensus sequence according to the base with the largest proportion at the same position of the sequence in each sequence cluster;
[0074] The reference sequence and the preliminary assembled sequence of the viral genome obtained by de novo splicing are combined into one data set, and divided into different sequence clusters according to the similarity between the sequences in the data set. For example, sequences with a similarity greater than 90% are divided into the same sequence cluster to ensure the diversity of splicing.
[0075] Each sequence cluster contains at least one sequence that was spliced from scratch, and through multiple sequence alignment within each sequence cluster, a consensus sequence is generated based on the most common bases at the same position. This not only improves the accuracy of the assembly, but also generates more consistent splicing results, which is suitable for the splicing of diverse virus strains.
[0076] Step 105, aligning the reads in the sequencing data after removing the host genome sequence with the consensus sequence, adjusting the consensus sequence according to the alignment result, and obtaining the final viral genome sequence.
[0077] The host reads were removed from the sequencing data, aligned to the generated consensus sequence, and the duplicate reads in the consensus sequence were marked to ensure the accuracy of the correction splicing results.
[0078] The consensus sequence is corrected according to the comparison results to ensure that the final output viral genome sequence is accurate. Specifically, the continuity of reads in the consensus sequence is determined based on repeated reads, and N bases are inserted into the consensus sequence with insufficient continuity; the base coverage in the consensus sequence is determined, and the bases of the consensus sequence with insufficient base coverage are replaced to generate the final viral genome sequence.
[0079] By accurately adding N bases, the regions that could not be sequenced or assembled accurately are replaced with N bases, thereby effectively connecting the fragmented assembly sequence. This strategy not only maintains the continuity of the sequence, but also provides clear guidance for subsequent further sequencing and genome improvement, ensuring that the generated sequence does not have splicing breaks or missing markers.
[0080] In this embodiment, by combining the reference sequence for extension and splicing, de novo splicing is first used to obtain the preliminary assembly sequence of the viral genome, and then a reference sequence with high similarity to the preliminary assembly sequence of the viral genome is automatically selected. By integrating the preliminary assembly sequence of the viral genome and the reference sequence, the base difference in the de novo splicing is reduced, ensuring that the spliced sequence is closer to the real viral genome, and significantly improving the accuracy of splicing. In addition, through the extension and splicing of the reference sequence, the length of the 5' and 3' non-coding regions of the virus can be accurately determined, which is of great significance for the functional research of the viral genome; the sequence clustering method based on similarity ensures that each cluster contains at least one de novo spliced preliminary assembly sequence of the viral genome, and a consensus sequence is generated by multiple sequence alignment. It not only improves the accuracy of assembly, but also generates more consistent splicing results, which is suitable for the splicing of diverse virus strains; by aligning the host-free reads to the generated consensus sequence and correcting the consensus sequence according to the alignment results, the base errors in splicing are further reduced; compared with the completely de novo splicing method, by combining the extended splicing of the reference sequence, the splicing time and computing resource requirements are significantly reduced, and high-quality splicing results can still be quickly generated when processing data with low sequencing depth or poor sequencing quality.
[0081] Based on the above embodiment, before selecting a viral genome sequence as a reference sequence from the virus library according to the virus classification annotation, this embodiment further includes:
[0082] The virus-related sequences in the preliminary assembled sequence of the viral genome are retained, and the sequences unrelated to the virus are removed.
[0083] On the basis of the above embodiment, this embodiment retains the virus-related sequences in the preliminary assembled sequence of the viral genome and removes the sequences unrelated to the virus, including:
[0084] Retaining sequences in the preliminary assembled sequence of the viral genome whose nucleotide similarity is greater than a second preset threshold and whose nucleotide coverage is greater than a third preset threshold as sequences related to the virus;
[0085] The sequences other than the virus-related sequences in the preliminary assembled sequence of the viral genome are removed as sequences unrelated to the virus.
[0086] In this embodiment, according to the customized nucleotide similarity (default is higher than 90%) and coverage (default is higher than 50%), sequences related to the virus are retained and other sequences not related to the virus are removed. That is, the second preset threshold value may be 90%, and the third preset threshold value may be 50%, but it is not limited thereto.
[0087] On the basis of the above embodiments, in this embodiment, a virus genome sequence is selected from a virus library as a reference sequence according to the virus classification annotation, including:
[0088] Selecting, from the virus library according to the virus classification annotation, a virus genome sequence whose similarity to the preliminary assembled virus genome sequence is greater than the first preset threshold;
[0089] After replacing the non-ATGC bases in the selected viral genome sequence with N bases, filtering out viral genome sequences in which the proportion of N bases exceeds a fourth preset threshold or the number of N bases is greater than a fifth preset threshold;
[0090] The filtered viral genome sequence is used as the reference sequence.
[0091] The fourth preset threshold may be 1%, and the fifth preset threshold may be 100, but are not limited thereto.
[0092] This embodiment provides a strict quality control mechanism for known reference sequences, such as replacing non-ATGC bases with N and filtering out sequences with a large proportion of N bases or a large number of N bases, to ensure the high quality of the reference sequence and effectively avoid the negative impact of low-quality sequences on the splicing results.
[0093] On the basis of the above embodiments, in this embodiment, the sequencing data of the viral genome sequence is quality controlled, including:
[0094] The adapter sequences, reads with quality values less than a sixth preset threshold, and sequences with consecutive repeated bases in the sequencing data are removed.
[0095] Perform quality control on the input sequencing data, remove adapter sequences, low-quality reads, and low-complexity sequences, and improve the accuracy of subsequent splicing.
[0096] Low-quality reads are reads whose quality values are less than a sixth preset threshold, and the sixth preset threshold may be Q20. Low-complexity sequences include sequences with consecutive repeated bases.
[0097] On the basis of the above embodiments, the sequencing data described in this embodiment is double-end sequencing data or single-end sequencing data.
[0098] This embodiment supports input of multiple sequencing formats, including paired-end and single-end sequencing data, ensures the compatibility of input data, and supports obtaining raw reads from paired-end and single-end sequencing data.
[0099] In practical applications, this embodiment is used to assemble the whole genome sequencing data of a new virus. The specific process is as follows Figure 2 As shown, the adapter sequences and low-quality regions in the sequencing data were removed by quality control, and the host contamination sequences were removed using the host removal module. The preliminary viral genome was generated by de novo assembly, and high-quality assembly results were obtained by extending the reference sequence.
[0100] Figure 3 How to obtain the viral genome by combining the preliminary assembly sequence of the viral genome and the reference sequence. Figure 3 The preliminary assembly results shown are composed of multiple discontinuous fragments with different lengths. By combining the reference genome information of different lengths that are highly similar to it, the preliminary assembled sequence can be extended and corrected to generate the final viral genome sequence. The final viral genome sequence can be obtained through a one-click command line operation.
[0101] Compared with traditional de novo splicing tools, this embodiment shows higher accuracy and efficiency in processing low-depth sequencing data. By extending the splicing of the reference sequence, the base differences in the splicing can be effectively avoided, and the length of the 5' and 3' non-coding regions of the viral genome can be accurately determined. The splicing results are highly consistent with the real viral genome (first-generation sequencing results), such as Figure 3 shown.
[0102] Figure 4 The lower line segment in the middle is the result of first-generation sequencing, and the upper line segment is the final viral genome sequence obtained. It can be seen that the 5' and 3' non-coding regions in the final viral genome sequence are longer than those in the first-generation sequencing.
[0103] The viral genome splicing system combining reference sequences and de novo splicing provided by the present invention is described below. The viral genome splicing system combining reference sequences and de novo splicing described below and the viral genome splicing method combining reference sequences and de novo splicing described above can be referenced to each other.
[0104] like Figure 5 As shown, the system includes a quality control module 501, a host removal module 502, a de novo assembly module 503, a classification annotation module 504, a reference sequence pull-down module 505, a consensus sequence generation module 506, a reads alignment module 507 and a sequence adjustment module 508, wherein:
[0105] The quality control module 501 is used to perform quality control on the sequencing data of the viral genome sequence;
[0106] The host removal module 502 is used to remove the host genome sequence in the sequencing data based on the k-mers strategy;
[0107] The de novo assembly module 503 is used to assemble the sequencing data from scratch after removing the host genome sequence to generate a preliminary assembly sequence of the viral genome;
[0108] The classification annotation module 504 is used to perform virus classification annotation on the preliminary assembled sequence of the virus genome based on the blastn tool;
[0109] The reference sequence pull-down module 505 is used to select a viral genome sequence from the virus library as a reference sequence according to the viral classification annotation, and the similarity between the reference sequence and the preliminary assembled sequence of the viral genome is greater than a first preset threshold;
[0110] The consensus sequence generation module 506 is used to generate a data set according to the reference sequence and the preliminary assembly sequence of the viral genome, divide the data set into multiple sequence clusters according to the similarity between the sequences in the data set, and each sequence cluster contains at least one preliminary assembly sequence of the viral genome, and generate a consensus sequence according to the base with the largest proportion at the same position of the sequence in each sequence cluster;
[0111] The reads comparison module 507 is used to compare the reads in the sequencing data after removing the host genome sequence with the consensus sequence;
[0112] The sequence adjustment module 508 is used to adjust the consensus sequence according to the comparison result to obtain the final viral genome sequence.
[0113] In this embodiment, by combining the reference sequence for extension and splicing, de novo splicing is first used to obtain the preliminary assembly sequence of the viral genome, and then a reference sequence with high similarity to the preliminary assembly sequence of the viral genome is automatically selected. By integrating the preliminary assembly sequence of the viral genome and the reference sequence, the base difference in the de novo splicing is reduced, ensuring that the spliced sequence is closer to the real viral genome, and significantly improving the accuracy of splicing. In addition, through the extension and splicing of the reference sequence, the length of the 5' and 3' non-coding regions of the virus can be accurately determined, which is of great significance for the functional research of the viral genome; the sequence clustering method based on similarity ensures that each cluster contains at least one de novo spliced preliminary assembly sequence of the viral genome, and a consensus sequence is generated by multiple sequence alignment. It not only improves the accuracy of assembly, but also generates more consistent splicing results, which is suitable for the splicing of diverse virus strains; by aligning the host-free reads to the generated consensus sequence and correcting the consensus sequence according to the alignment results, the base errors in splicing are further reduced; compared with the completely de novo splicing method, by combining the extended splicing of the reference sequence, the splicing time and computing resource requirements are significantly reduced, and high-quality splicing results can still be quickly generated when processing data with low sequencing depth or poor sequencing quality.
[0114] Figure 6 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 6 As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630 and a communication bus 640, wherein the processor 610, the communication interface 620 and the memory 630 communicate with each other through the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute a virus genome splicing method combining a reference sequence and de novo splicing, the method comprising: after quality control of the sequencing data of the virus genome sequence, removing the host genome sequence in the sequencing data based on the k-mers strategy; de novo splicing of the sequencing data to generate a preliminary assembly sequence of the virus genome; performing virus classification annotation on the preliminary assembly sequence of the virus genome based on the blastn tool, selecting a virus genome sequence with high similarity from the virus library as a reference sequence according to the virus classification annotation; generating a data set according to the reference sequence and the preliminary assembly sequence of the virus genome, dividing multiple sequence clusters according to the similarity between the sequences in the data set, and generating a consensus sequence; aligning the reads after removing the host genome sequence with the consensus sequence, adjusting the consensus sequence according to the alignment result, and obtaining the final virus genome sequence.
[0115] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk.
[0116] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the viral genome splicing method provided by the above methods that combines a reference sequence and de novo splicing, the method including: after quality control of the sequencing data of the viral genome sequence, removing the host genome sequence in the sequencing data based on the k-mers strategy; de novo splicing the sequencing data to generate a preliminary assembly sequence of the viral genome; annotating the preliminary assembly sequence of the viral genome with virus classification based on the blastn tool, and selecting a viral genome sequence with high similarity from the virus library as a reference sequence according to the virus classification annotation; generating a data set based on the reference sequence and the preliminary assembly sequence of the viral genome, dividing multiple sequence clusters according to the similarity between the sequences in the data set, and generating a consensus sequence; aligning the reads after removing the host genome sequence with the consensus sequence, adjusting the consensus sequence according to the alignment result, and obtaining the final viral genome sequence.
[0117] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the viral genome splicing method provided by the above-mentioned methods combining a reference sequence and de novo splicing, the method comprising: after quality control of the sequencing data of the viral genome sequence, removing the host genome sequence in the sequencing data based on the k-mers strategy; de novo splicing the sequencing data to generate a preliminary assembled sequence of the viral genome; annotating the preliminary assembled sequence of the viral genome by virus classification based on the blastn tool, and selecting a viral genome sequence with high similarity from the virus library as a reference sequence according to the virus classification annotation; generating a data set according to the reference sequence and the preliminary assembled sequence of the viral genome, dividing multiple sequence clusters according to the similarity between the sequences in the data set, and generating a consensus sequence; aligning the reads after removing the host genome sequence with the consensus sequence, adjusting the consensus sequence according to the alignment result, and obtaining the final viral genome sequence.
[0118] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0119] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for assembling viral genomes combining reference sequences and de novo assembly, characterized in that: include: After quality control of the sequencing data of the viral genome sequence, the host genome sequence in the sequencing data is removed based on the k-mers strategy; De novo assembly of the sequencing data after removal of the host genome sequence to generate a preliminary assembly sequence of the viral genome; Performing virus classification annotation on the preliminary assembled sequence of the viral genome based on the blastn tool, selecting a viral genome sequence from a virus library as a reference sequence according to the virus classification annotation, wherein the similarity between the reference sequence and the preliminary assembled sequence of the viral genome is greater than a first preset threshold; Generate a data set based on the reference sequence and the preliminary assembled sequence of the viral genome, divide the data set into multiple sequence clusters based on the similarity between the sequences in the data set, and each sequence cluster contains at least one preliminary assembled sequence of the viral genome, and generate a consensus sequence based on the base with the largest proportion at the same position of the sequence in each sequence cluster; The reads in the sequencing data after removing the host genome sequence are aligned with the consensus sequence, and the consensus sequence is adjusted according to the alignment result to obtain the final viral genome sequence.
2. The method for viral genome assembly combining reference sequence and de novo assembly according to claim 1, characterized in that: Before selecting a viral genome sequence as a reference sequence from a virus library according to the virus classification annotation, the method further comprises: The virus-related sequences in the preliminary assembled sequence of the viral genome are retained, and the sequences unrelated to the virus are removed.
3. The method for assembling viral genomes combining reference sequences and de novo assembly according to claim 2, characterized in that: The virus-related sequences in the preliminary assembled sequence of the virus genome are retained, and the sequences unrelated to the virus are removed, including: Retaining sequences in the preliminary assembled sequence of the viral genome whose nucleotide similarity is greater than a second preset threshold and whose nucleotide coverage is greater than a third preset threshold as sequences related to the virus; The sequences other than the virus-related sequences in the preliminary assembled sequence of the viral genome are removed as sequences unrelated to the virus.
4. The method for assembling viral genomes combining reference sequences and de novo assembly according to claim 1, characterized in that: Selecting a viral genome sequence as a reference sequence from a virus library according to the virus classification annotation includes: Selecting, from the virus library according to the virus classification annotation, a virus genome sequence whose similarity to the preliminary assembled virus genome sequence is greater than the first preset threshold; After replacing the non-ATGC bases in the selected viral genome sequence with N bases, filtering out viral genome sequences in which the proportion of N bases exceeds a fourth preset threshold or the number of N bases is greater than a fifth preset threshold; The filtered viral genome sequence is used as the reference sequence.
5. The method for assembling viral genomes by combining reference sequences and de novo assembly according to any one of claims 1 to 4, characterized in that: Perform quality control on the sequencing data of the viral genome sequence, including: The adapter sequences, reads with quality values less than a sixth preset threshold, and sequences with consecutive repeated bases in the sequencing data are removed.
6. The method for assembling viral genomes by combining reference sequences and de novo assembly according to any one of claims 1 to 4, characterized in that: The sequencing data is double-end sequencing data or single-end sequencing data.
7. A viral genome assembly system combining reference sequence and de novo assembly, characterized in that: include: A quality control module, used to perform quality control on sequencing data of viral genome sequences; A host removal module, used to remove host genome sequences in the sequencing data based on a k-mers strategy; A de novo assembly module, used to assemble the sequencing data from scratch after removing the host genome sequence to generate a preliminary assembly sequence of the viral genome; A classification annotation module, used for performing virus classification annotation on the preliminary assembled sequence of the virus genome based on the blastn tool; A reference sequence pull-down module, used to select a viral genome sequence from a virus library as a reference sequence according to the viral classification annotation, wherein the similarity between the reference sequence and the preliminary assembled sequence of the viral genome is greater than a first preset threshold; A consensus sequence generation module, used to generate a data set based on the reference sequence and the preliminary assembly sequence of the viral genome, divide the data set into multiple sequence clusters according to the similarity between the sequences in the data set, and each sequence cluster contains at least one preliminary assembly sequence of the viral genome, and generate a consensus sequence according to the base with the largest proportion at the same position of the sequence in each sequence cluster; A reads comparison module, used to compare the reads in the sequencing data after removing the host genome sequence with the consensus sequence; The sequence adjustment module is used to adjust the consensus sequence according to the comparison result to obtain the final viral genome sequence.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for assembling a viral genome combining a reference sequence and de novo assembly is implemented as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for assembling a viral genome combining a reference sequence and de novo assembly is implemented as claimed in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for assembling a viral genome combining a reference sequence and de novo assembly is implemented as claimed in any one of claims 1 to 6.
Citation Information
Patent Citations
Virus sequence assembly method based on low-depth siRNA data
CN111180014A
Optimal analysis method for macro virus group process
CN112750501A
Annotation method for sequencing data of single-bacterium DNA library and related equipment
CN114360647A
Analysis method for high-throughput prediction of phage hosts based on next-generation sequencing technology
CN115662516A
Virus genome identification and splicing method and application
CN116072222A
Cited By
Virus-host RNA sequence classification method and device based on AI
CN120977392A
An AI-based virus-host RNA sequence classification method and device
CN120977392B