Method for optimizing an assembled genome

By performing multiple assembly methods and characteristics optimization of the sequencing sequence set, the problem of difficulty in meeting high integrity and high continuity is solved in genome assembly, efficient optimization of the genome is achieved, and a genome version with high continuity and high intactness is obtained.

CN114496091BActive Publication Date: 2025-05-30ANNOROAD GENE TECHNOLOGY (BEIJING) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111660340.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-05-30
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

In the prior art, it is difficult for the assembled genome to meet the problems of high integrity and high continuity at the same time.

Method used

By assembling the sequencing sequence set in two or more ways, multiple initial genomes are obtained and optimized according to their continuity and integrity characteristics. The specific steps include traversing the continuity characteristics of each initial genome, selecting the genome with a better continuity as the base genome. When the integrity characteristics of other initial genomes are better than the base genome, use its dominant region to replace the corresponding region of the base genome to obtain an optimized genome.

Benefits of technology

Through this method, the genome integrity can be improved without damaging the continuity, thereby obtaining a genome version with high continuity and high intactness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003447372900000091
    Figure BDA0003447372900000091
  • Figure BDA0003447372900000101
    Figure BDA0003447372900000101
  • Figure HDA0003447372910000011
    Figure HDA0003447372910000011
Patent Text Reader

Abstract

The present invention provides a method for optimizing an assembled genome, and the method comprises the following steps: sequencing a sample to obtain a set of sequencing sequences; assembling the set of sequencing sequences in two or more ways to obtain two or more initial genomes, and obtaining a first feature and a second feature corresponding to each initial genome; traversing the first features of each initial genome, and taking the initial genome with the dominant first feature as the basic genome; when the second feature of any remaining initial genome except the basic genome is dominant relative to the second feature of the basic genome, using the dominant region of the remaining initial genome to replace the corresponding region of the basic genome, so as to obtain an optimized genome. The method of the present invention can use highly complete assembled sequences to correct low-complete assembled sequences, thereby obtaining a genome version with high continuity and high integrity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of sequencing technology, and specifically, relates to a method for optimizing an assembled genome. Background Art

[0002] The purpose of genome assembly is to obtain a genome with high continuity and high integrity. However, in the actual assembly process, different genome assembly software has differences in the assembly methods, and thus there are also differences in the sensitivity to heterozygous regions during use. Therefore, after assembly using different software, there are differences in the continuity and integrity of the genome. Some genomes have high integrity but low continuity; while some genomes have high continuity but low integrity.

[0003] Therefore, a method for obtaining a genome with high continuity and high integrity is needed. Summary of the Invention

[0004] Aiming at the problem that the assembled genome in the prior art of genome assembly cannot simultaneously meet the requirements of high integrity and high continuity, the present invention provides a method for optimizing an assembled genome.

[0005] Specifically, the present invention relates to the following aspects:

[0006] 1. A method for optimizing an assembled genome, characterized in that the method comprises the following steps:

[0007] Sequencing a sample to obtain a set of sequencing sequences;

[0008] Assembling the set of sequencing sequences in two or more ways to obtain two or more initial genomes, and obtaining a first feature and a second feature corresponding to each initial genome;

[0009] Traversing the first feature of each initial genome, and taking the initial genome with the dominant first feature as the basic genome;

[0010] When the second feature of any remaining initial genome other than the basic genome is dominant relative to the second feature of the basic genome, using the dominant region of the remaining initial genome to replace the corresponding region of the basic genome to obtain an optimized genome.

[0011] 2. The method according to item 1, characterized in that the first feature represents the continuity of the genome, and the second feature represents the integrity of the genome.

[0012] 3. The method according to item 2, characterized in that the first feature is the length of the sequence obtained by splicing in the genome, preferably Contig N50.

[0013] 4. The method according to item 2, wherein the second feature represents the integrity of the assembled sequence in the genome, preferably the C value of Busco.

[0014] 5. The method according to any one of items 1-4, wherein the sequence set is a sequence set obtained through quality control.

[0015] 6. The method according to any one of items 1-5, wherein replacing the corresponding region in the reference genome with the dominant region in the remaining initial genomes includes:

[0016] extending the dominant region in the remaining initial genomes,

[0017] extracting the extended region from the remaining initial genomes and performing sequence alignment with the reference genome to confirm the dominant extended region in the reference genome,

[0018] and replacing the dominant extended region with the corresponding region in the remaining initial genomes.

[0019] 7. The method according to item 6, wherein the length of the extension is 10 bp - 10 kbp, preferably 50 bp - 5 kbp, more preferably 500 bp - 1 kbp.

[0020] 8. The method according to item 6, wherein the sequence alignment is performed by Blat alignment.

[0021] 9. The method according to any one of items 1-8, wherein the dominant region is a region with dominant feature items.

[0022] 10. The method according to any one of items 1-9, wherein the source of the sample is an animal, a plant or a microorganism.

[0023] 11. The method according to any one of items 1-10, wherein the sequence set includes a collection of base sequence information and other sequence information.

[0024] 12. The method according to item 11, wherein the other sequence information includes base positions and sequence lengths.

[0025] 13. The method according to any one of items 1-12, wherein the sequencing is third-generation sequencing.

[0026] The method of the present invention can use highly complete assembled sequences to correct low-integrity assembled sequences, thereby obtaining a genome version with high contiguity and high integrity. Description of the Drawings

[0027] Figure 1 This is a schematic flowchart of an embodiment of the present invention. Detailed implementation manners

[0028] The present invention will be further described below in conjunction with embodiments. It should be understood that the embodiments are only used to further illustrate and explain the present invention, and are not used to limit the present invention.

[0029] Unless otherwise defined, the technical and scientific terms in this specification have the same meaning as commonly understood by those skilled in the art. Although methods and materials similar or identical to those described herein may be used in experiments or practical applications, the materials and methods are described below. In case of conflict, the present specification, including its definitions, shall prevail. In addition, the materials, methods, and examples are for illustrative purposes only and are not restrictive. The present invention will be further described below in conjunction with specific embodiments, but not to limit the scope of the present invention.

[0030] In view of the problems existing in the prior art, the present invention provides a method for optimizing an assembled genome, and the method includes the following steps:

[0031] Step 1: Sequencing a sample to obtain a set of sequencing sequences;

[0032] Step 2: Assembling the set of sequencing sequences in two or more ways to obtain two or more initial genomes, and obtaining a first feature and a second feature corresponding to each initial genome;

[0033] Step 3: Traversing the first feature of each initial genome, and taking the initial genome with the dominant first feature as the basic genome;

[0034] Step 4: When the second feature of any remaining initial genome other than the basic genome is dominant relative to the second feature of the basic genome, using the dominant region of the remaining initial genome to replace the corresponding region of the basic genome to obtain an optimized genome.

[0035] In Step 1, the source of the sample to be sequenced can be an animal, a plant, or a microorganism. The sequencing is a known and feasible sequencing technology, such as the second-generation high-throughput sequencing technology, the third-generation single-molecule sequencing technology, etc.

[0036] In a specific implementation manner, the sequencing is third-generation sequencing. Third-generation sequencing is a single-molecule sequencing technology that does not require PCR amplification and can achieve the technology of sequencing each DNA molecule separately. It has no GC bias and has a faster data reading speed. The application of third-generation sequencing technology is mainly in aspects such as genome sequencing, methylation research, and mutation identification (SNP detection).

[0037] Single-molecule sequencing refers to using DNA polymerase to synthesize a DNA strand complementary to the template, recording the template position and nucleotide sequence information in three-dimensional space, and then reverse-constructing the sequence of the DNA template. In addition to the three major elements of the DNA synthesis reaction (template, enzyme, nucleotide), the position where the template is located and the order of monochromatically fluorescently labeled nucleotides (such as A, C, G, T) in the reaction cycle are also key elements for the completion of the final DNA sequence. If the nucleotides used in the reaction are labeled with four different fluorescences, light of different wavelengths needs to be switched in each reaction cycle to record different bases.

[0038] Among the existing third-generation sequencing technologies, the single-molecule real-time sequencing system (Single Molecule Real Time, SMRT) developed by Pacific Biosciences and the nanopore single-molecule sequencing technology of Oxford Nanopore Technologies are relatively representative. Compared with the first-generation sequencing and the second-generation sequencing, their biggest feature is single-molecule sequencing, and PCR amplification is not required during the sequencing process.

[0039] In specific operations, for example, the third-generation sequencer PacBio Sequel can be used to sequence the sample.

[0040] The sequence set obtained by sequencing refers to the collection of sequences, for example, it can be a collection including base sequences and other sequence information. Among them, the other sequence information can be information such as base positions and sequence lengths.

[0041] The sequence set can be a sequence set directly obtained by sequencing or a sequence set obtained through quality control.

[0042] In a specific embodiment, the sequence set is a sequence set obtained through quality control. Specifically, the sequence set can be quality-controlled by filtering out low-quality sequences and removing adapter sequences. For example, for the PacBio sequencing platform, the off-machine data can be quality-controlled and data-converted by using the PacBio official quality control software SMRT Link to remove low-quality sequences.

[0043] In step two, the sequencing sequence set can be assembled in any known manner in the prior art. In a specific embodiment, the sequence assembly is de novo assembly, that is, Denovo assembly. The assembly software used can be related software known in the prior art, such as CANU, Flye, WTDBG2, hifiasm, etc. Using several methods for assembly can obtain several initial genomes. For example, when using two methods for assembly, two initial genomes can be obtained. When using three methods for assembly, three initial genomes can be obtained.

[0044] In the present invention, a feature may refer to an index capable of characterizing the quality of an assembled genome. The quality of an assembled genome can generally be evaluated using three principles as evaluation indices, namely the 3C principle: Contiguity, Correctness, and Completeness. Contiguity refers to obtaining a sequence (Contig) with sufficient length through splicing, which can be characterized by contig N50. Correctness means that the error rate of the assembled contig sequence is low. Completeness means that the assembled contig sequence contains as much of the entire genome information as possible. For example, BUSCO can be used for evaluation. The first feature and the second feature of the initial genome are indices for characterizing the quality of the initial genome, that is, they can reflect the quality of the assembled genome.

[0045] In a specific embodiment, the first feature represents the contiguity of the genome, and the second feature represents the completeness of the genome. The contiguity of the genome refers to whether the assembled genomic contigs are long enough. The completeness of the genome refers to whether the assembled genome contains all the sequence information of the species.

[0046] Both contiguity and completeness are important indices for evaluating the assembly effect of sequencing sequences. The purpose of genome assembly is to obtain a genome with high contiguity and high completeness. However, in the actual assembly process, it is often found that some genomes are highly complete but have low contiguity; while some genomes are highly contiguous but have low completeness. The method of the present invention is to obtain a genome with high contiguity and high completeness.

[0047] Therefore, through step two, the contiguity and completeness data of each initially assembled genome can be obtained respectively.

[0048] In a specific embodiment, the first feature representing the contiguity of the genome is the length of the sequence (Contig) obtained by splicing to characterize the assembly result, preferably Contig N50. Among them, Contig N50 means that by adding up the lengths of all Contigs, a total Contig length can be obtained, and then the lengths of all Contigs are sorted from long to short. For example, Contig 1, Contig 2, Contig 3... Contig 25 are obtained. The Contigs are added in this order successively. When the added length reaches half of the total Contig length, the length of the last added Contig is the Contig N50.

[0049] In a specific embodiment, the second feature is the C value of Busco.

[0050] Among them, as a method for evaluating transcriptome and genome integrity, BUSCO (Benchmarking Universal Single-Copy Orthologs) collects conserved sequences among related species, constructs gene sets for six major phylogenetic branches (Bacteria, Eukaryota, Protists, Metazoa, Fungi, Plants) using the OrthoDB ortholog database, and compares the assembled transcriptome and genome.

[0051] Although the genomes of each species are different, for species with close evolutionary relationships, there are always some conserved gene sequences among them. Based on this feature, BUSCO constructs a conserved gene database for major phylogenetic branches (OrthoDB database), and constructs core single-copy gene sets for several major phylogenetic branches respectively. After the initial assembly of the transcriptome or genome is completed, the assembly result can be compared with the core database of the major phylogenetic branch to which the species belongs, and the result is given according to whether the assembly result contains these core sequences, including single, multiple, partial or non-inclusion. Among them, for the genome, the BUSCO evaluation software first calls the Augustus software to predict the gene structure of the genome, and then uses HMMER3 to align to the reference gene set; for transcripts, after identifying the longest open reading frame, HMMER3 is used to align to the reference gene set. Finally, according to the proportion of aligned sequences, integrity, etc., the accuracy and integrity of the assembly result are evaluated. The BUSCO evaluation result will display values such as C (complete), S (singel-copy), D (duplicated), F (Fragmented), M (Missing). Generally, the value of S+D is the C value. Usually, the larger the C value, the better the integrity of the assembled sequence it reflects. If the D value is larger, it may mean that there is a greater possibility of redundancy in the assembled sequence, or it may be due to a recent whole-genome duplication event in the genome. When the C value in the BUSCO evaluation of the genome is relatively low, or when the redundancy of the original assembled genome leads to a large decrease in the C value of the BUSCO evaluation, it is necessary to increase the C value and retrieve the lost genes.

[0052] For multiple genes in the genome, they are evaluated as having or not having the gene in the BUSCO evaluation based on whether they are included in the busco library.

[0053] In step three, traverse the first feature of each initial genome, and use the initial genome with the dominant first feature as the basic genome. Among them, traversing the first feature of each initial genome includes searching for the first feature of each assembled initial genome and comparing to determine the initial genome with the dominant first feature.

[0054] Among them, the feature being dominant means that when multiple genomes are evaluated using a certain evaluation method, in a certain genome, this feature is evaluated as a dominant feature item relative to other genomes. Similarly, the first feature being dominant means that when multiple genomes are evaluated using a certain evaluation method, in a certain genome, its first feature is evaluated as a dominant feature item relative to other genomes. Specifically, for example, when using the BUSCO method to evaluate two genomes, when the first feature represents the completeness of the genome, in one genome, its completeness is evaluated as a dominant feature item relative to the other genome. A dominant feature item refers to a feature item that represents a genome having better quality when using a certain or certain features to describe the quality differences of a series of genomes. For example, when the feature is the completeness of the genome, the C value (Complete BUSCOs value) of BUSCO can be used to represent the completeness of the genome. At this time, a higher C value represents better genome quality, that is, a higher C value is a dominant feature item. Specifically, the quality of a genome with a completeness of 90% is better than that of a genome with a completeness of 80%, that is, the completeness of 90% is a dominant feature item at this time. Another example is when the feature is the D value (Duplicated BUSCOs value, the proportion of duplicated BUSCOs in all BUSCOs) of a general diploid genome, a lower D value represents better genome quality, that is, a lower D value is a dominant feature item at this time. Specifically, the quality of a genome with a D value of 50% is better than that of a genome with a D value of 80%, that is, the D value of 50% is a dominant feature item at this time.

[0055] In a specific embodiment, the continuity of the sequence obtained by splicing is used as the dominant feature item, and among the initially assembled genomes, the genome with a larger continuity of the sequence obtained by splicing is used as the basic genome.

[0056] In a specific embodiment, the Contig N50 of the sequence obtained by splicing is used as the dominant feature item, and among the initially assembled genomes, the genome with a larger Contig N50 of the sequence obtained by splicing is used as the basic genome.

[0057] In step four, when the second feature of any remaining initial genome other than the basic genome is dominant relative to the second feature of the basic genome, the dominant region of the remaining initial genome is used to replace the corresponding region of the basic genome to obtain an optimized genome. For example, the second feature value represents the completeness of the assembled sequence in the genome. Preferably, when the C value of Busco of other initial genomes is greater than the C value of the basic genome, the dominant region of this initial genome is used to replace the corresponding region of the basic genome.

[0058] The advantageous region refers to the region with advantageous characteristic items. Specifically, it refers to the region with characteristic items representing better sequence quality when using a certain or certain characteristics to describe the quality differences of a certain or certain regions in a series of genomes. Since the alignment position coordinates can be obtained after aligning the genome to be compared with the reference sequence set, the absolute position of all nucleotide sequences in each assembled genome relative to a certain reference genome can be marked, and the regions with the same absolute position in different assembled genomes are called corresponding regions. In some non-advantageous regions, if some characteristics of the genome to be compared in this region are superior to the reference genome or other genomes to be compared, then these non-advantageous regions are advantageous extended regions. Advantageous extended regions can be preferentially searched near the advantageous regions. Advantageous extended regions can include advantageous regions in some cases. Some characteristics of the non-non-advantageous regions can be characteristics representing the continuity and integrity of this region, such as Contig N50 and the C value of Busco, etc.

[0059] Further, replacing the corresponding region in the basic genome with the advantageous region in the remaining initial genomes includes the following steps:

[0060] Expand the advantageous region in the remaining initial genomes,

[0061] Extract the expanded region from the remaining initial genomes and perform sequence alignment with the basic genome to confirm its optimal alignment region in the basic genome,

[0062] And replace the optimal alignment region with the corresponding region in the remaining initial genomes.

[0063] Among them, when expanding the advantageous region, the expansion can be forward expansion, backward expansion, or simultaneous forward and backward expansion of the sequence. In a specific embodiment, the expansion is the forward and backward expansion of the sequence. The expansion length can be adjusted according to the length of the advantageous region, or can be expanded according to some characteristics of the non-advantageous region. For example, the expansion length can be 10bp - 10kbp, preferably 50bp - 5kbp, more preferably 500bp - 1kbp.

[0064] Sequence alignment can be carried out in a manner known in the prior art, for example, it can be through Blat alignment.

[0065] In a specific embodiment, the method for optimizing the assembled genome of the present invention is characterized in that the method includes the following steps:

[0066] Sequencing the sample to obtain a sequencing sequence set;

[0067] Two initial genomes are assembled from the sequencing sequence set in two ways, and the first feature and the second feature corresponding to each initial genome are obtained, where the first feature represents the continuity of the genome and the second feature represents the integrity of the genome;

[0068] Traverse the first feature of each initial genome, and use the initial genome with the dominant first feature as the basic genome;

[0069] When the second feature of any remaining initial genome other than the basic genome is dominant relative to the second feature of the basic genome, use the dominant region of the remaining initial genome to replace the corresponding region of the basic genome to obtain an optimized genome.

[0070] In a specific embodiment, the method for optimizing the assembled genome of the present invention includes the following steps:

[0071] Sequencing and quality control of the sample are performed to obtain a sequencing sequence set;

[0072] Two initial genomes are assembled from the sequencing sequence set in two ways, and the Contig N50 and the C value of Busco corresponding to each initial genome are obtained;

[0073] Traverse the Contig N50 of each initial genome, and use the initial genome with the larger Contig N50 as the basic genome;

[0074] When the C value of Busco of another initial genome is larger than the C value of Busco of the basic genome, expand the dominant region in the other initial genome,

[0075] Extract the expanded region from the other initial genome, and perform sequence alignment with the basic genome through Blat alignment to confirm the optimal alignment region in the basic genome,

[0076] And replace the optimal alignment region with the corresponding region in the remaining initial genome.

[0077] Example

[0078] The processes of the following examples are as Figure 1As shown, first, a library is constructed for the sample and sequenced. After data quality control and filtering, a sequencing sequence set is obtained. Then, the sequencing sequence set is assembled using Gene Assembly Version 1 and Gene Assembly Version 2 to obtain two initial genomes, each of which has a Contig N50 and a C value of BUSCO. One initial genome has a higher Contig N50 and a lower C value of BUSCO, while the other initial genome has a lower Contig N50 and a higher C value of BUSCO. Based on the genome with the higher Contig N50 among the two initial genomes as the reference genome, the dominant regions in the other initial genome are extended, the extended regions are extracted from the other initial genome, and sequence alignment is performed with the reference genome to confirm the optimal alignment regions in the reference genome, and the optimal alignment regions are replaced with the corresponding genomes in the remaining initial genome.

[0079] 1. Third-generation sequencing was performed on a certain marine organism using the Pacbio platform with a sequencing data volume of 327G subreads. The sequencing data was quality-controlled using the quality control software SMRT Link.

[0080] 2. Canu was used for long-read error correction, and the corrected reads were assembled using CANU (Software 1) and WTDBG2 (Software 2) respectively. The obtained genome assembly results were named Genome A and Genome B respectively.

[0081] 3. The assembly results were evaluated for N50 and BUSCO, and the results are shown in Table 1:

[0082] Table 1

[0083]

[0084] 4. The estimated genome size of the sample is 1.1G. Since Genome B is closer to the true size, has a longer N50 evaluation, and a lower BUSCO value, Genome B was used as the reference genome, and the nucleotide sequences of the dominant regions with a higher BUSCO value in Genome A were tried to replace the corresponding regions in the reference genome B to improve the integrity of Genome B.

[0085] 5. The software blat was used to align the genome to be replaced with the reference genome to obtain the dominant regions of Genome A, that is, the high N50 regions. The dominant regions were extended by 1 kbp in both the forward and reverse directions of the sequence to obtain the dominant extended regions (including the aforementioned dominant regions). According to the scheme in step 4, the nucleotide sequences of the corresponding regions of Genome B were replaced with the dominant extended regions of Genome A (the replacement process was written in Python) to obtain the improved Genome B (renamed as Genome B'), and the results are shown in Table 2:

[0086] Table 2

[0087]

[0088] As can be seen from Table 2, after the BUSCO optimization of genome B by genome A using the method of the embodiment, genome B’ is obtained, with a size of 1,125,786,841 bp and a BUSCO C value representing the completeness of 92.1%. Compared with the initial genome B, the genome size and Contig N50 are close. This indicates that the BUSCO C value reflecting the genome integrity has increased by 1.4%, and the proportion of Duplicated BUSCOs has not increased significantly. That is, by the method of this embodiment, without affecting (such as reducing) the continuity of the initial genome, an optimized genome with a better completeness than the initial assembled genome is obtained. That is, the method of this embodiment can effectively improve the quality of the assembled genome.

Claims

1. A method for optimizing an assembled genome, characterized in that, the method comprises the following steps: Sequencing a sample to obtain a sequencing sequence set; Assembling the sequencing sequence set in two or more ways to obtain two or more initial genomes, and obtaining a first feature and a second feature corresponding to each initial genome; Traversing the first feature of each initial genome, and taking the initial genome with the dominant first feature as the basic genome; When the second feature of any remaining initial genome other than the basic genome is dominant relative to the second feature of the basic genome, using the dominant region of the remaining initial genome to replace the corresponding region of the basic genome to obtain an optimized genome; The first feature is the length of the sequence obtained by splicing in the genome, the second feature represents the integrity of the assembled sequence in the genome, and the dominant region is the region with the dominant feature item; Using the dominant region in the remaining initial genome to replace the corresponding region in the basic genome includes: Expanding the dominant region in the remaining initial genome, Extracting the expanded region from the remaining initial genome, and performing sequence alignment with the basic genome to confirm its dominant expanded region in the basic genome, And replacing the dominant expanded region with the corresponding region in the remaining initial genome; The first feature is Contig N50, and the second feature is the C value of Busco.

2. The method according to claim 1, characterized in that, the sequence set is a sequence set obtained through quality control.

3. The method according to claim 1, characterized in that, the length of the expansion is 10bp - 10kbp.

4. The method according to claim 3, characterized in that, the length of the expansion is 50bp - 5kbp.

5. The method according to claim 3, characterized in that, the length of the expansion is 500bp - 1kbp.

6. The method according to claim 1, characterized in that, the sequence alignment is performed by Blat alignment.

7. The method according to any one of claims 1 - 3, characterized in that, the source of the sample is an animal, a plant or a microorganism.

8. The method according to any one of claims 1 - 3, characterized in that, the sequence set includes a collection of base sequence information and other sequence information.

9. The method according to any one of claims 1 - 3, characterized in that, the sequencing is third-generation sequencing.

Citation Information

Patent Citations

  • Reference genome and de novo assembly combination based next-generation sequencing data assembly method

    CN105303068A