A method for assembling a local goose t2t genome
By combining third-generation and second-generation sequencing technologies, assembling the goose genome using Hifiasm and NextDenovo software, and filling gaps using quarTeT and RepeatMasker software to identify telomeres and centromeres, the problem of missing regions in the goose genome was solved, achieving gap-free genome assembly and providing more comprehensive genetic research resources.
Patent Information
- Application Number
- CN202311854898.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2043-12-29
AI Technical Summary
The existing goose genome contains a large number of deleted regions, especially highly repetitive regions such as centromeres and telomeres, which affect the interpretation of genetic information and the integrity of the genome structure.
A combination of third-generation long-read sequencing and second-generation sequencing was used. Using Hi-C data, the genome was assembled using Hifiasm and NextDenovo software, and gaps were filled using quarTeT software. Gene structure prediction and repetitive sequence annotation were performed using BUSCO and RepeatMasker software. Interspecies genome collinearity alignment was performed using NGenmoesyn software to identify telomeres and centromeres.
The uninterrupted genome of the domestic goose was successfully assembled, filling the gaps in the autosomes and providing more comprehensive genetic research resources. A large number of genes and mRNAs were identified, providing important resources for studying the biological characteristics and growth and development of geese.
Smart Images

Figure CN117802091B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of molecular markers, and particularly relates to a T2T genome assembly method of local geese. BACKGROUND
[0002] Domestic goose (Anser cygnoides domesticus) is an important agricultural poultry, and its meat, egg and ornamental multiple purposes make it a widely bred species. About 6000 years ago, geese were domesticated together with chickens and ducks, becoming one of the earliest domesticated poultry. Geese have the characteristics of rapid growth, strong disease resistance and highly developed liver fat storage, and are suitable for roughage feeding environment. Compared with other terrestrial poultry (such as chickens), geese have unique biological characteristics, for example, low susceptibility to some avian viruses, although they may exist as virus carriers, but rarely show symptoms of infection, thus becoming a natural reservoir of avian viruses. In addition, the high fat accumulation ability of goose liver and the characteristics of not easy to occur liver fibrosis or necrosis suggest that it has unique lipid storage and metabolism characteristics, which provides an important reference for the study of human lipid metabolism disorders. With the development of genomics research, the assembly and analysis of the goose genome has become an important way to understand its genetic characteristics and biological functions. In the past few years, the progress of goose genome sequencing has enabled us to explore its genome structure and function more comprehensively.
[0003] Lu et al. (2015) first sequenced and analyzed the goose genome in 2015, using second-generation sequencing data and the SOAPdenovo software (Li R et al,. 2010) for assembly, resulting in a 1.12 Gb goose genome draft. The genome draft contained 1,049 Scaffold sequences, with a Scaffold N50 of 5.2 Mb. Subsequently, Gao et al. (2016) published the genome sequence map of a female Sichuan white goose. By resequencing the ancestor of the domestic goose, the whooper swan (Anser cygnoides), it was found that the two species diverged 3.4-6.3 million years ago (Gao et al,. 2016). In addition, in 2020, the chromosome-level genome of the goose was also released. Researchers published a 1.11 Gb Tianfu goose genome, with Contig N50 and Scaffold N50 values of 1.85 Mb and 33.12 Mb, respectively. The genome assembly contained 39 pseudo-chromosomes (2n = 78), accounting for about 88.36% of the size of the whole goose genome (Li et al., 2020). In the past two years, high-quality chromosome-level reference genomes of Xingguo grey goose (Ouyang et al., 2022), lion-headed goose (Zhao et al., 2023), and others have been released, providing valuable genetic resources and data foundations for promoting the breeding and biological research of geese.
[0004] However, due to limitations of past technologies, the existing goose genome still contains numerous missing regions, primarily involving centromeres, telomeres, and other highly repetitive regions containing much important genetic information. Telomeres are highly repetitive DNA sequences at the ends of chromosomes, protecting them from degradation (Shay et al., 2019). Centromeres are another unique chromosomal domain, serving as the assembly site of the centromere during chromosome separation (Wu et al., 2011). Centromere DNA sequences are typically composed of satellite DNA and represent the most rapidly evolving sequences in eukaryotic genomes (Francesca., 2022). With the development of sequencing technologies, ultralong Oxford nanopore (ONT) technology and high coverage depth (HiFi) data from PacBio have been widely used to fill gaps in plant and animal genomes. By integrating third-generation DNA sequencing technology and second-generation Hi-C data, a complete, gap-free assembly of the goose genome can be achieved. Recently, the T2T genome of domestic chickens was published, filling most of the gaps in previous genome studies and revealing the structural features of chicken telomeres and centromeres (Huang et al., 2023). However, there has been no report on the completion of a gap-free reference genome for geese. In this study, we assembled a gap-free domestic goose genome for the first time, employing multiple assembly strategies and utilizing high-coverage and accurate long-read sequence data. This assembly reveals for the first time the structural features of highly repetitive regions (centromeres and telomeres) in geese, providing a foundation for better analysis of the structural features and functions of the goose genome. Summary of the Invention
[0005] This invention aims to provide a method for assembling the T2T genome of local geese. The assembled T2T chromosome-level genome of geese lays an important research foundation for future genetic improvement and genetic mechanism analysis of geese.
[0006] This invention provides a method for assembling the T2T genome of a local goose, comprising the following steps:
[0007] Step 1: Sample collection and sequencing
[0008] (1) Collect blood samples from the wing vein, pectoral muscle, and six organ tissues from an adult female Taihu goose in the Taihu goose conservation population. Then extract DNA and RNA from the samples.
[0009] (2) DNA library construction and sequencing: The blood sample extracted in step (1) is used to obtain a complete genome fragment by combining third-generation long-read sequencing and second-generation sequencing.
[0010] (3) Hi-C sequencing library construction and sequencing: cross-linking reaction of the pectoral muscle tissue in step (1) in formaldehyde solution for Hi-C library construction and sequencing.
[0011] (4) RNA library construction and sequencing, six kinds of tissues in step (1) are subjected to second-generation transcriptome sequencing. In order to improve the accuracy of gene annotation, the six kinds of tissues are mixed in equal amounts and subjected to third-generation full-length transcriptome sequencing.
[0012] Step 2: Construction of genomic sequence map
[0013] (1) The K-mer method is used to estimate the genome size of Taihu goose based on second-generation short sequencing data.
[0014] (2) The genome assembly is performed by combining Hifiasm (v 0.18.5) and NextDenovo (v2.4.0) software.
[0015] (3) The quarTeT software is used to fill the gaps in the assembled scaffold sequence.
[0016] (4) BUSCO (v 5.4.5) is used to call metaeuk (v 6.a5d39d9) software for gene structure prediction, and HMMER (v3.3.2) is used to align the predicted gene sequence with the eukaryotic bird reference dataset. By analyzing the alignment degree and coverage of the predicted gene sequence with the reference sequence, the integrity of the Taihu goose genome assembly is evaluated, i.e. whether the genome contains these conserved gene sequences.
[0017] (5) RepeatMasker software (v 4.1.5) is used to annotate the repetitive sequences of the goose genome.
[0018] (6) In the process of identifying telomeres and centromeres in the goose genome, the animal "TTAGGG" is used as the telomere recognition sequence of the goose, and the TeloExplorer function of the quarTeT software (v 1.1.3) is used for telomere identification.
[0019] (7) In order to study the similarity of goose and duck, chicken in poultry at the karyotype level, NGenmoesyn software (v1.39) is used to perform collinear alignment of the assembled goose chromosome genome data with duck and chicken chromosome genomes.
[0020] Preferably, the six organ tissue samples in step 1 include brain, heart, liver, spleen, lung, and kidney.
[0021] Preferably, the sample DNA in step 1 is extracted by the root blood / cell / tissue genomic DNA extraction kit (TIANGEN® DP304). The total RNA extraction process of the sample tissue is strictly in accordance with the instructions of the TRNzol Universal Total RNA Extraction Kit (TIANGEN® DP424).
[0022] Preferably, the DNA library construction and sequencing in step 1 include the use of third-generation ultra-long sequencing, HiFi sequencing, and second-generation short-fragment sequencing library construction.
[0023] Preferably, the Hi-C sequencing library construction and sequencing process in step 1 includes resuspending the pellets and resuspending the cells using NEB buffer. Then, the cell nuclei are lysed with dilute SDS lysis solution, the DNA is digested with the four-base enzyme MboI, and the DNA ends are labeled with biotin-14-dctp. After labeling is complete, T4 DNA polymerase is used to remove biotin. Then, T4 DNA ligase is used for ligation. Finally, after DNA purification, double-end 150bp sequencing is performed on the Illumina Hiseq platform.
[0024] Preferably, the RNA library construction and sequencing in step 1 uses the EasyPure RNA Kit (Transgen) to isolate total RNA from organ tissues. Then, the NEBNext® UltraTM RNA Library Prep Kit for Illumina® (NEB, Ipswich, MA, USA) is used to prepare the sequencing library for the sample RNA. Finally, double-end (2x125bp) sequencing is performed on the Illumina HiSeq Xten platform. For the mixed sample full-length transcript library construction and sequencing, the Pacbio Sequel system (Pacific Biosciences, CA, USA) is used for full-length transcript sequencing.
[0025] Preferably, the genome size evaluation in step 2 is statistically analyzed by double-end sequencing library data, and the K-mer distribution is obtained using the Jellyfish tool. Then, GenomeScope (v 2.0) is used to model based on the K-mer distribution, thereby preliminarily revealing the characteristics of the Taihu goose genome.
[0026] Preferably, the genome assembly in step 2 is first assembled using Hifi data, Hifi+Hi-C data, and Hifi+ONT ultra-long read+Hi-C data, respectively. In addition, the assembly is performed using Hifi+ONT ultra-long read+Hi-C data using NextDenovo. To further improve the assembly quality, the run_purge_dups.py (v 1.2.4) tool is used to remove duplicate contigs. Finally, according to the evaluation of N50 value, the assembly result of Hifi+Ont+Hi-C is selected as the data for subsequent analysis. Considering the problem of low accuracy of ONT third-generation ultra-long sequencing, the Hifi data is used to correct the ONT data. Preferably, the gap filling in step 2 uses the following parameters: "-GapFiller -g *fasta -t 30 -l 5000 -i 60", and refers to the genome data assembled by multiple methods.
[0027] The present application has the following beneficial effects:
[0028] (1) The blank areas on most chromosomes in the existing goose reference genome are filled, and 33 autosomes reach the level of complete non-interstitial, providing a more comprehensive genome reference for genetic research of geese.
[0029] (2) A high-quality goose genome sequence is successfully assembled, including autosomes and sex chromosomes, providing an important basis for studying the sex determination and reproductive mechanisms of geese.
[0030] (3) Through genome annotation, a large number of genes and mRNAs are identified, providing important resources for studying the biological characteristics, growth and development process, and disease resistance of geese. BRIEF DESCRIPTION OF DRAWINGS
[0031] Figure 1 Sample photos and genome quality evaluation charts for Example 1: (A) Morphological photo of Taihu goose; (B) Genome size estimation GenomeScope2; (C) Taihu goose whole genome Hi-C heat map;
[0032] Figure 2 Circos plot of Taihu goose 41-chromosome genome assembly for Example 1: The rings from outside to inside represent (a) Taihu goose genome chromosomes, (b) GC density, (c) Exon density, (d) CDS density, (e) IncRNA density, (f) mRNA density, (g) Gene density, b-g is 100kb; The innermost circle is a collinear plot of homologous genes on different chromosomes;
[0033] Figure 3For the distribution map of centromere, telomere and gap of the assembled goose genome of Example 1, the chromosome heat map represents the gene density, and the wavy line represents the density of the repetitive region;
[0034] Figure 4 For the whole genome alignment of the goose chromosomes of Example 1 with duck and chicken genomes. DETAILED DESCRIPTION
[0035] The preferred embodiments of the present application will be described in detail below with reference to the drawings, in which the drawings constitute a part of this application and illustrate embodiments of the present application together with the principles of the present application, but are not intended to limit the scope of the present application.
[0036] A local goose T2T genome assembly method, comprising the following steps:
[0037] Step 1: sample collection and sequencing
[0038] An adult female Taihu goose in the Taihu goose conservation population was collected, and the wing vein blood, pectoral muscle, and brain, heart, liver, spleen, lung, and kidney tissue samples were collected. Then, the sample DNA and RNA were extracted. The sample DNA was extracted by the blood / cell / tissue genomic DNA extraction kit (DP304). The total RNA extraction process of the sample tissue was strictly performed according to the instructions of the TRNzol Universal total RNA extraction kit (DP424).
[0039] DNA library construction and sequencing adopt three generations of ultra-long sequencing, HiFi sequencing and two generations of short sequencing gene library construction. The extracted blood sample is used to obtain the complete genomic fragment by combining three generations of long read sequencing and two generations of sequencing.
[0040] The pectoral muscle tissue is cross-linked in a formaldehyde solution for Hi-C library sequencing.
[0041] The Hi-C sequencing library construction and sequencing process includes resuspending the ball in the lysis solution and resuspending the cells using the NEB buffer. Then, the cell nucleus is dissolved using the dilute SDS lysis solution, the DNA is cut using the four-base enzyme MboI, and the DNA ends are labeled using biotin-14-dctp. After the labeling is completed, the T4 DNA polymerase is used to remove the biotin. Then, the T4 DNA ligase is used for ligation. Finally, after the DNA purification treatment, the double-end 150bp sequencing is performed on the Illumina Hiseq platform.
[0042] To improve the accuracy of gene annotation, the six tissues were mixed in equal amounts and subjected to third-generation full-length transcriptome sequencing. Total RNA was isolated from the organ tissues using EasyPure RNA Kit (Transgen). Subsequently, the sample RNA was subjected to sequencing library preparation using NEBNext® UltraTM RNA Library Prep Kit for Illumina® (NEB, Ipswich, MA, USA). Finally, double-end (2x125bp) sequencing was performed on the Illumina HiSeq Xten platform. For full-length transcript library construction and sequencing of the mixed sample, the Pacbio Sequel system (Pacific Biosciences, CA, USA) was used for full-length transcript sequencing.
[0043] Step 2: Genome sequence mapping
[0044] The K-mer method was used to estimate the genome size of the Taihu goose based on second-generation short sequencing data. Statistical analysis was performed on the paired-end sequencing library data, and the distribution of K-mers was obtained using the Jellyfish tool. Subsequently, GenomeScope (v 2.0) was used to model based on K-mer distribution, thereby preliminarily revealing the characteristics of the Taihu goose genome.
[0045] Genome assembly was performed by combining Hifiasm (v 0.18.5) and NextDenovo (v2.4.0) software. First, genome assembly was performed using Hifi data, Hifi+Hi-C data, and Hifi+ONT ultra-long read+Hi-C data, respectively. In addition, Hifi+ONT ultra-long read+Hi-C data was used for assembly using NextDenovo. To further improve assembly quality, the run_purge_dups.py (v 1.2.4) tool was used to remove duplicate contigs. Finally, based on the evaluation of N50 values, the assembly results of Hifi+Ont+Hi-C were selected as the data for subsequent analysis. Considering the problem of low accuracy of ONT third-generation ultra-long sequencing, Hifi data was used to correct the ONT data.
[0046] The assembled scaffold sequence was subjected to gap filling using the quarTeT software. In the filling process, the following parameters were used: "-GapFiller -g *fasta -t 30 -l 5000 -i 60", and reference was made to the genome data assembled using multiple methods.
[0047] BUSCO (v 5.4.5) was used to call the metaeuk (v 6.a5d39d9) software for gene structure prediction, and HMMER (v3.3.2) was used to align the predicted gene sequences with the eukaryotic bird reference dataset. By analyzing the alignment degree and coverage of the predicted gene sequences with the reference sequences, the integrity of the Taihu goose genome assembly was evaluated, i.e., whether the genome contains these conserved gene sequences.
[0048] RepeatMasker software (v 4.1.5) was used to annotate the repetitive sequences of the goose genome.
[0049] In the process of identifying telomeres and centromeres in the goose genome, the animal "TTAGGG" was used as the telomere recognition sequence of the goose, and the TeloExplorer function of the quarTeT software (v 1.1.3) was used for telomere identification.
[0050] In order to study the similarity of goose, duck and chicken at the karyotype level in poultry, NGenmoesyn software (v1.39) was used to perform collinear alignment of the assembled goose chromosome genome data with duck and chicken chromosome genomes.
[0051] Example 1 A local goose T2T genome assembly method
[0052] 1. Sample collection and sequencing
[0053] 1.1 Collection and extraction of sample DNA and RNA
[0054] The research sample was selected from an adult female Taihu goose (Anser cygnoides) in the Taihu goose conservation population of the National Waterfowl Gene Library (Jiangsu). Figure 1A). The strategy adopted for goose genome assembly is shown in Fig. IB. Before slaughter, we used 5 ml anticoagulant blood collection tubes (BD Vacutainer ® EDTA) to extract sample blood from the wing vein, and then extracted DNA for subsequent sequencing analysis. In order to obtain complete fragments of genomic fragments, we used a combination of third-generation long-read sequencing and second-generation sequencing techniques. In addition, the sample pectoral muscle tissue was cut into small pieces and placed in a formaldehyde solution for cross-linking reaction for Hi-C library sequencing. At the same time, six kinds of tissue samples of brain, heart, liver, spleen, lung and kidney were collected and cut into small pieces, and then placed in 1.8 ml cryotubes (Nunc CryoTube) and quickly frozen in liquid nitrogen tank, and stored in -80 °C ultra-low temperature refrigerator (Hair DW-86L728J) for second-generation transcriptome sequencing. In addition, in order to improve the accuracy of gene annotation, the six kinds of tissue samples collected were mixed in equal amounts for third-generation full-length transcriptome sequencing. All the above sampling experimental operations conform to the rules and regulations of the Animal Welfare Committee of Jiangsu Animal Husbandry Vocational College (Animal Ethics Approval No. 22110313195050999).
[0055] The sample DNA extraction process strictly followed the operation instructions of the TIANGEN® DP304 blood / cell / tissue genomic DNA extraction kit. After DNA extraction, the quality of the DNA was detected using a Nanodrop 2000 spectrophotometer. The sample DNA quality qualified parameter settings were: OD value (260 / 280) between 1.8-2.0, and concentration greater than 100 ng / μl. Finally, the prepared 2% agarose gel was used for electrophoresis, and the sample DNA that passed the DNA band detection was stored in a -80 °C refrigerator (Hair DW-86L728J). The total RNA extraction process of the sample tissue was strictly operated according to the use instructions of the TIANGEN® DP424 TRNzol Universal total RNA extraction kit. After RNA extraction, the concentration and purity of the RNA were determined. After detection, the sample RNA was stored in a -80 °C refrigerator (Hair DW-86L728J).
[0056] 1.2 DNA library construction and sequencing
[0057] The sample genome third-generation ultra-long sequencing followed the standard protocol provided by Oxford Nanopore Technologies (ONT) company. First, the genomic DNA was randomly cut using Megaruptor (Diagenode, USA). Then, the Nanopore SQK-LSK 109 (Oxford Nanopore technologies, USA) kit was used for adapter preparation and ligation, and the ligated DNA library was detected again by Qubit 3.0 Fluorometer. Finally, the sample was loaded onto the Nanopore Flow cells R9.4 for sequencing on the PromethION platform. The final sequencing result statistics are shown in Table 1, a total of 577,228 reads were obtained, the total base number reached 52,490,712,237 bp, the average length of reads was 90,935.9 bp, the N50 length was 100,823 bp, and the GC content was 42.82%.
[0058] The sample HiFi sequencing used the PacBio single molecule real-time circular consistent sequencing (CCS) library preparation method. First, a total of 100 μg of high-quality genomic DNA was cut using Covaris g-TUBEs (Covaris) to obtain fragments with a target size of about 20 kb. Then, the size distribution of the cut genomic DNA was detected using an Agilent 2100 Bioanalyzer DNA 12000 chip (Agilent Technologies) to ensure that it met the requirements. Next, the PacBio DNA template preparation kit 2.0 (Pacific Biosciences of California, Inc., CA) was used to construct the sequencing library for HiFi sequencing on the PacBio RS II machine (Pacific Bioscences of California, Inc.). Finally, the constructed library was loaded onto a SMRT CELL for sequencing. A total of 4,261,430 reads of sequencing data were obtained (Table 1), the total base number reached 71,413,769,333 bp, the average length of reads was 16,758 bp, the N50 length was 16,838 bp, and the GC content was 42.61%.
[0059] The process of constructing the second-generation short-fragment sequencing genomic library of the sample is as follows: first, the high-quality genomic DNA is randomly cut using a Covaris ultrasonic instrument (Covaris, USA). Then, the Truseq nano DNA HT library preparation kit (Illumina, USA) is used to construct the Illumina sequencing library, and the target insertion size is 350 bp. Finally, the purified library is loaded onto the Illumina NovaSeq 6000 platform for sequencing. After sequencing is completed, a total of 385,826,042 sequences are obtained, with a total of 57,873,906,300 bp of sequencing data, and a GC content of 43.51%.
[0060]
[0061] 1.3 Hi-C sequencing library construction and sequencing
[0062] The construction and sequencing of the Hi-C sequencing library of the sample are based on the standard process with some modifications. First, the pectoral muscle tissue is cross-linked at room temperature using a 4% formaldehyde solution. Then, 20 μl of lysis buffer is used to resuspend the pellets, and 100 μl of NEB buffer is used to resuspend the nuclei. Next, the nuclei are dissolved using a dilute SDS lysis solution. Then, the DNA is cut using the four-base enzyme Mbol, and the DNA ends are labeled with biotin-14-dctp, and after labeling is completed, T4 DNA polymerase is used to remove the biotin. Then, T4 DNA ligase is used for ligation. Finally, after DNA purification, double-end 150 bp sequencing is performed on the Illumina Hiseq platform. The sequencing results are shown in Table 1: a total of 1,075,285,592 reads are obtained, with a total base data amount of 161,292,838,800 bp, an average read length of 90,935.80 bp, an N50 length of 100,823 bp, and an average GC content of 42.82%.
[0063] 1.4 RNA library construction and sequencing
[0064] For RNA sequencing library construction and sequencing of 6 samples, first, total RNA was isolated from brain, heart, liver, spleen, lung, pectoral muscle tissues, respectively, using EasyPure RNA Kit (Transgen). Subsequently, NEBNext® UltraTM RNA Library Prep Kit for Illumina® (NEB, Ipswich, MA, USA) was used for sequencing library preparation of sample RNA. Finally, double-end (2x125bp) sequencing was performed on the Illumina HiSeq Xten platform. The specific sequencing results are shown in Table 1. Among them, the heart tissue sequencing obtained the highest number of reads, reaching 45,882,692, while the spleen tissue obtained the lowest number of reads, 38,462,044. The average total reads data of the six tissues was 6,393,354,550bp, and the GC content was 46.04%.
[0065] For full-length transcript library construction and sequencing of mixed samples, Pacbio Sequel system (Pacific Biosciences, CA, USA) was used for full-length transcript sequencing. According to the Isoform Sequencing (Iso-Seq) protocol, first, NEBNext Single Cell / Low Input cDNA Synthesis & Amplification Module was used for cDNA synthesis and amplification of the sample. Then, PacBio SMRTbell Express Template Prep Kit 2.0 was used to process the sample, including adapter ligation and SMRTbell sequence addition. Next, size-selective purification was performed by ProNex® Size-Selective Purification System to remove low-quality and short-fragment sequences to complete Iso-Seq library preparation. Finally, full-length transcript sequencing was performed on the Sequel Sequel System (Pacific Biosciences) to obtain high-quality full-length transcript sequence information. A total of 48,373,842 reads were obtained, with a total data amount of 84,725,302,734bp, an average read length of 1,751.50bp, an N50 length of 2,447bp, and a GC content of 46.05%.
[0066] 2. Genome sequence map construction
[0067] 2.1 Genome size evaluation
[0068] In this study, K-mer method was used to estimate the genome size of Taihu goose based on the short reads sequencing data. By statistical analysis of the paired-end sequencing library data, the distribution of K-mer was obtained using Jellyfish tool. Subsequently, GenomeScope (v 2.0) was used to model according to the distribution of K-mer, thus the characteristics of Taihu goose genome were preliminarily revealed. In Figure 1 In Fig. 1C, the blue line represents the actual observed distribution of K-mer in the sequencing sequence of Taihu goose genome. At the same time, the brown line represents the K-mer in the sequence caused by sequencing errors. Since sequencing errors are random, these K-mers usually have a low frequency. Finally, GenomeScope models according to this information and estimates that the length of Taihu goose genome is about 1.12 Gb, and the genome heterozygosity is about 0.5%. Based on the de novo assembly results of the genome, Taihu goose belongs to high heterozygosity genome.
[0069] 2.2 Genome assembly
[0070] The present study combined Hifiasm (v 0.18.5) and NextDenovo (v2.4.0) software for assembly. The Hifiasm software was used for genome assembly. First, the genome was assembled using Hifi data, Hifi+Hi-C data, and Hifi+ONT ultra-long read+Hi-C data, respectively. In addition, Hifi+ONT ultra-long read+Hi-C data was used for assembly using NextDenovo. The assembly results are shown in Table 2, in which the assembly effect of NextDenovo is the best, with 244 contigs, N50 length of 33,928,929 bp, and selected for downstream analysis. In order to further improve the assembly quality, the run_purge_dups.py (v 1.2.4) tool was used to remove duplicate contigs. Finally, according to the evaluation of N50 value, the assembly results of Hifi+Ont+Hi-C were selected as the data for subsequent analysis. Considering the problem of low accuracy of ONT third-generation ultra-long sequencing, the Hifi data was used to correct the ONT data. The specific operation includes using meryl software (v 1.4) to count the number of kmer occurrences, using winnowmap software (v 2.03) to re-align the assembled genome with Hifi data, and then using falconc software (v 1.15.0) for secondary filtering and deletion of chimeric alignment fragments. Finally, using racon software (v1.5.0) for three rounds of error correction, the genome assembly sequence after HiFi error correction was obtained. Next, using Chromap software (v 0.2.5) and yahs (v 1.2a.1) software suite, combined with Hi-C data, the genome was assembled with high quality to obtain complete scaffold sequence. In order to identify and align the assembled scaffold sequence, it was aligned and analyzed with the known lion-headed goose genome (GCA_025388735.1). Through alignment, the correspondence of the scaffold sequence with each chromosome in the lion-headed goose genome was determined, and was renamed according to the matched autosomes 1-38 and Z chromosome.
[0071] To obtain more complete sequence information of goose W chromosome, we performed W chromosome assisted assembly. In the published version of goose genome, the sequence information of W chromosome was missing. Therefore, we used the genome of duck, a close relative of goose, as reference to assemble the W chromosome by using the "scaffold" module of ragtag.py software (v2.1.0). By this strategy, we successfully assembled a W chromosome with a length of 17.35 Mb. The W chromosome was composed of 18 scaffolds, among which scaffold_42 was the main part of W chromosome, accounting for 9.63% of the total length. Finally, we successfully assembled 38 autosomes and two sex chromosomes, W and Z, which is the most complete goose genome so far Figure 2 It is worth emphasizing that due to the complexity of sex chromosome structure, the assembly of sex chromosomes is much more difficult than that of autosomes. Therefore, we used the assisted assembly method and the relevant information of duck W chromosome genome to obtain a relatively complete sequence of goose W chromosome. This work has important academic value for further study of the sex determination mechanism and genetic characteristics of goose.
[0072]
[0073] 2.3 Gap filling
[0074] Gap filling was performed by using the quarTeT software (v 1.1.3). In the filling process, the following parameters were used: "-GapFiller -g *fasta -t 30 -l 5000 -i 60", and the data of the genome assembled by multiple methods were referred to. This tool uses tetrad alignment information to fill gaps and uses other related known genome information to improve the accuracy of gap filling. After gap filling, we successfully closed 33 autosomes completely, except for a small number of gaps on two sex chromosomes. Figure 2 The distribution of gaps on each chromosome is shown.
[0075] 2.4 Assessment of genome integrity
[0076] BUSCO (v 5.4.5) (Seppey et al., 2019) was used to call the metaeuk (v 6.a5d39d9) software for gene structure prediction, and the predicted gene sequences were aligned with the eukaryotic bird reference dataset using HMMER (v3.3.2). By analyzing the alignment degree and coverage of the predicted gene sequences with the reference sequences, the integrity of the Taihu goose genome assembly was evaluated, i.e., whether the genome contains these conserved gene sequences. According to the statistics of the alignment results, it was determined that there were single copy genes (S) and multiple copy genes (D) in the assembled genome. Among them, 96.5% of the single copy genes could be completely aligned to the genome, and 0.4% of the multiple copy genes were completely present in the genome. In addition, we also used Quast (v 5.2.0) software to evaluate the key indicators of the genome. The results showed that the Taihu goose genome size was 1,197,991,206 bp, and the scaffold N50 reached 81,007,908 bp. Compared with the published chromosome-level goose genome, the number of scaffolds in our assembly result was significantly the least, only 73. Notably, the scaffold N50 length of this assembly exceeded 80M, which was significantly better than the previous genome version. The detailed comparison results are shown in Table 3.
[0077]
[0078] 2.5 Gene annotation
[0079] The RepeatMasker software (v 4.1.5) was used to annotate the repetitive sequences of the goose genome. According to the results (see Table 4), among the annotated repetitive sequences, the interspersed repetitive sequences accounted for 8.92% of the entire goose genome length, with a total length of about 106.89 Mb. Among them, about 77.17 Mb (6.44%) were retrotransposons, and 3.66 Mb were DNA transposons. In addition, about 4.87% of the sequences on the Taihu goose genome belonged to long interspersed nuclear elements (LINEs), which were the largest proportion of repetitive sequence species in the genome. It is worth noting that among them, the abundance of chicken repeat 1 (CR1) was the highest, accounting for almost 100% of all LINEs. In addition, 1.49% of the Taihu goose genome sequences belonged to long terminal repeats (LTRs), and 0.08% belonged to small interspersed nuclear elements (SINEs). After masking the repetitive sequences, we used the Liftoff software (v 1.6.3) to annotate the coding genes and mRNAs of the Taihu goose genome, with reference to the NCBI goose genome (GCF_002166845.1) and its annotation information as well as the transcriptome dataset. The annotation results showed that a total of 34898 genes and 62248 mRNAs were annotated.
[0080]
[0081] 2.6 Telomere and centromere identification
[0082] In the process of identifying telomeres and centromeres in the goose genome, we used the animal "TTAGGG" as the telomere recognition sequence of the goose, and used the TeloExplorer function of the quarTeT software (v 1.1.3) to identify the telomeres. The results showed that there were the most telomere repeats in the 10000bp window at the ends of chromosome 3, with 1101 and 1793 respectively. The specific telomere distribution diagram can be seen in Figure 3For centromere identification, we used the Centromics software (https: / / github.com / ShuaiNIEgithub / Centromics) and utilized the ONT, HiFI, and Hi-C datasets to identify centromeres on the assembled genome. The location of centromeres on the chromosome was determined based on the peak values of the Hic and TR-CL2 (length sequencing captures chromosome conformation fixation) data. The centromere locations are marked on the chromosome schematic diagram. Figure 3 ).
[0083] 2.7 Interspecific genome collinearity
[0084] To investigate the karyotype similarity between geese, ducks, and chickens, we used NGenmoesyn software (v1.39) to perform collinearity alignment of the assembled goose chromosome genome data with that of ducks and chickens. For example... Figure 4 As shown, most long chromosome segments of ducks (chromosomes 1-9) have corresponding chromosomes in the goose genome, especially showing high similarity on chromosomes Z and W. This aligns with the similar lifestyles and taxonomic classifications of ducks and geese as waterfowl. However, compared to geese, the chicken genome shows only partial agreement with the goose genome in linear alignment results. Although both chickens and geese belong to the poultry group, they differ significantly in their lifestyles and evolutionary relationships. This suggests a closer kinship between ducks and geese.
[0085] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for assembling a local goose T2T genome, comprising the following steps: Step 1: Sample collection and sequencing: (1) Collect adult female Taihu goose in Taihu goose breeding population, collect wing vein blood, pectoral muscle and six kinds of organ tissue samples, and then extract DNA and RNA samples; (2) DNA library construction and sequencing, the blood sample extracted in step (1) is used to obtain the complete fragment of genome by combining the third generation long read sequencing and the second generation sequencing; (3) Hi-C sequencing library construction and sequencing, the pectoral muscle tissue in step (1) is cross-linked in formaldehyde solution for Hi-C library sequencing; (4) RNA library construction and sequencing, the six kinds of organ tissues in step (1) are subjected to second generation transcriptome sequencing, in order to improve the accuracy of gene annotation, the six kinds of organ tissues are mixed in equal amount, and third generation full-length transcriptome sequencing is carried out; Step 2: Genome sequence map construction: (1) The K-mer method is used to evaluate the genome size of Taihu goose based on the second generation short fragment sequencing data; (2) The genome is assembled by combining Hifiasm and NextDenovo software; (3) The quarTeT software is used to fill the gaps of the assembled scaffold sequence; (4) The BUSCO calling metaeuk software is used for gene structure prediction, and the predicted gene sequence is aligned with the eukaryotic bird reference data set by using HMMER; by analyzing the alignment degree and coverage information of the predicted gene sequence and the reference sequence, the integrity of the Taihu goose genome assembly is evaluated, that is, whether the genome contains these conserved gene sequences; (5) The RepeatMasker software is used to annotate the repeated sequences of goose genome; (6) In the process of identifying telomeres and centromeres in goose genome, animal "TTAGGG" is used as the telomere recognition sequence of goose, and the TeloExplorer function of quarTeT software is used for telomere identification; (7) In order to study the similarity of goose, duck and chicken in poultry at the karyotype level, NGenmoesyn software is used to align the colinearity of the assembled goose chromosome genome data with duck and chicken chromosome genome.
2. The method according to claim 1, wherein The six kinds of organ tissue samples in step 1 are brain, heart, liver, spleen, lung and kidney.
3. The method according to claim 1, wherein The sample DNA in step 1 is extracted by blood, cell, tissue genomic DNA extraction kit.
4. The method according to claim 1, wherein The DNA library construction and sequencing in step 1 includes the use of third generation ultra-long sequencing, HiFi sequencing and second generation short fragment sequencing gene library construction.
5. The method according to claim 1, wherein The Hi-C sequencing library construction and sequencing process in step 1 includes resuspension of the ball and resuspension of the cells using NEB buffer; then the cell nucleus is dissolved with dilute SDS lysis solution, the DNA is cut with four-base enzyme MboI, and the DNA end is labeled with biotin-14-dctp, and after the labeling is completed, T4 DNA polymerase is used to remove biotin; then, T4 DNA ligase is used for ligation operation; finally, after DNA purification treatment, double-end 150bp sequencing is carried out on Illumina Hiseq platform.
6. The method according to claim 1, wherein The RNA library construction and sequencing described in Step 1, total RNA was isolated from organ tissues using EasyPure RNA Kit; Subsequently, the sample RNA was subjected to sequencing library preparation using NEBNext Ultra RNA Library Prep Kit for Illumina; finally, double-end 2x125bp sequencing was performed on the Illumina HiSeq Xten platform; the mixed sample full-length transcript library was constructed and sequenced, and the Pacbio Sequel system was used for full-length transcript sequencing.
7. The method according to claim 1, wherein The genome size evaluation described in Step 2, statistical analysis was performed by double-end sequencing library data, and the distribution of K-mer was obtained using Jellyfish tool; subsequently, GenomeScope was used to model according to the K-mer distribution, thereby preliminarily revealing the characteristics of Taihu goose genome.
8. The method according to claim 1, wherein The genome assembly described in Step 2, first, Hifi data, Hifi+Hi-C data and Hifi+ONT ultra-long read+Hi-C data were used for genome assembly respectively; in addition, Hifi+ONT ultra-long read+Hi-C data was used for assembly using NextDenovo; in order to further improve the assembly quality, run_purge_dups.py tool was used to remove duplicate contigs; finally, according to the evaluation of N50 value, the assembly result of Hifi+Ont+Hi-C was selected as the data for subsequent analysis; Considering the problem of low accuracy of ONT third-generation ultra-long sequencing, Hifi data was used to correct the ONT data.
Citation Information
Patent Citations
Method for assembling and annotating Guide black fur sheep genome based on three-generation PacBio and Hi-C technologies
CN113005189A
Method for assembling and annotating Hu sheep genome based on three-generation PacBio and Hi-C technologies
CN113122642A