Chromosome level whole genome sequence of sea anemone sinensis
Through the third generation of high-throughput sequencing technology and chromosome mounting technology, the chromosome-level genome sequence of Chinese anemone was successfully obtained, solving the problems of data loss and splicing difficulty in the existing technology, and achieving high-integrity genome sequence acquisition, supporting subsequent molecular research and genetic analysis.
Patent Information
- Application Number
- CN202510172574.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the entire genome data of Chinese anemones are missing, resulting in incomplete coverage of functional gene mining and genetic information, and the high heterozygousness of the anemone genome sequence makes it difficult to splice.
The third-generation high-throughput sequencing technology combined with sequence assembly and chromosome mounting technology was used to obtain the chromosome-level genome sequence of Chinese anemone anemone. Specific steps include DNA extraction, library construction and sequencing, HiC library construction and sequencing, assembly and chromosome mounting.
The obtained genome-wide sequence has high integrity, and the BUSCO evaluation result is 95.9%, covering almost all gene sequences and containing a large number of non-coding sequences, supporting molecular marker development and population genetic structure analysis.
Smart Images

Figure CN120060237A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of molecular biology, and in particular relates to a chromosome-level full genome sequence of a Chinese anemone. Background Art
[0002] The Chinese nodular sea anemone (Paracondylactis sinensis), also known as sand garlic, is an economic species in the coastal areas of Zhejiang Province, my country. It has not yet been cultivated artificially, and is only supplied to the market through wild fishing. Studies have shown that due to overfishing of sand garlic, the wild sand garlic population has already experienced heterozygous deficiency. If relevant protection and artificial breeding measures are not taken, it will be detrimental to future sustainable development.
[0003] The whole genome sequence contains all the genetic information of an organism and is an important tool for the development of molecular markers, the mining of functional genes, and the study of the evolutionary process of species populations. Although some molecular data of the Chinese anemone have been published in public databases, the information provided by the published molecular data is limited compared to the whole genome data. Specifically, the following are the following: 1) The molecular information of the published partial gene data of the mitochondrial genome and the full-length data of the mitochondrial genome only involves the DNA information in the organelle genome, but cannot cover the nuclear DNA with richer genetic information, thus limiting the mining of functional genes; 2) Although the published second-generation transcriptome data contains the coding sequences of nuclear genes, the transcript sequences obtained after assembly are often not consistent due to the short read length of the second-generation sequencing, which is only 150bp. There are many assembly errors, which will also make many transcript sequences only obtain partial fragments, and it is impossible to obtain the full-length sequence of the transcript; in addition, the gene data obtained by the transcriptome only includes the transcript information contained in the sample tissue when sampling, and the gene information that was not expressed at that time cannot be obtained, and the genetic information contained in the second-generation transcriptome is incomplete; 3) Although the published third-generation transcriptome sequence can obtain the full-length sequence of the transcript, there is still a drawback of incomplete genetic information. After the data is evaluated by BUSCO software, its completeness is only 66.8%, which shows that the coverage of the genetic information of the sequence information for sand garlic is far from enough. In addition, the transcriptome data only includes the information of the protein coding sequence and its upstream and downstream regulatory regions, and it is impossible to cover most of the non-coding sequences in the genome. Finally, simply relying on transcriptome information cannot conduct relevant research on the gene distribution and chromosome evolution in the genome. These defects can only be solved after obtaining the genome information at the chromosome level.
[0004] There are generally two reasons why the chromosome-level genome of *Paracondylactis sinensis* has not been published yet: 1) There are relatively few experts in the taxonomy and genetic breeding of sea anemones in China. The research on its wild germplasm resources is relatively weak, and the government and scientific research departments have not paid enough attention to its research. Therefore, even though *Paracondylactis sinensis* has become a high-class dish in the restaurants of coastal cities in Jiangsu and Zhejiang, there is still a lack of access to its whole-genome information; 2) The genome sequence of sea anemones has the characteristic of high heterozygosity, so there is a certain difficulty in the assembly of its genome sequence. Summary of the Invention
[0005] In view of the deficiencies in the prior art, the present invention uses the third-generation high-throughput sequencing technology combined with sequence assembly and chromosome scaffolding to obtain the chromosome-level whole-genome sequence of *Paracondylactis sinensis*, laying a solid foundation for future research in the molecular field of sand garlic.
[0006] To achieve the above objectives, the solution of the present invention is as follows:
[0007] The chromosome-level whole-genome sequence of *Paracondylactis sinensis* is as shown in the ZENODO database, and the sequence storage address is https: / / zenodo.org / records / 14880344.
[0008] Furthermore, the method for obtaining the chromosome-level whole-genome sequence of *Paracondylactis sinensis* includes the following steps:
[0009] (1) Sample acquisition
[0010] Collect samples of *Paracondylactis sinensis*
[0011] (2) DNA extraction
[0012] Extract DNA from *Paracondylactis sinensis*
[0013] (3) Library construction and sequencing
[0014] Select high-quality DNA samples that pass the detection; randomly fragment the DNA into fragments; enrich and purify the DNA; perform damage repair and end repair on the fragmented DNA; ligate stem-loop sequencing adapters to both ends of the DNA fragments, and use exonuclease to remove the fragments with failed ligation; sequence the constructed library on the PacBio Revio platform in CCS mode to obtain HiFi data;
[0015] (4) HiC library construction and sequencing
[0016] The sea anemone tissue was ground in liquid nitrogen, and the chromatin was cross-linked with formalin solution; incubated at room temperature to terminate the formalin reaction, then incubated at room temperature and incubated on ice for more than 15 minutes; subsequently, the cells were lysed in pre-cooled lysis buffer; the chromatin was digested with the restriction enzyme DpnII and labeled with biotin residues, and then end-repaired; the HiC library with an insert size of 350 bp was prepared using the NEBNext DNA Library Construction Kit and sequenced with a read length of 150 bp to obtain the off-machine data;
[0017] (5) Assembly
[0018] The HiFi data after off-machine was assembled;
[0019] (6) Chromosome scaffolding
[0020] First, use the trimmomatic software to filter the HiC data after off-machine to remove adapter sequences and low-quality sequences;
[0021] Then use the fastuniq software to remove PCR duplicates during sequencing; construct the index, sequence alignment, and sorting;
[0022] Perform scaffolding and correction to obtain the chromosome-level genome sequence.
[0023] Preferably, the high-quality DNA sample described in step (3) is: a sample with a main band > 30 kb.
[0024] Preferably, the software used for the assembly in step (5) includes Hifiasm, Canu+purge, Flye, NextDenovo, and Spades software.
[0025] Preferably, the scaffolding and correction method in step (6) is:
[0026] Perform preliminary scaffolding to obtain the result of sequence scaffolding yahs.out_scaffolds_final.fa; process this result file with the following commands:
[0027] juicer pre - a - o out_JBAT yahs.out.bin yahs.out_scaffolds_final.agpcontig.fasta.fai
[0028] java - Xmx200G - jar juicer_tools.jar pre out_JBAT.txt out_JBAT.hicassembly 210673832
[0029] Among them, 210673832 is the genome size of *Paranemertes sinensis* calculated based on contig.fasta. After this step runs, out_JBAT.hic is obtained. This file and out_JBAT.assembly obtained from the previous command are used as input files for the juicebox software to manually adjust the mounting relationship of genomic fragments. After manual adjustment, the juicebox software outputs the file out_JBAT.review.assembly showing the arrangement relationship of each contig on the chromosome. This file is used as the input file for the juicer software to perform the final mounting of chromosome sequences. The command is as follows:
[0030] juicer post - o out_JBAT out_JBAT.review.assembly out_JBAT.liftover.agpcontig.fasta
[0031] The final output result file is out_JBAT.FINAL.fa, which serves as the final sequence file of the chromosome - level genomic sequence.
[0032] Preferably, the programs for mounting include: chromap + yahs process, HapHic software, juicer + 3d - DNA process, All - HiC software.
[0033] Preferably, the size of the DNA fragment described in step (3) is 15 - 18 kb.
[0034] Preferably, step (5) is to use the Hifiasm software to assemble the HiFi data after sequencing. For the software running parameters, except that - l is 3, all others are default. The running command is:
[0035] hifiasm - o contig.fasta - t 36 - l 3 hifi.fasta
[0036] Among them, contig.fasta is the genomic sequence at the contig level after assembly, and hifi.fasta is the data after PacBioRevio sequencing.
[0037] Preferably, the lysis buffer described in step (4) contains 10 mM NaCl, 0.2% IGEPAL CA - 630, 10 mM Tris–HCl, and 1× protease inhibitor solution.
[0038] Furthermore, the present invention provides the application of the chromosome-level whole genome sequence of *Paraneomeris sinensis* or the chromosome-level whole genome sequence of *Paraneomeris sinensis* obtained by the method in the analysis of the population genetic structure of *Paraneomeris sinensis*.
[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0040] The present invention uses the CCS sequencing technology in combination with sequence splicing and HiC chromosome mapping technology to obtain the whole gene sequence of *Paraneomeris sinensis* at the chromosome level. The chromosome-level genome is finally mapped to 19 haploid chromosomes. After BUSCO evaluation, its integrity is as high as 95.9%, indicating that the sequence already contains almost all the gene sequences of *Paraneomeris sinensis*. In addition, in addition to protein-coding sequences, the genome sequence also contains a large number of non-coding sequences (about 85.7%), which can be used for the development of molecular markers and the analysis of population genetic structure. The publication of this chromosome-level genome is an effective supplement to the gene resource information of *Paraneomeris sinensis*, which is conducive to the full development and utilization of biological gene resources in the later stage. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 : HiC interaction map of *Paraneomeris sinensis* chromosomes.
[0042] Figure 2 : Time tree of 12 sea anemone and 4 outgroup species. (The pink dots represent the time-calibrated nodes, and the numbers at the nodes represent the divergence time of the nodes, in millions of years).
[0043] Figure 3 : Dynamic change diagram of the effective population size of *Paraneomeris sinensis* in Taizhou and Dalian. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] In order to better understand the technical content of the present invention, specific embodiments are provided below to further illustrate the present invention.
[0045] Unless otherwise specified, the experimental methods used in the embodiments of the present invention are all conventional methods.
[0046] Unless otherwise specified, the materials, reagents, etc. used in the embodiments of the present invention can all be obtained from commercial channels.
[0047] Example 1:
[0048] (1) Sample acquisition
[0049] Samples of Peachia cathayensis were collected at low tide in the intertidal zone along the coast of Taizhou, Zhejiang. After sample collection, the mitochondrial genomes of the samples were sequenced. After obtaining the mitochondrial genome sequences of the samples, they were compared with the NCBI database, and it was found that the similarity between the sequenced mitochondrial genome and the mitochondrial genome sequence of Peachia cathayensis (OP903146) in the database was 100%, thus determining that the samples used for genome sequencing were of the species Peachia cathayensis.
[0050] (2) DNA extraction for third-generation sequencing
[0051] The extraction of high-quality DNA for third-generation sequencing was completed using the HMW DNAKit kit (Qiagen, Germany), and the operation steps were completed according to the kit instructions. The DNA content after extraction was measured using a Qubit 3.0 spectrophotometer, the integrity of the DNA after extraction was detected using agarose gel electrophoresis, and the purity of the DNA after extraction was detected using a NanoDrop instrument.
[0052] (3) Construction and sequencing of third-generation libraries
[0053] High-quality DNA samples that passed the detection (main band > 30 kb) were selected; randomly fragmented into fragments (15 - 18 kb) through a G-Tube; large-fragment DNA was enriched and purified using magnetic beads; the fragmented DNA was subjected to damage repair and end repair; stem-loop sequencing adapters were ligated to both ends of the DNA fragment, and the fragments with failed ligation were removed using exonuclease. The constructed library was sequenced on the PacBio Revio platform in the CCS (Circular Consensus Sequencing Mode) mode. Since the genome sizes of other related species of sea anemones are between 200M and 400M, in order to obtain sufficient sequencing depth, a total of 12.4G of HiFi data was measured in this invention for subsequent genome sequence assembly.
[0054] (4) Construction and sequencing of HiC libraries
[0055] The sea anemone tissue was ground in liquid nitrogen, and chromatin was cross-linked with 4% formalin solution. After incubating at room temperature for 5 minutes, 0.2 M glycine was added to terminate the formalin reaction, followed by incubation at room temperature for another 5 minutes and incubation on ice for more than 15 minutes. Subsequently, the cells were lysed in pre-cooled lysis buffer (10 mM NaCl, 0.2% IGEPAL CA-630, 10 mM Tris-HCl, and 1× protease inhibitor solution). The chromatin was digested with the restriction enzyme DpnII and labeled with biotin residues, and then end repair was performed. A Hi-C library with an insert size of 350 bp was prepared using the NEBNext DNA Library Construction Kit and sequenced on the NovaSeq 6000 platform with a read length of 150 bp, and a total of 25.1 Gb of raw data was obtained.
[0056] (5) Third-generation assembly
[0057] The Hifiasm software was used to assemble the HiFi data after sequencing. For the software running parameters, except for -l being 3, all others were default. The running command was:
[0058] hifiasm -o contig.fasta -t 36 -l 3 hifi.fasta
[0059] Among them, contig.fasta is the genomic sequence at the contig level after assembly, and hifi.fasta is the data after sequencing on PacBio Revio.
[0060] (6) Chromosome scaffolding
[0061] First, the trimmomatic software was used to filter the HiC data after sequencing to remove adapter sequences and low-quality sequences. The running command is as follows:
[0062] trimmomatic PE -threads 10 -phred33 XXX_1.fastq XXX_2.fastq XXX_R1_paired.fq XXX_R1_unpaired.fq XXX_R2_paired.fq XXX_R2_unpaired.fq ILLUMINACLIP:~ / sofware / trimmomatic -0.39 -2 / adapters / TruSeq3 -PE -2.fa:2:30:10 LEADING:15 TRAILING:15 SLIDINGWINDOW:4:15 MINLEN:40
[0063] Among them, XXX_1.fastq and XXX_2.fastq are the raw data after Illumina sequencing, where XXX is determined by the actual name of the data after sequencing, and XXX_R1_paired.fq and XXX_R2_paired.fq are the paired clean data after removing adapter sequences and low-quality sequences.
[0064] Then, use the fastuniq software to remove PCR duplicates during sequencing. The command is as follows:
[0065] fastuniq -i illumina.list -t q -o Hicrd_R1.fastq -p Hicrd_R2.fastq -c 1
[0066] Among them, illumina.list specifies the input file list, which contains the file name list of Illumina sequencing data, that is, XXX_R1_paired.fq and XXX_R2_paired.fq obtained in the previous step; Hicrd_R1.fastq and Hicrd_R2.fastq are the sequence data of the forward and reverse sequences after removing PCR duplicates, respectively.
[0067] Construct an index, perform sequence alignment and sorting. The command is as follows:
[0068] samtools faidx contig.fasta
[0069] chromap -i -r contig.fasta -o contigs.index
[0070] chromap --preset hic -r contig.fasta -x contigs.index --remove - pcr - duplicates -1 Hicrd_R1.fastq -2 Hicrd_R2.fastq --SAM -o aligned.sam -t 26
[0071] samtools view -bh aligned.sam|samtools sort -@50 -n > aligned.bam
[0072] Mount and visualize the preparation of corrected data. The command is as follows:
[0073] yahs contig.fasta aligned.bam
[0074] The output yahs.out_scaffolds_final.fa after running is the result of sequence scaffolding. To prevent errors in automatic scaffolding by the software, manual correction using the juicebox software is required. Therefore, this result file needs to be processed as the input file for juicebox. The commands are as follows:
[0075] juicer pre-a-o out_JBAT yahs.out.bin yahs.out_scaffolds_final.agpcontig.fasta.fai
[0076] java -Xmx200G -jar juicer_tools.jar pre out_JBAT.txt out_JBAT.hicassembly 210673832
[0077] Among them, 210673832 is the size of the Capitella capitata genome calculated based on contig.fasta. After this step runs, the obtained file out_JBAT.hic can be used together with the out_JBAT.assembly obtained in the previous step as the input files for the juicebox software to manually adjust the scaffolding relationship of genomic fragments. After manual adjustment, the juicebox software outputs the file out_JBAT.review.assembly showing the arrangement relationship of each contig on the chromosome. This file can be used as the input file for the juicer software for the final scaffolding of chromosome sequences. The commands are as follows:
[0078] juicer post-o out_JBAT out_JBAT.review.assembly out_JBAT.liftover.agpcontig.fasta
[0079] The final output result file is out_JBAT.FINAL.fa, which serves as the final sequence file for the Capitella capitata chromosome-level genome sequence.
[0080] Analyzing the scaffolding result, it is found that 93.414% of the contig.fasta sequences can be scaffolded onto the chromosomes, indicating a high final chromosome scaffolding rate. A total of 19 haploid chromosomes are obtained after scaffolding, and their interaction map is as follows ( Figure 1):The color of its diagonal is darker, indicating that there are more interaction sites among sequences within the same staining. After scaffolding, the chromosome-level N50 is 9.4 Mb, the contig N50 is 8.2 Mb, and the genome size is 210.68 Mb. The evaluation of the assembled genome using the BUSCO software shows that the sequence integrity is 95.9%.
[0081] (7) Data upload
[0082] The assembled chromosome sequence file has been uploaded to the ZENODO database as shown, and the sequence storage address is https: / / zenodo.org / records / 14880344. The md5 value of this sequence file is 744ce297bdb145f4cc45656f373f872f. Since each file has a unique md5 value, it can be used for subsequent patent examination and for users to check the sequence integrity after downloading.
[0083] (8) Data analysis
[0084] 1) Time tree
[0085] Based on the genome sequence of *Paranemactis sinica*, the protein-coding genes were predicted using the braker3 software, and a total of 22,897 protein-coding genes of *Paranemactis sinica* were predicted. The protein-coding genome of *Paranemactis sinica* was clustered with the protein sequences of other published sea anemone genomes (including: *Actinoscyphia liui* (GCA_041296415), *Anthopleura sola* (GCA_023349425), *Actinia tenebrosa* (GCA_009602425), *Alvinactis idsseensis* (https: / / doi.org / 10.57760 / sciencedb.06919), *Paraphelliactis xishaensis* (https: / / doi.org / 10.6084 / m9.figshare.13076387.v3), *Telmatactis stephensoni* (GCA_029948255), *Exaiptasia diaphana* (GCA_001417965), *Diadumene lineata* (GCA_918843875), *Nematostella vectensis* (GCA_932526225), *Scolanthus callimorphus* (GCA_033964015), *Actinernus sp.* (GCA_030867875)) using the orthofinder 3.0.1b1 software (the protein sequences of corals (*Acropora digitifera* (GCA_000222465), *Pocillopora damicornis* (GCA_003704095)) and corallimorpharians (*Amplexidiscus fenestrafer* (http: / / corallimorpharia.reefgenomics.org / download / ), *Discosoma sp.* (http: / / corallimorpharia.reefgenomics.org / download / )) were used as outgroups), and a total of 605 single-copy orthologous gene families were selected from the clustering results. The phylogenetic tree was constructed using the iqtree 1.5.5 software based on the protein sequences of the single-copy orthologous gene families.The obtained phylogenetic tree, together with the concatenated CDS sequences of 605 single-copy orthologous gene families, was used as the input file of MCMCTREE 4.9 software to infer the species differentiation time. The differentiation time of corals Acropora digitifera and Pocillopora damicornis (165-225 million years) and the differentiation time of sea anemones (424-608 million years) were used as time correction nodes in the inference process. After 200,000 Markov chain cycles, the species differentiation time tree (. Figure 2 ).
[0086] From the time tree results, we can know that the divergence time of the Chinese anemone Paracondylactis sinensis (the species marked in orange in the picture) is about 101.31 million years ago, which is the middle Cretaceous period. Due to global warming, rising sea levels and other factors, an environment suitable for the survival and evolution of many marine organisms has been formed, resulting in an increase in marine biodiversity.
[0087] 2) Dynamic changes in populations
[0088] Wild individuals of the Chinese anemone Trichodora sinensis living in Taizhou and Dalian were collected, and the extracted DNA was used for small fragment library construction and second-generation sequencing, with a sequencing volume of more than 4G for each individual. The genome sequence of the Chinese anemone Trichodora sinensis was used as a reference, and the measured second-generation data was aligned to the genome using BWA 0.7.17 software. PAML 4.9 software was used to combine the CDS sequences of the single-copy orthologous gene families in the beneficial effect 1 and the differentiation time results of each species ( Figure 3 ), the mutation rate of S. chinensis was calculated, and its base mutation rate was 1.23059×e-8 / year. Using psmc0.6.5 software, the base mutation rate of S. chinensis and the second-generation sequencing comparison results were used as input files to obtain the dynamic history of population size changes of wild species of S. chinensis in Taizhou and Dalian:
[0089] From the results, it can be seen that although the Chinese anemones in Taizhou and Dalian are the same species, they show different population size dynamics due to their different geographical locations. In general, the Taizhou population shows a higher effective population size than the Dalian population, indicating that the Taizhou population has a stronger adaptability to changes in the external environment. Although the population sizes of both species declined about 100,000 years ago, the effective population size in Taizhou declined earlier. This may be because 100,000 years ago was before the coldest stage of the last glacial period, and the temperature had already dropped. The Taizhou population was more adapted to warm environments, while the Dalian population had a stronger ability to adapt to low temperatures. Therefore, when the ambient temperature dropped sharply, the effective population size of the Taizhou population changed earlier and more dramatically than that of the Dalian population.
[0090] 3) Mining of functional proteins
[0091] The genes of *Paranemactis sinensis* encode 22,897 genes, which contain rich genetic resources for human beings to explore and utilize. The habitat of *Paranemactis sinensis* is relatively special compared with other sea anemones. It buries itself in sediment all year round, and usually only a small number of tentacles are exposed to seawater. Therefore, it is more vulnerable to invasion by pathogenic bacteria in the environment than other sea anemones. Through comparative genomics, we found that in the gene family of *Paranemactis sinensis*, there are 6 genes in the natterin gene family, showing an obvious expansion compared with other sea anemones.
[0092] Table 1
[0093]
[0094] Previous studies have shown that natterin is an important class of novel immune defense factors. The expansion of these genes plays an important role in the process of *Paranemactis sinensis* resisting the invasion of a large number of pathogenic bacteria in the buried environment. The 6 natterin protein sequences of *Paranemactis sinensis* are as follows, which can be used for the screening of potential drugs and the research and development of new anti-cancer drugs in the future.
[0095] >Psin_g10895_t1
[0096] MGRFGVMSSLLLAVFLLIETASAASNLKWVSASNGNIPAYAVAGGEDFPGEVLYVARIQLATGLTPGKVDASSKLAHSSWGGKEIYKSDYQVLTNPGWRSHLEWKKSSGHTPPANAVVGGSDNGKPLYVARFLYSDGHMIPGKASYVQGLAHIAYGGKEYYESEWYVLVEHTGSGRKRGLSSRFEYPDLPDVKPLKRAPVIDQ*
[0097] >Psin_g10897_t1
[0098] MGRFGVMSSILLAVLLLIETASAASNLKWVYATNGNIPAYAVAGGEDFPGEVLYVARIPLRTGLTPGKVDKSSKLAHSPWGGKEHYKSYYQVLTNPGWRSHLEWKKSSGHTPPANAVVGGSDNGKPLYVARALYSDGHMIPGKAAYFYGKAYIPYGGKEYEVKEWYVLVEHTGSGRKRGLSSRFEYPDLPDVKPLKRAPVIDQ*
[0099] >Psin_g10911_t1
[0100] MGRLEGKLSILLAVLTLMETASAASNLKWVSADRGNVPDHAVAGGEDYPGEVLYIARIYLDSGPTPGKMDESSKLAHASWGGKEHYNSEYQILTNPGWKSHLEWKKSSGQTPPANAVVGGHDHGKPLYVARALYSDGHMIPGKAAYFYGKAYIPYGGKEYEAKEWYVLVEHTRSRRKRGVSSRFEYPDLPDVKPLKRAPVIDQ*
[0101] >Psin_g8317_t1
[0102] MEGFGVMFCFLVLAGMFAIETTVDAQRSLPTKLIWVRPTQGQVPDFAVVGGTEASGEPLYIARASINGALIPGKMNPGYKKAYVANGKYENEFTNDNYWILTNPRYSTHLKWVQSSGSYKPLNAVVGGYENGNLLYVARETFNDGHFIPGKASYSDQRASFPHNGFANIRQKWEVLTEEPRVFDPPIVDQPTIEDLMNGHVIGKRNVSDSFDGLLMKTP*
[0103] >Psin_g8318_t1
[0104] MEGFGVKFCFLVLAGVLVIETTVEADSDLRWVAASSTTPFPKYAVAGGTDYSGEALYVARITLNNNVIPGKVTKSQKIAHATHNTREINADHYEILTNPRWTVKFEWTKSSGSTPPAKAIIGGHEDCTSLYVARMRHTKGNFIPGKANYVEQKGSYAHGYREEQRNEWYVLTEPRRPFGKRDALDHSLKSAKDLLMKSP*
[0105] >Psin_g8322_t1
[0106] MGGSWVKFFILLVVVFMTENLVDASKLKWVADSYGHVPDNPVVGGSEPDGSPLYVARIRLESGFVPGKMSPLYKKAYAPYNTYEKEGSHYEILTNPERVDLDWQESSGSHPPANAVAGGSDNDRVLYVARMSHGHHMIPGKAAYFYQKGYYSYHGREYREDTWHVLVEQSSPGRKRGLLSLLKFPKDVADAKPLKPAINQ*
[0107] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. The chromosome-level whole genome sequence of the Chinese anemone, It is obtained as follows: (1) Sample acquisition Collecting samples of the Chinese anemone (2) DNA extraction Extraction of DNA from the Chinese anemone (3) Library construction and sequencing Select high-quality DNA samples that meet the test requirements; randomly break them into DNA fragments; enrich and purify DNA; perform damage repair and end repair on the fragmented DNA; connect stem-loop sequencing adapters at both ends of the DNA fragments, and use exonucleases to remove fragments that fail to connect; sequence the constructed library using the PacBio Revio platform in CCS mode to obtain HiFi data; (4) HiC library construction and sequencing Anemone tissue was ground in liquid nitrogen and chromatin was cross-linked using formalin solution; the formalin reaction was terminated by incubation at room temperature, followed by incubation at room temperature and on ice for more than 15 minutes; Subsequently, the cells were lysed in pre-chilled lysis buffer; the chromatin was digested with restriction endonuclease DpnII and labeled with biotin residues, followed by end repair; A HiC library with an insert size of 350 bp was prepared using the NEBNext DNA library construction kit and sequenced with a read length of 150 bp to obtain offline data; (5) Assembly Assemble the HiFi data after leaving the machine; (6) Chromosome mounting First, the HiC data after being downloaded were filtered using trimmomatic software to remove adapter sequences and low-quality sequences; Then, fastuniq software was used to remove PCR duplications during sequencing; primers were constructed, sequences were aligned and sorted; Mount and correct to obtain chromosome-level genome sequence.
2. The chromosome-level whole genome sequence of the Chinese anemone Anemone cerana according to claim 1, characterized in that: The high-quality DNA sample in step (3) is a sample with a main band >30 kb, and the size of the DNA fragment is 15-18 kb.
3. The chromosome-level whole genome sequence of the Chinese anemone Anemone cerana according to claim 1, characterized in that: The software used for assembly in step (5) includes Hifiasm, Canu+purge, Flye, NextDenovo, and Spades software.
4. The chromosome-level whole genome sequence of the Chinese anemone Anemone cerana according to claim 1, characterized in that: The mounting and correction method of step (6) is: Perform a preliminary mount and obtain the result of the sequence mount, yahs.out_scaffolds_final.fa; process the result file with the following command: juicer pre-ao out_JBAT yahs.out.bin yahs.out_scaffolds_final.agpcontig.fasta.fai java-Xmx200G-jar juicer_tools.jar pre out_JBAT.txt out_JBAT.hic assembly210673832 210673832 is the genome size of S. sinensis calculated based on contig.fasta. This step is followed by out_JBAT.hic, which is used together with out_JBAT.assembly obtained from the previous step as input files of juicebox software to manually adjust the mounting relationship of genome fragments. After manual adjustment, the juicebox software outputs the arrangement relationship file of each contig on the chromosome, out_JBAT.review.assembly, which is used as the input file of the juicer software for the final mounting of the chromosome sequence. The command is as follows: juicer post-o out_JBAT out_JBAT.review.assembly out_JBAT.liftover.agpcontig.fasta The final output result file is out_JBAT.FINAL.fa, which is the final sequence file of chromosome-level genome sequence.
5. The chromosome-level whole genome sequence of the Chinese anemone Asterina sinensis according to claim 1, characterized in that: The programs used for mounting include: chromap+yahs process, HapHic software, juicer+3d-DNA process, and All-HiC software.
6. The chromosome-level whole genome sequence of the Chinese anemone Anemone cerana according to claim 1, characterized in that: Step (5) is to use Hifiasm software to assemble the HiFi data after logging off the machine. The software running parameters, except for -l which is 3, are all default. The running command is: hifiasm-o contig.fasta-t 36-l 3hifi.fasta Among them, contig.fasta is the genome sequence at the contig level after assembly, and hifi.fasta is the data after PacBio Revio is downloaded.
7. The chromosome-level whole genome sequence of the Chinese anemone Anemone cerana according to claim 1, characterized in that: The lysis buffer in step (4) contains 10 mM NaCl, 0.2% IGEPAL CA-630, 10 mM Tris-HCl and 1× protease inhibitor solution.
8. The chromosome-level whole genome sequence of the Chinese anemone Anemone cerana according to claim 1, characterized in that: The incubation time in step (4) is 5 min; the formalin reaction is terminated by adding 0.2 M glycine to terminate the formalin reaction.
9. The chromosome-level whole genome sequence of the Chinese anemone Anemone cerana according to claim 1, characterized in that: The chromosome-level whole genome sequence of the Chinese anemone is shown in the ZENODO database, and the sequence storage address is https: / / zenodo.org / records / 14880344.
10. Use of the chromosome-level whole genome sequence of the Chinese anemone Aspergillus niger according to any one of claims 1 to 9 in the analysis of the genetic structure of the Chinese anemone population.