A method for constructing a haploid genome of oyster

By constructing a whole oyster family line and using high-throughput sequencing and parent-specific SNP alignment, the haplotype genome of the parent and maternal haplotypes were successfully isolated and assembled, which solved the problem of difficulty in oyster genome phase and improved the integrity and continuity of genome sequences.

CN115992261BActive Publication Date: 2025-06-24INST OF OCEANOLOGY - CHINESE ACAD OF SCI
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211603430.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-13
Publication Date
2025-06-24
Estimated Expiration
2042-12-13

AI Technical Summary

Technical Problem

The prior art cannot effectively construct the high-quality haplotype genome of oysters, resulting in difficulty in genomic sequence phase, limiting the application of haplotype identification, structural variation detection and genetic analysis.

Method used

The whole oyster family line was constructed by artificially constructing, and the haplotype genomes of the parent-specific single nucleotide variant site were combined with the alignment of parent-specific sequencing, and the progeny sequencing sequences were grouped and assembled to obtain the haplotype genomes of the parent and maternal parent.

Benefits of technology

High-quality haplotype genome construction has been achieved, and the integrity and continuity of genomic sequences have been significantly improved, which is suitable for haplotype identification and genetic analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0003996287150000031
    Figure BDA0003996287150000031
  • Figure BDA0003996287150000041
    Figure BDA0003996287150000041
Patent Text Reader

Abstract

The present invention belongs to the field of marine organism genomes and molecular genetics, and particularly relates to a method for constructing a haploid genome of oysters. Select a male and a female oyster individual, artificially construct a full-sib family and conventionally breed the offspring; perform high-throughput sequencing on the paternal and maternal individuals respectively to identify paternal-specific single nucleotide variant sites and maternal-specific single nucleotide variant sites; after the offspring individuals are bred for one year, take one individual for high-throughput sequencing; use the parental-specific single nucleotide variant sites to group the sequencing sequences of the offspring individuals to obtain paternal-derived offspring sequences and maternal-derived offspring sequences; assemble the paternal-derived and maternal-derived offspring sequences respectively to obtain paternal-derived and maternal-derived haploid genomes. Applying this method to obtain high-quality haploid genomes of Crassostrea gigas and Crassostrea angulata, the BUSCO assessment shows that the genome integrity is greater than 93%, indicating that the haploid genomes of oysters obtained by this method have high quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field:

[0001] The present invention belongs to the field of marine organism genomes and molecular genetics, and particularly relates to a method for constructing a haploid genome of oyster. Background Art:

[0002] The genomes of marine shellfish such as oysters are relatively complex, and conventional genome assembly strategies cannot obtain high-quality haploid genomes. With the development of sequencing technology, the cost of sequencing has decreased rapidly, and the data output and the length of sequencing sequences have both increased. However, generally, the conventional splicing and assembly of the genomes of diploid organisms such as oysters cannot obtain two sets of complete haploid genome sequences. Instead, a mosaic chimeric form of two sets of complete haploid genome sequences is obtained, and the phasing of the whole genome sequence cannot be achieved, which restricts the application of genome sequences in genetic analyses such as haplotype identification, structural variation detection, comparative genomics, and linkage disequilibrium. Currently, there is no mature and effective method and strategy for haploid assembly of complex genomes of shellfish such as oysters. Summary of the Invention:

[0003] The purpose of the present invention is to provide a method for constructing a haploid genome of oyster.

[0004] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0005] A method for constructing a haploid genome of oyster,

[0006] (1) Select female and male individuals of oyster, artificially construct a full-sib family and conventionally breed the offspring;

[0007] (2) After the offspring individuals are bred for one year, take one individual for long-read high-throughput sequencing;

[0008] (3) Respectively perform short-read high-throughput sequencing on the female and male individuals of the full-sib family;

[0009] (4) Use the sequencing sequences of the offspring for preliminary assembly;

[0010] (5) Align the parental sequences to the preliminary assembly to identify parental-specific single nucleotide variant sites;

[0011] (6) Group the sequencing sequences of the offspring individuals using the parental-specific single nucleotide variant sites to obtain long sequences from the paternal source and the maternal source;

[0012] (7) Respectively assemble the long sequences from the two parental sources to obtain two haploid genome sequences.

[0013] In step (1), several first - or second - year oyster female and male individuals with a shell height > 5 cm are collected during the oyster breeding season for future use. Among them, the female and male oysters can be of the same species or different species.

[0014] In step (2), one individual is subjected to long - read high - throughput sequencing. Among them, the long - read high - throughput sequencing uses the SMART sequencing mode of PacBio to obtain ccs sequences, and the sequencing depth is not less than 60×.

[0015] In step (3), the short - read high - throughput sequencing depth is not less than 50×, and then the genome size, heterozygosity, and proportion of repetitive sequences of the female and male parental individuals are estimated. The estimated value of the genome size (Ge) provides a reference for the genome size (Go) obtained by subsequent sequence assembly, and there is no range requirement for the size of Ge. The estimated value of heterozygosity (HETe) is preferably > 1%. A lower heterozygosity is not conducive to the subsequent SNP - based sequencing sequence grouping. The proportion value of repetitive sequences (REp) is generally required to be < 70%. An excessive REp value will lead to fragmentation of the genome sequence assembly.

[0016] In step (4), using the offspring individual sequences obtained in step 2 as input, splicing is performed, and the integrity after splicing needs to reach a Complete BUSCOs (C) index value of not less than 85%.

[0017] The sequencing sequences of the parents in step (3) are respectively compared with the gene sequences obtained in step (4), and parental - specific single - nucleotide variant sites are identified based on the comparison information.

[0018] The parental - specific single - nucleotide variant sites are SNV genotype combinations that have different genotypes in the two parents and have parental - specific alleles in the offspring. Among them, the SNV genotype combination forms are:

[0019] (1) FF×MM, where F is specific to the father and M is specific to the mother;

[0020] (2) FF×FM, where the origin of F cannot be determined and M is specific to the mother;

[0021] (3) FM×MM, where F is specific to the father and the origin of M cannot be determined;

[0022] Among them, F and M represent the four constituent bases A, C, G, and T of genomic DNA.

[0023] The specific implementation process is as follows:

[0024] 1. For the oyster species to be tested, collect a number of one-year-old or two-year-old oyster individuals with a shell height greater than 5 cm during the oyster breeding season. After completely prying open the two shells of the oyster with an oyster knife, select individuals with fully developed gonads, and use an ordinary optical microscope to examine the gonad tissue cells to determine the sex of the oysters. Construct a full-sib family of oysters using the method of one male to one female artificial insemination and conduct conventional artificial breeding. At the same time, freeze the mantle tissue of the parental oyster individuals used for family construction at -80 °C for future use.

[0025] 2. Approximately one year after the full-sib family offspring individuals are cultured, randomly select one individual with a shell height greater than 5 cm, open the shell and dissect it, take about 300 - 500 mg of mantle tissue, extract large-fragment genomic DNA by the phenol-chloroform method. After passing the quality inspection, use the PACBIO SMART SEQUELII sequencer for sequencing. Process the bam file of the raw data downloaded from the machine using the ccs program to obtain ccs sequences. Accumulate the lengths of all ccs sequences and calculate the genomic coverage of the offspring sequences. If the coverage is less than 60×, additional data testing is required.

[0026] 3. Extract genomic DNA sequences from the mantle tissues of male and female parents. After passing the quality inspection, use sequencers such as ILLUMINAHISEQ series or NOVASEQ6000 for sequencing. Filter the fastq file of the raw data downloaded from the machine using the fastp program to remove sequences with an average quality value lower than 20. Accumulate the lengths of high-quality sequences and calculate the genomic coverage of the sequences, which is required to reach more than 50×, otherwise additional data testing is needed. Use the genomescope program to estimate the genomic size, heterozygosity, and repetitive sequence ratio of male and female parents and offspring individuals.

[0027] 4. Use programs such as hifiasm, with the ccs sequences of the offspring individuals obtained in step 2 as the input, for splicing. After the initial splicing is completed, use the PurgeHaplotigs program to remove redundant sequences in the assembled sequences. Use the BUSCO program to evaluate the integrity of the assembly with the metazoa_odb10 database as the reference, and the integrity needs to reach an index value of CompleteBUSCOs (C) not lower than 85%. If the C value is lower than 85%, it indicates that the sequencing data volume in step 2 is too low and additional data testing is required. If the index value of Complete and duplicated BUSCOs (D) is greater than 10%, it indicates that there are too many redundant sequences, and the PurgeHaplotigs program needs to be used to remove redundancy from the assembled sequences multiple times. After this step, a set of diploid fusion genome assembly sequences is obtained.

[0028] 5. Align the sequencing sequences of the parents to the genomic sequence obtained in step 4 using the bwa program to obtain the aligned bam file. Use the Samtools program to sort the sequences in the bam file, use the Picard program to remove the duplicate alignments caused by PCR amplification in the bam file, and use the gatk program to detect single nucleotide variations (SNVs) in the paternal bam and maternal bam files to identify parental-specific SNVs. The specific criteria are as follows: SNV genotype combinations that are different in genotype between the two parents and have parental-specific alleles in the offspring, including 3 patterns: FF×MM, FF×FM, and FM×MM. The specific genotype combinations are shown in Table 1 below.

[0029] Table 1 SNV genotype combinations of parental-specific alleles (paternal×maternal)

[0030]

[0031]

[0032] 6. Align the ccs sequences of the offspring individuals to the genomic sequence obtained in step 4 using the minimap2 program to obtain the aligned bam file. Locate the parental-specific SNVs obtained in step 5 in the offspring ccs sequences and count the allele base types of the SNVs in the offspring ccs sequences. According to the base combination forms of the parental genotypes and offspring genotypes of the SNVs in Table 1, count the number of paternal-specific bases and maternal-specific bases with a sequencing quality value greater than Q20 in each offspring ccs sequence to determine the sequence origin. Finally, obtain the paternal-specific sequence set and the maternal-specific sequence set.

[0033] 7. Use programs such as hifiasm, with the offspring ccs sequences from the two parents obtained in step 6 as input, for assembly. After the assembly is completed, if the genomic sequence length is significantly greater than the expected value, the PurgeHaplotigs program needs to be used to remove redundant sequences in the assembled sequence. Use the BUSCO program to evaluate the integrity of the assembly with the metazoa_odb10 database as a reference, and the integrity needs to reach a Complete BUSCOs (C) index value of not less than 90%. If the C value is lower than 90%, it indicates that the sequencing data volume in step 2 is too low and additional data needs to be sequenced. If the Complete and duplicated BUSCOs (D) index value is greater than 5%, it indicates that there are too many redundant sequences and the PurgeHaplotigs program needs to be used to remove redundancy from the assembled sequence multiple times. After this step, two haploid genomes are obtained.

[0034] Advantages of the present invention:

[0035] The haplotype genome construction method of the present invention can achieve the grouping and splicing and assembly of the paternal and maternal sequences of the sequenced individuals, so as to obtain high-quality paternal and maternal haplotype genomes simultaneously. Applying this method to obtain high-quality haplotype genomes of the Pacific oyster and the Fujian oyster, the BUSCO evaluation of the genome integrity is greater than 93%, indicating that the oyster haplotype genomes obtained by this method have high quality. Compared with the commonly used hybrid assembly strategy, higher sequence continuity and integrity can be obtained under the condition of lower sequencing depth.

[0036] The present invention adopts the strategy of identifying parental-specific SNPs and using them for the grouping and separate assembly of the sequencing sequences of offspring individuals, which can achieve the efficient grouping of the parental haplotype sequencing sequences, give full play to the advantages of the high heterozygosity of oysters, and obtain two completely phased haplotype genomes, namely the haplotype genome of the paternal individual and the haplotype genome of the maternal individual. Specific implementation mode:

[0037] The present invention is not limited to the various components described below, and various changes can be made within the scope of the invention claimed. The implementation modes and embodiments obtained by appropriately combining the technical means disclosed in different implementation modes and embodiments are also included in the technical scope of the present invention.

[0038] The construction method of the present invention is as follows: Select one male and one female oyster individual, artificially construct a full-sib family and conduct conventional breeding on the offspring; after one year of breeding the offspring individuals, take one individual for high-throughput sequencing; conduct high-throughput sequencing on the paternal and maternal individuals respectively, and identify the paternal-specific single nucleotide variant sites and the maternal-specific single nucleotide variant sites; use the parental-specific single nucleotide variant sites to group the sequencing sequences of the offspring individuals to obtain the paternal offspring sequences and the maternal offspring sequences; assemble the paternal and maternal offspring sequences respectively to obtain the paternal and maternal haplotype genomes.

[0039] Example 1

[0040] (1) In May of the current year, select the Pacific oyster cultured in the Qingdao sea area in northern China. During the oyster breeding season, select several one-year-old oyster individuals with well-developed gonads and a shell height > 5 cm to construct 5 full-sib families, and sample the mantle and gill tissues of the parents and store them at -80 °C for later use. Use the method of artificial insemination with one male and one female to construct the oyster full-sib family and conduct conventional artificial breeding. The offspring individuals are cultured in the Qingdao sea area until December of the following year. Randomly select 30 offspring individuals with a shell height > 5 cm from 1 family, and sample the mantle and gill tissues after dissection and store them.

[0041] (2) Select one offspring individual from the above selected offspring. After shelling, dissect it, take about 300 - 500 mg of mantle tissue, extract the mantle genomic DNA using the conventional phenol-chloroform method. After passing the quality inspection, use the PACBIO SEQUELII sequencer for sequencing. Process the bam file of the raw data downloaded from the machine using the ccs program (parameters: --min-passes 3 -j 6 --min-length 5000 --max-length 100000) to obtain the ccs sequences. Accumulate the lengths of all ccs sequences and calculate the genomic coverage of the offspring sequences; a total of 2,635,123 long sequences were obtained, and the coverage was 66×.

[0042] (3) Extract the genomic DNA sequences from the mantle tissues of the two selected parents above. After passing the quality inspection, perform second-generation whole-genome resequencing using sequencers such as NOVASEQ6000. Filter the fastq file of the raw data downloaded from the machine using the fastp program (parameters: -w 8 -q 20 -u 40 -n 0 -e 20) to obtain paired raw sequencing sequences with a length of 150 bp. After deleting low-quality sequences with an average sequencing quality value less than Q20 and more than 3 Ns, 249,410,590 and 255,611,428 sequences were obtained in the male and female parents respectively, and the sequencing coverages were 62× and 64× respectively. Then use the genomescope program to estimate (parameters: -i k21.histo -k 21) that the genomic sizes of the male and female parents are 582M and 579M respectively, and the heterozygosities are 2.8% and 2.7% respectively.

[0043] (4) Use hifiasm for preliminary assembly with the long sequences of the offspring in (2) (the ccs sequences of the offspring individuals) as the input. The obtained genomic size is 1026M, which is significantly larger than the expected value. After using the PurgeHaplotigs program to remove redundancy, the obtained fused genomic size is 631M. Use the BUSCO program to evaluate the assembly integrity with the metazoa_odb10 database as the reference (parameters: --cpu 6 --offline --mode genome --evalue 1e-5

[0044] --metaeuk_parameters = "--remove-tmp-files = 1"

[0045] --metaeuk_rerun_parameters = "--remove - tmp - files = 1"), after the BUSCO program evaluation, it was found that the integrity index value of Complete BUSCOs was 93.7%, and the index value of Complete and duplicated BUSCOs was 4.1%, thus obtaining the diploid fusion genome assembly sequence set.

[0046] (5) Align the second - generation sequencing sequences of the parental and offspring individuals to the fusion genome using the bwa program, and use the samtools, picard, and gatk programs to identify single - nucleotide base variations (SNVs) in the genome. The number of SNV genotype distributions obtained is shown in the following table; specifically: obtain the aligned bam file, use the sort command of the Samtools program to sort the sequences in the bam file, use the Picard program to remove the duplicate alignments caused by PCR amplification in the bam file, and use the gatk program to detect single - nucleotide variations (SNVs) in the paternal bam and maternal bam files (parameters are QD < 2 || FS > 60 || MQ < 30 || MQRankSum < - 8 || ReadPosRankSum < - 8"), and identify parental - specific SNVs. The specific criteria are: SNV genotype combinations that are different in genotype between the two parents and have parental - specific alleles in the offspring, including 3 patterns: FF×MM, FF×FM, and FM×MM.

[0047] SNV 亲本型 子代型 数目 SNV 亲本型 子代型 数目 A:C AA×AC AC 27683 C:A AA×CA CA 47049 A:C AA×CC AC 13641 C:A AA×CC CA 26149 A:C AC×AA AC 46034 C:A CA×AA CA 25170 A:C AC×CC AC 30144 C:A CA×CC CA 64454 A:C CC×AA AC 26281 C:A CC×AA CA 10694 A:C CC×AC AC 40493 C:A CC×CA CA 29958 BE AA×AG BE 72647 C:G CC×CG CG 15740 BE AA×GG BE 25979 C:G CC×GG CG 7119 BE AG×AA BE 142020 C:G CG×CC CG 26670 BE AG×GG BE 55424 C:G CG×GG CG 13262 BE GG×AA BE 55998 C:G GG×CC CG 9030 BE GG×AG BE 97826 C:G GG×CG CG 22007 A:T AA×AT AT 79711 C:T CC×CT CT 99233 A:T AA×TT AT 21834 C:T CC×TT CT 32387 A:T AT×AA AT 105529 C:T CT×CC CT 121656 A:T AT×TT AT 46313 C:T CT×TT CT 48851 A:T TT×AA AT 50906 C:T TT×CC CT 60923 A:T TT×AT AT 73512 C:T TT×CT CT 110312 G:A AA×GA GA 122334 T:A AA×TA TA 81680 G:A AA×GG GA 61449 T:A AA×TT TA 42343 G:A GA×AA GA 48978 T:A TA×AA TA 45586 G:A GA×GG GA 149182 T:A TA×TT TA 118425 G:A GG×AA GA 31056 T:A TT×AA TA 28694 G:A GG×GA GA 99967 T:A TT×TA TA 70588 G:C CC×GC GC 22904 T:C CC×TC TC 111852 G:C CC×GG GC 9068 T:C CC×TT TC 54825 G:C GC×CC GC 10396 T:C TC×CC TC 54443 G:C GC×GG GC 26173 T:C TC×TT TC 10728 G:C GG×CC GC 6524 T:C TT×CC TC 27353 G:C GG×GC GC 18308 T:C TT×TC TC 75121 G:T GG×GT GT 35141 T:G GG×TG TG 44564 G:T GG×TT GT 15020 T:G GG×TT TG 26061 G:T GT×GG GT 61866 T:G TG×GG TG 28589 G:T GT×TT GT 22228 T:G TG×TT TG 42135 G:T TT×GG GT 22628 T:G TT×GG TG 10891 G:T TT×GT GT 39996 T:G TT×TG TG 35674

[0048] (6) Align the long sequences of the offspring to the fusion genome using the minimap2 program, align to the genome sequence obtained in step 4 (parameters are - t 40 - c - N 2 - Y --eqx - x asm20), obtain the aligned bam file, locate the parental - specific SNVs obtained in step 5 in the offspring ccs sequences, and count the allele base types of the SNVs in the offspring ccs sequences. According to the SNV genotype distributions listed in the above table, determine the parental origin of each offspring sequence. Finally, 1,251,683 offspring sequences of paternal origin and 1,256,624 offspring sequences of maternal origin are obtained, and the genome coverage is approximately 33×.

[0049] (7) Using the hifiasm program, taking the offspring sequences from the two parental sources obtained in step 6 as input, perform splicing and assembly separately. The lengths of the obtained genomic sequences are 590M and 586M respectively, which are consistent with the evaluation results in step (3). Using the BUSCO program, evaluate the assembly integrity with the metazoa_odb10 database as the reference (parameters are --cpu 6 --offline --mode genome --evalue 1e-5 --metaeuk_parameters = "--remove-tmp-files = 1" --metaeuk_rerun_parameters = "--remove-tmp-files = 1"). The completeness of Complete BUSCOs is 93.3% and 92.5% respectively, and the values of Complete and duplicated BUSCOs are 1.2% and 0.9% respectively, indicating that two high-quality haploid genomes are obtained.

[0050] Example 2

[0051] (1) During the oyster breeding season of the current year, select two-year-old individuals of Crassostrea gigas, the main cultured oyster species in the northern sea area of China, and individuals of Crassostrea angulata, the main cultured oyster species in the southern sea area of China. Select individuals with well-developed gonads and a shell height > 5 cm. Use Crassostrea gigas as the female parent and Crassostrea angulata as the male parent to construct a Changfu hybrid family. Sample the mantle and gill tissues of the parents and store them at -80 °C for later use. Use the method of artificial insemination with one male and one female to construct a full-sib family of oysters, and conduct conventional artificial breeding. The offspring individuals are cultured in the Qingdao sea area until December of the following year. Randomly select 30 offspring individuals with a shell height > 5 cm from 1 family, and sample and preserve the mantle and gill tissues after dissection.

[0052] (2) From 1 offspring individual selected from the above-mentioned selected offspring, open the shell and dissect it. Take about 300 - 500 mg of mantle tissue, extract the genomic DNA of the mantle using the conventional phenol-chloroform method. After passing the quality inspection, use the PACBIO SEQUELII sequencer for sequencing. Process the bam file of the raw data downloaded from the machine using the ccs program (parameters are --min-passes 3 -j 6 --min-length 5000 --max-length 100000) to obtain ccs sequences, accumulate the lengths of all ccs sequences, and calculate the genomic coverage of the offspring sequences; a total of 2,709,684 long sequences are obtained, and the coverage is 76×.

[0053] (3) Extract the genomic DNA sequences from the mantle tissues of the two selected parents. After passing the quality inspection, use a sequencer such as NOVASEQ6000 to perform second-generation whole-genome resequencing. Filter the fastq files of the raw data downloaded from the sequencer using the fastp program (parameters: -w 8 -q 20 -u 40 -n 0 -e 20) to obtain paired raw sequencing sequences with a length of 150bp. After deleting low-quality sequences with an average sequencing quality value less than Q20 and an N count greater than 3, 366741548 and 38339924 sequences are obtained in the male and female parents respectively, and the sequencing coverages are 91× and 95× respectively. Then use the genomescope program to estimate (parameters: -i k21.histo -k 21) the genome sizes of the male and female parents to be 580M and 558M respectively, and the heterozygosities are 2.9% and 2.6% respectively.

[0054] (4) Use hifiasm to perform preliminary assembly with the long sequences of the offspring in (2) (the ccs sequences of the offspring individuals) as the input. The obtained genome size is 983M, which is significantly larger than the expected value. After using the PurgeHaplotigs program to remove redundancy, the obtained fused genome size is 730M. Use the BUSCO program to evaluate the assembly integrity with the metazoa_odb10 database as the reference (parameters: --cpu 6 --offline --mode genome --evalue 1e-5

[0055] --metaeuk_parameters = "--remove-tmp-files = 1"

[0056] --metaeuk_rerun_parameters = "--remove-tmp-files = 1"). After the BUSCO program evaluation, it is found that the integrity index value of Complete BUSCOs is 95.7%, and the index value of Complete and duplicated BUSCOs is 4.5%, thus obtaining the diploid fused genome assembly sequence set.

[0057] (5) Align the second-generation sequencing sequences of the parents and offspring individuals to the fused genome using the bwa program, and use the samtools, picard, and gatk programs to identify single nucleotide base variations (SNVs) in the genome (parameters: QD < 2 || FS > 60 || MQ < 30 || MQRankSum < -8 || ReadPosRankSum < -8”), and the number of SNV genotype distributions obtained is shown in the following table.

[0058] SNV 亲本型 子代型 数目 SNV 亲本型 子代型 数目 A:C AA×AC AC 31820 C:A AA×CA CA 40560 A:C AA×CC AC 12749 C:A AA×CC CA 22350 A:C AC×AA AC 51724 C:A CA×AA CA 22474 A:C AC×CC AC 25987 C:A CA×CC CA 58067 A:C CC×AA AC 23258 C:A CC×AA CA 13042 A:C CC×AC AC 47639 C:A CC×CA CA 36535 BE AA×AG BE 76471 C:G CC×CG CG 15282 BE AA×GG BE 31300 C:G CC×GG CG 5933 BE AG×AA BE 118350 C:G CG×CC CG 23813 BE AG×GG BE 62275 C:G CG×GG CG 11145 BE GG×AA BE 55444 C:G GG×CC CG 10380 BE GG×AG BE 115090 C:G GG×CG CG 20568 A:T AA×AT AT 68717 C:T CC×CT CT 86290 A:T AA×TT AT 25993 C:T CC×TT CT 31444 A:T AT×AA AT 106595 C:T CT×CC CT 133688 A:T AT×TT AT 46313 C:T CT×TT CT 54279 A:T TT×AA AT 45452 C:T TT×CC CT 54886 A:T TT×AT AT 87515 C:T TT×CT CT 103096 G:A AA×GA GA 102802 T:A AA×TA TA 87828 G:A AA×GG GA 54866 T:A AA×TT TA 45531 G:A GA×AA GA 54420 T:A TA×AA TA 46517 G:A GA×GG GA 134399 T:A TA×TT TA 106690 G:A GG×AA GA 32017 T:A TT×AA TA 25620 G:A GG×GA GA 86179 T:A TT×TA TA 68533 G:C CC×GC GC 20635 T:C CC×TC TC 115312 G:C CC×GG GC 10423 T:C CC×TT TC 55379 G:C GC×CC GC 11179 T:C TC×CC TC 61868 G:C GC×GG GC 24235 T:C TC×TT TC 117897 G:C GG×CC GC 5878 T:C TT×CC TC 31806 G:C GG×GC GC 15257 T:C TT×TC TC 76655 G:T GG×GT GT 36606 T:G GG×TG TG 47919 G:T GG×TT GT 13061 T:G GG×TT TG 23063 G:T GT×GG GT 57819 T:G TG×GG TG 26229 G:T GT×TT GT 22453 T:G TG×TT TG 51385 G:T TT×GG GT 22404 T:G TT×GG TG 12665 G:T TT×GT GT 40813 T:G TT×TG TG 31852

[0059] (6) Align the long sequences of the offspring to the fusion genome using the minimap2 program, aligning to the genomic sequence obtained in step 4 (parameters: -t 40 -c -N 2 -Y --eqx -x asm20) to obtain the aligned bam file. Locate the parental-specific SNVs obtained in step 5 in the offspring ccs sequences, and count the allele base types of these SNVs in the offspring ccs sequences. Based on the bam file generated by the alignment, use the samtools program to determine the SNVs, and determine the parental origin of each offspring sequence according to the SNV genotype distribution listed in the above table. Finally, 1,359,581 offspring sequences from the paternal source and 1,308,794 offspring sequences from the maternal source are obtained, with genome coverages of 38× and 36× respectively.

[0060] (7) Use the hifiasm program to perform splicing and assembly respectively with the offspring sequences from the two parental sources obtained in step 6 as input. The lengths of the obtained genomic sequences are 575M and 553M respectively, which are consistent with the evaluation results in step (3). Use the BUSCO program to evaluate the assembly integrity with the metazoa_odb10 database as the reference (parameters: --cpu 6 --offline --mode genome --evalue 1e - 5 --metaeuk_parameters = "--remove - tmp - files = 1"

[0061] --metaeuk_rerun_parameters = "--remove - tmp - files = 1"), the completeness of Complete BUSCOs is 96.3% and 96.1% respectively, and the values of the Complete and duplicated BUSCOs indicators are 1.1% and 1.5% respectively, indicating that two high-quality haploid genomes are obtained.

Claims

1. A method for constructing a haploid genome of oyster, characterized in that: (1) Select female and male oyster individuals, artificially construct a full-sib family and conventionally breed the offspring; (2) After the offspring individuals are cultured for one year, take one individual for long-read high-throughput sequencing; (3) Perform short-read high-throughput sequencing on the female and male parent individuals of the full-sib family respectively; (4) Use the sequencing sequences of the offspring for preliminary assembly; (5) Align the parental sequences to the preliminary assembly to identify parental-specific single nucleotide variant sites; (6) Use the parental-specific single nucleotide variant sites to group the sequencing sequences of the offspring individuals to obtain long sequences from the paternal and maternal sources; (7) Use the hifiasm program, with the two long sequences from the parental sources obtained in step (6) as input, perform splicing and assembly respectively to obtain two haploid genome sequences; In step (2), when taking one individual for long-read high-throughput sequencing, the long-read high-throughput sequencing adopts the SMART sequencing mode of PacBio company, and the bam file of the raw data is processed by the ccs program to obtain ccs sequences, and the sequencing depth is not less than 60×; In step (3), the fastq file of the raw data is filtered by the fastp program to obtain paired raw sequencing sequences, and the short-read high-throughput sequencing depth is not less than 50× for both, and then the genome size, heterozygosity and repetitive sequence ratio of the female and male parent individuals are estimated; Align the sequencing sequences of the parents in step (3) with the gene sequences obtained in step (4) respectively to obtain alignment bam files, use the sort command of the Samtools program to sort the sequences in the bam files, use the Picard program to remove the duplicate alignments caused by PCR amplification in the bam files, and use the gatk program to detect single nucleotide variants (SNVs) in the paternal and maternal bam files, and identify parental-specific single nucleotide variant sites according to the alignment information; Locate the parental-specific SNVs obtained in step (5) in the offspring ccs sequences, and count the allelic base types of the SNVs in the offspring ccs sequences. The parental-specific single nucleotide variant sites are SNV genotype combinations with different genotypes in the two parents and the presence of parental-specific alleles in the offspring. Among them, the SNV genotype combination forms are: (1) FF×MM, where F is paternal-specific and M is maternal-specific; (2) FF×FM, where the origin of F cannot be determined and M is maternal-specific; (3) FM×MM, where F is paternal-specific and the origin of M cannot be determined; Where F and M represent the four constituent bases A, C, G, and T of genomic DNA.

2. The method for constructing the oyster haplotype genome according to claim 1, characterized in that: In step (1), collect several one-year-old or two-year-old female and male oyster individuals with a shell height > 5 cm during the oyster breeding season for later use. Among them, the female and male oysters can be of the same species or different species.

3. The method for constructing a haploid genome of oyster according to claim 1, characterized in that: In step (4), use the offspring individual sequences obtained in step 2 as input for splicing, and the integrity after splicing needs to reach a Complete BUSCOs index value of not less than 85%.

Citation Information

Patent Citations

  • Fluid meter

    CA22474A

  • Genetic map construction method and device, haplotype analytical method and device

    CN102952855A

  • Methods and systems for haplotype determination

    US20140045706A1