Method for screening molecular markers based on apparent annotation information to carry out whole genome selective breeding

The cost and accuracy challenges of existing genome-wide selection breeding methods are addressed by screening functional SNPs using apparent annotation information and histone-modified regions, achieving more efficient and accurate breeding value estimation.

CN119920316AActive Publication Date: 2025-05-02SANYA INST OF OCEANOGRAPHY OCEAN UNIV OF CHINA +1

Patent Information

Application Number
CN202510408809.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-05-02
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

Existing genome-wide selection breeding methods have challenges in terms of cost and accuracy, especially in cross-group sports genome value estimation.

Method used

By screening molecular markers using apparent annotation information, combining histone-modified regions for ChromHMM analysis, functional SNPs with key histone-modified regions were screened for genome-wide selection breeding.

Benefits of technology

This method can use less number of SNPs with the same prediction accuracy, reduce genotyping costs, improve breeding efficiency, and significantly improve the accuracy of genetic breeding value estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920316A_ABST
    Figure CN119920316A_ABST
Patent Text Reader

Abstract

The invention discloses a method for screening molecular markers based on apparent annotation information to carry out whole genome selective breeding, which comprises the following steps: step A, carrying out phenotype determination and genome sequencing on growth traits of a litopenaeus vannamei breeding population to obtain phenotype data and SNP (Single Nucleotide Polymorphism) typing data; step B, performing GS analysis by using phenotype data and SNP typing data, predicting a genome breeding value of a breeding population through a mainstream breeding model, and screening out an optimal breeding model; and step C, classifying chromatin states through ChromHMM analysis by combining the optimal breeding model with the key histone modified region, and screening out the functional SNPs with the key histone modified region. The SNPs have relatively strong functional relevance, are used for different regional groups, and can provide more accurate molecular markers for growth character genetic analysis of bred animals by using less SNPs under the same prediction accuracy, so that the accuracy of genetic breeding value estimation and the cross-group effect are remarkably improved, and the breeding process is accelerated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of molecular breeding of farmed animals, and in particular relates to a method for whole genome selective breeding based on screening of molecular markers based on epigenetic annotation information. Background Art

[0002] Selective breeding is an important means of genetic improvement of farmed animals. In recent years, with the development of molecular biotechnology and the continuous reduction of high-throughput sequencing costs, single nucleotide polymorphism (SNP) molecular markers widely present in the genome have become a trend in breeding of farmed animals. Whole-genome selective breeding uses molecular markers, mainly single nucleotide polymorphism markers, in the whole genome to build statistical models for genetic breeding value estimation. At present, high-throughput typing technologies mainly include whole-genome resequencing and SNP chip technology. Whole-genome resequencing can obtain SNP site typing information in the whole genome when the sequencing depth is sufficient, but the cost is high. In actual breeding applications, in order to reduce the cost of genotyping and model calculation costs, SNP genotyping chips are usually developed by screening SNP sites related to target traits for genotyping of the test population. Although the universal high-density SNP chip has a wide genome coverage and a wide range of applications, the cost of use is still too high for production applications due to the large number of SNPs used, which seriously hinders the practical application of whole-genome selection methods. In addition, due to the changes in SNP genotype frequencies in different populations, the whole-genome selection model constructed through the training population can often only obtain good breeding value estimation results in the internal population, and the breeding value estimation across populations is often difficult to achieve satisfactory accuracy.

[0003] Litopenaeus vannamei is one of the most widely farmed aquatic species in the world and has important economic value. Growth, as a ubiquitous biological process, is one of the most concerned traits in the breeding of aquatic animals, including most other economic animals. At present, the research on genetic analysis of growth traits and genetic breeding mainly uses genome-wide association study (GWAS) to screen loci and genes related to growth traits, and genome-wide selection analysis (GS) to select breeding by using high-density SNPs covering the whole genome.

[0004] GS analysis uses high-density SNPs covering the entire genome for selective breeding. It is based on the LD state between each SNP and its adjacent quantitative trait locus (QTL) to trace the effects of all potential QTLs, and to predict individuals with unknown phenotypes while avoiding QTL positioning. Therefore, it requires a sufficiently high SNP density and genetic communication between the test population and the training population to ensure that the model breeding value prediction performance is optimal. This not only leads to high genotyping costs in subsequent applications, but also due to the randomness of the screened SNP sites (no functional association), some sites that have no actual functionality are retained in the breeding model, while other functional sites are eliminated.

[0005] Therefore, it is necessary to establish methods for screening relevant molecular markers and whole-genome selective breeding using epigenetic annotation information, to provide more accurate molecular markers for the genetic analysis of growth traits of farmed animals, significantly improve the accuracy of genetic breeding value estimation, and the effect of cross-population breeding value estimation, and lay the foundation for genetic breeding research of farmed animals and other economic traits. Summary of the invention

[0006] The purpose of the present invention is to apply histone modification information and epigenetic annotation information to whole-genome selection breeding, to ensure that the screened SNPs have strong gene expression functional relevance, and are suitable for different geographical groups, to use fewer SNPs and to improve breeding efficiency.

[0007] To solve the above technical problems: In one aspect, the present invention provides a method for molecular marker screening based on epigenetic annotation information, comprising the following steps: Step A: Perform phenotypic measurement of growth traits and genome sequencing on the breeding population of Penaeus vannamei to obtain phenotypic data and SNP typing data; Step B: GS analysis was performed using phenotypic data and SNP typing data, and the genomic breeding value of the vannamei breeding population was predicted through the mainstream breeding model to screen out the optimal breeding model; Step C: Using the optimal breeding model combined with key histone modification regions, the chromatin state is classified through ChromHMM analysis to screen out functional SNPs with key histone modification regions.

[0008] Furthermore, step A comprises: step A-1: ​​performing phenotypic measurement of growth traits on a breeding population of Penaeus vannamei to obtain phenotypic data; the growth trait is body length or weight; step A-2: extracting total DNA from the breeding population of Penaeus vannamei to obtain total DNA; constructing a DNA library using the total DNA; sequencing the DNA library to obtain SNP typing data; wherein the construction process of the DNA library comprises: breaking the total DNA into small fragments by ultrasonic treatment, and undergoing end modification, linker connection, PCR amplification and magnetic bead purification.

[0009] Furthermore, step B specifically includes: performing GS analysis using phenotypic data and SNP typing data, and predicting the genomic breeding value of the vannamei breeding population through mainstream breeding models; mainstream breeding models include GBLUP, BayesA, BayesB or BayesC models; the screening method of mainstream breeding models includes: randomly selecting different marker densities in the vannamei breeding population to systematically evaluate the prediction performance of different models at different marker densities, and selecting the mainstream breeding model with the highest prediction accuracy.

[0010] Furthermore, the key histone modification regions in step C include H3K4me1, H3K4me3, H3K27me3 and H3K27ac; the chromatin states include BivFlnk1, Quies, ReprPCWk, ReprPC, EnhWk, BivFlnk2, TxWk and TssFlnkD.

[0011] On the other hand, the present invention provides a method for whole genome selection breeding of Penaeus vannamei by using a method for molecular marker screening based on epigenetic annotation information, wherein a subset of SNP sites is selected from functional SNPs with key histone modification regions, and a genome breeding value of a breeding population of Penaeus vannamei is predicted using an optimal breeding model; a preferred genotype in the breeding population of Penaeus vannamei is determined based on the subset of SNP sites, and breeding is performed by retaining parents with the preferred genotype for pairing and breeding.

[0012] Furthermore, the mainstream breeding models are GBLUP, BayesA, BayesB or BayesC models; key histone modification regions include H3K4me1, H3K4me3, H3K27me3 and H3K27ac.

[0013] On the other hand, the present invention provides a computer device for a method for breeding Penaeus vannamei based on whole genome data, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a method for whole genome selective breeding of Penaeus vannamei by molecular marker screening based on epigenetic annotation information is implemented.

[0014] The present invention innovatively uses different histone modifications for epigenetic annotation and applies epigenetic annotation information to GS analysis. Unlike traditional breeding that randomly selects SNPs, it screens SNPs through regulatory elements. These SNPs have strong functional associations and are therefore more stable and applicable to different regional groups. They can also use fewer SNPs with the same prediction accuracy. This advantage will greatly reduce costs and improve breeding efficiency in the subsequent genotyping process. It provides more accurate molecular markers for the genetic analysis of growth traits of farmed animals, significantly improves the accuracy of genetic breeding value estimation, and the effect of cross-population breeding value estimation, lays the foundation for genetic breeding research on farmed animals and other economic traits, and accelerates the breeding process. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The above content and the following specific embodiments of the present invention will be better understood when read in conjunction with the accompanying drawings. It should be noted that the accompanying drawings are only examples of the technical solutions claimed for protection.

[0016] Figure 1 Box plot of the prediction accuracy of the breeding value of body length phenotype for different marker densities and different models of Penaeus vannamei (where the horizontal axis is the SNP marker density and the vertical axis is the prediction accuracy; purple is the GBLUP model; green is the BayesA model; orange is the BayesB model; blue is the BayesC model); Figure 2This is a chromatin state analysis diagram for the critical period of embryonic development of Penaeus vannamei (where A is the IGV map of chromatin state trajectories during embryonic development. ChromHMM learns the chromatin state definition based on these trajectories and assigns each genomic position to the corresponding state; B is the emission parameter heat map of 8 chromatin state extension models based on H3K4me1, H3K4me3, H3K27me3 and H3K27ac. The red depth of the heat map reflects the probability of histone marks appearing, and darker red indicates a higher probability of marks; C is an overlapping enrichment heat map of chromatin states and genomic annotated regions, showing the enrichment of each chromatin state in specific genomic features; D is a spatial enrichment heat map of chromatin states in the vicinity of transcription start sites (TSS), showing the enrichment pattern of chromatin states around transcription start sites; E is the annotation of each state and its abbreviation map; the annotations are: Flanking bivalent TSS / Enh 1, BivFInk1), Quiescent / low (Quies), Weak repressed Polycomb (ReprPCWk), Repressed Polycomb (ReprPC), Weak Enhancer (EnhWk), Flanking bivalent TSS / Enh 2 (BivFInk2), Weak transcription (TxWk), and Flanking TSS downstream (TssFlnkD); Figure 3 This is a box plot of the prediction accuracy of the SNP marker density in different chromosomal states during the critical period of embryonic development of Penaeus vannamei for the breeding value of the phenotypic trait of body length (wherein the horizontal axis is the SNP marker density and the vertical axis is the prediction accuracy; embryo_E1 is: SNPs located in the BivFlnk1 state; embryo_E2 is: SNPs located in the Quies state; embryo_E3 is: SNPs located in the ReprPCWk state; embryo_E4 is: SNPs located in the ReprPC state; embryo_E5 is: SNPs located in the EnhWk state; embryo_E6 is: SNPs located in the BivFlnk2 state; embryo_E7 is: SNPs located in the TxWk state; embryo_E8 is: SNPs located in the TssFlnkD state; random is a randomly selected SNP site). DETAILED DESCRIPTION

[0017] The detailed features and advantages of the present invention are described in detail in the specific implementation modes below, and the contents are sufficient to enable any person skilled in the art to understand the technical contents of the present invention and implement them accordingly. Moreover, according to the description, claims and drawings disclosed in this specification, those skilled in the art can easily understand the relevant objects and advantages of the present invention.

[0018] It should be noted that in this specification, similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings.

[0019] In order to make the purpose, technical scheme and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. The experimental methods described in the embodiments of the present invention are conventional methods unless otherwise specified, and the materials, reagents, etc. used in the following embodiments can be obtained from commercial channels unless otherwise specified.

[0020] Example A method for whole genome selective breeding by screening molecular markers based on epigenetic annotation information, comprising the following steps: S1. The phenotypic data collection of individual Litopenaeus vannamei and the extraction of total DNA from muscle tissue are as follows: (1) A total of 972 three-month-old Penaeus vannamei (purchased from Hebei Huanghua Hatchery) were randomly selected and divided into three batches (the first batch, the second batch, and the third batch). Among them, there were 313 Penaeus vannamei in the first batch, 331 in the second batch, and 328 in the third batch.

[0021] (2) The three batches of Penaeus vannamei sampled above were individually numbered, and their body length and weight were measured using an electronic vernier caliper with an accuracy of 0.02 mm and an electronic balance with an accuracy of 0.01 g (Quintix). The number of each individual Penaeus vannamei was repeated three times and the average value was taken.

[0022] Among them, the body length of the first batch of Penaeus vannamei individuals was between 160~280 mm, and the weight was between 36~108 g; the body length of the second batch of Penaeus vannamei individuals was between 160~240 mm, and the weight was between 24~84 g; the body length of the third batch of Penaeus vannamei individuals was between 120~200 mm, and the weight was between 12~48 g.

[0023] (3) DNA was extracted from the muscle of each Litopenaeus vannamei using a marine animal genomic DNA extraction kit (Tian Gen, model DP324-03, the following reagents are included in the kit).

[0024] The DNA extraction process is as follows: about 30 mg of tissue sample is added with GA buffer and Proteinase K (Protease K, used to decompose RNA and protein in tissue sample and protect DNA from damage), then treated and ground in a 56°C water bath until uniform, and RNase A is added to degrade RNA molecules in the sample. After adjusting the lysis time to optimize the DNA extraction effect, GB buffer is added for further treatment, and the DNA of muscle tissue of Penaeus vannamei is extracted using a marine animal genomic DNA extraction kit. The specific process includes tissue grinding, DNA precipitation, column elution, and DNA dissolution and elution. The total DNA of muscle tissue of Penaeus vannamei is obtained and stored at -20°C for future use.

[0025] The concentration and purity of DNA were tested by nucleic acid meter. The OD260 / OD280 of total DNA in muscle tissue of 972 individual Penaeus vannamei were between 1.8 and 2.0, indicating that total DNA with good purity was obtained. 1.0% agarose gel electrophoresis was used to detect that the total DNA extraction was complete and could be used for subsequent experiments.

[0026] S2. Construct DNA library and sequence. The specific steps are as follows: The total DNA was broken into small fragments by ultrasonic treatment, and a DNA library was constructed by end modification, adapter ligation, PCR amplification and magnetic bead purification; the DNA library was used to perform PE150 double-end sequencing on the Illumina NovaSeq 6000 platform, that is, the original data file was quality controlled using Trimmomatic v3.29 software; the reads were aligned to the genome using BWA v0.7.17 software to obtain a SAM file; the SAM format file was converted into a BAM format file using Samtools v1.17 software; site variation detection was performed using GATK v4.4.0.0 software; SNP site variation was extracted using GATK v4.4.0.0 software and VCFtools v0.1.16 software, and missing genotypes were supplemented using Beagle v5.5 software, and qualified SNP sites with MAF>0.05 and Harwin equilibrium test P value<0.0001 were screened to obtain the SNP typing data. The specific reference steps are as follows: (1) The qualified total DNA is randomly broken into small fragments by the Covaris ultrasonic processing system, and these small fragments are used for end modification. An A base is added to the 3' end of the small fragment by Klenow enzyme (T4 DNA polymerase & DNA polymerase I). After the addition of the A base, the original small fragment changes from a flat end to a sticky end, which is easier to connect to the subsequent primers and adapters; the end of the small fragment after end modification has a protruding A tail, and the adapter has a protruding T tail. T4 DNA ligase is used to connect the adapter to the two ends of the DNA fragment. After the adapter is successfully connected, the low-cycle technology is used to modify the adapter.

[0027] (2) The system after adding the adapter contains polymerase and ligase, etc., and there may be large fragments. Therefore, double screening is performed by magnetic beads to remove large fragments and impurities, so as to obtain library fragments with successful addition of adapters. The DNA fragments with adapters are amplified with primers complementary to the adapters. After PCR, magnetic bead purification is required again to separate the PCR products from impurities. The obtained PCR products are the constructed DNA library. The PCR products are quantified by Qubit DNA HS ASSAY KIT; 2100 High Sensitivity DNA Chip electrophoresis is performed to determine whether the fragment size meets the requirements of subsequent sequencing (the fragment size is generally about 400bp); the molar concentration is calculated based on the Qubit quantitative results and the fragment size detected by the 2100 chip.

[0028] (3) The PCR products obtained above were used to perform PE150 double-end sequencing on the Illumina NovaSeq platform with a sequencing depth of 2G / sample to generate raw sequencing data. The sequencing data was aligned with the unpublished reference genome of the Penaeus vannamei chromosome in our laboratory for SNP typing, and the haplotype library of our laboratory was used to fill in (fill in missing genotypes) using Beagle v5.5 software, filtering SNPs with a minimum allele frequency (MAF) greater than 0.05 and a Hawes-Wenzhou equilibrium test (HWE) P value less than 0.0001, and a total of 2,609,549 qualified SNP sites were obtained.

[0029] S2. GS analysis, the specific reference steps are as follows: (1) The heritability estimation results based on the GCTA algorithm and the 2,609,549 high-quality SNPs screened by sequencing showed that the body length and weight traits of Penaeus vannamei had medium to high heritability (body length 0.44, body weight 0.29), among which the genetic contribution of body length was significantly higher than that of body weight. The two are highly genetically correlated, and the heritability of body length is higher. Therefore, the body length phenotype was selected as the research indicator of growth traits for GS analysis (Plink v1.9 software).

[0030] The genotype and body length phenotype data of Penaeus vannamei were imported into the GBLUP model respectively, and the prediction accuracy was calculated using the sommer R package in the R language in the GBLUP model (Genomic Best Linear Unbiased Prediction, a best linear unbiased prediction method based on genomic information, which directly estimates the breeding value of individuals by constructing a genomic relationship matrix (G matrix) to replace the traditional kinship matrix (A matrix); or the prediction accuracy was calculated using the BGLR R package in the BayesA, BayesB or BayesC models.

[0031] The higher the correlation coefficient, the better the predictive ability of the method. The above prediction process randomly and uniformly extracts SNP markers of different densities, randomly extracts each density 5 times, and uses different models to perform five-fold cross-validation and repeat it 5 times to systematically evaluate the predictive performance of each model under different marker densities.

[0032] (2) The prediction results are as follows: Figure 1 As shown in the figure, with body length as the phenotype, GS analysis was performed on 2,609,549 qualified SNP sites of 972 different groups of vannamei shrimp. With the gradual increase of marker density, the prediction accuracy of the mainstream breeding models such as GBLUP, BayesA, BayesB and BayesC for the breeding value of traits with body length as the phenotype is also gradually increasing. When the marker density reaches 10k, the prediction performance of the model breeding value tends to be flat, among which BayesA has the highest prediction accuracy, and the prediction accuracy is 0.34±0.01 (mean±standard deviation) at 10k.

[0033] S3. ChromHMM analysis, the specific reference steps are as follows: (1) Well-developed Litopenaeus vannamei embryos at critical embryonic developmental stages were randomly selected for construction of the CUT&Tag library, covering a total of seven developmental stages (blastocyst stage, gastrula stage, limb bud larvae, intramembranous larvae, nauplii stage I, nauplii stage III, and nauplii stage VI).

[0034] The CUT&Tag library was constructed using the Illumina Hyperactive Universal CUT&TagAssay Kit (Vazyme, China, #TD903), and the sample was prepared using the nuclear extraction method. The basic principles and operation steps of the experiment were referred to CUT&Tag (Kaya-Okur HS et al. CUT&Tag for efficient epigenomic profiling of small samples and single cells. Nat Communications, 2019, full text), and the vannamei embryo samples were obtained for library construction.

[0035] (2) The vannamei embryo samples obtained for library construction were sequenced on the Illumina NovaSeq 6000 platform (Novogene, Tianjin). Sequencing data were obtained using Nova-PE150 paired-end sequencing (1G raw data / library). The sequencing data were aligned with the unpublished vannamei chromosome reference genome of our laboratory using the BWA-MEM algorithm (a new alignment algorithm used to align sequencing reads or assembled contigs to a large reference genome), and peak calling (a computational method used to identify regions in the genome where alignment reads obtained by sequencing are enriched) was performed using MACS2.

[0036] The key parameters include: macs2 callpeak -t A.bam -c IgG.bam -f BAMPE --outdir... / ... -B -n A -g 1.66e9 --keep-dup auto -q 0.05, the obtained bam file was subjected to ChromHMM analysis (a computational tool based on Hidden Markov Model (HMM), which is used to systematically identify and characterize chromatin states in the genome, annotate their distribution locations across the genome, and assist the biological interpretation of these states by automatically calculating the enrichment of external annotations). The key parameters include java -mx4000M -jarChromHMM.jar BinarizeBam -gzip -b 200 -f 0 -g 0 -p 0.0001 DUIXIA_chrom_sizes.txt bam sheet.txt binarization, java -mx4000M -jar ChromHMM.jarLearnModel -gzip -d 0.001 -color 129,0,0 -p 20 -i chrhmm binarizationlearnmodel 10 dx.

[0037] (3) The classification of chromatin states requires information about histone modifications, and there are many histone modifications with known functions. ChromHMM was used to analyze histone modifications in the key early embryonic development stage of Litopenaeus vannamei. Using the signal tracks marked by four histone modification regions (H3K4me1: a mark of enhancer regions, H3K4me3: commonly found in promoter regions, usually associated with gene transcription activation, H3K27me3: usually associated with gene expression repression, and H3K27ac: a mark of chromatin regions in an open state), ChromHMM first learned and defined a series of chromatin states, and then assigned each segment in the genome to one of these states. Using the LearnModel function, eight chromatin states were finally identified and named according to the ChromHMM protocol (reference: Ernst J et al. Chromatin-state discovery and genome annotation with ChromHMM. Nature protocols, 2017, full text).

[0038] It can be understood that the four histone modification regions (H3K4me1: a mark of enhancer regions, H3K4me3: commonly found in promoter regions, usually associated with gene transcription activation, H3K27me3: usually associated with gene expression inhibition, H3K27ac: a mark of chromatin regions in an open state) and their functions are conserved in farmed animals, and these four histone modification regions can be used in GS analysis of farmed animals.

[0039] It can be understood that key histone modification regions can also be other histone modification combinations, which can make chromatin more refined or obtain different annotation types.

[0040] It can be understood that the histone modification information in tissues related to the economic traits of Litopenaeus vannamei or other farmed animals can also be used to screen SNPs.

[0041] The results are as follows Figure 2 As shown in the figure, using the signal tracks of the four key histone modification regions H3K4me1, H3K4me3, H3K27me3, and H3K27ac in the genome during the critical period of embryonic development, a total of eight chromatin state information, including active transcription, enhancer, bivalent, repression, and quiescence, were annotated at the position 3506745-3731315 of chromosome 1 of Penaeus vannamei (chr1: 3506745-3731315), namely: bivalent TSS / enhancer region 1 (BivFlnk1), quiescent / low activity (Quies), weakly repressed by Polycomb (ReprPCWk), repressed by Polycomb (ReprPC), weak enhancer (EnhWk), bivalent TSS / enhancer region 2 (BivFlnk2), weak transcription (TxWk), and downstream TSS region (TssFlnkD), which can be visualized using IGV software.

[0042] S4, ChromHMM and GS joint analysis, the specific reference steps are as follows: (1) The match parameter in the R script was used to screen SNPs located in the peak interval of the ChromHMM bed file, and the --extract parameter of plink was used to retain the SNP sites in the genotype file. Then, the --thin-count parameter of plink was used to randomly extract SNP sites of different densities, and the --recodeA parameter of plink was used to convert the genotype file into a 012 format file. The prediction accuracy of the GBLUP model was calculated using the sommer R package, and the prediction accuracy of the BayesA / B / C model was calculated using the BGLR R package.

[0043] (2) The accuracy of the prediction was determined by using the sommer R package in the R language in the GBLUP model, or by using the BGLR R package in the BayesA, BayesB, or BayesC model. The higher the correlation coefficient, the better the predictive ability of the method. The above prediction process was performed by randomly and uniformly extracting SNP markers of different densities, with each density randomly extracted 5 times, and five-fold cross-validation was performed using different models and repeated 5 times to systematically evaluate the predictive performance of each model under different marker densities.

[0044] The results are as follows Figure 3 As shown, the epigenetic annotation information is used to carefully distinguish the chromatin state of the genome to screen SNPs, where the horizontal axis is the SNP marker density and the vertical axis is the prediction accuracy; embryo_E1 is: SNPs located in the BivFlnk1 state; embryo_E2 is: SNPs located in the Quies state; embryo_E3 is: SNPs located in the ReprPCWk state; embryo_E4 is: SNPs located in the ReprPC state; embryo_E5 is: SNPs located in the EnhWk state; embryo_E6 is: SNPs located in the BivFlnk2 state; embryo_E7 is: SNPs located in the TxWk state; embryo_E8 is: SNPs located in the TssFlnkD state; random is a randomly selected SNP site. As the density of SNPs located in different chromosomal states during the critical embryonic development period gradually increases, the prediction accuracy of the BayesA model for the breeding value of the phenotypic trait of body length is also gradually increasing. When the marker density is in the range of 1k~2.5k, the prediction accuracy of SNPs located on the E1, E5, E6 and E8 state Peaks is generally 6.25%~11.11% higher than that of randomly selected SNPs, and the prediction ability is more stable.

[0045] It can be concluded that the present invention innovatively utilizes different histone modifications for epigenetic annotation and applies epigenetic annotation information to GS analysis. Different from the random selection of SNPs in traditional breeding, the present invention screens SNPs through regulatory elements. These SNPs have strong functional associations and are therefore more stable and applicable to different regional groups. They can also use fewer SNPs at the same prediction accuracy. This advantage will greatly reduce costs and improve breeding efficiency in the subsequent genotyping process. At the same time, the screening method and breeding method are also applicable to other farmed animals, which also contain four histone modification regions (H3K4me1, H3K4me3, H3K27me3 and H3K27ac), such as aquatic farmed animals, such as fish, shrimps, crabs or shellfish; shrimps include Penaeidae, Neopenaeidae, Hawkclaw Shrimp, Shrimp or Sand Shrimp; crabs include Frog Crab, Mantou Crab, Jade Crab, Portunidae or Fan Crab. The method of screening molecular markers based on epigenetic annotation information can provide more accurate molecular markers for the genetic analysis of growth traits of farmed animals, significantly improve the accuracy of genetic breeding value estimation, and the effect of cross-population breeding value estimation, lay the foundation for genetic breeding research of farmed animals and other economic traits, and accelerate the breeding process.

[0046] The terms and expressions used herein are for descriptive purposes only and the present invention should not be limited to these terms and expressions. The use of these terms and expressions does not mean to exclude any equivalent features of the illustrations and descriptions (or parts thereof), and it should be recognized that various modifications that may exist should also be included in the scope of the claims. Other modifications, changes and substitutions may also exist. Accordingly, the claims should be deemed to cover all such equivalents.

[0047] Similarly, it should be pointed out that although the present invention has been described with reference to the current specific embodiments, ordinary technicians in this technical field should realize that the above embodiments are only used to illustrate the present invention, and various equivalent changes or substitutions may be made without departing from the spirit of the present invention. Therefore, as long as the changes and modifications to the above embodiments are within the scope of the essential spirit of the present invention, they will fall within the scope of the claims of the present invention.

Claims

1. A method for molecular marker screening based on epigenetic annotation information, characterized in that: The following steps are involved: Step A: Perform phenotypic measurement of growth traits and genome sequencing on the breeding population of Penaeus vannamei to obtain phenotypic data and SNP typing data; Step B: using the phenotypic data and the SNP typing data to perform GS analysis, predicting the genomic breeding value of the vannamei breeding population through a mainstream breeding model, and screening out the optimal breeding model; Step C: using the optimal breeding model in combination with key histone modification regions to classify chromatin states through ChromHMM analysis, and screening out functional SNPs with key histone modification regions.

2. The method for molecular marker screening based on epigenetic annotation information according to claim 1, characterized in that: The step A comprises: Step A-1: ​​performing phenotypic measurement of growth traits on the breeding population of Penaeus vannamei to obtain the phenotypic data; the growth trait is body length or weight; Step A-2: extracting total DNA from the Litopenaeus vannamei breeding population to obtain total DNA; constructing a DNA library using the total DNA; and sequencing the DNA library to obtain the SNP typing data; The construction process of the DNA library includes: breaking the total DNA into small fragments by ultrasonic treatment, and undergoing end modification, linker connection, PCR amplification and magnetic bead purification.

3. The method for molecular marker screening based on epigenetic annotation information according to claim 1, characterized in that: The step B is specifically as follows: GS analysis is performed using the phenotypic data and the SNP typing data, and the genomic breeding value of the vannamei shrimp breeding population is predicted by a mainstream breeding model; the mainstream breeding model includes a GBLUP, BayesA, BayesB or BayesC model; The screening method of the mainstream breeding model comprises: randomly selecting different marker densities in the vannamei shrimp breeding population, systematically evaluating the prediction performance of different models at different marker densities, and selecting the mainstream breeding model with the highest prediction accuracy.

4. The method for molecular marker screening based on epigenetic annotation information according to claim 1, characterized in that: The key histone modification regions in step C include H3K4me1, H3K4me3, H3K27me3 and H3K27ac; The chromatin states include BivFlnk1, Quies, ReprPCWk, ReprPC, EnhWk, BivFlnk2, TxWk and TssFlnkD.

5. A method for whole genome selective breeding of Penaeus vannamei using the method for molecular marker screening based on epigenetic annotation information as claimed in claim 1, characterized in that: Selecting a subset of SNP sites from functional SNPs with key histone modification regions, and predicting the genomic breeding value of the Litopenaeus vannamei breeding population using the optimal breeding model; The preferred genotype in the breeding population of Penaeus vannamei is determined according to the SNP locus subset, and breeding is performed by retaining parents with the preferred genotype for pairing and breeding.

6. The method for whole genome selective breeding of Penaeus vannamei according to the method for molecular marker screening based on epigenetic annotation information according to claim 5, characterized in that: The mainstream breeding model is GBLUP, BayesA, BayesB or BayesC model; The key histone modification regions include H3K4me1, H3K4me3, H3K27me3 and H3K27ac.

7. A computer device for a method of breeding Penaeus vannamei based on whole genome data, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a method for whole genome selective breeding of Penaeus vannamei is implemented according to the method for molecular marker screening based on epigenetic annotation information as described in claim 5 or 6.

Citation Information

Patent Citations

  • Method for improving aquatic animal whole-genome selective breeding efficiency

    CN110867208A

  • Data processing method and system for identifying enhancer and super enhancer

    CN115083517A

  • Method for excavating sheep skeletal muscle satellite cell proliferation and differentiation related genes based on Hi-C technology and multi-omics combined analysis

    CN118620995A

  • Methods and compositions for improving fertility

    WO2023213272A1

Cited By

  • Selective breeding method for prawns with high genotype deletion rate and storage medium

    CN121549303A

  • A method for breeding and selecting shrimp pairs for high genotype deletion rate and storage medium

    CN121549303B