A method for whole-genome selective breeding based on epigenetic annotation information screening of molecular markers
By applying epigenetic annotation information and histone modification regions to screen SNPs in whole-genome selective breeding, the problems of high cost and low accuracy in whole-genome selective breeding are solved, and more efficient and accurate breeding results are achieved, which is suitable for vannamei shrimp and other farmed animals.
Patent Information
- Application Number
- CN202510408809.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2045-04-02
AI Technical Summary
Existing technologies in whole-genome selective breeding have problems such as high genotyping costs and low accuracy in estimating cross-population breeding values. In particular, in the genetic analysis of growth traits of farmed animals such as the vannamei shrimp, the SNPs selected by traditional methods lack functional relevance, resulting in high costs and insufficient accuracy.
A molecular marker screening method based on epigenetic annotation information is used, combined with histone modification regions, to screen out functional SNPs through ChromHMM analysis, and the breeding value is predicted using the optimal breeding model, reducing the number of SNPs required and improving breeding efficiency and accuracy.
It significantly improves the accuracy of genetic breeding value estimation and the effect of cross-population breeding value estimation, reduces the cost of genotyping, improves breeding efficiency, and is suitable for farmed animals in different geographical groups.
Smart Images

Figure CN119920316B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of molecular breeding of farmed animals, and particularly relates to a method for whole genome selection breeding based on screening of molecular markers according to epigenetic annotation information. BACKGROUND
[0002] Selection breeding is an important means of genetic improvement of farmed animals. In recent years, with the development of molecular biology technology and the continuous reduction of high-throughput sequencing cost, assisted breeding through the widely existing single nucleotide polymorphism (SNP) molecular markers in the genome has become a trend in the breeding of farmed animals. Whole genome selection breeding estimates genetic breeding value through molecular markers in the whole genome, mainly single nucleotide polymorphism markers, and constructs a statistical model. At present, high-throughput genotyping technologies mainly include whole genome resequencing and SNP chip technology. Whole genome resequencing can obtain SNP site genotyping information in the whole genome when the sequencing depth is sufficient, but the cost is high. In order to reduce the cost of genotyping and model calculation in actual breeding application, SNP genotyping chips are usually developed by screening SNP sites related to target traits for genotyping of the test population. Although the universal high-density SNP chip has wide genome coverage and wide application range, the use cost is still too high for production application due to the large number of SNPs used, which seriously hinders the practical application of whole genome selection method. In addition, due to the change of SNP genotype frequency in different populations, the whole genome selection model constructed by training population can only obtain good breeding value estimation effect in the internal population, and the breeding value estimation in cross-population is often difficult to achieve satisfactory accuracy.
[0003] Litopenaeus vannamei is one of the most widely farmed aquatic species in the world, with important economic value. Growth, as a ubiquitous biological process, is one of the most concerned traits in the breeding of aquatic farming animals, including other most economic animals. At present, the research on genetic analysis and genetic breeding of growth traits mainly uses genome-wide association study (GWAS) to screen sites and genes related to growth traits, and genomic selection (GS) to use high-density SNPs covering the whole genome for selection breeding.
[0004] GS analysis uses high-density SNPs covering the whole genome for selection breeding, which is based on each SNP and the quantitative trait locus (QTL) adjacent to it in LD state to trace the effect of all potential QTLs, to realize the prediction of unknown phenotype individuals without QTL positioning, so as to ensure that the model breeding value prediction performance is optimal. This not only leads to higher genotyping cost in subsequent application, but also due to the randomness (non-functional association) of the screened SNP sites, a part of the sites without actual functionality is retained in the breeding model, while another part of the functional sites is excluded.
[0005] Therefore, it is necessary to establish a method for screening related molecular markers and whole genome selection breeding using epigenetic annotation information, to provide more accurate molecular markers for genetic analysis of growth traits of farmed animals, significantly improve the accuracy of genetic breeding value estimation, and cross-population breeding value estimation effect, and lay a foundation for genetic breeding research of farmed animals and other economic traits. SUMMARY
[0006] The purpose of the present application is to apply histone modification information and epigenetic annotation information to whole genome selection breeding, to ensure that the screened SNPs have strong gene expression functional association, and are suitable for different regional populations, to improve breeding efficiency with fewer SNPs.
[0007] To solve the above technical problems:
[0008] In one aspect, the present application provides a method for screening molecular markers based on epigenetic annotation information, comprising the following steps:
[0009] Step A: Phenotype determination and genome sequencing of Litopenaeus vannamei breeding population for growth traits, to obtain phenotype data and SNP typing data;
[0010] Step B: GS analysis using phenotype data and SNP typing data, predicting the genomic breeding value of Litopenaeus vannamei breeding population by mainstream breeding model, and screening the optimal breeding model;
[0011] Step C: Using the optimal breeding model combined with the key histone modification region to classify the chromatin state by ChromHMM analysis, and screening the functional SNPs with the key histone modification region.
[0012] Furthermore, step A includes: step A-1: performing phenotypic measurement of growth traits on a breeding population of Penaeus vannamei to obtain phenotypic data; the growth trait is body length or weight; step A-2: extracting total DNA from the breeding population of Penaeus vannamei to obtain total DNA; constructing a DNA library using the total DNA; sequencing the DNA library to obtain SNP typing data; wherein the construction process of the DNA library includes: breaking the total DNA into small fragments through ultrasonic treatment, and performing end modification, linker ligation, PCR amplification and magnetic bead purification.
[0013] Furthermore, step B specifically includes: using phenotypic data and SNP typing data to perform GS analysis, and predicting the genomic breeding value of the vannamei breeding population through mainstream breeding models; mainstream breeding models include GBLUP, BayesA, BayesB or BayesC models; the screening method of mainstream breeding models includes: randomly sampling different marker densities in the vannamei breeding population, systematically evaluating the prediction performance of different models at different marker densities, and selecting the mainstream breeding model with the highest prediction accuracy.
[0014] Furthermore, the key histone modification regions in step C include H3K4me1, H3K4me3, H3K27me3 and H3K27ac; the chromatin states include BivFlnk1, Quies, ReprPCWk, ReprPC, EnhWk, BivFlnk2, TxWk and TssFlnkD.
[0015] Another aspect of the present invention provides a method for genome-wide selective breeding of Penaeus vannamei by using a method for molecular marker screening based on epigenetic annotation information, wherein a subset of SNP sites is selected from functional SNPs in key histone modification regions, and the genomic breeding value of a breeding population of Penaeus vannamei is predicted using an optimal breeding model; a preferred genotype in the breeding population of Penaeus vannamei is determined based on the SNP site subset, and breeding is performed by retaining parents with the preferred genotype for pairing and reproduction.
[0016] Furthermore, the mainstream breeding models are GBLUP, BayesA, BayesB or BayesC models; key histone modification regions include H3K4me1, H3K4me3, H3K27me3 and H3K27ac.
[0017] Another aspect of the present invention provides a computer device for a method for breeding Penaeus vannamei based on whole genome data, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, a method for whole genome selection breeding of Penaeus vannamei is implemented by molecular marker screening based on epigenetic annotation information.
[0018] The present application carries out epigenetic annotation by innovatively using different histone modifications, and applies the epigenetic annotation information to GS analysis, which is different from the traditional breeding of randomly selecting SNPs, and the SNPs are screened by regulatory elements, and the SNPs have strong functional correlation, so that the stability is higher, and the SNPs are suitable for different regional populations, and the same prediction accuracy can be achieved by using less SNPs, and the advantage is that the cost is greatly reduced and the breeding efficiency is improved in the subsequent genotyping process. The present application provides more accurate molecular markers for genetic analysis of growth traits of farmed animals, significantly improves the accuracy of genetic breeding value estimation, and improves the estimation effect of cross-population breeding value, lays a foundation for genetic breeding research of farmed animals and other economic traits, and speeds up the breeding process. BRIEF DESCRIPTION OF DRAWINGS
[0019] The above and the following detailed description of the present application can be better understood when read in conjunction with the accompanying drawings. It should be noted that the drawings are only examples of the claimed technical solutions.
[0020] Figure 1 The prediction accuracy box plot of the body length as a phenotype of the breeding value of different marker densities and different models of Litopenaeus vannamei (wherein, the horizontal coordinate is the SNP marker density, and the vertical coordinate is the prediction accuracy; purple is the GBLUP model; green is the BayesA model; orange is the BayesB model; blue is the BayesC model);
[0021] Figure 2Figure S2. Chromatin state analysis of Litopenaeus vannamei at key periods of embryonic development (A) IGV plot of the chromatin state trajectories of Litopenaeus vannamei at key periods of embryonic development. ChromHMM learned chromatin state definitions from these trajectories and assigned each genomic location to a corresponding state. (B) Heatmap of emission parameters for the 8-state extension model based on H3K4mel, H3K4me3, H3K27me3, and H3K27ac. The intensity of the red color in the heatmap reflects the probability of the histone mark appearing, with darker red indicating a higher probability of the mark. (C) Overlap Enrichment Heatmap of chromatin states with annotated regions of the genome, showing the enrichment of each chromatin state in specific genomic features. (D) Spatial Enrichment Heatmap of chromatin states in the region adjacent to the transcription start site (TSS), showing the enrichment pattern of chromatin states around the transcription start site. (E) Annotation of each state and its abbreviation. The annotations are: Flanking bivalent TSS / Enh 1 (BivFInk1), Quiescent / low (Quies), Weak repressed Polycomb (ReprPCWk), repressed Polycomb (ReprPC), Weak Enhancer (EnhWk), Flanking bivalent TSS / Enh 2 (BivFInk2), Weak transcription (TxWk), and Flanking TSS downstream (TssFlnkD).
[0022] Figure 3 Figure S3. Prediction accuracy of SNP marker density in different chromatin states at key periods of embryonic development for the body length phenotype (A) SNP marker density in BivFlnk1 state at key periods of embryonic development. (B) SNP marker density in Quies state at key periods of embryonic development. (C) SNP marker density in ReprPCWk state at key periods of embryonic development. (D) SNP marker density in ReprPC state at key periods of embryonic development. (E) SNP marker density in EnhWk state at key periods of embryonic development. (F) SNP marker density in BivFlnk2 state at key periods of embryonic development. (G) SNP marker density in TxWk state at key periods of embryonic development. (H) SNP marker density in TssFlnkD state at key periods of embryonic development. (I) Randomly selected SNP sites. DETAILED DESCRIPTION
[0023] The detailed features and advantages of the present invention are described in detail below in the specific embodiments, and the content is sufficient to enable any person skilled in the art to understand the technical content of the present invention and implement it accordingly. Based on the description, claims and drawings disclosed in this specification, those skilled in the art can easily understand the relevant purposes and advantages of the present invention.
[0024] It should be noted that in this specification, similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0025] To make the objectives, technical solutions, and advantages of the present invention more apparent, embodiments of the present invention will be described in further detail below with reference to the accompanying drawings. The experimental methods described in the examples of the present invention are conventional methods unless otherwise specified. The materials, reagents, etc. used in the following examples are all commercially available unless otherwise specified.
[0026] Example
[0027] A method for whole-genome selective breeding based on epigenetic annotation information screening of molecular markers, comprising the following steps:
[0028] S1. Collection of phenotypic data and extraction of total DNA from muscle tissue of individual Litopenaeus vannamei. Specific steps are as follows:
[0029] (1) A total of 972 three-month-old Penaeus vannamei shrimp (purchased from Hebei Huanghua Hatchery) were randomly selected and divided into three batches (the first batch, the second batch, and the third batch). The first batch consisted of 313 Penaeus vannamei shrimp, the second batch consisted of 331 Penaeus vannamei shrimp, and the third batch consisted of 328 Penaeus vannamei shrimp.
[0030] (2) The three batches of Penaeus vannamei sampled above were individually numbered, and the body length and weight of Penaeus vannamei were measured using an electronic vernier caliper with an accuracy of 0.02 mm and an electronic balance with an accuracy of 0.01 g (Quintix, Sartorius). Each individual Penaeus vannamei was counted three times and the average value was taken.
[0031] Among them, the body length of the first batch of Penaeus vannamei individuals ranged from 160 to 280 mm, and the weight ranged from 36 to 108 g; the body length of the second batch of Penaeus vannamei individuals ranged from 160 to 240 mm, and the weight ranged from 24 to 84 g; the body length of the third batch of Penaeus vannamei individuals ranged from 120 to 200 mm, and the weight ranged from 12 to 48 g.
[0032] (3) Using marine animal genomic DNA extraction kit (Tiangen, DP324-03 type, the reagents are all provided in the kit) to extract DNA from the muscle of each individual white shrimp.
[0033] The DNA extraction process is as follows: about 30 mg of tissue sample is added to GA buffer and Proteinase K (proteinase K is used to decompose RNA and protein in the tissue sample, while protecting DNA from damage), then treated in a 56℃ water bath and ground to uniformity, and RNase A is added to degrade RNA molecules in the sample. After adjusting the lysis time to optimize the DNA extraction effect, GB buffer is added for further treatment, and the marine animal genomic DNA extraction kit is used to extract the white shrimp muscle tissue DNA, the specific process includes tissue grinding, DNA precipitation, column washing and elution, and DNA dissolution and elution, etc. Steps to obtain white shrimp muscle tissue total DNA, stored at -20℃ for standby.
[0034] The concentration and purity of the DNA are detected by a nucleic acid detector, and the OD260 / OD280 of the muscle tissue total DNA of 972 white shrimp individuals is between 1.8 and 2.0, that is, the total DNA with good purity is obtained, and the total DNA extraction is complete, which can be used for subsequent experiments, detected by 1.0% agarose gel electrophoresis.
[0035] S2, constructing a DNA library and sequencing, the specific steps are as follows:
[0036] The total DNA is broken into small fragments by ultrasonic treatment, and is subjected to end modification, linker connection, PCR amplification and magnetic bead purification to construct a DNA library; PE150 double-end sequencing is performed on the Illumina NovaSeq 6000 platform using the DNA library, that is, the original data file is quality controlled using Trimmomatic v3.29 software; reads are aligned to the genome using BWA v0.7.17 software to obtain a SAM file; the SAM format file is converted to a BAM format file using Samtools v1.17 software; site variation detection is performed using GATK v4.4.0.0 software; SNP site variation is extracted using GATK v4.4.0.0 software and VCFtools v0.1.16 software, missing genotypes are filled using Beagle v5.5 software, and qualified SNP sites with MAF>0.05 and Hardy-Weinberg equilibrium test P value<0.0001 are screened out to obtain the SNP typing data, the specific steps are as follows:
[0037] (1) The qualified total DNA is randomly broken into small fragments by a Covaris ultrasonic treatment system, and end modification is performed using these small fragments. A Klenow enzyme (T4 DNA polymerase & DNA polymerase I) is used to add an A base to the 3' end of the small fragments. After the A base is added, the original small fragments change from blunt ends to sticky ends, which makes it easier to connect the subsequent primers and adaptors. The end of the small fragments after end modification has a protruding A tail, while the adaptor has a protruding T tail. T4 DNA ligase is used to connect the adaptor to both ends of the DNA fragments. After successful connection of the adaptor, low-cycle technology is used to modify the adaptor.
[0038] (2) The above system with added adaptors contains polymerase and ligase, and there may be large fragments. Therefore, double screening is performed by magnetic beads to remove large fragments and impurities, thereby obtaining library fragments with successfully added adaptors. The DNA fragments with added adaptors are amplified using primers complementary to the adaptors. After PCR, magnetic bead purification is performed again to separate the PCR product from impurities, and the obtained PCR product is the constructed DNA library. The PCR product is quantified by Qubit DNA HS ASSAY KIT; 2100 High Sensitivity DNA Chip electrophoresis is performed to determine whether the fragment size meets the subsequent sequencing requirements (the fragment size is generally about 400 bp); and the molar concentration is calculated based on the Qubit quantification result and the fragment size detected by the 2100 chip.
[0039] (3) The above obtained PCR product is used for PE150 double-end sequencing on the Illumina NovaSeq platform, with a sequencing depth of 2G / sample, to generate raw sequencing data. The sequencing data is compared with the unpublished Litopenaeus vannamei mounted chromosome reference genome in the laboratory to perform SNP typing, and Beagle v5.5 software is used to fill in the missing genotypes using the haplotype library in the laboratory. The minimum allele frequency MAF is greater than 0.05, and the HWE P value of the SNPs is less than 0.0001, and a total of 2609549 qualified SNP sites are obtained.
[0040] S2, GS analysis, the specific steps are as follows:
[0041] (1) Based on the GCTA algorithm and the heritability estimation results of the 2,609,549 high-quality SNPs screened by sequencing, the body length and weight traits of Litopenaeus vannamei have medium to high heritability (body length 0.44, weight 0.29), with the genetic contribution of body length significantly higher than that of weight. The two traits are highly genetically correlated, and body length has a higher heritability. Therefore, the body length phenotype was selected as the research indicator of growth traits for subsequent GS analysis (Plink v1.9 software).
[0042] The genotype and body length phenotype data of Litopenaeus vannamei were imported into the GBLUP model, and the prediction accuracy of the sommer R package in the R language was used in the GBLUP model (Genomic Best Linear Unbiased Prediction, a best linear unbiased prediction method based on genomic information, which directly estimates the breeding value of individuals by constructing a genomic relationship matrix (G matrix) instead of the traditional kinship matrix (A matrix); or the prediction accuracy was used in the BGLR R package in the BayesA, BayesB, or BayesC model.
[0043] The higher the correlation coefficient, the better the predictive ability of the method. The above prediction process is carried out by randomly and uniformly extracting SNP markers of different densities, randomly extracting each density 5 times, and using different models for five-fold cross-validation and repeating it 5 times to systematically evaluate the predictive performance of each model under different marker densities.
[0044] (2) The prediction results are as follows Figure 1 As shown in the figure, GS analysis was performed on 2,609,549 qualified SNP sites of 972 different groups of vannamei shrimp using body length as the phenotype. With the gradual increase of marker density, the prediction accuracy of the mainstream breeding models such as GBLUP, BayesA, BayesB and BayesC for the breeding value of traits with body length as the phenotype also gradually increased. When the marker density reached 10k, the prediction performance of the model breeding value tended to be flat, among which BayesA had the highest prediction accuracy, with a prediction accuracy of 0.34±0.01 (mean±standard deviation) at 10k.
[0045] S3. ChromHMM analysis. The specific steps are as follows:
[0046] (1) Well-developed Litopenaeus vannamei embryos at the critical embryonic developmental stage were randomly selected to construct a CUT&Tag library, covering a total of seven developmental stages (blastula stage, gastrula stage, limb bud larvae, intramembranous larvae, nauplii stage I, nauplii stage III and nauplii stage VI).
[0047] The construction of CUT&Tag library adopts Illumina Hyperactive Universal CUT&Tag Assay Kit (Vazyme, China, #TD903), the sample preparation adopts nuclear extraction method, the basic principle and operation steps of the experiment refer to CUT&Tag (Kaya-Okur HS et al. CUT&Tag for efficient epigenomic profiling of small samples and single cells. Nat Communications, 2019, full text), and the library building shrimp embryo sample is obtained.
[0048] (2) The library building shrimp embryo sample obtained above is sequenced on an Illumina NovaSeq 6000 platform (Novogene, Tianjin), Nova-PE150 double-end sequencing (1G raw data / library) is adopted to obtain sequencing data, the sequencing data is aligned with the unpublished laboratory Litopenaeus vannamei mounted chromosome reference genome by using a BWA-MEM algorithm (a new alignment algorithm for aligning sequencing reads or assembled contigs to a large reference genome), and MACS2 is used for peak calling (a calculation method for identifying the region of the genome enriched by the aligned reads obtained by sequencing).
[0049] Key parameters include: macs2 callpeak -t A.bam -c IgG.bam -f BAMPE --outdir... / ... -B -n A -g 1.66e9 --keep-dup auto -q 0.05, the obtained bam file is subjected to ChromHMM analysis (a computational tool based on Hidden Markov Model (HMM), which is used to systematically identify and characterize chromatin states in the genome, annotate the distribution position in the genome range, and assist the biological interpretation of these states by automatically calculating the enrichment of external annotations), key parameters include: java -mx4000M -jar ChromHMM.jar BinarizeBam -gzip -b 200 -f 0 -g 0 -p 0.0001 DUIXIA_chrom_sizes.txt bam sheet.txt binarization, java -mx4000M -jar ChromHMM.jar LearnModel -gzip -d 0.001 -color 129,0,0 -p 20 -i chrhmm binarization learnmodel 10 dx.
[0050] (3) The classification of chromatin states requires information of histone modification, and there are many known functional histone modifications at present. By performing ChromHMM analysis on the histone modification of the key early embryonic development stage of Marsupenaeus japonicus, using the signal tracks of four histone modification regions (H3K4me1: marker of enhancer region, H3K4me3: commonly found in promoter region, usually related to gene transcription activation, H3K27me3: usually related to gene expression inhibition, H3K27ac: marker of open state chromatin region), ChromHMM first learns and defines a series of chromatin states, and then allocates each segment in the genome to one of these states. Using the LearnModel function, 8 kinds of chromatin states are finally identified, and are named according to the ChromHMM protocol (reference: Ernst J et al. Chromatin-state discovery and genome annotation with ChromHMM. Nature protocols, 2017, full text).
[0051] It can be understood that the four histone modification regions (H3K4me1: a mark of enhancer regions, H3K4me3: commonly found in promoter regions, usually associated with gene transcription activation, H3K27me3: usually associated with gene expression repression, H3K27ac: a mark of chromatin regions in an open state) and their functions are conserved in farmed animals, and these four histone modification regions can be used in GS analysis of farmed animals.
[0052] It can be understood that key histone modification regions can also be other histone modification combinations, which can make chromatin more refined or obtain different annotation types.
[0053] It can be understood that the histone modification information in tissues related to the economic traits of the breeding targets of Litopenaeus vannamei or other farmed animals can also be used to screen SNPs.
[0054] The results are as follows Figure 2 As shown in the figure, using the signal tracks of the four key histone modification regions H3K4me1, H3K4me3, H3K27me3, and H3K27ac in the genome during the critical period of embryonic development, eight chromatin state information, including active transcription, enhancer, bivalent, repression, and quiescence, were annotated at the 3506745-3731315 position of chromosome 1 of the shrimp Litopenaeus vannamei (chr1: 3506745-3731315), namely: bivalent TSS / enhancer region 1 (BivFlnk1), quiescent / low activity (Quies), weakly repressed by Polycomb (ReprPCWk), repressed by Polycomb (ReprPC), weak enhancer (EnhWk), bivalent TSS / enhancer region 2 (BivFlnk2), weakly transcribed (TxWk), and downstream TSS region (TssFlnkD), which can be visualized using IGV software.
[0055] S4, ChromHMM and GS joint analysis, the specific reference steps are as follows:
[0056] (1) The match parameter in the R script was used to select SNPs located in the peak interval of the ChromHMM bed file. The --extract parameter of plink was used to retain the SNP sites in the genotype file. Then, the --thin-count parameter of plink was used to randomly extract SNP sites of different densities. The --recodeA parameter of plink was used to convert the genotype file into a 012 format file. The prediction accuracy of the GBLUP model was calculated using the sommer R package, and the prediction accuracy of the BayesA / B / C model was calculated using the BGLR R package.
[0057] (2) The accuracy of the prediction was evaluated using the sommer R package in the GBLUP model, or using the BGLR R package in the BayesA, BayesB, or BayesC model. The higher the correlation coefficient, the better the predictive ability of the method. The above prediction process was performed by randomly and uniformly sampling SNP markers of different densities, with each density randomly sampled 5 times, and using different models for five-fold cross-validation and repeating it 5 times to systematically evaluate the prediction performance of each model under different marker densities.
[0058] The results are as follows Figure 3 As shown in the figure, the epigenetic annotation information is used to carefully distinguish the chromatin states of the genome to screen SNPs, where the horizontal axis is the SNP marker density and the vertical axis is the prediction accuracy; embryo_E1 is: SNPs located in the BivFlnk1 state; embryo_E2 is: SNPs located in the Quies state; embryo_E3 is: SNPs located in the ReprPCWk state; embryo_E4 is: SNPs located in the ReprPC state; embryo_E5 is: SNPs located in the EnhWk state; embryo_E6 is: SNPs located in the BivFlnk2 state; embryo_E7 is: SNPs located in the TxWk state; embryo_E8 is: SNPs located in the TssFlnkD state; random is a randomly selected SNP site. As the density of SNPs located in different chromosomal states during critical embryonic developmental periods gradually increases, the prediction accuracy of the BayesA model for breeding values of body length phenotypes also gradually increases. When the marker density is in the range of 1k~2.5k, the prediction accuracy of SNPs located on the E1, E5, E6 and E8 state peaks is generally 6.25%~11.11% higher than that of randomly selected SNPs, and the prediction ability is more stable.
[0059] Therefore, it can be concluded that the present application uses different histone modifications for epigenetic annotation innovatively, and applies the epigenetic annotation information to the GS analysis, which is different from the traditional breeding of randomly selecting SNPs, and the SNPs are screened by the regulatory elements, the SNPs have strong functional correlation, and thus are more stable, and are suitable for different regional populations, and can use less SNPs quantity under the same prediction accuracy, and the advantage will greatly reduce the cost and improve the breeding efficiency in the subsequent genotyping process. At the same time, the screening method and the breeding method are also suitable for other domestic animals, and the four histone modification regions (H3K4me1, H3K4me3, H3K27me3 and H3K27ac) are also contained in other domestic animals, for example, aquatic domestic animals, the aquatic domestic animals are fish, shrimp, crab or shellfish, etc.; the shrimps include penaeid shrimps, new penaeid shrimps, chelaeoconaridae, shrimps or sand shrimps, etc.; the crabs include frog crabs, monitocarcinidae, leucosidae, portunidae or xanthidae, etc. The method of screening molecular markers based on epigenetic annotation information can provide more accurate molecular markers for genetic analysis of growth traits of domestic animals, significantly improve the accuracy of genetic breeding value estimation and cross-population breeding value estimation effect, lay a foundation for genetic breeding research of domestic animals and other economic traits, and accelerate the breeding process.
[0060] The terms and expressions herein have been chosen to best describe the application, it is not intended that the application be limited thereto as set forth herein. Rather, it is intended to cover all alternatives, modifications, and equivalents falling within the spirit, or the scope of the application.
[0061] Also, it is pointed out that, although the present application has been described with reference to specific exemplifying embodiments, it is clear that the ordinary skilled person in the art will be able to devise many equivalents and / or alternatives to those embodiments without feeling departing from the scope of the present application. Therefore, the scope of the present application is not limited to the embodiments described above, but is defined by the appended claims.
Claims
1. A method for molecular marker screening based on epigenetic annotation information, characterized in that: The following steps are involved: Step A: Phenotyping of growth traits and genome sequencing of breeding populations of farmed animals to obtain phenotypic data and SNP typing data; Step B: performing GS analysis using the phenotypic data and the SNP typing data, predicting the genomic breeding value of the breeding population using a mainstream breeding model, and screening out the optimal breeding model; Step C: using the optimal breeding model in combination with key histone modification regions to classify chromatin states through ChromHMM analysis, and screening functional SNPs with key histone modification regions; The key histone modification regions include H3K4me1, H3K4me3, H3K27me3 and H3K27ac; The chromatin states include BivFlnk1, Quies, ReprPCWk, ReprPC, EnhWk, BivFlnk2, TxWk and TssFlnkD.
2. The method for molecular marker screening based on epigenetic annotation information according to claim 1, characterized in that: The step A comprises: Step A-1: performing phenotypic measurement of growth traits on the breeding population of farmed animals to obtain phenotypic data; the growth traits are body length or weight; Step A-2: extracting total DNA from the breeding population of farmed animals to obtain total DNA; constructing a DNA library using the total DNA; and sequencing the DNA library to obtain the SNP typing data; The construction process of the DNA library includes: breaking the total DNA into small fragments through ultrasonic treatment, and then undergoing end modification, linker ligation, PCR amplification and magnetic bead purification.
3. The method for molecular marker screening based on epigenetic annotation information according to claim 1, characterized in that: The step B is specifically as follows: Performing GS analysis using the phenotypic data and the SNP typing data to predict the genomic breeding value of the breeding population of the farmed animals using a mainstream breeding model; the mainstream breeding model includes a GBLUP, BayesA, BayesB or BayesC model; The screening method of the mainstream breeding model comprises: randomly selecting different marker densities in the breeding population of farmed animals to systematically evaluate the prediction performance of different models at different marker densities, and selecting the mainstream breeding model with the highest prediction accuracy.
4. A method for whole genome selective breeding of farmed animals using the method for molecular marker screening based on epigenetic annotation information as claimed in claim 1, characterized in that: A subset of SNP sites is selected from functional SNPs with key histone modification regions, and the genomic breeding value of the breeding population of farmed animals is predicted using the optimal breeding model; the preferred genotype in the breeding population of farmed animals is determined based on the subset of SNP sites, and breeding is performed by retaining parents with the preferred genotype for pairing and breeding.
5. The method for whole genome selective breeding of farmed animals according to claim 4, wherein: The mainstream breeding model is GBLUP, BayesA, BayesB or BayesC model.
6. A computer device for a method of breeding farmed animals based on whole genome data, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for whole genome selective breeding of farmed animals by performing molecular marker screening based on epigenetic annotation information as described in claim 4 or 5 is implemented.
Citation Information
Patent Citations
Method for improving aquatic animal whole-genome selective breeding efficiency
CN110867208A