A method for screening molecular markers related to economic traits of Litopenaeus vannamei and breeding using epigenetic annotation information

By combining histone modification information and GWAS analysis, SNPs significantly related to the growth trait of vannabinoid shrimp were screened, solving the problems of high genotyping costs and poor accuracy of cross-group sports value estimation in the prior art, and achieving more accurate genetic breeding value estimation and more efficient breeding process.

CN119899901BActive Publication Date: 2025-06-13SANYA INST OF OCEANOGRAPHY OCEAN UNIV OF CHINA +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510408469.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-06-13
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

When the prior art carries out genetic improvement of breeding animals, the genotyping costs are high, the model calculation costs are high, and the accuracy of the breeding value estimation of the whole genome selection model in cross-population is difficult to guarantee.

Method used

By using histone modification information and apparent annotation information, combined with GWAS analysis, SNPs significantly related to the growth trait of vannabinoid shrimp were screened out, and the expression regulatory elements of their corresponding genes were accurately located to provide more accurate molecular markers.

Benefits of technology

It significantly improves the accuracy of genetic breeding value estimation, especially the effect of breeding value estimation across populations, reduces the occurrence of false positive and false negative results, and thus accelerates the breeding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119899901B_ABST
    Figure CN119899901B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for screening molecular markers related to economic traits of Litopenaeus vannamei and breeding by using epigenetic annotation information, comprising the following steps: Step A: Conduct phenotypic determination of growth traits and genomic sequencing on a Litopenaeus vannamei breeding population to obtain phenotypic data and SNP genotyping data; Step B: Perform GWAS analysis by using the phenotypic data and SNP genotyping data to obtain the P value of each SNP; Step C: Use the P value of each SNP through the Bonferroni multiple correction method, with a P value < 1.9E-08 as the judgment of significant loci; screen out SNPs significantly related to growth traits; use key histone modification regions to accurately locate SNPs significantly related to growth traits and the expression regulatory elements of their corresponding genes. By applying epigenetic annotation information to GWAS analysis, it provides more accurate molecular markers for the genetic analysis of growth traits of Litopenaeus vannamei, lays a foundation for the genetic breeding research of other economic traits, and accelerates the breeding process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of molecular breeding of Litopenaeus vannamei, and particularly relates to a method for screening molecular markers related to economic traits of Litopenaeus vannamei and breeding by using epigenetic annotation information. Background Art

[0002] Selective breeding is an important means for genetic improvement of farmed animals. In recent years, with the development of molecular biotechnology and the continuous reduction of the cost of high-throughput sequencing, assisted breeding through single nucleotide polymorphism (SNP) molecular markers widely present in the genome has become a trend in the selection and breeding of farmed animals. Genome-wide selection breeding estimates genetic breeding values by constructing a statistical model through molecular markers across the whole genome, mainly single nucleotide polymorphism markers. Currently, the high-throughput genotyping technologies are mainly whole-genome resequencing and SNP chip technology. Whole-genome resequencing can obtain SNP locus genotyping information across the whole genome when the sequencing depth is sufficient, but the cost is relatively high. In practical breeding applications, in order to reduce the genotyping cost and model calculation cost, SNP genotyping chips are usually developed by screening SNP loci related to target traits for genotyping of the tested population. Although the general high-density SNP chip has a wide genome coverage and a large applicable range, due to the large number of SNPs used, the usage cost is still too high for production applications, seriously hindering the practical application of the genome-wide selection method. In addition, due to the change of SNP genotype frequencies in different populations, the genome-wide selection model constructed through the training population can often only obtain good breeding value estimation effects within the internal population, and it is often difficult to achieve satisfactory accuracy in breeding value estimation across populations.

[0003] Litopenaeus vannamei is one of the most widely farmed aquatic species globally and has important economic value. Growth, as a common biological process, is one of the most concerned traits in the breeding of aquaculture animals, including most other economic animals. Currently, the research on genetic analysis and genetic breeding of growth traits mainly uses genome-wide association studies (GWAS) to screen loci and genes related to growth traits, and genomic selection (GS) to conduct selective breeding by using high-density SNPs covering the whole genome.

[0004] GWAS analysis searches for molecular marker loci (such as SNPs) and genotypes that are significantly associated with specific phenotypes through the genotype data and corresponding phenotype data of a large number of individuals, and finds the strong linkage regions of significant SNPs based on Linkage Disequilibrium (LD). The genes located within the regions are used as candidate genes related to traits. However, due to the existence of LD, when a locus functionally related to a trait is tightly linked to another locus unrelated to the trait, it is often difficult to determine which locus is the real cause of trait differences, which may lead to false positive and false negative results in GWAS gene mapping studies. GS analysis uses high-density SNPs covering the whole genome for selective breeding. It traces the effects of all potential QTLs based on the fact that each SNP and the adjacent Quantitative Trait Locus (QTL) are in the LD state, and realizes the prediction of individuals with unknown phenotypes without QTL mapping. Therefore, it requires a sufficiently high density of SNPs and genetic exchange between the test population and the training population to ensure that the prediction performance of the model breeding value reaches the optimal. This not only results in a relatively high genotyping cost in subsequent applications, but also due to the randomness (no functional relevance) of the selected SNP loci, some loci without actual functionality are retained in the breeding model, while some functional loci are excluded.

[0005] Therefore, it is necessary to establish a method for screening relevant molecular markers and breeding using epigenetic annotation information, provide more accurate molecular markers for the genetic analysis of growth traits of Litopenaeus vannamei, significantly improve the accuracy of genetic breeding value estimation, and the estimation effect of cross-population breeding value, and lay a foundation for the genetic breeding research of other economic traits. Summary of the Invention

[0006] The object of the present invention is to apply histone modification information and epigenetic annotation information to molecular marker screening and breeding, provide more accurate molecular markers for the genetic analysis of growth traits of Litopenaeus vannamei aquaculture animals, significantly improve the accuracy of genetic breeding value estimation, and the estimation effect of cross-population breeding value, and lay a foundation for the genetic breeding research of aquaculture animals and other economic traits.

[0007] To solve the above technical problems:

[0008] On the one hand, the present invention provides a method for screening molecular markers related to the economic traits of Litopenaeus vannamei, comprising the following steps: Step A: performing phenotypic determination of growth traits and genomic sequencing on a Litopenaeus vannamei breeding population to obtain phenotypic data and SNP genotyping data; Step B: performing GWAS analysis using the phenotypic data and the SNP genotyping data to obtain the P-value of each SNP; Step C: using the P-value of each SNP through the Bonferroni multiple correction method, with a P-value < 1.9E-08 as the judgment of significant loci, screening out SNPs significantly related to growth traits; precisely localizing the SNPs significantly related to growth traits and the expression regulatory elements of their corresponding genes using key histone modification regions.

[0009] Further, Step A includes: Step A-1: performing phenotypic determination of growth traits on the Litopenaeus vannamei breeding population to obtain the phenotypic data; the growth traits are body length or body weight; Step A-2: extracting total DNA from the Litopenaeus vannamei breeding population to obtain total DNA; constructing a DNA library using the total DNA; sequencing the DNA library to obtain the SNP genotyping data; wherein, the construction process of the DNA library includes: fragmenting the total DNA into small fragments by ultrasonic treatment, and performing end modification, adapter ligation, PCR amplification and magnetic bead purification.

[0010] Further, Step C includes: Step C-1: sorting the P-values of each SNP in ascending order, through the Bonferroni multiple correction method, with a P-value < 1.9E-08 as the judgment of significant loci; screening out SNPs significantly related to growth traits; Step C-2: using CUT&Tag to screen the key histone modification regions, and precisely localizing the SNPs significantly related to growth traits and the expression regulatory elements of their corresponding genes using the key histone modification regions; the key histone modification regions are H3K4me1, H3K4me3, H3K27me3 and H3K27ac.

[0011] On the other hand, the present invention provides a method for breeding Litopenaeus vannamei based on whole-genome data, comprising the following steps: Step A: performing phenotypic determination and genome sequencing on the breeding population of Litopenaeus vannamei to obtain phenotypic data and SNP genotyping data; Step B: performing GWAS analysis using the phenotypic data and the SNP genotyping data to obtain the P-value of each SNP in the GWAS analysis; Step C: using the P-value of each SNP through the Bonferroni multiple correction method, and taking P-value < 1.9E-08 as the judgment of significant loci, screening out SNPs significantly related to growth traits; precisely locating the SNPs significantly related to growth traits and the expression regulatory elements of their corresponding genes in the key histone modification regions; Step D: determining the preferred genotypes in the breeding population of Litopenaeus vannamei according to the subset of SNP loci, and breeding by retaining the parent individuals with the preferred genotypes for paired reproduction.

[0012] On the other hand, the present invention provides a computer device for a method of breeding Litopenaeus vannamei based on whole-genome data, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the method of breeding Litopenaeus vannamei based on whole-genome data.

[0013] On the other hand, the present invention provides an application of an SNP molecular marker in the assisted breeding of Litopenaeus vannamei. The sequence of the SNP molecular marker is as shown in SEQ ID NO.1; the SNP locus is located at the 233rd position from the 5'-end of the sequence shown in SEQ ID NO.1; the polymorphism of the SNP molecular marker is of the A / G type; the application in the assisted breeding of Litopenaeus vannamei is the screening of the body length or body weight of Litopenaeus vannamei.

[0014] By innovatively using different histone modifications for epigenetic annotation, obtaining cis-regulatory element information and applying it to GWAS analysis, the present invention can obtain the expression regulatory element information of SNPs significantly related to growth traits, which is beneficial to a more in-depth study of its molecular mechanism. It provides more accurate molecular markers for the genetic analysis of the growth traits of Litopenaeus vannamei, lays a foundation for the genetic breeding research of other economic traits, and accelerates the breeding process. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The above content of the present invention and the following specific embodiments will be better understood when read in conjunction with the accompanying drawings. It should be noted that the drawings are only examples of the claimed technical solutions.

[0016] Figure 1Manhattan plot of genome-wide association analysis for body length trait of Litopenaeus vannamei (where the black line represents the significance threshold, and loci with P < 1.9E-08 are considered significantly associated after Bonferroni correction multiple test correction);

[0017] Figure 2 Manhattan plot of genome-wide association analysis for body length trait of Litopenaeus vannamei (where the black line represents the significance threshold, and loci with P < 1.9E-08 are considered significantly associated after Bonferroni correction multiple test correction);

[0018] Figure 3 Venn diagram of significant loci for genome-wide association analysis of body length / weight traits of Litopenaeus vannamei (the blue area length_SNPs is the number of SNP loci for body length trait; the pink area weigth_SNPs is the number of SNP loci for body weight trait; the intersection of the blue area and the pink area is the number of SNP loci for body length / weight traits);

[0019] Figure 4 Figure of significant loci and key genes related to growth traits modified by H3K4me3 at key stages of individual embryonic development of Litopenaeus vannamei (where the scale of the pfdn6-like gene is 2 kb; the black box line marks the SNP locus 44_1839577; the square brackets mark the H3K4me3 histone modification at the promoter; blas is the blastula stage; gast is the gastrula stage; lbe is the limb bud larva stage; lim is the intra-membrane larva stage; N1 is nauplius stage I; N3 is nauplius stage III; N6 is nauplius stage VI; Z1 is zoea stage I; Z3 is zoea stage III; M1 is mysis stage I; M3 is mysis stage III; the control is IgG);

[0020] Figure 5Conservative analysis of significant loci of Litopenaeus vannamei in this population and other populations (where A is the statistical chart of body length trait and genotype of SNP locus (44_1839577) in this population; B is the statistical chart of body weight trait and genotype of SNP locus (44_1839577) in this population; C is the statistical chart of body length trait and genotype of SNP locus (44_1839577) in other populations; D is the statistical chart of body weight trait and genotype of SNP locus (44_1839577) in other populations; the unit of body length is mm; the unit of body weight is g; group1 is the first batch of population; group2 is the second batch of population; group3 is the third batch of population; "*" represents p value < 0.05, with differences; "**" represents p value < 0.01, with significant differences; "****" represents p value < 0.001, with extremely significant differences; AA genotype is 0 / 0; AG genotype is 0 / 1; TAB represents other populations);

[0021] Figure 6 Box plot of the prediction accuracy of body length phenotypic breeding values of Litopenaeus vannamei with different marker densities and different models (where the abscissa is the SNP marker density and the ordinate is the prediction accuracy; purple is the GBLUP model; green is the BayesA model; orange is the BayesB model; blue is the BayesC model);

[0022] Figure 7Chromatin state analysis diagrams for key stages of Litopenaeus vannamei embryo development (wherein, A is the IGV diagram of chromatin state trajectories during embryo development. ChromHMM learns chromatin state definitions based on these trajectories and assigns each genomic position to the corresponding state; B is the heatmap of emission parameters of an 8-chromatin state expansion model based on H3K4me1, H3K4me3, H3K27me3, and H3K27ac. The depth of red in the heatmap reflects the probability of histone marks appearing, and a deeper red indicates a higher probability of the mark; C is the heatmap of the overlap enrichment of chromatin states with genomic annotation regions, showing the enrichment of each chromatin state in specific genomic features; D is the spatial enrichment heatmap of chromatin states in the region adjacent to the transcription start site (TSS), showing the enrichment pattern of chromatin states around the transcription start site; E is the diagram of the annotation and its abbreviation for each state, respectively annotated as: Flanking bivalent TSS / Enh 1 (BivFInk1), Quiescent / low (Quies), Weak repressed Polycomb (ReprPCWk), repressed Polycomb (ReprPC), Weak Enhancer (EnhWk), Flanking bivalent TSS / Enh 2 (BivFInk2), Weak transcription (TxWk), and Flanking TSS downstream (TssFlnkD)).

[0023] Figure 8 Box plot of the prediction accuracy of SNP marker density in different chromosome states of Litopenaeus vannamei at key stages of embryo development for the trait breeding value with body length as the phenotype (wherein, the abscissa is the SNP marker density, and the ordinate is the prediction accuracy; embryo_E1 are SNPs located in the BivFlnk1 state; embryo_E2 are SNPs located in the Quies state; embryo_E3 are SNPs located in the ReprPCWk state; embryo_E4 are SNPs located in the ReprPC state; embryo_E5 are SNPs located in the EnhWk state; embryo_E6 are SNPs located in the BivFlnk2 state; embryo_E7 are SNPs located in the TxWk state; embryo_E8 are SNPs located in the TssFlnkD state; random are randomly selected SNP sites). Detailed implementation mode

[0024] The detailed features and advantages of the present invention are described in detail in the following specific embodiments. The content is sufficient for any person skilled in the art to understand the technical content of the present invention and implement it accordingly. Based on the specification, claims and drawings disclosed in this specification, those skilled in the art can easily understand the related objectives and advantages of the present invention.

[0025] It should be noted that in this specification, similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0026] To make the objectives, technical solutions and advantages of the present invention clearer, the following will further describe in detail the embodiments of the present invention in conjunction with the drawings. The experimental methods described in the embodiments of the present invention are all conventional methods unless otherwise specified. The materials, reagents, etc. used in the following embodiments can all be obtained from commercial channels unless otherwise specified.

[0027] Example 1

[0028] A method for screening molecular markers related to economic traits of Litopenaeus vannamei by combining information on key histone modification regions with GWAS analysis, comprising the following steps:

[0029] S1. Collection of phenotypic data of Litopenaeus vannamei individuals and extraction of total DNA from muscle tissues. The specific reference steps are as follows:

[0030] (1) Randomly select a total of 972 Litopenaeus vannamei (purchased from a hatchery in Huanghua, Hebei) at 3 months old, and divide the 972 Litopenaeus vannamei individuals into three batch populations (the first batch, the second batch, and the third batch). Among them, there are 313 Litopenaeus vannamei individuals in the first batch, 331 Litopenaeus vannamei individuals in the second batch, and 328 Litopenaeus vannamei individuals in the third batch.

[0031] (2) Number the three batches of sampled Litopenaeus vannamei individuals, and use an electronic vernier caliper with a precision of 0.02 mm and an electronic balance with a precision of 0.01 g (Sartorius, Quintix) to measure the body length and weight data of Litopenaeus vannamei. After each Litopenaeus vannamei individual is counted 3 times repetitively, take the average value.

[0032] Among them, the body length of the Litopenaeus vannamei individuals in the first batch is between 160 and 280 mm, and the weight is between 36 and 108 g; the body length of the Litopenaeus vannamei individuals in the second batch is between 160 and 240 mm, and the weight is between 24 and 84 g; the body length of the Litopenaeus vannamei individuals in the third batch is between 120 and 200 mm, and the weight is between 12 and 48 g.

[0033] (3) Extract DNA from the muscle of each individual Litopenaeus vannamei using a Marine Animal Genomic DNA Extraction Kit (Tiangen, DP324-03 type, and the following reagents are all included in the kit). The DNA extraction process is as follows: Take about 30 mg of tissue sample, add GA buffer and Proteinase K (Proteinase K, which is used to decompose RNA and proteins in the tissue sample while protecting DNA from damage), then process and grind it evenly in a 56 °C water bath, and add RNase A to degrade RNA molecules in the sample. After adjusting the lysis time to optimize the DNA extraction effect, add GB buffer for further treatment, and use the Marine Animal Genomic DNA Extraction Kit to extract DNA from the muscle tissue of Litopenaeus vannamei. The specific process includes steps such as tissue grinding, DNA precipitation, column elution, and DNA dissolution and elution, to obtain the total DNA of the muscle tissue of Litopenaeus vannamei, and store it at -20 °C for later use. Then, use a nucleic acid analyzer to detect the concentration and purity of the DNA. The OD 260 / OD 280 of the total DNA of the muscle tissue of 972 individual Litopenaeus vannamei are all between 1.8 and 2.0, that is, total DNA with good purity is obtained, and it is detected by 1.0% agarose gel electrophoresis that the total DNA extraction is complete and can be used for subsequent experiments.

[0034] S2. Construct a DNA library and sequence

[0035] Fragment the total DNA into small fragments by ultrasonic treatment, and construct a DNA library through end modification, adapter ligation, PCR amplification, and magnetic bead purification; use the DNA library for PE150 paired-end sequencing on the Illumina NovaSeq 6000 platform, that is, use Trimmomatic v3.29 software to perform quality control on the original data file; use BWA v0.7.17 software to align the reads to the genome to obtain a SAM file; use Samtools v1.17 software to convert the SAM format file into a BAM format file; use GATK v4.4.0.0 software for site variant detection; use GATK v4.4.0.0 software and VCFtools v0.1.16 software to extract SNP site variants, use Beagle v5.5 software to fill in the missing genotypes, and screen out qualified SNP sites with MAF > 0.05 and P value of Hardy-Weinberg equilibrium test < 0.0001 to obtain the SNP genotyping data. The specific reference steps are as follows:

[0036] (1) Randomly fragment the above-mentioned qualified total DNA by Covaris ultrasonic treatment system, and use these small fragments for end modification. Add an A base to the 3' end of the small fragments through Klenow enzyme (T4 DNA polymerase & DNA polymerase I). After adding the A base, the original small fragments change from blunt ends to sticky ends, making it easier to ligate the subsequent primers and adapters. The ends of the small fragments after end modification have protruding A tails, while the adapters have protruding T tails. Use T4 DNA ligase to ligate the adapters to both ends of the DNA fragments. After successful adapter ligation, use low-cycle technology to modify the adapters.

[0037] (2) The system after adding adapters above contains polymerases and ligases, etc., and there may be large fragments. Therefore, perform double screening with magnetic beads to remove large fragments and impurities, thereby obtaining library fragments with successfully added adapters. For the DNA fragments with added adapters, use primers complementary to the adapters for amplification. After PCR, magnetic bead purification needs to be performed again to separate the PCR products from impurities. The obtained PCR products are the constructed DNA libraries. Quantify the PCR products using Qubit DNA HS ASSAY KIT; perform 2100 High Sensitivity DNA Chip electrophoresis to determine whether the fragment size meets the requirements for subsequent sequencing (the fragment size is generally about 400 bp); calculate the molar concentration based on the Qubit quantification results and the fragment sizes detected by the 2100 chip.

[0038] (3) Use the obtained PCR products for PE150 paired-end sequencing on the Illumina NovaSeq platform, with a sequencing depth of 2G / sample to generate raw sequencing data. Align the sequenced data with the unpublished Litopenaeus vannamei chromosome-mounted reference genome in this laboratory for SNP genotyping, and use the Beagle v5.5 software to fill in (complete the missing genotypes) using the haplotype library in this laboratory. Filter SNPs with a minor allele frequency MAF greater than 0.05 and a Hardy-Weinberg equilibrium test HWE P value less than 0.0001, and a total of 2,609,549 qualified SNP sites are obtained.

[0039] S3, GWAS analysis, the specific reference steps are as follows:

[0040] (1)Perform GWAS analysis on the 2,609,549 qualified SNP loci screened from the individuals of Litopenaeus vannamei in the above 972 different batches of populations through sequencing. Import the genotype and phenotype data of Litopenaeus vannamei respectively, and use the key parameters gemma -bfile DX -gk 2 -p phe.txt -o length, gemma -bfile DX -k length.sXX.txt -lmm 1 -p length.txt -c group.txt -o length_GWAS to perform GWAS analysis.

[0041] (2)According to the P - value of each SNP obtained from the GWAS analysis, in each SNP analysis data, sort the P - values of each SNP in ascending order. After multiple test correction by Bonferroni correction in traditional GWAS analysis, loci with P - value < 1.9E - 08 are considered to be significantly correlated. Among the 2,609,549 qualified SNP loci, taking body weight as the phenotype, a total of 883 significantly correlated SNP loci were screened (as Figure 1 shown); taking body length as the phenotype, a total of 493 significantly correlated SNP loci were screened (as Figure 2 shown). Among them, the significant loci of the body - length phenotype are all within the significant loci of the body - weight phenotype, that is, the two coincide (as Figure 3 shown). Therefore, subsequent analysis was performed on the 883 SNP loci significantly correlated with body weight.

[0042] S4. The histone modification information is applied to the GWAS analysis. The specific reference steps are as follows:

[0043] (1)Through CUT&Tag experiments and data processing (using fastp software to perform quality control on the original data files; using BWA v0.7.17 software to align reads to the genome to obtain a SAM file; using awk to perform fragment length filtering on the generated SAM file and retain the alignment results with the insert fragment length between 10 and 1000 bp; using Samtools v1.17 software to convert the SAM format file to a BAM format file; using MACS2 v2.1.4 software to perform PeakCalling analysis on the BAM format alignment file, using the IgG sample as a control and setting the significance threshold as q < 0.05; using IGV v2.19.1 software to perform visualization analysis on specific genes), the obtained peak files of different histone modifications, including peak chromosome, start site, and end site information, and the genotype bim file includes chromosome and SNP site information. First, use the command match <- peak[site,.(chr, peak.start, peak.end,peak.id), on =.(chr = chr,peak.start <= pos, peak.end >= pos)] in the data.table R package to screen SNPs located in different histone modification peak intervals.

[0044] The annotation results of SNP sites with key histone modification regions are shown in Table 1. Among them, among the 883 SNP sites, there are 43 significant sites located in H3K4me1 peaks, accounting for 4.87% of the total number of significant sites; there are 47 significant sites located in H3K4me3 peaks, accounting for 4.87% of the total number of significant sites; there are 25 significant sites located in H3K27me3 peaks, accounting for 2.83% of the total number of significant sites; there are 39 significant sites located in H3K27ac peaks, accounting for 4.42% of the total number of significant sites. The total number of SNP sites located in histone modifications is 104, accounting for 11.78% of the total number of significant sites. These SNPs located in histone modification regions mean they are located in genomic regulatory element regions, such as active or silent enhancer regions and promoter regions.

[0045] Table 1 Significant sites located in histone modification regions

[0046]

[0047] (2)Further screen SNPs and key genes significantly related to growth traits, and the results are as Figure 4As shown, taking the SNP locus 44_1839577 (SEQ ID NO.1, the 233rd position from the 5' end, polymorphism type A / G) significantly related to growth traits in the H3K4me3 histone modification region as an example, this locus is located in the pfdn6-like promoter region modified by H3K4me3. This promoter region (including this locus) is modified by H3K4me3 from the blastocyst stage to the mysis larva stage, proving that H3K4me3 is a marker of gene transcription activation. And PFDN6 (Prefoldin Subunit 6) has an important function in the cytoplasmic folding of actin and tubulin monomers during cytoskeleton assembly and is closely related to growth. Therefore, histone modification information can be used to effectively identify SNPs located in genomic regulatory regions and reduce the probabilities of false positives and false negatives.

[0048] (3) Further verify the conservation of the SNP locus 44_1839577 in this population and other populations (TAB). Conduct statistical analysis on the phenotypic data (body length, body weight) and SNP genotyping data of this SNP locus (44_1839577). Since the number of individuals with the GG genotype screened this time is too small and the GG genotype is not the preferred genotype, it is not included in the significant comparison.

[0049] The statistical analysis results of this population are shown in Table 2, Table 3, and Figure 5 as shown in A and B in it. The AA genotype is 0 / 0, and the AG genotype is 0 / 1. It can be seen that among the three batches of Litopenaeus vannamei in this population, the number of Litopenaeus vannamei individuals with the AA genotype is greater than that of those with the AG genotype.

[0050] In the first batch: The body length of the AA genotype is 224.35 ± 1.14 mm, and the body length of the AG genotype is 214.26 ± 1.83 mm, that is, there is a highly significant difference between the two; the body weight of the AA genotype is 69.13 ± 0.84 g, and the body weight of the AG genotype is 61.54 ± 1.33 g, and there is a highly significant difference between the two. In the second batch: The body length of the AA genotype is 201.04 ± 0.83 mm, and the body length of the AG genotype is 197.87 plus or minus 1.18 mm, and there is a significant difference between the two; the body weight of the AA genotype is 49.95 ± 0.51 g, and the body weight of the AG genotype is 47.24 ± 0.80 g, and there is a significant difference between the two. In the third batch: The body length of the AA genotype is 169.95 ± 0.71 mm, and the body length of the AG genotype is 169.79 ± 1.48 mm, and there is a difference between the two; the body weight of the AA genotype is 31.11 ± 0.34 g, and the body weight of the AG genotype is 31.11 ± 0.70 g, and there is a difference between the two.

[0051] Table 2 Statistical table of SNP genotyping results of growth traits of this population

[0052]

[0053] Table 3 Statistical table of phenotypic data corresponding to SNP genotyping of growth traits of this population

[0054]

[0055] The statistical analysis results of other populations (TAB) are shown in Table 4 and Figure 5 as shown in C and D below. The AA genotype is 0 / 0, and the AG genotype is 0 / 1. It can be seen that among the Litopenaeus vannamei in other populations, the number of Litopenaeus vannamei individuals with the AA genotype is 17, and the number of Litopenaeus vannamei individuals with the AG genotype is 14. The body length of the AA genotype is 113.22 ± 0.15 mm, and the body length of the AG genotype is 104.66 ± 0.22 mm, showing a significant difference between the two; the weight of the AA genotype is 10.55 ± 0.37 g, and the weight of the AG genotype is 8.67 ± 10.42 g, showing a significant difference between the two.

[0056] Table 4 Statistical table of phenotypic data corresponding to SNP genotyping of growth traits of other populations

[0057]

[0058] In both this population and other populations (TAB), it was found that the individuals with the AA genotype at the SNP locus 44_1839577 were larger than those with the AG genotype, indicating that the method combining histone modification and GWAS can effectively identify SNPs located in the genomic regulatory region and reduce the occurrence probabilities of false positives and false negatives.

[0059] Example 2

[0060] A method for breeding Litopenaeus vannamei by combining epigenetic annotation information and GS analysis, comprising the following steps:

[0061] S1. Collection of phenotypic data of Litopenaeus vannamei individuals and extraction of total DNA from muscle tissues, same as step S1 in Example 1, to obtain total DNA of Litopenaeus vannamei muscle tissues, which is stored at -20°C for later use. And the concentration and purity of the DNA are detected by a nucleic acid detector. The OD 260 / OD 280 of the total DNA of 972 Litopenaeus vannamei individuals are all between 1.8 and 2.0, that is, total DNA with good purity is obtained, and the integrity of the total DNA extraction is detected by 1.0% agarose gel electrophoresis, which can be used for subsequent experiments.

[0062] S2. Construct a DNA library and sequence it. The procedure is the same as step S2 in Example 1. Align the sequencing data with the unpublished reference genome of Litopenaeus vannamei assembled into chromosomes in our laboratory for SNP genotyping, and use the Beagle v5.5 software to impute (fill in the missing genotypes) using the haplotype library in our laboratory. Filter SNPs with a minor allele frequency (MAF) less than 0.05 and a Hardy-Weinberg equilibrium test (HWE) P-value less than 0.0001. A total of 2,609,549 qualified SNP loci are obtained.

[0063] S2. GS analysis. The specific steps are as follows:

[0064] (1) The heritability estimation results based on the GCTA algorithm and the 2,609,549 high-quality SNPs screened by sequencing above show that the body length and body weight traits of Litopenaeus vannamei have medium to high heritability (body length 0.44, body weight 0.29). Among them, the genetic contribution of body length is significantly higher than that of body weight. The two are highly genetically correlated, and at the same time, the heritability of body length is higher. Therefore, the body length phenotype is selected as the research index for growth traits for GS analysis (Plink v1.9 software) in the following.

[0065] Import the genotype and body length phenotype data of Litopenaeus vannamei into the GBLUP model respectively, and use the sommer R package in R language in the GBLUP model (Genomic Best Linear Unbiased Prediction, an optimal linear unbiased prediction method based on genomic information, which directly estimates the breeding value of an individual by constructing a genomic relationship matrix (G matrix) to replace the traditional kinship matrix (A matrix)) to predict the accuracy; or use the BGLR R package in the BayesA, BayesB or BayesC model to predict the accuracy. The higher the correlation coefficient, the better the prediction ability of the method. The above prediction process randomly and evenly extracts SNP markers with different densities. Each density is randomly extracted 5 times, and different models are used for five-fold cross-validation and repeated 5 times to systematically evaluate the prediction performance of each model under different marker densities.

[0066] (2) The prediction results are as Figure 6 shown. Taking body length as the phenotype, GS analysis is performed on 2,609,549 qualified SNP loci of 972 Litopenaeus vannamei from different populations. As the marker density gradually increases, the prediction accuracy of the breeding values of the traits with body length as the phenotype by the mainstream breeding models such as GBLUP, BayesA, BayesB, and BayesC also gradually increases. When the marker density reaches 10k, the prediction performance of the model breeding value tends to level off. Among them, the prediction accuracy of BayesA is the highest, and the prediction accuracy at 10k is 0.34 ± 0.01 (mean ± standard deviation).

[0067] S3. ChromHMM analysis, the specific reference steps are as follows:

[0068] (1) Randomly select well-developed Litopenaeus vannamei embryos at key embryonic development stages to construct CUT&Tag libraries, covering a total of seven developmental stages (blastula stage, gastrula stage, limb bud larva, intra-membrane larva, nauplius I stage, nauplius III stage, and nauplius VI stage). The CUT&Tag library was constructed using the Illumina Hyperactive Universal CUT&Tag Assay Kit (Vazyme, China, #TD903). The sample preparation used the nuclear extraction method. The basic principles and operation steps of the experiment refer to CUT&Tag (Kaya-Okur HS et al. CUT&Tag for efficient epigenomic profiling of small samples and single cells. 《Nat Communications》, 2019, full text) to obtain the Litopenaeus vannamei embryo samples for library construction.

[0069] (2)Sequencing was completed on the Illumina NovaSeq 6000 platform (Novogene, Tianjin) using the obtained Litopenaeus vannamei embryo samples for library construction. Nova-PE150 paired-end sequencing (1G raw data / library) was used to obtain sequencing data. The sequencing data was aligned with the unpublished Litopenaeus vannamei chromosome-level reference genome of our laboratory using the BWA-MEM algorithm (a new alignment algorithm for aligning sequencing reads or assembled contigs to a large reference genome), and MACS2 was used for peak calling (a computational method for identifying regions in the genome enriched with aligned reads obtained from sequencing). Key parameters include: macs2 callpeak -t A.bam -c IgG.bam -f BAMPE --outdir... / ... -B -n A -g 1.66e9 --keep-dup auto -q 0.05. The obtained bam file was subjected to chromHMM analysis (a computational tool based on the Hidden Markov Model (HMM) that is used to systematically identify and characterize chromatin states in the genome, annotate their distribution positions genome-wide, and assist in the biological interpretation of these states by automatically calculating the enrichment of external annotations). Key parameters include java -mx4000M -jar ChromHMM.jar BinarizeBam -gzip -b 200 -f 0 -g 0 -p 0.0001 DUIXIA_chrom_sizes.txt bam sheet.txt binarization, java -mx4000M -jar ChromHMM.jar LearnModel -gzip -d 0.001 -color 129,0,0 -p 20 -i chrhmm binarization learnmodel 10 dx.

[0070] (3) The classification of chromatin states requires the use of information on histone modifications, and there are currently many histone modifications with known functions. By performing ChromHMM analysis on histone modifications in the key early embryonic development stages of Litopenaeus vannamei, using the signal tracks marked by four histone modifications (H3K4me1: a marker of enhancer regions, H3K4me3: commonly found in promoter regions, H3K27me3: usually associated with gene expression repression, H3K27ac: a marker of open chromatin region maps), ChromHMM first learned and defined a series of chromatin states, and then assigned each segment in the genome to one of these states. Using the LearnModel function, 8 chromatin states were finally identified and named according to the ChromHMM protocol (reference: Ernst J et al. Chromatin-state discovery and genome annotation with ChromHMM. 《Nature protocols》, 2017, full text).

[0071] The results are as Figure 7 shown. Using the signal tracks of the four key histone modification regions of H3K4me1, H3K4me3, H3K27me3, and H3K27ac in the genome during the key period of embryonic development, at the position of 3506745 - 3731315 on chromosome 1 of Litopenaeus vannamei (chr1: 3506745 - 3731315), a total of eight chromatin state information, namely active transcription, enhancer, bivalent, repressed, and quiescent, were annotated, respectively: bivalent TSS / enhancer region 1 (BivFlnk1), quiescent / low activity (Quies), weakly repressed by Polycomb (ReprPCWk), repressed by Polycomb (ReprPC), weak enhancer (EnhWk), bivalent TSS / enhancer region 2 (BivFlnk2), weak transcription (TxWk), and downstream TSS region (TssFlnkD), can be visualized using the IGV software.

[0072] S4. Joint analysis of ChromHMM and GS, the specific reference steps are as follows:

[0073] (1)Use the match parameter in the R script to screen SNPs located in the peak regions of the ChromHMM bed file, use the --extract parameter of plink to retain the SNP sites in the genotype file, then use the --thin-count parameter of plink to randomly extract SNP sites with different densities, and use the --recodeA parameter of plink to convert the genotype file into a 012 format file. The sommer R package in R is used for the prediction accuracy of the GBLUP model, and the BGLR R package is used for the prediction accuracy of the BayesA / B / C models.

[0074] (2)And use the sommer R package in the R language in the GBLUP model to predict the accuracy, or use the BGLR R package in the BayesA, BayesB or BayesC models to predict the accuracy. The higher the correlation coefficient, the better the prediction ability of the method. In the above prediction process, SNP markers with different densities are randomly and evenly extracted. Each density is randomly extracted 5 times, and different models are used for five-fold cross-validation and repeated 5 times to systematically evaluate the prediction performance of each model under different marker densities.

[0075] The results are as Figure 8 shown. By using the epigenetic annotation information to carefully distinguish the genomic chromatin states to screen SNPs, as the marker density of SNPs in different chromosomal states during the key embryonic development period gradually increases, the prediction accuracy of the BayesA model for the breeding value of the body length phenotype also gradually increases. When the marker density is in the range of 1k - 2.5k, for the SNPs located on the Peaks of the E1, E5, E6 and E8 states, their prediction accuracy is generally 6.25% - 11.11% higher than that of randomly extracted SNPs, and the prediction ability is more stable.

[0076] Therefore, it can be concluded that the GWAS method provided by the present invention for screening significant SNPs and gene mapping, applying histone modification information to the genome-wide association analysis of the growth traits of Litopenaeus vannamei, is more conducive to screening SNPs and gene mapping related to traits, can effectively identify SNP sites located in the genomic regulatory regions, and obtain the expression regulatory element information of key SNPs related to traits, reduce the probability of false positives and false negatives, and is more conducive to a deeper study of its molecular mechanism. It provides more accurate molecular markers for the genetic analysis of the growth traits of Litopenaeus vannamei, lays a foundation for the genetic breeding research of Litopenaeus vannamei and other economic traits, and accelerates the breeding process.

[0077] The terms and expressions used herein are for descriptive purposes only, and the present invention should not be limited to these terms and expressions. The use of these terms and expressions does not mean excluding any equivalent features of the illustration and description (or parts thereof), and it should be recognized that various modifications that may exist should also be included within the scope of the claims. Other modifications, variations, and substitutions may also exist. Accordingly, the claims should be regarded as covering all such equivalents.

[0078] Similarly, it should be noted that although the present invention has been described with reference to current specific embodiments, those of ordinary skill in the art should recognize that the above embodiments are only used to illustrate the present invention, and various equivalent changes or substitutions can be made without departing from the spirit of the present invention. Therefore, as long as the changes and modifications of the above embodiments are within the scope of the spirit of the present invention, they will fall within the scope of the claims of the present invention.

Claims

1. A method for screening molecular markers related to economic traits of Penaeus vannamei, characterized in that: The following steps are involved: Step A: Perform phenotypic measurement of growth traits and genome sequencing on the breeding population of Penaeus vannamei to obtain phenotypic data and SNP typing data; Step B: performing GWAS analysis using the phenotypic data and the SNP typing data to obtain a P value for each SNP; Step C: Using the P value of each SNP after the Bonferroni multiple correction method, taking the P value <1.9E-08 as the significant site judgment, screen out SNPs significantly associated with growth traits; using CUT&Tag to screen key histone modification regions, and using the key histone modification regions to accurately locate SNPs significantly associated with growth traits and expression regulatory elements of their corresponding genes; the key histone modification regions are H3K4me1, H3K4me3, H3K27me3 and H3K27ac; The sequence of the SNP molecular marker screened by the screening method is shown in SEQ ID NO.1; the SNP site is located at the 233rd position from the 5' end of the sequence shown in SEQ ID NO.1; the polymorphism of the SNP molecular marker is A / G type; and the application of the method in assisted breeding of Penaeus vannamei is the screening of body length or weight of Penaeus vannamei.

2. The method for screening molecular markers related to economic traits of Penaeus vannamei according to claim 1, characterized in that: The step A comprises: Step A-1: ​​performing phenotypic measurement of growth traits on the breeding population of Penaeus vannamei to obtain the phenotypic data; the growth trait is body length or weight; Step A-2: extracting total DNA from the Litopenaeus vannamei breeding population to obtain the total DNA; constructing a DNA library using the total DNA; sequencing the DNA library to obtain the SNP typing data; The construction process of the DNA library includes: breaking the total DNA into small fragments by ultrasonic treatment, and undergoing end modification, linker connection, PCR amplification and magnetic bead purification.

3. The method for screening molecular markers related to economic traits of Penaeus vannamei according to claim 1, characterized in that: The step C comprises: Step C-1: sort the P value of each SNP in ascending order, and use the Bonferroni multiple correction method to judge the P value <1.9E-08 as a significant site; screen out SNPs that are significantly associated with growth traits; Step C-2: Use CUT&Tag to screen the key histone modification regions, and use the key histone modification regions to accurately locate SNPs significantly associated with growth traits and the expression regulatory elements of their corresponding genes; the key histone modification regions are H3K4me1, H3K4me3, H3K27me3 and H3K27ac.