Fine mapping of major QTL for growth traits in red sea bream (Pagrosomus major) and its application
Patent Information
- Application Number
- CN202611191118.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-07
- Publication Date
- 2026-09-04
AI Technical Summary
该方法通过整合低成本育种芯片靶向测序与全基因组重测序参考面板的高密度基因型填充,结合混合线性模型GWAS、条件-联合分析、连锁不平衡分析、贝叶斯统计精细定位和表型极端个体等位基因频率验证,能够将复杂的着丝粒周围QTL簇解析为多个相互独立的关联信号,并将候选致因变异锁定至单基因乃至单SNP水平,从而克服现有方法定位区间宽、无法区分连锁信号、缺乏候选变异概率支持的缺陷
[0030]1.本发明首次在红鳍东方鲀上建立了整合高密度基因型填充、混合线性模型GWAS、条件-联合分析、贝叶斯统计精细定位和表型极端个体等位基因频率验证的生长性状主效QTL精细定位框架,实现了从关联发现到候选致因基因锁定的完整技术路线。
Smart Images

Figure CN122686802A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of genomics and molecular breeding technology of aquaculture animals, specifically involving the fine mapping method and application of major QTLs for the growth traits of redfin pufferfish. Background Technology
[0002] Growth rate is one of the most important economic traits in aquaculture animals, directly affecting production efficiency, market cycle, feed utilization, and overall economic benefits. In recent years, selective breeding and genome-wide selection have made significant genetic progress in various aquaculture animals, and genomic tools are increasingly being used to improve traits such as growth, disease resistance, and environmental tolerance. However, the efficiency of marker-assisted breeding and genome-wide selection largely depends on the genetic structure of the target trait; that is, trait variation may be controlled by a large number of minor loci, or by a few major quantitative trait loci (QTLs), or both.
[0003] The redfin pufferfish (Takifugu rubripes) is a high-value marine aquaculture fish mainly cultivated in East Asia. Compared with aquaculture animals such as Atlantic salmon and tilapia, which have undergone long-term selective breeding, the selective breeding of redfin pufferfish is still in its early stages. Existing studies have shown that its growth, disease resistance, and ammonia nitrogen tolerance traits exhibit heritable variations, possessing additive genetic variations that can be selected and utilized. However, the genomic regions and candidate genes controlling its growth variation are still unclear.
[0004] Genome-wide association studies (GWAS) have been widely used to analyze growth traits in aquaculture animals. However, their localization resolution is often limited by marker density, linkage disequilibrium (LD) structure, reference genome quality, and population structure, often yielding only broad association intervals. This makes it difficult to distinguish between multiple tightly linked independent signals and to provide probabilistic evidence supporting candidate variants. Growth traits are generally considered quantitative traits, but in some species, major QTLs may also contribute significantly. When major QTLs are located in recombination-suppressed regions such as around the centromere, extended haplotypes and strong local LDs further complicate fine localization and the differentiation of candidate causative variants.
[0005] Current QTL studies in aquaculture animals largely rely on linkage analysis, medium-density GWAS, or location-based candidate gene interpretation. While these methods can identify broad association intervals, they typically cannot separate multiple linkage signals within the same region, nor can they assign posterior probability support to candidate causative variations. Therefore, there is an urgent need for a method capable of finely mapping major-effect QTLs of growth traits, distinguishing independent signals, and identifying candidate causative genes and molecular markers to support marker-assisted breeding and genome selection of growth traits in redfin pufferfish. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention provides a method and application for fine mapping of major-effect QTLs for growth traits in the red-finned pufferfish. This method integrates low-cost breeding chip targeted sequencing with high-density genotype filling from whole-genome resequencing reference panels, combined with hybrid linear model GWAS, conditional-joint analysis, linkage disequilibrium analysis, Bayesian statistical fine mapping, and allele frequency verification in phenotypic extreme individuals. This enables the resolution of complex pericentromere QTL clusters into multiple independent associated signals, and pinpoints candidate causative variants to the single-gene or even single-SNP level. This overcomes the shortcomings of existing methods, such as wide mapping intervals, inability to distinguish linkage signals, and lack of probability support for candidate variants.
[0007] This invention is achieved through the following technical solution:
[0008] The method for fine mapping of major-effect QTLs of growth traits in redfin pufferfish according to the present invention includes the following steps:
[0009] (1) Collect fin samples from redfin pufferfish families and record growth phenotypic data of the population in multiple consecutive culture stages or environments;
[0010] (2) Extract the genomic DNA of all individuals in the family population described in step (1), perform targeted sequencing and genotyping using a breeding chip, obtain the SNP genotypes of all individuals, and obtain a low-density skeleton SNP panel after quality control filtering.
[0011] (3) Using the whole genome resequencing SNPs of redfin pufferfish as a reference panel, the low-density backbone SNP panel described in step (2) was filled with genotypes and filtered to obtain a high-density SNP panel;
[0012] (4) Using the high-density SNP panel described in step (3) and the growth trait phenotypic data described in step (1), a mixed linear model is used to perform genome-wide association analysis to obtain genome-wide significant association SNPs for each growth trait;
[0013] (5) For growth traits that are significantly associated across the entire genome, conditional-joint analysis is performed to identify mutually independent conditional association signals on the same chromosome;
[0014] (6) Perform linkage disequilibrium analysis and Bayesian statistical fine localization on each conditionally independent association signal in step (5), calculate the posterior inclusion probability of each SNP and determine the credible set, and lock in the candidate causal SNP and the candidate genes in or near it.
[0015] (7) Verify the allele frequencies of the candidate causal SNPs described in step (6) by comparing the phenotypic extreme individuals. When the direction of the allele frequency difference is consistent with the effect direction of the conditional-co-analysis in step (5), the associated signal is identified as the major QTL of the growth trait of redfin pufferfish, and its candidate causal SNPs, candidate genes and molecular markers that can be used for molecular marker-assisted breeding are output.
[0016] Preferably, in step (1), the number of individuals in the family group is >800; the continuous culture stage or environment includes three culture stages carried out sequentially: seawater earthen pond, indoor cement pond and marine net cage; body weight (BW) is measured at the beginning and end of each stage, and the average daily weight gain (ADG), specific growth rate (SGR) and daily weight gain coefficient (DGC) of each stage and the whole period are calculated according to the generally accepted definition, and individuals with missing corresponding phenotypic records are removed in the corresponding trait analysis; the growth phenotypic data include body weight, average daily weight gain, specific growth rate and daily weight gain coefficient.
[0017] Preferably, in step (2), the number of SNPs contained in the breeding chip is >20,000 and the targeted sequencing depth is >100X; after quality control, the sequencing reads are aligned to the telomere-to-telomere (T2T) reference genome of redfin pufferfish without gaps, and after mutation detection and filtering, the following quality control conditions are adopted: SNPs with a minimum allele frequency <0.01 or a deletion rate >0.05 are removed, as are individuals with a genotype deletion rate >0.05, and duplicate individuals are removed.
[0018] Preferably, in step (3), the reference panel is obtained by whole-genome resequencing of the redfin pufferfish, containing ≥700 individuals and ≥2×10 high-quality biallelic SNPs. 6 One genotype was used; the software used for genotype filling was Beagle; after filling, sites with a dose determination coefficient DR2 < 0.8 or a minimum allele frequency < 0.05 were removed, and all genotypes were stored in PLINK binary format.
[0019] Preferably, before step (4), a genome kinship matrix is constructed using the VanRaden method with a high-density SNP panel. The basic heritability of each growth trait is estimated by the average information-constrained maximum likelihood method (AI-REML) under the framework of genome best linear unbiased prediction (GBLUP). The first 5 genomic principal components are used as fixed covariates to confirm that there are additive genetic variations available for the trait.
[0020] Preferably, in step (4), the mixed linear model uses the genome-wide SNP-based kinship matrix as a random multigene effect and the first 5 genome principal components as fixed covariates as a single-marker mixed linear association model, and the software used is GCTA; the genome-wide significance threshold is corrected using Bonferroni, i.e., 0.05 / the total number of SNPs involved in the test; and the genome expansion factor λ is calculated to assess the control of population structure and kinship.
[0021] Preferably, in step (5), the conditional-joint analysis is performed using stepwise conditional-joint analysis GCTA-COJO. The reference linkage disequilibrium is estimated by the study population, and its significance threshold is consistent with the single-marker whole-genome significance threshold. The search window is set to 10000kb, and the maximum collinearity threshold r 2 Set to 0.9; when a SNP remains significant across the entire genome after conditional analysis of the selected SNPs, it is retained as a conditionally independent association signal.
[0022] Preferably, after step (5), the proportion of phenotypic variation explained by each conditional independent signal is estimated by chi-square approximation using the conditional joint effect statistic of the conditional-joint analysis, and the phenotypic variation explanation rate of each conditional independent signal on the same chromosome is summed as the total phenotypic variation explanation rate of the major effect QTL cluster.
[0023] Preferably, in step (6), the chain imbalance analysis uses PLINK based on r 2 Pair linkage disequilibrium was calculated and cross-centromere linkage disequilibrium was assessed. When significant associations were found on both sides of the centromere of the same chromosome and the cross-centromere linkage disequilibrium between the guiding SNPs on both sides was low, the associated region was identified as a pericentromere QTL cluster, and the signals on its two arms were subjected to conditional-joint analysis and fine localization, respectively.
[0024] Preferably, in step (6), the Bayesian statistical fine localization includes two methods used in parallel: one is the Wakefield approximate Bayesian factor method, which sets the prior variance of the true allele effect W=0.04, only includes SNPs with nominal P<0.001 within the window to calculate the posterior inclusion probability, and uses the smallest SNP set with a cumulative posterior inclusion probability of 0.95 as the 95% confidence set; the other is SuSiE fine localization, which inputs the GWAS summary statistics of each signal and the linkage disequilibrium matrix estimated by the same filled genotype panel into the susie_rss function, and performs fine localization within a 300kb window upstream and downstream of the conditionally independent signal-guided SNP with the maximum single effect number L=5 and a confidence set coverage of 95%; the variants that are consistently prioritized by the two methods, or variants with high posterior inclusion probability and adjacent to the guiding SNP and with strong marginal association evidence, are used as candidate causative SNPs of the signal.
[0025] Preferably, after step (6), according to the genome annotation file of the reference genome, the candidate causative SNPs are annotated according to their genomic positions relative to neighboring protein-coding genes, and the protein-coding gene with the highest posterior containment probability variation in the main confidence set of each conditionally independent signal is taken as the main candidate causative gene of the signal; and the phenotypic variation explained rate (PVE) of each signal is estimated by chi-square approximation using the conditional joint effect statistic of the conditional-joint analysis, and the sum of the PVEs of each signal on the same chromosome is taken as the total PVE of the QTL cluster.
[0026] Preferably, in step (7), for each trait to be verified, individuals with non-missing phenotypic records are sorted by phenotypic value, the top 30% are the fast-growing group and the bottom 30% are the slow-growing group, and the frequency difference Δf of the alleles tracked by the candidate causative SNP is calculated as f_top30% - f_bottom30%. If the direction of Δf is consistent with the direction of the conditional combined effect, then the phenotypic verification is passed.
[0027] Preferably, the major QTL for the growth trait of redfin pufferfish determined in step (7) is located in the region around the centromere of chromosome 22, and the candidate causative genes include one or more of EGFR, PTPRF, PRKCI, KISS1R, TTLL12 and JPH1.
[0028] The application of candidate causative SNPs or candidate genes of major QTLs for growth traits of red-finned pufferfish determined by the method of the present invention in the preparation of molecular marker-assisted breeding or genomic selection products for growth traits of red-finned pufferfish.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] 1. This invention establishes for the first time a fine mapping framework for major QTLs of growth traits in redfin pufferfish, integrating high-density genotype filling, mixed linear model GWAS, conditional-joint analysis, Bayesian statistical fine mapping, and allele frequency verification of phenotypic extreme individuals, realizing a complete technical route from association discovery to candidate causative gene identification.
[0031] 2. This invention can resolve complex QTL clusters around centromeres into multiple independent correlation signals, overcoming the limitation of traditional methods that can only obtain a single broad correlation peak and cannot distinguish closely linked signals; in the embodiment, the significant correlation of 8 growth traits is resolved into 11 conditionally independent signals at 9 sites.
[0032] 3. This invention uses a parallel fine localization of two Bayesian methods, Wakefield approximate Bayesian factor and SuSiE, to provide posterior inclusion probabilities and confidence sets for candidate variants, thereby locking candidate causative variants down to the single gene or even single SNP level, and identifying high-confidence candidate causative genes such as EGFR, PTPRF, PRKCI, and KISS1R, overcoming the deficiency of existing methods in lacking probabilistic candidate evidence.
[0033] 4. This invention uses allele frequency comparisons of phenotypic extreme individuals to perform internal phenotypic verification of the mapping results, requiring that the direction of frequency difference be consistent with the direction of conditional combined effects, which significantly improves the reliability of candidate causative variants.
[0034] 5. This invention employs a strategy of low-cost breeding chip targeted sequencing combined with whole-genome resequencing reference panel filling, balancing cost and resolution. The resulting candidate causative SNPs and candidate genes can be directly used for molecular marker-assisted breeding and genome selection of growth traits in redfin pufferfish.
[0035] 6. This invention specifies the specific implementation requirements and parameters for each step. The absence of any key step or the parameters exceeding the specified range will result in the inability to accurately locate the major QTL or the generation of false positive results. For example, if high-density genotype filling is not performed, the QTL clusters around the centromere cannot be resolved into independent signals. If the sequencing depth is too low or the genotype filtering parameters exceed the specified range of this invention, it will lead to localization failure or false positives. Attached Figure Description
[0036] Figure 1 Manhattan plots were created for eight growth traits that showed significant signals on chromosome 22.
[0037] Figure 2 Regional association signal map of the QTL regions of chromosome 22 growth (ADG2, BW2, DGC2);
[0038] Figure 3 The diagram shows the linkage disequilibrium structure of the QTL region on chromosome 22, where (a) is a diagram showing the linkage disequilibrium decay of the left arm-guided SNP (Chr22: 4.65Mb), (b) is a diagram showing the linkage disequilibrium decay of the right arm-guided SNP (Chr22: 8.31Mb), and (c) is a bar chart comparing the degree of transcentromere linkage disequilibrium between the two arm-guided SNPs.
[0039] Figure 4The SuSiE statistical fine localization maps of high-confidence QTL signals on chromosome 22 are shown below. (a) is the fine localization map of the EGFR region (ADG2 signal 1, 2.5-3.2Mb), (b) is the fine localization map of the PTPRF-PRKCI-KISS1R cluster (ADG2 signal 2, 3.8-5.2Mb), (c) is the fine localization map of the JPH1 region (ADG2 signal 3, right arm, 7.5-8.4Mb), and (d) is a bar chart of the maximum SuSiE PIP value and the dominant confidence set for each signal window.
[0040] Figure 5 The diagrams show the allele frequencies of candidate causative SNPs in fast-growing and slow-growing individuals. (a) shows the allele frequencies of candidate causative SNPs for the ADG2 trait in fast-growing and slow-growing individuals, and (b) shows the allele frequencies of candidate causative SNPs for the BW2 trait in fast-growing and slow-growing individuals. Detailed Implementation
[0041] To enable those skilled in the art to better understand the technical content of this invention, the preferred embodiments of this invention are further described below in conjunction with specific examples. Where specific techniques or conditions are not specified in the examples, they are performed according to the techniques or conditions described in the literature or according to the product instructions; unless otherwise specified, all equipment and reagents used are commercially available.
[0042] Example 1: Implementation of the major-effect QTL fine mapping method for growth traits of redfin pufferfish according to the present invention in redfin pufferfish.
[0043] (1) Collection of family population samples and growth phenotype data
[0044] A family of redfin pufferfish, totaling 811 individuals, was used as the experimental group. They were cultured in three sequential stages / environments: seawater earthen ponds, indoor cement ponds, and marine net cages. Body weight was measured at the beginning of the experiment and at the end of each stage, resulting in BW0, BW1, BW2, and BW3. Average daily weight gain (ADG1-ADG3 and total ADG), specific growth rate (SGR1-SGR3 and total SGR), and daily weight gain coefficient (DGC1-DGC3 and total DGC) were calculated according to accepted definitions for each stage and the entire production period, analyzing a total of 16 growth traits. Individuals lacking corresponding phenotypic records were excluded from the corresponding trait analysis. Fin rays were collected at the final body weight measurement for DNA extraction.
[0045] (2) Breeding chip targeted sequencing and backbone SNP panel construction
[0046] Genomic DNA was extracted from the fin rays of all individuals and targeted sequencing was performed using the DNBSEQ-T7 platform with a 20K SNP breeding array for red-finned pufferfish, achieving a sequencing depth >100X. Sequencing reads were quality controlled using FASTP and aligned to the gap-free T2T red-finned pufferfish reference genome (BWA). Variation detection was performed using GATK, and filtering was conducted using VCFtools. Further quality control was performed using PLINK: SNPs with a minimum allele frequency <0.01 or a deletion rate >0.05, as well as individuals with a genotype deletion rate >0.05, were removed. After quality control, 67,868 high-quality SNPs were retained as a low-density backbone panel.
[0047] (3) High-density genotype filling
[0048] Using the whole-genome resequencing reference panel of *Putrachea rubra* (containing 721 individuals and 2.39 million high-quality biallelic SNPs) as a filling reference, genotyping of the backbone panel was performed using Beagle v5.5. After filling, loci with dose determination coefficient DR2 < 0.8 or minimum allele frequency < 0.05 were removed, resulting in 1,420,976 high-quality filling SNPs distributed across 22 chromosomes. The marker density was approximately 21-fold higher than that of the backbone panel, and the median DR2 of the retained loci was 0.99.
[0049] (4) SNP baseline heritability estimation
[0050] Using a filled high-density SNP panel, a genomic phylogenetic matrix was constructed using the VanRaden method. Within the HIBLUP GBLUP framework, AI-REML was used to estimate the basal heritability of SNPs for 16 growth traits, with the first five principal genomic components used as fixed covariates. The SNP heritability of the 16 traits ranged from 0.09 to 0.54; the heritability estimates obtained from the filled panel and the skeleton panel were almost identical (h between the two panels). 2 The mean absolute difference is only 0.015, indicating that the skeleton panel has captured most of the additive kinship information in the population, and the main value of the filler is to improve the localization resolution.
[0051] (5) Genome-wide association analysis
[0052] Genome-wide association analysis (GWA) was performed using a GCTA-based mixed linear model with filled high-density SNP panels and 16 growth trait phenotypes. The VanRaden genome phylogenetic matrix was used as a random polygenic effect, and the first five principal components of the genome were used as fixed covariates. The Bonferroni-corrected genome-wide significance threshold was 0.05 / 1420976 = 3.52 × 10⁻⁶. -8 The results showed that eight traits were significantly associated across the entire genome, and all significant SNPs were located on chromosome 22 (see Table 1 and 2018). Figure 1 Among them, ADG2 signal was the strongest, with 664 SNPs showing genome-wide significance, and the top SNP located at Chr22:7,856,970 (P=3.65×10). -18 The genomic expansion factor λ for each trait was close to or slightly below 1.0, indicating that population structure and kinship were well controlled.
[0053] Table 1. SNP information significantly associated with growth traits of redfin pufferfish
[0054]
[0055] (6) Linkage disequilibrium analysis and determination of QTL clusters around centromeres
[0056] Regional analysis revealed significant associations extending from approximately 0.96 Mb to 11.4 Mb on chromosome 22, and spanning the annotated centromere region (6.26–6.93 Mb). Figure 2 SNPs are scarce within the centromere region, forming local troughs between the correlation peaks of the two arms. Using PLINK, press r... 2 Calculate chain imbalances, such as Figure 3 The black vertical dashed lines mark the left arm guiding SNP (Chr22:4,649,905:C:T) and the right arm guiding SNP (Chr22:8,308,427:G:A). The mean r² and standard error of each position relative to the guiding SNP are calculated using a 100 kb window. The transcentromere linkage disequilibrium between the left and right arm guiding SNPs is relatively low (r²). 2 =0.148), intraarm linkage imbalance decreases with physical distance, while transcentromere linkage imbalance is generally low (median r). 2 <0.05). Based on this, the associated region was identified as a QTL cluster around the centromere, and the signals on its two arms were subjected to conditional-joint analysis and fine localization, respectively.
[0057] (7) Conditional-joint analysis identifies conditionally independent signals
[0058] For eight traits with significant genome-wide associations, GCTA-COJO stepwise conditional-joint analysis was used to identify conditionally independent signals, with a search window of 10000 kb and a maximum collinearity threshold r. 2=0.9, the significance threshold is the same as the single-marker threshold. A total of 11 conditionally independent signals were identified at 9 loci on chromosome 22 for the 8 traits (see Table 2). ADG2 had the most complex genetic structure, containing 3 independent signals (2.94 Mb, 4.43 Mb, and 7.86 Mb), with the 7.86 Mb signal being the only one located on the right arm; BW2 contained 2 independent signals (2.61 Mb and 4.65 Mb). The signals were concentrated in the cluster of approximately 2.6–5.7 Mb on the left arm and at approximately 7.86 Mb on the right arm, and multiple signals were shared among closely related growth traits (especially in the 4.0–4.7 Mb region), indicating that this QTL cluster has pleiotropic effects on multiple growth indicators.
[0059] Table 2 Conditionally Independent Association Signals of Chromosome 22 Identified by GCTA-COJO
[0060] (8) Bayesian statistical fine mapping and candidate causative gene identification
[0061] For each conditionally independent signal, fine localization was performed using the Wakefield approximate Bayesian factorization method (prior variance W=0.04, including only SNPs with nominal P<0.001 within the window, with the smallest SNP set having a cumulative posterior inclusion probability (PIP) of 0.95 as the 95% confidence set) and the SuSiE method (guiding 300kb windows upstream and downstream of the SNP, L=5, 95% confidence set coverage). The most concentrated evidence for fine localization appeared around 2.80Mb: variant 22:2,797,238:A:G had the highest PIP in all analyzed regions (0.923 for BW2 signal 1 and 0.907 for ADG2 signal 1), located within the EGFR gene, making EGFR the primary candidate causative gene for these signals. The signal in the 4.0-4.1 Mb region was localized within the PTPRF (LAR) gene (ADG3 signal 1 guided SNP 22:4,006,127:A:T SuSiE PIP=0.792); ADG2 signal 2 (approximately 4.43 Mb) was located within the main confidence set for 22:4,434,720:A:G (PIP=0.614, P=3.10×10⁻⁶). -14 The signal is located within the PRKCI gene and is considered the main candidate causative gene for this signal; another signal is located in the KISS1R gene (22:4,667,206:T:C). The DGC2-specific signal (5.70 Mb) guides the SNP to be located within the TTLL12 intron, and the right arm ADG2 signal 3 (approximately 7.86 Mb) is adjacent to JPH1, but both have wide confidence sets and low highest PIP values, resulting in limited resolution (see Table 3 and...). Figure 4 ).
[0062] Table 3 Candidate causative variants with fine localization preference for SuSiE (PIP≥0.4)
[0063]
[0064] (9) Estimation of the rate of explanation for phenotypic variation
[0065] The phenotypic variation explained by each conditionally independent signal was estimated using the chi-square approximation of the GCTA-COJO conditional combined effect statistic, and the sum of the PVEs of all signals for the same trait was taken as the total contribution of the QTL cluster on chromosome 22 (see Table 4, where bJ is the conditional combined effect of COJO on the A1 allele). The three independent signals of ADG2 explained a total of 17.5% of the phenotypic variation, with the PRKCI region contributing the most (PVE=7.0%), followed by the EGFR region (5.4%) and the right arm JPH1 region (5.1%); the two signals of BW2 explained a total of 15.7%. Compared with the heritability of genome-wide SNPs, the QTL cluster on chromosome 22 explained a large proportion of the heritable variation in the second-stage growth trait. For example, ADG2's 17.5% contribution is approximately equivalent to 48% of its SNP heritability, indicating that this QTL cluster is the main contributor to the growth performance of this population in the indoor cement pool stage.
[0066] Table 4. Phenotypic Variation Explained (PVE) of Condition-Independent Signals
[0067]
[0068] (10) Verification of allele frequencies in phenotypic extreme individuals
[0069] To validate the fine mapping results internally using phenotypes, taking ADG2 as an example, individuals with non-missing records were sorted phenotypes, with the top 30% (235 individuals, ADG2 ≥ 1.13 g / day) designated as the fast-growing group and the bottom 30% (234 individuals, ADG2 ≤ 0.85 g / day) as the slow-growing group. The frequency difference Δf between candidate causative SNPs and alleles was calculated as ftop 30% - fbottom 30%. The direction of Δf for all examined loci was consistent with the direction of their COJO conditional combined effects (see Table 5). The minor G allele of EGFR candidate SNP 22:2,797,238:A:G was only 3.2% in the fast-growing group and reached 26.0% in the slow-growing group (Δf = -0.228); the minor allele of the JPH1 locus on the right arm was almost absent in the fast-growing group (0.6%) and reached 18.7% in the slow-growing group (Δf = -0.180). Conversely, the minor G allele at PRKCI locus 22:4,434,720:A:G was associated with faster growth, at 55.3% and 20.3% in the fast and slow groups, respectively (Δf=+0.351), representing the largest absolute difference among the tested loci, consistent with its positive conditional effect. Figure 5The consistent direction of phenotypic extremes defined by BW2 and DGC2 indicates that the aforementioned allele frequency contrasts are robust among closely related growth traits.
[0070] Table 5. Comparison of allele frequencies of candidate causative SNPs in extreme ADG2 phenotype individuals (top 30% and bottom 30%)
[0071]
[0072] (11) Main effect QTL fine localization results
[0073] Based on the above analysis, this embodiment finely maps the major QTLs affecting the growth traits of *Putrachea rubra* to the pericentromere region of chromosome 22, and resolves this complex QTL cluster into 11 independent associated signals at 9 loci. Through Bayesian statistical fine mapping, candidate gene functional annotation, and validation of extreme allele frequencies, EGFR, PTPRF, PRKCI, and KISS1R were identified as high-confidence candidate causative genes (their preferential variants are located within or adjacent to genes, most have high PIP values, and are supported by conditionally independent effects and directional consistency in phenotypic validation). TTLL12 and JPH1 were identified as lower-confidence candidate causative genes. These candidate causative SNPs and candidate genes can be used as molecular markers for marker-assisted breeding and genome selection of growth traits in *Putrachea rubra*.
[0074] This embodiment demonstrates that only when the complete process of multi-stage phenotypic collection of family populations, breeding chip targeted sequencing, high-density filling of whole-genome resequencing reference panels, mixed linear model GWAS, conditional-joint analysis, linkage disequilibrium analysis, Wakefield approximate Bayesian factor and SuSiE Bayesian fine mapping, and verification of extreme allele frequencies of phenotypes is completed in sequence, and the parameters of each step are within the limits defined by this invention, can the complex QTL clusters around the centromere be finely resolved into multiple independent associated signals and high-confidence candidate causative genes and molecular markers be identified; without high-density filling or if the parameters exceed the limits, the above-mentioned fine mapping effect cannot be obtained.
Claims
1. A method for fine mapping of major-effect QTLs of growth traits in red-finned pufferfish, characterized in that, Includes the following steps: (1) Collect fin samples from redfin pufferfish families and record growth phenotypic data of the population in multiple consecutive culture stages or environments; (2) Extract the genomic DNA of all individuals in the family population described in step (1), perform targeted sequencing and genotyping using a breeding chip, obtain the SNP genotypes of all individuals, and obtain a low-density skeleton SNP panel after quality control filtering. (3) Using the whole genome resequencing SNPs of redfin pufferfish as a reference panel, the low-density backbone SNP panel described in step (2) was filled with genotypes and filtered to obtain a high-density SNP panel; (4) Using the high-density SNP panel described in step (3) and the growth trait phenotypic data described in step (1), a mixed linear model is used to perform genome-wide association analysis to obtain genome-wide significant association SNPs for each growth trait; (5) For growth traits that are significantly associated across the entire genome, conditional-joint analysis is performed to identify mutually independent conditional association signals on the same chromosome; (6) Perform linkage disequilibrium analysis and Bayesian statistical fine localization on each conditionally independent association signal in step (5), calculate the posterior inclusion probability of each SNP and determine the credible set, and lock in the candidate causal SNP and the candidate genes in or near it. (7) Verify the allele frequencies of the candidate causal SNPs described in step (6) by comparing the phenotypic extreme individuals. When the direction of the allele frequency difference is consistent with the effect direction of the conditional-co-analysis in step (5), the associated signal is identified as the major QTL of the growth trait of redfin pufferfish, and its candidate causal SNPs, candidate genes and molecular markers that can be used for molecular marker-assisted breeding are output.
2. The method for fine-tuning the major-effect QTL of growth traits in *Pueraria lobata* according to claim 1, characterized in that: In step (1), the number of individuals in the family group is greater than 800; the continuous breeding stage or environment includes three breeding stages carried out in sequence: seawater earthen pond, indoor cement pond and marine net cage; the body weight is measured at the beginning and end of each stage, and the average daily weight gain, specific growth rate and daily weight gain coefficient of each stage and the whole period are calculated. The growth phenotypic data include body weight, average daily weight gain, specific growth rate and daily weight gain coefficient.
3. The method for fine-tuning the major-effect QTL of growth traits in *Pueraria lobata* according to claim 1, characterized in that: In step (2), the number of SNPs contained in the breeding chip is >20,000 and the targeted sequencing depth is >100X; the quality control filtering conditions are: removing SNPs with a minimum allele frequency <0.01 or a deletion rate >0.05, and removing individuals with a genotype deletion rate >0.
05.
4. The method for fine-tuning the major-effect QTL of growth traits in *Pueraria lobata* according to claim 1, characterized in that: In step (3), the reference panel was obtained by whole-genome resequencing of the redfin pufferfish, containing ≥700 individuals and ≥2×10 high-quality biallelic SNPs. 6 One; the software used for genotype filling was Beagle; after filling, sites with a dose determination coefficient DR2 < 0.8 or a minimum allele frequency < 0.05 were removed.
5. The method for fine-tuning the major-effect QTL of growth traits in *Pueraria lobata* according to claim 1, characterized in that: Before step (4), the high-density SNP panel described in step (3) is used to construct a genome kinship matrix using the VanRaden method. The average information constraint maximum likelihood method under the optimal linear unbiased prediction framework of the genome is used to estimate the basic heritability of SNPs for each growth trait in order to confirm that the trait has available additive genetic variation. In step (4), the mixed linear model is a single-marker mixed linear association model that uses the genome-wide SNP-based kinship matrix as a random multigene effect and the first 5 genome principal components as fixed covariates. The software used is GCTA. The genome-wide significance threshold is corrected using Bonferroni, i.e., 0.05 / the total number of SNPs involved in the test.
6. The method for fine-tuning the major-effect QTL of growth traits in *Pueraria lobata* according to claim 1, characterized in that: In step (5), the conditional-joint analysis was performed using stepwise conditional-joint analysis GCTA-COJO. The reference linkage disequilibrium was estimated by the study population, and its significance threshold was consistent with the single-marker whole-genome significance threshold. The search window was set to 10000kb, and the maximum collinearity threshold r was set to 10000kb. 2 Set to 0.9; when a SNP remains significant across the entire genome after conditional analysis of the selected SNPs, it is retained as a conditionally independent association signal; After step (5), the proportion of phenotypic variation explained by each conditional independent signal is estimated by chi-square approximation using the conditional joint effect statistic of the conditional-joint analysis, and the phenotypic variation explained by each conditional independent signal on the same chromosome is summed to obtain the total phenotypic variation explained by the major QTL cluster.
7. The method for fine-tuning the major-effect QTL of growth traits in *Pueraria lobata* according to claim 1, characterized in that: In step (6), the chain imbalance analysis is performed using PLINK based on r 2 Calculate pairwise linkage disequilibrium and assess transcentromere linkage disequilibrium; when significant associations are found on both sides of the centromere of the same chromosome and transcentromere linkage disequilibrium between the guiding SNPs on both sides is low, the associated region is identified as a pericentromere QTL cluster, and the signals on its two arms are subjected to conditional-joint analysis and fine localization respectively. In step (6), the Bayesian statistical fine localization includes two methods used in parallel: one is the Wakefield approximate Bayesian factor method, which sets the prior variance of the true allele effect W=0.04, only includes SNPs with nominal P<0.001 within the window to calculate the posterior inclusion probability, and uses the smallest SNP set with a cumulative posterior inclusion probability of 0.95 as the 95% confidence set; the other is SuSiE fine localization, which inputs the GWAS summary statistics of each signal and the linkage disequilibrium matrix estimated by the same filled genotype panel into the susie_rss function, and performs fine localization within a 300kb window upstream and downstream of the conditionally independent signal-guided SNP with the maximum single effect number L=5 and a confidence set coverage of 95%; the variants that are consistently prioritized by the two methods, or variants with high posterior inclusion probability and adjacent to the guiding SNP and with strong marginal association evidence, are used as candidate causative SNPs of the signal; After step (6), the candidate causative SNPs are annotated according to their genomic positions relative to neighboring protein-coding genes based on the genome annotation file of the reference genome. The protein-coding gene with the highest posterior containment probability variation in the master confidence set of each conditionally independent signal is taken as the master candidate causative gene of that signal.
8. The method for fine-tuning the major-effect QTL of growth traits in *Pueraria lobata* according to claim 1, characterized in that: In step (7), for each trait to be verified, individuals with non-missing phenotypic records are sorted by phenotypic value, the top 30% of individuals are selected as the fast-growing group and the bottom 30% of individuals are selected as the slow-growing group. The frequency difference Δf of the alleles tracked by the candidate causative SNP in the two groups is calculated as f_front30% - f_back30%. If the direction of Δf is consistent with the direction of the conditional combined effect of the SNP, the candidate causative SNP is determined to have passed phenotypic verification.
9. The method for fine-tuning the major-effect QTL of growth traits in *Pueraria lobata* according to claim 1, characterized in that: The major QTL for the growth trait of redfin pufferfish determined in step (7) is located in the pericentromere region of chromosome 22, and the candidate causative genes include one or more of EGFR, PTPRF, PRKCI, KISS1R, TTLL12 and JPH1.
10. The application of the candidate causative SNP or candidate gene of the major QTL of the growth trait of redfin pufferfish determined by the method according to any one of claims 1-9 in the preparation of molecular marker-assisted breeding or genomic selection products of the growth trait of redfin pufferfish.