Method for locating pig daily weight gain related genes based on graph pan-genome
Patent Information
- Application Number
- CN202610661624.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-14
- Publication Date
- 2026-08-18
AI Technical Summary
然而,基于单一参考基因组的遗传变异分析存在固有局限性:单一参考基因组无法代表物种的全部遗传信息,对该物种基因组多样性的代表性不足,会直接削弱检测效力
[0039] This invention utilizes Minigraph-Cactus to construct a graph pan-genome, combining Minigraph's efficient structural variant capture capability with Cactus's high-precision multiple sequence alignment, significantly improving the detection accuracy of structural variants. The capture rate for SV fragments in the 50-500 bp range reaches 62.3%, an improvement of over 40% compared to traditional linear reference genomes.
Smart Images

Figure CN122598764A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of animal molecular biology and genetic breeding technology, and in particular to a method for locating genes related to daily weight gain in pigs based on a pan-genome mapping method. Background Technology
[0002] Daily weight gain is one of the most important economic traits in pig breeding, directly affecting the economic benefits of pig farming. Traditional breeding relies on phenotypic selection and pedigree information. With the development of molecular marker technology, genome-wide association analysis has become an important tool for elucidating the genetic basis of complex traits.
[0003] In existing technologies, genome-wide association studies (GWAS) primarily focus on the association between single nucleotide polymorphisms (SNPs) and daily weight gain, and rely on a single reference genome (such as Sscrofa11.1 in pigs). However, genetic variation analysis based on a single reference genome has inherent limitations: a single reference genome cannot represent all the genetic information of a species, resulting in insufficient representativeness of the species' genomic diversity and directly weakening the detection power. Furthermore, existing studies mainly focus on SNP markers, neglecting structural variation, an important type of genetic variation.
[0004] Structural variations, including deletions, insertions, duplications, inversions, and translocations, affect genomic regions far more extensively than SNPs, significantly influencing gene function and phenotypic variation. Studies have shown that structural variations play a crucial role in breed formation, environmental adaptation, and the determination of economic traits in pigs. However, due to the difficulty in detecting structural variations and their reliance on a single reference genome, their application in the genetic analysis of daily weight gain in pigs has lagged behind.
[0005] Graphical pangenomics is an emerging genomics approach that can capture all variations within a species, and it is particularly efficient at detecting structural variations longer than 50 bp. The FEZF2 gene has a function in regulating muscle growth, but its association with daily weight gain in pigs has not yet been reported.
[0006] The existing technologies have the following main defects: (1) They mainly rely on SNP markers and a single reference genome, ignoring the genetic contribution of structural variations and genomic diversity; (2) They lack methods for detecting systemic structural variations and analyzing trait associations based on graph pan-genomics; (3) They fail to effectively integrate graph pan-genomics strategies to locate key functional genes. Summary of the Invention
[0007] To address the technical problems mentioned in the background section, this invention provides a method for locating pig daily weight gain-related genes based on graph pangenome mapping.
[0008] This invention is achieved using the following technical solution: a method for locating pig daily weight gain-related genes based on a graph pangenome, comprising the following steps:
[0009] Step 1: Collect blood or tissue samples from different breeds of pigs, extract genomic DNA, perform whole-genome sequencing, and obtain sequencing data;
[0010] Step 2: Perform quality control and filtering on the raw sequencing data to remove low-quality reads and adapter sequences, obtaining high-quality data;
[0011] Step 3: Perform haplotype-based phase assembly on the high-quality data to obtain the primary assembled genome for each sample;
[0012] Step 4: Using the pig reference genome Sscrofa11.1 as a backbone, the primary assembled genome is mounted at the chromosome level to obtain the chromosome-level genome;
[0013] Step 5: Based on the chromosome-level genome, construct a pig graph pan-genome. The construction includes: using Sscrofa11.1 as the backbone genome, performing a global alignment of all sample genomes with the backbone genome, integrating variation information, and constructing a graph pan-genome structure containing nodes and edges.
[0014] Step 6: Based on the pig pangenome diagram, perform structural variant typing on the second-generation sequencing data of the target population to obtain a set of structural variants;
[0015] Step 7: Record the daily weight gain phenotypic data for each experimental pig, and perform normality tests and outlier handling on the phenotypic data;
[0016] Step 8: Perform genome-wide association analysis on the structural variation set and daily weight gain phenotypic data, using a mixed linear model to control for population stratification and kinship, and screen for structural variation sites that are significantly associated with daily weight gain;
[0017] Step 9: Perform functional annotation on the screened significant associated structural variations to determine the gene regions they affect. Through gene function annotation and pathway enrichment analysis, locate candidate functional genes related to daily weight gain.
[0018] Furthermore, in step 1, the third-generation sequencing technology is PacBio HiFi sequencing technology, with an average sequencing depth ≥30×, an average read length of 15-18kb, and a sequencing quality >99.9%.
[0019] Furthermore, in step 1, the sample includes 27 Chinese local pigs and 5 foreign pig breeds. The Chinese local pigs involve 24 representative breeds, and the foreign pig breeds include Large White, Duroc, and Yorkshire.
[0020] Furthermore, in step 5, the Minigraph-Cactus toolchain is used to construct the pig graph pangenome, specifically including:
[0021] Short sequences less than 10kb were filtered out, and a unified sequence naming rule was adopted to ensure consistency with the chromosome numbering of the reference genome Sscrofa11.1. Using Sscrofa11.1 as the backbone genome, the autosomal Chr1-18, X chromosome, and mitochondrial sequences were preserved, and the construction command was executed. All sample genomes were re-aligned to the pan-genome backbone. The gfatools tool was used to count the core indicators of the pan-genome.
[0022] Furthermore, in step 6, the PanGenie software is used to perform structural variant typing on the second-generation sequencing data of the target population, specifically including:
[0023] The gfatools tool was used to simplify the graph pan-genome map and remove redundant nodes; a graph pan-genome index was constructed; the second-generation sequencing data of the target population was compared with the graph pan-genome index; genotyping was performed based on the comparison results to obtain the structural variation genotype file for each sample; and the bcftools tool was used to merge the genotype files of all samples.
[0024] Furthermore, in step 7, the normality test and outlier handling of the phenotypic data specifically include:
[0025] A linear mixture model was constructed using DMU software to correct for environmental effects. The linear mixture model is as follows: , where y is the daily weight gain phenotypic value, μ is the population mean, Sex is the sex effect, HYS is the herd-year-season effect, BW is the initial weight effect, Litter is the litter effect, and ε is the random error; after correction, outliers were removed using the three-standard-deviation method.
[0026] Furthermore, in step 8, the genome-wide association analysis specifically includes:
[0027] Structural variations were quality controlled, and sites with a deletion rate greater than 10% and a minimum allele frequency less than 0.05 were removed. Association analysis was performed using the rMVP software package, with daily weight gain as the response variable, structural variation genotype as the fixed effect, and variety and kinship matrix as the random effect. A mixed linear model and a FarmCPU model were used for association analysis. Structural variation sites that were significantly associated with daily weight gain were screened based on significance thresholds.
[0028] Furthermore, the candidate functional gene was identified as the FEZF2 gene, with a significantly associated structural variant site being a 286 bp deletion at 44.95 Mb on porcine chromosome 13. The minimum allele frequency of the deletion was 0.211, with a significance level of P < 3.12 × 10⁻⁶. -6 .
[0029] Furthermore, the procedure also includes the following step 10: using quantitative real-time PCR to verify the expression differences of candidate functional genes in different daily weight gain pig populations;
[0030] In step 10, the quantitative real-time PCR verification specifically includes:
[0031] Pigs were divided into a high daily weight gain group and a low daily weight gain group based on their daily weight gain. The high daily weight gain group had a daily weight gain greater than 1050 g / d, while the low daily weight gain group had a daily weight gain less than 700 g / d. Total RNA was extracted from the muscle tissue of both groups and reverse transcribed to synthesize cDNA. The relative expression level of the FEZF2 gene was detected by real-time PCR. The difference in FEZF2 gene expression level between the two groups was compared to verify the correlation between FEZF2 expression level and daily weight gain.
[0032] Furthermore, step 11 includes the following: developing molecular markers based on the 286bp deletion site for assisted breeding of pig daily weight gain traits. The detection methods for these molecular markers include:
[0033] Design specific PCR primers;
[0034] The upstream primer sequence is 5'-AGCTGGTACCGAGCTCGGAT-3'.
[0035] The downstream primer sequence is 5'-TGCATGCCTGCAGGTCGACT-3';
[0036] PCR amplification was performed on the genomic DNA of the pigs to be tested;
[0037] Amplification products were detected by agarose gel electrophoresis to distinguish between deletion-type and non-deletion-type individuals.
[0038] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0039] This invention utilizes Minigraph-Cactus to construct a graph pan-genome, combining Minigraph's efficient structural variant capture capability with Cactus's high-precision multiple sequence alignment, significantly improving the detection accuracy of structural variants. The capture rate for SV fragments in the 50-500 bp range reaches 62.3%, an improvement of over 40% compared to traditional linear reference genomes.
[0040] This invention uses PanGenie for structural variant typing based on the graph pangenome, eliminating the need for read alignment and achieving accurate genotype inference through k-mer counting. It successfully obtained 438,256 high-quality SVs, of which approximately 290,000 were newly identified variants, greatly enriching the pig structural variant resource library.
[0041] This invention is the first to systematically integrate structural variation detection and genome-wide association analysis based on graph pangenome, locating the FEZF2 gene (a 286bp deletion at 44.95Mb on chromosome 13) which is significantly associated with daily weight gain in pigs. It was found that the expression level of FEZF2 is positively correlated with daily weight gain, providing new functional genes and molecular markers for the genetic improvement of pig growth traits.
[0042] This invention overcomes the limitations of traditional SNP association analysis and single reference genome dependence, and explores the genetic contribution of structural variations to daily weight gain traits, providing a key technical means for resolving the problem of "deletion heritability".
[0043] The FEZF2 gene intron deletion molecular marker developed in this invention has a simple and low-cost detection method, which can be directly applied to pig molecular marker-assisted breeding practices, promoting the transformation of the pig industry from traditional breeding to precision breeding.
[0044] This invention constructs a high-quality graph pangenome covering 32 pig individuals (including 24 local Chinese pig breeds), with a total length of 2.76 Gb and containing 266 Mb of non-reference sequences, providing a more complete reference for structural variation detection and solving the reference bias problem of traditional linear reference genomes. Attached Figure Description
[0045] Figure 1 This is a flowchart of the method proposed in the embodiments of the present invention;
[0046] Figure 2 Genetic structure analysis (genetic distance heatmap) of 32 experimental pigs in this embodiment of the invention.
[0047] Figure 3 This is a density distribution diagram of structural variations on chromosomes in the pan-genome in an embodiment of the present invention;
[0048] Figure 4 This is a histogram showing the distribution of daily weight gain phenotype data for 448 Large White pigs in this embodiment of the invention.
[0049] Figure 5 This is a Manhattan plot of genome-wide association analysis based on SV in an embodiment of the present invention (showing significant association sites on chromosome 13). Detailed Implementation
[0050] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments. It should be noted that, without conflict, the various embodiments or technical features described below can be arbitrarily combined to form new embodiments.
[0051] Example 1
[0052] like Figure 1As shown, a method for locating genes related to daily weight gain in pigs based on a graph pan-genome mapping method includes the following steps:
[0053] Step 1: Experimental Materials and Sequencing
[0054] Thirty-two pigs were selected as research subjects, including 27 Chinese local pigs (representing 24 representative breeds) and 5 foreign pig breeds (Large White, Duroc, Yorkshire, etc.). Seven of the Chinese local pigs (Bamei, Qingping, Fengjing, Enshi Black, Neijiang, Dongchuan, and Guizhou Xiang) underwent third-generation sequencing using the PacBio Revio platform, achieving an average sequencing depth ≥30× and obtaining high-quality HiFi reads (average read length 16,552 bp, average sequencing quality Q33.31). The third-generation sequencing data for the remaining 25 pigs were downloaded from public databases (covering PacBio HiFi, Nanopore, and other technology platforms).
[0055] like Figure 2 As shown in the figure (population genetic structure analysis of 32 experimental pigs - genetic distance heatmap), the 32 experimental pigs included Chinese local pigs and foreign pig breeds. The genetic distance heatmap shows the large genetic differences among Chinese local pigs and the significant differentiation between them and foreign pig breeds, providing a diverse genetic basis for constructing a high-quality graph pangenome.
[0056] Step 2: Sequencing data quality control and filtering
[0057] The raw sequencing data is subjected to quality control and filtering to remove low-quality reads and adapter sequences.
[0058] Step 3: Phased assembly of genome haplotypes
[0059] Hifiasm software (v0.16.1) was used to perform haplotype-based phase assembly on the filtered high-quality data to obtain the primary assembled genome for each sample. The parameters were set as follows: hifiasm -o samplename.asm -t 32 --primarysamplename.fastq.gz.
[0060] Step 4: Chromosome mounting
[0061] Using the pig reference genome Sscrofa11.1 as a backbone, the primary assembled genome was mounted to the chromosome level using Ragtag software (v2.1.0).
[0062] Assembly quality assessment showed that the genome assembly length ranged from 2.34 to 2.90 Gb, the Contig N50 ranged from 47.17 to 65.88 Mb, and the BUSCO integrity score was ≥95% (up to 97.8%), meeting the requirements for subsequent graphical pangenome construction.
[0063] Step 5: Construct the pig pangenome
[0064] Based on chromosome-level genome, a pig graph pangenome was constructed using the Minigraph-Cactus toolchain (v2.0.0).
[0065] 5.1 Input data preprocessing: Filter scaffolds with a length of <10kb, unify sequence naming rules, and ensure consistency with chromosome numbering of the reference genome Sscrofa11.1.
[0066] 5.2 Preliminary construction: Using Sscrofa11.1 as the backbone genome (preserving autosomal Chr1-18, X chromosome, and mitochondrial sequences), execute the command: cactus minigraph -outDir MCPanGenome -ref Sscrofa11.1.fa-threads 32 -batchSize 8 input.txt.
[0067] 5.3 Reassembly and realignment: The genomes of 32 pig individuals were realigned to the graphical pangenome backbone.
[0068] 5.4 Quality assessment: Use gfatools (v0.5) to calculate core pan-genome metrics.
[0069] The completed pig graph pangenome has a total length of 2.76 Gb, containing 185,997,078 nodes (representing genomic fragments) and 263,244,864 edges (representing connections between fragments), covering 266 Mb of non-reference sequences (NRS), of which Chinese local pigs contributed 78% of the NRS (204.36 Mb).
[0070] Step 6: Structural Variation Typing
[0071] Based on the constructed graph pangenome, PanGenie software (v1.0.3) was used to perform structural variant typing on the second-generation sequencing data of the target population to obtain a set of structural variants.
[0072] 6.1 Graph pangenome preprocessing: gfatools was used to simplify the graph and remove redundant nodes.
[0073] 6.2 Build the index: pangenie index -g simplifiedpangenome.gfa -opangenomeindex.
[0074] 6.3 Reads comparison: pangenie map -i pangenomeindex -1 sampleR1.fastq.gz -2sampleR2.fastq.gz -t 16 -o samplemapping.bam.
[0075] 6.4 Genotyping: pangenie genotype -i pangenomeindex -bsamplemapping.bam -v pangenomevariants.vcf -o samplegenotypes.vcf.
[0076] 6.5 Population integration: The genotype files of 448 samples were merged using bcftools.
[0077] A total of 438,256 high-quality structural variants were obtained, of which approximately 290,000 were newly identified variants. Length distribution characteristics: SVs of 50-500 bp accounted for 62.3% (273,034), SVs of 500 bp-1 kb accounted for 20.4% (89,404), SVs of 1-10 kb accounted for 12.3% (53,905), and large SVs >10 kb accounted for 5.0% (21,913).
[0078] like Figure 3 As shown in the figure (density distribution of structural variations on chromosomes in the pan-genome), the horizontal axis represents pig chromosomes (Chr1-Chr18), the vertical axis represents chromosome location, and the color intensity represents the density of structural variations. Structural variations are non-uniformly distributed across the entire genome, and there are significant differences in SV density among different chromosomes.
[0079] Step 7: Phenotypic Data Measurement and Processing
[0080] Record the daily weight gain phenotypic data for each experimental pig, and perform normality tests and outlier handling on the phenotypic data.
[0081] Forty-eight Large White pigs from the same core breeding group had their daily weight gain recorded from 30 kg to 100 kg. The average daily weight gain was calculated as (final weight - initial weight) / number of feeding days. A linear mixture model was constructed using DMU software to correct for environmental effects. The model formula was: y = μ + Sex + HYS + BW + Litter + ε, where y is the daily weight gain phenotypic value, μ is the population mean, Sex is the sex effect, HYS is the herd-year-season effect, BW is the initial weight effect, Litter is the litter effect, and ε is the random error. After correction, outliers were removed using the 3-standard-deviation method, resulting in valid phenotypic data for 448 individuals. The daily weight gain ranged from 580 to 1250 g / d, with an average daily weight gain of 876.8 ± 178.5 g / d.
[0082] like Figure 4 The histogram showing the distribution of daily weight gain phenotype data for 448 Large White pigs is shown. The horizontal axis represents daily weight gain (g / d), and the vertical axis represents the number of individuals. The daily weight gain phenotype data follows a normal distribution, with a mean of 876.8 g / d and a standard deviation of 178.5 g / d, ranging from 580 to 1250 g / d.
[0083] Step 8: Genome-wide association analysis (SV-GWAS)
[0084] Genome-wide association analysis was performed on the structural variations obtained in step 6 and the daily weight gain phenotypic data in step 7. A mixed linear model was used to control for population stratification and kinship to screen for structural variation sites that were significantly associated with daily weight gain.
[0085] Quality control was performed on 438,256 SVs: loci with deletion rates >10% and minimum allele frequencies (MAF) <0.05 were removed, resulting in 83,784 high-quality SVs for GWAS. SV-GWAS analysis was performed using the rMVP software package, with daily weight gain as the response variable, SV genotype as the fixed effect, and variety and kinship matrix as random effects. Mixed linear model (MLM) and FarmCPU models were used for association analysis. A total of 18 SV loci significantly associated with daily weight gain were detected (P < 1.18 × 10⁻⁶). -6 ), of which 15 are independent genetic effect sites.
[0086] like Figure 5 As shown in the Manhattan plot (genome-wide association analysis based on SV), the horizontal axis represents pig chromosomes (Chr1-Chr18), the vertical axis represents -log10 (P-value), and the red horizontal line represents the significance threshold (P=1.18×10⁻⁶). -6Significantly associated structural variation sites were found on chromosome 13, with the 286 bp deletion site at 44.95 Mb reaching a highly significant level (P = 3.12 × 10⁻⁶). -6 ).
[0087] Step 9: Functional annotation and gene localization
[0088] Functional annotation was performed on the screened significant associated structural variations to identify the gene regions they affect. Candidate functional genes related to daily weight gain were located through gene function annotation and pathway enrichment analysis.
[0089] A 286 bp deletion (DEL) was detected at 44.95 Mb on chromosome 13 and was significantly associated with daily weight gain (P = 3.12 × 10⁻⁶). -6 (MAF=0.211). A search of the DAVID and NCBI databases revealed that this structural variant is associated with the FEZF2 gene.
[0090] Step 10: Validation by Real-Time PCR
[0091] The expression differences of the FEZF2 gene in different daily weight gain pig populations were verified by real-time quantitative PCR.
[0092] The results of quantitative real-time PCR showed that the relative expression level of FEZF2 in the high daily weight gain recombinant (>1050g / d) (3.12±0.63) was significantly higher than that in the low daily weight gain recombinant (<700g / d) (1.08±0.28), with a highly significant difference (P<0.001), indicating that the expression level of FEZF2 is positively correlated with daily weight gain.
[0093] Step 11: Molecular marker development
[0094] Molecular markers were developed based on significantly associated structural variation sites for assisted breeding of pig daily weight gain traits.
[0095] Based on the 286bp deletion site at 44.95Mb on chromosome 13, specific PCR primers were designed:
[0096] Upstream primer: 5'-AGCTGGTACCGAGCTCGGAT-3'
[0097] Downstream primer: 5'-TGCATGCCTGCAGGTCGACT-3'
[0098] Individuals with and without deletion patterns could be distinguished by PCR amplification and agarose gel electrophoresis. The association between the marker and daily weight gain was examined in an independent validation population. The results showed that the mean daily weight gain of homozygous deletion (DD) individuals was significantly higher than that of heterozygous (ND) and wild-type (NN) individuals (P<0.01), validating the effectiveness of this marker for auxiliary selection of the daily weight gain trait in pigs.
[0099] like Figure 1 As shown, the technical route of this invention demonstrates a complete technology chain from sample collection, genome assembly, graph pangenome construction, structural variation detection and typing, phenotypic data determination and processing, genome-wide association analysis, FEZF2 gene localization and validation to molecular marker development.
[0100] like Figure 2 As shown, the 32 experimental pigs included 27 Chinese local pigs (representing 24 representative breeds) and 5 foreign pig breeds (Large White, Duroc, Yorkshire, etc.). The genetic distance heatmap revealed significant genetic differences among the Chinese local pigs, representing their rich genetic diversity; the foreign pig breeds were genetically closer together, showing significant genetic differentiation from the Chinese local pigs. This population structure provides a diverse genetic basis for constructing a high-quality graph pangenome.
[0101] like Figure 3 As shown, the horizontal axis represents pig chromosomes (Chr1-Chr18), the vertical axis represents chromosome location, and the color intensity represents the density of structural variations. Structural variations are non-uniformly distributed across the entire genome, with significant differences in SV density among different chromosomes.
[0102] like Figure 4 As shown, the horizontal axis represents daily weight gain (g / d), and the vertical axis represents the number of individuals. The daily weight gain phenotypic data follow a normal distribution, with a mean of 876.8 g / d and a standard deviation of 178.5 g / d, ranging from 580 to 1250 g / d.
[0103] like Figure 5 As shown, the horizontal axis represents the pig's chromosomes (Chr1-Chr18), the vertical axis represents -log10 (P-value), and the red horizontal line represents the significance threshold (P=1.18×10). -6 Significantly associated structural variation sites were found on chromosome 13, with the 286 bp deletion site at 44.95 Mb reaching a highly significant level (P = 3.12 × 10⁻⁶). -6 ).
[0104] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for locating genes related to daily weight gain in pigs based on a graph pangenome, characterized in that, Includes the following steps: Step 1: Collect blood or tissue samples from different breeds of pigs, extract genomic DNA, perform whole-genome sequencing, and obtain sequencing data; Step 2: Perform quality control and filtering on the raw sequencing data to remove low-quality reads and adapter sequences, obtaining high-quality data; Step 3: Perform haplotype-based phase assembly on the high-quality data to obtain the primary assembled genome for each sample; Step 4: Using the pig reference genome Sscrofa11.1 as a backbone, the primary assembled genome is mounted at the chromosome level to obtain the chromosome-level genome; Step 5: Based on the chromosome-level genome, construct a pig graph pangenome. The construction includes: using Sscrofa11.1 as the backbone genome, performing a global alignment of all sample genomes with the backbone genome, integrating variation information, and constructing a graph pangenome structure containing nodes and edges. Step 6: Based on the porcine pangenome, perform structural variant typing on the second-generation sequencing data of the target population to obtain a set of structural variants; Step 7: Record the daily weight gain phenotypic data for each experimental pig, and perform normality testing and outlier handling on the phenotypic data; Step 8: Perform genome-wide association analysis on the set of structural variations and the daily weight gain phenotype data, using a mixed linear model to control for population stratification and kinship, and screen for structural variation sites that are significantly associated with daily weight gain; Step 9: Perform functional annotation on the screened significant associated structural variations to determine the gene regions they affect. Through gene function annotation and pathway enrichment analysis, locate candidate functional genes related to daily weight gain.
2. The method according to claim 1, characterized in that, In step 1, the third-generation sequencing technology is PacBio HiFi sequencing technology, with an average sequencing depth ≥30×, an average read length of 15-18kb, and a sequencing quality >99.9%.
3. The method according to claim 1, characterized in that, In step 1, the sample includes 27 Chinese local pigs and 5 foreign pig breeds. The Chinese local pigs represent 24 representative breeds, and the foreign pig breeds include Large White, Duroc, and Yorkshire.
4. The method according to claim 1, characterized in that, In step 5, the pig graph pangenome is constructed using the Minigraph-Cactus toolchain, specifically including: Short sequences less than 10kb were filtered out, and a unified sequence naming rule was adopted to ensure consistency with the chromosome numbering of the reference genome Sscrofa11.
1. Using Sscrofa11.1 as the backbone genome, the autosomal Chr1-18, X chromosome, and mitochondrial sequences were preserved, and the construction command was executed. All sample genomes were re-aligned to the pan-genome backbone. The gfatools tool was used to count the core indicators of the pan-genome.
5. The method according to claim 1, characterized in that, In step 6, the structural variant typing of the second-generation sequencing data of the target population using PanGenie software specifically includes: The gfatools tool was used to simplify the graph pangenome map and remove redundant nodes; a graph pangenome index was constructed; the second-generation sequencing data of the target population was compared with the graph pangenome index; genotyping was performed based on the comparison results to obtain the structural variation genotype file for each sample; and the bcftools tool was used to merge the genotype files of all samples.
6. The method according to claim 1, characterized in that, In step 7, the normality test and outlier handling of the phenotypic data specifically include: A linear mixture model was constructed using DMU software to correct for environmental effects. The linear mixture model is as follows: , where y is the daily weight gain phenotypic value, μ is the population mean, Sex is the sex effect, HYS is the herd-year-season effect, BW is the initial weight effect, Litter is the litter effect, and ε is the random error; after correction, outliers were removed using the three-standard-deviation method.
7. The method according to claim 1, characterized in that, In step 8, the genome-wide association analysis specifically includes: Structural variations were quality controlled, and sites with a deletion rate greater than 10% and a minimum allele frequency less than 0.05 were removed. Association analysis was performed using the rMVP software package, with daily weight gain as the response variable, structural variation genotype as the fixed effect, and variety and kinship matrix as the random effect. A mixed linear model and a FarmCPU model were used for association analysis. Structural variation sites that were significantly associated with daily weight gain were screened based on significance thresholds.
8. The method according to claim 1, characterized in that, The candidate functional gene is the FEZF2 gene, and the significantly associated structural variant site is a 286 bp deletion at 44.95 Mb on porcine chromosome 13. The minimum allele frequency of the deletion is 0.211, with a significance level of P < 3.12 × 10⁻⁶. -6 .
9. The method according to claim 1, characterized in that, The method also includes step 10: using quantitative real-time PCR to verify the expression differences of the candidate functional genes in different daily weight gain pig populations; In step 10, the quantitative PCR verification specifically includes: Pigs were divided into a high daily weight gain group and a low daily weight gain group based on their daily weight gain. The high daily weight gain group had a daily weight gain greater than 1050 g / d, while the low daily weight gain group had a daily weight gain less than 700 g / d. Total RNA was extracted from the muscle tissue of both groups and reverse transcribed to synthesize cDNA. The relative expression level of the FEZF2 gene was detected by quantitative real-time PCR. The difference in FEZF2 gene expression levels between the two groups was compared to verify the correlation between FEZF2 expression level and daily weight gain.
10. The method according to claim 9, characterized in that, The method also includes the following step 11: developing molecular markers based on the 286bp deletion site for assisted breeding of pig daily weight gain traits, wherein the detection method for the molecular markers includes: Design specific PCR primers; The upstream primer sequence is 5'-AGCTGGTACCGAGCTCGGAT-3'. The downstream primer sequence is 5'-TGCATGCCTGCAGGTCGACT-3'; PCR amplification was performed on the genomic DNA of the pigs to be tested; Amplification products were detected by agarose gel electrophoresis to distinguish between deletion-type and non-deletion-type individuals.