Application of SNP haplotype molecular marker related to residual feed intake of large white pig

By detecting SNP haplotype molecular markers in Large White pigs, the problems of insufficient marker density and poor cross-population applicability in existing technologies have been solved, achieving efficient molecular breeding and improving feed utilization and selection accuracy of Large White pig offspring.

CN122128445APending Publication Date: 2026-06-02WUHAN POLYTECHNIC UNIVERSITY

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN POLYTECHNIC UNIVERSITY
Filing Date
2026-04-30
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing technologies, the screening of genetic markers related to pig feed utilization efficiency traits suffers from problems such as insufficient marker density, poor cross-population applicability, and insufficient marker discovery in domestically bred populations, making it difficult to achieve efficient molecular breeding practices.

Method used

A SNP haplotype molecular marker associated with the remaining feed intake of Large White pigs is provided, including four sites from SNP1 to SNP4. Through whole-genome sequencing and GWAS analysis, two SNP haplotype molecular markers that affect the remaining feed intake of Large White pigs were detected, forming an LDblock module for identifying the remaining feed intake of individuals and improving feed utilization in offspring.

Benefits of technology

It improves the accuracy of early selection of feed utilization efficiency traits, has good economic benefits and broad application prospects, can explain 7.41% of phenotypic variance, and significantly improves the feed utilization rate of Large White pig offspring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122128445A_ABST
    Figure CN122128445A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of animal molecular breeding technology, specifically involving the application of SNP haplotype molecular markers related to the remaining feed intake of Large White pigs. The SNP haplotype molecular markers include four loci, SNP1 to SNP4. When the combination of SNP1 to SNP4 is TTAT, the remaining feed intake of Large White pigs is lower than that of the CCGC combination. This invention provides four SNP loci that significantly affect the remaining feed intake of Large White pigs, forming an LDblock haplotype module, and detects two haplotype molecular markers affecting the remaining feed intake trait in Large White pigs. In actual breeding processes, retaining individuals with the H002 haplotype (TTAT) is beneficial for screening individuals with high feed utilization efficiency for breeding, improving the accuracy of early selection for the feed utilization efficiency trait in Large White pigs, and has good economic benefits and broad application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of animal molecular breeding technology, specifically relating to the application of a SNP haplotype molecular marker related to the remaining feed intake of Large White pigs. Background Technology

[0002] Feed costs are the largest single input in pig farming, accounting for 60% to 75% of total farming costs, far exceeding other expenditures such as veterinary drugs, labor, and facility depreciation. The growth, development, reproduction, body maintenance, and even the functioning of the immune system in pigs are highly dependent on a continuous and sufficient supply of feed. Feed quality and feed utilization efficiency directly determine the economic benefits of the pig farming industry. my country is the world's largest producer and consumer of pork. The massive consumption base keeps the number of pigs in stock and slaughtered at a consistently high level, resulting in a substantial total demand for feed. Pig farming commonly uses compound feed with corn and soybean meal as the main energy and protein sources, leading to a year-on-year increase in the consumption of corn, soybeans, and other grain crops. Therefore, obtaining higher livestock product output per unit of feed input, i.e., improving feed efficiency (FE), has become one of the most direct and economical ways to alleviate the supply and demand imbalance of raw materials and reduce farming costs. Furthermore, improving feed efficiency also helps reduce the amount of feed required per unit of weight gain, thereby reducing the total amount of manure generated and potential greenhouse gas emissions.

[0003] From a genetic breeding perspective, feed utilization efficiency is a typical quantitative trait, simultaneously regulated by genetic factors and non-genetic factors such as feeding management, environmental temperature, and diet composition. Compared to improvements brought about by environmental management measures, which are often temporary and limited by specific feeding conditions, the improvements obtained through genetic selection are cumulative and transferable, continuing to play a role in offspring populations. Therefore, it is recognized as the most fundamental technical approach to improving feed utilization efficiency. Commonly used indicators of feed utilization efficiency include feed conversion ratio (FCR) and residual feed intake (RFI). RFI represents the difference between actual feed intake and the predicted feed intake required to sustain life and growth, reflecting the metabolic differences determined by an individual's genetic background. A negative RFI indicates that the individual consumes less feed to achieve the same weight gain and body composition goals, with feed utilization efficiency higher than the population average; conversely, a positive RFI indicates that the individual has problems with feed waste or low metabolic efficiency. RFI eliminates basal metabolic requirements related to growth and body weight composition, and can more purely reflect the genetic differences in feed utilization among individual animals. In recent years, it has been widely used in international pig breeding practices.

[0004] Genomic variation is a crucial genetic basis for individual differences in pigs. Single nucleotide polymorphisms (SNPs), the most common type of mutation, account for over 90% of genomic polymorphisms and are a major factor contributing to phenotypic differences. Feed utilization efficiency in pigs is also influenced by genetic variation. Screening and identifying relevant candidate genes or gene mutation sites can help improve feed utilization efficiency through marker-assisted selection, accelerating the process of genetic improvement in pigs.

[0005] Currently, methods for locating genetic variations in complex traits in pigs mainly rely on candidate gene methods and genome-wide association studies (GWAS). Candidate gene methods primarily target the discovery of single nucleotide polymorphisms (SNPs) associated with the target trait in known functional genes. This mainly employs an RFLP-PCR strategy, involving PCR amplification and enzyme digestion of DNA samples from each pig, followed by association analysis based on linkage disequilibrium principles to obtain relevant loci. However, this method has relatively low screening efficiency. GWAS, on the other hand, is a powerful method for discovering SNP markers associated with the target trait across the entire genome and has become the mainstream approach for genetic analysis of important economic traits in pigs. GWAS can identify SNP markers in known genes and also screen for SNP markers in new, unknown genes.

[0006] There are still many problems and shortcomings in the current field of screening genetic markers related to pig feed utilization efficiency, mainly in the following three aspects: (1) The accumulation of known genetic markers is far from sufficient to support efficient breeding. In recent years, scholars at home and abroad have reported several candidate genes and molecular marker loci related to feed utilization efficiency in different varieties and hybrid combinations. However, feed utilization efficiency, as a typical quantitative trait, is controlled by a large number of minor genes, and the contribution of each locus to the phenotype is extremely limited. In comparison, the candidate markers that have been successfully located and validated so far are far from sufficient to support precise and efficient molecular breeding practices in terms of quantity, effect size of a single locus, and cross-population applicability. It is still very urgent to continue to carry out in-depth research and validation of relevant genetic markers.

[0007] (2) The resolution of genotyping strategies based on conventional SNP chips is insufficient. Although chips (such as 60K or 80K SNP chips) widely used in current pig genome selection breeding have played an important role in genome-wide association analysis and genome prediction, their marker density is still sparse relative to the pig genome. The low marker density means that a large number of genomic regions are in coverage gaps, especially some low-frequency variations, structural variations, and important sites located in regions whose functions were unknown during chip design. This leads to insufficient statistical power of association analysis, and many truly causal genetic variations are difficult to capture and locate effectively.

[0008] (3) Genetic markers are highly specific to breeds and populations, especially in domestically bred populations where their discovery is insufficient. The effectiveness of genetic markers often depends heavily on the linkage disequilibrium structure and allele frequency distribution of a specific breed or population. Even within the same breed, there are significant differences in breeding history, population genetic structure, and inbreeding degree between different countries or regions. This results in a significant reduction in the applicability and predictive accuracy of molecular markers identified in imported populations in domestically bred populations, making it difficult to directly serve the molecular genetic improvement of feed utilization efficiency in domestic pig breeding populations.

[0009] Therefore, developing an application technology for tracking the remaining feed intake of Large White pigs has become a key technical problem that urgently needs to be solved in this field. Summary of the Invention

[0010] The purpose of this invention is to provide an application of SNP haplotype molecular markers related to the remaining feed intake of Large White pigs, thereby solving the problems existing in the prior art.

[0011] The technical solution adopted in this invention is: This invention provides an application of SNP haplotype molecular markers related to the remaining feed intake of Large White pigs. The SNP haplotype molecular markers include four sites, SNP1 to SNP4, and the nucleotide sequences containing SNP1 to SNP4 are shown in SEQ ID NO.1 to SEQ ID NO.4. SNP1 is located at 101 bp of SEQ ID NO.1, and its nucleotide is T or C; SNP2 is located at 101 bp of SEQ ID NO.2, and its nucleotide is C or T; SNP3 is located at position 101 of SEQ ID NO.3, and its nucleotide is G or A; SNP4 is located at 101 bp of SEQ ID NO.4, and its nucleotide is C or T; When the combination of SNP1 to SNP4 is TTAT, the remaining feed intake of Large White pigs is lower than that of the CCGC combination. The application refers to any one of the following (1) and (2): (1) Determine the remaining feed intake of the Large White pig; (2) Improve the feed utilization rate of offspring of Large White pigs.

[0012] Preferably, the method for determining the remaining feed intake of Large White pigs is as follows: Genomic DNA was extracted from the Large White pigs to be tested and sequenced. Determine the nucleotides of SNP1 to SNP4 in this Large White pig; When the combination of SNP1 to SNP4 is TTAT, the remaining feed intake of Large White pigs is no higher than -0.12 kg.

[0013] Preferably, the method for improving feed utilization in Large White pig offspring is as follows: Genomic DNA was extracted from the Large White pigs to be tested and sequenced. Determine the nucleotides of SNP1 to SNP4 in this Large White pig; If the combination of SNP1 to SNP4 is TTAT, then selecting Large White pigs carrying this SNP haplotype molecular marker as parents for breeding can improve the feed utilization rate of Large White pig offspring.

[0014] Preferably, the genomic DNA is derived from any one of the ear tissue, hair follicles, and blood of Large White pigs.

[0015] Compared with the prior art, the beneficial effects of the present invention are: This invention provides an application of SNP haplotype molecular markers related to the remaining feed intake of Large White pigs. The SNP haplotype molecular markers include four sites, SNP1 to SNP4, and the nucleotide sequences of SNP1 to SNP4 are shown in SEQ ID NO.1 to SEQ ID NO.4. SNP1 is located at the 101st bp of SEQ ID NO.1, and its nucleotide is T or C. SNP2 is located at the 101st bp of SEQ ID NO.2, and its nucleotide is C or T. SNP3 is located at the 101st bp of SEQ ID NO.3, and its nucleotide is G or A. SNP4 is located at the 101st bp of SEQ ID NO.4, and its nucleotide is C or T. When the combination of SNP1 to SNP4 is TTAT, the remaining feed intake of Large White pigs is lower than that of the combination CCGC. The application refers to any one of the following (1) and (2): (1) identifying the remaining feed intake of Large White pigs; (2) improving the feed utilization rate of Large White pig offspring.

[0016] This invention provides four SNP loci that significantly influence the residual feed intake of Large White pigs, forming an LDblock haplotype module, and detects two haplotype molecular markers affecting the residual feed intake trait in Large White pigs. Specifically, these are located between loci 88658287 and 88663062 on chromosome 8, spanning 4.776 kb. Haplotype H002 individuals have low residual feed intake, indicating high feed utilization efficiency; haplotype H001 individuals have high residual feed intake, indicating low feed utilization efficiency. The dominant haplotype explains 7.41% of the residual feed intake phenotypic variance. In actual breeding, retaining haplotype H002 individuals is beneficial for screening individuals with high feed utilization efficiency for breeding, improving the accuracy of early selection for the feed utilization efficiency trait in Large White pigs, and has good economic benefits and broad application prospects. Attached Figure Description

[0017] Figure 1 Genomic DNA was extracted from Large White pig ear tissue and detected by agarose gel electrophoresis. Lanes 1 to 21 represent 21 different samples.

[0018] Figure 2 Linkage disequilibrium analysis for Large White pig populations.

[0019] Figure 3 GWAS analysis of residual feed intake traits in Large White pigs. A: Manhattan plot; B: QQ plot.

[0020] Figure 4 Haplotype analysis of SNPs on chromosome 8 of Large White pigs that are associated with the trait of residual feed intake. A: LDblock obtained based on GWAS signal and linkage disequilibrium analysis; B: Correlation analysis of haplotypes H001 and H002 with residual feed intake.

[0021] Figure 5 Association analysis was performed on the genotypes of four SNP loci and the trait of remaining feed intake. A to D are chr8_88658287, chr8_88659448, chr8_88659856, and chr8_88663062, respectively. Detailed Implementation

[0022] The present invention will be further illustrated below with specific embodiments, but these embodiments do not limit the scope of the invention. Modifications or substitutions to the details and form of the technical solutions of the present invention may be made without departing from the spirit and scope of the invention, but all such modifications or substitutions fall within the protection scope of the present invention.

[0023] The inventive concept of this invention is as follows: In the identification of genetic markers for feed utilization efficiency, the number of genetic markers for Large White pigs remains insufficient, traditional microarray technology screening strategies struggle to capture all key variations, and genetic markers exhibit strong breed and population specificity, particularly in domestically bred populations where their discovery is inadequate. This invention aims to provide a genetic marker for the residual feed intake trait in domestic Large White pig populations and its application. Specifically, this application aims to supplement or solve the following problems:

[0024] 1. How to improve the accuracy of identifying genetic markers for pig feed utilization efficiency traits by using a genome sequencing full-coverage strategy to fill in SNPs across the entire genome and then performing GWAS analysis.

[0025] 2. How to perform haplotype analysis on multiple identified SNP markers, solve the problem of low effect of a single SNP marker, and screen and identify haplotypes of traits related to remaining feed intake.

[0026] This invention provides four SNP loci that significantly affect the residual feed intake trait in Large White pigs, forming an LDblock module. Two SNP haplotype molecular markers affecting residual feed intake in Large White pigs were detected on chromosome 8 of the pig genome (Sus_scrofa.11.1), as described below: (1) The chr8_88658287 locus is located at the base of chromosome 88658287 in the pig genome (Sus_scrofa.11.1). The base of the reference genome at this locus is T. The variation here is T mutated to C. The remaining feed intake of the TT genotype is greater than that of the TC genotype, indicating that the feed utilization efficiency of the TC genotype individuals is higher.

[0027] (2) The chr8_88659448 locus is located at the base of chromosome 88659448 in the pig genome (Sus_scrofa.11.1). The base of the reference genome at this locus is C. The variation here is C to T. The remaining feed intake of the CC genotype is greater than that of the CT genotype, indicating that the feed utilization efficiency of individuals with the CT genotype is higher.

[0028] (3) The chr8_88659856 locus is located at the base of chromosome 88659856 in the pig genome (Sus_scrofa.11.1). The base of the reference genome at this locus is G. The variation here is G mutated to A. Among them, the remaining feed intake of the GG genotype is greater than that of the GA genotype, indicating that the feed utilization efficiency of individuals with the GA genotype is higher.

[0029] (4) The chr8_88663062 locus is located at the base of chromosome 88663062 in the pig genome (Sus_scrofa.11.1). The base of the reference genome at this locus is C. The mutation here is C to T. The remaining feed intake of the CC genotype is greater than that of the CT genotype, indicating that the feed utilization efficiency of individuals with the CT genotype is higher.

[0030] The haplotypes H001 (CCGC) and H002 (TTAT), which are formed by SNP loci on chromosome 8, are located between loci 88658287 and 88663062, spanning 4.776 kb. The residual feed intake of haplotype H002 is less than that of haplotype H001, indicating that individuals with haplotype H002 have higher feed utilization efficiency.

[0031] In actual breeding processes, retaining the H002 haplotype individuals is beneficial for breeding new lines with low residual feed intake (i.e., screening for high feed utilization efficiency), improving the accuracy of early selection of feed utilization efficiency traits, and has good economic benefits and broad application prospects.

[0032] To enable those skilled in the art to better understand and implement the technical solutions of this invention, the invention will be further described below with reference to specific embodiments. In the description of this invention, unless otherwise specified, all reagents used are commercially available, and all methods used are conventional techniques in the art.

[0033] Example 1: Application of SNP haplotype molecular markers related to remaining feed intake in Large White pigs, as detailed below: 1. Method.

[0034] 1.1 Sample Collection and Genomic DNA Preparation Phenotypic data collection and cleaning were performed on feed utilization efficiency traits of the Large White pig population, and DNA was extracted from the collected ear tissue.

[0035] 1.1.1 Phenotypic data collection.

[0036] The Bos Intelligent Measurement Station is used to measure the breeding pigs in the pig farm. When the pigs enter the measurement equipment, the equipment door closes automatically, the equipment recognizes the pig's electronic ear tag, and then opens the feed trough door. The measurement equipment will automatically record the weight of feed (g), feeding time (s), and remaining feed weight (g) of each pig's feeding each time. The feeding records of the pigs are statistically analyzed every day, and all data is automatically uploaded to the cloud system corresponding to the measurement equipment and stored.

[0037] 1.1.2. Phenotypic data cleaning.

[0038] Quality control was conducted based on factors such as feed intake (18g), feeding duration (60s), feeding rate (18g / min~66g / min), body weight (median corrected for daily weight), and number of feedings per day (2 times / day~15 times / day).

[0039] 1.1.3 Calculation of remaining feed intake.

[0040] The mean daily feed intake (DFI), mean daily gain (ADG), midpoint metabolic weight (MBW), and fixed effects were estimated using the following linear regression model: 𝑅𝐹𝐼=𝐷𝐹𝐼−Station−𝛽1×𝑀𝐵𝑊−𝛽2×𝐴𝐷𝐺.

[0041] Wherein, 𝑅𝐹𝐼: remaining feed intake (kg); 𝐷𝐹𝐼: average daily feed intake (kg); Station: the measurement station where each individual is located; 𝛽1 and 𝛽2 are partial regression coefficients; 𝑀𝐵𝑊: the mean weight at the beginning and end of the measurement to the power of 0.75; 𝐴𝐷𝐺: average daily weight gain (kg).

[0042] 1.1.4 DNA extraction.

[0043] Genomic DNA was extracted from porcine ear tissue using a fully automated nucleotide extractor (Bayer) and its accompanying kit. The extracted DNA was then examined using 2% agarose gel electrophoresis to detect degradation, and DNA concentration was determined using a NanoDrop 2000 ultra-micro spectrophotometer. Samples that passed the above DNA quality tests were then used for library construction and sequencing.

[0044] 1.2 Sequencing of pig genomic DNA samples.

[0045] Samples that passed quality control were sent to Wuhan Shadow Gene Technology Co., Ltd. for 2× low-depth resequencing (BGI T7 genome sequencing platform). Raw sequencing data were quality controlled using fastp (v0.23.4), with the main quality control parameter being "-q 20".

[0046] 1.3 SNP typing and filling.

[0047] The quality-controlled fastq files were aligned to the pig reference genome (Sus_scrofa.Sscrofa.11.1) using BWA (v0.7.17) software. Next, samtools (v1.22.1) is used to sort, deduplicate, remove redundancy, and create an index on the obtained BAM file, thus obtaining a preliminarily processed BAM file; Genetic variations were detected and genotyping and merging were performed using the Haplotype Caller, Genotype GVCFs, and Combine GVCFs tools in GATK 4.5. The Haplotype Caller tool was used to detect variants and generate GVCF files for each sample. The GVCF files record variant information for all sites, which prepares for joint typing. The GVCF files of all samples were merged using GLnexus software, and joint typing was performed to obtain VCF files containing the original SNP variant information. SNP variant sites were extracted from the original VCF file using the SelectVariants tool, and the SNPs were strictly filtered using the VariantFiltration tool with the following filter parameters: "QUAL<30.0 || QD<2.0 || FS>60.0 || SOR>3.0 || MQRankSum<-12.5 || ReadPosRankSum<-8.0". Population SNP filtering was performed using PLINK (v1.9) software. The filtering conditions were as follows: the SNP detection rate reached 90% or higher; the minimum allele frequency (MAF) threshold was set to 0.01.

[0048] Then, using Beagle 5.5 software and high-depth (>10×) Large White pig population genome data downloaded from the NCBI database, genotyping was performed. The filled data retained R... 2 Sites with a value ≥0.6 were then filtered using the same criteria in PLINK, and the resulting high-quality SNP sites were used for subsequent analysis.

[0049] 1.4 Chaining Disequilibrium (LD) Analysis.

[0050] PopLDdecay software was used to evaluate the degree of genome-wide linkage disequilibrium (LD) in a population, aiming to assess the efficiency and accuracy of association analysis.

[0051] 1.5 Genome-wide association analysis (GWAS).

[0052] GEMMA was used to perform genome-wide association analysis (GWAS) on feed utilization efficiency traits, with batch number, age at measurement, and weight gain during the measurement period as covariates. Boferroni correction was used to adjust the p-values ​​of the GWAS analysis. Because Boferroni correction is very strict, the effective number N of independent tests was calculated using the R package simpleM (https: / / github.com / LTibbs / SimpleM), with a threshold of 1 / N for the GWAS signal. The significance of each SNP was assessed using a likelihood ratio test. The genomic inflation factor (λ) for the test statistics was calculated using R (v4.3) software as the ratio between the median of the observed p-value distribution and the theoretical median. It was ensured that there was no significant stratification in the population, and that the results were not false positives due to population structure. The mixed linear model is as follows:

[0053] .

[0054] In the above model, y Indicates phenotypic value; W It is an n x (w + 1) matrix that includes the intercept and covariates; Then it is a vector of (w + 1) × 1, representing the effect size of the covariate; Gs It is an n × 1 vector representing the genotype of a certain locus. The value of each item is usually 0, 1, or 2 (the copy number of the allele). gamma This is a scalar, representing the effect size of the genotype at the target locus; g This is due to the accumulation effect; epsilon This is the residual.

[0055] 1.6. Haplotype construction.

[0056] Linkage disequilibrium association analysis was used to identify haplotypes associated with specific phenotypic traits. Data from genome-wide significant loci obtained through GWAS were statistically analyzed and used to construct a linkage disequilibrium module. LDBlockShow was used for linkage disequilibrium module analysis and result visualization. The plink software package was used for haplotype inference and sample genotyping, and R software was used for result visualization analysis.

[0057] 2. Results.

[0058] 2.1 Phenotypic data collection and ear tissue sample DNA preparation.

[0059] (1) 234,377 feeding records of 673 Large White pigs (weight range of 25kg to 115kg during fattening period) were collected using the Bos Intelligent Measurement Station. Quality control was carried out based on conditions such as feed intake (18g), feeding duration (60s), feeding rate (18g / min to 66g / min), weight (median corrected for daily weight), and number of feedings per day (2 times / day to 15 times / day). After cleaning, 138,908 data points were obtained, and the number of samples that met the conditions was 360.

[0060] (2) The final residual intake (RFI) ranges from -0.42 kg to 0.59 kg, as shown in Table 1 below.

[0061] Table 1. Descriptive statistics of remaining feed intake of 360 Large White pigs (3) Genomic DNA was extracted from pig ear tissue using a fully automated nucleotide extractor (Zhongke Bayer) and its accompanying reagent kit. Since low-quality DNA can affect sequencing results, 2% agarose gel electrophoresis was used to detect DNA degradation. The electrophoresis results are shown below. Figure 1 The DNA sample bands were clear and undegraded. Detection using an ultra-micro spectrophotometer showed that the lowest concentration of DNA in the extracted ear tissue samples was 108.75 ng / μL, the highest was 1559.51 ng / μL, and the average was 526.84 ng / μL. OD 260 / 280 The minimum value was 1.65, the maximum value was 2.11, and the average value was 1.83. These test results indicate that the DNA extracted from the ear tissue of the Large White pig population is of good quality and meets the requirements for library construction and sequencing.

[0062] 2.2 Sample sequencing.

[0063] 360 qualified samples were sent to Wuhan Shadow Gene Technology Co., Ltd. for 2× low-depth resequencing (BGI T7 genome sequencing platform). The raw sequencing data were quality controlled using fastp (v0.23.4), with the main quality control parameter being "-q 20".

[0064] 2.3 SNP typing.

[0065] A total of 2,274,704 high-quality SNP loci were obtained for subsequent analysis.

[0066] 2.4 Chaining Imbalance Analysis.

[0067] PopLDdecay software was used to evaluate the degree of genome-wide linkage disequilibrium (LD) in a population, aiming to assess the efficiency and accuracy of association analysis.

[0068] 2.5 Genome-wide association analysis.

[0069] Calculations showed that the λ value for the population residual feed intake trait of the present invention was 0.97, indicating that the population did not have obvious stratification and the results did not produce false positives due to population structure.

[0070] 2.6. Haplotype construction.

[0071] Linkage disequilibrium association analysis was used to identify SNP haplotypes associated with specific phenotypic traits. Data from genome-wide significant loci obtained through GWAS were statistically analyzed and used to construct LDblock (linkage disequilibrium module). LDBlockShow was used for linkage disequilibrium module analysis and the results were visualized. The plink software package was used for haplotype inference and sample genotyping, and R software was used for result visualization analysis.

[0072] 2.7 Chaining Imbalance Analysis Depend on Figure 2 It can be seen that in linkage disequilibrium (LD) analysis, as the marker distance between paired SNPs increases, r... 2 The value tends to decrease, and r is observed in the first 100Kb range. 2 The value shows a rapid downward trend. In the study population, the linkage disequilibrium was at approximately 70 kb when r... 2 The value decayed to 0.2.

[0073] 2.8 GWAS signal identification.

[0074] GWAS analysis results show that ( Figure 3 Four significant SNP signals were focused on chromosome 8 of the Large White pig, and the QQ plot further confirmed the reliability of the analysis results. Detailed information on the four loci is shown in Table 2.

[0075] Table 2. Remaining feed intake signal peak locations in the Large White pig population. Note: "-" in Table 2 indicates that this item is not available.

[0076] The nucleotide sequences containing the four SNP sites in Table 2 are shown in SEQ ID NO.1 to SEQ ID NO.4, with the bolded parts in SEQ ID NO.1 to SEQ ID NO.4 indicating the mutation sites.

[0077] SEQ ID NO.1: AAAAAGTCTTGATATAATGAGTCCCTTCATTCTTTTGTTAGTTAGGGTAATGCCAGCTGCTATAACAAATAAACCACACATTTTTGATGGCTTTAAACAGTTGCACATTTAAGTTTTAAATTATATAACAGTCCACTGTTGGTGTCCCTGGCTGAGTGGTGGGCTTTCTTCCACATGTTGATTCTTTAGTCCAGATTCCTT。

[0078] SEQ ID NO.2: TTGTTTATGGAGTTCTTCATCTGTGGACAGTGGCAATTGGTGAGATTTAGATTAAACATCAGTAGTATTGAAATTTAAAATGTCTCAGTGAAAAAACTCCCATCACATTGCCATGGTACTGGAAGAAGCAGAAGGCAGGTGTTTAATGCTTGAAATCAGTGGGAACACTGTGTACCTTTATGTGACAGTCTGCTTGTCACA。

[0079] SEQ ID NO.3: CCCACTCACGGTGTGCTAGAGAGACAGTGATGGCACTAAATGACAATGAAGCAAAAATCACTGAAAACATACAAAGCAATTCATATGATTCCACTAATCTGTAATATGATTGTTTCAAGTCTGGAAAACTATTAAGAACATATATGTGTAGTTCATTTCCTGCTAGAAATGTCAGAAATCACATGCCCAGGTAGTACAGTT。

[0080] SEQ ID NO.4: AAAGCTTTCCAATAAACATTATATTCCACAATTACATTTTCTCTGAAAATTAATGGAGTTTGAGGAGGCACATGCATAAGAGAGAGAGGGTACACTCTACACTTCAGAGGCTTAATGCTCTGACTTCAAGGTAGGATAGATTGGTGGCTGGGTTTTGGGTTTCCACTTACCAGCTGTGCATCCTTGAACAGTTATTAAA。

[0081] 2.9 Haplotype Identification and Module Analysis.

[0082] Based on the attenuation distance of GWAS signals and chain imbalance analysis, this invention detects one LDblock ( Figure 4 (A), and detected two haplotypes associated with the remaining feed intake trait ( Figure 4 Haplotypes H001 (CCGC) and H002 (TTAT) are located between loci 88658287 and 88663062, spanning 4.776 kb, and are composed of SNP1 to SNP4. Haplotype H002 is significantly associated with the trait of low residual feed intake (P < 0.0001). Figure 4 (B). As shown in Table 3, the average residual feed intake of Large White pigs with haplotype H002 was -0.15 kg, significantly lower than that of Large White pigs with haplotype H001. This indicates that the feed utilization efficiency of haplotype H002 individuals was significantly higher than that of haplotype H001 individuals. Haplotype H002 can serve as a potential molecular marker and can be used for breeding pigs with high feed utilization efficiency.

[0083] Table 3. Association analysis between different haplotypes of RFI_block and residual feed intake phenotype in Large White pigs. 2.10 Correlation analysis of SNP loci and pig residual feed intake phenotype.

[0084] Association analysis between genotypes and residual feed intake at four significant loci was performed using R software. The results are shown in Table 4. However, since the number of homozygous mutants in this resource population was only 1-2, statistical analysis was not feasible, making it impossible to stably evaluate their trait effects. To ensure the accuracy and reliability of the experimental data, this invention only used wild-type homozygous and heterozygous genotypes with sufficient samples for phenotypic association analysis. The association analysis diagram between genotypes and phenotypes at significant loci is shown below. Figure 5 As shown.

[0085] Table 4. Association analysis of genotypes at four SNP loci with residual feed intake trait. Table 4 shows that the correlation between the four SNP loci and the remaining feed intake of pigs was extremely significant (P<0.0001). The remaining feed intake of the mutant genotypes at each SNP locus was lower than that of the reference genotype. The four loci chr8_88658287_T>C, chr8_88659448_C>T, chr8_88659856_G>A, and chr8_88663062_C>T showed high linkage and formed a haplotype block. The remaining feed intake of haplotype H001 (CCGC) was 0.02±0.01, and that of haplotype H002 (TTAT) was -0.15±0.03. The remaining feed intake of haplotype H001 was significantly higher than that of haplotype H002, making haplotype H002 the dominant haplotype.

[0086] Using R software, we analyzed the genetic effects of dominant haplotypes on residual feed intake (RFI) phenotypic data of haplotypes and 360 individuals using an additive linear model. After matching haplotype data with individual phenotypes, we constructed an additive regression model to test the overall genetic effect of dominant haplotypes and calculate the phenotypic variance explained. The results showed that dominant haplotypes had a highly significant effect on residual feed intake and could effectively explain phenotypic variance. The specific model is as follows:

[0087] .

[0088] in, Let be the phenotypic value of the i-th individual; mu It is the group mean; G i For the genotype effect of the i-th individual; e i Let be the random residual effect of the i-th individual.

[0089] The results showed that the additive effect value of the dominant haplotype H002 was -0.089. Under the additive linear model, the dominant haplotype explained a sum of squares of 0.9333, a residual sum of squares of 11.6579, and a total phenotypic sum of squares of 12.5912. The calculated variance of the residual feed intake phenotypic explained by this dominant haplotype was 7.41% (0.9333 / 12.5912*100%). In practical breeding, retaining individuals carrying superior haplotypes is beneficial for screening individuals with high feed utilization efficiency, improving the accuracy of early selection, and has good economic benefits and broad application prospects.

[0090] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0091] The embodiments described above are merely examples of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.

Claims

1. The application of a SNP haplotype molecular marker associated with residual feed intake in Large White pigs, characterized in that, The SNP haplotype molecular markers include four sites, SNP1 to SNP4, and the nucleotide sequences containing SNP1 to SNP4 are shown in SEQ ID NO.1 to SEQ ID NO.

4. SNP1 is located at 101 bp of SEQ ID NO.1, and its nucleotide is T or C; SNP2 is located at 101 bp of SEQ ID NO.2, and its nucleotide is C or T; SNP3 is located at position 101 of SEQ ID NO.3, and its nucleotide is G or A; SNP4 is located at 101 bp of SEQ ID NO.4, and its nucleotide is C or T; When the combination of SNP1 to SNP4 is TTAT, the remaining feed intake of Large White pigs is lower than that of the CCGC combination. The application refers to any one of the following (1) and (2): (1) Determine the remaining feed intake of the Large White pig; (2) Improve the feed utilization rate of offspring of Large White pigs.

2. The application as described in claim 1, characterized in that, The method for determining the remaining feed intake of Large White pigs is as follows: Genomic DNA was extracted from the Large White pigs to be tested and sequenced. Determine the nucleotides of SNP1 to SNP4 in this Large White pig; When the combination of SNP1 to SNP4 is TTAT, the remaining feed intake of Large White pigs is no higher than -0.12 kg.

3. The application as described in claim 1, characterized in that, The following methods can be used to improve feed utilization in Large White pig offspring: Genomic DNA was extracted from the Large White pigs to be tested and sequenced. Determine the nucleotides of SNP1 to SNP4 in this Large White pig; If the combination of SNP1 to SNP4 is TTAT, then selecting Large White pigs carrying this SNP haplotype molecular marker as parents for breeding can improve the feed utilization rate of Large White pig offspring.

4. The application as described in claim 2 or claim 3, characterized in that, The genomic DNA was derived from any one of the ear tissue, hair follicles, or blood of the Large White pig.