Shanxi white pig whole genome SNP chip and application thereof

By developing a whole-genome SNP chip for Jinfen White Pig, the problem of insufficient accuracy of existing chips in detecting local pig breeds and breeding pig breeds in my country has been solved, achieving efficient and accurate genome selection and breeding results, and improving the genetic progress of pig breeds.

CN121852547APending Publication Date: 2026-04-14SHANXI AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-06-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing genome detection chips are not accurate enough when detecting local and domestic pig breeds in my country, which affects the efficiency and accuracy of molecular breeding.

Method used

A whole-genome SNP chip for Jinfen White pigs was developed. By resequencing the whole genome of Jinfen White pigs and their four parents, 55,957 Jinfen White pig-specific SNP loci were screened. A high-performance liquid phase gene chip detection kit was designed and manufactured for the genome selection, parentage identification, pig cluster analysis, kinship analysis, linkage map construction and gene localization of Jinfen White pigs.

Benefits of technology

It has improved the efficiency and accuracy of testing for local and bred pig breeds in my country, and enhanced the accuracy of breeding, especially in the accuracy of breeding values ​​for traits such as the number of teats in sows, number of piglets per litter, initial litter weight of piglets, age at which piglets reach 100 kg body weight, and backfat thickness, which has greatly accelerated genetic progress.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_1
    Figure SMS_1
  • Figure SMS_2
    Figure SMS_2
  • Figure SMS_3
    Figure SMS_3
Patent Text Reader

Abstract

The invention discloses a Shanxi white pig whole genome SNP chip and application thereof, and on the basis of Shanxi white pig whole genome re-sequencing data, through GWAS screening verification, invalid SNP markers are removed, and 55957 Shanxi white pig specific SNP sites are determined. The Shanxi white pig whole genome SNP can be used for paternity test when Shanxi white pigs and Shanxi white pigs are parent varieties, pig clustering analysis when Shanxi white pigs and Shanxi white pigs are parent varieties, pig genetic relationship analysis, pig linkage map construction and gene localization and pig molecular breeding material background selection, and the pig selection and breeding process in China can be greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gene molecular breeding, specifically to a whole genome SNP chip of Jinfen White Pig and its application. Background Technology

[0002] SNPs, as the third generation of DNA genetic markers following restriction fragment length polymorphisms and simple sequence repeats, are characterized by their wide distribution, large number, high genetic stability, and ease of rapid and automated detection. High-throughput molecular marker technologies based on SNPs mainly include sequencing-based molecular marker technologies and gene chip-based molecular marker technologies. While sequencing-based molecular marker technologies offer high sequencing depth and wide sequencing range, sequencing short sequence fragments requires a reference genome sequence. Sequences located in repetitive regions or not matched to the reference genome are difficult to detect and analyze, and the cost of sequencing hinders its widespread adoption in molecular breeding.

[0003] Gene chip-based molecular technology, due to its high efficiency, high throughput, miniaturization, and automation, is widely used in species evolution, gene mapping, and molecular breeding. Gene chips, also known as biochips, were first proposed in 1991. In 1996, a laboratory at Stanford University in the United States synthesized a DNA array and launched the world's first commercially available gene chip, sparking a research boom in gene chips. The principle of gene chips is to perform molecular hybridization between the target sequence of the DNA sample to be tested and a known nucleic acid probe immobilized on a substrate, and determine the target sequence of the sample through the resulting hybridization pattern. When a fluorescently labeled DNA sample to be tested pairs with a nucleic acid probe on the chip, the DNA chip produces detectable fluorescent spots; the more bases paired, the stronger the fluorescence.

[0004] With the advent of commercially available microarrays for livestock and poultry, genomic selection has been widely applied in animal breeding. Currently available genomic detection chips (such as Illumina's Porcine SNP60 60K chip, GeneSeek's Porcine 80 KXT SNP chip, and Affymetrix's Axiom PorcineHD 70K chip) are mainly products developed based on imported breeds such as Duroc, Landrace, and Large White. Due to differences between domestic and foreign pig breeds, the accuracy of genomic selection using these chips for local or bred breeds in my country is significantly reduced, severely impacting the accuracy of genomic selection. To ensure the detection effectiveness for local and bred pig breeds in my country and improve the efficiency of molecular breeding, selecting suitable SNP chip detection kits has become an urgent priority to continuously meet the needs of large-scale pig breeding. Summary of the Invention

[0005] The purpose of this application is to provide a whole genome SNP chip for Jinfen White Pig to address the urgent issue of selecting a suitable SNP chip in order to ensure the detection effect of local and bred pig breeds in my country and improve the efficiency of molecular breeding.

[0006] The technical solution of the present invention is: a whole genome SNP chip of Jinfen White Pig, the SNP chip is composed of 55957 Jinfen White Pig-specific SNP sites, the 55957 Jinfen White Pig-specific SNP sites are: SEQ ID NO.1-SEQ ID NO.55957.

[0007] The Jinfen White Pig, mentioned above, is a national-level new breed independently developed by Shanxi Agricultural University over more than 20 years. It was approved by the National Livestock and Poultry Genetic Resources Committee in 2014 and selected as a national leading breed twice, in 2015 and 2016. This breed contains 6.25% Masculine pig bloodline, 3.125% Taihu pig bloodline, 40.625% Landrace pig bloodline, and 50% Large White pig bloodline. The Jinfen White Pig combines the characteristics of all four parent breeds, resulting in high litter size, excellent meat quality (significantly higher meat color and marbling than Duroc, Landrace, and Large White, with finer muscle fibers and higher intramuscular fat content), rapid growth (comparable to other Duroc, Landrace, and Large White breeds), strong disease resistance (adapting from the northernmost part of Shanxi (winter temperatures -20°C) to the southernmost part (summer temperatures 35°C) during the breeding process), tolerance to roughage, and good disease resistance.

[0008] The successful development of the Jinfen White Pig genome selection chip can effectively solve the problem of difficulties in developing "Chinese chips" for domestically bred pig breeds, laying a solid foundation for the breeding of Chinese pig breeds represented by the Jinfen White Pig.

[0009] This invention also provides a method for fabricating a whole-genome SNP chip of Jinfen White Pig, comprising the following steps:

[0010] S1. Whole genome resequencing detection:

[0011] Whole genome resequencing was performed on Jinfen White Pig and its four parent breeds to obtain whole genome resequencing data of Jinfen White Pig and its four parent breeds;

[0012] S2, Quality Control of Whole Genome Resequencing Data:

[0013] Quality control was performed on the whole genome resequencing data of Jinfen White pigs and their four parents to obtain quality control data of Jinfen White pigs and their four parents. Specifically, low-quality reads, undetected reads, and adapter sequence data in the whole genome resequencing data of Jinfen White pigs and their four parents were filtered out using the FastP software. The filtering conditions were the default parameters of the FastP software.

[0014] S3, Quality control data compared with the pig Susscrofa 11.1 reference genome:

[0015] The quality control data of Jinfen White Pig and its four parents were compared with the pig Susscrofa11.1 reference genome using minimap2 software. Then, duplicate gene loci were removed from the quality control data of Jinfen White Pig and its four parents using GATK4 software to obtain the gene sequences of Jinfen White Pig and its four parents.

[0016] S4. Obtain the set of SNP mutation sites of Jinfen White Pig and its four parents.

[0017] After re-aligning the quality control data of Jinfen White pigs and their four parents using the BaseRecalibration module of GATK4 software, GVCF data of Jinfen White pigs and their four parents were generated using the HaplotypeCaller module of GATK4 software. Locus re-alignment was then performed on the GVCF data of Jinfen White pigs and their four parents. Finally, locus quality filtering was performed on the GVCF data of Jinfen White pigs and their four parents using the VariantFilteration module of GATK4 software to obtain the set of SNP variant loci of Jinfen White pigs and their four parents. The filtering conditions were: SQR <= 10, QD >= 2, MQ >= 40.0, QUAL >= 30.0, MQRandSum >= 12.5, ReadPosRankSum >= -8.0, InbreedingCoeff >= 0.8, and FS <= 60.0.

[0018] S5. Screening for Jinfen White pig-specific SNP loci that appear only in the Jinfen White pig population:

[0019] The SNP variant sites of Jinfen White pigs and their four parents were screened to obtain Jinfen White pig-specific SNP sites. The screening criteria were: allele frequency > 0.1. The Jinfen White pig-specific SNP sites were compared with sites in the NCBI SNP database. Sites were optimized through GWAS association analysis of Jinfen White pig production performance. Finally, 55,957 Jinfen White pig-specific SNP sites that appeared only in the Jinfen White pig population were obtained.

[0020] S6. Integrate the 55,957 Jinfen White Pig-specific SNP loci obtained in step S5 to obtain the whole genome SNP chip of Jinfen White Pig.

[0021] Another objective of this invention is to provide an application of the whole genome SNP chip of Jinfen White Pig in parentage identification of Jinfen White Pig and when Jinfen White Pig is the parent breed.

[0022] Another objective of this invention is to provide an application of the whole genome SNP chip of Jinfen White pig in pig cluster analysis and pig kinship analysis when Jinfen White pig is the parent.

[0023] Furthermore, the parent breeds of Fenbai pig are: Mashen pig, Taihu pig, Large White pig, and Landrace pig.

[0024] Another objective of this invention is to provide an application of the whole genome SNP chip of Jinfen White Pig in the construction of pig linkage maps and gene localization.

[0025] Another objective of this invention is to provide an application of the whole genome SNP chip of Jinfen White Pig in the background selection of pig molecular breeding materials.

[0026] The beneficial effects of this invention are as follows:

[0027] 1. The Jinfen White Pig SNP chip of this invention uses four parent populations—Jinfen White Pig and its broodstock (Masshin Pig, Taihu Pig, Landrace Pig, and Large White Pig)—as the detection population. The chip detection kit is designed to be breed-specific and representative. This chip detection kit mainly includes the Jinfen White Pig-specific SNP sites detected by this invention, effectively improving the detection efficiency of local and bred pig breeds in my country, and has significant application value in the field of pig molecular breeding.

[0028] 2. The Jinfen White Pig SNP chip of this invention effectively improves the accuracy of Jinfen White Pig breeding. Using this chip for Jinfen White Pig genome selection, the breeding values ​​for traits such as sow teat number, litter size, piglet birth weight, age at 100kg weight, and backfat thickness reached accuracy values ​​of 0.5432, 0.6101, 0.4827, 0.5375, and 0.6755, respectively, greatly improving breeding accuracy and accelerating genetic progress. Attached Figure Description

[0029] Figure 1 This is a distribution map of SNPs / genes on chromosomes in Example 1;

[0030] Figure 2 This is a diagram illustrating the high-performance liquid chromatography gene chip detection kit for Jinfen white pigs used in Example 1.

[0031] Figure 3 This is a flowchart of the method for fabricating the whole genome SNP chip of Jinfen White Pig in Example 2;

[0032] Figure 4 This is a graph showing the results of agarose gel electrophoresis detection of genomic DNA in Example 3;

[0033] Figure 5 This is a distribution diagram of SNPs on each chromosome before and after quality control in Example 3;

[0034] Figure 6 This is the minimum allele frequency distribution map of SNP markers in Example 3;

[0035] Figure 7 This is a graph showing the ROH analysis results of the Jinfen White Pig population in Example 3;

[0036] Figure 8 This is a graph showing the principal component analysis results of the Jinfen White Pig population in Example 3;

[0037] Figure 9 This is a G-matrix heatmap of Jinfen White Pig in Example 3;

[0038] Figure 10 This is the IBS matrix heatmap of Jinfen White Pig in Example 3. Detailed Implementation

[0039] Example 1

[0040] This embodiment is a whole genome SNP chip of Jinfen White Pig. The SNP chip consists of 55,957 Jinfen White Pig-specific SNP sites, which are: SEQ ID NO.1-SEQ ID NO.55957.

[0041] The Jinfen White Pig is a new breed developed through 20 years of complex crossbreeding using the Ma Shen Pig, Erhualian Pig, Landrace Pig, and Large White Pig as parents. Currently, commonly used genome sequencing chips, both domestically and internationally, are mainly products developed based on genomic marker sites from imported breeds such as Duroc, Landrace, and Large White. Due to differences between domestic and international pig breeds, the accuracy of using these chips for genomic selection of local or cultivated breeds in my country is severely affected, significantly reducing the accuracy of genomic selection.

[0042] This embodiment, based on the genetic background of Jinfen White pigs, uses genome resequencing technology to perform whole-genome resequencing analysis on Jinfen White pigs and their parents. It screens for germplasm-specific SNPs of Jinfen White pigs by allele frequency, develops a high-performance liquid gene chip detection kit for Jinfen White pigs, and applies it to the genome selection work of Jinfen White pigs, providing a foundation for germplasm identification, genome selection, germplasm improvement, and new breed breeding of Jinfen White pigs. Specifically, it includes the following steps:

[0043] S1. Screening of specific variant sites in Jinfen White pigs;

[0044] DNA was collected from ear tissues of 30 Jinfen white pigs for genome resequencing analysis. Libraries were constructed using the TruSeq Library Construction Kit and sequenced using the Illumina HiSeq platform. After quality control of the raw data, genome alignment and annotation analysis were performed on the clean data.

[0045] S2. Genetic diversity analysis of Jinfen White Pigs:

[0046] The resequencing yielded 25,421,850 SNPs, with a number of polymorphic markers (NSNPs) of 15,862,511, a polymorphic marker ratio (PN) of 0.62, an observed heterozygosity (Ho) of 0.31, and an expected heterozygosity (He) of 0.30.

[0047] Statistical analysis of resequencing data from five varieties yielded 47.32 Gb of raw reads and 5772.82 Gb of raw data. The average Q20 and Q30 of the raw data were 94.75% and 88.70%, respectively, with an average GC content of 42.72%. After quality control, 45.63 Gb of clean reads and 5584.17 Gb of clean data were obtained. The average Q20 and Q30 of the clean data were 96.26% and 90.40%, respectively, with an average GC content of 42.48% and an average effective value of 95.76%. Based on whole-genome resequencing, 15,862,511 SNP markers were obtained after quality control, and 2,486 ROHs were detected, with an average incrossing coefficient of 0.0558.

[0048] S3. Methods for mining specific SNPs in Jinfen White Pigs:

[0049] Based on the whole-genome sequencing results of Mashan pigs and Jinfen white pigs, resequencing data from 17 Large White pigs, 30 Landrace pigs, and 40 Taihu pigs were downloaded from NCBI, with an average sequencing depth of 9.81X. Specific loci were obtained based on gene frequency and specificity. Enrichment functional analysis and KEGG pathway analysis were performed on genes annotated with missense SNPs in CDS using the DAVID analysis platform (https: / / david.ncifcrf.gov / ). The results are shown in Table 1.

[0050] Table 1. Statistical results of resequencing data of Jinfen White pigs and their parents.

[0051]

[0052]

[0053] S4. Acquisition of specific SNP information of Jinfen White Pig:

[0054] Based on allele frequency screening, 87,366 specific SNPs of Jinfen White pigs were obtained. Annotation of these SNPs revealed that they were mainly distributed in intergenic regions and intronic regions, with 30,779 (35.23%) and 40,535 (46.40%) NPs respectively; followed by upstream and downstream regions, totaling 8,295 (9.50%) NPs.

[0055] The results of SNP / gene distribution on chromosomes are shown below. Figure 1 (The outermost ring represents chromosomes, with values ​​in Mb, indicating the location of SNPs / genes on the chromosome. A window size of 500kb is used to screen for SNP / gene density; the intensity of the color represents the number of SNPs / genes within the window.) Figure 1 It can be seen that the specific SNPs / genes of Jinfen White Pigs are evenly distributed on the chromosomes, with good SNP / gene coverage. Specific SNP loci are more abundant on chromosomes 1, 2, 6, and 15.

[0056] Table 2 shows the statistical information on the number of SNPs / genes. The table indicates that the total length of the pig chromosome is approximately 2.44 Gb. A total of 87,366 SNPs were obtained, averaging one specific SNP per 28.44 kb. The gene density distribution is even, with specific SNPs located on 25,606 genes, averaging one gene per 95.17 kb.

[0057] Table 2. Statistical information on different chromosomes

[0058]

[0059]

[0060] S5. Development of a high-performance liquid chromatography gene chip detection kit for Jinfen white pigs:

[0061] The selected specific SNP loci were annotated on the pig genome (Sscrofa 11.1). Optimization was performed based on the accuracy of the SNP effect results and their distribution in the genome, resulting in 55,957 specific SNP loci for Jinfen White pigs. Based on this, a high-performance liquid chromatography-mass spectrometry (HPLC-MS / MS) detection kit for Jinfen White pigs was developed. Figure 2 As shown.

[0062] Example 2

[0063] This embodiment describes a whole-genome SNP chip for Jinfen White pigs, such as... Figure 3 As shown, the method for fabricating a whole-genome SNP chip of Jinfen White Pig includes the following steps:

[0064] S1. Whole genome resequencing detection:

[0065] Whole genome resequencing was performed on Jinfen White Pig and its four parent breeds to obtain whole genome resequencing data of Jinfen White Pig and its four parent breeds;

[0066] S2, Quality Control of Whole Genome Resequencing Data:

[0067] Quality control was performed on the whole genome resequencing data of Jinfen White pigs and their four parents to obtain quality control data of Jinfen White pigs and their four parents. Specifically, low-quality reads, undetected reads, and adapter sequence data in the whole genome resequencing data of Jinfen White pigs and their four parents were filtered out using the FastP software. The filtering conditions were the default parameters of the FastP software.

[0068] S3, Quality control data compared with the pig Susscrofa 11.1 reference genome:

[0069] The quality control data of Jinfen White Pig and its four parents were compared with the pig Susscrofa11.1 reference genome using minimap2 software. Then, duplicate gene loci were removed from the quality control data of Jinfen White Pig and its four parents using GATK4 software to obtain the gene sequences of Jinfen White Pig and its four parents.

[0070] S4. Obtain the set of SNP mutation sites of Jinfen White Pig and its four parents.

[0071] After re-aligning the quality control data of Jinfen White pigs and their four parents using the BaseRecalibration module of GATK4 software, GVCF data of Jinfen White pigs and their four parents were generated using the HaplotypeCaller module of GATK4 software. Locus re-alignment was then performed on the GVCF data of Jinfen White pigs and their four parents. Finally, locus quality filtering was performed on the GVCF data of Jinfen White pigs and their four parents using the VariantFilteration module of GATK4 software to obtain the set of SNP variant loci of Jinfen White pigs and their four parents. The filtering conditions were: SQR <= 10, QD >= 2, MQ >= 40.0, QUAL >= 30.0, MQRandSum >= 12.5, ReadPosRankSum >= -8.0, InbreedingCoeff >= 0.8, and FS <= 60.0.

[0072] S5. Screening for Jinfen White pig-specific SNP loci that appear only in the Jinfen White pig population:

[0073] The SNP variant sites of Jinfen White pigs and their four parents were screened to obtain Jinfen White pig-specific SNP sites. The screening criteria were: allele frequency > 0.1. The Jinfen White pig-specific SNP sites were compared with sites in the NCBI SNP database. Sites were optimized through GWAS association analysis of Jinfen White pig production performance. Finally, 55,957 Jinfen White pig-specific SNP sites that appeared only in the Jinfen White pig population were obtained.

[0074] S6. Integrate the 55,957 Jinfen White Pig-specific SNP loci obtained in step S5 to customize a whole-genome SNP chip detection kit for Jinfen White Pig.

[0075] Example 3

[0076] This embodiment describes the application of a whole genome SNP chip of Jinfen White Pig in parentage identification when Jinfen White Pig is the parent breed, based on the whole genome SNP chip detection kit of Jinfen White Pig in Example 2.

[0077] The above application includes the following steps:

[0078] S1. Experimental Animals and Sample Collection:

[0079] Use ear clippers to cut 2-3 pieces of ear tissue and place them in 2mL cryovials. Place the cryovials in liquid nitrogen and bring them back to the laboratory. Transfer them to a -80°C freezer for storage in preparation for DNA extraction.

[0080] The Jinfen White pigs used in this embodiment were obtained from Datong Breeding Pig Farm, Pingding Huayi Breeding Pig Farm, Jiangxian Jialv Breeding Cooperative, and Taigu Riyichang Technology Breeding Co., Ltd. All experimental animals were in good health, raised under conventional conditions, with free access to water and feed, and vaccinated according to the farm's immunization program.

[0081] S2. Gene chip detection and SNP genotyping, including the following steps:

[0082] S2-1. Use PLINK software to calculate genetic diversity indices, including: minimum allele frequency (MAF), observed heterozygosity (Ho), expected heterozygosity (He), and polymorphic marker ratio (P). N ),

[0083] Minimum allele frequency calculation:

[0084] Gene frequency refers to the proportion of a particular gene at a specific locus in a biological population. Minimum allele frequency, on the other hand, refers to the frequency of an uncommon allele at a particular locus within a population, ranging from 0 to 0.5. It reflects the genetic diversity of a population and is calculated using the following formula:

[0085]

[0086] P i It is the frequency of allele i, N i / i N represents the number of samples carrying the i / i genotype. i / m It represents the number of samples carrying the i / m genotype, where m = 1 to n represents the n distinct multiple alleles of allele i.

[0087] Heterozygosity calculation:

[0088] Heterozygosity is an important indicator for measuring genetic variation in a population. Expected heterozygosity (He) refers to the probability that any individual in the population is heterozygous at any locus; observed heterozygosity (Ho) refers to the proportion of individuals in the population that are heterozygous at a particular locus, and its value ranges from 0 to 1. The calculation formula is as follows:

[0089]

[0090]

[0091] n is the total number of individuals in the population, N is the total number of loci, and H is the total number of loci. k P represents the number of heterozygous individuals at locus k. ki The frequency of allele i at locus k.

[0092] Polymorphic marker ratio (P) N )calculate:

[0093] Polymorphic marker ratio (P) N The term "polymorphic locus" refers to the proportion of polymorphic loci detected in the target population out of the total number of loci. A polymorphic locus is defined as a gene locus with a minimum allele frequency (MAF) greater than or equal to 0.05. This study first used PLINK software to calculate the MAF for each locus, and then used a self-developed R script to calculate P... N We calculate P using the following formula. N :

[0094]

[0095] M represents the number of polymorphic sites, and N represents the total number of sites.

[0096] S2-2, Principal Component Analysis:

[0097] The main purpose of Principal Component Analysis (PCA) is to transform multiple complex variables into a few principal comprehensive variables. The analytical approach involves reducing data dimensionality by converting multiple differential components into principal components, thus simplifying the data. PCA is performed using PLINK software. The `--pca4` command is used to analyze the quality-controlled SNP data. The top two eigenvector values ​​(PC1 and PC2) are plotted, and a scatter plot is created using the R package `ade4`.

[0098] S2-3, Construction of G matrix:

[0099] The G matrix, first proposed by Vanraden, is a genomic relationship matrix estimated based on SNP information obtained from SNP microarrays. It can replace the A matrix for calculating genomic relationships. With appropriate marker density, the G matrix more accurately reflects the relationships between individuals, providing a more realistic picture compared to genomic relationship analysis based solely on pedigree information. The population G matrix was constructed using Gmatrix software, and heatmaps were generated using R language to illustrate the genetic relationships between samples. The calculation formula is as follows:

[0100]

[0101] P i It is the frequency of the i-th allele.

[0102] S2-4, Family Lineage Construction:

[0103] A high-quality SNP dataset was selected, and the phylogenetic tree was constructed using PLINK software. First, a Perl script was used to convert the dataset into the .ped file format used by PLINK, and then into a binary .bed file. Then, an In-State Homologous Distance (IBS) matrix was constructed based on the binary .bed file to estimate the allele sharing rate between two samples. For the constructed matrix, a Neighbor Joining Tree (NJ tree) was built using MEGA software. The IBS calculation formula is as follows:

[0104]

[0105] IBS1 is a locus that shares two alleles, IBS2 is a locus that shares one allele, and N is the total number of SNP loci used to calculate IBS among individuals.

[0106] S3. Results and Analysis, including the following steps:

[0107] S3-1. Analysis of Genomic DNA Detection Results:

[0108] Figure 4 These are the gel electrophoresis results of a portion of the DNA sample, from... Figure 4 The DNA bands are clearly visible, with no obvious banding. DNA concentration and purity tests show that all sample concentrations are >50 ng / L, and OD... 260 / OD 280 The pH value was between 1.8 and 2.1, with no RNA or protein contamination and no sample degradation. This indicates that the DNA quality was good and suitable for subsequent experiments.

[0109] S3-2, Genotype quality control results:

[0110] Genotyping of 101 samples was performed using the Illumina CAUPorcine 50K SNP chip, yielding a total of 43,832 SNP loci. The genotyping detection rate ranged from 0.88 to 0.98, with an average of 0.98. After removing one sample with a CallRate less than 0.9, a total of 34,351 SNP loci were obtained after quality control, and 100 samples were used for further analysis. Detailed quality control results are shown in Table 3. The distribution of SNP loci on each chromosome before and after quality control is shown in Table 3. Figure 5 As shown, from Figure 5 The results show that the proportion of invalid SNP loci is roughly equal across all chromosomes. Due to differences in chromosome length, the number of SNPs varies considerably among different chromosomes, with chromosome 1 having the most SNPs at 4045 and chromosome 18 having the fewest at 923.

[0111] Table 3 SNP Quality Control Statistics

[0112]

[0113] S3-3, Genetic Diversity Analysis:

[0114] Genetic diversity indices for each individual were calculated using the obtained SNP genotyping data. The number of polymorphic markers (N) was also considered. SNP The polymorphic marker ratio (P) was 34351. N The mean value was 0.84. The minimum allele frequency (MAF) for each marker ranged from 0.05 to 0.50, with a mean of 0.29 ± 0.13. The observed heterozygosity (Ho) ranged from 0.05 to 0.73, with a mean of 0.39 ± 0.12; the expected heterozygosity (He) ranged from 0.09 to 0.5, with a mean of 0.38 ± 0.11. Figure 6 This is a map showing the minimum allele frequency (MAF) distribution of SNPs, from... Figure 6 As can be seen, the MAF distribution of SNPs is relatively uniform.

[0115] S3-4, ROH analysis:

[0116] The number, length, and distribution of ROHs in an animal's genome contain rich genetic background information, including population history and inbreeding levels. Detecting ROHs in an animal's genome can help infer population history, assess inbreeding, optimize mating strategies to reduce inbreeding, identify harmful mutations, assess genetic diversity, and promote the preservation of livestock and poultry genetic resources.

[0117] In this embodiment, a total of 1629 ROH fragments were detected, mainly concentrated below 20 Mb. ROH fragments below 10 Mb and between 10 Mb and 20 Mb accounted for 35.05% and 39.53%, respectively. Figure 7 a). The shortest ROH is 5.72 Mb in length, located on chromosome 14, and contains 101 SNPs. The longest ROH is 165.33 Mb in length, located on chromosome 13, and contains 2447 SNPs. The number of ROHs in pigs is positively correlated with chromosome length; chromosome 1 has the most ROHs (286), while chromosome 12 has the fewest (26). Figure 7 (b) The number of ROHs varied considerably among the individuals tested, ranging from 3 to 33, with an average of 18.3. The largest number of individuals contained 15 ROHs. Figure 7 c). The total ROH length of individual individuals ranged from 27.06 Mb to 1123.81 Mb, with an average length of 303.04 Mb. The largest number of individuals (89%) had a total ROH length of less than 400 Mb. Figure 7 d). By statistically analyzing the ROH of each individual in the population, the inbreeding coefficient of each individual was calculated. The average inbreeding coefficient of all individuals is currently 0.1236, and the average inbreeding coefficient of the 16 boars is 0.1198.

[0118] S3-5, Results of Individual Kinship Analysis:

[0119] Principal component analysis results:

[0120] Principal component analysis was performed on 100 samples using quality-controlled SNP data to preliminarily understand the genetic structure of the Jinfen White Pig population. The results are shown below. Figure 8 . Figure 8 Individuals with serial numbers are boars, and those without serial numbers are sows. The results show that all samples are divided into two groups. In the upper circle, JB268M, JB656M, and JB657M clearly cluster together, consistent with the three samples in the original pedigree belonging to the same family. In the lower circle, JB4703M, JB1200M, and JB22000M, all from the same breeding farm, are more widely distributed. Overall, the sample distribution is relatively random, and the kinship is distant.

[0121] Analysis of G matrix and IBS distance matrix:

[0122] When analyzing kinship within a population, one can use kinship coefficients from the G matrix and genetic distance analysis based on the IBS (Individual Spectrum of the Baseline). The G matrix results (...) Figure 9 Each square represents the kinship value between a sample and other samples. The closer the color of the square is to red, the larger the value, that is, the closer the kinship between the two individuals.

[0123] A total of 4950 relationship pairs were formed from 100 samples, with an average cohort of 0.00128. Among these, 57.78% of the individuals were distantly related (cohort < 0), 25.70% were closely related (cohort 0–0.1), and 16.52% were very closely related (cohort > 0.1), with an average cohort of 0.20. Mating these individuals is highly likely to cause inbreeding depression. Combined with the PCA results, it can be seen that most individuals are distantly related, with only a small percentage being closely related. Extra caution is needed when selecting partners for these closely related individuals.

[0124] Genetic distance (IBS) among 100 samples is represented using a heatmap. Figure 10 The IBS values ​​among all individuals ranged from 0.1396 to 0.3892, with a mean genetic distance of 0.3011 ± 0.1248. The genetic distance among the 16 boars ranged from 0.1874 to 0.3703, with a mean genetic distance of 0.3117 ± 0.0274.

[0125] Example 4

[0126] This embodiment describes the application of a whole-genome SNP chip of Jinfen White pigs in cluster analysis and kinship analysis of pigs, specifically when Jinfen White pigs are the parent breeds. It is based on the whole-genome SNP chip of Jinfen White pigs from Embodiment 2. This application has the advantage of conducting association analyses of important economic traits in pigs and obtaining key genes and loci related to pig production performance.

[0127] This embodiment uses GWAS analysis combined with SNP chip detection to screen for gene loci related to the reproduction of Jinfen White Pigs, including the following steps:

[0128] S1, Reference Group Setup:

[0129] Excellent breeding pigs were selected from the core farm of Pingding Huayi lean-type pig breeding farm to form a reference group. The pedigrees and reproductive records of the reference group were compiled, and finally the reproductive records and pedigree information of 400 sows and 1380 litters were selected.

[0130] S2. Genotyping and Quality Control:

[0131] SNP microarray analysis was performed on ear tissue DNA, and genotyping was determined using GenomeStudio software. Quality control of the SNP genotyping data was performed using Plinkv1.90 software. SNP loci with a detection rate greater than 0.9 were retained; SNP loci with an individual genotype deletion rate less than 0.1 were retained; SNP loci with a minimum allele frequency (MAF) greater than 0.05 were retained; and SNP loci significantly deviating from Hardy-Weinberg equilibrium (HWE) (P-value < 1*E-6) were removed. Missing genotypes were filled using Beagle software.

[0132] S3 and GWAS correlation analysis

[0133] Genome-wide association analysis (FDI) was performed on SNP loci with reproductive traits such as total litter size, live litter size, birth weight, and litter weight using the LMM model in EMMAX software. EMMAX, based on the LMM model, controlled for genetic background by setting the genotype relationship matrix as a random effect and by setting fixed-effect covariates to control for other confounding factors such as population structure, thereby reducing the free-difference ratio (FDR) of the results and obtaining reliable data.

[0134] The LMM model is as follows:

[0135] Y=μ+Kα+Xβ+Qγ+ε

[0136] Where Y is the sow phenotypic vector; μ is the mean vector; K is the genotype relationship matrix (i.e., the kinship matrix), α is the coefficient vector of the genotype relationship matrix; X is the genotype matrix; β is the genotype effect vector; Q is other covariates such as population structure, γ is its effect; and ε is the random residual vector.

[0137] Genotype relationship matrices were calculated using the algorithm provided by EMMAX software. After association analysis, 100 permutation tests were performed to determine the p-value for genome-wide significance, followed by the creation of Manhattan plots using a self-compiled function in R. After obtaining the association analysis results, independently significant QTLs were determined according to the following criteria: the distance between QTLs was greater than 2 Mb; the distance between QTLs was less than 2 Mb, but the coefficient of determination between the most significant SNPs (PeakSNPs) of the association between the two QTLs was less than 0.3.

[0138] S4. Genetic analysis was performed on 400 Jinfen White pig samples using a whole-genome SNP chip. The results showed that the data sample alignment rate ranged from 98.02% to 98.67%, with an average alignment rate of 98.42%; the sequencing depth ranged from 18.02 to 35.9X, with an average sequencing depth of 22.48X. A total of 1115.3G of raw data was obtained, and after quality control, 1111.8G of quality control data was obtained, with a base error rate of approximately 0.03%. All samples had Q20 ≥ 95.68%, Q30 ≥ 89.39%, an average alignment rate of 99.68%, and GC content ranging from 42.85% to 43.81%.

[0139] S4 and GWAS Result Analysis:

[0140] GWAS analysis was performed on reproductive phenotypic data using EMMAX software to obtain SNPs that were significantly associated at the genome-wide level. Manhattan plots and quantile plots (QQ plots) for different reproductive traits were then plotted. The Manhattan plot can more intuitively show the distribution of SNPs on different chromosomes and detect SNPs associated with reproductive traits.

[0141] S5. Candidate gene functional enrichment analysis:

[0142] The SNPs obtained in step S4 were used as candidate genes, and GO enrichment analysis was performed on the candidate genes using KOBAS to obtain significant functional classifications (GO terms).

[0143] S6. Screening of candidate genes for reproductive traits:

[0144] Based on gene function descriptions and previous research findings, key genes and loci related to the production performance of Jinfen White Pigs were further screened and identified.

[0145] Example 5

[0146] This embodiment describes the application of a whole-genome SNP chip of Jinfen White pig in pig linkage map construction and gene localization, based on the whole-genome SNP chip of Jinfen White pig in Embodiment 2.

[0147] The difference between this embodiment and Embodiment 4 lies in the location of gene loci, i.e., the analysis of GWAS results, including the following:

[0148] Using EMMAX software, GWAS analysis was performed on reproductive phenotypic data such as total litter size, live litter size, birth weight, and litter weight to obtain SNPs that were significantly associated at the whole genome level. Manhattan plots and quantile quantile plots (QQ plots) of different reproductive traits were then plotted. The Manhattan plot can more intuitively show the distribution of SNPs on different chromosomes.

[0149] Loci and candidate gene locations affecting total pig litter size:

[0150] After identifying SNPs that are significantly associated with the total number of piglets born, the chromosomes containing the SNPs were obtained, and candidate genes that significantly affect the total number of piglets born were screened through genomic locus annotation.

[0151] Loci and candidate gene locations affecting piglet live birth rate:

[0152] After identifying SNPs that are significantly associated with the number of live piglets produced in pigs, the chromosomes containing the SNPs were obtained, and candidate genes that significantly affect the number of live piglets produced in pigs were screened through genomic locus annotation.

[0153] Loci and candidate genes affecting litter weight at birth in piglets:

[0154] After identifying SNPs that are significantly associated with the birth litter weight of piglets, the chromosomes containing the SNPs were obtained, and candidate genes that significantly affect the birth litter weight of piglets were screened through genomic locus annotation.

[0155] Example 6

[0156] This embodiment describes the application of a whole-genome SNP chip of Jinfen White Pig in the background selection of pig molecular breeding materials, based on the whole-genome SNP chip of Jinfen White Pig in Example 2.

[0157] The above application includes the following steps:

[0158] S1. Core Group Setup:

[0159] Collect pedigree and production data of Jinfen White pigs, and conduct performance testing of the breeding pigs based on the results of breed and electronic pedigree identification.

[0160] S2, Performance Testing:

[0161] Body size traits were measured using a pig face recognition system, and the data were analyzed in conjunction with actual measurements. Reproductive performance and growth and development were measured according to NY / T 820. Finishing performance was measured according to NY / T 822. Carcass traits were measured according to NY / T 825. Muscle quality was measured according to NY / T 821.

[0162] S3, Jinfen White Pig Whole Genome SNP Chip Detection Kit

[0163] Ear tissue samples were collected from individuals in the core breeding group, ear tissue DNA was extracted, and the whole genome SNP chip detection kit of Jinfen white pig was used to detect individual SNP data.

[0164] S4. Breeding Value Estimation:

[0165] Based on the breeding objectives, total litter size, daily weight gain, and backfat thickness were identified as the main selected traits. The ssGBLUP method was used to analyze the three traits of the pigs.

[0166] ssBLUP analysis model

[0167] y = Xb + Za + e

[0168] In the formula, y is the phenotypic value vector; b is the fixed effect vector; and a is the random additive genetic effect vector. For additive genetic variance, H is the relation matrix formed by combining the pedigree relation matrix A and the genomic relation matrix G; X and Z are the corresponding association matrices for b and a; e is the residual effect vector. This represents the residual variance.

[0169]

[0170] In the formula, A 11 A is a submatrix of matrix A containing no individuals with the specified genotype; A 22 This is a submatrix of matrix A containing individuals with specific genotypes; A 12 Or A 21 This is a kinship matrix of individuals with and without genotypes. G ω =(1-ω)G+ωA 22 ω is a parameter for adjusting the genome relation matrix to ensure that the genome information is compatible with the pedigree information (set to 0.1 in this study), and G is the genotype relation matrix.

[0171] S5. Successive generation selection and breeding:

[0172] Phase 1: Before weaning, piglets are initially selected based on phenotypic scores, with 30% of boars and 80% of sows retained. The selected best individuals form a testing group, and one boar and two sows from the same litter are randomly selected to collect ear tissue samples for genomic testing.

[0173] Phase Two: Estimating the GEBV breeding value of replacement gilts using ssGBLUP. Boars, after genomic selection, are introduced into the replacement gilt herd and undergo training and semen collection. Selection is based on a comprehensive evaluation of boar libido, semen quality, and the reproductive performance of the sows to be bred. Sows, after genomic selection, are also introduced into the replacement gilt herd and selected based on factors such as age at first mating, first-parity farrowing information, and litter size. The boar retention rate is approximately 15%, and the sow retention rate is approximately 25%.

[0174] S6. Breeding Progress:

[0175] Through continuous selective breeding, the reproductive performance of Jinfen White Pigs has been significantly improved. As shown in Table 4, primiparous sows had a total litter size of 11.13 piglets, a live birth of 10.8 piglets, an average birth weight of 1.41 kg, a litter weight of 14.61 kg, a weaning count of 10.68 piglets, a weaning litter weight of 73.69 kg, and an average weaning weight of 6.90 kg. Multiparous sows had a total litter size of 13.32 piglets, a live birth of 12.93 piglets, an average birth weight of 1.42 kg, a litter weight of 18.36 kg, a weaning count of 11.75 piglets, a weaning litter weight of 82.25 kg, and an average weaning weight of 7.256 kg.

[0176] Table 4 Results of breeding selection for reproductive performance of Jinfen White Pigs

[0177]

[0178] Fattening performance:

[0179] Forty piglets (20 males and 20 females, castrated boars and undcastrated females) with similar weights and normal growth and development were selected from the offspring for a fattening experiment. The experiment started at a weight of 30 kg and ended at a weight of 100 kg. The main measurements were the age at which the target weight was reached, the average daily weight gain, and the feed conversion ratio. The results are shown in Table 5. As can be seen from the table, the average daily weight gain of Jinfen White pigs during the fattening period was 863.00 g / d, the feed consumption per kg of weight gain was 2.77 kg, and the age at which the pigs reached a weight of 100 kg was 164.35 days. The results indicate that purebred Jinfen White pigs have the characteristics of rapid growth and high feed conversion ratio.

[0180] Table 5 Results of the selection and breeding of fattening performance of Jinfen White Pigs

[0181]

[0182] Slaughter performance and meat quality:

[0183] After the fattening experiment, castrated individuals were slaughtered, and carcass and meat quality were measured. The results are shown in Table 6. As can be seen from the table, when Jinfen White Pigs were slaughtered at a weight of 100kg, the individual dressing percentage was 75.67%, the lean meat percentage of the carcass was 63.70%, the meat color score was 3.4, the shear force value of the longissimus dorsi muscle was 3.57 lbs, and the intramuscular fat content was 2.79%.

[0184] Table 6. Results of breeding programs for slaughter performance and meat quality of Jinfen White Pigs

[0185]

[0186]

Claims

1. A whole-genome SNP chip for Jinfen White pigs, characterized in that, The SNP chip consists of 55,957 Jinfen White Pig-specific SNP sites, which are SEQ ID NO.1-SEQ ID NO.55957.

2. The whole genome SNP chip of Jinfen White Pig as described in claim 1, characterized in that, The method for fabricating the whole genome SNP chip of Jinfen White Pig includes the following steps: S1. Whole genome resequencing detection: Whole genome resequencing was performed on Jinfen White Pig and its four parent breeds to obtain whole genome resequencing data of Jinfen White Pig and its four parent breeds; S2, Quality Control of Whole Genome Resequencing Data: Quality control was performed on the whole genome resequencing data of Jinfen White pigs and their four parents to obtain quality control data of Jinfen White pigs and their four parents. Specifically, low-quality reads, undetected reads, and adapter sequence data in the whole genome resequencing data of Jinfen White pigs and their four parents were filtered out using the FastP software. The filtering conditions were the default parameters of the FastP software. S3, Quality control data compared with the pig Susscrofa 11.1 reference genome: The quality control data of Jinfen White Pig and its four parents were compared with the pig Susscrofa11.1 reference genome using minimap2 software. Then, duplicate gene loci were removed from the quality control data of Jinfen White Pig and its four parents using GATK4 software to obtain the gene sequences of Jinfen White Pig and its four parents. S4. Obtain the set of SNP mutation sites of Jinfen White Pig and its four parents. After re-aligning the quality control data of Jinfen White pigs and their four parents using the BaseReealibration module of GATK4 software, GVCF data of Jinfen White pigs and their four parents were generated using the HaplotypeCaller module of GATK4 software. Locus re-alignment was then performed on the GVCF data of Jinfen White pigs and their four parents, and then locus quality filtering was performed on the GVCF data of Jinfen White pigs and their four parents using the VariantFilteration module of GATK4 software. This yielded the SNP variant locus set of Jinfen White pigs and their four parents. The filtering conditions were: SQR <= 10, QD >= 2, MQ >= 40.0, QUAL >= 30.0, MQRandSum >= 12.5, ReadPosRankSum >= -8.0, InbreedingCoeff >= 0.8, and FS <= 60.

0. S5. Optimize loci through GWAS association analysis of Jinfen White pig production performance, and screen for Jinfen White pig-specific SNPs that appear only in the Jinfen White pig population: The SNP variant sites of Jinfen White pigs and their four parents were screened to obtain Jinfen White pig-specific SNP sites. The screening criteria were: allele frequency > 0.

1. The Jinfen White pig-specific SNP sites were compared with sites in the NCBI SNP database. Sites were optimized through GWAS association analysis of Jinfen White pig production performance. Finally, 55,957 Jinfen White pig-specific SNP sites that appeared only in the Jinfen White pig population were obtained. S6. Integrate the 55,957 Jinfen White Pig-specific SNP loci obtained in step S5 to obtain the whole genome SNP chip of Jinfen White Pig.

3. The whole genome SNP chip of Jinfen White Pig as described in claim 2, characterized in that, The application of the SNP chip in parentage testing of Jinfen White Pigs and when Jinfen White Pigs are the parent breeds.

4. The whole genome SNP chip of Jinfen White Pig as described in claim 2, characterized in that, The application of the SNP chip in pig clustering analysis and pig kinship analysis when Jinfen White pigs are used as parents.

5. A whole-genome SNP chip for Jinfen White pigs as described in any one of claims 4-5, characterized in that, The parent breeds of the Jinfen White Pig are: Ma Shen Pig, Taihu Pig, Large White Pig, and Landrace Pig.

6. The whole genome SNP chip of Jinfen White Pig as described in claim 2, characterized in that, The application of the SNP chip in pig linkage map construction and gene localization.

7. The whole genome SNP chip of Jinfen White Pig as described in claim 2, characterized in that, Application of the SNP chip in background selection of pig molecular breeding materials.