Chinese local breed pig whole genome 30k snp breeding chip and application thereof
Patent Information
- Application Number
- CN202611134677.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-29
- Publication Date
- 2026-09-25
AI Technical Summary
这些芯片在中国地方猪品种中存在两个核心问题:其一,芯片上相当比例的SNP位点在中国地方猪群体中呈低多态性或单态,有效标记数量大幅缩水,无法满足基因组评估所需的标记密度;其二,现有芯片的位点以非功能性的中性标记为主,其育种信息依赖于与因果突变的连锁不平衡关系,而不同品种乃至同一品种不同群体间的LD结构存在显著差异,导致标记效应难以跨群体稳定迁移,芯片的通用性和评估准确性受到根本性制约
[0017]1.本发明提供了一种中国地方品种猪30K SNP液相育种芯片,SNP位点设计涵盖了基因组注释类型、次要等位基因频率、保守元件、品种特异性QTL及通用QTL多种信息。所述33,492个SNP位点覆盖了外显子、UTR、非编码RNA、剪接位点等8种基因组注释类型,其中基因编码及调控区内(exonic、UTR、ncRNA、splicing)的位点共计12,777个,占总位点数的26.5%,改变了传统商业芯片以中性标记为主的位点筛选策略;同时,15.4%的位点与猪经济性状QTL区域(±10kb)关联,其中8个地方品种猪品种特异性QTL位点2,531个(5.2%),非8品种猪QTL位点4,897个(10.2%),为地方品种猪高繁殖力等特色性状的基因组选择和分子标记辅助育种提供了直接的功能标记支撑。
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention belongs to the field of animal breeding gene chip technology, specifically relating to a 30K SNP breeding chip for the whole genome of a Chinese local breed of pig and its application. Background Technology
[0002] Single nucleotide polymorphisms (SNPs) are the most widespread type of genetic variation in the genome, characterized by their large number, genetic stability, and ease of high-throughput automated detection. They are currently the most commonly used molecular marker type in livestock and poultry genomic selection breeding and population genetics research. In recent years, SNP microarray technology based on liquid-phase capture sequencing has gradually become an important genotyping platform for livestock and poultry molecular breeding due to its advantages such as flexibility in probe site design, short customization cycle, and the ability to adjust site combinations as needed.
[0003] In pig genome breeding practices, the quality of SNP microarrays directly affects the accuracy of genome breeding value predictions and breeding efficiency. However, currently mainstream commercial pig SNP microarrays internationally, including Illumina PorcineSNP60, GeneSeek GGP Porcine series, and Axiom Porcine microarrays, primarily rely on population genetic data from Western lean-type pig breeds such as Duroc, Landrace, and Large White for their SNP site selection strategies. These microarrays present two core problems with local Chinese pig breeds: First, a significant proportion of SNP sites on the microarrays exhibit low polymorphism or monomorphism in local Chinese pig populations, resulting in a substantial reduction in the number of effective markers and failing to meet the marker density required for genome assessment. Second, existing microarrays primarily use non-functional neutral markers, and their breeding information depends on linkage disequilibrium relationships with causal mutations. Significant differences in the LD structure between different breeds and even between different populations of the same breed make it difficult for marker effects to migrate stably across populations, fundamentally limiting the universality and assessment accuracy of the microarrays.
[0004] The Meishan pig (Jiangsu), Erhualian pig (Jiangsu), Rongchang pig (Chongqing / Sichuan), Jinhua pig (Zhejiang), Laiwu pig (Shandong), Bamei pig (Shaanxi, Gansu, Ningxia), Luchuan pig (Guangxi), and Wuzhishan pig (Hainan) – these eight pig breeds cover five major regions of my country: East China, Southwest China, North China, Northwest China, and South China. They represent the core genetic resources of local Chinese pig breeds in terms of high fertility, excellent meat quality, and stress resistance. Among them, the Meishan and Erhualian pigs are world-renowned for their high fertility, the Rongchang and Jinhua pigs are famous in the market for their excellent meat quality, the Laiwu and Bamei pigs are known for their stress resistance and adaptability, and the Luchuan and Wuzhishan pigs have unique value in terms of heat tolerance, tolerance to roughage, and experimental animal models. However, for a long time, there has been no universal breeding chip for these Chinese local pig breeds, either domestically or internationally. Developing a breeding chip that deeply integrates the genetic characteristics and functional genomic information of multiple breed populations is of great significance for the systematic protection and scientific utilization of the excellent genetic resources of local Chinese pigs and for promoting germplasm innovation. Summary of the Invention
[0005] In view of the above problems, the purpose of this invention is to provide a 30K SNP breeding chip for the whole genome of Chinese local breed pigs and its application. Specifically, based on the whole genome resequencing data of 141 pigs from 8 Chinese local breeds and the pig reference genome Sscrofa11.1, chip loci are designed by combining SNP genome annotation type, minor allele frequency, conserved elements, QTLs of 8 Chinese local breed pigs and QTLs of non-8 Chinese local breed pigs, thereby obtaining a 30K SNP breeding chip for the whole genome of Chinese local breed pigs.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] This invention provides a molecular probe array for the specific detection of 33,492 SNP sites located in the porcine reference genome Sscrofa11.1 as shown in Table 1.
[0008] This invention provides a 30K SNP breeding chip for the whole genome of a Chinese local breed of pig, wherein the SNP breeding chip includes a probe combination targeting the above-mentioned SNP molecular markers.
[0009] The present invention provides a kit comprising the above-described molecular probe combination or a 30K SNP breeding chip.
[0010] This invention provides applications of the above-described molecular probe combinations, SNP breeding chips, or kits in any of the following aspects:
[0011] (1) Application in genotyping detection of local Chinese pig breeds;
[0012] (2) Application in molecular marker-assisted breeding or whole-genome selection breeding of local Chinese pig breeds;
[0013] (3) Application in genome-wide association analysis of local Chinese pig breeds;
[0014] (4) Application in the identification or genetic diversity analysis of local Chinese pig breeds;
[0015] The local pig breeds mentioned include Meishan pig, Erhualian pig, Rongchang pig, Jinhua pig, Laiwu pig, Bamei pig, Luchuan pig, and Wuzhishan pig.
[0016] Through the above technical solution, the present invention has the following beneficial effects:
[0017] 1. This invention provides a 30K SNP liquid-phase breeding chip for Chinese local pig breeds. The SNP loci design encompasses various information including genome annotation types, minor allele frequencies, conserved elements, breed-specific QTLs, and general QTLs. The 33,492 SNP loci cover eight genome annotation types, including exons, UTRs, non-coding RNAs, and splicing sites. Among them, 12,777 loci are located within gene coding and regulatory regions (exonic, UTR, ncRNA, splicing), accounting for 26.5% of the total loci. This changes the traditional commercial chip's locus selection strategy, which mainly relies on neutral markers. Simultaneously, 15.4% of the loci are associated with QTL regions (±10kb) of economically important pig traits, including 2,531 breed-specific QTL loci (5.2%) for the eight local pig breeds and 4,897 QTL loci (10.2%) for non-eight breed pigs. This provides direct functional marker support for genome selection and marker-assisted breeding of distinctive traits such as high fertility in local pig breeds.
[0018] 2. This invention has significant advantages in terms of breed specificity. The chip loci are obtained based on the screening of whole-genome resequencing data of local pig breeds. All loci have been verified by MAF≥0.05 quality control in various pig breed populations, which effectively avoids the problem of low polymorphism rate (usually only 50%~70%) of general commercial chips in Chinese local pig breeds, and can meet the actual needs of precision breeding of local pig breeds.
[0019] 3. The microarray loci of this invention are evenly distributed across the entire genome. Using a 40 Kb window as a basis for genome-wide locus screening, representative loci are evenly distributed across all 18 autosomes, with an average spacing of approximately 52 Kb between loci. This avoids marker loss in large genomic regions and ensures full utilization of linkage disequilibrium information in genome-wide association studies and genome selection.
[0020] 4. This invention can be widely applied to scenarios such as genetic diversity assessment and germplasm resource identification of local pig breeds, genome-wide association analysis, genome selection breeding, molecular marker-assisted breeding, and kinship identification. It can scientifically guide the improvement of local pig breeds and the protection of genetic resources in China, significantly shorten the breeding generation interval, and improve the accuracy and efficiency of breeding. Attached Figure Description
[0021] Figure 1 The distribution map of SNP sites on different chromosomes of local breed pigs detected by the 30K SNP chip of the whole genome of the pig provided by the present invention.
[0022] Figure 2 The distribution map of SNP sites detected by the 30K SNP chip of the whole genome of local pig breeds provided by this invention in different genome annotation regions.
[0023] Figure 3 The results of PCA analysis of eight local pig breeds using the 30K SNP chip of the whole genome of local pig breeds provided in this invention. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0025] Example 1: Design of 30K SNP breeding chip sites and probes for the whole genome of Chinese local pig breeds
[0026] This embodiment provides a method for site screening and probe design using a 30K SNP liquid phase breeding chip for local Chinese pig breeds. The specific steps are as follows:
[0027] 1. Data collection and quality control
[0028] Eight breed-specific VCF files for Chinese local pig breeds (including Meishan, Erhualian, Rongchang, Jinhua, Laiwu, Bamei, Luchuan, and Wuzhishan) were downloaded from the pig genome filling database PGIDB (http: / / animalbreedinglab.hzau.edu.cn / AGIDB / Pig / Pig_home.php). These files were then phased and genotype-optimized using Beagle v5.4, ultimately retaining 20,954,035 autosomal SNP loci from 141 high-quality individuals. The reference genome was pig Sscrofa11.1 (Ensembl release 104). The VCF files were quality-controlled using bcftools (--mac 1) and MAF filtered using PLINK v1.9 (--maf 0.05). The retained SNPs after quality control were used for subsequent microarray site selection.
[0029] 2. Genome window segmentation and haplotype detection
[0030] Based on the chromosome length index file of the pig reference genome Sscrofa11.1, the genome was divided into 40 Kb windows using the bedtoolsmakewindows tool, resulting in 62,827 windows covering 18 autosomes and the X chromosome. PLINK v1.9 was used to perform haplotype detection on SNPs within each 40 Kb window, with parameters set to --maf0.05 --show-tags all --tag-kb 100 --tag-r2 0.8. The tags.list file for each window was output, obtaining the linkage disequilibrium (LD) block structure and tagSNP information for each window.
[0031] 3. SNP weighting assessment
[0032] Each SNP locus was comprehensively scored based on the following five dimensions:
[0033] (1) Genomic region annotation weights: Genome-wide SNPs were annotated using ANNOVAR software. An annotation database was constructed based on the pig reference genome Sscrofa11.1 and the Ensembl release 104 GTF annotation file (Sus_scrofa.Sscrofa11.1.104.gtf). The GTF file was converted to GenePred format using the gtfToGenePred tool, the reference sequence was extracted using retrieve_seq_from_fasta.pl, and gene-based annotation was completed using annotate_variation.pl. Differential weights are assigned based on the genomic region in which the SNP is located: splicing site weight is 5, premature termination / stop loss weight is 4, UTR3 / UTR5 / ncRNA weight is 4, synonymous mutation weight is 3, exon weight is 3, nonsynonymous mutation weight is 2, and intergenic / intronic / upstream / downstream weight is 1.
[0034] (2) Minor allele frequency (MAF): Extract AC (Allele Count) and AN (Allele Number) information from the INFO field of VCF, calculate the allele frequency of each SNP in 8 Chinese local variety populations by AC / AN, and directly include the minor allele frequency in the weight.
[0035] (3) Conserved element weights: Download the pig Sscrofa11.1 conserved element annotation file ConsElements.0-based.bed from the UCSC Genome Explorer (http: / / hgdownload.cse.ucsc.edu / ), and use bedtoolsintersect to determine whether the SNP is located in a conserved element region. The weight of SNPs located in a conserved element region is increased by 4.
[0036] (4) QTL weights for local breed pigs: 52,840 QTL records for all pigs were downloaded from the Animal QTLdb database (https: / / www.animalgenome.org / QTLdb / , release 59, built in SS11). 6,235 QTL records containing "Meishan", "Rongchang", "Jinhua", "Laiwu", "Bamei", "Luchuan", "Wuzhishan", or "Erhualian" were filtered. The QTL coordinates were extended 10 Kb upstream and downstream, and bedtools intersect was used to determine if the SNP was located within the QTL region of the 8 local breed pigs. For each QTL record matching the 8 local breed pigs, the weight was multiplied by 8.
[0037] (5) Weighting of QTLs for other pig breeds: For the 34,096 QTL records of other pig breeds that do not include the 8 local breeds in the QTLdb, SNP overlap was determined after expanding by ±10 Kb. For each non-local breed pig QTL record matched, the weight was multiplied by 6.
[0038] The final overall weight of a single SNP = Annotation type score + MAF + Conservative element score + Number of QTL hits for 8 local pig breeds × 8 + Number of QTL hits for non-8 local pig breeds × 6.
[0039] 4. Selection of the optimal SNP within the window
[0040] For each haplotype within a window, the sum of the comprehensive scores of all SNPs contained within it is calculated as the score for that haplotype. The haplotype with the highest score in each window is selected, and the SNP with the highest comprehensive score from that haplotype is selected as the first candidate locus for that window, resulting in 56,487 first candidate SNP loci. At the same time, the 2nd to Nth SNPs in the window, sorted by weight, are retained as backup candidate loci.
[0041] 5. Probe design and multi-round iterative screening
[0042] (1) Genome probe splitting: Using the blockParse module of OligoMiner software, the pig reference genome Sscrofa11.1 was split into chromosomes, and probe sequences were split with a window length of 120 bp and a Tm range of 40-200 (parameters: -l 120 -L 120 -t 40 -T 200).
[0043] (2) Genome uniqueness filtering: Use bowtie2 v2.5.5 to post candidate probe sequences back to the reference genome (parameter: --very-sensitive-local -k 100), and use outputClean.py (-u parameter) to retain probes that have only one unique match in the genome.
[0044] (3) kmer-specific filtering: A genome-wide 25mer frequency database was constructed using Jellyfish v2.2.10 (parameters: -s 3G -m 25 --out-counter-len 1 -L 2). kmerFilter.py was used to retain probes whose 25 bp subsequences appeared less than 3 times in the whole genome.
[0045] (4) SNP and probe matching: The filtered probe BED files and candidate SNP location BED files were overlapped using the bedtools tool to determine whether each candidate SNP was covered by a specific probe. The first round yielded 26,602 valid sites (coverage rate 47.1%), and the remaining 29,885 windows had no matching probes.
[0046] (5) Multiple rounds of iterative remediation: For probe-free windows, backup candidate sites (the 2nd to Nth largest weight SNPs) are selected sequentially, and the probe design steps described above are repeated. In each round, the SNPs with newly obtained probes are merged into the cumulative probe set. After multiple rounds of iteration, a total of 33,492 valid sites are obtained. Site information is detailed in Table 1. For example, 1:71589_A / G represents the SNP site at position 71589 bp on pig chromosome 1 with a reference base of A and a mutated base of G.
[0047] Table 1 SNP locus information
[0048]
[0049]
[0050]
[0051]
[0052]
[0053]
[0054]
[0055]
[0056]
[0057]
[0058]
[0059]
[0060]
[0061]
[0062]
[0063]
[0064]
[0065]
[0066]
[0067]
[0068]
[0069]
[0070]
[0071]
[0072]
[0073]
[0074]
[0075]
[0076]
[0077]
[0078]
[0079]
[0080]
[0081]
[0082]
[0083]
[0084]
[0085]
[0086]
[0087]
[0088]
[0089]
[0090]
[0091]
[0092]
[0093]
[0094]
[0095]
[0096]
[0097]
[0098]
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128] Figure 1 and Figure 2The distribution of 30K SNPs in liquid-phase breeding microarrays for eight local pig breeds was shown across different chromosomes and genomic annotation regions. Of the final 33,492 SNPs, the genomic region distribution was as follows: intergenic regions 19,955 (41.4%), intronic regions 16,416 (34.0%), ncRNA intronic regions 3,326 (6.9%), UTR regions 4,160 (8.6%), exon regions 2,062 (4.3%), ncRNA exon regions 1,385 (2.9%), upstream / downstream regions 884 (1.8%), and splicing sites 34 (0.1%). The QTL region distribution was as follows: 2,531 (5.2%) QTL regions for the eight local pig breeds, 4,897 (10.2%) QTL regions for other pig breeds, and 170 (0.4%) QTL regions located in conserved element regions.
[0129] Example 2: Method of using the 30K SNP breeding chip for the whole genome of Chinese local pig breeds
[0130] This embodiment provides a method for using the 30K SNP chip of the whole genome of Chinese local breed pigs in Embodiment 1 above, specifically including the following:
[0131] 1. Extraction and quality control of genomic DNA
[0132] DNA was extracted from local pig tissue samples (blood, ear tissue, or muscle tissue) using a magnetic bead method. The concentration of DNA samples was determined using a Qubit real-time fluorescence analyzer, and the integrity of the DNA samples was assessed using 1% agarose gel electrophoresis. DNA samples meeting the following quality criteria were selected for library preparation: DNA concentration ≥20 ng / μL, OD... 260 / 280 The ratio is between 1.8 and 2.0, OD 260 / 230 ≥1.8, the main band in the electrophoresis is clear and there is no obvious degradation.
[0133] 2. Construction and quality control of cGPS liquid phase capture library
[0134] (1) Take 50-150 ng of qualified genomic DNA, use fragmentation enzyme to cut the DNA into fragments of 100-500 bp, repair the enzyme ends and add an A base at the 3' end, and use agarose gel electrophoresis to detect the fragment size distribution.
[0135] (2) The sequencing adapters were ligated to the end-repaired DNA fragments using T4 DNA ligase, and the ligation products were purified using magnetic beads. The concentration of the purified ligation products was detected using a Qubit real-time fluorescence instrument, and the size of the ligation product fragments was detected by agarose gel electrophoresis.
[0136] (3) The purified ligation products were subjected to PCR amplification, and the amplified products were screened for fragments using magnetic beads. The concentration of the products after fragment screening was detected using a Qubit fluorescence quantitative PCR instrument, and the fragment size was detected by agarose gel electrophoresis to confirm that the main peak of the library was distributed in the range of 250-450 bp.
[0137] (4) Take 200 ng of the constructed sequencing library, add the biotin-labeled DNA probe combination corresponding to the chip described in Example 1 and hybridization buffer, and incubate at 50°C for 16-24 hours to complete the liquid-phase hybridization capture reaction. Use streptavidin magnetic beads to capture the target segment and wash away non-specifically bound fragments. Perform PCR amplification and enrichment on the captured library, and purify it with magnetic beads to complete the construction of the cGPS sequencing library. Use Qubit quantification and Agilent 2100 Bioanalyzer to detect the library concentration and fragment distribution.
[0138] 3. High-throughput sequencing and data analysis
[0139] (1) The prepared cGPS capture library was sequenced at both ends of 150bp (PE150) using a high-throughput sequencing platform.
[0140] (2) Use FASTP software to perform quality control on the raw data after the machine, remove adapter sequences and low-quality reads (Q<20) to obtain high-quality Clean Reads. Use the BWA-MEM algorithm to align the Clean Reads to the pig reference genome Sscrofa11.1, and use Samtools to sort and remove duplicates to obtain the BAM alignment file for each sample.
[0141] (3) Use GATK or bcftools to detect variants in the sequencing results and extract genotyping results for 33,492 target SNP loci. Further format and quality control of the genotyping data are performed based on PLINK or a self-written script.
[0142] 4. Downstream Application Analysis
[0143] The following analyses can be performed using the obtained genotyping results:
[0144] (1) Genome-wide association analysis (GWAS): Combined with phenotypic records of the target trait (such as litter size, age at puberty, daily weight gain, backfat thickness, etc.), GWAS analysis is performed using a mixed linear model to identify SNP loci that are significantly associated with important traits.
[0145] (2) Genome selection breeding: Using the best linear unbiased prediction of genome (GBLUP) or Bayesian method, a genome prediction model is constructed based on microarray genotyping data to calculate the genome estimated breeding value (GEBV) of candidate individuals, so as to achieve early and accurate breeding.
[0146] Example 3: Application of the 30K SNP breeding chip for the whole genome of Chinese local pig breeds in breed identification.
[0147] This embodiment uses the 30K SNP breeding chip of the whole genome of Chinese local pig breeds provided in Example 1 to perform genotyping detection on 141 pig samples. The samples cover 8 Chinese local pig breeds: Meishan pig (64 pigs), Erhualian pig (25 pigs), Jinhua pig (12 pigs), Rongchang pig (12 pigs), Laiwu pig (8 pigs), Bamei pig (7 pigs), Wuzhishan pig (7 pigs), and Luchuan pig (6 pigs). Principal component analysis (PCA) was performed on 33,492 SNP loci of the 141 samples using Plink software. The first three principal components were extracted to construct a PCA scatter plot. Each point in the scatter plot represents a sample. The greater the distance between two points, the greater the difference in genetic background. Individuals with similar genetic backgrounds are clustered together. Comparison of the PCA analysis results with the actual breed grouping revealed that the 8 breeds showed obvious breed clustering in the space formed by the first principal component (PC1), the second principal component (PC2), and the third principal component (PC3), such as... Figure 3 As shown in the figure, this chip can effectively distinguish eight local Chinese pig breeds, verifying its application value in breed identification and genetic diversity assessment.
[0148] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.
Claims
1. A molecular probe assembly, characterized in that, The molecular probe combination is used to detect the SNP site combinations in Table 1 of the instruction manual.
2. A 30K SNP breeding chip for the whole genome of Chinese local pig breeds, characterized in that... The chip is loaded with the molecular probe assembly as described in claim 1.
3. A reagent kit, characterized in that, The kit includes the molecular probe combination of claim 2 or the breeding chip of claim 3.
4. The application of the molecular probe combination of claim 1, the breeding chip of claim 2, or the kit of claim 3 in the genotyping detection of local Chinese pig breeds.
5. The application of the molecular probe combination of claim 1, the breeding chip of claim 2, or the kit of claim 3 in the identification or genetic diversity analysis of local Chinese pig breeds, characterized in that... The local pig breeds mentioned include Meishan pig, Erhualian pig, Rongchang pig, Jinhua pig, Laiwu pig, Bamei pig, Luchuan pig, and Wuzhishan pig.
6. The application of the molecular probe combination of claim 1, the breeding chip of claim 2, or the kit of claim 3 in genome-wide association analysis of Chinese local pig breeds, characterized in that, The local pig breeds mentioned include Meishan pig, Erhualian pig, Rongchang pig, Jinhua pig, Laiwu pig, Bamei pig, Luchuan pig, and Wuzhishan pig.
7. The application of the molecular probe combination of claim 1, the breeding chip of claim 2, or the kit of claim 3 in the whole-genome breeding of local Chinese pig breeds, characterized in that, The local pig breeds mentioned include Meishan pig, Erhualian pig, Rongchang pig, Jinhua pig, Laiwu pig, Bamei pig, Luchuan pig, and Wuzhishan pig.