Specific SNP site combination for identifying bai xing wuzi mountain pig strain and application thereof
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SANYA RESEARCH INSTITUTE OF HAINAN ACADEMY OF AGRICULTURAL SCIENCES (HAINAN EXPERIMENTAL ANIMAL RESEARCH CENTER)
- Filing Date
- 2026-04-01
- Publication Date
- 2026-06-23
Smart Images

Figure CN121951083B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of molecular biology identification technology, specifically relating to a specific SNP site combination for identifying the white Wuzhishan pig breed and its application. Background Technology
[0002] Hainan Island's unique geographical and ecological conditions have fostered numerous local pig genetic resources, including Wenchang pigs, Tunchang pigs, Lingao pigs, Ding'an pigs, Duntou pigs, and Wuzhishan pigs. Among them, the Wuzhishan pig is a rare small local pig breed in my country, which can be divided into three strains based on coat color: black with snow-white coat, pure black, and white. The white Wuzhishan pig was once on the verge of extinction due to its long-term scarcity. In recent years, researchers have successfully revived its germplasm resources through biotechnology such as somatic cell cryopreservation and cloning, and its population size is gradually recovering.
[0003] However, in the process of resource restoration and subsequent preservation, breeding and industrialization, a key technical bottleneck is faced: the lack of a method to conduct efficient and accurate genetic identification of the restored population and its offspring to ensure the purity of the breed and to distinguish it from other Wuzhishan pig breeds and other pig breeds with similar appearances.
[0004] The core of establishing a reliable genetic identification system lies in developing a set of highly discriminative single nucleotide polymorphism (SNP) loci. This set must meet three levels of identification requirements: first, it must be able to distinguish Wuzhishan white pigs from imported commercial pigs such as Duroc, Landrace, and Large White; second, it must be able to distinguish them from other regional local pig breeds in China; and, more importantly, it must be able to specifically distinguish them from Hainan local pig breeds with similar genetic backgrounds (such as Duntou pigs and Lingao pigs).
[0005] Genome-wide association analysis (GWAS) and selection signal analysis are key technical approaches for screening such differentially expressed loci. GWAS can systematically screen gene loci significantly associated with the phenotypic characteristics of the White Wuzhishan pig; selection signal analysis can reveal genomic regions that have undergone strong selection during long-term breeding. Combining these two methods allows for the identification of highly discriminative SNP loci from both phenotypic association and evolutionary selection dimensions. This provides a core tool for constructing a "molecular identity card" for the White Wuzhishan pig and establishing a standardized system for breed purity identification and traceability, and has significant application value for germplasm resource protection, breeding, and brand development. Summary of the Invention
[0006] In order to overcome the shortcomings and disadvantages of the existing technology, the primary objective of this invention is to provide a specific SNP locus combination for identifying the white Wuzhishan pig breed. This specific SNP locus combination can achieve rapid and accurate identification of purebred white Wuzhishan pigs at the gene level.
[0007] Another object of the present invention is to provide the application of the above-mentioned specific SNP site combination for identifying the white Wuzhishan pig breed.
[0008] Another objective of this invention is to provide a method for rapidly identifying the white Wuzhishan pig breed.
[0009] The objective of this invention is achieved through the following technical solution:
[0010] A specific SNP locus combination for identifying the white Wuzhishan pig breed comprises 34 SNP loci, named SNP loci 1-34. The physical locations of these 34 SNP loci were determined based on sequence alignment of the pig reference genome Sscrofa11.1. The locations and variation information of these 34 SNP loci are shown below:
[0011] SNP site 1 is located at 72220198 bp on chromosome 3, and its mutation type is C / T.
[0012] SNP site 2 is located at 95365104 bp on chromosome 3, and its mutation type is C / T.
[0013] SNP site 3 is located at 116456303 bp on chromosome 3, and its mutation type is C / T.
[0014] SNP site 4 is located at 40551892 bp on chromosome 6, and its mutation type is A / C.
[0015] SNP site 5 is located at 40552273 bp on chromosome 6, and its mutation type is G / A.
[0016] SNP locus 6 is located at 40552360 bp on chromosome 6, and its mutation type is A / G.
[0017] SNP site 7 is located at 40553929bp on chromosome 6, and its mutation type is G / A.
[0018] SNP locus 8 is located at 40553931 bp on chromosome 6, and its mutation type is A / C.
[0019] SNP locus 9 is located at 164810557bp on chromosome 6, and its mutation type is G / A.
[0020] SNP locus 10 is located at 56,997,264 bp on chromosome 7, and its mutation type is T / C.
[0021] SNP locus 11 is located at 57005177bp on chromosome 7, and its mutation type is G / A.
[0022] SNP locus 12 is located at 57016196 bp on chromosome 7, and its mutation type is G / A.
[0023] SNP locus 13 is located at 57021566 bp on chromosome 7, and its mutation type is C / G.
[0024] SNP locus 14 is located at 57029163 bp on chromosome 7, and its mutation type is G / A.
[0025] SNP locus 15 is located at 57031495 bp on chromosome 7, and its mutation type is A / C.
[0026] SNP locus 16 is located at 57032411 bp on chromosome 7, and its mutation type is C / T.
[0027] SNP locus 17 is located at 57032554 bp on chromosome 7, and its mutation type is T / C.
[0028] SNP locus 18 is located at 57035290 bp on chromosome 7, and its mutation type is C / T.
[0029] SNP locus 19 is located at 57035378 bp on chromosome 7, and its mutation type is G / A.
[0030] SNP 20 is located at 19703658 bp on chromosome 10, and its mutation type is A / G.
[0031] SNP site 21 is located at 19,717,443 bp on chromosome 10, and its mutation type is G / A.
[0032] SNP site 22 is located at 19,718,589 bp on chromosome 10, and its mutation type is G / C.
[0033] SNP site 23 is located at 19719575 bp on chromosome 10, and its mutation type is C / T.
[0034] SNP site 24 is located at 19719621 bp on chromosome 10, and its mutation type is C / T.
[0035] SNP site 25 is located at 19719631 bp on chromosome 10, and its mutation type is G / A.
[0036] SNP locus 26 is located at 19722528 bp on chromosome 10, and its mutation type is G / A.
[0037] SNP site 27 is located at 19723319 bp on chromosome 10, and its mutation type is G / A.
[0038] SNP site 28 is located at 19723330bp on chromosome 10, and its mutation type is G / C.
[0039] SNP 29 is located at 19724019 bp on chromosome 10, and its mutation type is A / G.
[0040] SNP site 30 is located at 19,757,672 bp on chromosome 10, and its mutation type is A / G.
[0041] SNP site 31 is located at 56826128 bp on chromosome 11, and its mutation type is G / A.
[0042] SNP site 32 is located at 36198591 bp on chromosome 13, and its mutation type is C / T.
[0043] SNP site 33 is located at 110336387bp on chromosome 13, and its mutation type is A / G.
[0044] SNP site 34 is located at 207921057 bp on chromosome 13, and its mutation type is G / A.
[0045] The application of the specific SNP locus combination in identifying the white Wuzhishan pig breed or in the genetic relationship analysis of the white Wuzhishan pig.
[0046] Application of a product that detects the above-mentioned specific SNP locus combinations in the identification of white Wuzhishan pig breeds or in the analysis of genetic relationships of white Wuzhishan pigs.
[0047] The products mentioned are reagents, kits, or gene chips, etc.
[0048] The reagents include at least one of primers, probes, etc.
[0049] A molecular probe array for identifying the white Wuzhishan pig breed, the molecular probe array targeting the aforementioned specific SNP site combination, is used to detect the aforementioned specific SNP site combination.
[0050] A gene chip for identifying the white Wuzhishan pig breed, the chip being loaded with the aforementioned combination of molecular probes.
[0051] A method for rapid identification of white Wuzhishan pig breeds preferably includes the following steps:
[0052] (1) Collect tissue samples from the pigs to be tested, extract DNA from the tissue samples, and obtain DNA samples from the pigs to be tested;
[0053] (2) Perform gene sequencing on the DNA sample of the pig to be tested obtained in step (1) to obtain the genotype data of the pig to be tested;
[0054] (3) The genotype data of the pigs to be tested obtained in step (2) are merged with the genotype data of the known white Wuzhishan pig population and the control population. The genotype data of the above-mentioned specific SNP locus combination (34 SNP loci) are extracted. The merged and extracted genotype data are subjected to principal component analysis, and it is determined whether the pigs to be tested are white Wuzhishan pigs.
[0055] The DNA sample mentioned in step (2) needs to be quality inspected first, and gene sequencing can be performed after the quality inspection is qualified.
[0056] The absorbance ratio of the DNA sample nucleic acid purity detection A260 / 280 in step (2) is between 1.8 and 2.0, and the concentration is ≥50 ng / μL.
[0057] The known white Wuzhishan pig population mentioned in step (3) is a population composed of multiple (≥15) purebred white Wuzhishan pigs. It can be the white Wuzhishan pig population in the example, or a white Wuzhishan pig population selected by the individual, or a white Wuzhishan pig population in a public database.
[0058] The control group mentioned in step (3) is a group of non-white Wuzhishan pigs (≥5 pigs in a single breed group, ≥30 pigs in the whole group). It can be the control group in the example, or other non-white Wuzhishan pig groups selected by oneself, or non-white Wuzhishan pig groups in public databases.
[0059] The control group is preferably at least one of the following: Wuzhishan pig, Duntou pig, Lingao pig, Bama fragrant pig, Diannan small-eared pig, Anqing pig, Erhualian pig, Meishan pig, Jinhua pig, Min pig, Landrace pig, Large White pig, and Duroc pig.
[0060] The merging and extraction in step (3) are preferably performed using Plink software.
[0061] The principal component analysis described in step (3) is preferably visualized using the ggplot2 package in R language.
[0062] The method for judgment described in step (3) is as follows:
[0063] When the principal component analysis results show that the pig to be tested is genetically close to the known white Wuzhishan pig population, i.e., clustered together, then the pig to be tested is a white Wuzhishan pig.
[0064] The present invention has the following advantages and effects compared with the prior art:
[0065] (1) Based on resequencing data, this invention screened a specific SNP locus combination for identifying the white Wuzhishan pig breed by combining genome-wide association analysis with selection signal analysis. This screening method significantly improves the specificity and accuracy of screening specific SNP locus combinations for identifying the white Wuzhishan pig breed, and provides more reliable molecular marker support for the breed identification of the white Wuzhishan pig.
[0066] (2) The SNP locus combination provided by the present invention consists of only 34 SNP loci, which has the advantages of simple operation, low detection cost, fast analysis speed, strong specificity and high accuracy. It can be used for scientific and accurate identification and genetic relationship analysis of white Wuzhishan purebred pigs, providing an efficient technical means for the protection and development of local characteristic pig breeds.
[0067] (3) The SNP locus combination provided by this invention has good universality and identification accuracy. It is suitable for accurately identifying multiple ecological groups such as white Wuzhishan pigs and local pig breeds in Hainan, South China, Jianghai, Central China, North China and introduced Western breeds (e.g., Wuzhishan pigs, Duntou pigs, Lingao pigs, Bama fragrant pigs, Yunnan small-eared pigs, Anqing pigs, Erhualian pigs, Meishan pigs, Jinhua pigs, Min pigs, Landrace pigs, Large White pigs and Duroc pigs). It has a wide coverage, high identification accuracy and a wider range of applications and scenarios.
[0068] (4) The present invention further establishes a method for rapid identification of white Wuzhishan purebred pig breeds. This method uses genomics and bioinformatics technology to achieve accurate identification of white Wuzhishan pigs. The operation is simple and convenient, and the results are clear and easy to read. It can intuitively, quickly and efficiently reveal the overall relationship structure and relationship between samples / groups.
[0069] (5) This invention provides a scientific basis for the protection of the white Wuzhishan pig breed at the genomic level, which helps to promote the precise management and sustainable utilization of local pig breed resources and has important theoretical and application value. Attached Figure Description
[0070] Figure 1 This is the principal component analysis diagram of the genotype data in Embodiment 1 of the present invention.
[0071] Figure 2 This is an evolutionary tree analysis diagram of genotype data in Embodiment 1 of the present invention.
[0072] Figure 3 This is a Manhattan plot of the GWAS analysis results of the white Wuzhishan pig in Example 1 of the present invention.
[0073] Figure 4 This is a Manhattan plot of the selection signal analysis results of the white Wuzhishan pig in Embodiment 1 of the present invention.
[0074] Figure 5This is a genotype distribution map of the specific SNP locus combination (C2 group) provided by the present invention in different populations.
[0075] Figure 6 This is a graph showing the validation results of principal component analysis of the specific SNP loci combination (C1 group, threshold 1E-20) of the white Wuzhishan pig and other pig breeds.
[0076] Figure 7 This is a graph showing the validation results of principal component analysis of the specific SNP loci combination (C2 group, threshold 7E-09) of the white Wuzhishan pig and other pig breeds. Detailed Implementation
[0077] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0078] Unless otherwise specified, the technical means used in the embodiments are conventional means well known to those skilled in the art. Unless otherwise stated, the reagents, methods, and equipment used in this invention are conventional reagents, methods, and equipment in this technical field.
[0079] Example 1: Screening of SNP loci specific to the white Wuzhishan pig breed
[0080] 1. Genotype Data Acquisition
[0081] (1) The experimental group consisted of 725 pigs from 14 different breeds (strains). Among them, 11 were local Chinese pig breeds, including Hainan local breeds: Wuzhishan White Pig (n=20), Wuzhishan Pig (n=73) (a nationally recognized breed with a coat color of black and white, https: / / zypc.nahs.org.cn / pzml / classify.html), Duntou Pig (n=20), and Lingao Pig (n=24); South China type pig breeds: Bama Fragrant Pig (n=6) and Yunnan Small-eared Pig (n=33); Jianghai type pig breeds: Anqing Pig (n=20), Erhualian Pig (n=24), and Meishan Pig (n=44); Central China type pig breeds: Jinhua Pig (n=12); and North China type pig breeds: Min Pig (n=18). The introduced breeds were: Landrace Pig (n=58), Large White Pig (n=198), and Duroc Pig (n=175).
[0082] Table 1. Statistics on the Varieties, Types, and Quantities of the Experimental Population
[0083]
[0084] (2) The genotype data of the white Wuzhishan pig, Duntou pig and Lingao pig populations were obtained directly through resequencing, and the specific method is as follows:
[0085] Ear tissue samples were collected from a white Wuzhishan pig herd, and total DNA was extracted from the samples using a tissue DNA extraction kit. The integrity and purity of the DNA in each sample were checked by agarose gel electrophoresis and UV spectrophotometry. The absorbance ratio (A260 / 280) of the DNA samples was required to be between 1.8 and 2.0, and the concentration was ≥50 ng / μL. The DNA samples that passed the quality check were then subjected to genome resequencing using the DNBSEQ sequencing platform (Shenzhen BGI Genomics Co., Ltd.).
[0086] (3) Genotypic data for other pig breeds were obtained from the PGIDB database.
[0087] 2. Genotype Data Processing
[0088] (1) Genotype data merging and quality control
[0089] Genotype data from each population were merged using Plink software according to standard methods. The merged dataset was then subjected to quality control, and SNPs with a detection rate of less than 95% and individuals with a genotyping rate of less than 90% were removed. Loci with a minimum allele frequency of less than 1% and sex chromosome loci were also removed. A total of 725 individuals and 6,733,908 high-quality SNP loci were retained for subsequent analysis.
[0090] (2) Principal Component Analysis (PCA)
[0091] PCA analysis was performed on the above SNP dataset using Plink software. The output results were prefixed with "pca" and generated files ending in ".eigenval" and ".eigenvec". ".eigenval" represents the weight of each PCA, while the other file (.eigenvec) records the eigenvectors used for plotting the PCA graph. The relevant code is shown below:
[0092] plink --bfile data --pca 3 --out pca
[0093] plink: Used to invoke Plink;
[0094] --bfile data: Used to read binary format genotype files;
[0095] --pca 3: Used to perform principal component analysis and obtain the eigenvector values of the first 3 columns;
[0096] --out pca: Outputs the principal component analysis results, with the file prefixed with pca.
[0097] (3) Evolutionary tree analysis
[0098] Phylogenetic analysis was performed on the above SNP dataset using VCF2Dis and FastMe software. Specifically, VCF2Dis was used to calculate genetic distances, and FastMe was used to construct the phylogenetic tree based on the distance matrix. The output results were prefixed with "data" and generated files ending in ".mat" and ".nwk". The ".mat" file represents the genetic distance between each individual, and the ".nwk" file records the feature vectors used for drawing the phylogenetic tree. The relevant code is shown below:
[0099] VCF2Dis -InPut data.vcf -OutPut data_dis.mat
[0100] fastme -i data_dis.mat -m N -o data.nwk
[0101] VCF2Dis\fastme: Calls the corresponding software;
[0102] -InPut data.vcf: Reads the input file data (vcf format);
[0103] -OutPut data_dis.mat: Outputs genetic distance results, with the file prefixed with "data";
[0104] -i data_dis.mat: Reads the output file from the previous step;
[0105] -m N: The adjacency-based method is selected to construct the tree;
[0106] -o data.nwk: Outputs a phylogenetic tree file.
[0107] (4) Results visualization
[0108] ①PCA Analysis Visualization: The output results of step (2) are formatted using Excel. Specifically, the first column is set as breed, the second column as eigenvector value 1 (PC1), and the third column as eigenvector value 2 (PC2). The formatted data (the data is named pca) is read into R language, and the ggplot2 package is used to plot the PCA analysis chart based on the eigenvector values of the first two columns (PC1 and PC2). The X-axis is set as PC1, the Y-axis as PC2, and different breeds are represented by different colors. The relevant code is shown below:
[0109] ggplot(pca,aes(x=PC1,y=PC2,color=Breed))+
[0110] geom_point(size=3)+
[0111] theme_bw() +
[0112] theme(panel.border=element_blank(), panel.grid.major=element_blank(), panel.grid.minor=element_blank(), axis.line=element_line(colour="black"))
[0113] ggplot(): Creates a base layer;
[0114] geom_point(): Draws a scatter plot layer;
[0115] theme_bw(): Draws the theme.
[0116] The PCA analysis results are shown below. Figure 1 As can be seen from the figure, the white Wuzhishan pig shows a relatively close genetic distance with Wuzhishan pig, Duntou pig, Lingao pig, Bama fragrant pig, etc., and even has some overlap with Wuzhishan pig. This indicates that the genetic characteristics of the white Wuzhishan pig and other local pigs are relatively similar, making identification difficult.
[0117] ② Visualization of phylogenetic tree analysis: The output of step (3) is visualized using iTOL. Specifically, the result file with the suffix .nwk is uploaded, and different colors are used to annotate the individual species information.
[0118] See the results of the phylogenetic tree analysis. Figure 2 As can be seen from the figure, the branches of the white Wuzhishan pig and the Wuzhishan pig are relatively close, indicating that the genetic characteristics of the two groups are highly similar, which also shows that it is difficult to identify the two.
[0119] 3. Genome-wide association analysis was used to screen for SNP loci specific to the white Wuzhishan pig breed.
[0120] Based on step 2 (1), the GEMMA software was used to conduct a genome-wide association analysis based on a mixed linear model. The specific method was as follows: the white Wuzhishan pig was used as the experimental group, and other breeds were uniformly set as the control group. The genotype data in step 2 was converted into a GEMMA-compatible Plink binary format file. First, the GEMMA software was used to calculate the genetic relationship matrix between samples. Then, the phenotype file was specified (the experimental group was assigned a value of 1, and the control group was assigned a value of 0) for association analysis. The threshold for significant sites was set to 1E-20 or 7E-09.
[0121] like Figure 3As shown, based on the significant site threshold setting, a total of two SNP site combinations were obtained. Among them, the significant SNP site combination above the threshold line 1E-20 was named the White Wuzhishan Pig strain-specific SNP site group A1, and the significant SNP site combination above the threshold line 7E-09 was named the White Wuzhishan Pig strain-specific SNP site group A2. Figure 3 ).
[0122] 4. Select signal analysis to screen for specific SNP sites in white Wuzhishan pigs.
[0123] Based on step 2, the quality-controlled genotype data were converted to VCF format using Plink software, and then the fixed index (F) was calculated using VCFtools software. ST The specific method is as follows: White Wuzhishan pigs were used as the experimental group, and other breeds were used as the control group. The sliding window size was set to 50,000 bp, and the sliding step size was 20,000 bp. The F... ST Sort the values from largest to smallest.
[0124] like Figure 4 As shown, filter F ST The SNP sites included in the top 1% window are designated as Group B of SNP sites specific to the Bai-type Wuzhishan pig.
[0125] 5. Genome-wide association analysis and selection signal analysis: Screening for common-specific SNP sites.
[0126] The shared SNP loci (30 SNP loci) between groups A1 and B were retained as group C1, and the shared SNP loci (34 SNP loci) between groups A2 and B were retained as group C2. The combination of SNP loci in groups C1 and C2 is the specific SNP locus combination of the Bai Wuzhishan pig.
[0127] The specific SNP loci screened in group C2 are shown in Table 2. The distribution of each specific SNP locus in the genome and the genotype distribution of each specific SNP locus in different populations are shown in Table 2. Figure 5 The rows in the figure represent different SNP loci, and the columns represent the experimental individuals. SNP loci are annotated in the vcf genotype file: homozygous wild type 0 / 0, gray; heterozygous 1 / 0, yellow; homozygous mutant 1 / 1, red. (Example) Figure 5 The results showed that the genotypes of specific SNP loci in the white Wuzhishan pig population differed significantly from those in other populations.
[0128] Table 2 Information on 34 specific SNP molecular marker combinations of Bai-type Wuzhishan pigs
[0129]
[0130] Example 2: PCA Visualization Analysis for Identifying White Wuzhishan Pigs Using Specific SNP Locus Combinations
[0131] 1. Processing of genotype data
[0132] Referring to Example 1, using Plink software, the SNP locus genotype data of the 14 pig breeds / strains in Example 1 (data of 725 individuals merged in step 2) were extracted to the corresponding C1 and C2 groups (Table 2) (named D1 and D2 respectively).
[0133] 2. Application in variety identification
[0134] (1) PCA analysis
[0135] PCA analysis was performed on the genotype data D1 and D2 using Plink software. The output results were prefixed with "pca", resulting in two files ending with ".eigenval" and ".eigenvec". ".eigenval" represents the weight of each PCA, while the other file (.eigenvec) records the eigenvectors used for plotting the PCA graph. The relevant code is shown below:
[0136] plink --bfile D --pca 3 --out pca
[0137] plink: Used to launch Plink;
[0138] --bfile D: Used to read the binary format genotype file D;
[0139] --pca 3: Used to perform principal component analysis and obtain the eigenvector values of the first 3 columns;
[0140] --out pca: Outputs the principal component analysis results, with the file prefixed with pca.
[0141] (2) Visualization
[0142] The output of step (1) is formatted using Excel. Specifically, the first column is set to breed, the second column to eigenvector value 1 (PC1), and the third column to eigenvector value 2 (PC2). The formatted data (named pca) is then read into R. The ggplot2 package is used to plot a PCA analysis chart based on the eigenvector values of the first two columns (PC1 and PC2). The X-axis is set to PC1, the Y-axis to PC2, and different breeds are represented by different colors. The relevant code is shown below:
[0143] ggplot(pca,aes(x=PC1,y=PC2,color=Breed))+
[0144] geom_point(size=3)+
[0145] theme_bw() +
[0146] theme(panel.border=element_blank(), panel.grid.major=element_blank(), panel.grid.minor=element_blank(), axis.line=element_line(colour="black"))
[0147] ggplot(): Creates a base layer;
[0148] geom_point(): Draws a scatter plot layer;
[0149] theme_bw(): Draws the theme.
[0150] The PCA analysis results of the genotype data D1 corresponding to the SNP locus combinations in group C1 are shown below. Figure 6 As can be seen from the figure, although the SNP site combination of group C1 corresponds to fewer SNP sites, its clustering effect is poor and it cannot clearly separate the white Wuzhishan pig from other pig breeds, indicating that the SNP site combination of group C1 is not suitable for identifying the white Wuzhishan pig breed.
[0151] The PCA analysis results of the genotype data corresponding to the SNP loci combinations in group C2 for D2 are shown below. Figure 7 As can be seen from the figure, the C2 group of SNP loci combination has better application potential for identifying white Wuzhishan pigs. The white Wuzhishan pigs in Table 1 are clustered together, showing a relatively close genetic distance; while they are clearly separated from other pig breeds, indicating that they have a relatively far genetic distance from other pig breeds.
[0152] The above results demonstrate that PCA analysis using the combination of 34 specific SNP sites in group C2 can accurately identify the white Wuzhishan pig. Therefore, the 34 specific SNP site combinations provided by this invention can be used to identify the white Wuzhishan pig.
[0153] Example 3: Application of specific SNP locus combinations of white Wuzhishan pigs in the identification of white Wuzhishan pigs
[0154] 1. Genotype Data Acquisition
[0155] Tissue samples were collected from the pigs to be tested, and DNA was extracted from the tissue samples to obtain DNA samples from the pigs to be tested.
[0156] 2. Obtaining SNP data from the pigs to be tested
[0157] The DNA samples of the pigs to be tested undergo quality inspection. Once the quality inspection is passed, the samples are sent to a sequencing company for genotyping to obtain the genotype data of the pigs to be tested.
[0158] 3. Processing of genotype data
[0159] Referring to Example 1, the genotype data of the pigs to be tested obtained in step 2 were merged with the genotype data of the known white Wuzhishan pig population (Table 1) and the control population (the non-white Wuzhishan pig population in Table 1) using Plink software, and the genotype data of the corresponding 34 specific SNP loci were extracted in conjunction with Table 2.
[0160] 4. Application in variety identification
[0161] (1) PCA analysis
[0162] The specific method is the same as in Example 2.
[0163] (2) Visualization
[0164] The specific method is the same as in Example 2.
[0165] (3) Result determination
[0166] When the principal component analysis results show that the genetic distance between the pig to be tested and the known white Wuzhishan pig population (white Wuzhishan pig population in Table 1) is close, that is, they cluster together, the pig to be tested is a white Wuzhishan pig; otherwise, it is another breed of pig.
[0167] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.
Claims
1. The application of a product for detecting and identifying specific SNP locus combinations of the white Wuzhishan pig breed in the identification of the white Wuzhishan pig breed or in the genetic relationship analysis of white Wuzhishan pigs, characterized in that... It contains 34 SNP loci, named SNP loci 1-34, and the physical locations of these 34 SNP loci are based on the pig reference genome. Sscrofa11.1 The locations and variation information of the 34 SNP sites, determined through sequence comparison, are shown below: SNP site 1 is located at 72220198 bp on chromosome 3, and its mutation type is C / T. SNP site 2 is located at 95365104 bp on chromosome 3, and its mutation type is C / T. SNP site 3 is located at 116456303 bp on chromosome 3, and its mutation type is C / T. SNP site 4 is located at 40551892 bp on chromosome 6, and its mutation type is A / C. SNP site 5 is located at 40552273 bp on chromosome 6, and its mutation type is G / A. SNP locus 6 is located at 40552360 bp on chromosome 6, and its mutation type is A / G. SNP site 7 is located at 40553929bp on chromosome 6, and its mutation type is G / A. SNP locus 8 is located at 40553931 bp on chromosome 6, and its mutation type is A / C. SNP locus 9 is located at 164810557bp on chromosome 6, and its mutation type is G / A. SNP locus 10 is located at 56,997,264 bp on chromosome 7, and its mutation type is T / C. SNP locus 11 is located at 57005177bp on chromosome 7, and its mutation type is G / A. SNP locus 12 is located at 57016196 bp on chromosome 7, and its mutation type is G / A. SNP locus 13 is located at 57021566 bp on chromosome 7, and its mutation type is C / G. SNP locus 14 is located at 57029163 bp on chromosome 7, and its mutation type is G / A. SNP locus 15 is located at 57031495 bp on chromosome 7, and its mutation type is A / C. SNP locus 16 is located at 57032411 bp on chromosome 7, and its mutation type is C / T. SNP locus 17 is located at 57032554 bp on chromosome 7, and its mutation type is T / C. SNP locus 18 is located at 57035290 bp on chromosome 7, and its mutation type is C / T. SNP locus 19 is located at 57035378 bp on chromosome 7, and its mutation type is G / A. SNP 20 is located at 19703658 bp on chromosome 10, and its mutation type is A / G. SNP site 21 is located at 19,717,443 bp on chromosome 10, and its mutation type is G / A. SNP site 22 is located at 19,718,589 bp on chromosome 10, and its mutation type is G / C. SNP site 23 is located at 19719575 bp on chromosome 10, and its mutation type is C / T. SNP site 24 is located at 19719621 bp on chromosome 10, and its mutation type is C / T. SNP site 25 is located at 19719631 bp on chromosome 10, and its mutation type is G / A. SNP locus 26 is located at 19722528 bp on chromosome 10, and its mutation type is G / A. SNP site 27 is located at 19723319 bp on chromosome 10, and its mutation type is G / A. SNP site 28 is located at 19723330bp on chromosome 10, and its mutation type is G / C. SNP 29 is located at 19724019 bp on chromosome 10, and its mutation type is A / G. SNP site 30 is located at 19,757,672 bp on chromosome 10, and its mutation type is A / G. SNP site 31 is located at 56826128 bp on chromosome 11, and its mutation type is G / A. SNP site 32 is located at 36198591 bp on chromosome 13, and its mutation type is C / T. SNP site 33 is located at 110336387bp on chromosome 13, and its mutation type is A / G. SNP site 34 is located at 207921057 bp on chromosome 13, and its mutation type is G / A.
2. The application according to claim 1, characterized in that: The products mentioned are reagents, kits, or gene chips.
3. The application according to claim 2, characterized in that: The reagents include at least one of primers and probes.
4. A molecular probe array for identifying the white Wuzhishan pig breed, characterized in that... The molecular probe combination targets the specific SNP site combination described in claim 1, and is used to detect the specific SNP site combination described in claim 1.
5. A gene chip for identifying the white Wuzhishan pig breed, characterized in that... The chip load has the molecular probe assembly as described in claim 4.
6. A method for rapid identification of the white Wuzhishan pig breed, characterized in that... It includes the following steps: (1) Collect tissue samples from the pigs to be tested, extract DNA from the tissue samples, and obtain DNA samples from the pigs to be tested; (2) Perform gene sequencing on the DNA sample of the pig to be tested obtained in step (1) to obtain the genotype data of the pig to be tested; (3) The genotype data of the pig to be tested obtained in step (2) is merged with the genotype data of the known white Wuzhishan pig population and the control population, and the genotype data of the specific SNP site combination described in claim 1 is extracted. The merged and extracted genotype data is subjected to principal component analysis, and it is determined whether the pig to be tested is a white Wuzhishan pig. The principal component analysis described in step (3) is implemented using the ggplot2 package in R language to visualize the principal components; The method for judgment described in step (3) is as follows: When the principal component analysis results show that the pig to be tested is genetically close to the known white Wuzhishan pig population, i.e., clustered together, then the pig to be tested is a white Wuzhishan pig.
7. The method for rapid identification of the white Wuzhishan pig breed according to claim 6, characterized in that: The known white Wuzhishan pig population mentioned in step (3) is a group composed of multiple purebred white Wuzhishan pigs; The control group mentioned in step (3) is other pig groups of non-white Wuzhishan pigs.
8. The method for rapid identification of the white Wuzhishan pig breed according to claim 6, characterized in that: The control group is at least one of the following: Wuzhishan pig, Duntou pig, Lingao pig, Bama fragrant pig, Diannan small-eared pig, Anqing pig, Erhualian pig, Meishan pig, Jinhua pig, Min pig, Landrace pig, Large White pig, and Duroc pig.
Citation Information
Patent Citations
Specific SNP (Single Nucleotide Polymorphism) site combination for identifying temporary high pig variety and application
CN119061167A
Specific SNP (Single Nucleotide Polymorphism) site combination for identifying Wuzhishan pig variety and application
CN121160887A