SNP site combination for identifying jinfen white pigs and application thereof
Patent Information
- Application Number
- CN202410631037.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-21
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2044-05-21
AI Technical Summary
现有技术中,基于全基因组重测序或高密度SNP芯片数据进行品种鉴别的成本相对较高,因此,有必要研发一种新的适合用于鉴别晋汾白猪的方法
本发明通过大量研究筛选得到了包括26个SNP位点的SNP位点组合,这26个SNP位点和晋汾白猪品种高度相关,基于这些SNP位点可以将晋汾白猪从大白猪、长白猪、杜洛克猪、马身猪和梅山猪等多个群体中鉴别出来。本发明提供的SNP位点组合可以快速、准确地实现晋汾白猪的鉴别,可作为分子标记物用于猪种质资源鉴定。此外,本发明提供的SNP位点组合还可以用于疑似晋汾白猪个体的品种鉴定,对晋汾白猪群体规模的扩大及后续合理开发利用有重要意义。
Smart Images

Figure CN118516467B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular biology, and in particular to a combination of SNP sites for identifying Jinfen White Pigs and its application. Background Technology
[0002] Jinfen White Pig is a new breed of pig developed through complex crossbreeding and population-based successive selection combined with molecular marker-assisted selection, using Ma Shen Pig, Erhua Lian Pig, Landrace Pig, and Large White Pig as breeding materials. Jinfen White Pigs have white, glossy coats, a compact and sturdy build, a moderately sized head with a slightly concave face, medium-sized, slightly erect ears tilted forward and to the side, a relatively long body, a wide back, a straight back and loin, a broad and deep chest, a tapering abdominal line, and a full rump. They have strong limbs and sturdy hooves and toes. Their teats are evenly and neatly arranged, well-developed, with at least seven pairs of effective teats. They are characterized by high litter size, rapid growth, good carcass and meat quality, good adaptability, and strong disease resistance.
[0003] SNP (Single Nucleotide Polymorphism) refers to a single nucleotide variation in the entire genome, caused by transformation, transversion, insertion, or deletion. With the decreasing cost of whole-genome sequencing technology, SNPs are increasingly used in animal breed identification. Compared to other DNA molecular markers, SNPs have several advantages, including abundant abundance in the genome, diverse types of variations, low mutation frequency, and relatively high genetic stability. Due to the large number and wide distribution of SNPs in the genome, and with the advancement of gene sequencing technology enabling high-throughput and automated detection, SNPs have become one of the most popular molecular markers, widely used in research across many fields such as biology, agriculture, medicine, and evolutionary biology. Currently, breed identification based on whole-genome resequencing or high-density SNP microarray data is relatively expensive; therefore, it is necessary to develop a new method suitable for identifying Jinfen White Pigs. Summary of the Invention
[0004] To address the problems existing in the prior art, this invention provides a combination of SNP sites for identifying Jinfen White Pigs and its application.
[0005] In a first aspect, the present invention provides a combination of SNP sites based on the Sscrofa11.1 genome version, including the following 26 SNP sites: The SNP locus located at 11992266 bp on chromosome 1 has a polymorphism of G / A. The SNP locus located at 11992282 bp on chromosome 1 has a polymorphism of C / T. The SNP locus located at 220534672 bp on chromosome 1 has a polymorphism of T / C. The SNP locus located at 250998626 bp on chromosome 1 has a polymorphism of C / T. The SNP locus at 257859658 bp on chromosome 1 has a polymorphism of C / G. The SNP locus located at 261177743 bp on chromosome 1 has a polymorphism of A / G. The SNP locus at 2672797bp on chromosome 2 has a polymorphism of G / A. The SNP locus located at 36044769 bp on chromosome 2 has a polymorphism of G / A. The SNP locus located at 129051455 bp on chromosome 3 has a polymorphism of T / C. The SNP locus at 26920894 bp on chromosome 4 has a polymorphism of C / T. The SNP locus located at 42811324 bp on chromosome 5 has a polymorphism of G / A. The SNP locus located at 77909528 bp on chromosome 5 has a polymorphism of C / A. The SNP locus located at 135404148 bp on chromosome 6 has a polymorphism of C / T. The SNP locus located at 42966924 bp on chromosome 7 has a polymorphism of G / T. The SNP locus located at 117482672 bp on chromosome 7 has a polymorphism of T / C. The SNP locus located at 138027774 bp on chromosome 9 has a polymorphism of G / A. The SNP locus located at 60481456 bp on chromosome 10 has a polymorphism of A / G. The SNP locus located at 19,469,269 bp on chromosome 10 has a polymorphism of C / A. The SNP locus at 60572614 bp on chromosome 10 has a polymorphism of C / G. The SNP locus at 15288071 bp on chromosome 14 has a polymorphism of C / T. The SNP locus is located at 841647bp on chromosome 14, and the polymorphism is A / T. The SNP locus at 78065845 bp on chromosome 14 has a polymorphism of C / T. The SNP locus at 99209752 bp on chromosome 14 has a polymorphism of G / A. The SNP locus at 99209753 bp on chromosome 14 has a polymorphism of C / G. The SNP locus at 119628839 bp on chromosome 14 has a polymorphism of A / T. The SNP locus is located at 15430487 bp on chromosome 3, and the polymorphism is C / T.
[0006] Furthermore, the 26 SNP sites mentioned above correspond to the following in order: The SNP site located at position 319 of the sequence shown in SEQ ID NO.1 has a polymorphism of G / A; The SNP site located at position 335 of the sequence shown in SEQ ID NO.2 has a polymorphism of C / T; The SNP site located at position 138 of the sequence shown in SEQ ID NO.3 has a polymorphism of T / C; The SNP site located at position 156 of the sequence shown in SEQ ID NO.4 has a polymorphism of C / T; The SNP site located at position 265 of the sequence shown in SEQ ID NO.5 has a polymorphism of C / G; The SNP site located at position 392 of the sequence shown in SEQ ID NO.6 has a polymorphism of A / G; The SNP site located at position 244 of the sequence shown in SEQ ID NO.7 has a polymorphism of G / A; The SNP site located at position 400 of the sequence shown in SEQ ID NO.8 has a polymorphism of G / A; The SNP site located at position 277 of the sequence shown in SEQ ID NO.9 has a polymorphism of T / C; The SNP site located at position 381 of the sequence shown in SEQ ID NO.10 has a polymorphism of C / T; The SNP site located at position 250 of the sequence shown in SEQ ID NO.11 has a polymorphism of G / A; The SNP site located at position 389 of the sequence shown in SEQ ID NO.12 has a polymorphism of C / A; The SNP site located at position 153 of the sequence shown in SEQ ID NO.13 has a polymorphism of C / T; The SNP site located at position 389 of the sequence shown in SEQ ID NO.14 has a polymorphism of G / T; The SNP site located at position 354 of the sequence shown in SEQ ID NO.15 has a polymorphism of T / C; The SNP site located at position 355 of the sequence shown in SEQ ID NO.16 has a polymorphism of G / A; The SNP site located at position 346 of the sequence shown in SEQ ID NO.17 has a polymorphism of A / G; The SNP site located at position 370 of the sequence shown in SEQ ID NO.18 has a polymorphism of C / A; The SNP site located at position 222 of the sequence shown in SEQ ID NO.19 has a polymorphism of C / G. The SNP site located at position 285 of the sequence shown in SEQ ID NO.20 has a polymorphism of C / T; The SNP site located at position 327 of the sequence shown in SEQ ID NO.21 has a polymorphism of A / T; The SNP site located at position 152 of the sequence shown in SEQ ID NO.22 has a polymorphism of C / T; The SNP site located at position 399 of the sequence shown in SEQ ID NO.23 has a polymorphism of G / A; The SNP site located at position 400 of the sequence shown in SEQ ID NO.24 has a polymorphism of C / G; The SNP site located at position 104 of the sequence shown in SEQ ID NO.25 has a polymorphism of A / T; The SNP site located at position 353 of the sequence shown in SEQ ID NO.26 has a polymorphism of C / T.
[0007] The present invention further provides a primer combination, the primer combination comprising primers for amplifying all SNP sites in the SNP site combination: 1.2F:ATAGAATCTGCATATTGAGACCCA, 1.2R:CGTAAAAGTGCAATCATGAAGCC; 3F:TCAAATGCCCTCCACTACTTCC, 3R: GGGGTTTATTGATTGTTGGGGAC; 4F: AACCTGCGCTGAACTTCTGA, 4R: AACAGACCTCCCCTTCCTGT; 5F:AATCCTCCCCATCCACTAAAAGC, 5R: TGCAGTTTGGGACTCTTCCTG; 6F: TCATGTCATTTCCATTCAATAGGC, 6R: AATTTTCCAGAGCCAGCCCTC; 7F: GACACTGTGGCACTTCCTGA, 7R: CCCCTCCTGAAAGGCTGATG; 8F:GCTGTCTTGCTTCCAATTGTCT, 8R:GGCACGACATATAGGCTGCT; 9F: GCCAGCAAAGGACAAAGC, 9R:TGGGGCACTTGAAAGGGTT; 10F:GTCATAGGTGTCAGGCCTTTCT, 10R:TTTCGAGCACTCTCCATCATT; 11F:TTTGCTGCGTTTTAGCTCCAC, 11R:TGAGTAGCTGTTCGGTGCAA; 12F: CCTCCTCAAAGCCTGCTTCA, 12R:GCTTTGCTTCCCCAAACCTG; 13F: TGATGCAGGGAGAGGAGGAA, 13R: ATGCATTCAGTTGTACCGGC; 14F:GGAGGCAGTAGATACCCAGATAC, 14R:CCCTCATTTCCAGGCTAATGC; 15F:TTAAGGCCGCCAACTTCTACC, 15R: TAAGAAGCGTGGGCGACAA; 16F: ATCGTCTCCCGCGGATTAGG, 16R:ACAGAGTCAGTGTATCCATCACC; 17F: CTTCTGGGGACAGGTCAGGA, 17R:CATGAGCAGGAGATCAGGGC; 18F: TCCACCTCGAGGGAAACAGA, 18R:CCCTTGTACTGGTCTGCGAG; 19F: ATCAGCCCGCAATAACCACA, 19R: GGACGGAACGAAGACAAGGT; 20F: CATAGCCAGACAGCAGGACC, 20R: TCAAAGGAGGTGAGGGTG; 21F:GTCCTGGCAAGACTCTCAGG, 21R:AGACCAACCGCCTTCTCTTG; 22F:TTCTGCTGCCCTTACCCTTG, 22R:AGGGAGACATGAGCAAAGCC; 23.24F:AGTTGGATCTTCAGCACTAGGAA, 23.24R:TACCATTGGCACAGTTGCAC; 25F:TAGAGCGAAAATTTCAGAGCCAC, 25R:TTTCCCTTCAACTCTAGACCACTTT; 26F:GAGCCCCCGAATGTCCTCTA, 26R:GTGATGGGGTTACTTGACCG。
[0008] gagcccccgaatgtcctctacctcccagcaaggcatagaaggaaggcttggagcaagcgaagaggcagtgctaggcaagtataatgtcagagcagaggtctgggggggacatacagcaggttcaggggacatcagcacagtggctgaggcatgtggtccagagtgatcaatagagacagggtcctgagagcccaggatggagggtcagcctgcccacccagcatcctcaggtgcagcccaggggacttcaagccaccaggaacttgccacaaatgctgattccttgagtctaaccccagacctactagggtggggcctgcatactagtaggcatctccaaagacattgttttc tcagtccgcacaatgtcttgagcatcagttatgggcaggacaatgtgttagatttaaaggtactgatcatgcgttcctgtttgcccaaggcaggcttaggttacgtctagtatcccaacctaattt ttttttaatttataattctattttagtttttatataatttgattttcacttgtaaggttaatttcactctttttaatgtagctttgcaacttttgacaaacacatacggtcaagtaaccccatcac gccagcaaagaggacaaagcaaagcttcgtgtccttatctccttccttcgacgaagaaagcagcgtttctggaaggggcatcagccctttgccgggtgctggtttctattctgtcttgctgcccccgtggggtgacgtttgaaata aggcccatgtaattgactgcatcctgattcatcagcttcgtcaatggtgcataggtagtctgactgtcactctcaaggcatgcttgtcagacgagggaggataattccaagggagtctggcttcctccactgcgtgtacacatgatc acagctcctccccaaagtggtgacgagcccgggtcgctgtccgacctcccttccttcttggtgatttacactctgggcctctgatatcctttcccctgagacctttgtcagaaccctcttggtccatcaaaatggttctaataata ttgaaccatagaatgggaagggttagagtgtcaatcttttacctccgggaaggggccttcacacatctggaatcgcactgcttctcggcagaccagctgtgcagacaaaggaggtactctgatgggtaacccttttcaagtgcccca The present invention further provides a gene chip for identifying the Jinfen White Pig breed, wherein the gene chip includes nucleic acid probes for detecting the SNP site combinations.
[0009] Furthermore, the nucleic acid probe comprises the following sequence: The sequence consisting of the first (15~50) bp and the last (15~50) bp of each SNP site.
[0010] Secondly, the present invention provides the application of the SNP site combination, the primer combination, the gene chip, or the kit in the breed identification of Jinfen White Pig.
[0011] Thirdly, the present invention provides a method for identifying the breed of Jinfen White Pig, comprising: Genomic DNA is extracted from the pigs to be tested, and the genotype information of the SNP locus combination of the pigs to be tested is detected. The breed of the pigs to be tested is identified based on the detection results.
[0012] Furthermore, the genotype information of the SNP loci can be detected using the primer combination, the gene chip, or the kit described above.
[0013] Furthermore, the step of identifying the breed of the pig to be tested based on the test results includes: If the test results are as follows, the pig to be tested is determined to be a Jinfen White Pig: The SNP locus located at 11992266 bp on chromosome 1 was detected as AA / AG. The SNP locus located at 11992282bp on chromosome 1 showed a result of TT / TC. The SNP locus located at 220534672bp on chromosome 1 was detected as CC / CT. The SNP locus located at 250998626 bp on chromosome 1 was detected as TT / TC. The SNP locus located at 257859658 bp on chromosome 1 was detected as GG / GC. The SNP locus located at 261177743bp on chromosome 1 showed a result of GG / GA. The SNP locus located at 2672797bp on chromosome 2 showed a result of AA / AG. The SNP locus located at 36044769bp on chromosome 2 was detected as AA / AG. The SNP locus located at 129051455bp on chromosome 3 was detected as CC / CT. The SNP locus located at 26920894bp on chromosome 4 was detected as TT / TC. The SNP locus located at 42811324bp on chromosome 5 was detected as AA / AG. The SNP locus located at 77909528 bp on chromosome 5 showed a result of AA / AC. The SNP locus located at 135404148bp on chromosome 6 was detected as TT / TC. The SNP locus located at 42966924bp on chromosome 7 showed a result of TT / TG. The SNP locus located at 117482672bp on chromosome 7 was detected as CC / CT. The SNP locus located at 138027774bp on chromosome 9 was detected as AA / AG. The SNP locus located at 60481456 bp on chromosome 10 showed a result of GG / GA. The SNP locus located at 19469269bp on chromosome 10 showed a result of AA / AC. The SNP locus located at 60572614bp on chromosome 10 showed a result of GG / GC. The SNP locus located at 15288071 bp on chromosome 14 was detected as TT / TC. The SNP locus located at 841647bp on chromosome 14 was detected as TT / TA. The SNP locus located at 78065845bp on chromosome 14 was detected as TT / TC. The SNP locus located at 99209752 bp on chromosome 14 was detected as AA / AG. The SNP locus located at 99209753bp on chromosome 14 was detected as GG / GC. The SNP locus located at 119628839 bp on chromosome 14 was detected as TT / TA. The SNP locus located at 15430487bp on chromosome 3 was detected as TT / TC.
[0014] The accuracy rate of this method can reach over 97%.
[0015] Fourthly, the present invention provides a method for screening SNP sites, comprising: (1) After aligning the whole genome SNP data of the sample to be tested to the pig reference genome, the first set of SNP sites was obtained by using GATK hard filtering for variant identification and screening. (2) Select SNP sites from the first SNP site set that have significant differences in allele frequencies between the target population and other populations to obtain the second SNP site set; (3) For the second set of SNP sites, site selection is performed using a support vector machine.
[0016] Further, step (2) includes: In the first SNP locus set, the allele frequency of each SNP in different varieties is obtained; the sum of the absolute values of the differences in allele frequencies of each SNP in the target population and all other populations is calculated, and the top (2.5~4.5) N SNP loci with an adjacent distance greater than 2.5Mb are selected from the largest to the smallest to form the second SNP locus set. N represents the number of mutations.
[0017] Furthermore, the sum of the absolute values of the allele frequency differences between each SNP in the target population and all other populations is calculated using the following formula: Where i represents the test group i, and j represents the test group j. The allele frequency of SNP f in population i is given. Let f be the allele frequency of SNP f in population j, and p be the total population frequency.
[0018] The sum of allele frequencies of SNP f in population i and other populations is used to define the magnitude of genotypic difference between SNP f in population i and other populations. This involves considering the frequencies of m SNPs... Sort and select SNPs with higher values (which can be defined by the user) are selected.
[0019] The present invention has the following beneficial effects: This invention, through extensive research and screening, has yielded a combination of 26 SNP loci, which are highly correlated with the Jinfen White pig breed. Based on these SNP loci, the Jinfen White pig can be distinguished from multiple populations, including Large White, Landrace, Duroc, Mason, and Meishan pigs. The SNP loci combination provided by this invention can rapidly and accurately identify the Jinfen White pig and can serve as a molecular marker for pig germplasm resource identification. Furthermore, the SNP loci combination provided by this invention can also be used for breed identification of suspected Jinfen White pig individuals, which is of great significance for expanding the Jinfen White pig population and its subsequent rational development and utilization. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in this invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention.
[0021] Figure 1The results of principal component analysis of 226 individuals from 29 populations, including Jinfen White Pig, Large White Pig, Landrace Pig, Duroc Pig, Mashan Pig, and Meishan Pig, based on 15,639,137 SNP loci, provided in Embodiment 2 of the present invention.
[0022] Figure 2 The results are based on principal component analysis of 226 individuals from 29 populations, including Jinfen White Pig, Large White Pig, Landrace Pig, Duroc Pig, Mashan Pig, and Meishan Pig, using 26 SNP loci provided in Embodiment 2 of this invention.
[0023] Figure 3 This is the result of the individual genetic distance analysis of 226 individuals from 29 populations, including Jinfen White Pig, Large White Pig, Landrace Pig, Duroc Pig, Mashan Pig, and Meishan Pig, based on 26 SNP loci provided in Example 2 of this invention.
[0024] Figure 4 This is a heatmap of genotype clustering of 226 individuals from 29 populations, including Jinfen White Pig, Large White Pig, Landrace Pig, Duroc Pig, Mashan Pig, and Meishan Pig, based on 26 SNP loci provided in Embodiment 2 of the present invention. Red indicates homozygous mutation, yellow indicates heterozygous mutation, blue indicates no mutation, and gray indicates missing genotype. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0026] Example 1 This embodiment provides a method for screening SNP sites, including the following steps: 1. Screening of SNP sites (1) Fourteen Jinfen White Pig individuals and 28 Mashen Pig individuals were selected for second-generation resequencing at a depth of 30× to obtain whole-genome SNP data and 184 reported whole-genome SNP data from 27 wild boar and local pig populations were processed. (2) Use Fastqc software (v0.11.9) to remove potential adapter sequences and low-quality bases at both ends of the sequencing reads; (3) Subsequently, BWA-MEM2 was used to align all remaining high-quality reads to the pig Sscrofa11.1 (GCA_000003025.6) reference genome; (4) Use samtools rmdup to remove PCR repetitive sequences, then convert the alignment results to bam format and sort them according to the aligned genomic positions; (5) Merge all files and use GATK hard filtering to identify and screen for mutations. The screening parameters are as follows: QD<2.0, FS>60.0, SOR>3.0, MQ<40.0, MQRankSum<-12.5, ReadPosRankSum<-8.0; (6) Perform HWE equilibrium test using PLINK software, and set the HWE test p-value to less than 1e. -6 The mutations were filtered, and then SNPs with a missing percentage greater than 10% were filtered to obtain the final mutation set, which yielded a total of 15,639,137 SNPs.
[0027] 2. Further screening of SNP sites and design of corresponding gene chips 2.1 Characteristics of candidate sites on the chip: SNPs that can be used for variety identification have large differences in frequency distribution across different populations. They can distinguish different varieties through genotype clustering and can also distinguish varieties from hybrid offspring.
[0028] 2.2 Calculation of Population Allele Frequency and Calculation of SNP Allele Frequency Differences Between Populations: Based on the above variation filtering, the allele frequencies of different varieties are statistically analyzed for each SNP. Then, the sum of the absolute values of the differences in SNP frequencies between the target population and all other populations is calculated for each population. The SNPs are then sorted from largest to smallest based on this sum. Details are as follows: Where i represents the test group i, and j represents the test group j. The allele frequency of SNP f in population i is given. Let f be the allele frequency of SNP f in population j, and p be the population total. The sum of allele frequencies of SNP f in population i and other populations is used to define the magnitude of genotypic difference between SNP f in population i and other populations. This involves considering the frequencies of m SNPs... Sort and select The top N SNPs with the highest values are selected.
[0029] 2.3 Population-Specific SNP Screening: Based on the sum of frequencies of each population and all other populations, the top 3N SNPs (N being the target variant number, N = 32 in this validation) were selected as candidate SNPs. To ensure that SNPs linked to the same LD interval are not selected repeatedly, the distance between different candidate SNPs should be no less than 2.5 Mb.
[0030] 2.4 Locus Evaluation: After obtaining SNP loci, machine learning (support vector machine) methods were used for locus selection and effect evaluation, specifically as follows: 2.4.1 Using the sequenced pig population as the training set, the wild boar population was removed, and individuals belonging to the target population were coded as 1, while other pigs were coded as 0. Support vector machine (SVM) was used to screen SNP sites and evaluate accuracy based on four types of SVM regression kernels (linear kernel, multiple normal distribution kernel, sigmoid function kernel, and radial kernel) and 10-fold cross-validation. The SNP set for chip design was determined based on whether the frequency of the SNP site having an effect greater than 0 in the 10-fold cross-validation process reached 50%. Finally, the top 72 sites with the largest effects were obtained as candidate sites.
[0031] 2.4.2 Due to the large number of selected sites, SNP site densities of 32, 48, 60 and 72 were subsequently set. The effectiveness of SNPs in distinguishing the target population from other pig populations was evaluated using the support vector machine and 10-fold cross-validation method described above. The smallest set of SNPs with an average prediction effect of greater than 97.5% after 10-fold cross-validation was selected as the final design site for the chip.
[0032] The support vector machine method is as follows: Assumption set Let i be a set of genotype-population information, where i is an individual, x is a SNP, y is population information, and n is the total number of SNPs. Individuals belonging to a certain population have their own unique genotype information on n SNPs.
[0033] Suppose there exists a hyperplane b such that: Where w is the weight of x. Then, the objective function can be optimized as follows: } This allows us to obtain b and w, which can then be used to distinguish which group y an individual belongs to. In the calculation process, we often assume a relationship between x and y, and then calculate... The relationships between x and y are determined by the relationships between x during the process; this is called the kernel SVM method. In SVM, several kernel functions are commonly used for calculation, as follows: , in, , and Kernel functions define the relationships between x and y, between x and x, or between y and y. The distributions of x and y are defined separately. After obtaining the hyperplane, the relationships between x and x, and between y and y, are defined using kernels to obtain the other unknown x and y. Therefore, k can be a Gaussian kernel, a linear kernel, a sigmoid function kernel, a multi-Gaussian kernel, or a radiometric function kernel, etc.
[0034] 2.5 Microarray Locus Probe Synthesis and Evaluation: After identifying the SNP microarray loci, probes were designed based on their genomic flanking information, and the SNP genotypes were detected using the Sequenom method. The microarray was applied to M individuals from known populations to verify the probe's effectiveness and accuracy. The specific methods are as follows: (1) Model training: Using the above chip design background population and the SNP genotype on the chip as the training set, individuals belonging to the target population are coded as 1 and other pigs are coded as 0. Model screening and model parameter estimation are performed using SVM and 10-fold cross-validation. The model with the highest average accuracy of 10-fold cross-validation and its corresponding parameters are used for subsequent validation. (2) Prediction using the trained model: The model and corresponding parameters selected in the first step are used to predict the population to which the M validated individuals belong based on their genotypes; (3) Individual significance: Due to the randomness of 10-fold cross-validation, some individuals may not be accurately predicted in a certain cross-validation sampling due to the bias between the training set and the prediction set. Repeating the model training-prediction process can overcome this problem. To overcome the lack of robustness of the model caused by its randomness, the above model training-validation process is repeated 20 times. Based on the frequency of the target group individuals in the background group in step (1) being identified as belonging to the target group in the 20 times, it is determined whether the M validation individuals are significantly belonging to the target group in the model prediction. For example, if the individuals in the background group in step (1) are identified as target group individuals on average 18 times (90% probability) in the 20 model training-validation process, then at least 18 of the M validation individuals are identified as target group individuals, which is a significant individual belonging to the target group.
[0035] 3. Screening Results Ultimately, 26 SNP loci were obtained that can be used for the identification of Jinfen White Pigs, including the following: The SNP locus located at 11992266 bp on chromosome 1 has a polymorphism of G / A. The SNP locus located at 11992282 bp on chromosome 1 has a polymorphism of C / T. The SNP locus located at 220534672 bp on chromosome 1 has a polymorphism of T / C. The SNP locus located at 250998626 bp on chromosome 1 has a polymorphism of C / T. The SNP locus at 257859658 bp on chromosome 1 has a polymorphism of C / G. The SNP locus located at 261177743 bp on chromosome 1 has a polymorphism of A / G. The SNP locus at 2672797bp on chromosome 2 has a polymorphism of G / A. The SNP locus located at 36044769 bp on chromosome 2 has a polymorphism of G / A. The SNP locus located at 129051455 bp on chromosome 3 has a polymorphism of T / C. The SNP locus at 26920894 bp on chromosome 4 has a polymorphism of C / T. The SNP locus located at 42811324 bp on chromosome 5 has a polymorphism of G / A. The SNP locus located at 77909528 bp on chromosome 5 has a polymorphism of C / A. The SNP locus located at 135404148 bp on chromosome 6 has a polymorphism of C / T. The SNP locus located at 42966924 bp on chromosome 7 has a polymorphism of G / T. The SNP locus located at 117482672 bp on chromosome 7 has a polymorphism of T / C. The SNP locus located at 138027774 bp on chromosome 9 has a polymorphism of G / A. The SNP locus located at 60481456 bp on chromosome 10 has a polymorphism of A / G. The SNP locus located at 19,469,269 bp on chromosome 10 has a polymorphism of C / A. The SNP locus at 60572614 bp on chromosome 10 has a polymorphism of C / G. The SNP locus at 15288071 bp on chromosome 14 has a polymorphism of C / T. The SNP locus is located at 841647bp on chromosome 14, and the polymorphism is A / T. The SNP locus at 78065845 bp on chromosome 14 has a polymorphism of C / T. The SNP locus at 99209752 bp on chromosome 14 has a polymorphism of G / A. The SNP locus at 99209753 bp on chromosome 14 has a polymorphism of C / G. The SNP locus at 119628839 bp on chromosome 14 has a polymorphism of A / T. The SNP locus is located at 15430487 bp on chromosome 3, and the polymorphism is C / T.
[0036] The genotypic distribution of these 26 SNP loci in Jinfen White pigs, Large White pigs, Landrace pigs, Duroc pigs, Eurasian pigs, and Meishan pigs is shown in Table 1 (due to the large amount of data, only a few representative groups were selected for display). In the table, JFW represents Jinfen White pigs, LW represents Large White pigs, LL represents Landrace pigs, DD represents Duroc pigs, MS represents Eurasian pigs, and MM represents Meishan pigs. The predominant genotype of each specific SNP locus in Jinfen White pigs differs from that of other pig breeds. Therefore, by combining the 26 specific SNP loci, Jinfen White pigs can be identified from 29 other groups, including Large White pigs, Landrace pigs, Duroc pigs, Eurasian pigs, and Meishan pigs.
[0037] Table 1. Distribution of 26 specific SNP loci (1-13) in 6 varieties Locus 1 2 3 4 5 6 7 8 9 10 11 12 13 Chr 1 1 1 1 1 1 2 2 3 4 5 5 6 JFW1 -- -- CC -- GG -- -- -- TT TT -- AA TT JFW2 -- -- CC -- CG -- -- -- TT -- -- AA -- JFW3 -- -- CC -- -- -- AA -- TT -- -- AA TT JFW4 AA TT CC -- -- -- -- -- TT -- AA AA TT JFW5 -- -- TC -- GG -- -- AA TT TT AA CA TT JFW6 AA TT CC -- GG -- -- -- TT TT -- CA TT JFW7 AA TT CC -- GG GG -- -- TT TT AA -- TT JFW8 -- -- CC TT -- -- -- -- TT TT AA AA TT JFW9 AA TT CC -- GG -- -- -- CT TT GA AA TT JFW10 -- -- CC -- GG GG -- -- TT -- -- AA -- JFW11 -- -- CC -- GG -- -- -- TT -- -- AA -- JFW12 -- -- CC TT GG GG -- -- TT TT -- AA TT JFW13 -- -- CC -- GG -- -- -- TT -- AA AA -- JFW14 -- -- CC TT GG -- -- -- TT TT -- AA -- LL1 GG CC TT CC CG AA GG GG CC CC GG CC CC LL2 GG CC TT CC CC AA GG GG CC CC GA CC CC LL3 GG CC TT CC CC AA AA GG CC CT GA CC CC LL4 GG CC TT CC CC AA GG GG CC CC GG CC CC LL5 GG CC TT CC CC AG GG GG CC CT GA CC TT LW1 GG CC TT CC CC AA AA GG CC CC GG CC TT LW2 GG CC TT CC CC AA GG GG CC CT GG CC CT LW3 GG CC TT CC CC AA GG GG CC TT GG CC CC LW4 GG CC TT CC CC AA AA GG CC TT GG CC CC LW5 GG CC TT CC CC AA GG GG CT CT GG CC CC LW6 GG CC TT CC CC AA GG GG CC CC GG CC CT LW7 GG CC TT CC CC AA GG GG CC CT GG CC CT LW8 GG CC TT CC CC AA GG GG CC CT GG CC CT LW9 GG CC TT CC CC AA GG GG CC CC GG CC TT LW10 GG CC TT CC CC AA GG GG CC CT GG CC CT LW11 GG CC TT CC CC AA AA GG CC CC GG CC TT LW12 GG CC TT CC CC AA GG GG CC [[ID= -- -- -- -- -- -- -- -- -- AA TT TT CC CC AA GG GG CC -- GG -- -- MS5 -- -- TC CC CC AA GG GG CC -- GG CC CC MS6 GA CT TT CC CC AA GG GG CC CC GG -- -- MS7 -- -- TT CC CC AA -- GG CC CC -- CA -- MS8 AA TT TT CC CC AA GG GG CC -- GA CC -- MS9 GA TT TC CC CC AA GG GG CC CT GG CC -- MS10 GA CT TC CC CC AA GG GG CC -- GG AA -- MS11 AA TT TC CC CC AA GG GG CC CC AA -- -- MS12 AA TT TC CC CC AA GG GG CC CC GG CA -- MS13 -- -- TT CC CC AA GG GG CC -- GG -- -- MS14 AA TT TT CC CC AA GG GG CC -- GG AA -- MS15 GA CT TT CC CC AA GG GG CC -- GG -- -- MS16 -- -- TT CC CC AA GG GG CC -- GG -- -- MS17 GG CC TC CC CC AA GG -- CT CC GG AA CC MS18 -- -- TC CC CC AA GG GG CC -- AA CC -- MS19 -- -- TC CC CC AA GG GG CC CC GG CC CC MS20 GA CT TC CC CC AA GG GG CC -- GA CA CC MS21 -- -- TC CC CC AA GG GG CC -- -- AA CC MS22 GA CT TC CC CC AA -- GG CT -- GA CC CC MS23 -- -- TT CC CC AA -- GG CT -- GG AA TT MS24 GG CC TC CC CC AA GG GG CC -- GA CA CC MS25 -- TT TC CC CC AA GG GG CC -- GG CA CC MS26 -- -- TC CC CC AA GG GG CC -- GG AA -- MS27 -- -- TC CC CC AA GG GG CC -- GG -- CC MS28 AA TT TC CC CC AA GG GG CC -- GA CC CC MM1 GG CC TT CC CC AA GG GG CC CC GG CC CC MM2 GG CC TT CC CC AA GG GG CC CC GG CC CC MM3 GG CC TT CC CC AA GG GG CC CC GG CC CC MM4 GG CC TT CC CC AA GG GG CC CC GG CC CC MM5 GG CC TT CC CC AA [[ID=3 Table 2. Distribution of 26 specific SNP loci (14-26) in 6 varieties 14 15 16 17 18 19 20 21 22 23 24 25 26 7 7 9 10 10 10 14 14 14 14 14 14 18 -- -- -- -- AA -- TT TT -- AA GG TT TT JFW3 TT -- AA -- AA -- TT TT -- AA GG -- TT JFW4 -- CC AA GG AA -- TT TT -- AA GG TT TT JFW5 TT CC AA GG AA -- TT TT -- AA GG -- TT JFW6 -- -- AA -- CA GG TT TT -- AA GG -- TT JFW7 -- CC AA -- AA -- TT TT TT -- -- -- TT JFW8 -- TC AA -- AA -- TT TT TT AA GG -- TT JFW9 -- -- -- -- AA -- TT TT TT -- -- -- CT JFW10 TT CC -- -- AA -- TT TT TT AA GG TT TT JFW11 TT CC AA -- AA -- TT TT TT -- -- -- TT JFW12 TT -- AA -- AA -- TT TT -- AA GG -- TT JFW13 -- CC AA -- AA -- TT TT -- AA GG -- TT JFW14 TT -- -- -- AA -- TT TT -- AA GG -- TT LL1 GG TT GA AA CC CC CC AA CT GG CC AA CC LL2 GG TT GG AA CC CG CC AA TT GG CC AA CC LL3 GG TT GG AA CC CC CC AA CT GG CC AA CC LL4 GG TT GG AA CC CC CC AA CC AA CC AA CC LL5 GG TT GG AA CC CC CC AA TT GG GG AA TT LW1 GG TT GG AA CC CC CC AA TT GG CC AA CC LW2 GG TT GG AA CC CC CC AA CC GG CC AA CC LW3 GG TT GG AA CC CC CC AA CT GG CC AA TT LW4 GG TT GG AA CC CC CC AA CT GG CC AA CC LW5 GG TT GG AA CC CC CC AA CT GG CC AA CC LW6 GG TT GG AA CC CC CT AA CC GG CC AA TT LW7 GG TT GG AA CC CC CC AA CT GG CC AA CC LW8 GG TT GG AA CC CC CC AA CT GG CC AA CC LW9 GG TT GG AA CC CC CC AA CT GG CC AA CC LW10 GG TT GG AA CC CC CC AA CT GG CC AA CC LW11 GG TT GG AA CC CC CC AA CT GG CC AA CC LW12 GG TT GA AA CC CC CC AA CC GG CC AA CC LW13 GG TT GG AA CC CC CC AA CT GA CC AA CC LW14 GG TT GG AA CC CC CC AA CT GG CC AA CT DD1 GG TT GG AA CC CC CT AA CC GG CC AA CC DD2 GG TT GG AA CC CC CC AA CC GG CC AA CC DD3 GG TT GG AA CC CC CC AA CC GG CC AA CC DD4 GG TT GG AA CC CC CC TT CC GG CC AA CC MS1 GG TT -- AA CA -- CT AT -- -- -- AA CC MS2 GG CC AA AA CA GG CT -- CC GG CC AA CC MS3 GG TC GA AA CA CC CT TT CC GG CC AA CC MS4 -- TT GG AA CA CC TT -- -- GG CC AA CC MS5 GG CC GG AA CA -- TT -- CC GG CC AA CC MS6 -- TT GG AA CA CC TT -- -- -- -- AA CC MS7 GG CC GG AA CA -- CT TT -- GG CC AA CC MS8 GG TT AA AA CA -- CT TT CC -- -- AA CC MS9 -- TT GG AA CA -- CT TT -- GG CC AA CC MS10 GG CC GG AA CA GG CT -- CC GG CC AA CC MS11 GG TT GG AA CA -- CT -- -- -- -- AA CC MS12 GG CC AA -- CA CC CT -- CC GG CC AA CC MS13 GG TT GG AA CA -- TT -- -- GG CC AA CC MS14 GG TT GG AA CA CC CT TT CC GG CC AA CC MS15 GG TT GG AA CA -- CT -- CC GG CC AA CC MS16 GG TT GG AA CA CC CT TT CC GG CC AA CC MS17 GG TT GA AA CA -- TT AA -- GG CC AA CC MS18 GG TC GG AA CA -- CT -- -- GG CC AA CC MS19 -- TT GG AA CA GG TT -- -- -- -- AA CC MS20 GG TT GG AA CA GG TT AT CC -- -- AA CT MS21 GG CC GG AA CA GG CT -- CC GG CC AA CC MS22 -- CC GA -- CA -- CT AA -- GG CC AA CC MS23 -- -- GA AA CA -- TT AA -- AA GG AA CC MS24 GG TC AA AA CA GG TT AA -- GG CC AA CC MS25 GG TT GG AA CA GG TT AA CC GG CC AA CC MS26 -- TC GA AA CA GG TT AT CC GG CC AA CC MS27 -- TC GG AA CA GG CT -- CC -- -- AA CC MS28 -- TC AA AA CA GG CT -- CC GG CC -- CT MM1 GG TT GG AA CC CC CC AA CC GG CC AT CC MM2 GG TT GG AA CC CC CC AA CC GG CC AT CC MM3 GG TT GG AA CC CC CC AA CC GG CC AA CC MM4 GG TT GG AA CC CC CC AA CC [[ID=3 Example 2 This embodiment further uses the SNP locus combinations obtained from Example 1 for variety identification, specifically including the following process: 1. Experimental subjects A total of 226 pigs were involved, including 14 Jinfen White Pigs, 28 Ma Shen Pigs, 14 Large White Pigs, 5 Landrace Pigs, 4 Duroc Pigs, 12 Meishan Pigs, 6 Bama Pigs, 5 Erhualian Pigs, 5 Ganzi Tibetan Pigs, 2 Hanpuxia Pigs, 6 Hetao Pigs, 3 Jinhua Pigs, 6 Laiwu Pigs, 6 Luchuan Pigs, 6 Min Pigs, 3 Neijiang Pigs, 2 Pengzhou Pigs, 5 Pietrain Pigs, 4 Gansu Tibetan Pigs, 6 Sichuan Tibetan Pigs, 29 Tibetan Pigs, 6 Yunnan Tibetan Pigs, 13 Wild Boars, 8 European Wild Boars, 3 Wujin Pigs, 6 Wuzhishan Pigs, 2 Xiang Pigs, 3 Yan'an Pigs, and 14 other pigs of unknown breeds.
[0038] 2. Experimental Methods All pigs underwent whole genome sequencing. Then, based on the 26 SNP loci screened in Example 1, the breeds of these 226 pigs were identified. A PCA diagram was constructed using plink to observe the population structure. A genotype distance matrix was calculated using plink. Then, an NJ-tree was constructed using the R package ape to observe genetic distances. Cluster heatmaps were used to observe the mutation status of the background population at the 26 SNP loci screened in Example 1.
[0039] 3. Experimental Results This invention uses 15,639,137 SNPs from step 1 (6) of Example 1 to perform principal component analysis on the aforementioned 226 pigs, and obtains the following results: The results show that Jinfen White Pigs and other white pig groups such as Large White Pigs and Landrace Pigs are mainly clustered in the upper left corner and are difficult to distinguish, while some black pig groups such as Horse-Shouldered Pigs and Meishan Pigs are clustered in the lower left corner and can be distinguished.
[0040] This invention further employs the 26 SNP loci obtained from step 3 of Example 1 to perform principal component analysis, individual genetic distance analysis, and cluster analysis on the aforementioned 226 pigs. The principal component analysis yielded the following results: The results show that all Jinfen White pigs are on the right side, while other pig breeds are clustered on the left side. Jinfen White pigs can be well distinguished from all other breeds, proving that the combination of these 26 specific SNP sites can distinguish Jinfen White pigs from 29 groups including Large White, Landrace, Duroc, Horse-faced Pigs and Meishan Pigs.
[0041] The results of individual genetic distance analysis are as follows As shown: 14 Jinfen White Pig individuals are grouped into a separate large category, which can be well distinguished from other pig breeds.
[0042] Genotype clustering heatmap as follows As shown in the figure: red indicates homozygous mutations, yellow indicates heterozygous mutations, and blue indicates no mutations. The figure shows that Jinfen White pigs exhibit homozygous mutations at these 26 specific SNP sites, while Mashan pigs exhibit partially heterozygous mutations, demonstrating a certain blood relationship between the two breeds. Other breeds show essentially no mutations. This proves that the combination of these 26 specific SNP sites can effectively distinguish Jinfen White pigs from 29 other groups, including Large White, Landrace, Duroc, Mashan, and Meishan pigs.
[0043] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A molecular marker combination, characterized in that, Including SEQ ID NO.1-26: As shown in SEQ ID NO.1, position 319 is an SNP site with a polymorphism of G / A; As shown in SEQ ID NO.2, position 335 is an SNP site with a polymorphism of C / T; As shown in SEQ ID NO.3, position 138 is an SNP site with a polymorphism of T / C; As shown in SEQ ID NO.4, position 156 is an SNP site with a polymorphism of C / T; As shown in SEQ ID NO.5, position 265 is an SNP site with a polymorphism of C / G; As shown in SEQ ID NO.6, position 392 is an SNP site with a polymorphism of A / G; As shown in SEQ ID NO.7, position 244 is an SNP site with a polymorphism of G / A; As shown in SEQ ID NO.8, position 400 is an SNP site with a polymorphism of G / A; As shown in SEQ ID NO.9, position 277 is an SNP site with a polymorphism of T / C; As shown in SEQ ID NO.10, position 381 is an SNP site with a polymorphism of C / T; As shown in SEQ ID NO.11, position 250 is an SNP site with a polymorphism of G / A; As shown in SEQ ID NO.12, position 389 is an SNP site with a polymorphism of C / A; As shown in SEQ ID NO.13, position 153 is an SNP site with a polymorphism of C / T; As shown in SEQ ID NO.14, position 389 is an SNP site with a polymorphism of G / T; As shown in SEQ ID NO.15, position 354 is an SNP site with a polymorphism of T / C; As shown in SEQ ID NO.16, position 355 is an SNP site with a polymorphism of G / A; As shown in SEQ ID NO.17, position 346 is an SNP site with a polymorphism of A / G; As shown in SEQ ID NO.18, position 370 is an SNP site with a polymorphism of C / A; As shown in SEQ ID NO.19, position 222 is an SNP site with polymorphism C / G; As shown in SEQ ID NO.20, position 285 is an SNP site with a polymorphism of C / T; As shown in SEQ ID NO.21, position 327 is an SNP site with a polymorphism of A / T; As shown in SEQ ID NO.22, position 152 is an SNP site with a polymorphism of C / T; As shown in SEQ ID NO.23, position 399 is an SNP site with a polymorphism of G / A; As shown in SEQ ID NO.24, position 400 is an SNP site with a polymorphism of C / G; As shown in SEQ ID NO.25, position 104 is an SNP site with a polymorphism of A / T; As shown in SEQ ID NO.26, position 353 is an SNP site with a polymorphism of C / T.
2. A primer combination for identifying SNP sites of the Jinfen White Pig breed, characterized in that, include: 1.2F:ATAGAATCTGCATATTGAGACCCA, 1.2R:CGTAAAAGTGCAATCATGAAGCC; 3F:TCAAATGCCCTCCACTACTTCC, 3R:GGGGTTTATTGATTGTTGGGGAC; 4F:AACCTGCGCTGAACTTCTGA, 4R:AACAGACCTCCCCTTCCTGT; 5F:AATCCTCCCCATCCACTAAAAGC, 5R:TGCAGTTTGGGACTCTTCCTG; 6F:TCATGTCATTTCCATTCAATAGGC, 6R:AATTTCCAGAGCCAGCCCTC; 7F:GACACTGTGGCACTTCCTGA, 7R:CCCCTCCTGAAAGGCTGATG; 8F:GCTGTCTTGCTTCCAATTGTCT, 8R:GGCACGACATATACGCTGCT; 9F:GCCAGCAAAGAGGACAAAGC, 9R:TGGGGCACTTGAAAAGGGTT; 10F:GTCATAGGTGTCAGGCCTTTCT, 10R:TTTCGAGCACTCTCTCCATCATT; 11F:TTTGCTGCGTTTTAGCTCCAC, 11R:TGAGTAGCTGTTCGGTGCAA; 12F:CCTCCTCAAAGCCTGCTTCA, 12R:GCTTTGCTTCCCCAAACCTG; 13F:TGATGCAGGGAGAGGAGGAA, 13R:ATGCATTCAGTTGTACCGGC; 14F:GGAGGCAGTAGATACCCAGATAC, 14R:CCCTCATTTCCAGGCTAATGC; 15F:TTAAGGCCGCCAACTTCTACC, 15R:TAAGAAGCGTGGGCGACAA; 16F:ATCGTCTCCCGCGGATTAGG, 16R:ACAGAGTCAGTGTATCCATCACC; 17F:CTTCTGGGGACAGGTCAGGA, 17R:CATGAGCAGGAGATCAGGGC; 18F:TCCACCTCGAGGGAAACAGA, 18R:CCCTTGTACTGGTCTGCGAG; 19F:ATCAGCCCGCAATAACCACA, 19R:GGACGGAACGAAGACAAGGT; 20F:CATAGCCAGACAGCAGGACC, 20R:TCAAAGGAGGTGAGGGGTGA; 21F:GTCCTGGCAAGACTCTCAGG, 21R: AGACCAACCGCCTTCTCTTG; 22F: TTCTGCTGCCCTTACCCTTG, 22R: AGGGAGACATGAGCAAAGCC; 23.24F:AGTTGGATTCTTCAGCACTAGGAA, 23.24R: TACCATTGGCACAGTTGCAC; 25F:TAGAGCGAAAATTTCAGAGCCAC, 25R:TTTCCCTTCAACTCTAGACCACTTT; 26F: GAGCCCCCGAATGTCCTCTA 26R: GTGATGGGGTTACTTGACCG; The site is: The SNP site located at position 319 of the sequence shown in SEQ ID NO.1 has a polymorphism of G / A; The SNP site located at position 335 of the sequence shown in SEQ ID NO.2 has a polymorphism of C / T; The SNP site located at position 138 of the sequence shown in SEQ ID NO.3 has a polymorphism of T / C; The SNP site located at position 156 of the sequence shown in SEQ ID NO.4 has a polymorphism of C / T; The SNP site located at position 265 of the sequence shown in SEQ ID NO.5 has a polymorphism of C / G; The SNP site located at position 392 of the sequence shown in SEQ ID NO.6 has a polymorphism of A / G; The SNP site located at position 244 of the sequence shown in SEQ ID NO.7 has a polymorphism of G / A; The SNP site located at position 400 of the sequence shown in SEQ ID NO.8 has a polymorphism of G / A; The SNP site located at position 277 of the sequence shown in SEQ ID NO.9 has a polymorphism of T / C; The SNP site located at position 381 of the sequence shown in SEQ ID NO.10 has a polymorphism of C / T; The SNP site located at position 250 of the sequence shown in SEQ ID NO.11 has a polymorphism of G / A; The SNP site located at position 389 of the sequence shown in SEQ ID NO.12 has a polymorphism of C / A; The SNP site located at position 153 of the sequence shown in SEQ ID NO.13 has a polymorphism of C / T; The SNP site located at position 389 of the sequence shown in SEQ ID NO.14 has a polymorphism of G / T; The SNP site located at position 354 of the sequence shown in SEQ ID NO.15 has a polymorphism of T / C; The SNP site located at position 355 of the sequence shown in SEQ ID NO.16 has a polymorphism of G / A; The SNP site located at position 346 of the sequence shown in SEQ ID NO.17 has a polymorphism of A / G; The SNP site located at position 370 of the sequence shown in SEQ ID NO.18 has a polymorphism of C / A; The SNP site located at position 222 of the sequence shown in SEQ ID NO.19 has a polymorphism of C / G. The SNP site located at position 285 of the sequence shown in SEQ ID NO.20 has a polymorphism of C / T; The SNP site located at position 327 of the sequence shown in SEQ ID NO.21 has a polymorphism of A / T; The SNP site located at position 152 of the sequence shown in SEQ ID NO.22 has a polymorphism of C / T; The SNP site located at position 399 of the sequence shown in SEQ ID NO.23 has a polymorphism of G / A; The SNP site located at position 400 of the sequence shown in SEQ ID NO.24 has a polymorphism of C / G; The SNP site located at position 104 of the sequence shown in SEQ ID NO.25 has a polymorphism of A / T; The SNP site located at position 353 of the sequence shown in SEQ ID NO.26 has a polymorphism of C / T.
3. A gene chip for identifying the Jinfen White Pig breed, characterized in that, The gene chip includes nucleic acid probes for detecting the molecular marker combination of claim 1.
4. A reagent kit, characterized in that, The kit comprises the primer combination of claim 2 or the gene chip of claim 3.
5. The application of the primer combination of claim 2, the gene chip of claim 3, or the kit of claim 4 in the breed identification of Jinfen White Pig.
6. A method for identifying the breed of Jinfen White Pig, characterized in that, include: Genomic DNA is extracted from the pig to be tested, and the genotype information of the molecular marker combination as described in claim 1 is detected. Based on the detection results, the breed of the pig to be tested is identified.
7. The variety identification method according to claim 6, characterized in that, The breed of the pig to be tested is identified based on the test results. If the test results are as follows, the pig to be tested is determined to be a Jinfen White Pig: The SNP site located at position 319 of the sequence shown in SEQ ID NO.1 was detected as AA / AG; The SNP site located at position 335 of the sequence shown in SEQ ID NO.2 was detected as TT / TC; The SNP site located at position 138 of the sequence shown in SEQ ID NO.3 was detected as CC / CT. The SNP site located at position 156 of the sequence shown in SEQ ID NO.4 was detected as TT / TC; The SNP site located at position 265 of the sequence shown in SEQ ID NO.5 was detected as GG / GC; The SNP site located at position 392 of the sequence shown in SEQ ID NO.6 was detected as GG / GA; The SNP site located at position 244 of the sequence shown in SEQ ID NO.7 was detected as AA / AG; The SNP site located at position 400 of the sequence shown in SEQ ID NO.8 was detected as AA / AG; The SNP site located at position 277 of the sequence shown in SEQ ID NO.9 was detected as CC / CT. The SNP site located at position 381 of the sequence shown in SEQ ID NO.10 was detected as TT / TC; The SNP site located at position 250 of the sequence shown in SEQ ID NO.11 was detected as AA / AG; The SNP site located at position 389 of the sequence shown in SEQ ID NO.12 was detected as AA / AC. The SNP site located at position 153 of the sequence shown in SEQ ID NO.13 was detected as TT / TC; The SNP site located at position 389 of the sequence shown in SEQ ID NO.14 was detected as TT / TG; The SNP site located at position 354 of the sequence shown in SEQ ID NO.15 was detected as CC / CT. The SNP site located at position 355 of the sequence shown in SEQ ID NO.16 was detected as AA / AG; The SNP site located at position 346 of the sequence shown in SEQ ID NO.17 was detected as GG / GA; The SNP site located at position 370 of the sequence shown in SEQ ID NO.18 was detected as AA / AC. The SNP site located at position 222 of the sequence shown in SEQ ID NO.19 was detected as GG / GC; The SNP site located at position 285 of the sequence shown in SEQ ID NO.20 was detected as TT / TC; The SNP site located at position 327 of the sequence shown in SEQ ID NO.21 was detected as TT / TA; The SNP site located at position 152 of the sequence shown in SEQ ID NO.22 was detected as TT / TC; The SNP site located at position 399 of the sequence shown in SEQ ID NO.23 was detected as AA / AG; The SNP site located at position 400 of the sequence shown in SEQ ID NO.24 was detected as GG / GC; The SNP site located at position 104 of the sequence shown in SEQ ID NO.25 was detected as TT / TA. The SNP site located at position 353 of the sequence shown in SEQ ID NO.26 was detected as TT / TC.
Citation Information
Patent Citations
Preparation method and genetic typing method of 15K liquid phase chip for Min pig breeding
CN117363750A