Specific molecular ID card for identifying Huainan pig breeds and its application

By screening the collection of SNP sites with high allele frequency as specific molecular ID cards, combining gene chip technology and multiple analytical methods, the false negative and false positive problems in genome-wide association analysis were solved, and efficient and accurate identification of Huainan pig breeds was achieved.

CN115011700BActive Publication Date: 2025-05-23HENAN AGRICULTURAL UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210206533.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-04
Publication Date
2025-05-23
Estimated Expiration
2042-03-04

AI Technical Summary

Technical Problem

The prior art is prone to false negative and false positive SNP sites in genome-wide association analysis, affecting the identification accuracy of Huainan pig breeds.

Method used

By screening the collection of SNP sites with high allele frequency as specific molecular ID cards, and combining gene chip technology and a variety of analysis methods, the identification efficiency and accuracy are improved.

Benefits of technology

The efficient and accurate identification of Huainan pig breeds has been achieved, revealing the landmark genetic differences between breeds, and helping to protect and innovatively utilize local pig genetic resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115011700B_ABST
    Figure CN115011700B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of pig breed identification, and relates to the identification of Huainan pig germplasm resources, and in particular refers to a specific molecular ID card and application for identifying Huainan pig breeds. This application has discovered a specific molecular ID card for identifying Huainan pig breeds through molecular biology technology, which can reveal the iconic genetic differences between breeds. The use of biotechnology to explore the characteristics of precious genetic resources of local pigs is conducive to promoting the protection and innovative use of local pig genetic resources in Henan Province, and promoting the high-quality development of the seed industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of pig breed identification, relates to the identification of Huainan pig germplasm resources, and particularly refers to a specific molecular identity card and application for identifying Huainan pig breeds. Background Art

[0002] The formation of excellent characteristics of Henan local pigs is closely related to the geographical environment and human selection. Therefore, breed-specific molecular markers screened at the genome level by molecular biology technology may reveal the iconic genetic differences between breeds. In recent years, with the development of bioinformatics technology, it has become possible to form a molecular identity card of local pigs based on genomic information.

[0003] When applying whole-genome association analysis, false negative and false positive SNP sites often appear in the results, which greatly affects the final judgment. In order to avoid situations that affect the effectiveness of whole-genome association analysis, genome control methods, principal component analysis methods, and multiple hypothesis testing methods can be used. At the same time, a combination of multiple analysis methods can also be used to reduce the error of the results. For example, the combination strategy of whole-genome association analysis and selection signals can improve the efficiency and accuracy of gene positioning to achieve the effect of mutual verification.

[0004] Analyzing and differentiating the germplasm resources of local pig breeds in Henan Province at the molecular level and exploring a molecular identity card that can efficiently and accurately identify Huainan pigs will help to further understand breed-specific variations and genetic structure characteristics, formulate specific protection plans and improve breed identification work. Summary of the invention

[0005] To achieve the above-mentioned purpose, the present invention proposes a specific molecular identification card and application for identifying Huainan pig breeds.

[0006] The technical solution of the present invention is achieved in this way:

[0007] A specific molecular ID card for identifying Huainan pig breeds, wherein the specific molecular ID card is a collection of SNP sites with higher allele frequencies.

[0008] Furthermore, the specific molecular ID card is located on the pig reference genome Ensembl Sscrofa 11.1 version.

[0009] Further, the set of SNP sites includes CNC10010500, CNC10010865, CNC10010911, CNC10013894, CNC10014112, CNC10014662, CNC10014900, CNC10014919, CNC10015519, CNC10020827, CNC10021138, CNC10021522, CNC10021729, CNC10021730, CNC10023047, CNC10030056, CNC10030704, CNC10031464, CNC10040602, and CNC10030603. 41374, CNC10042724, CNC10050348, CNC10060818, CNC10060867, CNC10061010, CNC10061259, CNC10061391, CNC10062505, CNC10062590, CNC10062745 , CNC10062777, CNC10062926, CNC10070739, CNC10070797, CNC10070799, CNC10071021, CNC10071232, CNC10071354, CNC10071586, CNC10071639, CNC1 0072325, CNC10072326 / CNC10080556, CNC10080573, CNC10080681, CNC10081051, CNC10081763, CNC10081844, CNC10081880, CNC10090326, CNC100906 34. CNC10091624, CNC10091676, CNC10092648, CNC10100053, CNC10100634, CNC10110699, CNC10110782, CNC10111310, CNC10111520, CNC10130866, CN CNC10131089, CNC10131950, CNC10133684, CNC10133697, CNC10133908, CNC10133909, CNC10134079, CNC10134173, CNC10142116, CNC10142634, CNC10150394, CNC10150460, CNC10150774, CNC10152162, CNC10152163, CNC10152316, CNC10152925, CNC10152958, CNC10170796, CNC10170918, and CNC10171012.

[0010] Further, the mutation type of the CNC10010500 site is G / A, the mutation type of the CNC10010865 site is G / A, the mutation type of the CNC10010911 site is C / T, the mutation type of the CNC10013894 site is G / A, the mutation type of the CNC10014112 site is G / A, the mutation type of the CNC10014662 site is T / A, the mutation type of the CNC10014900 site is C / T, the mutation type of the CNC10014919 site is A / G, the mutation type of the CNC10015519 site is A / C, the mutation type of the CNC10020827 site is C / T, the mutation type of the CNC10021138 site is G / T, C The mutation type of NC10021522 is C / G, the mutation type of CNC10021729 is T / C, the mutation type of CNC10021730 is T / C, the mutation type of CNC10023047 is A / G, the mutation type of CNC10030056 is G / C, the mutation type of CNC10030704 is T / G, the mutation type of CNC10031464 is T / C, the mutation type of CNC10040602 is A / G, the mutation type of CNC10041374 is G / A, the mutation type of CNC10042724 is G / A, the mutation type of CNC10050348 is G / A, and the mutation type of CNC10060 The mutation type of CNC10060818 is T / C, the mutation type of CNC10060867 is C / T, the mutation type of CNC10061010 is C / G, the mutation type of CNC10061259 is C / G, the mutation type of CNC10061391 is C / T, the mutation type of CNC10062505 is T / G, the mutation type of CNC10062590 is T / A, the mutation type of CNC10062745 is T / G, the mutation type of CNC10062777 is T / C, the mutation type of CNC10062926 is A / G, the mutation type of CNC10070739 is T / C, and the mutation type of CNC10070797 is The mutation type of the CNC10070799 site is G / A, the mutation type of the CNC10071021 site is T / C, the mutation type of the CNC10071232 site is T / C, the mutation type of the CNC10071354 site is T / C, the mutation type of the CNC10071586 site is G / C, the mutation type of the CNC10071639 site is T / C, the mutation type of the CNC10072325 site is T / C, the mutation type of the CNC10072326 site is A / G, the mutation type of the CNC10080556 site is T / C, the mutation type of the CNC10080573 site is T / C, the mutation type of the CNC10080681 site is T / C,The mutation type of CNC10081051 is G / C, the mutation type of CNC10081763 is C / A, the mutation type of CNC10081844 is T / C, the mutation type of CNC10081880 is G / C, the mutation type of CNC10090326 is G / C, the mutation type of CNC10090634 is G / A, the mutation type of CNC10091624 is C / T, the mutation type of CNC10091676 is T / C, the mutation type of CNC10092648 is G / C, and the mutation type of CNC10091664 is G / C. The mutation type of the 100053 site is C / T, the mutation type of the CNC10100634 site is G / A, the mutation type of the CNC10110699 site is A / T, the mutation type of the CNC10110782 site is G / C, the mutation type of the CNC10111310 site is A / G, the mutation type of the CNC10111520 site is C / A, the mutation type of the CNC10130866 site is C / G, the mutation type of the CNC10131089 site is C / T, the mutation type of the CNC10131950 site is T / G, and the mutation type of the CNC10133684 The mutation type of the site is C / T, the mutation type of the site CNC10133697 is C / T, the mutation type of the site CNC10133908 is T / C, the mutation type of the site CNC10133909 is T / G, the mutation type of the site CNC10134079 is T / C, the mutation type of the site CNC10134173 is G / C, the mutation type of the site CNC10142116 is G / T, the mutation type of the site CNC10142634 is C / A, the mutation type of the site CNC10150394 is G / A, and the mutation type of the site CNC10150460 is The mutation type of the CNC10150774 site is T / C, the mutation type of the CNC10152162 site is C / T, the mutation type of the CNC10152163 site is T / G, the mutation type of the CNC10152316 site is G / A, the mutation type of the CNC10152925 site is A / C, the mutation type of the CNC10152958 site is C / T, the mutation type of the CNC10170796 site is A / T, the mutation type of the CNC10170918 site is G / C, and the mutation type of the CNC10171012 site is C / T.

[0011] A gene chip used to identify the above-mentioned specific molecular ID card.

[0012] The application of the above gene chip in identifying Huainan pig breeds.

[0013] The steps for the above application are:

[0014] (1) Collect tissue samples from the pigs to be tested and extract genomic DNA;

[0015] (2) Using a gene chip to perform SNP typing on the genomic DNA of step (1) to obtain genotype data of the pig to be tested;

[0016] (3) The genotype data of the pigs to be tested and the genotype data of the Huainan pigs were merged using PLINK software, the SNP loci genotypes on the specific molecular ID card were extracted, and then principal component analysis was performed.

[0017] Preferably, in step (1), the light absorption ratio of the genomic DNA at A260 / 280 is between 1.8 and 2.0, and the concentration is ≥ 50 ng / μl.

[0018] Preferably, when the result of principal component analysis of SNP typing in step (3) is close to the genetic distance of the Huainan pig population and clusters into one cluster, it is the Huainan pig.

[0019] The present invention has the following beneficial effects:

[0020] 1. This application has discovered a specific molecular ID card for identifying Huainan pig breeds through molecular biology technology, which can reveal the iconic genetic differences between breeds. The use of biotechnology to explore the characteristics of precious genetic resources of local pigs will help promote the protection and innovative use of local pig genetic resources in Henan Province and promote the high-quality development of the seed industry.

[0021] 2. Genotype filling is an important part of genome-wide association analysis. Its purpose is to predict SNPs that have not been typed in the research samples, increase the number of SNPs that can be used to detect associations, thereby improving the detection capability of GWAS, and the combined strategy of genome-wide association analysis and selection signal analysis, which further improves the efficiency and accuracy of identifying molecular markers to achieve the role of mutual verification. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0023] Figure 1 Manhattan plot (left) and QQ plot (right) of the GWAS analysis results for Huainan pig.

[0024] Figure 2 Manhattan plot of the selection signal analysis results for Huainan pigs.

[0025] Figure 3This is the principal component analysis validation diagram for the Huainan pig breed-specific loci. DETAILED DESCRIPTION

[0026] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. Example

[0027] The acquisition of the specific molecular ID of the Huainan pig breed includes the following steps:

[0028] (1) Ear sample collection

[0029] The experimental population consisted of 1117 pigs of 10 pig breeds, including 7 Chinese pig breeds: Nanyang Black Pig (n=10, NY), Huainan Pig (n=10, HN), Yunong Black Pig (n=1036, YN), Queshan Black Pig (n=10, QS), Laiwu Pig (n=10, LWH), Erhualian Pig (n=10, EHL), Min Pig (n=6, MIN); and 3 Western commercial pig breeds: Duroc Pig (n=10, DU), Large White Pig (n=10, LW), Landrace Pig (n=5, LR). The ears of the pigs were cleaned with 75% alcohol, and a small amount of ear tissue was cut with ear sample forceps, placed in a 2ml centrifuge tube filled with 75% alcohol, and stored in a -20℃ refrigerator.

[0030] (2) Total DNA extraction, quality testing and genotyping

[0031] ①. Use animal tissue genomic DNA extraction kit to extract total DNA, use DYY-6C electrophoresis instrument, use 1% agarose gel electrophoresis detection; use Nanodrop-2000 UV spectrophotometer to detect DNA concentration, retain the light absorption ratio (A260 / 280) between 1.8 and 2.0, and the concentration of genomic DNA samples ≥ 50 ng / μl. Use IlluminaPorcine SNP50 BeadChip (Beijing Compson Biotechnology Co., Ltd., Zhongxin No. 1) for whole genome chip typing. The specific operation is as follows:

[0032] a. Use Tn5 transposase to establish a gene library for the sample pigs and perform a 50 K gene chip scan.

[0033] b. Use Beagle to fill in genotypes for the 50 K chip and whole genome resequencing results in step a.

[0034] c. Perform genome-wide association analysis and selection signal analysis on the genotype filling data of all individuals obtained through step ③.

[0035] d. For the significant loci obtained in step c, calculate the allele frequencies among breeds, retain the SNP loci with higher allele frequencies of Nanyang black pigs, and use the collection of SNP loci as the breed-specific molecular ID card of Nanyang black pigs.

[0036] (3) Genotype data filling and quality control

[0037] A total of 1,117 heads and 51,315 SNPs were obtained by chip sequencing. The chip data were quality controlled using PLINK software. The genotype data were filtered using the following parameters: individual genotype detection rate (--mind) > 90%, marker genotype detection rate (--geno) > 95%, minimum allele frequency (--maf) > 1%, minimum Hardy-Weinberg equilibrium (--hwe) of 10-6, and located on autosomes. The missing genotypes were filled in using the Hidden Markov Model (HMM) algorithm in BEAGLE software.

[0038] (4) Screening of SNP-specific loci by genome-wide association analysis

[0039] GEMMA software was used for genome-wide association analysis. The experimental group consisted of 10 Huainan pigs (case) and the control group consisted of the other 9 breeds (control). The Manhattan plot of Huainan pigs is shown in Figure 1 See the left and QQ diagrams Figure 1 Right, there are two threshold lines in the Manhattan plot, where the solid line threshold is 0.05 / N (N is the number of chip sites used), and the sites above the solid line are at the genome-wide significant level; the dotted line threshold is 1 / N, and the sites above the dotted line are at the chromosome significant level. The closer the λ value in the QQ plot is to 1, the more reliable the results of the genome-wide association analysis are; the Banferonni correction method is used to identify SNPs that are significantly associated with the variety, and the set of significant SNPs is group A.

[0040] (5) Select signal analysis to screen SNP-specific sites

[0041] The genetic differentiation index (Fst) was calculated using VCFtools software, using the sliding window mean calculation method. The results are shown in Figure 2 , Figure 2The threshold line in is the top 1% of the largest Fst values ​​after sorting, and the sites above this threshold line are significant sites (marked in red). The specific parameters are as follows: the size of the sliding window (--fst-window-size) is 100,000 bp, and the step length of the sliding window (--fst-window-step) is 40,000 bp. The windows are sorted from large to small by Fst value, and the top 1% of windows are defined as significant windows. Then, PLINK software is used to extract SNPs in the significant windows. The set of significant SNPs is group B.

[0042] (6) Allele frequency screening of SNP-specific sites

[0043] PLINK software was used to merge the significant SNPs in group A and group B, and the allele frequency of each SNP in the 10 breeds was calculated. The SNP set with a higher allele frequency in the experimental group than in the other 9 breeds was selected as the specific molecular identity card of the Huainan pig breed.

[0044] Table 1 Specific molecular marker set for Huainan pig breed

[0045]

[0046] (7) PLINK software was used to extract the above 82 SNPs from 10 varieties and perform principal component analysis. The results are as follows: Figure 3 As shown by Figure 3 It can be seen that the 82 SNP sites can be used to identify the Huainan pig breed, and the SNP sites are located in the genome version Ensembl Sscrofa 11.1.

[0047] Application Examples

[0048] A method for identifying pig germplasm to be tested, specifically comprising the following steps:

[0049] 1. Extract ear tissue samples from the pigs to be tested, extract genomic DNA from the tissue samples, and send the genomic DNA that has passed the quality and concentration tests to Beijing Compson Biotechnology Co., Ltd. (the following genotyping operation is performed by this company) for SNP typing on the "Axiom" chip for local pigs. The experimental principle of SNP typing on the chip is based on the ligation reaction, in which two probes play a role. The first is the capture probe on the chip, which is 30 bp in length and plays the role of fixing the target DNA fragment to the surface of the chip. The second is the colorimetric probe, which is responsible for coloring the SNP chip (red and green fluorescence). The experiment is carried out in two rounds of hybridization. In the first round of hybridization, the target DNA is hybridized with the chip, and the capture probe will capture the matching target DNA fragment; the colorimetric probe hybridizes to the DNA fragment in the second round of hybridization. Then, using the recognition effect of the ligase, only the colorimetric probe complementary to the target DNA fragment will be connected to the capture probe. Through fluorescent labeling staining, SNP typing is performed under laser scanning to obtain the genotype data of the pigs to be tested.

[0050] 2. Use PLINK software to merge the genotype data of the pig to be tested and the Huainan pig, and extract the above 82 loci in the data. Then perform principal component analysis and R language result visualization. When the genetic distance between the individual pig to be tested and the Huainan pig population is close and they are clustered into one cluster, it can be determined that the pig to be tested is a Huainan pig.

[0051] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for identifying Huainan pig breeds, It is characterized in that The method is achieved by identifying the following set of SNP sites; the SNP sites are located in the pig reference genome EnsemblSscrofa 11.1 version; the set of SNP sites is:

2. Gene chip for identifying a collection of SNP sites, It is characterized in that The SNP sites are located in the pig reference genome EnsemblSscrofa 11.1 version; the set of SNP sites is:

3. Use of the gene chip according to claim 2 in identifying Huainan pig varieties.

4. The use according to claim 3, It is characterized in that The steps are: (1) Collecting tissue samples from the pigs to be tested and extracting genomic DNA; (2) performing SNP typing on the genomic DNA of step (1) using a gene chip to obtain genotype data of the pig to be tested; (3) The genotype data of the pigs to be tested and the genotype data of the Huainan pigs were merged using PLINK software, the SNP loci genotypes on the specific molecular ID card were extracted, and then principal component analysis was performed.

5. The use according to claim 4, Features: In the step (1), the genomic DNA has an A260 / 280 light absorption ratio between 1.8 and 2.0, and a concentration of ≥50 ng / μl.

6. The use according to claim 5, Features: When the result of the principal component analysis of the SNP typing in the step (3) is close to the genetic distance of the Huainan pig population and clusters into one cluster, it is the Huainan pig.