A molecular identity card for identifying tianhua mutton sheep breed and application thereof
Patent Information
- Application Number
- CN202510296212.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-03-13
AI Technical Summary
目前,尚无针对天华肉羊品种鉴别的SNP分子标记技术
本发明通过全基因组重测序获得55只天华肉羊、11只南非肉用美利奴羊和11只甘肃高山细毛羊的基因组原始数据,通过与已公开的其他7个与天华肉羊外貌相近或甘肃省其他绵羊遗传资源的119个个体的全基因组基因型数据相比较分析,筛选得到天华肉羊的36个特异性SNP组合,作为天华肉羊特异分子身份证。本发明的天华肉羊特异分子身份证具有明显的天华肉羊品种特异性,对于所述的36个SNP,构建遗传进化树可以将天华肉羊与非天华肉羊群体明显区分。也可使用机器学习分类器模型对测试集进行品种预测,所用的支持向量机、随机森林、朴素贝叶斯和逻辑回归模型中准确率、精确率及F1分数均达到了95%以上。说明此分子身份证能够利用较少的基因型信息简单明了区分甘肃省及临近地区的天华肉羊群体和非天华肉羊群体,为天华肉羊的种质资源鉴定和保护提供科学有力的依据,为天华肉羊的进一步改良和在绵羊养殖业的推广奠定基础。
Smart Images

Figure CN119979722B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of sheep breed identification technology, and relates to a molecular identification card for identifying the Tianhua meat sheep breed. This invention also relates to the application of the molecular identification card in the identification of Tianhua meat sheep germplasm resources. Background Technology
[0002] Tianhua mutton sheep is a new breed of fine-wool sheep adapted to high-altitude and cold climates, developed over 15 years using imported South African Merino sheep as the sire and Gansu alpine fine-wool sheep as the dam. This breed boasts excellent meat and wool performance, high fertility, and strong adaptability, meeting the current market demand for dual-purpose (meat and wool) sheep, while also possessing the dam's adaptability to the local high-altitude and cold environment. The development of this breed addresses the shortcomings of existing breeds in high-altitude and surrounding areas, such as slow growth, low fertility, long fattening cycles, poor meat performance, and declining wool performance. It effectively ensures that valuable fine-wool sheep resources are not lost due to fluctuations in the wool market and the difference in the comparative benefits of wool and mutton, meeting the significant demand for preserving wool and increasing meat production in high-altitude and cold regions, especially fine-wool sheep producing areas, and providing a good source of breeding stock for the upgrading and replacement of sheep breeds in these areas.
[0003] Currently, sheep breed identification relies heavily on physical appearance or pedigree analysis. These methods are time-consuming, labor-intensive, and susceptible to subjective or environmental factors, making accurate identification difficult. However, with the rapid development of genomics, molecular marker technology has made significant progress in genetic breeding. New-generation molecular markers, such as Specialized Numerical NPs (SNPs), are characterized by their large number, wide distribution, genetic stability, high specificity, and ease of standardization and automation. By selecting signals and combining them with machine learning classifiers, breed-specific SNPs can be identified, enabling accurate, rapid, and efficient sheep breed identification. Currently, there is no SNP molecular marker technology specifically for identifying the Tianhua meat sheep breed. Summary of the Invention
[0004] The purpose of this invention is to provide a molecular identification card for identifying the Tianhua meat sheep breed, which is used for the identification of Tianhua meat sheep germplasm resources. This is beneficial for the further promotion and breeding of the breed, and prevents the loss of genetic resources or economic losses to herders due to breed imitation or hybrid substitution, thereby promoting the long-term development of Tianhua meat sheep in the sheep breeding industry.
[0005] To achieve its purpose, the present invention adopts the following technical solution: This invention provides a molecular identity card for identifying the Tianhua meat sheep breed. The molecular identity card consists of 36 characteristic SNP sites located in the sheep reference genome ARS-UI_Ramb_v2.0. These SNP sites are a set of Tianhua meat sheep-specific SNP sites selected by selection signal and maximum correlation minimum redundancy (mRMR). The specific information is shown in Table 1 below.
[0006] Table 1. The aforementioned molecular identification card for Tianhua mutton sheep can be used for the identification of Tianhua mutton sheep germplasm resources. The application method is as follows: (1) Whole genome resequencing was performed on the sheep samples of unknown breeds to be tested to obtain the raw genome data; (2) The raw data obtained in step (1) are compared with the sheep reference genome ARS-UI_Ramb_v2.0 and genotyped to obtain the genotype data of the sheep to be tested; (3) The genotype data from step (2) is filtered and then self-filled using Beagle software. Subsequently, the SNP sites corresponding to the molecular identity cards of Tianhua meat sheep are extracted. (4) Merge the genotype data of the SNP sites on the molecular ID card of the Tianhua meat sheep population, and construct a genetic evolution tree to determine whether it is a Tianhua meat sheep genetic resource; or convert its genotype format to 012 format and input it into a trained Python machine learning classifier model, and determine whether it is a Tianhua meat sheep genetic resource based on its predict function.
[0007] The beneficial effects of this invention are as follows: This invention obtained the raw genome data of 55 Tianhua mutton sheep, 11 South African merino sheep, and 11 Gansu alpine fine-wool sheep through whole-genome resequencing. By comparing and analyzing the whole-genome genotype data of 119 individuals from 7 other publicly available sheep genetic resources similar in appearance to Tianhua mutton sheep or other sheep from Gansu Province, 36 specific SNP combinations for Tianhua mutton sheep were screened and identified as their unique molecular identification. This unique molecular identification of Tianhua mutton sheep exhibits significant breed specificity. For the 36 SNPs, a phylogenetic tree can be constructed to clearly distinguish Tianhua mutton sheep from non-Tianhua mutton sheep populations. Machine learning classifier models can also be used to predict breed characteristics on the test set. The accuracy, precision, and F1 score of the support vector machine, random forest, Naive Bayes, and logistic regression models used all exceeded 95%. This molecular identification method can clearly distinguish between Tianhua sheep populations and non-Tianhua sheep populations in Gansu Province and neighboring areas using relatively little genotypic information. It provides a strong scientific basis for the identification and protection of Tianhua sheep germplasm resources and lays the foundation for further improvement of Tianhua sheep and its promotion in the sheep farming industry. Attached Figure Description
[0008] Figure 1 Phylogenetic tree diagrams for 10 sheep breeds; Figure 2 To improve the accuracy of SNPs acquired from selected signals in machine learning classifiers; Figure 3 The accuracy of machine learning classifiers with 10-100 different numbers of SNPs selected for mRMR; Figure 4 The accuracy of machine learning classifiers with 30-40 different numbers of SNPs selected for mRMR; Figure 5 A phylogenetic tree constructed based on the 36 molecular identity SNPs of Tianhua meat sheep; Figure 6 This is a diagram showing the genetic evolution analysis results of an application example of the present invention. Detailed Implementation
[0009] To better illustrate the purpose, technical solution, and advantages of this invention, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.
[0010] The Tianhua mutton sheep described in this invention comes from the Tianhua mutton sheep core breeding farm of Gansu Lantian Tonghe Agricultural Co., Ltd. in Tianzhu Tibetan Autonomous County, Wuwei City, Gansu Province.
[0011] Example 1 A method for screening specific molecular identifiers for Tianhua meat sheep germplasm resources, comprising the following steps: I. Sample Data Acquisition Blood samples were collected from 55 Tianhua mutton sheep, 11 South African Merino sheep, and 11 Gansu alpine fine-wool sheep. Genomic DNA was extracted using the phenol-chloroform method, and the DNA samples were sent to Beijing Berry Genomics Co., Ltd. to construct DNA libraries. Sequencing was performed using the Illumina PE150 strategy, and fastq.gz data for each sample were obtained. In addition, fastq.gz data of 119 sheep breeds in Gansu Province, including Australian Merino (n=7), Chinese Merino (n=37), German Merino (n=17), Lanzhou fat-tailed sheep (n=4), Minxian black fur sheep (n=9), Oula sheep (n=10), and Tan sheep (n=35), as well as breeds similar in appearance to Tianhua mutton sheep, were collected from the NCBI database.
[0012] II. Data Format Conversion BaseNumber DNA sequencing data analysis software was used for quality control, alignment, and SNP detection. The reference genome version used for alignment was the sheep reference genome ARS-UI_Ramb_v2.0, resulting in a total of 196 vcf.gz format files.
[0013] III. Site Filtering The vcf.gz file was filtered using bcftools and plink software to retain autosomal and biallelic SNPs, while filtering out sites with a minor allele frequency (MAF) less than 0.05 and removing all missing sites. This ensured the sites could be used as features input to the subsequent machine learning classifier. Specifically, plink software was used to remove linked sites with the following parameters: window size of 50 SNPs, step size of 10 SNPs, and a threshold for low allele frequency (LD) of 0.2.
[0014] IV. Population Genetic Analysis and Selection of Characteristic SNPs Genetic phylogenetic trees were constructed for all samples using the filtered loci. A genetic distance matrix was built using PLINK software, followed by phylogenetic tree construction using MEGA software and enhancement using the iTOL website. The phylogenetic tree results are shown below. Figure 1As shown, the 10 selected breeds exhibit good clustering and significant differentiation. Feature SNP selection employed a strategy combining two selection signals. First, the Pi value (nucleotide diversity) of the 10 breeds was calculated using vcftools software, with a window size of 10000. Then, a self-written script was used to calculate the logarithmic value of the Pi ratio between Tianhua sheep and the other 9 breeds, i.e., ln-Piratio. The top 2% of the ln-Piratio values between each breed pair were obtained (i.e., the top 1% and bottom 1% windows), and SNPs were extracted from these windows using bcftools. Second, vcftools software was used to calculate the genetic differentiation index (Fst) for each locus between Tianhua sheep and the other breeds within the aforementioned windows, and the top 1% of SNP loci for each breed pair were extracted. The intersection of these loci was then used as the characteristic SNP loci for Tianhua sheep.
[0015] V. Testing for Feature SNPs Extract the SNP loci obtained in the previous step and convert the SNP locus information into 012 encoding format using plink software, where 0 represents homozygous at the reference genome locus, 1 represents heterozygous, and 2 represents homozygous at the mutation locus. Use each SNP locus information as a feature and train a classifier model using support vector machine, random forest, Naive Bayes, and logistic regression models from the sklearn library in Python. Label all samples as either Tianhua sheep or non-Tianhua sheep, with 0 representing Tianhua sheep and 1 representing non-Tianhua sheep. Use 80% of the samples as the training set to train the classification model, and the remaining 20% as the test set to test the model's accuracy. Model training employs a parameter grid search and 5-fold cross-validation to obtain the optimal model. The training set validation results are as follows. Figure 2 As shown: This illustrates that the 164 SNP loci obtained in step four can effectively distinguish between Tianhua mutton sheep and non-Tianhua mutton sheep.
[0016] VI. Simplification of Characteristic SNPs The SNPs obtained in the previous step are further filtered using the Maximum Relevance Minimum Redundancy (mRMR) algorithm. This algorithm identifies a set of features in the original feature set that has the highest correlation with the final output but the lowest correlation between the features themselves. Specifically, this is done using the Python mrmr library. Different numbers of SNPs are set to observe the accuracy of the machine learning model classifier under different SNP counts. While ensuring good accuracy, combinations of fewer SNP sites are selected. The accuracy, precision, and F1 score of four machine learning classifiers with different numbers of SNPs are compared to determine the most suitable combination of SNP counts. The three evaluation metrics for different numbers of SNPs are as follows: Figure 3 and 4As shown, 36 SNPs were ultimately selected as the final molecular identifiers for Tianhua mutton sheep.
[0017] Example 2 Method 1 for identifying Tianhua meat sheep germplasm resources using Tianhua meat sheep molecular identification cards, the steps are as follows: I. Acquisition of SNP loci data for sheep molecular identity cards Blood or tissue samples were collected from the sheep to be tested, DNA was extracted, and sent to a sequencing company for sequencing. The sequencing data was then analyzed using BaseNumber DNA sequencing data analysis software or BWA+GATK to perform genome alignment and genotyping, resulting in a vcf.gz file. Further, bcftools software was used to extract the SNP loci (molecular identifiers) of Tianhua sheep based on their location information. For missing loci, the genotyping software Beagle was used for self-filling.
[0018] II. SNP locus genotype format conversion For the filled vcf.gz file, the SNP site information is converted into a 012 encoding format using Plink software, where 0 represents homozygous at the reference genome site, 1 represents heterozygous, and 2 represents homozygous at the mutation site. This ensures that the final characteristic SNPs are consistent with the molecular identity card of Tianhua mutton sheep.
[0019] III. Machine Learning Classifier Prediction The machine learning classifier prediction relies on Python. First, NumPy and Pandas are used for data processing, and joblib is used to save and load the model and other objects. sklearn is used to load the classifier model, and the machine learning model uses its predictive function to determine whether a sheep is a Tianhua mutton sheep genetic resource. An output of 0 indicates a Tianhua mutton sheep, and an output of 1 indicates a non-Tianhua mutton sheep.
[0020] Example 3 Method two for identifying Tianhua meat sheep germplasm resources using Tianhua meat sheep molecular identification cards, the steps are as follows: I. Acquisition of SNP loci data for sheep molecular identity cards Blood or tissue samples were collected from the sheep to be tested, DNA was extracted, and sent to a sequencing company for sequencing. The sequencing data was then analyzed using BaseNumber DNA sequencing data analysis software or BWA+GATK to perform genome alignment and genotyping, resulting in a vcf.gz file.
[0021] II. SNP locus data merging and phylogenetic tree construction The bcftools software was used to merge the vcf.gz file of the sheep to be tested with the existing vcf.gz file of Tianhua meat sheep, and 36 molecular identification loci of Tianhua meat sheep were extracted from the data. The genetic distance matrix was calculated using the plink software, and the genetic phylogenetic tree was constructed using the MEGA software. When the genetic distance of the sheep to be tested is close to that of the Tianhua meat sheep population, the sheep to be tested can be identified as Tianhua meat sheep.
[0022] Application examples The application of molecular identification for Tianhua mutton sheep in identifying Tianhua mutton sheep germplasm resources involves the following steps: I. Eighty-one new Tianhua sheep were selected as test subjects. Blood samples were collected from these sheep, and DNA was extracted and sent to Borui Biotechnology Co., Ltd. for sequencing. The sequencing data returned by the sequencing company were analyzed using BaseNumberDNA sequencing data analysis software to obtain the vcf.gz file. The reference genome used was ARS-UI_Ramb_v2.0. Further, bcftools software was used to extract 36 specific SNP loci for the Tianhua sheep molecular identity card based on their location information, and these loci were merged with existing populations used to construct the Tianhua sheep molecular identity card.
[0023] II. The genetic distance matrix was calculated using PLINK software, the phylogenetic tree was constructed using MEGA software, and the phylogenetic tree was beautified using the iTOL website. The results of the genetic evolution analysis are as follows: Figure 6 As shown, the sheep to be tested (i.e., the newly collected Tianhua meat sheep) has a close genetic relationship with the original Tianhua meat sheep, and it can be determined that the sheep to be tested belongs to the Tianhua meat sheep germplasm genetic resources.
[0024] Third, for the vcf.gz file obtained in the first step, missing loci were filled using the Beagle genotyping software. Then, the SNP locus information was converted to 012 encoding format using Plink software, where 0 represents homozygous at the reference genome locus, 1 represents heterozygous, and 2 represents homozygous at the mutation locus. Raw data processing was performed using Python's NumPy and Pandas, and data and model saving and loading were handled using joblib. The classifier model was loaded using sklearn. Taking the support vector machine (SVM) model, which performed best during training, as an example, the `predict` and `predict_proba` functions of the trainer model were used to judge the test sheep. The model output results are shown in Table 2.
[0025] Table 2. In Table 2, an output of 0 indicates Tianhua mutton sheep, while an output of 1 indicates non-Tianhua mutton sheep. All 81 newly collected sheep tested had an output of 0, and the prediction accuracy exceeded 90%. Therefore, it can be concluded that all newly collected sheep are of the Tianhua mutton sheep breed.
[0026] The above embodiments only illustrate several implementation methods of the present invention. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. The application of reagents for detecting SNP locus combinations in Tianhua meat sheep in the identification of Tianhua meat sheep germplasm resources, characterized by: The SNP locus combination consists of 36 characteristic SNP loci located in the sheep reference genome ARS-UI_Ramb_v2.0, and the SNP locus combination includes ID1-ID36, wherein: ID1-ID4 are located at positions 96522898, 139740899, 153767171, and 213201360 on chromosome 1, respectively. ID5 and ID6 are located at positions 71425831 and 210125476 on chromosome 2, respectively. ID7-ID10 are located at positions 49175961, 118709694, 132448684, and 218813552 on chromosome 3, respectively. ID11-ID16 are located at positions 77459709, 80420769, 90323194, 92536278, 104178591, and 116757608 on chromosome 4, respectively. ID17 and ID18 are located at positions 83142360 and 95932782 on chromosome 5, respectively. ID19 is located at position 66431083 on chromosome 6; ID20-21 are located at positions 20949969 and 98180525 on chromosome 7, respectively. ID22 is located at position 42520240 on chromosome 8; ID23-24 are located at positions 34397831 and 40091099 on chromosome 9, respectively; ID25-27 are located at positions 25822711, 33041271, and 42569697 on chromosome 14, respectively. ID28-29 are located at positions 7988627 and 15198939 on chromosome 15, respectively; ID30 is located at position 24325191 on chromosome 16; ID31 is located at position 19258771 on chromosome 17; ID32 is located at position 21,749,026 on chromosome 18; ID33 is located at position 37662890 on chromosome 20; ID34 is located at position 16,097,845 on chromosome 21; ID35-ID36 are located at positions 37701101 and 39398635 on chromosome 25; The mutation types at sites ID1, ID4, ID6, ID8, ID13, ID14, ID23, ID29, and ID30 are G / A; at sites ID2, ID3, ID27, and ID28, the mutation type is T / G; at sites ID5, ID10, ID16, ID17, ID19, ID20, ID21, ID24, ID25, and ID33, the mutation type is C / T; at site ID7, the mutation type is G / T; at sites ID9 and ID32, the mutation type is A / G; at sites ID11 and ID34, the mutation type is T / A; at sites ID12, ID18, ID26, ID35, and ID36, the mutation type is T / C; at sites ID15 and ID22, the mutation type is C / A; and at site ID31, the mutation type is A / T.
2. The application of the reagent for detecting SNP locus combinations in Tianhua mutton sheep as described in claim 1 in the identification of Tianhua mutton sheep germplasm resources, characterized in that: Includes the following steps: (1) Whole genome resequencing was performed on the sheep samples of unknown breeds to be tested to obtain the raw genome data; (2) The raw data obtained in step (1) are compared with the sheep reference genome ARS-UI_Ramb_v2.0 and genotyped to obtain the genotype data of the sheep to be tested; (3) Filter the genotype data in step (2), then use Beagle software to fill in the data, and then extract the genotype of the SNP site combination of Tianhua meat sheep. (4) Merge the genotype data of the sheep population to be tested and the SNP loci combination of Tianhua meat sheep, and determine whether it is a Tianhua meat sheep genetic resource by constructing a genetic evolution tree; or convert its genotype format to 012 format and input it into the trained Python machine learning classifier model, and determine whether it is a Tianhua meat sheep genetic resource based on its predict function.
Citation Information
Patent Citations
SNP site combination for identifying Mongolian sheep and non-Mongolian sheep and application of SNP site combination
CN116426649A
SNP (Single Nucleotide Polymorphism) site combination for quickly identifying Chinese merino sheep
CN119410780A