An SNP chip used to distinguish five types of native Hubei yellow cattle
By combining whole-genome sequencing and selection signal analysis with random forest method to screen SNP sites, an SNP chip for distinguishing Hubei native yellow cattle was developed, which solves the problem of the difficulty in accurately identifying Hubei yellow cattle breeds in existing technologies, and realizes efficient and low-cost breed identification and breeding support.
Patent Information
- Application Number
- CN202411786768.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2044-12-05
AI Technical Summary
Existing technologies are insufficient to efficiently and accurately distinguish the five native yellow cattle breeds in Hubei Province. Traditional pedigree identification methods are inaccurate and complex, molecular biology methods are difficult to operate and have certain toxicity, and high-throughput sequencing technology has not yet been widely used in cattle breed identification.
By combining whole-genome sequencing, selection signal analysis, and random forest methods to screen out SNP loci, and by scoring using a random forest model, at least 285 loci were selected for the development of SNP chips to identify five native Hubei yellow cattle breeds.
This study enabled high-precision identification of five native Hubei yellow cattle breeds, reduced sequencing costs, provided important technical support for Hubei yellow cattle breeding, and improved the accuracy and efficiency of breed identification.
Smart Images

Figure BDA0005174249370000061 
Figure BDA0005174249370000071 
Figure BDA0005174249370000081
Abstract
Description
Technical Field
[0001] This invention belongs to the field of animal genetics and breeding technology, specifically relating to an SNP chip for distinguishing five breeds of native Hubei yellow cattle. Background Technology
[0002] Cattle are an important livestock in my country, and Hubei Province boasts many native yellow cattle breeds. Hubei yellow cattle are characterized by their tolerance to roughage, high reproductive rate, and rapid weight gain in summer and autumn. Among them, Enshi cattle, Yiling cattle, and Yunba cattle are listed in the national breed catalog as livestock genetic resources, and "Enshi Yellow Beef" has obtained national geographical indication product protection. Yiling snowflake beef sells for 1200 yuan per kilogram, becoming a major beef cattle breed developed in Hubei Province. The five Hubei yellow cattle breeds—Yiling cattle, Dabie Mountain cattle, Yunba cattle, Enshi cattle, and Sanjiaoshan cattle—share very similar appearance characteristics, including similar coat color and size, making breed identification difficult based on appearance alone. In recent years, Hubei Province has continuously encouraged and developed beef cattle farming, introducing large foreign beef cattle breeds as breeding stock, which has formed a certain scale in the Hubei region. However, due to Hubei Province's location in southern my country, with a predominantly hot and humid climate, the introduced beef cattle breeds are poorly adapted to the hot and humid climate of the south, which is not conducive to production. At the same time, the large-scale introduction of foreign breeds has threatened the protection of my country's genetic resources, impacting the scale of Hubei Yellow Cattle farming and breeding to some extent. Furthermore, there is currently limited research on the genetic resources of Hubei Yellow Cattle. To fully explore the genetic resources of Hubei Yellow Cattle, it is necessary to conduct crossbreeding between Hubei Yellow Cattle and foreign breeds, and crossbreeding urgently requires a breed identification method. Although traditional pedigree identification methods can identify breeds through statistical pedigree data, they still face certain difficulties in terms of accuracy and the complexity of data collection. In addition, breed identification can also be performed by identifying microsatellite sequences in the genetic sequence. This method belongs to the field of molecular biology. Although it has advantages in accuracy, the reagents required for the experiment have a certain degree of toxicity, and the experimental cycle is long and the operation is difficult.
[0003] In recent years, high-throughput sequencing has developed rapidly, with second-generation sequencing and third-generation sequencing emerging one after another. Sequencing costs have further decreased, and sequencing accuracy and speed have further increased. Selection signal analysis plays a key role in population genetics, aiming to identify the impact of natural and artificial selection on variations in the genome, compare genetic variations in different populations, and reveal the relationship between adaptive traits and selection pressure. In specific applications, Chen et al. used FST, XP-EHH, and πratio selection signals to study the heat tolerance trait of zebu cattle (Chen, N., Xia, X., Hanif, Q. et al. Global genetic diversity, introgression, and evolutionary adaptation of indicine cattle revealed by whole genome sequencing. Nat Commun 14, 7803 (2023). https: / / doi.org / 10.1038 / s41467-023-43626-z). Random Forest (RF) is a machine learning method used for classification and regression. Gao et al. applied this method for breed identification in Chinese pigs (Gao, J.; Sun, L.; Zhang, S.; Xu, J.; He, M.; Zhang, D.; Wu, C.; Dai, J. Screening Discriminating SNPs for Chinese Indigenous Pig Breeds Identification Using a Random Forests Algorithm. Genes 2022, 13, 2207. https: / / doi.org / 10.3390 / genes13122207), but it has not yet been applied to cattle. This invention combines whole-genome sequencing and selection signal analysis to obtain SNP (single nucleotide polymorphism) sites. Further, by finding the taxonomic branch with the lowest error rate, the SNP sites are scored to obtain breed-specific SNP sites, which are then further developed into an SNP chip for distinguishing five breeds of native Hubei yellow cattle. Compared with traditional selection signal methods, combining random forests with selection signals can further reduce the number of sites required, achieving more accurate breed identification with fewer sites, and providing options for the breeding of native Hubei yellow cattle.
[0004] In summary, we developed an SNP chip for distinguishing five native Hubei cattle breeds (Yiling cattle, Dabie Mountain cattle, Yunba cattle, Enshi cattle, and Sanjiaoshan cattle) by combining second-generation high-throughput sequencing technology, selected signals, and random forest methods, providing important technical support for breed identification and improvement of Hubei cattle. Summary of the Invention
[0005] In order to overcome the shortcomings of the prior art, the present invention provides an SNP chip for distinguishing five types of native Hubei cattle (Yiling cattle, Dabie Mountain cattle, Yunba cattle, Enshi cattle, and Sanjiaoshan cattle).
[0006] The present invention adopts the following technical solution:
[0007] This invention uses random forest and selection signal methods to screen SNP locus combinations for the identification of five Hubei Yellow Cattle breeds. The specific screening method includes:
[0008] S1. Collect raw second-generation whole genome sequencing data of cattle, compare it with the reference genome of cattle ARS-UCD1.2, and identify and screen out SNP variant sites;
[0009] S2. Based on S1, select SNP variant sites, use Tajima's D, FST, and πratio selection signals to take the union, and then use LD trimming to obtain the selected sites;
[0010] S3. Using a random forest model, the corresponding loci and variety information are input to score the loci, and the minimum number of loci required to achieve 100% accuracy for the five varieties is selected. Specifically, the loci are input into the random forest model, scored, and the Gini and accuracy scores are obtained. The loci data are then divided into a test set and a training set. At the same time, the variety information is input into the random forest model. Using the loci with Gini scores, the number of loci is gradually increased in descending order of scores to find the minimum 285 loci that can achieve 100% accuracy in identifying the five varieties and are stable.
[0011] S4. Using a random forest model, the corresponding loci and variety information are input to score the loci, and the minimum number of loci required to achieve 100% accuracy for the five varieties is selected. Specifically, the loci are input into the random forest model, scored, and the Gini and accuracy scores are obtained. The loci data are then divided into a test set and a training set. At the same time, the variety information is input into the random forest model. Using the loci with Gini scores, the number of loci is gradually increased in descending order of scores to find the minimum 285 loci that can achieve 100% accuracy for the identification of the five varieties and are stable. The specific loci information is shown in Table 1.
[0012] A molecular probe array is used to detect the SNP sites described in Table 1. Probes are designed based on the SNP sites and their preceding and following sequences (60 bp) as described in Table 1.
[0013] SNP chips were used to distinguish five native Hubei yellow cattle species (Yiling cattle, Dabie Mountain cattle, Yunba cattle, Enshi cattle, and Sanjiaoshan cattle). The gene chips were loaded with molecular probe combinations for detecting the SNP sites listed in Table 1.
[0014] A method for distinguishing between Yiling cattle, Dabie Mountain cattle, Yunba cattle, Enshi cattle, and Sanjiaoshan cattle includes the following steps:
[0015] (1) Obtain the SNP locus genotype of the cattle to be tested. The SNP loci are shown in Table 1.
[0016] (2) Perform cluster analysis or principal component analysis on the SNP locus genotype of the cattle to be tested and the known genotypes of Yiling cattle, Dabieshan cattle, Yunba cattle, Enshi cattle and Sanjiaoshan cattle, and determine the breed of the cattle to be tested based on the analysis results.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0018] 1. The five Hubei yellow cattle breeds involved in this invention—Yiling cattle, Dabieshan cattle, Yunba yellow cattle, Enshi cattle, and Sanjiaoshan cattle—encompass the main Hubei yellow cattle breeds and are of great representative importance for the selection and breeding of Hubei yellow cattle. They are also of great application value in the molecular breeding of Hubei yellow cattle.
[0019] 2. In addition to the conventional single selection signal, this invention uses three selection signals—Tajima's D, FST, and πratio—to select SNP variant sites, further increasing the accuracy of the selected sites. At the same time, random forest scoring is applied to the selected sites, further reducing the number of sites required and improving the classification effect, thereby achieving lower sequencing costs and finding the most representative sites.
[0020] 3. This invention can provide a good research foundation and data support for the germplasm identification, selection breeding and other research of Hubei Yellow Cattle, and further reduce the cost of bovine genome selection, accelerate the genetic progress of the improvement of high-quality cattle breeds in my country, and has good social value and promotion value. Attached Figure Description
[0021] Figure 1 Flowchart for screening SNP combinations of five native Hubei yellow cattle species.
[0022] Figure 2 Line graph showing the number of SNP sites and the accuracy of species differentiation in a random forest model.
[0023] Figure 3Figure 1 shows the results of three-dimensional PCA analysis of different SNP locus combinations on five native Hubei yellow cattle species. Figure A shows the PCA analysis results of 17,579,873 loci obtained based on sequencing data alignment, and Figure B shows the PCA analysis results of 285 SNP loci screened based on selection signal analysis and random forest method.
[0024] Figure 4 The results of cluster analysis of different SNP locus combinations on five native Hubei yellow cattle species are shown in Figure A, which shows the cluster analysis results of 17,579,873 loci obtained based on sequencing data alignment, and Figure B shows the cluster analysis results of 285 SNP loci selected based on selection signal analysis and random forest method. Detailed Implementation
[0025] To better illustrate the purpose, technical solution, and advantages of the present invention, the present invention will be described in detail below with reference to specific embodiments.
[0026] This invention uses random forest and signal selection methods to screen SNP locus combinations for the identification of five native Hubei yellow cattle breeds. The specific screening and application are as follows:
[0027] S1. Collection of bovine second-generation whole-genome sequencing data
[0028] Liver tissues from 63 Hubei Yellow Cattle of 5 breeds (including Yiling cattle, Dabie Mountain cattle, Yunba cattle, Enshi cattle, and Sanjiaoshan cattle) were collected and sent to BGI for second-generation whole-genome sequencing to obtain FASTQ data of paired-end sequencing.
[0029] S2. Comparison of sequencing data
[0030] The sequence information of the FASTQ format sequencing data files was aligned to the reference genome ARS-UCD1.2 of domestic cattle using the MEM algorithm of the BWA software with default parameters. The obtained BAM files were sorted according to chromosome position using SAMtools, and redundancy was removed and indexed using the default parameters of the SAMtools software.
[0031] S3. Construct a mutation map and perform quality control on it.
[0032] S31. Using the bam file preprocessed in S2, identify SNP variant sites using the sentieon driver parameter "--algoHaplotyper". One gvcf file is obtained for each individual, resulting in a total of 63 gvcf files and their corresponding indexes.
[0033] S32. Merge all the gvcf files obtained in S31 using the sentieon driver parameter "--algo GVCFtyper" to obtain a vcf file containing 63 Hubei yellow cattle SNPs and InDels.
[0034] S33. Use GATK to perform hard filtering on the obtained VCF file with the following parameters: "QUAL<30.0||QD<2.0||MQ<40.0||FS>60.0||SOR>3.0||MQRankSum<-12.5||ReadPosRankSum<-8.0\". This filters the corresponding loci. These parameters mainly filter the quality value of the locus variant, the quality-to-depth ratio of the variant, the mapping quality, and the mapping quality ranking. The main function of these parameters is to filter false positive variants and improve the reliability of the detected variants.
[0035] S34. Perform quality control on the VCF file, using GATK's SelectVariants to remove InDel. Our main focus of analysis is SNP sites.
[0036] S35. Using the SelectVariants tool in GATK with the parameters set to "--exclude-filtered--restrict-alleles-to BIALLELIC", only biallelic loci are retained, resulting in 133,388,567 loci.
[0037] S36. Use Plink to filter the data for missing loci and minimum allele frequency, with the specific parameters "--geno 0.05--maf 0.05". This step reduces invalid loci and ensures the validity of the analyzed data, yielding 17,579,873 loci.
[0038] S4. Use three selection signals, take their intersection, and then perform LD trimming to filter sites.
[0039] S41. Use VCFtools to input the vcf file obtained in S36, and then input the grouping data of the five Hubei Yellow Cattle breeds that we need to distinguish. Calculate the following three selection signals: Tajima's D (a selection signal within a population that assesses genomic variation through nucleotide diversity and haplotype frequency in a sliding window), FST (population divergence index, a measure of genetic differentiation between different populations, its main principle is to compare genetic variation within a population with genetic variation between populations), and πratio (nucleotide diversity ratio, used to measure genetic diversity within a population, calculated as the average nucleotide difference between each pair of alleles; this ratio can be used to compare the diversity levels of different populations or different genomic regions).
[0040] S42. Organize the results of the three selection signals. After sorting FST, Tajima's D and πratio in descending order, take the first 200 sites of each. Take the union of the sites obtained from the three selection signals and remove duplicate sites to obtain 3978 SNPs. Taking the union of the three selection signals can retain the results of the three selection signals to the greatest extent.
[0041] S43. Use Plink to perform LD trimming on the data obtained in S42 with the parameter "--indep-pairwise 500500.5", which will eventually yield 763 sites. LD trimming can remove sites that are linked by LD, ensuring the simplicity of the sites and further improving the accuracy of the sites.
[0042] S5. Scoring the loci using a random forest model, and calculating the minimum number of loci required to identify five Hubei Yellow Cattle breeds based on the loci scores.
[0043] S51. Input the 763 sites obtained in S43 into the random forest model, score the sites, and obtain the gini and accuracy scores of the 763 sites. The sites with higher scores in the gini model will be ranked using the gini scores in subsequent analysis.
[0044] S52. Arrange the loci in the Gini model in descending order of score, and increase the selected loci combinations from 5 to 5 each time. For each variety, randomly divide individuals into a test set and a training set, with a test set:training set ratio of 7:3. Simultaneously, input the variety information corresponding to each individual into the random forest model. As the number of loci increases, the identification accuracy of the five varieties gradually changes. Figure 2 Finally, the minimum number of loci combinations that can achieve a 100% accuracy rate in identifying the five breeds and remain stable were identified. The minimum number of loci combinations obtained included 285 loci, and the specific locus information is shown in Table 1 (the physical location of the loci was determined based on the ARS-UCD 1.2 version of the bovine reference genome).
[0045] Table 1 shows the SNP locus combinations used to distinguish five species of native Hubei yellow cattle.
[0046]
[0047]
[0048]
[0049]
[0050] S6. Test the accuracy of SNP locus combinations.
[0051] S61. The 285 loci obtained in S52 were used to test the identification effect of 63 samples of five native Hubei yellow cattle breeds using 3D PCA. The PCA value was calculated using Plink "--pca 10", and the 3D PCA plot was drawn using the R package scatterplot3d based on PC1, PC2, and PC3. At the same time, in order to compare the results, the data obtained in S3 (17,579,873 loci) was plotted using the same method to draw a 3D PCA plot.
[0052] Figure 3 A represents the PCA analysis results of 17,579,873 sites. Figure 3 B represents the PCA analysis results of the 285 selected sites, and... Figure 3 Compared to A, Figure 3 The clustering effect of the five varieties in B is more significant. The 285 sites selected by three selection signals and random forest scoring can effectively distinguish the five Hubei Yellow Cattle breeds.
[0053] S62. The 285 loci obtained in S52 were used to test the identification effect of 63 samples from five Hubei Yellow Cattle breeds using clustering graphs. The distance matrix was calculated using VCF2Dis, and the data was converted to .tre format using the FastMe website (http: / / www.atgc-montpellier.fr / fastme / ). The file was then input into the ITOL website (https: / / itol.embl.de / tree) for visualization. Simultaneously, to compare the results, the data obtained in S3 (17,579,873 loci) was used to create a clustering graph using the same method.
[0054] Figure 4 A represents the cluster analysis results for 17,579,873 sites. Figure 4 B represents the cluster analysis results of the 285 selected sites. Figure 4 Compared to A, Figure 4The five breeds in group B have greater genetic distance and more obvious differentiation. The 285 loci selected by three selection signals and random forest scoring can increase the differentiation of the five Hubei Yellow Cattle breeds.
[0055] In summary, based on the results of three-dimensional PCA analysis and cluster analysis, the 285 loci obtained through three selection signals and random forest can effectively distinguish five native Hubei yellow cattle breeds: Yiling cattle, Dabieshan cattle, Yunba cattle, Enshi cattle, and Sanjiaoshan cattle. Based on the SNP loci and their preceding and following sequences (60 bp) as described in Table 1, probes can be designed to further develop an SNP chip for distinguishing these five native Hubei yellow cattle breeds.
Claims
1. The application of SNP locus combinations in distinguishing five species of native Hubei yellow cattle, characterized by: The information of the SNP loci is shown in Table 1 of the specification. The five native Hubei yellow cattle species are Yiling cattle, Dabieshan cattle, Yunba cattle, Enshi cattle, and Sanjiaoshan cattle. The physical location of the loci is determined based on the ARS-UCD version 1.2 bovine reference genome.
2. A molecular probe assembly, characterized in that, The molecular probe combination detects the SNP sites shown in Table 1 of the instruction manual, the physical location of which is determined based on the bovine reference genome version ARS-UCD 1.
2.
3. The application of the molecular probe combination according to claim 2 in distinguishing five types of native Hubei yellow cattle, characterized in that, The five native Hubei cattle breeds are Yiling cattle, Dabie Mountain cattle, Yunba cattle, Enshi cattle, and Sanjiaoshan cattle.
4. Distinguishing the SNP chips of five types of local Hubei yellow cattle, characterized by: The gene chip is loaded with the molecular probe combination as described in claim 2, and the five native Hubei cattle species are Yiling cattle, Dabie Mountain cattle, Yunba cattle, Enshi cattle, and Sanjiaoshan cattle.
5. A method for distinguishing five types of native Hubei yellow cattle, namely Yiling cattle, Dabie Mountain cattle, Yunba cattle, Enshi cattle, and Sanjiaoshan cattle, characterized in that, Includes the following steps: (1) Obtain the SNP locus genotype of the cattle to be tested. The SNP loci are shown in Table 1 of the instruction manual. The physical location of the loci is determined based on the bovine reference genome version ARS-UCD1.
2. (2) Perform cluster analysis or principal component analysis on the SNP locus genotype of the cattle to be tested and the known genotypes of Yiling cattle, Dabieshan cattle, Yunba cattle, Enshi cattle and Sanjiaoshan cattle, and determine the breed of the cattle to be tested based on the analysis results.
Citation Information
Patent Citations
Liquid-phase breeding chip for Hainan cattle and application of liquid-phase breeding chip
CN115198023A