Cattle SNP (Single Nucleotide Polymorphism) chip and application thereof

By designing a liquid-phase SNP chip for Tibetan local cattle and using whole-genome resequencing and chip data integration, specific SNP sites were screened out, solving the problem of poor typing results in Tibetan local cattle populations and achieving efficient genetic resource protection and breed identification.

CN120776005APending Publication Date: 2025-10-14CHINA AGRI UNIV

Patent Information

Application Number
CN202511206111.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

Existing commercial solid-phase and liquid-phase chips have poor typing effects in Tibetan local cattle populations, making it difficult to meet the needs of breed identification and genetic resource protection. There is a lack of efficient typing tools suitable for local cattle.

Method used

Based on liquid phase chip technology, a SNP chip was designed for Tibetan local cattle. Through whole genome resequencing and chip data integration, 7,829 specific SNP sites were screened and a liquid phase chip was constructed to achieve accurate typing and identification of local cattle populations.

Benefits of technology

It has achieved accurate identification of local cattle populations in Tibet, improved typing consistency and identification efficiency, and provided technical support for genetic resource protection and breed selection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120776005A_ABST
    Figure CN120776005A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of biology, and discloses a cattle SNP chip and application thereof, the detection object of the liquid phase chip comprises 7,829 SNP sites located on a common cattle reference genome AR-UCD 1.2, and the position information of the SNP sites is shown in the specification table 3. The liquid chip provided by the invention highlights the specificity of the local cattle variety in the Tibet autonomous region, and by analyzing the local cattle population structure, sites capable of effectively distinguishing the population are screened, so that the problems of strong universality and poor specificity of a commercial chip are solved. The liquid-phase chip can also realize flexible selection and upgrading of sites, and breaks through the limitation that a solid-phase chip cannot flexibly supplement or update the sites. The whole method can be stably operated on the premise of limited sample size and complex population structure, and is suitable for genetic resource protection and application scenarios of rare varieties. The method can also be expanded into a standardized typing tool, and is suitable for various applications such as local cattle germplasm registration, variety identification, region identification and the like.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biotechnology, and relates to a SNP chip for cattle and application thereof, in particular to a SNP chip for Tibetan local cattle and application thereof. BACKGROUND

[0002] At present, genome selection technology has been widely used in genetic evaluation of economic cattle breeds such as dairy cattle in China, and the core is the genotyping of a large number of known SNP sites. The mainstream genotyping method is a commercial solid-phase chip, which relies on the hybridization principle of fixed fluorescent labeled probes, and has the advantages of high throughput and standardization. However, this technology has certain limitations in site selection flexibility and later upgrading, especially when facing local cattle breeds with complex genetic background and scarce sample resources, the genotyping effect is poor, which is difficult to meet the actual demand.

[0003] Liquid chip is a kind of genotyping technology that has developed rapidly in recent years, based on the principle of targeted sequencing (Genotype By Targeted Sequencing, GBTS), with the advantages of flexible site selection, strong expansion, controllable cost, etc., especially suitable for developing personalized site set for specific breeds or populations to meet the demand of accurate genotyping. In the field of liquid chip, there are practical applications based on GBTS technology in China. For example, the 85K liquid chip for dairy cattle integrates four types of key genetic markers: (1) the core sites in the existing 50K, 80K and 150K solid-phase chips; (2) the economic trait functional gene sites discovered in Chinese Holstein cattle research; (3) known genetic defect sites of dairy cattle; (4) core sites required for parentage identification. The chip has good genotyping performance and application value in mainstream dairy cattle populations.

[0004] Although the existing commercial solid-phase chip and liquid chip have been widely used in dairy cattle breeding and genome selection, their main design is based on the genetic background of European and American dairy cattle or Chinese Holstein cattle, which lacks specific support for Tibetan local cattle populations, and it is difficult to meet the genotyping needs of diversified local breeds. Tibetan local cattle breeds include Apeijiazha cattle, Chayu yellow cattle, Camel Peak cattle, Tibetan cattle and Zhangmu cattle, etc., with complex population structure and scarce sample resources. Conventional chips have problems such as low individual detection rate, low genotyping accuracy and poor population recognition ability in the genotyping process. At the same time, the existing technology cannot effectively meet the demand for accurate breed identification in practical application scenarios such as local cattle germplasm resource protection and traceability, and there is a lack of efficient and popular genotyping tools suitable for local cattle.

[0005] Therefore, the present application combines the germplasm resources of Tibetan local cattle, and constructs a specific genetic marker site set based on liquid chip technology, which is the first to realize accurate genotyping of Tibetan local cattle populations, and provides an innovative solution for local cattle breed identification and traceability. SUMMARY

[0006] The present application aims at the above-mentioned deficiencies, and provides a SNP chip for Tibetan local cattle and an application thereof. Based on a liquid chip technology and a set of genetic marker sites specially oriented to Tibetan local cattle population, a set of efficient and scalable local cattle breed typing and identification scheme is constructed. By accurately selecting SNP sites with representativeness and distinguishing power, accurate identification of local cattle population is realized, and typing consistency and identification efficiency are improved, thereby providing technical support and application guarantee for genetic resource protection, germplasm evaluation, breed selection and traceability management of local cattle.

[0007] In order to achieve the above-mentioned purpose, the present application adopts the following technical scheme:

[0008] The present application provides a liquid chip for Tibetan local cattle species identification, wherein the detection object of the liquid chip includes 7,829 SNP sites located on the bovine reference genome ARS-UCD 1.2, and the position information of the SNP sites is shown in Table 3 of the specification.

[0009] Preferably, the Tibetan local cattle species includes but is not limited to Apei Jiazha cattle, camel peak cattle, Tibetan cattle, Zhangmu cattle and Chayu yellow cattle.

[0010] Preferably, the identified cattle species includes not only the Tibetan local cattle species, but also Holstein cattle, Jersey cattle and tumor cattle.

[0011] The present application also provides a design method of the above-mentioned liquid chip, including the following steps:

[0012] S1, obtaining whole genome resequencing data, and obtaining Tibetan local cattle specific SNP sites by processing;

[0013] S2, obtaining 100K chip data, and obtaining SNP sites related to Tibetan local cattle breed specificity by processing;

[0014] S3, integrating the specific SNP sites obtained from resequencing and chip data, and screening 7,829 SNP sites;

[0015] S4, designing probes based on the screened 7,829 SNP sites, and preparing a liquid chip.

[0016] Preferably, the integration method includes one or more of removing duplicate sites and removing sites with a detection rate lower than 90% by actual sample detection.

[0017] Preferably, step S1 includes the following steps:

[0018] S11, sample collection

[0019] Collecting biological samples of A'pei Jiazu cattle, Lufeng cattle, Tibetan cattle and Zhangmu cattle, extracting DNA;

[0020] S12, library construction and sequencing

[0021] Resequencing the DNA sample obtained in step S11 to obtain raw sequencing data;

[0022] S13, specific site screening

[0023] Quality control, SNP detection and specificity screening are performed on the raw sequencing data to obtain Tibetan local cattle specific SNP sites.

[0024] Preferably, in step S11, the biological sample is collected in the form of blood, tissue, semen, hair follicle or feces.

[0025] Preferably, the blood is from the tail root vein or the jugular vein.

[0026] Preferably, in step S13, the method of quality control includes one or more of the following: using Trimmomatic v0.38 to remove low-quality sequences and adapter sequences, aligning to the common cattle reference genome ARS-UCD 1.2, using Picard v2.18.2 to sort and remove PCR duplicates.

[0027] Preferably, in step S13, the method of specificity screening includes one or more of the following: using an allele frequency difference threshold of 90%-10% to screen, using a Python script to identify sites with significant differences between different breeds, and combining gene function annotation for analysis and minimum allele frequency filtering.

[0028] Preferably, step S2 includes the following steps:

[0029] S21, sample collection and chip detection

[0030] Using the DNA of A'pei Jiazu cattle, Lufeng cattle, Tibetan cattle, Zhangmu cattle, plateau Holstein cattle, plateau Janan cattle, Yunnan humpbacked cattle and Brahman cattle, and obtaining 770K chip data of American Brahman cattle from WIDDE database, converting the reference genome version to ARS-UCD 1.2;

[0031] S22, chip data processing

[0032] Genotyping, quality control and merging are performed on all sample DNA obtained in step S21 to obtain quality-controlled genotype data;

[0033] S23, population genetic analysis

[0034] The genotype data after quality control is subjected to principal component analysis, a phylogenetic tree is constructed, population component inference is carried out, and SNP sites related to the specificity of the Tibetan local cattle breed are obtained.

[0035] The application also provides application of the liquid chip, including application in Tibetan local cattle genotyping or application in Tibetan local cattle species identification or application in Tibetan local cattle breeding.

[0036] The application also provides a method for identifying the Tibetan local cattle species by using the liquid chip, comprising the following steps:

[0037] Step 1, extracting DNA of the sample to be tested

[0038] The DNA sample to be tested is extracted, the purity and integrity of the extracted DNA sample are detected, and the DNA concentration is accurately quantified;

[0039] Step 2, constructing a library

[0040] After constructing the DNA library, Qubit 2.0 is used for preliminary quantification, and the effective concentration of the library is accurately quantified by using the qPCR method to ensure the quality of the library;

[0041] Step S3, sequencing

[0042] The biotin-labeled probe is hybridized with the target region of the genome in a liquid system to form a double strand, the streptavidin-coated magnetic beads are used to adsorb the molecules carrying biotin, so that the target hybridized with the probe is captured, the raw sequencing data is obtained, the fastQC software is used for quality control, and finally the genotype of the SNP of the sample to be tested is obtained.

[0043] Step 4, breed typing and population identification analysis

[0044] Based on the specific SNP sites, principal component analysis, population evolution tree construction and Admixture analysis are used to classify the target sample; by comparing with the known breed information, the breed attribution judgment and population structure recognition of the unknown sample are realized.

[0045] Preferably, the quality control screening step comprises: removing the adapter sequence; removing paired reads with more than 10 N bases in the sequencing data read; removing paired reads with more than 40% of the length of the read of low-quality bases with Q<20 in the sequencing data read to obtain Clean data; then mapping the Clean data to the common cattle reference genome ARS-UCD 1.2 using bwa software to generate a bam file, sorting the bam file using Picard software, and removing all PCR duplicate reads, and finally detecting SNPs and genotyping using the standard process of GATK software.

[0046] Compared with the prior art, the present application has the beneficial effects that:

[0047] 1. The present application highlights the specificity of local cattle breeds in Tibet Autonomous Region in chip design, solves the problem of strong universality and poor specificity of commercial chips by analyzing the population structure of local cattle and screening sites that can effectively distinguish the population.

[0048] 2. The liquid chip prepared by the present application can realize flexible selection and upgrading of sites, breaking through the limitation of solid-phase chips that cannot flexibly supplement or update sites. The entire method can stably operate under the premise of limited sample size and complex population structure, and is suitable for genetic resource protection and rare breed application scenarios. It can also be expanded as a standardized typing tool, suitable for local cattle germplasm registration, breed identification, regional identification and other applications. BRIEF DESCRIPTION OF DRAWINGS

[0049] Figure 1 Figure 7 is the distribution of 7,829 SNP sites on the common cattle genome (set the window to 1Mb);

[0050] Figure 2 Figure 5 is the functional annotation and frequency statistics results of 7,829 specific SNP sites (A is the minimum allele frequency distribution graph of 7,829 SNP sites, and B is the statistical graph of site functional annotation results);

[0051] Figure 3 Figure 3 is the genetic structure analysis, population structure analysis PCA and phylogenetic tree graph of 36 local cattle populations in Tibet Autonomous Region (A is the population genetic structure analysis graph, B is the population structure analysis PCA graph, and C is the phylogenetic tree graph). DETAILED DESCRIPTION

[0052] For those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. It should be understood that the specific embodiments described herein are intended to explain the present application only, not to limit the present application. The present application can be implemented without some of these specific details for those skilled in the art. The following description of the embodiments is merely to provide a better understanding of the present application by showing examples of the present application.

[0053] The experimental methods used in the following examples are conventional methods unless otherwise specified.

[0054] The materials, reagents, etc. used in the following examples can be obtained from commercial channels unless otherwise specified.

[0055] Example 1

[0056] I. Screening of specific SNP sites and chip design

[0057] The present embodiment provides a specific SNP site screening method for Tibetan local cattle based on whole genome resequencing combined with chip detection data. Through systematic analysis of whole genome data and chip data of different local breed cattle populations in Tibet Autonomous Region, specific genetic markers capable of distinguishing different cattle breeds are screened, and a customized liquid chip candidate site set is constructed in combination with reference genome information. The specific steps are as follows:

[0058] 1. Acquisition and processing of whole genome resequencing data

[0059] Step 1. Sample collection

[0060] Four breeds, Apeijiazha cattle, Tuofeng cattle, Tibetan cattle and Zhangmu cattle, were selected from local breed cattle in Tibet Autonomous Region, a total of 29 individuals. Tail root venous blood was collected, and after DNA extraction, it was used for library construction and sequencing. The quality of sequencing data of each breed is shown in Table 1.

[0061] Table 1 Quality of whole genome resequencing data of 4 local cattle breeds

[0062]

[0063] Step 2. Library construction and sequencing

[0064] After the genomic DNA is broken by ultrasonic, it is subjected to fragmentation treatment, and then the fragment purification, end repair, 3' end A addition, and adapter ligation are completed in turn. The appropriate fragment length is selected by agarose gel electrophoresis. Then PCR amplification is performed to obtain the sequencing library. After the library passes the quality inspection, double-end high-throughput sequencing is performed using the Illumina platform.

[0065] Step 3. Data preprocessing

[0066] The raw sequencing data was stored in FASTQ format. Quality control was performed with Trimmomatic v0.38 to remove low quality sequences and adapter sequences. The qualified data was aligned to the Bos taurus reference genome ARS-UCD 1.2 by BWA v0.7.17 (mem algorithm). File conversion and indexing were performed by Samtools v1.8, and bam files were sorted and PCR duplicates were removed by Picard v2.18.2. Further local realignment and base quality recalibration (BaseRecalibrator, ApplyBQSR modules) were performed by GATK 3.8, and variant calling was performed by HaplotypeCaller. Finally, 41,390,691 SNP sites were detected.

[0067] Step 4, specific site screening

[0068] To obtain marker sites with significant differences between breeds, a difference threshold of 90%-10% in allele frequency was used for preliminary screening. Python scripts were used to identify sites with significant differences between different breeds, and gene function annotations were used for analysis. Further filtering was performed on the obtained sites based on the minimum allele frequency (MAF). Finally, 11,984 Tibetan local cattle-specific SNP sites derived from resequencing data were obtained.

[0069] 2, Acquisition and processing of 100K chip data

[0070] Step 1, sample collection and chip detection

[0071] In Lhasa City, Nyingchi City, Gyamda County, and Nyalam County of Tibet, 177 Tibetan local cattle were collected, including 61 Tibetan cattle, 39 Apei Jiazha cattle, 57 Tuofeng cattle, and 20 Zhangmu cattle. In addition, 22 highland Holstein cattle, 20 highland Jersey cattle, 46 Yunnan hump cattle, and 11 Brahman cattle were collected. Meanwhile, 770K chip data of 45 American Brahman cattle were obtained from the WIDDE database, and the reference genome version was converted to ARS-UCD 1.2 by the liftOver tool. Finally, chip sample information of 321 cattle from 9 breeds was formed, as shown in Table 2.

[0072] Table 2 Chip sample information of 9 breeds of cattle

[0073]

[0074] Step 2, chip data processing

[0075] All sample DNA was genotyped using Illumina GGP100K chips, and 95,256 SNP sites were initially obtained. The Plink software was used to filter sites with a minimum allele frequency (MAF) less than 0.05 and a detection rate lower than 90%, and the genotype data after quality control was obtained. The genotype data of 177 Tibetan local cattle was combined with the data of the reference population (Holstein, Jersey and Zebu population), and the common SNP sites were used for subsequent analysis.

[0076] Step 3, population genetic analysis

[0077] The combined 321 sample genotype data was subjected to principal component analysis (PCA, Plink--pca implementation), the MEGA X software was used to construct a phylogenetic tree (neighbor-joining method based on maximum composite likelihood genetic distance matrix), and the population component inference was performed by Admixture software. By this method, 41,230 SNP sites related to the specificity of Tibetan local cattle breeds were obtained.

[0078] 3, integration and screening of final sites

[0079] The specific SNP sites obtained by resequencing and chip data were integrated, and after removing the duplicates, a total of 8,603 sites were obtained. Through actual detection of 36 Tibetan local cattle samples (covering five groups of Tibetan cattle, Lufeng cattle, Zhangmu cattle, Apei Jiazha cattle and Chayu yellow cattle, which can fully reflect the genetic background of local cattle in Tibet Autonomous Region), sites with a detection rate lower than 90% were removed, and finally 7,829 sites were determined as the final specific SNP site set of the liquid phase chip of the application, see Table 3.

[0080] Table 3 Specific SNP site information in the liquid phase chip

[0081]

[0082]

[0083]

[0084]

[0085]

[0086]

[0087]

[0088]

[0089]

[0090]

[0091]

[0092]

[0093]

[0094]

[0095]

[0096]

[0097]

[0098]

[0099]

[0100]

[0101] 4. Synthesis of probes

[0102] Probes were designed and synthesized according to the information of 7,829 SNP sites, and the probes were tested and adjusted to determine the final probes and capture sites.

[0103] According to the principle of DNA complementarity, one or more probes covering the target SNP were designed at each site to be tested. These probes modified with biotin can hybridize with the target region in the denatured resequencing library to form double-stranded DNA. Streptavidin-coated magnetic beads are used to adsorb the molecules carrying biotin. After elution, amplification and sequencing, the genotype of the target SNP is finally obtained.

[0104] (1) Investigation of the distribution of 7,829 SNP sites on the common cattle genome

[0105] The CMplot software package in R language was used to statistically and visually analyze the positions of 7,829 sites on the common cattle reference genome ARS-UCD 1.2. In order to show the distribution characteristics of the sites on the whole genome, the common cattle reference genome was divided into 1 Mb window intervals. The reference genome was arranged in order of chromosome number, and the number of SNPs contained in each window was counted. The SNP distribution map was generated by CMplot software.

[0106] The results showed that 7,829 SNP sites were evenly distributed in the whole genome, and there was no obvious aggregation or deletion region. Each chromosome was well covered (see Figure 1). By 7,829 SNP sites evenly distributed in the whole genome, the liquid chip provided by the application can ensure comprehensive coverage of the bovine genome and representativeness of genetic information, and provide a reliable basis for subsequent molecular identification, population structure analysis and molecular breeding application.

[0107] (2) Functional annotation and frequency statistics of the screened 7,829 specific SNP sites

[0108] The calculation of the minor allele frequency (MAF) is based on the chip detection population data, and the Plink software is used to statistically analyze the population genotype data (allele frequency) of 7,829 sites in the tested 36 sample population, so as to obtain the MAF distribution in the whole genome. The functional annotation of the site is based on the ARS-UCD 1.2 annotation information of the common cattle reference genome, and the ANNOVAR software is used to annotate the gene region of each site.

[0109] The results show that the MAF of most sites is greater than 0.1, and has good population polymorphism (see Figure 2 A). These sites are distributed in multiple functional regions, covering untranslated regions, non-coding RNA related regions, coding region variations (synonymous / non-synonymous / stop mutations), splicing regions, intron regions, and upstream and downstream regulatory regions (see Figure 2 B), which indicates that the screened SNP sites can be evenly covered in functionally related and structurally related regions, and can comprehensively reflect the genetic information related to function and structure. Therefore, the practicability and representativeness of the chip are ensured, and the stability and functional correlation of genetic markers are considered, thereby providing a reliable tool for molecular breeding and genetic improvement.

[0110] (3) Population genetic structure analysis, population structure analysis PCA and phylogenetic tree diagram of 36 local cattle in Tibet Autonomous Region

[0111] The liquid chip containing 7,829 sites constructed by the application is used for genotyping detection of 36 local cattle in Tibet Autonomous Region (test samples). First, the Admixture software is used to analyze the population genetic structure of the test samples under different ancestral population settings (K=3, K=4, K=5), and the population component diagram is drawn. Second, the Plink software (—pca 20) is used for principal component analysis, the genotype matrix is extracted, and the genetic differences between different cattle populations are displayed. Finally, based on the genotype data of 7,829 sites, the Plink software (—distance square) is used to extract the distance matrix, the phylogenetic tree is constructed by the R package ape, and the phylogenetic relationship between different individuals is depicted, and then the R language ggplot2 is used for visualization.

[0112] The results showed that at different K values, the genetic components of each cattle group showed obvious differentiation characteristics, which could clearly distinguish different breeds or groups (see Figure 3 A in ); samples from different groups formed relatively independent clusters in the PC1 and PC2 coordinate systems, and the genetic relationships between groups were significantly different (see Figure 3 B in ); Individuals from different groups clustered into independent branches on the phylogenetic tree, which was highly consistent with their species classification (see Figure 3 C).

[0113] Through population structure analysis, principal component analysis, and phylogenetic tree verification, the 7,829 loci provided by this invention can accurately distinguish different populations of local cattle in the Tibet Autonomous Region, demonstrating excellent genetic resolution and representativeness. This demonstrates that the liquid phase array designed by this invention can not only be used for genetic structure analysis of cattle populations and breed identification, but also provides a reliable basis for molecular breeding practices, and has important practical significance for the protection and utilization of genetic resources.

[0114] Example 2 Tibetan cattle SNP chip detection and analysis process

[0115] 1. Sample Collection and DNA Extraction

[0116] Blood was collected from local cattle breeds in Tibet Autonomous Region (Apai Jiazha cattle, Bamai cattle, Zayu cattle, Camelback cattle, Tibetan cattle and Zhangmu cattle, etc.) through the tail root vein for DNA extraction.

[0117] use DNA samples were extracted using a DNA extraction kit (magnetic bead method). The extracted DNA samples were subjected to two tests: (1) DNA purity and integrity analysis using 1% agarose gel electrophoresis; and (2) DNA concentration was accurately quantified using Qubit. DNA samples with a DNA content >250 ng qualified the quality control and proceeded to the next step of the experiment.

[0118] 2. Library Construction

[0119] use DNA libraries were constructed using the DNA Library Prep Kit. After library construction, preliminary quantification was performed using Qubit 2.0. qPCR was then used to accurately quantify the effective concentration of the library to ensure library quality. Once the library passed the assay, sequencing began.

[0120] 3. Sequencing

[0121] The technology utilizes a biotin-labeled probe to hybridize with a target region of a genome in a liquid system to form a double strand, and a streptavidin-coated magnetic bead is used to adsorb the biotin-carrying molecule, so as to capture the target hybridized with the probe. The captured target library can be sequenced by a sequencer. or sequencing analysis.

[0122] The obtained raw data is subjected to quality control by using fastQC software. The quality control screening steps include: removing the adapter sequence; removing the paired reads with more than 10 N bases in the sequencing data read; removing the paired reads with more than 40% of the low-quality base number of Q<20 in the sequencing data read, to obtain clean data. Then, the clean data is mapped to the common cattle reference genome ARS-UCD 1.2 by using bwa software to generate a bam file. The bam file is sorted by using Picard software, and all PCR repeated reads are removed. Finally, the standard process of GATK software is used to detect SNP and perform genotyping.

[0123] IV. Breed typing and population identification analysis

[0124] Based on specific SNP sites, principal component analysis, population phylogenetic tree construction and admixture analysis are used to classify the target samples. By comparing with the known breed information, the breed attribution judgment and population structure identification of unknown samples are realized.

[0125] Example 3: Performance test experiment of Tibet local cattle SNP chip

[0126] In order to comprehensively evaluate the performance of the designed cattle liquid SNP chip in practical application, the present application carries out a systematic function verification experiment. 36 DNA samples of Tibet local cattle are selected for detection by using a unified chip and standard experimental process. The result shows that the average detection rate of individuals reaches 99.01%, indicating that the chip has good detection ability in the target species.

[0127] In order to evaluate the species specificity of the chip, the same detection is carried out by using horse DNA samples. The result shows that the proportion of effective typing sites is only 0.02%, which is significantly lower than that of cattle samples, verifying that the chip has high species recognition specificity.

[0128] In addition, in order to detect the repeatability of the chip, 2 cattle are randomly selected from the same batch of samples for independent repeated detection. The result shows that the average repeatability reaches 99.04%.

[0129] The results fully demonstrate that the SNP chip has high detection rate, high accuracy, good specificity and high repeatability in the detection of Tibetan local cattle, and can be widely applied to precise identification and typing analysis of bovine germplasm resources.

[0130] The embodiments are only used to explain the present application, and are not used to limit the present application, and those skilled in the art can make modifications to the embodiments without creative contribution, as long as the modifications are within the scope of the present application.

Claims

1. A liquid phase chip for identifying Tibetan cattle species, characterized by: The detection targets of the liquid phase chip include 7,829 SNP sites located on the common cattle reference genome ARS-UCD 1.

2. The position information of the SNP sites is shown in Table 3 of the specification.

2. The liquid phase chip according to claim 1, characterized in that The Tibetan local cattle species include but are not limited to Apaijiazha cattle, camelback cattle, Tibetan cattle, Zhangmu cattle and Zayu cattle.

3. The method for designing a liquid phase chip according to claim 1 or 2, wherein: The steps include: S1. Obtain whole genome resequencing data, perform quality control processing, and obtain Tibetan cattle-specific SNP sites; S2. Obtain 100K chip data and process them to obtain SNP sites specifically associated with Tibetan local cattle breeds; S3. Integrate the specific SNP sites obtained from resequencing and chip data to screen out 7,829 SNP sites; S4. Design probes based on the 7,829 SNP sites screened out and prepare liquid phase microarrays.

4. The method for designing a liquid phase chip according to claim 3, wherein: The integration method includes one or more of removing duplicate sites and eliminating sites with a detection rate lower than 90% through actual sample detection.

5. The design method according to claim 3, characterized in that: Step S1 includes the following steps: S11. Sample Collection Collect biological samples from Apaijiaza cattle, Camelback cattle, Tibetan cattle, and Zhangmu cattle, and extract DNA; S12. Library construction and sequencing Resequencing the DNA sample obtained in step S11 to obtain original sequencing data; S13. Specific site screening The original sequencing data were subjected to quality control, SNP detection and specificity screening to obtain the Tibetan cattle-specific SNP positions.

6. The design method according to claim 5, characterized in that: In step S11 , the biological sample is collected in the form of blood, tissue, semen, hair follicles or feces.

7. The design method according to claim 6, characterized in that: The blood was obtained from the caudal vein or the jugular vein.

8. The design method according to claim 5, characterized in that: In step S13, the quality control method includes one or more of using Trimmomatic v0.38 to remove low-quality sequences and adapter sequences, aligning to the common bovine reference genome ARS-UCD1.2 to remove duplicates, using Picard v2.18.2 to sort, and removing PCR duplicates.

9. The design method according to claim 5 or 8, characterized in that: In step S13, the method for specific screening quality control includes using an allele frequency difference threshold of 90%-10% to screen sites with significant differences between different varieties, and combining gene function annotation for analysis and performing one or more of minimum allele frequency filtering on the sites.

10. The design method according to claim 3, characterized in that: Step S2 includes the following steps: S21. Sample collection and chip testing DNA from Apaijiaza cattle, Camelback cattle, Tibetan cattle, Zhangmu cattle, Plateau Holstein cattle, Plateau Jersey cattle, Yunnan Zebu cattle, and Brahman cattle was used, and 770K chip data of American Brahman cattle was obtained from the WIDDE database. The reference genome version was converted to ARS-UCD 1.

2. S22. Chip data processing Perform genotyping, quality control, and merging of all sample DNA obtained in step S21 to obtain quality-controlled genotype data; S23. Population genetic analysis Principal component analysis was performed on the quality-controlled genotype data to construct a phylogenetic tree, infer population composition, and obtain SNP sites specifically associated with Tibetan local cattle breeds.

11. The use of the liquid phase chip according to claim 1 or 2, characterized in that: Including the application in Tibetan cattle genotyping, the application in Tibetan cattle species identification, or the application in Tibetan cattle breeding.

12. A method for identifying Tibetan endemic cattle species using the liquid phase chip according to claim 1 or 2, characterized in that: The steps include: Step 1: Extract DNA from the sample to be tested Extract the bovine DNA sample to be tested, test the purity and integrity of the extracted DNA sample, and accurately quantify the DNA concentration; Step 2: Build the library After constructing the DNA library, Qubit2.0 was used for preliminary quantification, and the effective concentration of the library was accurately quantified using qPCR to ensure library quality; Step S3: sequencing A biotin-labeled probe is used to hybridize with the target region of the genome in a liquid system to form a double-stranded sequence. The biotin-carrying molecules are adsorbed by streptavidin-coated magnetic beads, thereby capturing the target hybridized with the probe and obtaining the original sequencing data. FastQC software is used for quality control to finally obtain the genotype of the SNP in the sample to be tested. Step 4: Variety typing and population identification analysis Based on specific SNP sites, principal component analysis, population evolutionary tree construction and Admixture analysis are used to classify the target samples into varieties; by comparing with known variety information, the variety affiliation and population structure identification of unknown samples are achieved.

13. The method according to claim 12, characterized in that The quality control screening step includes: removing adapter sequences; eliminating paired reads with an N base content of more than 10 in the sequencing data reads; removing paired reads in which the number of low-quality bases with Q < 20 accounts for more than 40% of the read length in the sequencing data reads to obtain clean data; then using bwa software to map the clean data to the common bovine reference genome ARS-UCD 1.2 to generate a bam file, then using Picard software to sort the bam file, and deduplicating all PCR duplicate reads, and finally using the standard process of GATK software to detect SNPs and perform genotyping.

Citation Information

Patent Citations

  • Cattle 12K SV liquid phase chip and design method and application thereof

    CN116144794A

  • Paternity test method and system based on liquid-phase biochip

    CN117316272A

  • Cow plateau adaptive breeding 10K liquid phase chip and application

    CN117701722A

  • Genomic selection (GS) breeding chip of huaxi cattle and use thereof

    US20240043912A1

  • Method for screening molecular marker of cattle adapting to high altitude hypoxia and application thereof

    WO2020206896A1

Cited By

  • Gene chip for bolete species identification and application thereof

    CN121183032A