SNP and InDel site combination for identifying inner mongolia cattle breeds, 5k targeted capture gene chip and application thereof
By developing SNP and InDel locus combinations and 5K targeted capture gene chips for Inner Mongolian cattle breed identification, the problem of efficient and low-cost genotyping for Inner Mongolian cattle breed identification has been solved, achieving efficient and accurate breed identification and supporting the development of the Inner Mongolian cattle industry and the protection of germplasm resources.
Patent Information
- Application Number
- CN202510548146.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2045-04-28
AI Technical Summary
Existing technologies are insufficient for efficient and low-cost genotyping and breed identification of Inner Mongolian cattle breeds, and there is a lack of effective gene chip tools.
To develop a combination of SNP and InDel loci and a 5K targeted capture gene chip for breed identification of Inner Mongolian cattle, the SNP and InDel loci for breed identification were screened by whole-genome resequencing, targeted capture probes were designed, and gene chips were constructed to achieve accurate breed identification of Mongolian cattle, Sanhe cattle, and Sunite cattle.
It has achieved efficient, accurate, and low-cost genotyping for cattle breed identification in Inner Mongolia, providing technical support for promoting the identification and evaluation of cattle germplasm resources, and supporting industrial development and policy support.
Smart Images

Figure CN120158524B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of bioinformatics and molecular breeding technology. More specifically, it relates to SNP and InDel locus combinations for the identification of Inner Mongolian cattle breeds, 5K targeted capture gene chips, and their applications. Background Technology
[0002] Inner Mongolia boasts a diverse range of cattle breeds, including Mongolian cattle, Sanhe cattle, Simmental cattle, Charolais cattle, Sunite cattle, and Grassland Red cattle. Each breed possesses unique characteristics, adapting to the diverse natural environments and livestock needs of Inner Mongolia. Mongolian cattle are robust and hardy, exhibiting strong adaptability, resistance to wind and sand, drought and cold tolerance, tolerance to roughage and extensive management, and excellent meat quality. They maintain good vitality and reproductive capacity even in challenging environments, making them an important and superior cattle breed in northern China, used for milk, meat, and draft purposes. Other breeds, such as Simmental cattle, are widely popular due to their good adaptability and excellent meat quality. Breed identification helps clarify the genetic characteristics and lineage of each breed, providing a scientific basis for the protection and utilization of germplasm resources. This is crucial for ensuring breed diversity and the sustainability of genetic resources. Breed identification clarifies the uses and value of each breed, providing guidance for cattle selection, mating, and sales. This contributes to the high-quality development of Inner Mongolia's cattle industry, enhancing its competitiveness and economic benefits. In conclusion, Inner Mongolian cattle are diverse in species, each with its own unique characteristics. Breed identification is of great significance for protecting germplasm resources, promoting industrial development, and enjoying policy support.
[0003] Gene chips can accurately genotype target sites, and are a highly efficient and low-cost genotyping method. Therefore, it is extremely necessary to develop a targeted capture gene chip for the identification of Inner Mongolian cattle breeds and to detect and analyze the lineage of three breeds, including Mongolian cattle, Sanhe cattle and Sunite cattle. Summary of the Invention
[0004] One objective of this invention is to provide a combination of SNP and InDel loci and a 5K targeted capture gene chip for the identification of Inner Mongolian cattle breeds. By performing whole-genome sequencing on different Inner Mongolian cattle breed populations, breed identification SNPs and InDel loci are screened, and targeted capture probes are designed based on these loci to construct a gene chip, so as to accurately complete the breed identification of Mongolian cattle, Sanhe cattle and Sunite cattle.
[0005] The second objective of this invention is to provide an application of the aforementioned 5K targeted capture gene chip.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] This invention first provides a set of SNP and InDel loci combinations for the identification of Inner Mongolian cattle breeds. The SNP and InDel loci combinations include 5,000 SNP and InDel loci. The physical location information of the SNP and InDel loci in the bovine reference genome with version number ARS-UCD1.2 is shown in Table 1.
[0008] The present invention further provides a 5K targeted capture gene chip for the identification of Inner Mongolian cattle breeds, wherein the 5K targeted capture gene chip includes probes that target the combination of the above-mentioned SNP and InDel sites.
[0009] The present invention further provides a method for synthesizing the above-mentioned 5K targeted capture gene chip, the method comprising the following steps:
[0010] S1. Perform whole-genome resequencing on samples of different breeds of Inner Mongolian cattle to obtain high-quality SNP and InDel sites of the Inner Mongolian cattle;
[0011] S2. Screen purebred Inner Mongolian cattle samples of different breeds and construct a reference library containing purebred Inner Mongolian cattle samples of different breeds.
[0012] S3. Based on the reference library and high-quality SNP and InDel sites, a set of candidate SNPs and InDel sites is selected.
[0013] S4. Evaluate and score the selected candidate SNP and InDel site sets to determine the likelihood of successful probe design for SNP and InDel sites, and include the SNP and InDel sites with successfully designed probes into the gene chip site set.
[0014] S5. Probes from the SNP and InDel sites of the gene chip locus set are synthesized to obtain a 5K targeted capture gene chip for the identification of Inner Mongolian cattle breeds.
[0015] In a specific embodiment of the present invention, step S1 includes the following steps:
[0016] S11. Extract genomic DNA from samples of different breeds of Inner Mongolian cattle and perform whole-genome resequencing to obtain raw sequence data;
[0017] S12. Perform quality control on the original sequence data to obtain Clean Reads, and align them to the bovine reference genome with version number ARS-UCD1.2 for SNP and InDel site detection;
[0018] S13. After index filtering and quality control, high-quality SNP and InDel sites of the Inner Mongolia cattle were obtained.
[0019] Furthermore, the quality control standards for the raw sequence data are as follows:
[0020] 1) Remove reads with connectors;
[0021] 2) When the N content in a sequencing read exceeds 1% of the total number of bases in that read, the paired read is removed;
[0022] 3) If the number of low-quality (Q<=5) bases in a sequencing read exceeds 50% of the total number of bases in that read, remove the paired read.
[0023] Furthermore, the metric filtering includes QD, FS, MQ, SOR, MQRankSum, and ReadPosRankSum, with the following criteria:
[0024] 1)hardfilter_SNP and InDel Expression1="QD<2.0||FS>60.0||MQ<40.0||SOR>3.0";
[0025] 2)hardfilter_SNP and InDel Expression2="MQRankSum<-12.5||ReadPosRankSum<-8.0";
[0026] 3)hardfilter_indelExpression1="QD<2.0||FS>200.0||SOR>10.0";
[0027] 4)hardfilter_indelExpression2="ReadPosRankSum<-20.0".
[0028] Furthermore, the quality control standards are as follows:
[0029] 1) Only autosomal loci are retained;
[0030] 2) Detection rate >= 90%;
[0031] 3) Sample detection rate >= 90%;
[0032] 4) MAF >= 0.05;
[0033] 5) The P-value of the Hawes-Wen equilibrium test is greater than or equal to 10e-6.
[0034] In a specific embodiment of the present invention, when constructing the reference library in step S2, individual likelihood values are used to calculate and screen purebred Inner Mongolian cattle samples of different breeds.
[0035] In a specific embodiment of the present invention, the candidate SNP and InDel site set to be selected in S3 includes functional SNPs and InDel sites for variety identification and background SNPs and InDel markers for variety identification; wherein, the functional SNPs and InDel sites are obtained by performing population differentiation index analysis among varieties based on a reference library and high-quality SNPs and InDel sites to obtain the Fst value of each site, and SNPs and InDel sites with higher Fst values are selected as functional SNPs and InDel sites for variety identification; the background SNPs and InDel markers are selected based on high-quality SNPs and InDel sites and annotated as stopgain / loss, alternative splicing sites, and non-synonymous mutations as background SNPs and InDel sites for variety identification.
[0036] In a specific embodiment of the present invention, the S4 evaluation score is evaluated based on indicators such as the specificity, complexity, and GC content of the upstream and downstream sequences of the target SNP and InDel sites. Priority is given to placing the target SNP and InDel sites in the middle of the probe, and the designed probe is 120bp in length.
[0037] This invention further provides the application of the above-mentioned SNP and InDel site combinations or the above-mentioned gene chips in the genotyping, pedigree analysis or breed identification of Inner Mongolia cattle.
[0038] The present invention further provides a method for identifying Inner Mongolian cattle breeds, comprising the following steps:
[0039] S1. Based on the genomic DNA of the sample to be tested, construct a whole genome library and hybridize it with the probes in the 5K targeted capture gene chip to construct a hybridization capture library.
[0040] S2. Analyze the data after sequencing the hybridization capture library to obtain the genotyping results of the target SNPs and InDel sites of the sample to be tested;
[0041] S3. Use the genotyping results to infer the lineage of the sample to be tested and complete the breed identification of Inner Mongolia cattle.
[0042] In a specific embodiment of the present invention, step S2 includes the following steps:
[0043] S21. Filter the data after sequencing the hybridization capture library to obtain Clean Reads;
[0044] S22. Align the Clean Reads to the bovine reference genome with version number ARS-UCD1.2 to obtain the alignment results of the Clean Reads on the reference genome;
[0045] S23. Based on the alignment results of Clean Reads with the reference genome, perform genotyping of the target SNPs and InDel sites to obtain the genotyping results of the target SNPs and InDel sites of the sample to be tested.
[0046] In a specific embodiment of the present invention, step S3, which uses genotyping results to infer the ancestral composition of the sample to be tested, is based on the reference library constructed above, and the ancestral composition of the sample to be tested is analyzed using the admixture software.
[0047] In this invention, the Inner Mongolian cattle include Mongolian cattle, Sanhe cattle, and Sunite cattle.
[0048] The beneficial effects of this invention are as follows:
[0049] This invention obtained the whole-genome genetic information of 149 individuals from three Inner Mongolian cattle breeds—Mongolian, Sanhe, and Sunite—using whole-genome resequencing technology. Functional SNPs and InDel loci, as well as background SNPs and InDel loci, for breed identification were identified through screening. Based on these loci, targeted capture probes were designed, leading to the development of a 5K targeted capture gene chip for Inner Mongolian cattle breed identification. This 5K targeted capture gene chip uses targeted probes to capture and performs genotyping using high-throughput sequencing technology, offering advantages such as low cost, high efficiency, and high accuracy. Furthermore, it can utilize genotypic data to achieve breed identification for Mongolian, Sanhe, and Sunite cattle, providing crucial technical support and specific methods for the identification and evaluation of Inner Mongolian cattle germplasm resources. Attached Figure Description
[0050] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0051] Figure 1 This is the result of principal component analysis (PCA) of all samples in the reference library of this invention.
[0052] Figure 2 This is a distribution diagram of the number of SNP and InDel sites on different chromosomes in the 5K targeted capture gene chip used for the identification of Inner Mongolian cattle breeds in this invention.
[0053] Figure 3 This is a distribution diagram of SNPs and InDel sites MAF in the 5K targeted capture gene chip used for the identification of Inner Mongolian cattle breeds in this invention.
[0054] Figure 4 This invention illustrates the distribution of SNP and InDel sites in different functional elements within a 5K targeted capture gene chip used for the identification of Inner Mongolian cattle breeds. Detailed Implementation
[0055] To more clearly illustrate the present invention, the following description, in conjunction with preferred embodiments and accompanying drawings, further explains the invention. Similar components in the drawings are indicated by the same reference numerals. Those skilled in the art should understand that the specific description below is illustrative rather than restrictive and should not be construed as limiting the scope of protection of the present invention.
[0056] Example 1: Development of Identification Chip Sites and Probe Design for Inner Mongolia Cattle 5K Breed
[0057] I. Selection of Inner Mongolian Cattle Breeds
[0058] Three Inner Mongolian cattle breeds were selected for population resequencing. The breed names and numbers are shown in Table 2.
[0059] Table 2. Inner Mongolian cattle breed information used for population resequencing.
[0060] Mongolian cattle MGN 49 Sanhe cattle SHN 50 Sunite cattle SNT 50
[0061] II. Sample Sequencing and Data Quality Control
[0062] 1. DNA extraction and quality control
[0063] Genomic DNA (gDNA) was extracted from whole blood using a magnetic bead DNA extraction kit (CW2361S). Genomic integrity was assessed using a 1.5% agarose gel electrophoresis system. DNA purity was determined using a NanoDrop 2000 nucleic acid and protein analyzer, measuring the A260 / A280 ratio. Precise quantification of gDNA samples was performed using a Qubit 2.0 analyzer.
[0064] 2. Sequencing library construction and sequencing
[0065] use The following are the specific steps for constructing sequencing libraries using the Universal DNA Library Construction Kit (for MGI):
[0066] gDNA is fragmented into 300-350 bp fragments using enzymes. After end repair, A-tailing, and ligation with sequencing adapters, DNA fragments of approximately 300-350 bp are screened, amplified by PCR, and the PCR products are purified again to obtain the sequencing library. After the sequencing library is constructed, preliminary quantification is performed using Qubit 2.0, followed by detection of the inserted fragments in the sequencing library. If the results meet expectations, sequencing is performed, i.e., library testing.
[0067] After the library passes inspection, pooling is performed based on the effective concentration of the sequencing library and the required amount of data for sequencing. The DNBSEQ-T7 sequencer is then used with the PE150 sequencing strategy selected for sequencing. The raw image data obtained from sequencing is converted into raw sequence data (raw reads) using base calling software and stored in FASTQ file format.
[0068] 3. SNP and InDel detection and filtering
[0069] FastP is used to perform a series of quality control (QC) checks on raw reads to obtain Clean Reads. The quality control standards are as follows:
[0070] 1) Remove reads with adapters;
[0071] 2) When the N content in a sequencing read exceeds 1% of the total number of bases in that read, the paired read is removed;
[0072] 3) If the number of low-quality (Q<=5) bases in a sequencing read exceeds 50% of the total number of bases in that read, remove the paired read.
[0073] Clean Reads were aligned to the bovine reference genome (version ARS-UCD1.2) using BWA 07.17 (men) software. Bam files were generated by sorting and indexing using samtools 1.7. The Bam files were then deduplicated using modules included in GATK 4.1.8.0 software. SNP, InDel, and INEDL detection were performed using GATK 4.1.8.0. Finally, the VariantFiltration module was used for rigorous index filtering of SNPs, InDels, and INEDL, resulting in filtered SNP and InDel loci. The filtering indexes included QD, FS, MQ, SOR, MQRankSum, and ReadPosRankSum, with the following criteria:
[0074] 1)hardfilter_SNP and InDel Expression1="QD<2.0||FS>60.0||MQ<40.0||SOR>3.0"
[0075] 2)hardfilter_SNP and InDel Expression2="MQRankSum<-12.5||ReadPosRankSum<-8.0"
[0076] 3)hardfilter_indelExpression1="QD<2.0||FS>200.0||SOR>10.0"
[0077] 4)hardfilter_indelExpression2="ReadPosRankSum<-20.0"
[0078] 4. SNP and InDel quality control
[0079] To retain high-quality SNPs and InDel sites, quality control was performed on the filtered SNPs and InDel sites before analysis. After quality control, 13,582,392 SNPs and InDel sites remained. The quality control criteria are as follows:
[0080] 1) Only autosomal loci are retained
[0081] 2) Detection rate >= 90%
[0082] 3) Sample detection rate >= 90%
[0083] 4) MAF >= 0.05
[0084] 5) The P-value of the Hawes-Wen equilibrium test is greater than or equal to 10e-6.
[0085] III. Building a Reference Library
[0086] The principle for selecting samples for the reference library was that the samples had high purity and could represent the genetic diversity of the breed. Purebred samples were identified using individual likelihood values. After screening 149 samples, 128 purebred samples were obtained, of which 122 were used to construct the reference library. The number of individuals for each breed in the reference library is shown in Table 3. The remaining 6 samples (2 each of Mongolian cattle, Sanhe cattle, and Sunite cattle) served as positive controls to verify the accuracy of the breed identification chip. Dimensionality reduction and clustering analysis were performed on all purebred samples in the reference library using PCA. The results are shown below. Figure 1 As shown, the three breeds of Mongolian cattle (MGN), Sanhe cattle (SHN), and Sunite cattle (SNT) can be distinguished relatively well.
[0087] Table 3 shows the purebred samples used to construct the reference library.
[0088] Mongolian cattle 38 Sanhe cattle 41 Sunite cattle 43
[0089] IV. SNP and InDel site screening for variety identification
[0090] Based on the reference library samples and high-quality SNPs and InDel sites after quality control, the population differentiation index (Fst) analysis was performed among the three varieties to obtain the Fst value of each site. Considering that some sites will be lost in subsequent probe design, it is necessary to select some additional sites. Finally, 7000 sites with high Fst values (top 7000) were selected as functional SNPs and InDel sites for variety identification.
[0091] To fill sparse regions in the genome, sites annotated as stopgain / loss, alternative splicing sites, and non-synonymous mutations were selected from high-quality SNPs and InDel sites after quality control as background SNPs and InDel sites for variety identification.
[0092] The aforementioned functional SNPs and InDel sites, along with background SNPs and InDel sites, are used as the proposed set of candidate SNPs and InDel sites.
[0093] V. Probe Design for SNP and InDel Sites in Variety Identification
[0094] The selected candidate SNPs and InDel sites were evaluated and scored, primarily assessing the specificity, complexity, and GC content of the upstream and downstream sequences of the target SNPs and InDel sites. Priority was given to placing the target SNPs and InDel sites in the middle of the probe, and the designed probes were 120 bp in length. The likelihood of successful SNP and InDel site targeting and capture probe design was determined based on the evaluation scores. Ultimately, 4690 functional SNPs and InDel sites and 310 background SNPs and InDel sites were successfully designed (see Table 4).
[0095] Table 4. Composition of SNPs and InDel sites in successfully designed targeted capture probes.
[0096] functional sites 4690 Background sites 310
[0097] Finally, 5,000 SNPs and InDel sites that were successfully designed with targeted capture probes were included in the final microarray site set, with version number ARS-UCD1.2 as the bovine reference genome. The physical location information of the SNPs and InDel sites is shown in Table 1.
[0098] Table 1. Physical location information of SNP and InDel sites.
[0099]
[0100]
[0101]
[0102]
[0103]
[0104]
[0105]
[0106]
[0107]
[0108]
[0109]
[0110]
[0111]
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126]
[0127]
[0128]
[0129]
[0130]
[0131]
[0132]
[0133]
[0134]
[0135]
[0136]
[0137]
[0138]
[0139]
[0140]
[0141]
[0142]
[0143]
[0144]
[0145]
[0146] Note: Taking SNP locus number 1 as an example, in 1:1103859, "1" represents the chromosome where the SNP locus is located, 1103859 indicates the physical location of the SNP locus on the chromosome, G represents the locus information before the mutation, and C represents the locus information after the mutation; taking InDel locus number 4 as an example, in 1:2230883, "1" represents the chromosome where the InDel locus is located, indicating that the AA sequence is deleted after position 2230883 on chromosome 1; taking InDel locus number 19 as an example, in 1:12164703, "1" represents the chromosome where the InDel locus is located, indicating that the AT sequence is inserted after position 12164703 on chromosome 1; in the case of mutation, "..." is included. underline "like A>G, represents a background SNP site, and the others are functional SNPs or InDel sites.
[0147] The target capture probes for the above 5000 SNP and InDel sites were synthesized to obtain the 5K target capture gene chip for the identification of Inner Mongolia cattle breeds.
[0148] The distribution of all SNPs and InDel sites on chromosomes is as follows: Figure 2 As shown in the figure, the SNP and InDel sites of this invention are distributed on different chromosomes. Therefore, the sites captured by the targeted capture probe designed according to this invention can cover the entire genome.
[0149] The distribution of minor allele frequencies (MAF) for all SNPs and InDel loci is shown in the figure below. Figure 3 As shown in the figure, the MAF values are mainly distributed between 0.4 and 0.5.
[0150] The distribution map of all SNPs and InDel sites in different genomic functional elements is shown below. Figure 4 As shown in the figure, all SNPs and InDel sites are more abundant in intergenic regions and introns.
[0151] VI. Effect Verification
[0152] For the six positive control purebred samples of Mongolian cattle, Sanhe cattle, and Sunite cattle identified in "III. Construction of the Reference Library," genotyping results for the aforementioned 5000 loci were extracted from their sequencing data. Admixture software was used to infer the pedigree composition of these six control samples, and the inference results are shown in Table 5. In Table 5, each row represents the proportion of each breed's pedigree in a sample, summed to 1, and each column represents a single breed's pedigree. A higher value indicates a higher proportion of the corresponding pedigree in the sample. The table shows that the pedigree analysis results are basically consistent with the breed information of the sample, demonstrating the accuracy of the selected pedigree identification loci and the reliability of the 5K targeted capture gene chip development for Inner Mongolian cattle breed identification.
[0153] Table 5. Lineage composition of positive control samples
[0154]
[0155] Note: The variety is the same as the variety of the positive control sample;
[0156] The sample number is an identification code used to identify a sample;
[0157] The Sanhe cattle lineage, Sunite cattle lineage, and Mongolian cattle lineage represent the percentages of the corresponding breed lineages in the tested samples.
[0158] Example 2: Sequencing of a 5K targeted capture gene chip for Inner Mongolian cattle breed identification
[0159] The principle of targeted capture gene chip sequencing is as follows: Based on the principle of DNA complementarity, probes covering each target SNP and InDel site are designed. Biotin-modified probes hybridize with sequencing libraries containing target SNPs and InDel sites to form double strands. Streptavidin-coated magnetic beads are used to adsorb the biotin-modified probes, thereby capturing the hybridization capture library. Finally, the hybridization capture library is eluted, enriched, and sequenced to obtain the genotyping results for the target SNPs and InDel sites. The specific steps are as follows:
[0160] I. Constructing a Hybrid Capture Library
[0161] 1. Genomic DNA extraction
[0162] DNA was extracted from 48 Inner Mongolian cattle samples using a magnetic bead DNA extraction kit (CW2361S). The extracted DNA samples were analyzed using the following three methods:
[0163] 1) Agarose gel electrophoresis analysis: to detect the degree of DNA degradation; to determine whether there is RNA contamination;
[0164] 2) Nanodrop: Detects the purity of DNA samples (OD260 / OD280 ratio);
[0165] 3) Qubit: Precisely quantifies DNA concentration.
[0166] 2. Genomic DNA fragmentation, end repair, and A addition
[0167] Take 300 ng of DNA sample and add 4 μL of Smearase Buffer and 2 μL of Smearase Enzymes, for a total volume of 24 μL. Place the sample in a PCR instrument and run the following program: 4℃ for 1 min → 30℃ for 10 min → 72℃ for 20 min → store at 4℃. Use streptavidin-coated magnetic beads to screen the DNA fragments after fragmentation, end repair, and A addition, removing excessively large and small fragments to concentrate the DNA fragments in the 200-300 bp range. The final fragment screening volume is 25 μL, which is then stored in a 96-well PCR plate.
[0168] 3. Connector connection and enrichment
[0169] Add 5 μL of CAGT Universal Adapters and 20 μL of Ligation MasterMix to a 96-well PCR plate, vortex to mix, and briefly centrifuge to collect the reaction solution at the bottom of the tube. Incubate at 20°C for 15 min in a PCR instrument to complete the sequencing adapter ligation and obtain the ligation product. After purifying the ligation product, perform PCR amplification and enrichment. Take 15 μL of the purified ligation product, add 10 μL of CAGT UDI Primer and 25 μL of Equinox Library Amp Mix (2x), mix well, and aspirate 35 μL of the reaction solution into a PCR instrument for library amplification to construct a whole-genome library. Quantify the whole-genome library using Qubit 2.0. At the same time, check whether the main peak of the fragment in the whole-genome library is in the range of 300-500 bp by electrophoresis.
[0170] 4. Chip Hybrid Capture
[0171] Whole-genome libraries from each sample were pooled, concentrated, and then hybridized with biotin-modified targeting probes from a targeting capture gene chip. Fragments of the target region were captured from the pooled library. An elution step removed excess targeting probes, hybridization reagents, and other reagent components. Finally, the target region was enriched by post-hybridization PCR amplification to obtain the hybridization capture library.
[0172] 5. Quality control and sequencing of hybridization capture libraries
[0173] After the hybridization capture library was constructed, it was quantified using Qubi 2.0; simultaneously, electrophoresis was used to check whether the main peak size of the library was within the range of 300-500 bp. The constructed hybridization capture library was sequenced using a DNBSEQ-T7 sequencer to obtain the raw data after sequencing of the hybridization capture library.
[0174] II. Data Analysis
[0175] The post-sequencing analysis workflow for hybrid capture libraries mainly includes data filtering and statistics, alignment with the reference genome, and analysis of target SNPs and InDel sites, detailed below:
[0176] 1. Data filtering and statistics
[0177] After the raw data is processed, it will contain reads with adapters or low quality. Before performing alignment with the reference genome and analysis of target SNPs and InDel sites, FastP is needed to filter the raw data and calculate the amount of data before and after filtering to obtain Clean Reads. The filtering conditions are as follows:
[0178] 1) Remove reads with adapters;
[0179] 2) When the N content in a sequencing read exceeds 10% of the total number of bases in that read, the paired read is removed;
[0180] 3) If the number of low-quality (Q<=5) bases in a sequencing read exceeds 50% of the total bases in that read, remove the paired read.
[0181] The Fastp parameters are set to -u 50-n 1-q 5-l 30.
[0182] 2. Alignment with reference genome
[0183] Clean Reads were aligned to the bovine reference genome (ARS-UCD version 1.2) using BWA 0.7.17 software. Samtools 1.7 was used to sort and build an index to obtain Bam files. The Bam files were deduplicated using the built-in module of GATK 4.1.8.0 software. Then, based on the Bam files, the sequencing depth, genome coverage, and other information of each sample were statistically analyzed to obtain the alignment results of Clean Reads to the reference genome.
[0184] 3. Target SNP and InDel site analysis
[0185] Based on the alignment results of Clean Reads with the reference genome, the GenotypeGVCFs module in GATK 4.1.8.0 software was used to genotype the target SNPs and InDel sites. Then, minDP=5X filtering and 100% deletion were used to mark the sites that did not meet the requirements as ". / .". The site detection rate and site depth were then statistically analyzed, and ANNOVAR was used to annotate the genotyping results of the target SNPs and InDel sites.
[0186] Based on the sequencing depth of the target SNPs and InDel sites, the sample detection rate and site detection rate were statistically analyzed. Table 6 shows the sample detection rate statistics, with average detection rates of 98.57% (sequencing depth greater than 5×) and 98.00% (sequencing depth greater than 10×), respectively. Table 7 shows the site detection rate statistics, with site detection rates between 90% and 100% accounting for 96.64% (sequencing depth greater than 5×) and 95.78% (sequencing depth greater than 10×), respectively, indicating good chip detection performance.
[0187] Table 6. Statistical analysis of sample detection rate
[0188]
[0189]
[0190] Note: Sample ID is the sample name;
[0191] dp0 is the number of sites with a sequencing depth greater than 0×;
[0192] dp0% is the proportion of sites with a sequencing depth greater than 0× out of all sites;
[0193] The same applies to other column headers.
[0194] Table 7. Detection rate statistics for loci
[0195] (-0.001,10.0] 0 0 2 0.04 32 0.64 45 0.9 54 1.08 (10.0,20.0] 0 0 9 0.18 11 0.22 13 0.26 10 0.2 (20.0,30.0] 0 0 20 0.4 13 0.26 7 0.14 6 0.12 (30.0,40.0] 1 0.02 9 0.18 6 0.12 5 0.1 8 0.16 (40.0,50.0] 0 0 11 0.22 13 0.26 14 0.28 24 0.48 (50.0,60.0] 1 0.02 17 0.34 11 0.22 22 0.44 16 0.32 (60.0,70.0] 7 0.14 20 0.4 35 0.7 34 0.68 33 0.66 (70.0,80.0] 22 0.44 37 0.74 36 0.72 39 0.78 46 0.92 (80.0,90.0] 53 1.06 43 0.86 54 1.08 58 1.16 62 1.24 (90.0,100.0] 4916 98.32 4832 96.64 4789 95.78 4763 95.26 4741 94.82
[0196] Note: The site detection rate % is a range of site detection rates;
[0197] dp0 is the number of sites detected when the depth is greater than 0;
[0198] dp0% is the percentage of detected sites when the depth is greater than 0;
[0199] The same applies to other column headers.
[0200] Example 3: Application of 5K Targeted Capture Gene Chip in Inner Mongolia Cattle Breed Identification
[0201] Using the 5K targeted capture gene chip sequencing method described in Example 2 for Inner Mongolia cattle breed identification, 10 samples were sequenced to complete the genotyping results of the target SNP and InDel loci. The genotyping results of the target SNP and InDel loci for each sample were obtained. These samples were randomly selected from purebred and hybrid populations of Mongolian cattle, Sanhe cattle, Sunite cattle and their hybrids.
[0202] Based on the reference library constructed in Example 1, the pedigree composition of 10 test samples was analyzed using Admixture software. The pedigree content of each sample is shown in Table 8. Each row represents the proportion of each breed's pedigree in a sample (the sum is 1), and each column represents a breed's pedigree. It can be seen that Mongolian cattle contain Sunite or Sanhe cattle pedigree; Sanhe cattle contain Sunite cattle pedigree; and Sunite cattle contain Sanhe or Mongolian cattle pedigree. This may be related to long-term crossbreeding.
[0203] Table 8. Lineage composition of the samples to be tested
[0204]
[0205] Note: The sample number is an identification code used to identify the sample;
[0206] The Sanhe cattle lineage, Sunite cattle lineage, and Mongolian cattle lineage represent the proportions of the corresponding breed lineages in the samples being tested.
[0207] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. For those skilled in the art, other variations or modifications can be made based on the above description. It is impossible to exhaustively list all the implementation methods here. All obvious variations or modifications derived from the technical solutions of the present invention are still within the protection scope of the present invention.
Claims
1. The application of a set of SNP and InDel locus combinations in pedigree analysis or breed identification of Inner Mongolian cattle, characterized in that, The SNP and InDel site combination includes 5000 SNP and InDel sites, and the physical location information of the SNP and InDel sites in the bovine reference genome with version number ARS-UCD1.2 is shown in the table below. The Inner Mongolian cattle mentioned refer to Mongolian cattle, Sanhe cattle, and Sunite cattle; 。 2. A 5K targeted capture gene chip for identifying Inner Mongolian cattle breeds, characterized in that, The 5K targeted capture gene chip includes probes that target the combination of SNP and InDel sites as described in claim 1; The Inner Mongolian cattle mentioned refer to Mongolian cattle, Sanhe cattle, and Sunite cattle.
3. The application of the 5K targeted capture gene chip as described in claim 2 in pedigree analysis or breed identification of Inner Mongolian cattle, characterized in that, The Inner Mongolian cattle mentioned refer to Mongolian cattle, Sanhe cattle, and Sunite cattle.
4. A method for identifying Inner Mongolian cattle breeds, characterized in that, Includes the following steps: S1. Based on the genomic DNA of the sample to be tested, construct a whole genome library and perform a hybridization reaction with the probe in the 5K targeted capture gene chip as described in claim 2 to construct a hybridization capture library; S2. Analyze the data after sequencing the hybridization capture library to obtain the genotyping results of the target SNPs and InDel sites of the sample to be tested; S3. Use the genotyping results to infer the lineage of the sample to be tested and complete the breed identification of Inner Mongolia cattle; The Inner Mongolian cattle mentioned refer to Mongolian cattle, Sanhe cattle, and Sunite cattle.
Citation Information
Patent Citations
SNP molecular marker combination for identifying different groups of Mongolian cattle
CN110396548A
Liquid phase chip for beef cattle bloodline identification, design method and application
CN118957105A