Genome-wide liquid chip for holstein cows in south china
By designing a whole-genome liquid-phase chip for Holstein dairy cattle in southern China, the problem of insufficient representativeness of existing chips in southern applications has been solved, enabling efficient genetic evaluation and breeding screening, and improving breeding results and production performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN ACADEMY OF AGRI SCI
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-29
AI Technical Summary
Existing gene chip designs are based on northern populations, resulting in insufficient representativeness of the reference population and weak regional specificity of loci when applied to Holstein dairy cattle populations in southern China. This leads to a decrease in the accuracy of genetic assessment, affecting breeding effectiveness and production performance.
A whole-genome liquid-phase chip was developed. The genotyping targets of the probe set were based on whole-genome resequencing and multi-omics data of Holstein dairy cattle in southern China. The designed loci focused on key traits adapted to the hot and humid climate of southern China, covering regional adaptability such as heat tolerance and disease resistance, while also taking into account lactation performance and reproductive performance.
It significantly improves the accuracy of genetic assessment in Southern Holstein dairy cattle populations, shortens the breeding cycle, accurately screens out individuals with superior genes, and improves breeding efficiency and production performance.
Smart Images

Figure CN122104930A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of molecular breeding and biotechnology, specifically relating to a whole-genome liquid-phase chip for Holstein dairy cows in southern China. Background Technology
[0002] Due to resource conditions and climate, Holstein dairy farming in my country has long been characterized by a predominance in the north and a smaller presence in the south. The dairy market in southern China is active, with continuously growing demand. Localized dairy farming plays an irreplaceable role in ensuring the supply of fresh milk and the stability of dairy processing in the region. However, the hot and humid climate of southern China poses a severe challenge to the production performance of Holstein dairy cows, becoming one of the main factors limiting key performance indicators such as milk yield and conception rate. Furthermore, due to factors such as suppressed immunity and environmental stress, the incidence of diseases such as mastitis and hoof diseases in dairy cows has increased significantly. This leads to a fundamental difference in breeding goals between southern and northern regions, with southern regions focusing more on selecting for regionally adaptable traits such as heat tolerance and disease resistance.
[0003] Gene chips, as one of the key tools of this technology, are a high-throughput SNP genotyping technology. The principle is to immobilize a large number of nucleic acid probes designed for specific SNP sites on a solid-phase carrier or place them in a liquid-phase reaction system. Through sample DNA hybridization with probes and signal detection, rapid and accurate genotyping of tens of thousands to hundreds of thousands of SNP sites in the whole genome can be achieved.
[0004] Currently, although there are reports on microchips for improving dairy cow productivity, longevity, heat stress, udder health, and high-altitude adaptation, the design of existing microchips is largely based on northern populations due to the fact that my country's main dairy-producing areas are located in the north. Breeding objectives are also centered on northern production performance, limiting their application in southern populations. Specifically, this manifests as insufficient regional representativeness of the reference population, weak regional specificity of the loci, and low correlation with actual production traits. More importantly, if breeding tools and strategies focused on northern production performance are used to select dairy cows in the south, the selected superior alleles may not fully cover the key genetic loci adapted to the hot and humid climate of the south. Offspring bred in this way may have their genetic potential fully expressed in the northern environment, but under the special climatic conditions of the south, they may experience increased maladaptation, leading to lower-than-expected production performance and frequent health problems. This not only affects breeding effectiveness but also poses a potential economic risk to actual production. Summary of the Invention
[0005] To address the problems of insufficient representativeness of reference populations, weak regional specificity of loci, and decreased accuracy of genetic assessment when existing chips are designed based on northern populations and applied to Holstein dairy cow populations in southern China, this invention provides a whole-genome liquid phase chip for Holstein dairy cows in southern China.
[0006] The whole-genome liquid-phase chip provided by this invention uses probe sets for genotyping based on whole-genome resequencing and multi-omics data screening of Holstein dairy cattle populations in southern China (represented by Hubei Province). The site design of this whole-genome liquid-phase chip is highly focused on key traits adapted to the hot and humid climate of southern China, systematically covering genomic regions closely related to regional adaptability such as heat tolerance and disease resistance, while also considering lactation and reproductive performance.
[0007] The technical solution provided by this invention is as follows: In a first aspect, the present invention provides a whole-genome liquid-phase chip, including a probe set, wherein the genotyping target of the probe set includes all SNP sites listed in Table 2.
[0008] Secondly, this invention provides the application of the above-mentioned whole-genome liquid phase chip in molecular marker-assisted selection of Holstein dairy cows in southern China.
[0009] In conjunction with the second aspect of the invention, in some embodiments, the application is for at least one of the following purposes: (1) A genomic genetic assessment of lactation, reproduction, health and adaptability traits of Holstein dairy cows in southern China; (2) Locate genes or genomic regions associated with important traits in Holstein dairy cows in southern China; (3) Analyze the genetic diversity, kinship or population structure of Holstein dairy cattle populations in southern China.
[0010] In conjunction with the second aspect of the invention, in some embodiments, the application specifically refers to marker-assisted selection of Holstein dairy cows in southern China.
[0011] In conjunction with the second aspect of the present invention, in some embodiments, the southern region of China is Hubei Province.
[0012] Thirdly, the present invention provides a combination of SNP sites, including all SNP sites listed in Table 2.
[0013] Fourthly, the present invention provides a probe set, wherein the genotyping target of the probe set includes all SNP sites listed in Table 2.
[0014] Fifthly, the present invention provides applications of the above-mentioned SNP site combinations and probe sets, wherein the applications include at least one of the following: (1) Prepare a genotyping chip or kit for breeding Holstein dairy cows in southern China for high yield, disease resistance and heat resistance traits; (2) Assess the genetic diversity and population structure of Holstein dairy cattle populations in southern China; (3) Select high-yield, disease-resistant, and heat-resistant traits in Holstein dairy cattle in southern China through marker-assisted selection or genomic selection; (4) Marker-assisted selection of Holstein dairy cows in southern China; (5) Detect the genotypes of Holstein dairy cows in southern China that are associated with high yield, disease resistance and heat tolerance.
[0015] In a sixth aspect, the present invention provides a kit comprising the above-described probe set.
[0016] Seventhly, the present invention provides a method for genetic evaluation of Holstein dairy cows in southern China, comprising the following steps: Using the whole genome liquid chip described above, the dairy cows to be tested were genotyped to obtain genotype data; Based on the genotype data, the estimated genomic breeding values of Holstein dairy cows in southern China for lactation, reproduction, health and adaptability traits were calculated using either the best linear unbiased genomic prediction method or the single-step best linear unbiased genomic prediction method. Based on the genotype data and the phenotypic data of the dairy cows to be tested, a mixed linear model was used to perform genome-wide association analysis to locate genomic regions associated with lactation, reproduction, health, and adaptive traits.
[0017] Eighthly, the present invention provides a genetic evaluation system for Holstein dairy cows in southern China, used to implement the above-mentioned method for genetic evaluation of Holstein dairy cows in southern China based on whole-genome resequencing data, the system comprising: processor; A memory on which computer programs are stored; When the processor executes the computer program, it performs the functions of the following modules: The genotype data acquisition module is used to control the above-mentioned whole genome liquid phase chip to perform genotyping on the dairy cows to be tested, and to receive and store the generated genotype data; The genome breeding value calculation module is used to call the genome best linear unbiased prediction or single-step genome best linear unbiased prediction algorithm program to calculate the genome estimated breeding value based on the genotype data. The gene mapping analysis module is used to run a mixed linear model analysis program to perform genome-wide association analysis based on the genotype and phenotypic data.
[0018] Compared with the prior art, the present invention has at least the following beneficial effects: Compared to directly applying general-purpose chips developed based on other geographical populations, this invention can more accurately reflect the genetic characteristics of Holstein dairy cows in southern China, effectively reducing assessment bias caused by differences in genetic background, thereby significantly improving the accuracy and efficiency of breeding in this specific population. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 : Density distribution map of core target sites.
[0021] Figure 2 Detection rate and comparison rate of chip sites in 30 dairy cow samples from southern China.
[0022] Figure 3 Detection rate and comparison rate of chip sites in 40 samples of dairy cows from northern China. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with the embodiments of this invention. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0024] In the terminology of this invention, GBLUP represents best linear unbiased prediction of the genome; SSGBLUP represents best linear unbiased prediction of the genome in a single step.
[0025] To address the limitations of existing microarrays in the application of genetic testing in Holstein dairy cattle populations in southern China, this invention provides a whole-genome liquid-phase microarray comprising a probe set whose genotyping targets include all SNP loci listed in Table 2. This regionally customized microarray significantly improves the accuracy of genetic assessment in southern Holstein dairy cattle.
[0026] This invention also provides the application of the above-mentioned whole-genome liquid-phase chip in molecular marker-assisted selection of Holstein dairy cows in southern China; the application is at least one of the following uses: (1) A genomic genetic assessment of lactation, reproduction, health and adaptability traits of Holstein dairy cows in southern China; (2) Locate genes or genomic regions associated with important traits in Holstein dairy cows in southern China; (3) Analyze the genetic diversity, kinship or population structure of Holstein dairy cattle populations in southern China.
[0027] Example 2 demonstrates that the chip can perform genomic genetic assessment of target traits with high accuracy, comparable to the much more expensive whole-genome sequencing; and shows that the data generated using the chip of this invention can be used to locate genes or regions related to important traits using mature GWAS methods. Example 4 demonstrates that the data from the chip of this invention can be practically applied to kinship analysis and breeding decisions.
[0028] Traditional breeding relies primarily on observing physical appearance and recording production performance (phenotype) for selection. The aforementioned whole-genome liquid microarray is used for marker-assisted selection of Holstein dairy cows in southern China, directly screening at the gene level. Microarray testing can predict a cow's future genetic potential in milk production, heat tolerance, and disease resistance after birth, eliminating the need to wait until adulthood and significantly shortening the breeding cycle. It allows for more accurate selection of individuals carrying superior genes (such as heat tolerance genes and high milk protein genes) as breeding stock, avoiding incorrect selection. By analyzing kinship and genotype, the most suitable bull can be selected for mating cows, ensuring offspring inherit the superior genes from both parents. Preferably, southern China refers to Hubei Province. All SNP loci listed in Table 2 are from the Holstein dairy cow population in Hubei Province, and this microarray is most effective when applied to Holstein dairy cows in Hubei Province.
[0029] This invention also provides a combination of SNP loci including all SNP loci listed in Table 2, and a probe set whose genotyping target includes all SNP loci listed in Table 2. All of these SNP loci are derived from Holstein dairy cattle populations in southern China. The core loci directly target key traits such as heat stress resistance, high humidity tolerance, and disease resistance, while also considering reproductive and milk production performance. Furthermore, the design integrates other functional loci, enabling a more accurate reflection of the genetic characteristics of Holstein dairy cattle in southern China and effectively reducing assessment bias caused by differences in genetic background.
[0030] This invention also provides applications of the above-mentioned SNP site combinations and probe sets, wherein the applications include at least one of the following: (1) Prepare a genotyping chip or kit for breeding Holstein dairy cows in southern China for high yield, disease resistance and heat resistance traits; (2) Assess the genetic diversity and population structure of Holstein dairy cattle populations in southern China; (3) Select high-yield, disease-resistant, and heat-resistant traits in Holstein dairy cattle in southern China through marker-assisted selection or genomic selection; (4) Marker-assisted selection of Holstein dairy cows in southern China; (5) Detect the genotypes of Holstein dairy cows in southern China that are associated with high yield, disease resistance and heat tolerance.
[0031] The present invention also provides a kit comprising the above-described probe set. This kit is a standardized, ready-to-use detection tool that, using the included probe set, can specifically, rapidly, and accurately detect bovine DNA samples, thereby obtaining genotypic information for all SNP loci in a single test.
[0032] This invention also provides a method for genetic evaluation of Holstein dairy cows in southern China, comprising the following steps: Using the whole genome liquid chip described above, the dairy cows to be tested were genotyped to obtain genotype data; Based on the genotype data, the estimated genomic breeding values of Holstein dairy cows in southern China for lactation, reproduction, health and adaptability traits were calculated using the GBLUP or SSGBLUP method. Based on the genotype data and the phenotypic data of the dairy cows to be tested, a mixed linear model was used to perform genome-wide association analysis to locate genomic regions associated with lactation, reproduction, health, and adaptive traits.
[0033] This invention also provides a genetic evaluation system for Holstein dairy cows in southern China, used to implement the above-mentioned method for genetic evaluation of Holstein dairy cows in southern China based on whole-genome resequencing data. The system includes: processor; A memory on which computer programs are stored; When the processor executes the computer program, it performs the functions of the following modules: The genotype data acquisition module is used to control the above-mentioned whole genome liquid phase chip to perform genotyping on the dairy cows to be tested, and to receive and store the generated genotype data; The genome breeding value calculation module is used to call the GBLUP or SSGBLUP algorithm program to calculate the estimated genome breeding value based on the genotype data. The gene mapping analysis module is used to run a mixed linear model analysis program to perform genome-wide association analysis based on the genotype and phenotypic data.
[0034] The genetic assessment methods and systems described above use statistical models such as GBLUP or SSGBLUP to convert genotypic data obtained from microarray analysis into GEBV values for specific traits such as lactation, reproduction, health, and adaptability. By comparing the GEBV values of different individuals, breeders can precisely select individuals with higher genetic potential in heat tolerance, disease resistance, or milk production performance as breeding cows after birth, significantly shortening generation intervals. Simultaneously, this provides objective and quantitative evidence for breeding selection, avoiding subjective judgments based solely on appearance or pedigree.
[0035] The technical solution of the present invention will be described in detail below through specific embodiments.
[0036] Example 1: Whole Genome Liquid Chip This embodiment provides a whole-genome liquid microarray related to the overall production performance and regional adaptability of Holstein dairy cows in southern regions. The development process of the whole-genome liquid microarray is as follows: 1. Collection of dairy cow blood samples and performance data In three Holstein dairy farms in Hubei Province with different operating models, 348 individuals with at least one complete calving record were selected, ensuring pedigree independence. Blood samples were collected from the tail vein of these Holstein cows, stored in EDTA anticoagulant tubes at 4°C, and all individual production performance (DHI) data were provided by the dairy farms. The 2120i blood analyzer tests the following in each cow: Eleven white blood cell parameters: white blood cell count, lymphocyte count, monocyte count, neutrophil count, basophil count, eosinophil count, lymphocyte percentage, monocyte percentage, neutrophil percentage, basophil percentage, and eosinophil percentage; Seven red blood cell parameters: red blood cell count, hemoglobin, hematocrit, mean corpuscular volume, mean corpuscular hemoglobin content, mean corpuscular hemoglobin concentration, and red blood cell distribution width. There are four platelet parameters: platelet count, mean platelet volume, platelet distribution width, and platelet hematocrit.
[0037] 2. Total DNA extraction and quality control from Holstein cow blood samples Total DNA was extracted from the samples using a genomic DNA extraction kit. DNA purity and integrity were analyzed by 1% agarose gel electrophoresis. DNA concentration was precisely determined using a Qubit 4.0 high-precision fluorescence quantitative PCR instrument.
[0038] 3. Library construction and sequencing Sequencing libraries were constructed using a DNA library preparation kit. The libraries were then quality-checked using a Fragment Analyzer 5300 (Agilent). After passing the quality check, the libraries were sequenced using a DNBSEQ-T7 sequencer with a PE150 sequencing strategy to obtain raw genome sequencing results (RAW DATA).
[0039] 4. Data filtering and preprocessing A 10× depth whole-genome resequencing strategy was used to sequence 348 individuals. The raw sequencing results were filtered using Fastp v0.23.2 software to remove adapters and low-quality bases. The bovine reference genome ARS-UCD 2.0 was downloaded and aligned using BWA v0.7.17 software. The data was then sorted according to alignment positions to generate BAM format alignment result files. SNP genotyping and hard filtering were performed using GATK v4.2 software, followed by genotyping using Beagle v5.4 software. SNP sites were further filtered using PLINK v1.9 software. The filtering criteria were as follows: individuals with a callrate < 97% were excluded; individuals with a callrate < 95%, minimum allele frequency (MAF) < 0.05, and a Hardy-Weinberg equilibrium significance level of P < 10 were excluded. -6 The SNP sites were identified. After filtering, 1,030,098 high-quality SNPs were obtained and could be used for subsequent analysis.
[0040] 5. Screening for relevant SNP loci through genome-wide association analysis. Based on 1,030,098 high-quality SNP loci, principal component analysis (PCA) was performed on the population using PLINK software to calculate inter-individual genetic distances and genetic diversity parameters. ADMIXTURE v1.3.0 was used to extrapolate the population's genetic structure and genetic introgression. PopLDdecay software was used to evaluate the degree of genome-wide linkage disequilibrium (LD) in the population, and the linkage disequilibrium parameter (r) was used to determine the degree of disequilibrium. 2 The decay rate of the ) is used to assess the efficiency and accuracy of the correlation analysis.
[0041] Using HIBLUP 1.6.0 software, the estimated breeding value (EBV) of each lactation trait (including milk yield, milk protein percentage, milk fat percentage, somatic cell count, peak lactation day, peak milk yield, etc.), reproductive traits (such as age at first litter, calving interval, etc.), health traits (white blood cell parameters, platelet parameters, etc.), and fitness (red blood cell parameters, etc.) was calculated using a single-trait animal model. The model is as follows: y = Xb + Zg + e, where y is the phenotypic value of each trait, X is the fixed effects matrix, b is the fixed effects variable such as batch (including first and second batches, a total of 2 levels), parity (including 1st, 2nd, 3rd, and 4th or more parities, a total of 4 levels), and lactation stage (including dry period, pre-lactation, mid-lactation, and post-lactation, a total of 4 levels); Z is the random effects matrix, g is the individual additive effect; and e is the residual effect.
[0042] GWAS analysis of lactation, reproductive, health, and adaptive EBV was performed using the FarmCPU algorithm in rMVP software. The kinship matrix and four principal components were used as covariates. To control for false positives, the Bonferroni method was used for multiple test correction. The significance threshold (P-value) for the GWAS signal was set to 0.05 / N, where N is the number of SNPs used in the analysis, which is 1,030,098; therefore, the P-value was 4.85 × 10⁻⁶. -8 Significant SNP sites were selected as candidate sites for Category I.
[0043] 6. Based on significant gene-related SNP loci identified in dairy cow multi-omics studies, this study screened the reproductive and growth axis positive selection signals of five common herbivorous livestock (cattle, buffalo, horses, goats, and sheep) and the resulting SNPs. It also analyzed the transcriptomics of endometrial cells from dairy cows in southern China under three different hormone combinations (published paper DOI: 10.3389 / fvets.2024.1344259), the proteomics of semen from four bulls under heat stress / non-heat stress conditions, and the proteomics of blood from 40 pregnant / non-pregnant cows (patent ZL20151039859). The results of studies such as 5.5) were cross-validated with GWAS screening studies of SNP loci related to economic traits and blood parameters published in databases such as PubMed and Elservier. Genes related to dairy cow lactation performance, reproductive performance, nutritional metabolism, immune response, and stress resistance were screened. The positions of these genes on the dairy cow genome ARS-UCD 2.0 were extracted using the BioMart online service on the Ensembl website (https: / / useast.ensembl.org / biomart / martview / ), and 5kb regions were added upstream and downstream based on the degree of genome-wide linkage disequilibrium. The intersect procedure of the bedtools software was used to locate SNP loci within these gene regions in the resequencing results as Class II candidate loci.
[0044] 7. Screening for relevant SNP loci using published QTLs associated with health traits. Download the Cattle QTLdb based on the ARS-UCD 2.0 version genome from the Animal QTL database (https: / / www.animalgenome.org / cgi-bin / QTLdb / ). Extract QTL loci associated with common health problems in dairy cows, such as mammary gland disease, reproductive disorders, metabolic disorders, longevity, and stress resistance. Use the intersect procedure of bedtools software to locate SNP loci in the QTLdb that overlap with the resequencing results as candidate loci for category III.
[0045] 8. Filling sites Chromosomal regions lacking candidate SNP sites were filled. Based on publicly published QTLs related to dairy cow lactation, reproduction, health, and adaptation in the Animal QTL database (reference genome version ARS-UCD 2.0), the intersect procedure of bedtools software was used to extract QTLs located in chromosomal regions lacking the first three types of SNP markers in Cattle QTLdb. These QTLs were then merged with the resequencing results, and functional SNP sites that did not have linkage relationships with each other were selected as Class IV candidate sites, i.e., filling sites.
[0046] 9. Genotyping targets of probe groups To ensure a uniform distribution of loci across the genome, a sliding window strategy was employed for screening. First, each locus in the set consisting of all four categories of candidate loci obtained in steps 5 to 8 was scored according to the scoring criteria in Table 1. Then, the genome was traversed within a set window size, and within each window, the loci were ranked based on their scores. Finally, the highest-ranking locus within each window was selected.
[0047] Table 1: Weighting Table of SNP Locus Annotation Information SNP site type Weight SNP site type Weight SNP site type Weight downstream 3 stoploss 5 UTR5 3 intergenic l synonymous SNV 2 UTR5; UTR3 3 intronic 1 unknown 0.5 Epigenomic domains 4 nonsynonymous SNV 5 upstream 3 Conserved domains 4 splicing 5 upstream; downstream 3 Functional.sites 30 stopgain 6 UTR3 3 MAF 2
[0048] Subsequently, candidate capture probes were designed using Oligominer v1.7 software for the top-ranked sites selected by the sliding window strategy. The designed probes had to meet the following technical requirements: (1) probe length 120 bp; (2) GC content 30%–70%; (3) the probe sequence was back-aligned to the genome using bowtie2 and could only be aligned to a unique position; (4) all 25 bp k-mers contained in the probe could not be repeated more than 3 times in the whole genome.
[0049] To ensure the final chip covers the predetermined number of key sites, for candidate sites where probes cannot be directly designed according to the four principles, a linkage method is used to select other closely linked sites for supplementary design. Linkage blocks (R-square >= 0.8) are calculated using plink (v1.90) in 100kb units. Secondary weighted sites with close linkage to sites that do not meet the above design principles are selected for design. If design is still not possible, other secondary weighted sites with close linkage are selected sequentially until the design is completed or all sites are traversed. The final SNP site combinations are shown in Table 2. Table 2 shows a total of 17,051 SNP sites, evenly distributed across 29 autosomes and the X chromosome. Figure 1 ).
[0050] Table 2: Physical location and genotyping information of SNP loci
[0051]
[0052] Note: The bolded sites are new sites included in this chip. For example, X_132716198_G_A, X represents the chromosome number, 132716198 represents the rs number, G_A represents the SNP genotyping information, and the version number of the dairy cow reference genome is ARS-UCD2.0.
[0053] After the above screening process, the final chip site combination (Table 2) has the following characteristics: Table 2 contains a total of 9,752 functional sites (i.e., Class I, Class II, and Class III sites), accounting for 57.19% of the total sites, of which a total of 6,621 are new sites, accounting for 67.89% of the functional sites. Of all the new loci, 841 were related to lactation performance (12.70%), used to assist in the selection of lactation traits; 769 were related to adaptability (11.61%), used to assist in the selection of adaptability traits such as heat stress resistance, oxidative stress resistance, and metabolic capacity; 1,327 were related to health (20.04%), used to assist in the selection of health traits such as disease resistance, high immunity, and longevity; 3,427 were related to reproductive performance (51.76%), used to assist in the selection of reproductive traits; and 257 were related to growth performance (3.88%), used to assist in the selection of growth and meat quality traits (Table 3). This locus combination fully reflects the design focus on key traits of Holstein dairy cattle in southern China.
[0054] Based on Table 2, this invention synthesized corresponding capture probes to prepare liquid-phase microarrays. Probes were synthesized using all SNP sites in Table 2 as gene analysis targets (repeated probes were designed for the 275 top-ranked sites to ensure their detection rate), and these probes were prepared into whole-genome liquid-phase microarrays containing 17,326 probes (SEQ ID NO. 1–17,326), covering 17,051 SNP sites in Table 2.
[0055] To highlight the innovation of the novel chip, we compared the 17,051 SNP loci in Table 2 with existing commercial chips and patents. This was achieved by comparing the chip with two representative commercial chips from Illumina (including the high-density BovineHD Genotyping Beadchip and the medium-density BovineSNP50 v3 DNA Analysis BeadChip) and four chip patents (SNP locus combination and chip preparation for identifying heat stress tolerance in Holstein cattle, CN202110923967.7; a low-density whole-genome chip for breeding high-yield and longevity traits in dairy cows and its application, CN202411812454.9; and a low-density 30K whole-genome chip for dairy cows). SNP chip and its application, CN202310266276.3; A low-density SNP locus combination, liquid phase chip and its application in dairy cows, CN202311046265.0), the results showed that among the 441 class I loci mined based on GWAS, 426 were novel loci, accounting for 96.60%; among the 6,391 class II loci mined based on multi-omics, 6,195 were novel loci, accounting for 96.93%; and 2,901 class III loci were mined based on health trait-related QTLs. For loci that could not be designed, strong linkage method (based on r) was used. 2 (With a threshold of ≥0.8), a total of 19 additional design sites were created, of which 18 were novel sites, accounting for 94.74%. For regions on chromosomes where the first three types of SNP markers were absent, after delinking, 7,299 filling sites (i.e., type IV sites) were successfully designed (Table 2). This result fully demonstrates the novelty and regional specificity of the chip sites in this invention.
[0056] Table 3 Functional classification of new loci
[0057] Example 2: The detection effect of the chip on Holstein dairy cows in southern China 2.1 DNA Sample Extraction and Detection Thirty healthy dairy cows were selected from three representative dairy farms in Hubei Province. Blood samples were collected via tail vein and extracted using a genomic DNA extraction kit (Suzhou Cretaceous Biology, CNA01990F64). DNA purity and integrity were analyzed by 1.5% agarose gel electrophoresis. DNA concentration was precisely determined using a Qubit 4.0 quantitative PCR instrument.
[0058] 2.2 DNA Library Construction and Sequencing The database construction process follows YZSeq TM Library preparation was performed according to the instructions of the Tn5 library preparation kit (Wuhan Yingzi, T1012-008). 50 ng of DNA from the quality control described in section 2.1 was randomly fragmented using the YZ-Tn5 Mix fragmentation enzyme in the kit at 55°C for 10 minutes, resulting in fragment sizes of 180–250 bp. The fragmented DNA was then mixed with a sequencing adapter PCR primer mixture (10 μM of the P5 sequence: 5′-AATGATACGGCGACCACCGAGATCTACAC-3′ (SEQ ID NO. 17327), and 10 μM of the P7 sequence: 5′-CAAGCAGAAGACGGCATACGAGAT-3′ (SEQ ID NO. 17328), respectively), and amplification primers containing the tag sequences Y5XX and Y7XX, in the following system, ensuring that the i5 and i7 tag sequences in each sample were a unique combination.
[0059] Table 4: Amplification System Components Volume (μL) Fragmentation products 24 PCR Primer Mix 5 Y5XX 5 Y7XX 5 YZ-S7 buffer 10 YZ-S7 enzyme 1 total 50
[0060] The nucleotide sequence of Y5XX is: 5′-AATGATACGGCGACCACCGAGATCTACAC-[i5 tag sequence]-TCGTCGGCAGCGTC-3′ (SEQ ID NO.17329); The nucleotide sequence of Y7XX is: 5′-CAAGCAGAAGACGGCATACGAGAT-[i7 tag sequence]-GTCTCGTGGGCTCGG-3′ (SEQ ID NO.17330) (TruePrep Tag Sequence Kit V2 for Illumina, Vazyme, TD202 is also available). PCR amplification was performed using the YZ-S7 high-fidelity DNA polymerase in the kit following the procedure. After completion, the library was double-selected with magnetic beads at a ratio of 0.5× / 0.15× (VAHTS DNA Clean Beads, catalog number: N411-01 / 02 / 03).
[0061] Table 5: Amplification Procedure
[0062] 2.3 The prepared gene chip detection kit containing the library hybridization and enrichment includes the following reagents: Table 6: Components of the Gene Chip Detection Kit (24 samples)
[0063] Prepare a library blocking mixture by mixing 500 ng of DNA library and 5.6 μL of PHB reagent. For multi-library hybridization, mix 2-24 libraries, 500 ng each, with a total library volume not exceeding 6 μg. Vortex thoroughly to mix the above components, then briefly centrifuge to collect the liquid in a 1.5 mL low-absorption centrifuge tube. Centrifuge the tube in a preheated vacuum concentrator at 30°C until no liquid remains. Prepare a hybridization reaction mixture by mixing 2 μL of 10KP and 15 μL of HB. Pipette the mixture into a dried 1.5 mL low-absorption centrifuge tube, mix thoroughly 20 times to avoid air bubbles, briefly centrifuge, and incubate at room temperature for 5-10 minutes. Transfer the incubated hybridization reaction system to a 0.2 mL low-absorption centrifuge tube and run the hybridization program on a PCR amplification instrument: hot cap temperature 105°C, reaction temperature 95°C for 40 seconds, 65°C for 16-18 hours.
[0064] Gently vortex the completely dissolved HB and 2xCBWB to minimize bubble formation. Prepare a 1× working solution (1×Capture Beads Wash Buffer, hereinafter referred to as "1×CBWB") by mixing 150 μL of 2×CBWB and 150 μL of enzyme-free water. Prepare a magnetic bead suspension by mixing 15 μL of HB and 2 μL of nuclease-free water. Gently vortex the CB, and transfer 20 μL of CB to a new 0.2 mL low-adsorption centrifuge tube. Add 100 μL of 1×CBWB to the 0.2 mL low-adsorption centrifuge tube, mix gently, and briefly centrifuge to collect the liquid at the bottom of the tube. Place the tube on a magnetic rack for 3 minutes until the solution is clear, and discard the supernatant. Repeat this step twice. Then, remove the 0.2 mL low-adsorption centrifuge tube, add 17 μL of the magnetic bead suspension, and gently mix the magnetic beads with a pipette, avoiding bubble formation.
[0065] After the 16-18h hybridization program, briefly centrifuge the hybridization reaction system to collect the liquid at the bottom of the tube, and completely transfer it to the preheated magnetic bead suspension. Gently mix 10-15 times with a pipette, avoiding the formation of air bubbles. Perform the elution program (I) on another PCR amplification instrument: heat the lid to 80°C, maintain the reaction temperature at 65°C for 45 minutes, and mix 10 times every 15 minutes using a pipette to ensure thorough mixing of the magnetic beads.
[0066] Prepare 1×WB S by mixing 30 μL of 10×WB S with 270 μL of enzyme-free water; prepare 1×WB I by mixing 50 μL of 5×WB I with 200 μL of enzyme-free water; and prepare 1×WB II by mixing 15 μL of 10×WB II with 135 μL of enzyme-free water. Preheat the prepared 1×WB I and 1×WB S solutions at 65°C and 68°C, respectively, for at least 15 minutes. After the 45-minute elution program (I), add 100 μL of preheated 1×WB I to the capture system and gently mix with a pipette. Place a 0.2 mL low-adsorption centrifuge tube on a magnetic rack for 1-2 minutes until the solution becomes clear, then discard the supernatant. Remove the 0.2 mL low-adsorption centrifuge tube and add 150 μL of preheated 1×WB I. S. Gently mix with a pipette to minimize the generation of air bubbles. Place the mixture in a PCR amplification instrument and perform the elution procedure (II): keep the lid at 80°C and maintain the reaction temperature at 68°C for 5 minutes. Repeat the process of letting the mixture stand, discarding the supernatant, and performing the elution procedure (II) once.
[0067] Place a 0.2 mL low-adsorption centrifuge tube on a magnetic rack for 1-2 minutes until the solution becomes clear, then discard the supernatant. Remove the 0.2 mL low-adsorption centrifuge tube and add 150 μL of 1×WB I at room temperature. Gently mix with a pipette and incubate at room temperature for 2 minutes, vortexing every 30 seconds. Place the 0.2 mL low-adsorption centrifuge tube on a magnetic rack for 1 minute until the solution becomes clear, then discard the supernatant. Remove the 0.2 mL low-adsorption centrifuge tube and add 150 μL of 1×WB I at room temperature. II. Gently mix with a pipette and incubate at room temperature for 2 minutes, vortexing every 30 seconds. Place the 0.2 mL low-adsorption centrifuge tube on a magnetic rack for 1 minute until the solution is clear, then flush with the supernatant. Remove the 0.2 mL low-adsorption centrifuge tube, centrifuge briefly, place on a magnetic rack for 1 minute, and remove any remaining liquid with a pipette. Remove the 0.2 mL low-adsorption centrifuge tube, add 48 μL of enzyme-free water, and gently mix with a pipette to obtain a magnetic bead solution containing DNA. Store at 4°C.
[0068] Prepare the PCR reaction mixture in a new 0.2 mL low-adsorption centrifuge tube: 10 μL YZ B, 3 μL PPM, 1 μL YZ E, and 36 μL magnetic bead solution containing DNA. Place the tube in a PCR chamber and run the PCR program as shown in Table 7. Table 7: Amplification Procedure
[0069] After the reaction, place the 0.2 mL low-adsorption centrifuge tube on a magnetic rack and let it stand for 1 minute until the solution is clear. Transfer the supernatant to a new 1.5 mL low-adsorption centrifuge tube. Add enzyme-free water to make up the volume to 100 μL, then add 100 μL of PB that has been equilibrated at room temperature for 30 minutes. Vortex to mix and let stand at room temperature for 5 minutes. Briefly centrifuge the 1.5 mL low-adsorption centrifuge tube to collect the mixture to the bottom of the tube. Place it on a magnetic rack for 5 minutes until the solution is clear and discard the supernatant. Keep the 1.5 mL low-adsorption centrifuge tube on the magnetic rack and add 200 μL of freshly prepared... Prepare 80% ethanol, do not disturb the magnetic beads, incubate at room temperature for 30 seconds, discard the supernatant, and repeat this step once; keep the 1.5 mL low-adsorption centrifuge tube in the magnetic rack, open the cap and dry at room temperature for 5 minutes until the surface of the magnetic beads is frosted; remove the 1.5 mL low-adsorption centrifuge tube from the magnetic rack, add 22 μL of enzyme-free water, vortex to mix, let stand at room temperature for 5 minutes, and then briefly centrifuge to collect the mixture to the bottom of the tube; place the 1.5 mL low-adsorption centrifuge tube on the magnetic rack, let stand for 5 minutes until the solution is clear, and use a pipette to transfer 20 μL of supernatant into a new 1.5 mL centrifuge tube.
[0070] Take 1 μL of the above library and quantify it using the QubitdsDNA HS Assay Kit, recording the library concentration. Dilute 1 μL of the sample to 2 ng / μL and determine the library fragment length using the Agilent 5300 Bioanalyzer system (High Sensitivity DNA Kit) or Q-seq. The library length is approximately 350 bp. After quality control, use DNBSEQ. TM Sequencing was performed using the T7 sequencing platform with the PE150 sequencing strategy.
[0071] 2.4 SNP typing The raw sequencing data (RAW DATA) was quality controlled using the software FastP v0.23.1. The main steps included: (1) Remove the connector sequence from the read segments (--detect_adapter_for_pe); (2) Remove low-quality bases at the 5' end and 3' end; (--cut_front-cut_front_window_size 1-cut_front_mean_quality 10--cut_tail-cut_tail_window_size 1-cut_tail_mean_quality10--cut_right_window_size 4-cut_right_mean_quality 20); (3) Set the minimum length of the read segment and remove read segments with a length < 40bp; (4) Remove the unqualified bases with a base quality value <20 to finally obtain clean reads for further analysis.
[0072] The clean reads after quality control were aligned with the ARS-UCD 2.0 reference genome using BWA v0.7.17-r1188 software (using default parameters) to determine the genomic location of each read. The reads were then sorted according to their alignment locations, generating a SAM format result file. Picard v2.20.7 software was used to remove duplicate reads, improving alignment quality. HaplotypeCaller in GATK v4.1.8.1 software was used for variant detection in individual samples. Then, CombineGVCFs and GenotypeGVCFs in GATK software were used to merge the GVCF results from all samples, ultimately obtaining the variant information file. Picard and GATK software were used to mark low-quality SNP genotyping results (GQ < 20) and filter them (genotypes with GQ < 20 were set as missing genotypes). The detection rate of 17,051 SNPs for each sample was calculated.
[0073] 2.5 Evaluation of the predictive accuracy of the chip Based on loci detected by microarrays and loci detected by 10× whole genomes, the genomic estimated breeding value (GEBV) for each lactation trait (including milk yield, milk protein percentage, milk fat percentage, somatic cell count, peak lactation day, peak milk yield, etc.), reproductive traits (such as age at first litter, calving interval, etc.), health traits (white blood cell parameters, platelet parameters, etc.), and fitness (red blood cell parameters, etc.) was calculated using a single-trait animal model in HIBLUP 1.6.0 software. The GEBV calculation formula is: y = Xb + Zg + e, where y is the phenotypic value of each trait, X is the fixed effects matrix, b is the fixed effects variables such as batch (including first and second batches, a total of 2 levels), parity (including 1st, 2nd, 3rd, and 4th or more parities, a total of 4 levels), and lactation stage (including dry period, pre-lactation, mid-lactation, and post-lactation, a total of 4 levels); Z is the genotype relationship matrix, g is the additive genetic random effect; and e is the residual effect. The accuracy of predictions was assessed using the Pearson correlation coefficient between GEBVs and EBVs for each trait (Step 5 of Example 1).
[0074] 2.6 The effectiveness of the chip in detecting dairy cows in southern regions In the test population, loci with a minimum allele frequency > 0.1 accounted for 66.15% of the 17,051 SNPs. Validation with 30 new samples showed that the genotyping success rate (detection rate) for the 17,051 SNPs exceeded 99.26%, with a maximum of 99.44%. Figure 2 Repeated testing of the same sample showed a genotyping consistency of 99.83% between the two tests; the average genotyping consistency of the same sample's gene chip test and resequencing (10×) of 17,051 SNPs was 96.76%.
[0075] 2.7 Evaluation of the Predictive Accuracy of the Chip Table 8 shows a comparison of the prediction accuracy of the microarray and 10× whole-genome sequencing in the test population. The correlation coefficients between the GEBVs calculated based on this microarray and the EBVs calculated based on pedigree records are both greater than 0.95, indicating a high degree of consistency between the trait values predicted by the microarray and the actual observed values. Furthermore, there was no significant difference between the GEBVs predicted by the microarray and the results of 10× whole-genome sequencing (P > 0.05), and the Pearson correlation coefficients were both greater than 0.95, indicating high reliability of the microarray predictions.
[0076] Table 8: Pearson correlation coefficients between GEBVs and EBVs for each trait under the two prediction methods Group chip 10× Whole Genome Sequencing lactation characteristics 0.9786 0.9837 Reproductive traits 0.9804 0.9935 health traits 0.9758 0.9782 Adaptability 0.9667 0.9767
[0077] Note: The predictive accuracy of lactation traits, reproductive traits, health traits, and fitness is the mean of the predictive accuracy of each indicator and parameter.
[0078] Example 3: The effect of bovine whole-genome liquid microarray on the detection of dairy cows in northern China 3.1 DNA Sample Extraction and Detection Blood samples were collected from 40 dairy cows at a dairy farm in Ningxia Hui Autonomous Region as representative samples for evaluating the effectiveness of the microarray. Genomic DNA was extracted using a genomic DNA extraction kit. The purity and integrity of the DNA were analyzed by 1.5% agarose gel electrophoresis. The DNA concentration was accurately determined using a Qubit 4.0 quantitative PCR instrument.
[0079] 3.2 DNA Library Construction and Sequencing Same as step 2.2.
[0080] 3.3 SNP typing Same as step 2.3.
[0081] 3.4 The effectiveness of the chip in detecting dairy cows in northern regions In a herd of 40 individual northern dairy cows, the average target site detection rate was 99.04%, with a maximum of 99.61%; the technical consistency rate of repeated sample typing reached 99.76%. Figure 3 After filtering and quality control, a total of 14,888 high-quality SNPs were obtained, of which 64.60% had a minimum allele frequency > 0.1. The results suggest that this chip can improve the applicability and selection effect of genomic selection technology for Holstein dairy cattle in southern regions, while also taking into account dairy cattle in northern regions.
[0082] Example 4: Application of molecular marker-assisted selection based on this gene chip in the localized breeding of dairy cows Ten individuals were randomly selected from a Holstein replacement gilt population in southern China, and their DNA samples were collected. Genotyping was performed using the liquid-phase chip described in Example 1. Following the steps in Example 2, SNP genotyping data were obtained, and the genomic kinship among these ten cattle was calculated. The number of favorable alleles at relevant SNP loci in different individuals was counted for heat tolerance and disease resistance. Based on the production performance of a reference population, GEBV was calculated using a single-trait animal model. With the goal of "breeding healthy, heat-tolerant dairy cows," individuals were ranked comprehensively based on the number of favorable alleles and GEBV, with those carrying more favorable alleles and higher GEBV ranking higher. The results showed that individuals 42AD01210063 and 420385210106 carried more favorable alleles at heat tolerance and health-related SNP loci, and their GEBV was significantly higher than the population average, indicating that they possess excellent genetic potential in heat tolerance and disease resistance and can be considered high-quality replacement gilts. It is recommended to use the semen of verified bulls with distant kinship and dominant alleles in genes related to lactation performance for breeding, in order to increase the probability that offspring will simultaneously possess genes beneficial to high productivity, heat resistance and health.
[0083] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A whole-genome liquid-phase chip, characterized in that: This includes a probe set, whose genotyping targets include all SNP sites listed in Table 2.
2. The application of the whole-genome liquid phase chip as described in claim 1 in molecular marker-assisted selection of Holstein dairy cows in southern China.
3. The application according to claim 2, characterized in that, The application is for at least one of the following purposes: (1) A genomic genetic assessment of lactation, reproduction, health and adaptability traits of Holstein dairy cows in southern China; (2) Locate genes or genomic regions associated with important traits in Holstein dairy cows in southern China; (3) Analyze the genetic diversity, kinship or population structure of Holstein dairy cattle populations in southern China.
4. The application according to claim 3, characterized in that, The specific application is molecular marker-assisted selection of Holstein dairy cows in southern China.
5. A combination of SNP sites, characterized in that: Includes all SNP sites listed in Table 2.
6. A probe assembly, characterized in that: The genotyping targets of the probe set include all SNP sites listed in Table 2.
7. The application of the SNP site combination of claim 5 and the probe set of claim 6, characterized in that: The application includes at least one of the following: (1) Prepare a genotyping chip or kit for breeding Holstein dairy cows in southern China for high yield, disease resistance and heat resistance traits; (2) Assess the genetic diversity and population structure of Holstein dairy cattle populations in southern China; (3) Select high-yield, disease-resistant, and heat-resistant traits in Holstein dairy cattle in southern China through marker-assisted selection or genomic selection; (4) Marker-assisted selection of Holstein dairy cows in southern China; (5) Detect the genotypes of Holstein dairy cows in southern China that are associated with high yield, disease resistance and heat tolerance.
8. A kit comprising the probe set of claim 6.
9. A method for genetic evaluation of Holstein dairy cows in southern China, characterized in that: Includes the following steps: Using the whole genome liquid chip as described in claim 1, genotyping was performed on the dairy cows to be tested to obtain genotype data; Based on the genotype data, the estimated genomic breeding values of Holstein dairy cows in southern China for lactation, reproduction, health and adaptability traits were calculated using either the best linear unbiased genomic prediction method or the single-step best linear unbiased genomic prediction method. Based on the genotype data and the phenotypic data of the dairy cows to be tested, a mixed linear model was used to perform genome-wide association analysis to locate genomic regions associated with lactation, reproduction, health, and adaptive traits.
10. A genetic evaluation system for Holstein dairy cows in southern China, used to implement the genetic evaluation method of claim 9, characterized in that, The system includes: processor; A memory on which computer programs are stored; When the processor executes the computer program, it performs the functions of the following modules: The genotype data acquisition module is used to control the whole genome liquid phase chip of claim 1 to perform genotyping on the dairy cows to be tested, and to receive and store the generated genotype data; The genome breeding value calculation module is used to call the GBLUP or SSGBLUP algorithm program to calculate the estimated genome breeding value based on the genotype data. The gene mapping analysis module is used to run a mixed linear model analysis program to perform genome-wide association analysis based on the genotype and phenotypic data.