A GBS genome-wide association study method for buffalo

By employing the GBS genome-wide association analysis method, the problems of efficiency and economy in genome-wide association analysis of economic traits in buffalo were solved, realizing the application of high-throughput sequencing technology in buffalo breeding and providing precise breeding candidate targets.

CN117095746BActive Publication Date: 2025-11-14GUANGXI ZHUANG AUTONOMOUS REGION BUFFALO INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311086801.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-28
Publication Date
2025-11-14
Estimated Expiration
2043-08-28

AI Technical Summary

Technical Problem

Existing technologies make it difficult to perform genome-wide association analysis of economic traits in buffalo efficiently and economically. Traditional chip typing is costly, cannot discover new SNP sites, and is complex to operate.

Method used

The GBS genome-wide association analysis method was used, including sequencing data quality control, alignment with the reference genome, SNP detection and annotation, population stratification analysis, and genome-wide association analysis. High-throughput sequencing technology was used to detect SNP sites in the buffalo population, and population stratification was performed through phylogenetic tree and principal component analysis. Trait association analysis was then performed using a mixed linear model.

Benefits of technology

It enables efficient and low-cost detection of unknown variant sites on the buffalo genome, with high SNP marker transformation success rate, accurate trait association analysis, and provides more precise breeding candidate targets. It also reduces the cost per SNP marker site and is simple and stable to operate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117095746B_ABST
    Figure CN117095746B_ABST
Patent Text Reader

Abstract

This invention discloses a GBS genome-wide association analysis method for buffalo, belonging to the field of genome association analysis technology. The key technical points are: the method includes several steps such as sequencing data quality control, reference gene alignment, SNP detection and annotation, population stratification analysis, and genome-wide association analysis. This method can detect novel SNPs at unknown variant sites in the genome, with a high SNP marker transformation success rate; it obtains millions of SNP sites in a single sequencing run, achieving high density; the cost of each obtained SNP marker site is reduced by an order of magnitude compared to traditional microarray technology; the method provides accurate data, is technically stable, simple to operate, and highly reproducible; utilizing high-throughput sequencing to obtain buffalo SNP markers and trait association analysis allows for more comprehensive and precise localization of genes or molecular modules related to target traits, providing more accurate candidate targets for molecular selection and genetic improvement in buffalo breeding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of genome-wide association analysis technology, and more specifically, to a GBS genome-wide association analysis method for buffalo. Background Technology

[0002] Improving the genetic quality of a buffalo population through breeding is key to enhancing the production level and efficiency of the industry. Molecular breeding technology, with genomic selection at its core, offers an opportunity to significantly increase the rate of genetic improvement and production efficiency compared to traditional breeding techniques.

[0003] Milk production, health, growth, and reproduction are among the most important economic traits in buffalo, directly impacting the buffalo industry. For many years, traditional breeding methods have yielded some success in genetically improving these economic traits. However, due to the long breeding cycle, the complexity of these traits, and their control by numerous genes, traditional methods have struggled to achieve significant genetic progress. In recent years, with rapid technological advancements, molecular marker-assisted breeding has emerged as a new method for improving genetic traits.

[0004] Currently, there are two main methods for genome-wide SNP genotyping: genotyping microarrays and sequencing. While genotyping microarrays are technically stable and have high reproducibility, they are costly to genotype a single experimental sample. Furthermore, population genotyping is even more expensive in population genetics research. In addition, due to technological limitations, SNP polymorphisms have poor universality across different populations, low marker density, and cannot perform fine functional gene localization or genome-wide association analysis.

[0005] Currently, a new high-throughput sequencing-based technology, GBS (Genotyping-by-sequencing), has been developed. This technology involves genotyping through sequencing and constructing SNP molecular markers by selecting appropriate restriction endonucleases and combining them with high-throughput population sequencing. It can be used in molecular marker development, ultra-high-density genetic map construction, population genetic analysis, and population GWAS analysis. Compared to microarrays, this technology is simpler and less costly; it yields a large number of SNP loci in a single sequencing run with high density; it can detect novel SNPs at unknown variant sites on the genome; it is suitable for species with or without a reference genome; the sequencing fragments are complete; and the SNP marker conversion success rate is high.

[0006] Genome-wide association study (GWAS) is a method for analyzing the overall association of common genetic variations (single nucleotide polymorphisms and copy number) across the entire genome. Using natural populations as the research object, this method is based on linkage disequilibrium (LD) between genes (locuses) retained after long-term recombination. It combines the diversity of target phenotypic phenotypes with gene (or marker) polymorphisms for analysis, directly identifying gene loci or marker loci closely related to phenotypic variations and possessing specific functions. GWAS technology, used across the entire genome, can locate multiple traits simultaneously, making it suitable for trait association studies, functional gene research, trait selection, and functional marker research. GWAS technology has been widely applied in animal breeding as a novel method. GWAS aims to identify single nucleotide polymorphisms (SNPs) associated with traits across the entire genome, yielding more reliable results. In recent years, GWAS has been applied in bovine molecular breeding as an assisted selection method, while its application in buffalo molecular breeding is still in the experimental research stage. Currently, most GWAS studies are based on microarray genotyping technology, which can only detect known SNP polymorphic sites, cannot discover new sites, and is complex and costly to operate. For these reasons, there is an urgent need to develop a universal, cost-effective, and easy-to-operate GWAS genome-wide association analysis method suitable for buffalo, providing technical support for molecular selection and genetic improvement in buffalo breeding. Summary of the Invention

[0007] The purpose of this invention is to provide a GBS genome-wide association analysis method for buffalo to analyze the economic traits of buffalo populations.

[0008] The above-mentioned technical objective of the present invention is achieved through the following technical solution: a GBS genome-wide association analysis method for buffalo, the genome-wide association analysis method comprising the following steps:

[0009] S1. Sequencing data quality control;

[0010] S2. Align with the reference genome;

[0011] S3. SNP detection and annotation;

[0012] S4. Group Stratification Analysis;

[0013] S5. Genome-wide association analysis.

[0014] The present invention is further configured such that the sequencing data quality control method is as follows:

[0015] 1) Filter out buffalo sequencing sequences containing adapter sequences;

[0016] 2) When the number of undetected bases in a single-end sequencing sequence exceeds 10% of the sequence length, this pair of bases needs to be removed.

[0017] 3) When the number of low-quality (<=5) bases in a single-end sequencing sequence exceeds 50% of the sequence length, this pair of bases needs to be removed.

[0018] 4) After the above-mentioned strict filtering of buffalo sequencing data, high-quality and effective data were obtained.

[0019] The present invention is further configured such that: the comparison reference gene is obtained by comparing the effective data obtained in S1 with the reference genome to obtain the comparison rate, average sequencing depth and other relevant data.

[0020] The present invention is further configured such that the SNP detection and annotation operations are as follows:

[0021] (1) Detect SNP sites in a buffalo population and filter the obtained polymorphic sites to obtain high-quality SNP sites;

[0022] (2) Perform population SNP annotation on the obtained high-quality SNP sites.

[0023] The present invention is further configured such that the population stratification analysis can employ two analysis methods: population phylogenetic tree analysis and population principal component analysis.

[0024] The present invention is further configured such that the whole genome association analysis is divided into two steps: trait association analysis and gene function annotation of the target trait related regions.

[0025] In summary, this invention has the following beneficial effects: This method can detect novel SNPs at unknown mutation sites in the genome, with a high SNP marker conversion success rate; it obtains millions of SNP sites in a single sequencing run, resulting in high density; the cost of each obtained SNP marker site is reduced by an order of magnitude compared to traditional microarray technology; the method provides accurate data, is technically stable, simple to operate, and highly reproducible; and by using high-throughput sequencing to obtain buffalo SNP markers and trait association analysis, it can more comprehensively and accurately locate genes or molecular modules related to target traits, providing more precise candidate targets for molecular selection and genetic improvement in buffalo breeding. Attached Figure Description

[0026] Figure 1 This is the GBS experimental procedure in Embodiment 1 of the present invention;

[0027] Figure 2This is the phylogenetic tree of different breeds of buffalo in Embodiment 4 of the present invention.

[0028] Figure 3 This is a two-dimensional diagram of the PCA results in Embodiment 4 of the present invention;

[0029] Figure 4 This is a three-dimensional diagram of the PCA results in Embodiment 4 of the present invention;

[0030] Figure 5 This refers to the trait association analysis results in Example 5 of the present invention;

[0031] Figure 6 This is a flowchart of the genome-wide association analysis of this invention. Detailed Implementation

[0032] The following is in conjunction with the appendix Figure 1-6 The present invention will be described in further detail below.

[0033] Example 1: Sequencing to obtain raw data

[0034] DNA testing was performed on 182 samples (1 replicate) from different buffalo populations (48 Mora buffalo, 29 Nile-Raffi buffalo, 12 Mediterranean buffalo, 23 local buffalo, and 70 Mora and Nile-Raffi crossbred buffalo, all aged 24–36 months) using the following three methods. The testing steps are as follows:

[0035] (1) DNA was extracted from buffalo blood according to the DNA extraction kit instructions, and the purity and integrity of the DNA were analyzed by 1% agarose gel electrophoresis.

[0036] (2) Nanodrop detection of DNA purity (OD260 / 280 ratio);

[0037] (3) Qubit accurately quantifies DNA concentration.

[0038] like Figure 1 As shown, after the detection was completed, a library was constructed. For GBS library construction, the genome was first digested with restriction endonucleases. 0.1-1 μg of genomic DNA was digested with restriction endonucleases to obtain a suitable marker density. P1 and P2 adapters (complementary to the nicks in the digested DNA) were added to both ends of the digested fragments. PCR amplification was performed using tag sequences containing P1 and P2 adapters at both ends. DNA fragment pooling was then performed, and the desired DNA regions were recovered by electrophoresis. Paired-end 150° sequencing was then performed using the Illumina HiSeq sequencing platform.

[0039] The enzyme digestion data of 182 buffaloes were statistically analyzed, and the data of 3 buffaloes were randomly selected, as shown in Table 1.

[0040] Table 1. Enzyme digestion capture statistics

[0041]

[0042] Statistical analysis was performed on the output data of 182 buffaloes (Table 2 shows the data of 3 randomly selected buffaloes), including sequencing data output, sequencing error rate, Q20 content, Q30 content, GC content, etc.

[0043] Table 2. Statistics of buffalo sequencing data output

[0044]

[0045] Q20: The percentage of bases with a quality value of 20 or higher (error rate below 1%);

[0046] Q30: The percentage of bases with a quality value of 30 or higher (error rate below 0.1%);

[0047] This project sequenced a total of 182 buffalo samples from different species, with a total sequencing data volume of 131.00 Gb, averaging 719.78 Mb per sample; high-quality clean data amounted to 130.99 Gb, averaging 719.71 Mb per sample. The sequencing quality was high (Q20 ≥ 93.60%, Q30 ≥ 85.00%), with normal GC distribution. None of the 182 buffalo samples were contaminated, and the library construction and sequencing were successful.

[0048] After the library was constructed, it was first preliminarily quantified using Qubit 2.0 to dilute the library to 1 ng / μl. Then, the insert size of the library was detected using Agilent 2100. After the insert size met the expectations, the effective concentration of the library was accurately quantified using Q-PCR (effective concentration of library > 2 nM) to ensure the quality of the library.

[0049] Example 2: Comparison with reference gene

[0050] Effective, high-quality sequencing data were aligned to the reference genome using BWA software (parameter: mem-t4-k32-M).

[0051] The genome size was 2,836,166,969 bp, with an average alignment rate of 95.25%–99.67% for the population samples. The average sequencing depth ranged from 7.33X to 26.46X, and the 1X coverage (coverage of at least one base) was above 2.26%. The alignment rate reflects the similarity between the sample sequencing data and the reference genome, while coverage depth and coverage directly reflect the uniformity of the sequencing data and homology with the reference sequence. The alignment results of each sample showed that their similarity to the reference genome met the requirements for resequencing analysis, while also exhibiting very good coverage depth and coverage. Detailed statistical results for some samples are shown in Table 3.

[0052] Table 3. Statistics on sequencing depth and coverage

[0053]

[0054] 1X refers to a site in the reference genome that is covered by at least one base;

[0055] 4X refers to a site in the reference genome that is covered by at least 4 bases.

[0056] Example 3: SNP Detection and Annotation

[0057] SNPs (Single Nucleotide Polymorphisms) mainly refer to DNA sequence polymorphisms caused by variations in a single nucleotide at the genomic level, including single base transitions and transversions. Software such as SAMTOOLS is used to detect population SNPs. Bayesian models are used to detect polymorphic sites in the population.

[0058] SAMTOOLS software detected a total of 2,528,010 SNP sites. The obtained SNPs were filtered to obtain high-quality SNPs. The filtering conditions were dp2, Miss0.2, and Maf0.01. Finally, a total of 263,946 SNP sites were obtained for subsequent analysis.

[0059] The obtained high-quality SNPs were used for population SNP annotation using ANNOVAR software. ANNOVAR is a highly efficient software tool that can functionally annotate gene variations detected from multiple genomes using the latest information. Given the chromosome, start site, stop site, reference nucleotide, and variant nucleotide, ANNOVAR can perform gene-based annotation, region-based annotations, filter-based annotation, and other functionalities. Given ANNOVAR's powerful annotation capabilities and international recognition, it was used to annotate the SNP detection results. The detection results are shown in Table 4.

[0060] Table 4. SNP Statistical Information and Annotation Results

[0061]

[0062] Example 4: Population Stratification Analysis

[0063] Population stratification refers to the phenomenon of subpopulations within a population, where the relationships between individuals within a subpopulation are greater than the average kinship between individuals within the entire population. Different subpopulations may exhibit different allele frequencies at certain loci, leading to false positives when two subpopulations are combined for association analysis. Therefore, population stratification analysis must be performed before conducting association analysis. Population genetic diversity analysis, including phylogenetic tree analysis and principal component analysis, can infer the origin and degree of differentiation of each subpopulation; the results of both can be cross-validated.

[0064] (1) Population phylogenetic tree analysis

[0065] A phylogenetic tree (also known as an evolutionary tree) is a branching graph or tree that describes the evolutionary order among populations, representing their evolutionary relationships. Based on the similarities or differences in physical or genetic characteristics, the closeness of kinship between populations can be inferred; that is, the relationships between individuals within a population arising from a common ancestor. We construct phylogenetic trees using the neighbor-joining method.

[0066] After SNP detection, the obtained individual SNPs can be used to calculate the distance between populations. The p-distance between two individuals i and j is calculated using the following formula.

[0067]

[0068] In the formula, L is the length of the high-quality SNP region, and the allele at position 1 is A / C. Therefore:

[0069]

[0070] The distance matrix was calculated using TreeBest software, and a phylogenetic tree was constructed based on this matrix using the neighbor-joining method. Bootstrap values ​​were obtained through up to 1000 calculations. The phylogenetic tree analysis results are as follows: Figure 2 As shown in the figure, the tree-like topology visually illustrates the evolutionary relationships among different species of buffalo. Evolutionary branches of closely related species often cluster together and are marked with the same color. From the figure, we can see that the red, green, and yellow groups have relatively obvious grouping patterns.

[0071] (2) Principal Component Analysis

[0072] Principal component analysis (PCA) is a purely mathematical method that uses linear transformations to select a smaller number of important variables from multiple related variables. PCA is widely used across multiple disciplines. In genetics, it is primarily used for cluster analysis. Based on the degree of SNP differences in an individual's genome, it clusters individuals into different subpopulations according to different trait characteristics, and is also used for cross-validation with other methods. PCA only applies to autosomal data with n=XX individuals, ignoring loci with more than two alleles and mismatches. The PCA analysis method is as follows:

[0073] The SNP at position i, k in individual is denoted by dik. If individual i is homozygous with the reference allele, then dik = 0; if it is heterozygous, then dik = 1; if individual i is homozygous with a non-reference allele, then dik = 2. M is an n×S matrix containing the standard genotype.

[0074]

[0075] In the formula, E(dk) is the average value of dk, and the individual sample covariance n×n matrix is ​​calculated by X=MMT / S.

[0076] The eigenvectors and eigenvalues ​​were calculated using GCTA software, and the PCA distribution plot was drawn using R software. The PCA analysis results are as follows: Figure 3 , 4 As shown in the figure, the horizontal and vertical axes represent principal component 1 and principal component 2, respectively. Different colors in the figure represent different populations. The results are largely consistent with the phylogenetic tree results of the buffalo population.

[0077] Example 5: Genome-wide association analysis

[0078] (1) Association analysis of growth traits

[0079] Nine body measurements of buffalo were taken (including height at hip cross, chest width, chest depth, body length (BL), hip width, rump length (RL), ischial end width (PBW), and loin angle width (HW)). Buffalo birth weight information was also consulted.

[0080] In GWAS analysis, individual kinship and population stratification are the main factors causing spurious associations. Therefore, a mixed linear model is used for trait association analysis, with population genetic structure as a fixed effect and individual kinship as a random effect, to correct for the influence of population structure and individual kinship.

[0081] y = Xα + Zβ + Wμ + e

[0082] y represents the phenotypic trait, X is the indicator matrix for fixed effects, α is the estimated parameter for fixed effects; Z is the indicator matrix for the SNP, β is the effect of the SNP; W is the indicator matrix for random effects, μ is the predicted random individual, and e is the random residual, following the order e ~ (0, δe). 2 ).

[0083] Given that kinship among individuals may influence group stratification, a QQ-plot of the population under a mixed linear model was plotted. Figure 5 The QQ-plot shows that the observed values ​​(vertical axis) are basically consistent with the expected values ​​(horizontal axis), indicating that the association analysis did not produce false negatives due to population stratification, and the association analysis results are reliable.

[0084] The mixed linear model analysis results showed that a total of 69 SNPs significantly associated with 10 growth traits of buffalo were screened, along with 81 most recently associated genes (detailed results are shown in Table 5). The Manhattan plot obtained from the mixed linear model analysis is shown below. Figure 5 .

[0085] Table 5. Significant SNP sites and number of candidate genes screened by GWAS

[0086]

[0087] (2) Gene function annotation of target trait-related regions

[0088] Based on the analysis results, functional annotations were performed on related genes within a certain region upstream and downstream of the physical location of significant SNP sites. The annotation results are shown in Table 6.

[0089] Table 6. Functional annotations of GWAS-associated genes.

[0090]

[0091] This specific embodiment is merely an explanation of the present invention and is not intended to limit the invention. After reading this specification, those skilled in the art can make modifications to this embodiment without contributing any inventive step, but such modifications are protected by patent law as long as they are within the scope of the claims of the present invention.

Claims

1. A GBS genome-wide association study method for buffalo, characterized by: The genome-wide association analysis method includes the following steps: S1. Sequencing data quality control: 1) Filter out buffalo sequencing sequences containing adapter sequences; 2) When the number of undetected bases in a single-end sequencing sequence exceeds 10% of the single-end sequencing sequence length, this pair of bases needs to be removed. 3) When the number of low-quality bases in a single-end sequencing sequence exceeds 50% of the length of the single-end sequencing sequence, this pair of bases needs to be removed. 4) After the above rigorous filtering of the buffalo sequencing data, high-quality and effective data were obtained; S2. Alignment with reference gene: The alignment with reference gene involves comparing the valid data obtained in S1 with the reference genome to obtain alignment rate, average sequencing depth, and other relevant data; the average alignment rate of the population sample is 95.25%–99.67%, the average sequencing depth of the genome is 7.33X–26.46X, the 1X coverage is above 2.26%, and the sequencing quality Q20 ≥ 93.60% and Q30 ≥ 85.00%; S3.SNP detection and annotation: (1) Detect SNP sites in the population and filter the obtained polymorphic sites with the filtering conditions being dp2, Miss0.2, and Maf0.01 to obtain high-quality SNP sites; (2) Perform population SNP annotation on the obtained high-quality SNP sites; S4. Population Stratification Analysis: The population stratification analysis can employ two methods: population phylogenetic tree analysis and population principal component analysis. S5. Genome-wide association analysis: The genome-wide association analysis consists of two steps: trait association analysis and gene function annotation of regions related to the target trait.

Citation Information

Patent Citations

  • Method for researching genetic genes of argali and hybrid offspring thereof

    CN110093406A