Application of functional haplotype marker in genetic feature identification and agronomic trait gene localization of cotton germplasm resources
By constructing functional haplotype markers of cotton genes and combining with mixed linear models, the linkage imbalance and distal regulation of cotton gene localization are solved, the genetic characteristics identification of cotton germplasm resources and the efficient localization of causal genes of agronomic traits are achieved, and the accuracy of breeding is improved.
Patent Information
- Application Number
- CN202510631123.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-26
AI Technical Summary
In existing cotton genomics research, gene localization methods have linkage imbalance and distal regulation problems, which lead to difficulty in localization of causal variants and lack of gene-level variant markers, making it difficult to efficiently identify the genetic characteristics and agronomic trait causal genes of cotton germplasm resources.
Functional haplotype markers are used to construct gene functional haplotypes through non-synonymous mutation combinations at the gene level, and genome-wide association analysis is carried out in combination with mixed linear models to achieve genetic characteristics identification of cotton germplasm resources and efficient localization of causal genes of agronomic traits.
It improves the interpretability of genetic analysis and the localization accuracy of agronomic trait genes, provides a solid foundation for cotton breeding, and can accurately judge the genetic relationship and agronomic trait performance of cotton varieties.
Smart Images

Figure CN120536616A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics, relates to a method for identifying gene-level functional haplotype markers in cotton, and also relates to the application of this type of marker in the identification of genetic characteristics of cotton germplasm resources and the efficient positioning of causal genes of agronomic traits. Background Art
[0002] Cotton is an important cash crop, a major source of natural fiber and a primary raw material for the fiber industry. China is the world's largest cotton producer and consumer. Cotton production not only significantly impacts the development of my country's agriculture and national economy, but also plays a pivotal role in the global cotton trade.
[0003] Since the last century, China's cotton varieties have evolved from imported varieties to domestically bred varieties, and breeding methods have evolved from conventional hybridization to genetically modified varieties. Since 2012, through genome assembly of representative cotton varieties and large-scale sequencing of germplasm populations, my country's cotton breeding has entered the era of functional genomics. In 2017, Fang Lei et al. conducted whole genome sequencing (WGS) and genome-wide association study (GWAS) on 318 local cotton varieties and modern improved varieties, and identified 119 key genomic variation sites, of which 71 were related to yield traits, 45 were related to fiber quality, and 3 were related to resistance to Verticillium wilt (Fang et al., 2017). In 2018, Ma Zhiying et al. conducted WGS on 419 cotton core germplasm resources and GWAS analysis on 13 fiber-related phenotypes, identifying 7383 genomic variations that were highly correlated with fiber phenotypes. Through genome annotation, these variations were associated with 4820 genes (Ma et al., 2018). et al., 2018); In April 2021, He Shoupu et al. conducted WGS on 1756 cotton germplasm resources, combined with 1492 cotton varieties published in previous studies, and conducted population differentiation studies on a total of 3248 cotton varieties, and conducted GWAS analysis on 1245 core cotton variety resources, revealing the genetic basis of geographical differentiation of upland cotton and key quantitative trait loci (QTL) affecting fiber quality (He et al., 2021); In May, Yuan Daojun et al. conducted WGS on 643 cotton populations containing a large number of wild and semi-wild upland cotton species. The results of population genetics research have greatly promoted our understanding of the origin, differentiation and trait improvement of cotton (Yuan et al., 2021); In August of the same year, Ma Zhiying et al. (2021) conducted WGS on 1081 cotton varieties, and GWAS analysis revealed 446 structural variants associated with seven traits related to cotton fiber quality and yield (Structural Based on a certain linkage disequilibrium interval, the researchers identified 907 candidate genes related to fiber quality and yield and 60 candidate genes related to resistance to Verticillium wilt; in 2022, Han Zegang et al. conducted WGS on 486 representative cotton varieties from Xinjiang, China (Han et al., 2022), and carried out GWAS analysis on phenotypes including fiber quality, yield, disease resistance and seed germination rate. The results showed that 502 variant sites significantly associated with the phenotype were detected.
[0004] Nearly a decade of large-scale population sequencing efforts has provided an extremely rich genetic resource for cotton functional genomics. Numerous GWAS results have revealed thousands of causal variants, but pinpointing the genes responsible for these causal variants remains a significant scientific challenge. First, due to linkage disequilibrium (LD), particularly in crop populations that have undergone extensive artificial breeding, LD levels are extremely high. This leads to high linkage between causal variants and other unrelated variants, making GWAS ineffective in distinguishing the two. Furthermore, due to the potential for remote regulation between variants and genes, the underlying gene may be located far from the causal variant. Therefore, even within a specific low LD region, the actual location of a gene may lie outside the expected LD region. Furthermore, many crop species, particularly polyploid crops such as wheat and cotton, typically have a high gene density, with dozens or even hundreds of genes often located within a single candidate region. These factors hinder gene mapping within the existing GWAS framework, which relies on short-variant sequencing. Furthermore, existing methods and technologies focus on basic genomic sequence variation and lack markers for measuring gene-level variation. Summary of the Invention
[0005] In response to the shortcomings of existing research on gene variation itself and the inefficiency of traditional genetic mapping methods under the current framework, the present invention aims to provide a new genetic information encoding method, which converts multiple non-synonymous mutations distributed in a gene into gene-based variation markers, called functional haplotype markers (FH), and provides the application of the functional haplotype markers in the identification of genetic characteristics of cotton germplasm resources and the efficient positioning of causal genes of agronomic traits.
[0006] To achieve the above-mentioned objectives, the present invention provides the following technical solutions: As a first aspect, a functional haplotype marker at the gene level is provided, which is constructed by the following steps: obtaining non-synonymous mutation sites on genes of different varieties of the same organism, taking genes as units, and identifying the functional haplotype of the gene in the population through the genotype combination of the non-synonymous mutation sites contained in each gene.
[0007] Furthermore, the non-synonymous mutation sites on genes of different species of the same organism are obtained by aligning genome sequencing information of a certain organism population with a reference genome of the organism.
[0008] As a second aspect, the application of the functional haplotype marker in the identification of genetic characteristics of cotton germplasm resources is provided, including any of the following:
[0009] (1) The ratio of the number of genes with inconsistent functional haplotypes to the total number of genes in two samples was used as the genetic distance to determine the genetic relationship between cotton varieties;
[0010] (2) Use specific functional haplotype markers to identify specific cotton varieties.
[0011] As a third aspect, the application of the functional haplotype markers in the positioning of cotton agronomic trait genes is provided, and whole-genome association analysis is performed. A mixed linear model is established based on phenotypic data, fixed effects, functional haplotypes of samples in a single gene to be tested, and kinship effects. Each gene is tested, and genes with a p-value lower than a threshold value in the test results are selected as candidate genes that significantly affect the performance of agronomic traits.
[0012] In some embodiments, the mixed linear model is as follows:
[0013] y=Xβ+Zv+r+ε
[0014] Where y is an N×1 phenotypic value vector, N is the number of samples, X is an N×p matrix containing p covariates, β is a p×1 covariate effect vector, Z is an N×q functional haplotype matrix, the i-th row represents the functional haplotype of the i-th sample in the gene to be tested (i is an integer from 1 to N), q is the total number of functional haplotypes of the gene to be tested in the population, υ is a q×1 functional haplotype effect vector, and is normally distributed. A random vector, I is the q×q identity matrix, is the effect variance of the functional haplotype; r is the kinship effect vector, which is normally distributed A random vector of K, where K is an N×N matrix, representing the kinship matrix reflected by the functional haplotype. is the effect variance of kinship effect; ε is the error term, which obeys normal distribution is the error term variance.
[0015] Furthermore, in the functional haplotype matrix, the functional haplotype category of the i-th sample is converted into a 1×q dummy vector through dummy variable encoding as the i-th row of the functional haplotype matrix.
[0016] Furthermore, the covariates are obtained by performing principal component analysis on the genotype matrix of conventional additive coding of SNPs, and are used to reflect the potential population structure and population stratification characteristics of the biological population.
[0017] Furthermore, in the kinship matrix reflected by the functional haplotype, the kinship correlation coefficient k between sample i and sample j is ij (The elements of the kinship matrix K) are calculated using the following formula:
[0018] kij =1-d ij ;
[0019] Among them, d ij It represents the ratio of the number of genes with inconsistent functional haplotypes between sample i and sample j to the total number of genes.
[0020] The beneficial effects of the present invention are as follows:
[0021] The functional haplotypes provided by this invention are novel markers for gene analysis. They represent direct protein sequence changes caused by non-synonymous mutations and offer high coverage and specificity. This analytical method for functional haplotypes allows for the study of species domestication and improvement targets from the perspective of genes, key biological elements, enhancing the interpretability of genetic analysis results. Improved association analysis methods also enable the efficient localization of key genes for agronomic traits, providing a solid foundation for molecular biology research and precision crop breeding. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 The following is a method and flow chart for constructing functional haplotypes;
[0023] Figure 2 This is an example diagram of using functional haplotypes to identify the genetic characteristics of cotton germplasm resources; the heat map shows the functional haplotype types of genes located on chromosome A08, and different colors in the heat map represent different functional haplotypes; the two bar charts on the left of the heat map show the geographical origins of different germplasm resources and the clustering results based on functional haplotypes, respectively, and the pie chart on the right of the heat map shows the proportion of geographical origins of samples with different functional haplotype blocks. The geographical origin types represented by different colors in the bar charts and pie charts are shown in the legend below. DETAILED DESCRIPTION
[0024] The present invention provides a functional haplotype marker for cotton genes, and the construction method thereof is as follows: first, combining the annotation information of the gene coding region in the reference genome, the non-synonymous mutation sites of each gene are extracted as the subsequent construction objects. For each gene, the functional haplotype type of the gene is defined according to all the non-synonymous mutations in the gene (e.g. Figure 1 For example, a gene has two non-synonymous mutations (A / a, B / b) in its coding region. For each non-synonymous mutation, it has three possible genotypes (0, 1, 2, corresponding to genotypes AA, Aa, aa or BB, Bb, bb). Then the two non-synonymous mutations have 3 2That is, there are 9 possible combinations. Therefore, there are 9 possible functional haplotypes of this gene in the population, namely FH-1 (AABB), FH-2 (AaBB), FH-3 (aaBB), FH-4 (AABb), FH-5 (AaBb), FH-6 (aaBb), FH-7 (AAbb), FH-8 (Aabb), FH-9 (aabb). Similarly, when a gene has m non-synonymous mutations, there are 3 possible functional haplotypes of this gene in the population. m There are possible functional haplotypes, but the number of functional haplotypes of a gene will not exceed the number N of samples in the population.
[0025] Based on the functional haplotype concept and its construction method proposed in the present invention, 1,368,685 functional haplotypes were identified in 60,154 genes in 3,724 global cotton germplasm resource populations, with an average of 22 different types of functional haplotypes per gene in the population.
[0026] The present invention also provides a method for applying the above-mentioned functional haplotype markers to identify the genetic characteristics of cotton germplasm resources, characterized in that: for a cotton variety of unknown geographic origin or subspecies, the functional haplotypes of all genes in the variety are identified using the above-mentioned method. For example, the functional haplotype of gene 1 is FH-3, the functional haplotype of gene 2 is FH-7, and so on. Ultimately, a gene functional haplotype vector of the variety is obtained, with a vector size of 1×M, where M corresponds to the number of genes, which is 60154 in this example. Pairwise comparisons are made with the functional haplotypes of 3724 cotton germplasm resource populations of known origin to obtain the genetic distance d between the cotton variety of unknown origin and a cotton variety of known origin, calculated as follows:
[0027] d=D / M
[0028] Where D is the number of genes with discrepant functional haplotypes between the two samples being compared, and d ranges from 0 to 1. A larger genetic distance, d, indicates a greater number of genes with different sequence types in the two samples, and a greater genetic distance. When d is 1, the sequence types of all genes in the two samples are discrepant. Conversely, a smaller d value indicates greater genomic similarity and a closer genetic distance. When d is 0, the sequences of all genes in the two samples are identical. This method can be used to determine the genetic relationships between cotton varieties and clarify the genetic composition and subspecies affiliation of unknown cotton varieties. Furthermore, the use of highly specific functional haplotypes can also facilitate the genetic characterization of cotton germplasm resources. For example, if the functional haplotype FH-10 of a gene is only found in cotton varieties from Xinjiang, then another cotton variety that also has FH-10 for the same gene is highly likely to be from Xinjiang.
[0029] The present invention also provides a method for applying the above-mentioned functional haplotype markers to efficiently locate causal genes for agronomic traits. Since functional haplotypes describe sequence variation information at the gene level, association analysis based on functional haplotypes can directly locate causal genes for agronomic traits. However, unlike traditional numerical markers, functional haplotypes at the gene level are factor-type markers, so existing association analysis models need to be improved to meet the requirements of variable regression. The new association analysis model features are:
[0030] Define a single-marker mixed linear model:
[0031] y=Xβ+Zv+r+ε
[0032] Where y is an N×1 phenotypic value vector, X is an N×p matrix containing p covariates, and the covariates usually include the first p principal component vectors obtained by principal component analysis of the genotype matrix of conventional SNP additive coding (N×H matrix, H is the number of SNP sites), reflecting the potential population structure and population stratification characteristics of the biological population, Z is the N×1 functional haplotype vector corresponding to a gene, which contains the functional haplotype information of the gene in all samples; β is a p×1 covariate effect vector, which is a fixed effect; v is the effect vector of the functional haplotype, which is a normal distribution. A random vector of is the effect variance of the functional haplotype, I is the identity matrix; r is the kinship effect vector, which is normally distributed. A random vector of K, where K is an N×N matrix, representing the kinship matrix reflected by the functional haplotype. is the effect variance of kinship effect; ε is the error term, which obeys normal distribution is the error term variance.
[0033] In the kinship matrix, the kinship correlation coefficient k between sample i and sample j is ij The calculation is as follows:
[0034] k ij =1-d ij
[0035] where d ij is the genetic distance between sample i and sample j, ranging from 0 to 1. The greater the genetic distance, the closer the d ij The larger the value is, the more distant the relationship between samples is. ij The smaller the genetic distance, the closer the relationship between samples, and k ij The bigger.
[0036] For the Z vector, since the functional haplotype is a categorical variable, it cannot be used for direct regression and needs to be coded as a dummy variable. Assuming that a gene has q types in the population, a vector z with q all-zeros is defined:
[0037] z=[01,02,03,…,0 q-1 ,0 q ]
[0038] Assuming that the functional haplotype type of individual i in this gene is FH-s, then
[0039] z i =[01,02,…,0 s-1 ,1 s ,0 s+1 ,…,0 q ]
[0040] That is, the functional haplotype dummy variable vector for this sample is a vector of length q, with all elements 0 except the sth element being 1. Therefore, Z is recoded into an N × q matrix, with υ being a q × 1 effect vector.
[0041] Finally, define the null model H0:
[0042] y=Xβ+r+ε
[0043] and the alternative model H1:
[0044] y=Xβ+Zv+r+ε
[0045] Using likelihood ratio test to test the effect variance of functional haplotype Whether it is greater than 0, the likelihood ratio test degree of freedom is 1, and genes with a p-value less than 1e-5 are identified as candidate genes that significantly affect the performance of agronomic traits.
[0046] Example 1
[0047] A core cotton germplasm population with a broad genetic base and numerous breeding-advantaged loci was selected for planting. Young leaves were collected at the seedling stage and DNA was extracted using the CTAB (Cetyltrimethylammonium Bromide) method. DNA extraction was performed as follows: two sterilized 3mm grinding balls were placed in a 1.5mL centrifuge tube containing fresh tissue, along with 600μL of DNA lysis buffer (1% v / v β-mercaptoethanol). The sample was then ground in a sample shaker at 60Hz for 1 minute until thoroughly ground. After grinding, the sample was lysed in a 65°C waterbath for one hour, with the tube shaken manually several times to ensure adequate contact between the lysis buffer and the sample. After lysis, 600μL of chloroform was added to the tube for extraction. After thoroughly mixing the sample and reagents, the tube was centrifuged at 13,000 rpm for ten minutes. After centrifugation, the supernatant was transferred to a new centrifuge tube. 400 μL of pre-chilled isopropanol was added, and the tube was gently inverted and shaken at least 50 times until a white, flocculent precipitate was visible. The tube was then placed in a -20°C refrigerator for half an hour. The tube was then centrifuged again at 13,000 rpm for ten minutes. After centrifugation, the supernatant was discarded, the precipitate was retained, and 1 mL of ethanol was added and repeatedly pipetted to mix the precipitate with the reagent. The tube was then centrifuged again at 13,000 rpm for ten minutes, and this process was repeated at least once. After multiple centrifugations, the supernatant was discarded, the tube was placed in a fume hood to dry, 100 μL of TE buffer was added, and the precipitate was thoroughly shaken to dissolve. DNA was verified to be non-degradable by 2% agarose gel electrophoresis, and DNA quality was verified using a NanoDrop 2000 spectrophotometer. Finally, qualified DNA samples were stored at -80°C. After DNA extraction from all samples, the samples were sent for library construction and high-throughput whole-genome sequencing.
[0048] DNA samples that met high-quality standards were sequenced using the Illumina HiSeq 2500 platform with paired-end sequencing of 300-500 base pairs (bp) in length, achieving an average sequencing depth of at least 15x. The raw sequencing files were first quality-controlled and low-quality data cleaned using fastp software. High-quality sequencing reads were then aligned to the TM-1 reference genome, a genetic standard for upland cotton, using BWA software. The alignment results were then PCR-duplicate labeled using sambamba software. Finally, variant detection was performed using the mpileup and call commands in bcftools software. Variants with a call quality above 30 were selected for subsequent analysis. Next, all samples were grouped according to their geographic origin, and genotypes were imputed for each population with the same geographic origin using beagle software. Sites with an imputation accuracy greater than 80% were retained for subsequent analysis.
[0049] Based on the result files obtained from the above mutation detection, snpEFF software was first used to annotate the mutation types, extract non-synonymous mutations located in the exon region of the gene, and use the genotype combination of different non-synonymous mutations located in the same gene to identify the functional haplotype markers of the target germplasm resource population at the gene level, such as Figure 1 .
[0050] After obtaining the functional haplotypes of the target germplasm population, the functional haplotypes of the genes on chromosome A08 were extracted and a heat map was drawn, such as Figure 2 Based on the clustering of different functional haplotypes, six functional haplotype blocks were identified. Block 1 (A08-Hapblock 1) and the functional haplotype types contained in this block are unique to cotton varieties in Northwest China, indicating the unique genomic domestication imprint in cotton varieties in this region and the effective application of this functional haplotype block in variety identification.
[0051] Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A functional haplotype marker at the gene level, characterized in that: It is constructed by the following steps: Obtain the non-synonymous mutation sites on the genes of different varieties of the same organism, and identify the functional haplotype of the gene in the population through the genotype combination of the non-synonymous mutation sites contained in each gene, taking the gene as the unit.
2. The functional haplotype marker according to claim 1, characterized in that The non-synonymous mutation sites on genes of different species of the same organism are obtained by aligning the genome sequencing information of a certain organism population with the reference genome of the organism.
3. The use of the functional haplotype marker according to claim 1 in the identification of genetic characteristics of cotton germplasm resources, characterized in that: The application includes any of the following: (1) The ratio of the number of genes with inconsistent functional haplotypes to the total number of genes in two samples was used as the genetic distance to determine the genetic relationship between cotton varieties; (2) Use specific functional haplotype markers to identify specific cotton varieties.
4. Use of the functional haplotype marker according to claim 1 in the mapping of cotton agronomic trait genes, characterized in that: Genome-wide association analysis was performed, and mixed linear models were built based on phenotypic data, fixed effects, functional haplotypes of samples at individual genes to be tested, and kinship effects to test each gene.
5. The use according to claim 4, characterized in that The mixed linear model is as follows: y=Xβ+Zυ+r+ε Where y is an N×1 phenotypic value vector, N is the number of samples, X is an N×p matrix containing p covariates, β is a p×1 covariate effect vector, Z is an N×q functional haplotype matrix, the i-th row represents the functional haplotype of the i-th sample in the gene to be tested, q is the total number of functional haplotypes of the gene to be tested in the population, υ is a q×1 functional haplotype effect vector, and is normally distributed. A random vector, I is the q×q identity matrix, is the effect variance of the functional haplotype; r is the kinship effect vector, which is normally distributed A random vector of K, where K is an N×N matrix, representing the kinship matrix reflected by the functional haplotype. is the effect variance of kinship effect; ε is the error term, which obeys normal distribution is the error term variance.
6. The use according to claim 5, characterized in that In the functional haplotype matrix, the functional haplotype category of the i-th sample is converted into a 1×q dummy vector through dummy variable coding as the i-th row of the functional haplotype matrix.
7. The use according to claim 5, characterized in that The covariates are obtained by performing principal component analysis on the genotype matrix of conventional additive coding of SNPs, and are used to reflect the potential population structure and population stratification characteristics of the biological population.
8. The use according to claim 5, characterized in that In the kinship matrix reflected by the functional haplotype, the kinship correlation coefficient k between sample i and sample j is ij Calculated by the following formula: k ij =1-d ij ; Among them, d ij It represents the ratio of the number of genes with inconsistent functional haplotypes between sample i and sample j to the total number of genes.