SNP marker related to Tabebuia chrysantha seed morphology and application thereof
By using genome-wide association analysis technology to mine SNP markers related to seed morphology in Huanghua Fengchi, the problem of difficult to achieve efficient and accurate seed morphology classification in the existing technology is solved, and the genotype distinction and prediction of related gene expression are achieved, providing a scientific basis for the genetic improvement and classification of Huanghua Fengchi.
Patent Information
- Application Number
- CN202510687197.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The prior art is difficult to achieve efficient and accurate morphological classification and genetic analysis of yellow-flowered chinchilla seeds, and there is a lack of classification methods based on molecular markers.
Through genome-wide association analysis technology, SNP markers related to seed morphology, including SNP1, SNP4, SNP8, SNP18 and SNP24, were excavated from Huanghua Fengchi genomic DNA, to typify seed morphology traits and predict the expression of related genes.
The genotype-based distinction of seed morphological traits has been achieved, providing a scientific basis for the genetic improvement, classification and breeding of Huanghuafengchi, and improving the accuracy of the identification of seed morphological characteristics.
Smart Images

Figure CN120210423A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of molecular markers of Handroanthus chrysanthus. Specifically, it is about SNP markers related to the seed morphology of Handroanthus chrysanthus and their applications. Background Art
[0002] The morphological differences among some species of Handroanthus are small, which easily leads to species confusion and there has always been controversy in their classification. Moreover, in the existing literature, there are few reports on the genetic relationship and taxonomic research of Handroanthus chrysanthus ( Handroanthus chrysanthus ), especially there are still large gaps in systematically sorting out the germplasm resources of Handroanthus chrysanthus and revealing its genetic diversity.
[0003] As an important reproductive organ of plants, the morphological characteristics of seeds are of great significance in taxonomy. Although there are many successful cases of using seed morphology for taxonomic research in plant taxonomy at home and abroad, there is still a lack of research on classifying based on the seed morphology of Handroanthus chrysanthus.
[0004] The traditional method of classifying plants based on seed morphology relies on morphological observation and measurement, and it is difficult to achieve efficient and accurate classification and genetic analysis. Therefore, there is an urgent need to develop a classification method based on molecular markers to associate the seed morphology of Handroanthus chrysanthus with SNP markers and achieve the distinction of seed morphological traits based on genotypes. Summary of the Invention
[0005] For this reason, the technical problem to be solved by the present invention is to provide SNP markers related to the seed morphology of Handroanthus chrysanthus and their applications. Through genome-wide association analysis technology, SNP markers related to seed morphology have been mined from the genomic DNA of Handroanthus chrysanthus, realizing the distinction of seed morphological traits based on genotypes, and providing a scientific basis for the genetic improvement, classification and breeding of Handroanthus chrysanthus.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: SNP markers related to the seed morphology of Handroanthus chrysanthus, the SNP markers include SNP1, SNP4, SNP8, SNP18 and SNP24, and the SNP molecular markers are used for genotyping the seed morphological traits of Handroanthus chrysanthus Handroanthus chrysanthus ; the seed morphological traits are at least one of winged length, winged width, seed length and seed width; wherein: SNP1 is located at 3975 bp on chromosome NKXS01000048.1, the reference base is G, and the mutated base is T; SNP4 is located at 4051 bp on chromosome NKXS01000048.1, the reference base is T, and the mutated base is C; SNP8 is located at 3134 bp on chromosome NKXS01002837.1, the reference base is T, and the mutated base is G; SNP18 is located at 11016bp on chromosome NKXS01004877.1, the reference base is C, and the mutated base is T; SNP24 is located at 92361 bp on chromosome NKXS01001449.1, the reference base is G, and the mutated base is C. The position of each SNP on the chromosome was calculated based on the genome data of Campanula chinensis.
[0007] The application of SNP markers related to seed morphology of Trumpetwood, the application is to use the above-mentioned SNP markers to type the seed morphological traits of Trumpetwood; or, the application is to use the above-mentioned SNP markers to predict the expression levels of key genes related to seed morphology in Trumpetwood.
[0008] In the above application, when the SNP marker is used to type the seed morphological traits of Trumpetia chrysantha, the genotype of the Trumpetia chrysantha at the site where at least one SNP marker among SNP1, SNP4, SNP8, SNP18 and SNP24 is located is obtained from the genomic DNA of the Trumpetia chrysantha to be tested.
[0009] In the above application, when the SNP marker is used to type the seed length, the genotype of the tested Trumpetwood at the site SNP4 or SNP8 is obtained from the genomic DNA of the tested Trumpetwood; when the genotype at the site SNP4 is C, or when the genotype at the site SNP8 is G, the seeds of the tested Trumpetwood are long types; when the genotype at the site SNP4 is T, or when the genotype at the site SNP8 is T, the seeds of the tested Trumpetwood are short types.
[0010] For a plant of Tabebuia chrysantha to be tested, as long as one of the alleles at the site of SNP4 is a mutant type C, the genotype at the site of SNP4 is considered to be C, that is, a mutant type; when all the alleles at the site of SNP4 in the genome of Tabebuia chrysantha to be tested are all reference type T, the genotype at the site of SNP4 is considered to be T, that is, a reference type. The same applies to other SNPs.
[0011] In the above application, when the SNP marker is used to type the seed width, the genotype of the site where SNP18 of the tested Trumpetwood is located is obtained from the genomic DNA of the tested Trumpetwood; when the genotype of the site where SNP18 is located is T, the seeds of the tested Trumpetwood are of a wide type; when the genotype of the site where SNP18 is located is C, the seeds of the tested Trumpetwood are of a narrow type.
[0012] In the above application, when the SNP marker is used to type the wing length, the genotype of the yellow-flowered Trumpet tree to be tested at the site where SNP24 is located is obtained from the genomic DNA of the yellow-flowered Trumpet tree to be tested; when the genotype at the site where SNP24 is located is C, the seeds of the yellow-flowered Trumpet tree to be tested are long-winged type, and when the genotype at the site where SNP24 is located is G, the seeds of the yellow-flowered Trumpet tree to be tested are short-winged type.
[0013] In the above application, when the SNP marker is used to type the wing width, the genotype of the tested Trumpetwood at the site where SNP1 is located is obtained from the genomic DNA of the tested Trumpetwood; when the genotype at the site where SNP1 is located is T, the seeds of the tested Trumpetwood are wide-winged type, and when the genotype at the site where SNP1 is located is G, the seeds of the tested Trumpetwood are narrow-winged type.
[0014] In the above application, when the SNP marker is used to simultaneously genotype seed length, seed width and wing width, the genotype of the tested Campanula truncatula at the site of SNP1 or SNP4 is obtained from the genomic DNA of the tested Campanula truncatula; When the genotype of the site where SNP1 is located is T, the seeds of the tested yellow bell tree are long, wide and wide-winged, and when the genotype of the site where SNP1 is located is G, the seeds of the tested yellow bell tree are short, narrow and narrow-winged; Alternatively, when the genotype of the site where SNP4 is located is C, the seeds of the tested Tabebuia chrysantha are long, wide and have wide wings; when the genotype of the site where SNP4 is located is T, the seeds of the tested Tabebuia chrysantha are short, narrow and have narrow wings.
[0015] In the above application, the key gene related to seed morphology in Campanula lutea is CDL12_15522; When predicting the expression level of the CDL12_15522 gene, the genotype of the tested Trumpetwood at the site of SNP8 is obtained from the genomic DNA of the tested Trumpetwood; when the genotype at the site of SNP8 is G, the CDL12_15522 gene is highly expressed; when the genotype at the site of SNP8 is T, the CDL12_15522 gene is lowly expressed.
[0016] In the above application, when obtaining the genotype of the tested Trumpetwood at any SNP marker site among SNP1, SNP4, SNP8, SNP18 and SNP24, PCR is performed using the genomic DNA of the tested Trumpetwood as the template DNA, and the PCR product is sequenced to obtain the genotype of the tested Trumpetwood at the site where the corresponding SNP marker is located.
[0017] The CDL12_15522 gene is located in the flanking region of SNP8, and they are tightly linked. Experimental results have shown that when the genotypes at SNP8 are different, the expression levels of the CDL12_15522 gene are also different; that is to say, by measuring the expression level of the CDL12_15522 gene, the seed morphology of Tabebuia chrysantha can also be predicted.
[0018] The technical solution of the present invention has achieved the following beneficial technical effects: 1. The present invention first proposed SNP markers related to the seed morphology of Tabebuia chrysantha, providing an important basis for the research on the molecular genetic mechanism of the seed morphological traits of Tabebuia chrysantha. These SNP molecular markers contribute to the rapid and accurate identification of the seed morphological characteristics of Tabebuia chrysantha, and the accurate identification of seed morphological characteristics provides a basis for the taxonomic research of Tabebuia chrysantha.
[0019] 2. The present invention has also screened candidate genes related to the seed morphology of Tabebuia chrysantha. Among them, the CDL12_15522 gene is located in the flanking region of the SNP marker related to the seed morphology of Tabebuia chrysantha. qRT-PCR experiments have shown that among the individuals of Tabebuia chrysantha with significant differences in seed morphology, the expression levels of the candidate gene CDL12_15522 also vary greatly, that is, CDL12_15522 may have the function of regulating the seed morphology of Tabebuia chrysantha, providing a new clue for the research on the genes regulating the seed morphology of Tabebuia chrysantha. Description of the Drawings
[0020] Figure 1 Frequency distribution diagram of the measured results of the seed lengths of 126 individuals of Tabebuia chrysantha in the embodiment of the present invention; Figure 2 Frequency distribution diagram of the measured results of the seed widths of 126 individuals of Tabebuia chrysantha in the embodiment of the present invention; Figure 3 Frequency distribution diagram of the measured results of the winged lengths of 126 individuals of Tabebuia chrysantha in the embodiment of the present invention; Figure 4 Frequency distribution diagram of the measured results of the winged widths of 126 individuals of Tabebuia chrysantha in the embodiment of the present invention; Figure 5 Correlation analysis result diagram of the seed length, seed width, winged length and winged width of Tabebuia chrysantha in the embodiment of the present invention; Figure 6 Optimal K value analysis result of the population composed of 126 individuals of Tabebuia chrysantha in the embodiment of the present invention; Figure 7 Column chart of the genetic composition of the sample composed of 126 individuals of Tabebuia chrysantha in the embodiment of the present invention; Figure 8Neighbor-joining phylogenetic tree of 126 individuals of Tabebuia chrysantha in the embodiments of the present invention; Figure 9 Principal component analysis distribution map of 126 individuals of Tabebuia chrysantha in the embodiments of the present invention; Figure 10 QQ plot of four models for seed length of 126 individuals of Tabebuia chrysantha in the embodiments of the present invention; Figure 11 QQ plot of four models for seed width of 126 individuals of Tabebuia chrysantha in the embodiments of the present invention; Figure 12 QQ plot of four models for winged length of 126 individuals of Tabebuia chrysantha in the embodiments of the present invention; Figure 13 QQ plot of four models for winged width of 126 individuals of Tabebuia chrysantha in the embodiments of the present invention; Figure 14 Manhattan plot of GWAS analysis results for seed length of 126 individuals of Tabebuia chrysantha in the embodiments of the present invention; Figure 15 GWAS analysis results for seed width of 126 individuals of Tabebuia chrysantha in the embodiments of the present invention; Figure 16 GWAS analysis results for winged length of 126 individuals of Tabebuia chrysantha in the embodiments of the present invention; Figure 17 GWAS analysis results for winged width of 126 individuals of Tabebuia chrysantha in the embodiments of the present invention; Figure 18 QQ plot of the MLM(QK) model for seed length of 126 individuals of Tabebuia chrysantha in the embodiments of the present invention; Figure 19 QQ plot of the MLM(QK) model for seed width of 126 individuals of Tabebuia chrysantha in the embodiments of the present invention; Figure 20 QQ plot of the MLM(QK) model for winged length of 126 individuals of Tabebuia chrysantha in the embodiments of the present invention; Figure 21 QQ plot of the MLM(QK) model for winged width of 126 individuals of Tabebuia chrysantha in the embodiments of the present invention; Figure 22 Venn diagram of SNPs controlling different seed morphological traits in the embodiments of the present invention; Figure 23 Venn diagram of candidate genes within the flanking regions of each SNP in the embodiments of the present invention; Figure 24Comparison chart of the expression levels of 9 key candidate genes in four individuals of Tabebuia chrysantha, namely F1, F4, M1, and M6, in the embodiments of the present invention. Detailed implementation manners
[0021] 1. Materials and methods 1.1 Experimental materials 126 germplasm resources of Tabebuia chrysantha collected from 12 regions in Guangdong Province were used as the associated population materials. These materials were from Zhanjiang, Maoming, Yangjiang, Zhaoqing, and Jiangmen. The specific information of the associated population is shown in Table 1.
[0022] Table 1 Geographic information of 126 Tabebuia chrysantha samples 1.2 Determination of phenotypic data When the seeds of the above 126 germplasm resources (126 Tabebuia chrysantha plants) were mature, 200 plump and healthy seeds were collected from each Tabebuia chrysantha plant, and 6 seeds were randomly selected from these 200 seeds to measure the seed morphological traits. The seeds of Tabebuia chrysantha have a thin film-like wing on the periphery of the seeds. When measuring, the following four indexes were measured with a vernier caliper: wing length (WL, corresponding classification: long-wing type / short-wing type), wing width (WW, corresponding classification: wide-wing type / narrow-wing type), seed length (Non Wing Length, NWL, that is, the seed length after removing the wing, corresponding classification: long-seed type / short-seed type), and seed width (Non Wing Width, NWW, that is, the seed width after removing the wing, corresponding classification: wide-seed type / narrow-seed type). Each seed was measured independently three times, and the average value was taken as the final data. The R software (version 3.5.0) was used to perform relevant statistical analysis on the measured phenotypic data, and further perform variance analysis and correlation analysis on the phenotypic data.
[0023] 1.3 Population structure and kinship analysis Since the whole genome sequencing of Tabebuia chrysantha has not been carried out, and only Handroanthus impetiginosus ( Handroanthus impetiginosus ) in the Tabebuia plants has been sequenced for the whole genome, the molecular determination (SNP, transcriptome, etc.) of the Tabebuia plants uses the genome of Handroanthus impetiginosus as the reference genome for alignment. 126 individuals of Tabebuia chrysantha were genotyped by the GBS (Genotyping-by-Sequencing) technology. The specific method is as follows: high-throughput sequencing was carried out on 126 individuals of Tabebuia chrysantha, and the obtained original image data files were converted into original sequencing sequences through base recognition analysis. After quality control of the original sequences and filtering of low-quality sequences, high-quality sequences were obtained. The obtained high-quality sequences were aligned with the published reference genome of Handroanthus impetiginosus to obtain genome-wide SNP markers.
[0024] Based on the genome-wide SNP markers filtered by linkage disequilibrium (LD), population structure analysis was performed using Admixture software (version 1.3). To determine the optimal number of subpopulations (K value), the range of the assumed number of subpopulations of the samples was set from 1 to 9, and clustering analysis was carried out. The K value with the minimum cross-validation error (CV error) was used as the optimal number of subgroupings. Subsequently, the Admixture software was run again. The Admixture software calculated the genetic composition of each sample based on the selected optimal K value and the SNP markers filtered by linkage disequilibrium (LD), and obtained the proportion of the genetic composition of each sample in each subgroup (that is, the proportion of the genetic material inherited by the sample from each subgroup).
[0025] The Admixture software will output a population structure matrix (Q matrix). Each row of the Q matrix represents a sample, each column represents a subgroup, and the elements in the matrix (i.e., Q values) represent the proportion of the genetic material inherited by the sample from the corresponding subgroup.
[0026] Finally, using pophelper software (version 2.2.7), according to the output results of Admixture, a bar chart of the genetic composition of each sample in each subgroup was drawn.
[0027] Using the filtered SNP markers (i.e., the above-mentioned genome-wide SNP markers), a phylogenetic tree based on the neighbor-joining (NJ) method (model: p-distance; bootstrap: 500 times) was constructed through MEGA-X software. In addition, GCTA software (version 1.93.2) was used for principal component analysis (PCA) to obtain the variance explanation rate of each principal component (PC) and the score matrix of the samples in each principal component. Finally, TASSEL 5.0 software was used for kinship analysis to obtain the kinship matrix between pairwise samples.
[0028] 1.4 Genome-wide association analysis Using the whole-genome SNP molecular markers obtained by aligning with the reference genome of Handroanthus impetiginosus described above, the SNP loci with a minor allele frequency (MAF) ≥ 0.05 were used for genome-wide association study (GWAS). The GWAS was performed using the GEMMA software (version 0.98.1). Common models in GWAS include the simple generalized linear model (GLM), the generalized linear model with the Q matrix as a covariate [GLM(Q)], the mixed linear model with the K matrix as a covariate [MLM(K)], and the mixed linear model with the Q matrix and the K matrix as covariates [MLM(QK)], etc. In this example, the population structure matrix corresponding to the optimal K value obtained by the Admixture software analysis in "1.3 Population Structure and Kinship Analysis" was used as the Q matrix, and the kinship matrix between samples was used as the K matrix. The above four models were respectively used to perform GWAS on various seed morphology-related traits, and the QQ plot was used to compare the distribution of the actual P values and the theoretical P values under different models to determine the optimal model.
[0029] After determining the optimal model, the Bonferroni multiple test correction method was used to determine the significance threshold of the P values, and the significantly associated regions were screened. Further, the SNP with the strongest association signal (i.e., the top-associated SNP) was screened from the significantly associated regions for subsequent association site analysis.
[0030] 1.5 Candidate Gene Analysis of Related Loci After performing the genome-wide association study, the published Handroanthus impetiginosus genome database (https: / / www.ncbi.nlm.nih.gov / datasets / taxonomy / 429701 / ) was used in combination with the BLAST software for alignment to screen the SNP markers significantly associated with the seed morphology traits. Under the condition of E-value (E-Value) ≤ 1e -10 The significantly associated SNP markers were aligned and mapped to the Handroanthus impetiginosus genome, that is, the significantly associated SNP markers were located on the Handroanthus impetiginosus genome, and the related genes were searched within the flanking regions of 50 kb upstream and downstream of them to obtain candidate genes. GO and KEGG analyses were performed on the candidate genes. Functional annotations were performed on the variations such as SNPs in the coding regions of the candidate genes.
[0031] 1.6 qRT-PCR Verification Select the leaves of four individuals, namely F1, F4, M1, and M6, with significantly different seed morphological phenotypes (F and M are location numbers, and the numbers represent individual numbers). Use an RNA extraction kit (TaKaRa MiniBEST Universal RNA Extraction Kit) to extract RNA, and use a reverse transcription kit (TUREscript 1stStand cDNA SYNTHESIS Kit) to transcribe the RNA into cDNA. Design primers for the above candidate genes and perform qRT-PCR detection using an ABI 7500 real-time quantitative PCR instrument. Set three technical replicates for each sample (i.e., each individual). Use the 2 −ΔΔCt method to calculate the relative expression levels of the candidate genes. When performing qRT-PCR detection, use the 18S gene as the internal reference gene.
[0032] 2. Results and Analysis 2.1 Sequencing Quality Perform high-throughput sequencing on 126 individuals of Tabebuia chrysantha using the GBS technique. The obtained raw image data files are converted into raw sequencing sequences through base recognition analysis. After quality control of the raw sequences and filtering of low-quality sequences, the number of high-quality sequences obtained ranges from 4,677,442 to 2,378,244. After quality control, the Q30 value range of the high-quality sequences is 87.90% to 94.13%, the Q20 value range is 95.01% to 98.08%, and the average GC content is 38.19%. The alignment rate of the population samples with the reference genome of Tabebuia pentaphylla ranges from 81.04% to 95.11%, and the average alignment rate is 88.65%. The above results indicate that there is a high similarity between the samples used in the experiment and the reference genome (Tabebuia pentaphylla genome), and the sequencing quality is high. After filtering low-quality sequences (sites), a total of 131,559 high-quality SNPs are finally retained for subsequent analysis.
[0033] 2.2 Phenotypic Analysis Descriptive statistical analysis was performed on four seed morphology-related traits in the association population composed of 126 individuals of *Tabebuia chrysantha* (see Table 2). The analysis results showed that significant differences were exhibited among different individuals for each trait, and the variation was relatively small. Specifically, the variation range of seed length was from 8.14 mm to 15.98 mm, with an average value of 11.96 mm and a coefficient of variation of 7.13%; the variation range of seed width was from 5.51 mm to 10.00 mm, with an average value of 7.77 mm and a coefficient of variation of 8.94%; the variation range of winged length was from 20.10 mm to 31.66 mm, with an average value of 26.52 mm and a coefficient of variation of 10.94%; the variation range of winged width was from 6.60 mm to 11.11 mm, with an average value of 8.86 mm and a coefficient of variation of 9.41%. Further analysis found that there were highly significant correlations among the four traits of NWL, NWW, WL, and WW (see Figure 5 ). As Figure 1 , Figure 2 , Figure 3 and Figure 4 are the frequency distribution diagrams of the four traits of NWL, NWW, WL, and WW respectively. The unit of the horizontal axis in the figure is mm. It can be seen from the trait frequency distribution diagrams that these four seed morphology traits show a continuous normal distribution, indicating that these traits are quantitative traits controlled by minor polygenes and are suitable for genome-wide association analysis.
[0034] Table 2 Statistical and variance analysis of seed morphology traits of 126 *Tabebuia chrysantha* plants 2.3 Population structure and genetic relationship Before performing genome-wide association analysis, 131,559 SNP markers that densely cover the whole genome of *Tabebuia chrysantha* (i.e., the SNP markers finally screened in "2.1 Sequencing quality") were used to analyze the population structure and genetic relationship among population materials. As Figure 6 is the result of population structure analysis. It can be seen from the figure that when K = 2, the ΔK value is the smallest. Therefore, the 126 samples were finally divided into two subpopulations. Among them, one subpopulation contains 113 samples, while the other smaller subpopulation only contains 13 samples. Figure 7 is the bar chart of the genetic composition of 126 samples. The result of constructing the neighbor-joining phylogenetic tree is consistent with the population structure analysis, and the 126 samples are divided into two clusters (see Figure 8 ). The results of principal component analysis (PCA) showed that the first two principal components explained 10.03% and 2.43% of the genetic variance, and the 126 *Tabebuia chrysantha* samples were divided into two subpopulations (see Figure 9), reflecting a certain degree of genetic differentiation between these two groups. Therefore, in the subsequent association analysis of seed morphological traits and SNP markers, the Q matrix generated when K = 2 was used as population structure information.
[0035] In addition, Table 3 shows the results of the kinship analysis.
[0036] Table 3 Statistical analysis of kinship among samples As can be seen from the table, among the 126 individuals of Tabebuia chrysantha, 80.30% of the kinship coefficients are distributed between 0 and 0.2, indicating that the kinship among most individuals is relatively distant. 17.45% of the kinship coefficients are distributed between -0.6 and -0.4, indicating that there is a weak kinship among some individuals. In addition, 2.22% of the kinship coefficients are greater than 0.6 or less than -0.6, indicating that there is a relatively high similarity and close kinship among a few individuals. Generally speaking, among the 126 individuals of Tabebuia chrysantha, very few individuals show a high degree of similarity, while most individuals have a relatively distant kinship or even no kinship. This result meets the requirements for performing genome-wide association analysis.
[0037] 2.4 Selection of association analysis models In this embodiment, four models, namely GLM, GLM(Q), MLM(K), and MLM(QK), were used to perform genome-wide association analysis on seed morphology-related traits. Figure 10 、 Figure 11 、 Figure 12 and Figure 13 are the QQplot graphs of four traits, namely NWL, NWW, WL, and WW, respectively. Through the analysis of the QQ plot graphs of each trait, it can be seen that for the four traits of NWL, NWW, WL, and WW, the GLM model and the MLM(K) model have poor control effects on false positives; while the GLM(Q) model and the MLM(QK) model have better control effects on false positives after incorporating population structure and kinship among materials into the analysis, although the control of false positives for some traits is slightly strict. Further comparison found that the distribution of -log(P) values of the MLM(QK) model is closest to the predicted value (i.e., closest to the dotted line in the figure). Therefore, the MLM(QK) model was finally selected as the final analysis model for genome-wide association analysis to locate significantly associated loci and perform subsequent analysis.
[0038] 2.5 Association analysis of SNP markers and phenotypic traits Taking P ≤ 5.0×10⁻ 5 as the significant association threshold, GWAS analysis was performed on the phenotypes (observed values) of each trait and SNP markers in 126 Tabebuia chrysantha materials. As Figure 14 、 Figure 15, Figure 16 and Figure 17 are the Manhattan plots of the GWAS analysis results for the four traits of NWL, NWW, WL, and WW, respectively. The red dashed line in the figure represents the significance threshold, and the blue dashed line represents the suggested significance level. Figure 18 , Figure 19 , Figure 20 and Figure 21 are the QQplot plots of the MLM(QK) model for the GWAS analysis of the four traits of NWL, NWW, WL, and WW, respectively. The red diagonal line in the figure represents the diagonal line when the expected -log10(p) is used as the horizontal and vertical coordinates. When the points deviate upward from the red diagonal line, it indicates that there are loci non - randomly related to the trait. The points (in the Manhattan plot) that exceed the given threshold in the non - random correlation can be considered significantly associated loci.
[0039] As shown in Tables 4 - 1 to 4 - 3, in the association analysis of the four traits related to seed morphology, a total of 29 significant SNP loci were detected (including 9 pleiotropic SNPs. In the table, "-" indicates that there are many heterozygous peaks in the sequencing results, making it difficult to accurately determine the base at this locus). The range of phenotypic variation (R²) explained by a single associated marker locus is from 6.61% to 95.05%. Among them, 14 loci are significantly associated with seed length (NWL); 16 loci are significantly associated with seed width (NWW); 2 loci are significantly associated with winged length (WL); 9 loci are significantly associated with winged width (WW). Among them: The SNP locus most significantly associated with NWL is NKXS01000048.1__4051, located at 4051 bp of NKXS01000048.1, with a P - value of 8.66E - 08, which can explain 31.92% of the phenotypic variation. The reference base at NKXS01000048.1__4051 is T, and the mutant base is C.
[0040] The SNP locus most significantly associated with NWW is NKXS01004877.1__11016, located at 11016 bp of NKXS01004877.1, with a P - value of 7.12E - 06, which can explain 10.64% of the phenotypic variation. The reference base at NKXS01004877.1__11016 is C, and the mutant base is T.
[0041] The SNP locus most significantly associated with WL is NKXS01001449.1__92361, located at 92361 bp of NKXS01001449.1, with a P - value of 1.77E - 05, which can explain 26.95% of the phenotypic variation. The reference base at NKXS01001449.1__92361 is G, and the mutant base is C.
[0042] The SNP locus most significantly associated with WW is NKXS01000048.1__3975, which is located at 3975 bp of NKXS01000048.1. The P-value is 1.93E-06, and it can explain 29.38% of the phenotypic variation. The reference base at NKXS01000048.1__3975 is G, and the mutant base is T.
[0043] By comparing the associated SNPs of four seed morphology-related traits, 9 pleiotropic SNPs were found. Among them, 3 pleiotropic SNPs were found in total between NWL, NWW, and WW; 4 pleiotropic SNPs were found in total between NWL and NWW; 1 pleiotropic SNP was found between NWL and WW; and 1 pleiotropic SNP was also found between NWW and WW ( Figure 22 ). After Bonferroni multiple test correction, a total of 2 SNPs were detected to be significantly associated with NWL, namely NKXS01000048.1__4051 and NKXS01002837.1__3134. These two loci have strong association signals with NWL and may be the major loci regulating this trait.
[0044] Table 4-1 SNP loci significantly associated with the seed morphology trait NWL detected in the association analysis Table 4-2 SNP loci significantly associated with the seed morphology trait NWW detected in the association analysis Table 4-3 SNP loci significantly associated with the seed morphology traits WL and WW detected in the association analysis SNP1 and SNP4 are located on the same chromosome, and the physical distance between them is 76 bp. Further, the relationship between the genotypes of some individuals at SNP1 and SNP4 and the measurement results of NWL, NWW, and WW was statistically analyzed. The statistical results showed that the mean value of NWL of mutant individuals (both are mutant at SNP1 and SNP4) was 14.36 mm, the mean value of NWW was 9.21 mm, and the mean value of WW was 10.15 mm. While the mean value of NWL of wild-type individuals (both are reference type at SNP1 and SNP4) was 11.39 mm, the mean value of NWW was 7.57 mm, and the mean value of WW was 8.47 mm. That is, the mean values of NWL, NWW, and WW of individuals with mutant type at SNP1 and SNP4 loci are greater than those of wild-type individuals. The relationship between the genotype at SNP1 and the measurement result of WW was statistically analyzed separately, and the relationship between the genotype at SNP4 and the measurement result of NWL was statistically analyzed separately, and the results are the same as above, which will not be elaborated here.
[0045] The relationship between the genotypes of some individuals at SNP8 and the NWL measurement results was statistically analyzed. The statistical results showed that the average NWL of individuals with a mutant genotype at SNP8 was 14.88 mm, and the average NWL of individuals with a wild-type genotype was 11.39 mm. That is, the NWL of individuals with a mutant genotype at SNP8 was greater than that of wild-type individuals.
[0046] Similar results were also obtained at SNP18 and SNP24. The average NWW of individuals with a mutant genotype at SNP18 was 9.34 mm, and the average NWW of individuals with a wild-type genotype was 7.57 mm; the average WL of individuals with a mutant genotype at SNP24 was 29.82 mm, and the average WL of individuals with a wild-type genotype was 26.66 mm.
[0047] The above results indicate that based on the genotypes at the SNP1, SNP4, SNP8, SNP18, and SNP24 loci, the genotyping of Tabebuia seeds can be completed relatively accurately.
[0048] 2.6 Candidate gene mapping and GO function analysis Through genome-wide association study (GWAS) of four seed morphological traits of Tabebuia chrysantha, combined with the published genome sequencing results of Tabebuia chrysantha, the SNP markers significantly associated with seed morphological traits were mapped to the genome of Tabebuia pentaphylla. Within the linkage disequilibrium (LD) decay distance range of 50 kb (R² = 0.1), a search for related genes was conducted for the 29 SNP loci significantly related to seed morphology in Tables 4-1 to 4-3, and a total of 68 related genes (these 68 related genes are candidate genes) were screened out. Among them, 29 candidate genes with functional annotations are shown in Tables 5-1 to 5-3.
[0049] Table 5-1 Candidate genes corresponding to SNP loci significantly associated with NWL Table 5-2 Candidate genes corresponding to SNP loci significantly associated with NWW Table 5-3 Candidate genes corresponding to SNP loci significantly associated with WL and WW As can be seen from Tables 5-1 to 5-3, a total of 31 candidate genes were screened out among the 14 significant SNP loci related to the seed length (NWL), and 9 of these genes had functional annotations. Among the 16 significant SNP loci related to the seed width (NWW), a total of 28 candidate genes were screened out, and 18 of these genes had relatively clear functional annotations. Among the 2 significant SNP loci related to the winged length (WL), a total of 7 candidate genes were screened out, and 3 of these genes had functional annotations. Among the 9 significant SNP loci related to the winged width (WW), a total of 6 candidate genes were screened out, and 3 of these genes had functional annotations.
[0050] In addition, 1 common candidate gene was mapped in each of NWL, NWW, and WW; 2 common candidate genes were mapped in each of NWL and NWW ( Figure 23 ).
[0051] In this example, a batch of candidate genes associated with the seed morphology of Tabebuia chrysantha were first discovered using genome-wide association analysis, providing a basis for further studying the molecular regulation mechanism of the seed morphology of Tabebuia chrysantha.
[0052] 2.7 qRT-PCR verification results The measurement results of the seed morphological traits of four individuals, F1, F4, M1, and M6, and the genotypes of the sites where each SNP marker is located are shown in Table 6.
[0053] Table 6 Seed morphology and SNP types of F1, F4, M1, and M6 To further verify the expression of candidate genes in Tabebuia chrysantha with different seed morphological phenotypes, 9 key candidate genes (6 functional genes and 3 genes without annotated functions) were selected, including CDL12_00443, CDL12_15521, CDL12_15522, CDL12_15254, CDL12_04662, CDL12_04666, CDL12_11100, CDL12_11105, and CDL12_03077. Using the primers shown in Table 7, qRT-PCR was used to measure the expression levels of each gene.
[0054] Among these 9 key candidate genes, CDL12_00443 is located in the flanking regions of both NKXS01000048.1__3975 (SNP1) and NKXS01000048.1__4051 (SNP4), CDL12_15521 and CDL12_15522 are located in the flanking region of NKXS01002837.1__3134 (SNP8), and CDL12_15254 is located in the flanking region of NKXS01002782.1__29157 (SNP24). These SNPs are very likely to regulate the expression of genes in their respective flanking regions. And SNP1, SNP4, SNP8, and SNP24 are all SNP loci closely associated with the seed morphology of Tabebuia chrysantha. Each SNP locus is tightly linked to the gene in its flanking region. The purpose of qRT-PCR is to verify whether the genotypes at these SNP loci are closely related to the expression levels of genes in their flanking regions.
[0055] The results of qRT-PCR showed that the expression levels of different candidate genes showed expression differences in different phenotypes. Among them, the expression level of the candidate gene CDL12_15522 showed significant differences in seeds of different morphologies (see Figure 24 , P<0.05). According to the expression level of the candidate gene CDL12_15522, the seed morphology of Tabebuia chrysantha can also be predicted. This candidate gene may have the function of regulating the seed morphology of Tabebuia chrysantha.
[0056] Table 7 Primer information for each gene 3. Results and Discussion Seed morphological traits are relatively stable genetic characteristics of plants and are of great value in plant classification and genetics research. Seed morphology plays an important role in classification at the genus, species, and even subspecies levels.
[0057] According to the statistical and analysis results of the seed morphology-related traits of 126 individuals of Tabebuia chrysantha from 12 regions, the seed morphology differences among the 126 individuals were extremely significant, and the coefficient of variation of each shape was relatively small. Therefore, seed morphology can be used as a relatively stable morphological characteristic and important basis in the taxonomy of Tabebuia chrysantha.
[0058] Under natural conditions, due to the influence of genetic factors and environmental conditions, the four morphological indicators of the winged length (WL), winged width (WW), seed length (NWL), and seed width (NWW) of Tabebuia chrysantha seeds showed large differences among plant individuals. Correlation analysis showed that there was an extremely significant positive correlation among these four morphological indicators, indicating that these traits influenced each other and changed synergistically, reflecting the close consistency in genetics and environmental adaptation.
[0059] When using seed morphology as a classification basis for genera or species, it is necessary to accurately distinguish individuals with relative traits. Most of the shapes related to seed morphology are quantitative traits, and it is difficult to accurately distinguish individuals with relative traits only through naked-eye observation or ordinary measurement methods. Therefore, developing molecular markers related to seed morphology, associating seed morphological traits with the genotypes of molecular markers, and distinguishing individuals with relative traits through genotypes can enable seed morphological traits to be better applied to classification.
[0060] In the above examples, based on the MLM(QK) model, 131,559 SNP markers were used to perform a genome-wide association analysis on four seed morphological traits (NWL, NWW, WL, WW), and 29 SNP loci significantly associated with seed morphological traits were detected. Further analysis found that 9 SNPs showed pleiotropy, and three SNP loci, namely NKXS01000048.1__3975, NKXS01000048.1__4022, and NKXS01000048.1__4051, were simultaneously detected in the three traits of NWL, NWW, and WW.
[0061] After Bonferroni multiple test correction, it was found that two SNPs, NKXS01000048.1__4051 and NKXS01002837.1__3134, had strong association signals with NWL, and these loci may be the major loci regulating the NWL trait.
[0062] A total of 68 candidate genes were detected in the 50bp regions upstream and downstream of the 29 SNP loci, and 29 of these genes had functional annotations. The study found that CDL12_00443, CDL12_15521, and CDL12_15522 showed pleiotropy. The CDL12_00443 gene was detected in all three traits of NWL, NWW, and WW, and the CDL12_15521 and CDL12_15522 genes were detected in both traits of NWL and NWW.
Claims
1. SNP markers related to the seed morphology of Tabebuia chrysantha, characterized in that, The SNP markers include SNP1, SNP4, SNP8, SNP18, and SNP24, and the SNP molecular markers are used for genotyping the seed morphological traits of Tabebuia chrysantha Handroanthus chrysanthus ; the seed morphological traits are at least one of the winged length, winged width, seed length, and seed width; wherein: SNP1 is located at 3975 bp on chromosome NKXS01000048.1, with the reference base being G and the mutated base being T; SNP4 is located at 4051 bp on chromosome NKXS01000048.1, with the reference base being T and the mutated base being C; SNP8 is located at 3134 bp on chromosome NKXS01002837.1, with the reference base being T and the mutated base being G; SNP18 is located at 11016 bp on chromosome NKXS01004877.1, with the reference base being C and the mutated base being T; SNP24 is located at 92361 bp on chromosome NKXS01001449.1, with the reference base being G and the mutated base being C.
2. Application of SNP markers related to the seed morphology of Handroanthus chrysanthus, characterized in that, The application is to use the SNP markers as described in claim 1 for genotyping the seed morphological traits of Handroanthus chrysanthus; or, the application is to use the SNP markers as described in claim 1 for predicting the expression levels of key genes related to seed morphology in Handroanthus chrysanthus.
3. The application according to claim 2, characterized in that, When using the SNP markers for genotyping the seed morphological traits of Handroanthus chrysanthus, obtain the genotype at the locus of at least one of the SNP markers SNP1, SNP4, SNP8, SNP18, and SNP24 from the genomic DNA of the Handroanthus chrysanthus to be tested.
4. The application according to claim 3, wherein When using the SNP markers for genotyping the seed length, obtain the genotype at the locus of SNP4 or SNP8 from the genomic DNA of the Handroanthus chrysanthus to be tested; when the genotype at the locus of SNP4 is C, or when the genotype at the locus of SNP8 is G, the seeds of the Handroanthus chrysanthus to be tested are of the long-seed type; when the genotype at the locus of SNP4 is T, or when the genotype at the locus of SNP8 is T, the seeds of the Handroanthus chrysanthus to be tested are of the short-seed type.
5. The application according to claim 3, wherein When using the SNP markers for genotyping the seed width, obtain the genotype at the locus of SNP18 from the genomic DNA of the Handroanthus chrysanthus to be tested; when the genotype at the locus of SNP18 is T, the seeds of the Handroanthus chrysanthus to be tested are of the wide-seed type; when the genotype at the locus of SNP18 is C, the seeds of the Handroanthus chrysanthus to be tested are of the narrow-seed type.
6. The application according to claim 3, wherein When using the SNP markers for genotyping the wing length, obtain the genotype at the locus of SNP24 from the genomic DNA of the Handroanthus chrysanthus to be tested; when the genotype at the locus of SNP24 is C, the seeds of the Handroanthus chrysanthus to be tested are of the long-wing type, and when the genotype at the locus of SNP24 is G, the seeds of the Handroanthus chrysanthus to be tested are of the short-wing type.
7. The application according to claim 3, characterized in that, When using the SNP markers for genotyping the wing width, obtain the genotype at the locus of SNP1 from the genomic DNA of the Handroanthus chrysanthus to be tested; when the genotype at the locus of SNP1 is T, the seeds of the Handroanthus chrysanthus to be tested are of the wide-wing type, and when the genotype at the locus of SNP1 is G, the seeds of the Handroanthus chrysanthus to be tested are of the narrow-wing type.
8. The application according to claim 3, characterized in that, When the SNP marker is used to simultaneously genotype seed length, seed width and wing width, the genotype of the tested Campanula truncatula at the site of SNP1 or SNP4 is obtained from the genomic DNA of the tested Campanula truncatula; When the genotype of the site where SNP1 is located is T, the seeds of the tested yellow bell tree are long, wide and wide-winged, and when the genotype of the site where SNP1 is located is G, the seeds of the tested yellow bell tree are short, narrow and narrow-winged; Alternatively, when the genotype of the site where SNP4 is located is C, the seeds of the tested Tabebuia chrysantha are long, wide and have wide wings; when the genotype of the site where SNP4 is located is T, the seeds of the tested Tabebuia chrysantha are short, narrow and have narrow wings.
9. The application according to claim 2, wherein The key gene related to seed morphology in Campanula lutea is CDL12_15522; When predicting the expression level of the CDL12_15522 gene, the genotype of the tested Trumpetwood at the site of SNP8 is obtained from the genomic DNA of the tested Trumpetwood; when the genotype at the site of SNP8 is G, the CDL12_15522 gene is highly expressed; when the genotype at the site of SNP8 is T, the CDL12_15522 gene is lowly expressed.
10. The application according to claim 3 or 9, characterized in that, When obtaining the genotype of the yellow-flowered trumpet tree to be tested at any SNP marker site among SNP1, SNP4, SNP8, SNP18 and SNP24, PCR is performed using the genomic DNA of the yellow-flowered trumpet tree to be tested as the template DNA, and the PCR product is sequenced to obtain the genotype of the yellow-flowered trumpet tree to be tested at the site where the corresponding SNP marker is located.
Citation Information
Patent Citations
Urban landscape tree species tabebuia pruning and shaping method
CN115067102A
SNP (Single Nucleotide Polymorphism) molecular marker set, primer and application of SNP molecular marker set and primer in identifying local variety of non-heading Chinese cabbage'arrow shaft white '
CN115961072A
Method for increasing 2n pollen occurrence rate of tabebuia chrysantha
CN117530172A
Molecular means to extend desiccation tolerance during seed germination
WO2025061878A1