Cotton grain size gene identification method and device based on multiple omics strategies
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies for identifying cotton seed size genes suffer from problems such as low localization accuracy, numerous candidate genes, long verification cycles, insufficient explanation of molecular mechanisms, and weak applicability in breeding.
Employing a multi-omics strategy that combines high-throughput phenomics, genomics, and transcriptomics, we identified phenotypic data through image data, performed principal component analysis and genome-wide association analysis to identify quantitative trait loci, conducted transcriptional prediction to obtain hub genes, and performed co-localization to obtain key gene information.
It has achieved high-precision localization and rapid identification of cotton seed size genes, improving the accuracy and efficiency of gene identification and providing comprehensive and precise breeding technology support.
Smart Images

Figure CN121789776A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the integrated technical field of crop molecular breeding and bioinformatics, and in particular to a method and apparatus for identifying cotton seed size genes based on a multi-omics strategy. Background Technology
[0002] Cotton is an important global economic crop, and the size of its cotton seeds is a key agronomic trait affecting yield and quality. Cotton seed size is not only related to seed germination and seedling vigor, but also closely related to the growth and development of cotton fibers. For example, large-seed varieties often have longer, stronger fibers and higher fiber length uniformity.
[0003] In existing technologies, the identification of genes related to cotton seed size primarily relies on forward genetics. Forward genetics traces genotypes from phenotypes; however, its effectiveness is limited by several factors. Firstly, the accuracy of phenotypic scoring is difficult to guarantee. Due to the irregularity of seed traits, manual measurement of seed phenotypes is not only time-consuming and labor-intensive, requiring meticulous and sequential evaluation of each seed, but also highly subjective; different operators may obtain different results when measuring the same group of seeds, leading to significant errors in phenotypic data. Secondly, the density of genetic markers and the size of the mapping population also affect the effectiveness of forward genetics. In relatively small populations, the phenotypic traits of a limited sample may not fully represent genetic diversity, increasing the possibility of false positives or missed associations.
[0004] Single-omics strategies for gene identification often fall short of comprehensively elucidating the relationship between genes and traits, especially when dealing with quantitative traits like cotton seed size, which are regulated by multiple genes and interact in complex ways with environmental factors. Recent advancements in phenotyping intelligence, high-throughput omics, and computational methods have brought new hope to the systematic analysis of complex traits and molecular design breeding, theoretically offering the possibility for more comprehensive and in-depth research into biological genetic mechanisms. However, in practical applications, how to accurately identify genes using multi-omics approaches remains a pressing challenge. Summary of the Invention
[0005] In view of this, the purpose of this invention is to provide a method and apparatus for identifying cotton seed size genes based on a multi-omics strategy, which can overcome the problems of low accuracy in locating complex traits, numerous candidate genes and long verification cycles, insufficient explanation of molecular mechanisms, and weak breeding applicability of existing technologies.
[0006] In a first aspect, the present invention provides a method for identifying cotton seed size genes based on a multi-omics strategy, comprising: Based on the image data corresponding to cotton germplasm materials, identify multiple phenotypic data corresponding to cotton germplasm materials; Based on multiple phenotypic data, composite phenotypic data were obtained to characterize cotton seed size. Genome-wide association analysis was then performed based on the composite phenotypic data to identify quantitative trait loci affecting cotton seed size under linkage disequilibrium. The Hub gene that influences cotton seed size was obtained based on multiple phenotypic data. Co-localization of quantitative trait loci and Hub genes yielded key gene information affecting cotton seed size.
[0007] In one implementation, composite phenotypic data for characterizing cotton seed size is obtained based on multiple phenotypic data, including: Principal component analysis was performed on multiple phenotypic data to obtain the principal component PC used to characterize cotton seed size, and the principal component PC was used as composite phenotypic data.
[0008] In one implementation, genome-wide association analysis is performed based on composite phenotypic data to identify quantitative trait loci affecting cotton seed size under linkage disequilibrium, including: Obtain SNP data corresponding to cotton germplasm materials. SNP data is obtained by whole-genome resequencing and variant assessment of cotton germplasm materials. A correlation analysis was performed between SNP data and composite phenotype data to obtain the probability value of SNP data influencing composite phenotype data. Significant SNP data were screened from the SNP data based on probability values, and quantitative trait loci affecting cotton seed size were determined based on the pairwise linkage disequilibrium values of the significant SNP data.
[0009] In one implementation, determining the quantitative trait loci affecting cotton seed size based on pairwise linkage disequilibrium values of significant SNP data includes: Candidate genomic regions are delineated based on the number of significant SNPs and the spacing between adjacent SNPs. Determine the pairwise linkage disequilibrium values of significant SNP data contained within candidate genomic regions to identify linkage disequilibrium blocks based on the pairwise linkage disequilibrium values; For adjacent chain imbalance blocks, the following operation is performed: if the pairwise chain imbalance value of significant SNP data at the boundary between adjacent chain imbalance blocks is greater than a preset threshold, the adjacent chain imbalance blocks are merged to obtain a composite chain imbalance region. Gene information contained in complex linkage disequilibrium regions is used as quantitative trait loci affecting cotton seed size.
[0010] In one implementation, transcriptional prediction is performed on cotton germplasm based on multiple phenotypic data to obtain Hub genes that affect cotton seed size, including: Unsupervised clustering of cotton germplasm materials was performed based on multiple phenotypic data to divide the cotton germplasm materials into multiple clusters and obtain representative cotton germplasm materials selected from each cluster. Differentially expressed genes were obtained from transcriptome sequencing of representative cotton germplasm materials; Weighted gene co-expression network analysis based on differentially expressed genes was performed to obtain the Hub gene that affects cotton seed size.
[0011] In one implementation, weighted gene co-expression network analysis is performed based on differentially expressed genes to obtain Hub genes that affect cotton seed size, including: A weighted gene co-expression network was constructed based on differentially expressed genes. Based on the weighted gene co-expression network, co-expression modules that affect cotton seed size were identified, and Hub genes were screened from the co-expression modules based on the continuity of characteristic genes.
[0012] In one implementation, co-localization of quantitative trait loci and the Hub gene yields key gene information affecting cotton seed size, including: The intersection of quantitative trait loci and Hub genes contains genetic information, which is identified as the key genetic information affecting cotton seed size.
[0013] Secondly, the present invention also provides an apparatus for identifying cotton seed size genes based on a multi-omics strategy, comprising: The phenotypic recognition module is used to identify multiple phenotypic data corresponding to cotton germplasm materials based on the image data corresponding to the cotton germplasm materials. The quantitative trait locus determination module is used to obtain composite phenotypic data for characterizing cotton seed size based on multiple phenotypic data, and to perform genome-wide association analysis based on the composite phenotypic data to obtain the quantitative trait loci affecting cotton seed size under linkage disequilibrium. The Hub gene identification module is used to predict the transcription of cotton germplasm based on multiple phenotypic data in order to obtain Hub genes that affect cotton seed size. The co-localization module is used to co-localize quantitative trait loci and Hub genes to obtain key gene information affecting cotton seed size.
[0014] Thirdly, the present invention also provides an electronic device including a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement any of the methods provided in the first aspect.
[0015] Fourthly, the present invention also provides a computer-readable storage medium storing computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement any of the methods provided in the first aspect.
[0016] This invention provides a method and apparatus for identifying cotton seed size genes based on a multi-omics strategy. First, based on image data corresponding to cotton germplasm materials, multiple phenotypic data corresponding to the cotton germplasm materials are identified. Then, based on the multiple phenotypic data, composite phenotypic data for characterizing cotton seed size is obtained, and genome-wide association analysis is performed on the composite phenotypic data to obtain quantitative trait loci affecting cotton seed size under linkage disequilibrium. Furthermore, based on the multiple phenotypic data, transcriptional prediction is performed on the cotton germplasm materials to obtain hub genes affecting cotton seed size. Finally, co-localization of quantitative trait loci and hub genes is performed to obtain key gene information affecting cotton seed size. The above method proposes a comprehensive multi-omics integration strategy that combines high-throughput phenomics, gene combinatorial analysis, and transcriptomics, overcoming the problems of low accuracy in complex trait localization, numerous candidate genes and long validation cycles, insufficient explanation of molecular mechanisms, and weak breeding applicability of existing technologies.
[0017] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0018] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0019] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0020] Figure 1 A flowchart illustrating a method for identifying cotton seed size genes based on a multi-omics strategy, provided in an embodiment of the present invention; Figure 2 A technical framework diagram of a method for identifying cotton seed size genes based on a multi-omics strategy provided in an embodiment of the present invention; Figure 3This is a schematic diagram illustrating PCA and genome-wide association analysis of phenotypic data provided in an embodiment of the present invention; Figure 4 A K-means clustering diagram of cotton seed size is provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of a WGCNA-related hub gene for seed size identification, provided as an embodiment of the present invention. Figure 6 A schematic diagram illustrating the positive regulation of seed size and weight using GhCSS5, provided in an embodiment of the present invention; Figure 7 A schematic diagram of the structure of a device for identifying cotton seed size genes based on a multi-omics strategy, provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Currently, how to accurately identify genes using multi-omics approaches remains a pressing problem in practical applications. Based on this, the present invention provides a method and device for identifying cotton seed size genes based on a multi-omics strategy, which can overcome the problems of low accuracy in locating complex traits, numerous candidate genes and long verification cycles, insufficient explanation of molecular mechanisms, and weak breeding applicability of existing technologies.
[0023] To facilitate understanding of this embodiment, a method for identifying cotton seed size genes based on a multi-omics strategy, as disclosed in this embodiment of the invention, will first be described in detail. (See [link to relevant documentation]). Figure 1 The diagram shows a flowchart of a method for identifying cotton seed size genes based on a multi-omics strategy. This method mainly includes the following steps S102 to S108: Step S102: Identify multiple phenotypic data corresponding to cotton germplasm materials based on the image data corresponding to the cotton germplasm materials.
[0024] In one example, a high-throughput intelligent seed phenotyping platform can be used to acquire and identify images of cotton seed size in cotton germplasm materials, and extract one or more phenotypic data from at least seed length, seed width, seed area, seed roundness, seed perimeter, length-to-width ratio, and seed finger.
[0025] Step S104: Based on multiple phenotypic data, composite phenotypic data for characterizing cotton seed size is obtained, and genome-wide association analysis is performed based on the composite phenotypic data to obtain quantitative trait loci affecting cotton seed size under linkage disequilibrium.
[0026] In one implementation, principal component analysis (PCA) is performed on the acquired multiple phenotypic data to obtain principal components (PCs) that can characterize cotton seed size. The principal components (PCs) are then used as composite phenotypic data. The PC scores of the composite phenotypic data are used to replace the original phenotypic data for genome-wide association study (GWAS). Quantitative trait loci (QTLs) are determined according to the linkage disequilibrium (LD) range.
[0027] Step S106: Based on multiple phenotypic data, transcriptional prediction is performed on cotton germplasm materials to obtain Hub genes that affect cotton seed size.
[0028] In one implementation, unsupervised K-means clustering is performed on multiple phenotypic data to identify several cotton germplasm materials into clusters. Representative cotton germplasm materials are selected from each cluster. For each cotton germplasm material, samples are taken during ovule development and transcriptome sequencing is performed. A weighted gene co-expression network (WGCNA) is constructed based on differentially expressed genes to obtain Hub genes that are significantly associated with cotton seed size.
[0029] Step S108 involves co-localization of quantitative trait loci and Hub genes to obtain key gene information affecting cotton seed size.
[0030] In one instance, the genetic information contained in the intersection of quantitative trait loci and Hub genes can be identified as the key genetic information affecting cotton seed size.
[0031] The method provided in this invention proposes a comprehensive multi-omics integration strategy that combines high-throughput phenomics, gene combinatorial analysis, and transcriptomics. This strategy overcomes the problems of low accuracy in locating complex traits, numerous candidate genes and long validation cycles, insufficient explanation of molecular mechanisms, and weak applicability in breeding in existing technologies.
[0032] To facilitate understanding, this invention provides a specific implementation of a method for identifying cotton seed size genes based on a multi-omics strategy. This method begins with high-throughput seed phenotypic acquisition, including seven phenotypic data: seed length, width, area, roundness, perimeter, aspect ratio, and seed index. Principal component analysis (PC) is used to reduce dimensionality and obtain principal components (PCs) that represent variations in cotton seed size. Genome-wide association analysis (GWAS) is then performed using these PCs as phenotypes to obtain quantitative trait loci (QTLs) related to cotton seed size under linkage disequilibrium (LD) constraints. The phenotypic data are further clustered to select representative materials. Ovule samples from key developmental stages are collected for transcriptome sequencing. Weighted co-expression network analysis (WGCNA) is used to identify Hub genes highly correlated with cotton seed size. Key gene information is obtained by co-localizing the QTLs obtained from GWAS with the Hub genes obtained from WGCNA, and functional verification is performed using transgenic materials (overexpression and RNA interference). The embodiments of this invention achieve high-precision localization of complex quantitative traits and rapid identification of key genes, which is applicable to cotton seed size improvement, variety selection and molecular design breeding scenarios, and has the advantages of accurate localization and efficient verification.
[0033] The specific implementation process is as follows: (a) Phenotypic Measurement: After the cotton matures, the cotton bolls are harvested manually and then ginned and delinted to obtain smooth cotton seeds. ≥ 30 healthy, plump and uniformly sized cotton seeds are selected from each germplasm for image acquisition and identification. The phenotype includes at least seven indicators: seed length, seed width, seed area, seed roundness, seed circumference, length-to-width ratio, and seed index.
[0034] For example, 413 upland cotton (Gossypium hirsutum L.) germplasms with rich genetic diversity were selected and planted in designated areas. After the cotton matured, the bolls were harvested by hand, and the smooth cotton seeds were obtained through ginning and delinting. From each germplasm, 30 healthy, plump, and uniformly sized cotton seeds were selected, and images were acquired and identified using an intelligent seed phenotypic analysis platform. Morphological data such as seed length, seed width, seed area, seed perimeter, length-to-width ratio, roundness, and seed finger were accurately measured and recorded.
[0035] (ii) PCA dimensionality reduction: Principal component analysis is performed on multiple phenotypic data to obtain the principal component PC used to characterize cotton seed size, and the principal component PC is used as composite phenotypic data.
[0036] Using the "FactoMineR" package in R, principal component analysis (PCA) was performed on the collected phenotypic data based on the correlation matrix of seven phenotypic data using the PCA function (with scale.unit=TRUE). The "factoextra" R package was used to visualize the PCA results. The correlation between eigenvalues and variables was extracted using standard functions. The prerequisite for using principal components (PCA) as composite phenotypic data to replace the original phenotypic data in genome-wide association analysis is that the cumulative phenotypic variance explained by the first two selected principal components is ≥80%, for example, 84.47%. In the correlation matrix, the rows and columns represent cotton germplasm materials and phenotypic terms, respectively. Each element represents the specific value of the cotton germplasm material corresponding to its row relative to the phenotypic term corresponding to its column; this value is the phenotypic data.
[0037] (III) Genome-wide association analysis: (3.1) Obtain SNP data corresponding to cotton germplasm materials. SNP data is obtained by whole-genome resequencing and variation assessment of cotton germplasm materials. For example, 3,665,030 SNP data were obtained from whole-genome resequencing and variation assessment of 419 upland cotton germplasm accessions. From these, SNP data corresponding to 413 cotton germplasm materials actually used in the embodiments of this invention were extracted. Furthermore, the SNP data can be screened. For example, using PLINK software, SNP data that meet preset conditions (e.g., minor allele frequency > 0.05, deletion rate < 10%) are retained, and SNP data that do not meet the preset conditions are removed, resulting in 1,941,408 high-quality SNPs.
[0038] (3.2) Perform correlation analysis on SNP data and composite phenotype data to obtain the probability value of SNP data affecting composite phenotype data. In one example, EMMAX software was used to conduct a correlation study on SNP data and composite phenotype data to obtain the probability value P of SNP data affecting composite phenotype data, and the corresponding Manhattan plot was generated using the "CMPlot" R package.
[0039] (3.3) Significant SNP data are screened from the SNP data based on probability values, and quantitative trait loci affecting cotton seed size are determined based on the pairwise linkage disequilibrium values of the significant SNP data. In one example, SNP data with probability values P higher than a preset threshold can be considered as significant SNP data. The process of determining quantitative trait loci specifically includes: (3.31) Candidate genomic regions are delineated based on the number of significant SNPs and the spacing between adjacent SNPs. In one example, QTLs are first defined by requiring at least three significant SNPs (-Log10(p)≥5) and that the spacing between adjacent SNPs does not exceed 200 kb.
[0040] (3.32) Determine the pairwise linkage disequilibrium values of significant SNP data contained within candidate genomic regions to identify linkage disequilibrium blocks based on the pairwise linkage disequilibrium values. In one example, to further define QTL boundaries, embodiments of the present invention introduce the concept of association peaks, which represent clusters of closely linked significant SNPs, and use the LDBlockShow tool to calculate the pairwise linkage disequilibrium (LD) values (R²) of all significant SNP data, and then identify linkage disequilibrium (LD) blocks according to the default criteria of the LDBlockShow tool.
[0041] (3.33) For adjacent chain imbalance blocks, the following operation is performed: if the pairwise chain imbalance value of significant SNP data at the boundary between adjacent chain imbalance blocks is greater than a preset threshold, the adjacent chain imbalance blocks are merged to obtain a composite chain imbalance region. For example, if the significant SNP data at the boundary between adjacent chain imbalance (LD) blocks exhibits moderate LD, that is, its significant chain imbalance (LD) value R²≥0.5, then the adjacent chain imbalance (LD) blocks are merged into a larger composite chain imbalance (LD) region.
[0042] (3.34) Gene information contained in the complex linkage disequilibrium region is used as a quantitative trait locus affecting cotton seed size.
[0043] (iv) K-means clustering: Unsupervised clustering of cotton germplasm materials based on multi-phenotypic data is performed to divide the cotton germplasm materials into multiple clusters, and representative cotton germplasm materials are selected from each cluster. In one example, the K-means clustering algorithm is applied for unsupervised clustering analysis. The sum of squares of different k values is compared, and the optimal number of clusters with no significant change in slope is selected to divide the cotton germplasm materials into multiple clusters. Two representative cotton germplasm materials are selected from each cluster.
[0044] (V) Transcriptome and Weighted Gene Co-expression Network: Differentially expressed genes were obtained from transcriptome sequencing of representative cotton germplasm materials. Weighted gene co-expression network analysis was then performed based on these differentially expressed genes to identify hub genes affecting cotton seed size. The process of identifying hub genes includes: constructing a weighted gene co-expression network based on the differentially expressed genes; identifying co-expression modules affecting cotton seed size based on this network; and screening hub genes from these co-expression modules based on the continuity of characteristic genes. Specifically: Six representative germplasms selected from different cotton seed size groups after K-means clustering analysis were sampled at 3, 10, 20, 30, and 45 days of ovule development and immediately stored at -80°C. Ovule tissues were ground in liquid nitrogen, and total RNA was extracted using an RNA (Ribonucleic Acid) purification kit. RNA-seq libraries were constructed, and sequencing was performed using paired-end 150 bp (PE150) reads. Clean reads were aligned to a reference genome using HISAT2 to identify differentially expressed genes (DEGs).
[0045] Using the FPKM values of differentially expressed genes (padj < 0.05, |log2FC| > 2, FPKM > 1) as input, a weighted gene co-expression network was constructed using the R package "WGCNA". The soft threshold parameter was set to 12, the minimum module size parameter to 30, and the CutHeight parameter to 0.2. Quantitative trait loci must be included in the DEG. For the seven phenotypic associated genes of cotton seed size, the genes in the module with the highest correlation were selected for co-expression network construction. Genes with high connectivity in the co-expression network were designated as Hub genes.
[0046] (vi) Co-location of key genes: Gene information contained in the intersection of quantitative trait loci and Hub genes is identified as key gene information affecting cotton seed size.
[0047] (vii) Functional verification: Construct transgenic materials with overexpression and RNA interference of the core candidate genes, or obtain edited materials by gene editing, and measure the phenotypes related to cotton seed size to verify their regulatory role.
[0048] The coding sequences of candidate genes, such as GhCSS5, were obtained from the CottonGen database. Gene-specific primers were designed to amplify the gene fragment from the ZM24 cDNA library. The amplified fragment was cloned into the pCAMBIA-2300 binary vector and overexpressed (GhCSS5-OE) under the control of the CaMV35S promoter. Simultaneously, a partial fragment of GhCSS5 was amplified and inserted into the pBI121 vector to construct an RNA interference (GhCSS5-RNAi) plasmid. ZM24 cotton seeds were delinted, sterilized, and germinated sequentially. Germinated seeds were then transferred to MS medium and cultured for 7 days. Under aseptic conditions, hypocotyl fragments were excised and inoculated with Agrobacterium tumefaciens (OD600=0.8) for 1 minute, followed by tissue culture. Regenerated plants were screened using target gene-specific primers to exclude non-transgenic lines. Confirmed transgenic lines were propagated to the T3 generation for phenotypic analysis.
[0049] In summary, the method provided by the embodiments of the present invention has at least the following characteristics: First, it enables highly efficient and accurate data acquisition: Through a self-developed intelligent seed phenotypic analysis platform, high-throughput and automated seed phenotypic parameter acquisition is achieved, significantly improving the efficiency and accuracy of data acquisition and laying a reliable foundation for gene identification. Second, it enables integrated analysis of multi-omics data: Principal component analysis is used to reduce the dimensionality of phenotypic data, and GWAS analysis is conducted in conjunction with genomic data to effectively identify more potential genetic loci, improving the accuracy and resolution of gene localization. Third, it accurately screens key candidate genes: A weighted gene co-expression network is constructed through transcriptomics analysis and co-localized with GWAS results, systematically identifying key genes regulating cotton seed size and providing an efficient and reliable means for new gene identification. Fourth, it provides comprehensive technical support: This method provides a comprehensive, accurate, and efficient technical system for cotton genetic improvement and molecular breeding, possessing significant application value and promising prospects for wider application.
[0050] For ease of understanding, this invention provides an application example of a method for identifying cotton seed size genes based on a multi-omics strategy. See [link to relevant documentation]. Figure 2 The diagram shows a technical framework for a method for identifying cotton seed size genes based on a multi-omics strategy, including: A, phenotypic measurement; B, PCA dimensionality reduction and genome-wide association analysis (PCA-based WAS); C, K-means clustering and weighted gene co-expression network; and D, function verification.
[0051] For example: (1) Seven seed-size related traits of 413 materials were measured, including seed length, seed width, seed area, seed roundness, seed circumference, length-to-width ratio, and seed finger. See [link to relevant documentation] Figure 3 The diagram illustrates PCA and genome-wide association analysis of phenotypic data. PCA analysis was performed on all seven traits, and PC1 explained 70.2% of the phenotypic variation. Figure 3 (A) Except for aspect ratio and seed finger, the other five traits all showed high positive loading on PC1 (0.87–0.99), while aspect ratio showed a negative loading (-0.23). Figure 3 (C in the text). Seeds with high PC1 values exhibit greater length, width, area, roundness, and circumference, but a smaller length-to-width ratio, and vice versa. PC2 explains 14% of the phenotypic variation (C in the text). Figure 3 B), the grain index has a high loading (0.97) on PC2 (in the B category). Figure 3 The C in the figure indicates that PC2 mainly represents the seed finger.
[0052] (2) A linear mixture model was used to perform GWAS analysis on PC1 and PC2 to correct for population structure and kinship bias. Figure 3 The results identified 187 and 15 SNPs significantly associated with PC1 and PC2, respectively. Linkage disequilibrium (LD) was used to define the boundaries of quantitative trait loci (QTLs). Eleven QTLs associated with cotton seed size (named qCSS1 to qCSS11) were identified from the GWAS signal of PC1, while two QTLs associated with grain index (named qSI1 and qSI2) were identified from the GWAS signal of PC2. Among these, the qCSS5 locus showed the strongest association signal in the GWAS of PC1 (…). Figure 3 The D in the middle). Its LD block reveals a complex genomic region composed of multiple closely linked small LD blocks (D). Figure 3 The F in the F group contains 68 candidate genes.
[0053] (3) K-means cluster analysis was performed on the seven phenotypic traits. See [link / reference] Figure 4 The image shows a K-means clustering diagram of cotton seed size, and the optimal number of clusters is determined to be 3. Figure 4 The AB in the data indicates that there is a clear differentiation among the three clusters. Figure 4 (C in the text). Clusters I, II, and III contained 119, 147, and 147 cotton materials, respectively. Evaluation of seed length, seed width, and seed index showed that cluster II exhibited the highest values for all three phenotypic traits, while cluster I had the lowest values, and cluster III was at an intermediate level. Figure 4 (DF in the text).
[0054] Two representative materials were selected from each cluster, and RNA sequencing was performed on the ovules of the six representative materials at six key developmental stages (3, 10, 15, 20, 25, and 35 DPA). The average alignment rate of clean reads to the TM-1 reference genome was 96.7%.
[0055] (4) To accurately identify the Hub gene set regulating seed size, the following criteria were used to screen DEGs: corrected p-value (padj) < 0.05, |log2FC| > 2, and FPKM > 1. See [link to relevant documentation] Figure 5 The diagram shown illustrates a WGCNA-based identification of hub genes related to seed size. A total of 16,417 genes were screened for weighted co-expression network analysis (WGCNA), and 27 co-expression modules were identified based on similar expression patterns. Figure 5 (B in the original text). Module-trait correlation results showed that a certain module (135 genes) exhibited the strongest positive correlation (0.73) with the seed size trait. <r<0.81)( Figure 5In the C section, a co-expression network was constructed for the genes within the module, and 17 hub genes with high connectivity were identified. Figure 5 (E in the text).
[0056] (5) GWAS identified 68 QTL candidate genes, and WGCNA identified 17 hub genes. See also Figure 6 The diagram illustrates the positive regulation of seed size and weight by GhCSS5. Co-location analysis of these two gene sets revealed an overlapping gene, Gh_A10G245100. Figure 6 (A in the text). This gene showed significant dominant expression in ovules of large-grained materials at different developmental stages. Figure 5 (F in the text). Functional annotation indicates that Gh_A10G245100 encodes the cytochrome P450 monooxygenase CYP90A1. This gene is a potential regulator of cotton seed size and has been named GhCSS5 for further functional validation.
[0057] (6) Stable GhCSS5 overexpression (GhCSS5-OE) and GhCSS5-RNA interference (GhCSS5-RNAi) lines were constructed. Quantitative RT-PCR analysis confirmed that, compared with the wild type (WT: ZM24), the transcription level of this gene was significantly increased in the OE line, while the transcription level was decreased in the RNAi line. Figure 6 (B in the text). All GhCSS5-RNAi lines showed smaller grains ( Figure 6 CD in the middle), including grain length ( Figure 6 E), kernel width ( Figure 4 F in the middle), grain area ( Figure 6 G in the middle), grain circumference ( Figure 6 H) and grain index ( Figure 6 The decrease in K (in the WT) resulted in a reduction in the grain index by 36.83%, 35.58%, and 27.22% compared to the WT. Figure 6 (K in the text). Conversely, all GhCSS5-OE lines exhibited the opposite seed size pattern. Compared to WT, the grain index increased by 27.85%, 29.6%, and 31.9%, respectively (K in the text). Figure 6 The K in the text is notable. It is noteworthy that, compared to WT, the seed roundness and aspect ratio of the GhCSS5-RNAi or GhCSS5-OE lines did not change significantly. Figure 6 The results (IJ) indicate that GhCSS5 regulates seed size without affecting seed shape. These findings suggest that GhCSS5 is a positive regulator of cotton seed size.
[0058] This invention provides a method for identifying cotton seed size genes based on a multi-omics strategy. Principal component analysis (PC) is used to reduce dimensionality and obtain principal components (PCs) that represent seed size variations. These PCs are then used as phenotypes for genome-wide association analysis (GWAS). Hub genes obtained from a weighted co-expression network (WGCNA) are then co-located with GWAS candidate genes to obtain candidate genes. By simply using the multi-omics gene identification model provided in this approach, high-precision localization of complex quantitative traits and rapid identification of key genes can be achieved.
[0059] Based on the foregoing embodiments, this invention provides a device for identifying cotton seed size genes based on a multi-omics strategy, see [link to related document]. Figure 7 The diagram shows a device for identifying cotton seed size genes based on a multi-omics strategy. The device mainly includes the following parts: Phenotypic recognition module 702 is used to identify multiple phenotypic data corresponding to cotton germplasm materials based on image data corresponding to cotton germplasm materials; The quantitative trait locus determination module 704 is used to obtain composite phenotypic data for characterizing cotton seed size based on multiple phenotypic data, and to perform genome-wide association analysis based on the composite phenotypic data to obtain the quantitative trait loci affecting cotton seed size under linkage disequilibrium. Hub gene identification module 706 is used to predict the transcription of cotton germplasm based on multiple phenotypic data in order to obtain Hub genes that affect cotton seed size. Co-localization module 708 is used to co-localize quantitative trait loci and Hub genes to obtain key gene information affecting cotton seed size.
[0060] The device provided in this invention proposes a comprehensive multi-omics integration strategy that combines high-throughput phenomics, gene combinatorial analysis, and transcriptomics. This strategy overcomes the problems of low accuracy in locating complex traits, numerous candidate genes and long verification cycles, insufficient explanation of molecular mechanisms, and weak applicability in breeding in existing technologies.
[0061] In one embodiment, the quantitative trait locus determination module 704 is specifically used for: Principal component analysis was performed on multiple phenotypic data to obtain the principal component PC used to characterize cotton seed size, and the principal component PC was used as composite phenotypic data.
[0062] In one embodiment, the quantitative trait locus determination module 704 is specifically used for: Obtain SNP data corresponding to cotton germplasm materials. SNP data is obtained by whole-genome resequencing and variant assessment of cotton germplasm materials. A correlation analysis was performed between SNP data and composite phenotype data to obtain the probability value of SNP data influencing composite phenotype data. Significant SNP data were screened from the SNP data based on probability values, and quantitative trait loci affecting cotton seed size were determined based on the pairwise linkage disequilibrium values of the significant SNP data.
[0063] In one embodiment, the quantitative trait locus determination module 704 is specifically used for: Candidate genomic regions are delineated based on the number of significant SNPs and the spacing between adjacent SNPs. Determine the pairwise linkage disequilibrium values of significant SNP data contained within candidate genomic regions to identify linkage disequilibrium blocks based on the pairwise linkage disequilibrium values; For adjacent chain imbalance blocks, the following operation is performed: if the pairwise chain imbalance value of significant SNP data at the boundary between adjacent chain imbalance blocks is greater than a preset threshold, the adjacent chain imbalance blocks are merged to obtain a composite chain imbalance region. Gene information contained in complex linkage disequilibrium regions is used as quantitative trait loci affecting cotton seed size.
[0064] In one implementation, the Hub gene determination module 706 is specifically used for: Unsupervised clustering of cotton germplasm materials was performed based on multiple phenotypic data to divide the cotton germplasm materials into multiple clusters and obtain representative cotton germplasm materials selected from each cluster. Differentially expressed genes were obtained from transcriptome sequencing of representative cotton germplasm materials; Weighted gene co-expression network analysis based on differentially expressed genes was performed to obtain the Hub gene that affects cotton seed size.
[0065] In one implementation, the Hub gene determination module 706 is specifically used for: A weighted gene co-expression network was constructed based on differentially expressed genes. Based on the weighted gene co-expression network, co-expression modules that affect cotton seed size were identified, and Hub genes were screened from the co-expression modules based on the continuity of characteristic genes.
[0066] In one implementation, the co-positioning module 708 is specifically used for: The intersection of quantitative trait loci and Hub genes contains genetic information, which is identified as the key genetic information affecting cotton seed size.
[0067] The device provided in this embodiment of the invention has the same implementation principle and technical effect as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment.
[0068] This invention provides an electronic device, specifically, the electronic device includes a processor and a memory; the memory stores a computer program, which, when run by the processor, executes the method described in any of the above embodiments.
[0069] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 100 includes: a processor 80, a memory 81, a bus 82, and a communication interface 83. The processor 80, the communication interface 83, and the memory 81 are connected through the bus 82. The processor 80 is used to execute executable modules, such as computer programs, stored in the memory 81.
[0070] The memory 81 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 83 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.
[0071] Bus 82 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 8 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0072] The memory 81 is used to store programs. After receiving an execution instruction, the processor 80 executes the program. The method executed by the device for defining the flow process disclosed in any of the foregoing embodiments of the present invention can be applied to the processor 80 or implemented by the processor 80.
[0073] The processor 80 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 80 or by instructions in software form. The processor 80 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 81. The processor 80 reads the information in memory 81 and, in conjunction with its hardware, completes the steps of the above method.
[0074] The computer program product of the readable storage medium provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For specific implementation, please refer to the foregoing method embodiments, which will not be repeated here.
[0075] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0076] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for identifying cotton seed size genes based on a multi-omics strategy, characterized in that, include: Based on the image data corresponding to the cotton germplasm materials, identify multiple phenotypic data corresponding to the cotton germplasm materials; Based on multiple phenotypic data, composite phenotypic data for characterizing cotton seed size is obtained, and genome-wide association analysis is performed based on the composite phenotypic data to obtain quantitative trait loci affecting cotton seed size under linkage disequilibrium. Based on multiple phenotypic data, transcriptional prediction was performed on the cotton germplasm to obtain the Hub gene that affects the size of the cotton seed. Co-localization of the quantitative trait loci and the Hub gene yielded key gene information affecting cotton seed size.
2. The method for identifying cotton seed size genes based on a multi-omics strategy according to claim 1, characterized in that, Based on multiple phenotypic data, composite phenotypic data for characterizing cotton seed size is obtained, including: Principal component analysis was performed on multiple phenotypic data to obtain the principal component PC used to characterize cotton seed size, and the principal component PC was used as composite phenotypic data.
3. The method for identifying cotton seed size genes based on a multi-omics strategy according to claim 1, characterized in that, Genome-wide association analysis was performed based on the composite phenotypic data to identify quantitative trait loci affecting cotton seed size under linkage disequilibrium, including: Obtain the SNP data corresponding to the cotton germplasm material, wherein the SNP data is obtained by whole-genome resequencing and variant assessment of the cotton germplasm material; A correlation analysis is performed on the SNP data and the composite phenotype data to obtain the probability value of the SNP data affecting the composite phenotype data; Significant SNP data are screened from the SNP data based on the probability values, and quantitative trait loci affecting cotton seed size are determined based on the pairwise linkage disequilibrium values of the significant SNP data.
4. The method for identifying cotton seed size genes based on a multi-omics strategy according to claim 3, characterized in that, Based on the pairwise linkage disequilibrium values of the significant SNP data, the quantitative trait loci affecting cotton seed size are determined, including: Candidate genomic regions are defined based on the number of significant SNP data and the spacing between adjacent SNP data. Determine the pairwise linkage disequilibrium values of the significant SNP data contained within the candidate genomic region, in order to identify linkage disequilibrium blocks based on the pairwise linkage disequilibrium values; For adjacent chain imbalance blocks, the following operation is performed: if the pairwise chain imbalance value of the significant SNP data at the boundary between adjacent chain imbalance blocks is greater than a preset threshold, the adjacent chain imbalance blocks are merged to obtain a composite chain imbalance region. The genetic information contained in the complex linkage disequilibrium region is used as the quantitative trait loci affecting the size of cotton seeds.
5. The method for identifying cotton seed size genes based on a multi-omics strategy according to claim 1, characterized in that, Based on multiple phenotypic data, transcriptional prediction was performed on the cotton germplasm to obtain Hub genes that affect the size of the cotton seeds, including: Unsupervised clustering is performed on the cotton germplasm materials based on multiple phenotypic data to divide the cotton germplasm materials into multiple clusters and obtain representative cotton germplasm materials selected from each cluster. Differentially expressed genes were obtained by transcriptome sequencing of the representative cotton germplasm materials. Weighted gene co-expression network analysis was performed based on the differentially expressed genes to obtain the Hub gene that affects the size of the cotton seeds.
6. The method for identifying cotton seed size genes based on a multi-omics strategy according to claim 5, characterized in that, Weighted gene co-expression network analysis was performed based on the differentially expressed genes to obtain the Hub genes that affect the size of the cotton seeds, including: A weighted gene co-expression network is constructed based on the differentially expressed genes. Based on the weighted gene co-expression network, co-expression modules that affect the size of cotton seeds are determined, and Hub genes are screened from the co-expression modules according to the continuity of characteristic genes.
7. The method for identifying cotton seed size genes based on a multi-omics strategy according to claim 1, characterized in that, Co-location of the quantitative trait loci and the Hub gene yielded key gene information affecting cotton seed size, including: The genetic information contained in the intersection of the quantitative trait loci and the Hub gene is identified as the key genetic information affecting the size of the cotton seed.
8. A device for identifying cotton seed size genes based on a multi-omics strategy, characterized in that, include: The phenotypic recognition module is used to identify multiple phenotypic data corresponding to the cotton germplasm based on the image data corresponding to the cotton germplasm. The quantitative trait locus determination module is used to obtain composite phenotypic data for characterizing cotton seed size based on multiple phenotypic data, and to perform genome-wide association analysis based on the composite phenotypic data to obtain the quantitative trait loci affecting cotton seed size under linkage disequilibrium. The Hub gene identification module is used to perform transcriptional prediction on the cotton germplasm based on multiple phenotypic data to obtain the Hub gene that affects the size of the cotton seed. The co-localization module is used to co-localize the quantitative trait loci and the Hub gene to obtain key gene information affecting the size of cotton seeds.
9. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the method of any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions that, when invoked and executed by a processor, cause the processor to perform the method according to any one of claims 1 to 7.