Compositions and methods for enhancing plant yield and resistance to drought

By identifying and incorporating specific SNPs in the DREB2A-A and HSFA6B-D binding sites of IPS1-A, cotton plants exhibit enhanced drought resistance and heat tolerance, addressing yield reduction in semi-arid conditions.

WO2025255420A1PCT designated stage Publication Date: 2025-12-11BOYCE THOMPSON INSTITUTE FOR PLANT RESEARCH INC +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/032570
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-05
Filing Date
2025-06-05
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Cotton crops in semi-arid regions face significant challenges due to drought and heat stress, leading to reduced productivity and yield, with limited understanding of the genetic and molecular mechanisms underlying drought tolerance and fiber yield in allopolyploid species like upland cotton.

Method used

Identification and utilization of specific single nucleotide polymorphisms (SNPs) in the DREB2A-A and HSFA6B-D binding sites of IPS1-A, particularly the C allele at nucleotide 19, which enhance drought resistance and heat tolerance, along with methods for detecting these SNPs using techniques such as hybridization and sequencing, and incorporating these SNPs into cotton plants to create recombinant varieties.

Benefits of technology

The presence of the C allele at nucleotide 19 in the DREB2A-A or HSFA6B-D binding site of IPS1-A confers enhanced drought resistance and heat tolerance, resulting in improved lint yield and resistance to adverse environmental conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000017_0001
    Figure IMGF000017_0001
  • Figure IMGF000051_0001
    Figure IMGF000051_0001
  • Figure IMGF000054_0001
    Figure IMGF000054_0001
Patent Text Reader

Abstract

Compositions and methods for creation of plants having improved agricultural traits, particularly drought resistant cotton plants, are disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]Compositions and Methods for Enhancing Plant Yield and Resistance to Drought Grant Support Statement This invention was made with government support under Grant No.2102120 awarded by the National Science Foundation (NSF). The government has certain rights in the invention. Cross-Reference to Related Applications This application claims the benefit of the filing date of U.S. Provisional Application No. 63 / 656,209, filed June 5, 2024, the entire contends of which is incorporated by reference herein. Incorporation-by-Reference of Material Submitted in Electronic Form The contents of the electronic sequence listing (BTCO-101-PCT.xml; Size: 64,232 bytes; and Date of Creation: June 5, 2025) is herein incorporated by reference in its entirety. Field of the Invention The present invention relates to the generation of plants having desirable characteristics. More specifically, the invention provides compositions and method for creation of plants, e.g., plants, such as cotton plants, having a SNP in the HSFA6B binding site of IPS1-A exhibiting enhanced plant yield and resistance to drought. Background of the Invention Several publications and patent documents are cited throughout the specification in order to describe the state of the art to which this invention pertains. Each of these citations is incorporated by reference herein as though set forth in full. Upland cotton (Gossypium hirsutum L.) is the world’s top renewable textile fiber that supports a multibillion-dollar industry with a global production of 120.2 million bales of cotton (~26 million metric tons, USDA, 2022). In particular, it is an economically important crop for the United States and for Arizona, where upland cotton is planted on ∼50,000 hectares (USDA 2021). This land mainly consists of the semi-arid environment of the low desert. Due to limited precipitation in this land, the plants are watered using surface irrigation to complement the precipitation. Cotton productivity in semi-arid areas of the Southwestern United States is severely threatened by global climate change. Increasing climatic variability is responsible for hotter summers, which causes day and night temperatures far above the thermal optimum (30 / 22◦C) for the crop, and lower, erratic rainfall patterns which expose the crop to an increasing risk of drought (Alizadeh et al., 2020; Burke and Wanjura, 2010; Zhang et al., 2021). In addition to being a critical fiber crop, cotton serves as an excellent model polyploid system for studying the impacts of interspecific hybridization on various agronomic traits. The two cotton species predominantly cultivated for fiber production, G. hirsutum (upland cotton) and Gossypium barbadense (Pima cotton), are New World allotetraploids that are believed to have formed ~ 1-2 million years ago from a transoceanic hybridization of two progenitor diploids. These two progenitors are an A genome diploid originating from Africa or Asia (e.g., Gossypium arboreum, tree cotton) and a D genome diploid from Central or South America (e.g., Gossypium raimondii). This unique combination of homologous gene pairs resulted in superior fiber yield and quality over diploid progenitors that has since undergone additional selection in both G. hirsutum and G. barbadense. Despite high levels of sequence conservation and collinearity between these two species and their diploid progenitors, both an altered epigenetic landscape and homeolog expression divergence, have contributed to G. hirsutum’s capacity to maintain lint yield under a wide range of environments. Like other crops, drought tolerance in cotton involves complex signaling pathways and transcriptional networks orchestrated by a number of transcription factors (TFs) and signaling proteins (Mahmood et al., 2019; Takahashi et al., 2020; Shinozaki and Yamaguchi-Shinozaki, 2007). These regulatory proteins are typically either upstream of, or transcriptionally interconnected with, the genes necessary for coping with stress and maintaining crucial metabolic pathways. Examples of typical regulatory targets include genes encoding for enzymes related to the production of protective metabolites, transporters, chaperones, and lipid biosynthesis proteins (Malhotra and Sowdhamini, 2014; Singh and Laxmi, 2015; Gupta et al., 2020). In cotton, multiple transcription factor families, such as GhNACs, GhDREBs, GhERFs, and GhWRKYs, have been associated with the drought stress response (Huang et al., 2009; Chu et al., 2015; Huang et al., 2013; Ma et al., 2017). Additionally, Genome Wide Association Studies (GWAS)-based approaches have identified numerous candidate genes related to lint yield, fiber quality, and other agronomically important traits (Fang et al., 2014; Sun et al., 2021; Ma et al., 2018). Indeed, while germplasm exists with the ability to maintain fiber growth and quality under heat and drought conditions, little is known about how these traits arose in the cotton genome or the regulatory factors that connect them. Clearly, a better understanding of the physio-genetic mechanisms that regulate cotton’s response to arid conditions is urgently needed. Summary of the Invention In accordance with the present invention, methods for identifying a drought resistant plant are provided. In certain embodiments, the methods comprise, a) obtaining a nucleic acid sample from said plant; and b) determining whether said sample contains at least one SNP containing nucleic acid in a DREB2A-A or HSFA6B-D binding site of IPS1-A, the presence of said SNP containing nucleic acid conferring at least one desirable plant characteristic selected from increased drought resistance, heat tolerance and lint yield in said plant. In certain embodiments, the SNP is present in nucleotide 19 of SEQ ID NO: 33. In certain embodiments, the SNP is a C allele, G allele, or an A allele. In certain embodiments, the C allele is present and the presence of the C allele is associated with enhanced drought resistance when compared to plants lacking a C allele. In another embodiment, the methods comprise, a) obtaining a nucleic acid sample from said plant; and b) determining whether said sample contains at least one SNP provided in Figure 8 or Figure 13, the presence of said SNP containing nucleic acid conferring at least one desirable plant characteristic selected from increased drought resistance, heat tolerance and lint yield in said plant. In certain embodiments, the target nucleic acid is amplified prior to detection. In certain embodiments, the step of detecting the presence of said SNP is performed using a process selected from the group consisting of detection of specific hybridization, measurement of allele size, restriction fragment length polymorphism analysis, allele-specific hybridization analysis, single base primer extension reaction, and sequencing of an amplified polynucleotide. In certain embodiments, the target nucleic acid is DNA. In certain embodiments, the SNP is in the DREB2A-A or HSFA6B-D binding site of IPS1-A. In certain embodiments, the SNP is present in tgtttccagaatattcccca(SEQ ID NO: 33) at nucleotide 19. In certain embodiments, the SNP is an A allele, a G allele, or a C allele. In certain embodiments, the SNP is a C allele and the presence of the C allele is associated with enhanced drought resistance when compared to plants lacking a C allele. In certain embodiments, the SNP is a C allele and the presence of the C allele is associated with enhanced drought resistance when compared to plants having a T allele. In certain embodiments, the SNP is an A allele and the presence of the A allele is associated with enhanced drought resistance when compared to plants lacking an A allele. In certain embodiments, the SNP is an A allele and the presence of the A allele is associated with enhanced drought resistance when compared to plants having a T allele. In certain embodiments, the SNP is a G allele and the presence of the G allele is associated with enhanced drought resistance when compared to plants lacking a G allele. In certain embodiments, the SNP is a G allele and the presence of the G allele is associated with enhanced drought resistance when compared to plants having a T allele. In certain embodiments, the plant is a cotton plant is selected from Gossypium hirsutum, Gossypium barbadense, Gossypium arboretum, and Gossypium raimondii. Another aspect of the invention comprises multiplex SNP panels. In certain embodiments, the multiplex SNP panel comprises nucleic acids informative of the presence of increased drought resistance, wherein said panel contains nucleic acids comprising at least one SNP in a DREB2A-A or HSFA6B-D binding site of IPS1-A. In certain embodiments, the SNP is present in nucleotide 19 of SEQ ID NO: 33. In certain embodiments, the SNP is a C allele, G allele, or an A allele. In certain embodiments, the C allele is present and the presence of the C allele is associated with enhanced drought resistance when compared to plants lacking a C allele. In certain embodiments, the multiplex SNP panel comprises nucleic acids informative of the presence of increased drought resistance, wherein said panel contains nucleic acids comprising SNPs as provided in Figure 8 or Figure 13. In another aspect of the invention, vectors comprising nucleic acids encoding a IPS1-A protein having at least one SNP in the DREB2A-A or HSFA6B binding site, the SNP comprising an allele that enhances drought resistance when compared to proteins without the allele, are provided. In certain embodiments, the target nucleic acid is DNA. In certain embodiments, the SNP is at least one SNP as provided in Figure 8, or 13E.In certain embodiments, the SNP is in the DREB2A-A or HSFA6B-D binding site of IPS1-A. In certain embodiments, the SNP is present in tgtttccagaatattcccca(SEQ ID NO: 33) at nucleotide 19. In certain embodiments, the SNP is an A allele, a G allele, or a C allele. In certain embodiments, the SNP is a C allele and the presence of the C allele is associated with enhanced drought resistance when compared to plants lacking a C allele. In certain embodiments, the SNP is a C allele and the presence of the C allele is associated with enhanced drought resistance when compared to plants having a T allele. In certain embodiments, the SNP is an A allele and the presence of the A allele is associated with enhanced drought resistance when compared to plants lacking an A allele. In certain embodiments, the SNP is an A allele and the presence of the A allele is associated with enhanced drought resistance when compared to plants having a T allele. In certain embodiments, the SNP is a G allele and the presence of the G allele is associated with enhanced drought resistance when compared to plants lacking a G allele. In certain embodiments, the SNP is a G allele and the presence of the G allele is associated with enhanced drought resistance when compared to plants having a T allele. Another aspect of the invention comprises a host cell comprising the vectors described herein. In another aspect of the invention, solid supports comprising the drought resistant SNP described herein is provided. Also provided herein are recombinant plants, including cotton plants, comprising nucleic acids harboring at least one SNP as provided in Figure 8, or Figure 13. Also provided herein are recombinant plants, including cotton plants, comprising nucleic acids harboring at least one SNP in a DREB2A-A or HSFA6B-D binding site of IPS1-A. In certain embodiments, the SNP is present in tgtttccagaatattcccca(SEQ ID NO: 33) at nucleotide 19. In certain embodiments, the SNP is an A allele, a G allele, or a C allele. In certain embodiments, the SNP is a C allele and the presence of the C allele is associated with enhanced drought resistance when compared to plants lacking a C allele. In certain embodiments, the SNP is a C allele and the presence of the C allele is associated with enhanced drought resistance when compared to plants having a T allele. In certain embodiments, the SNP is an A allele and the presence of the A allele is associated with enhanced drought resistance when compared to plants lacking an A allele. In certain embodiments, the SNP is an A allele and the presence of the A allele is associated with enhanced drought resistance when compared to plants having a T allele. In certain embodiments, the SNP is a G allele and the presence of the G allele is associated with enhanced drought resistance when compared to plants lacking a G allele. In certain embodiments, the SNP is a G allele and the presence of the G allele is associated with enhanced drought resistance when compared to plants having a T allele. In another aspect of the invention, methods of increasing drought resistance in a plant, such as a cotton plant, as compared to the wild type are provided. An exemplary method comprises transforming a plant cell with the vectors described herein are provided. In another aspect of the invention, methods of producing a transgenic plant with increased drought resistance as compared to the wild type, the method comprising transforming a plant cell with the vectors described herein and regenerating a plant from the plant cell are provided. Another aspect of the invention comprises methods of producing a transgenic plant with an increased drought resistance gene / protein expression profile as compared to the wild type, the method comprising crossing the recombinant plants described herein with a wild type plant. In certain embodiments, the plant is a cotton plant. In certain embodiments, the methods further comprise detecting in the plant at least one of the SNPs described herein. In certain embodiments, the second plant is selected from the cotton plant, a wild-type cotton plant, or a recombinant plant. In certain embodiments, the methods further comprise detecting at least one drought-resistant phenotype in said plant. In certain embodiments, the drought resistant phenotype is selected from increased lint yield, increased length uniformity (UI), scaled photochemical reflective index (sPRI), carotenoid reflectance index (CRI), and water index / normalized difference vegetation index (WI / NDVI). In another aspect of the invention, seeds, plants, or plant cells produced by the methods described herein are provided. Still other aspects and advantages of these compositions and methods for making the compositions and using the compositions are described further in the following detailed description of the preferred embodiments thereof. Brief Description of the Drawings Figures 1A-1E: Experiment information and genomic, metabolomic, and transcriptomic profiling of the cotton panel used in this work. Fig.1A) Overview of the watering regime and timing of data collection for the experiment in the 2019 summer field season at Maricopa, AZ. Water levels were reduced to 50% of normal for the treatment plots at the early flowering stage. Metabolomic and transcriptomic data were collected at late flowering / early boll development, with phenotypic measurements taken throughout. Fig.1B) A phylogeny of the accessions used in this study based on filtered SNPs derived from the transcriptomic data. The three groups, “Foreign”, “U.S.”, and “Mixed”, reflect historical breeding information obtained from USDA GRIN-Global. Fig.1C) Overview of the phenotype profiles of 21 upland cotton accessions under water-limited (WL) and well-watered (WW) conditions: A cluster heatmap and indicated the ratio of phenotypic data under WL condition relative to the WW condition (WL value / WW value, range: 0.8 - 1.2). Group coloring denotes accessions belonging to the three phylogenetic groups from Fig.1B. Fig.1D) Three-dimensional Principal Component Analysis (PCA) displaying the impact that treatment has on large-scale changes in normalized metabolite content between accessions. The different treatments have been denoted with light red (water-limited) and light blue (well-watered) ovals. Fig.1E) Three- dimensional PCA was used to display the top 10% (N = 3768) most variably expressed genes based on normalized read counts, with treatment groups denoted similarly to Fig.1D. Figure 2A-2B: Comparison of six production-related traits and four vegetative indices from two watering conditions. The 21 upland cotton accessions were split into “Foreign”, “U.S.”, and “Mixed” groups for pairwise comparisons among 10 traits from two water conditions. For each pair-wise comparison, a t-test was performed to effects of combined heat-drought stress. Figure 3A-3C: Transcriptome comparison of expression profiles in the three phylogenetic groups. Fig.3A) Upset plot of differentially expressed genes (DEGs) in all representatives within a particular phylogenetic group. The vertical black bars stand for the numbers of unique DEGs from each group, intersection between each two group, and intersection among three groups. Fig.3B and 3C) Enriched gene ontology terms of commonly down-regulated (top) and up-regulated (bottom) genes among 21 accessions. The calculation is based on -log10 Q values (cutoff:3). Figure 4A-4C: Weighted Co-expression network and trait-module membership. Fig. 4A) The left panel represents the frequency of gene connection revealed by connectivity (k). The right panel shows the log-log plot of the connectivity, the correlation coefficient represents the scale of free topology. The 0.86 here indicated the high approximation of a scale-free of network of network derived from top 10% median absolute deviation (MAD) score. Fig.4B) The left panel is the relation between approximation of a scale-free of network and different soft threshold (power). The seven is the least numbers to ensure a scale free network. The right panel stands for the mean connectivity of network when different power score was applied. Fig.4C) A heatmap plot depicting the topological overlap matrix (TOM) supplemented by hierarchical clustering dendrograms and the module colors was created by TOMplot function from the weighted gene co-expression analysis (WGCNA) package. Each cell corresponds to a single gene, the light color indicated the high topological overlap, and the darker orange and red colors represent the lower topological overlap. Figures 5A-5C: Trait and co-expression association in response to water deficit conditions Fig.5A) Modules of genes derived from a weighted gene co-expression analysis were correlated with phenotypic and metabolomic traits, with traits passing the significance threshold (P < 0.05) shown along the bottom. Modules are named according to color, with the number of genes within each module shown (gene counts). Positive (red hues) and negative (blue hues) correlation values are shown for all significant trait-module correlations with scale bar below. gene ontology (GO) term enrichment was performed on all genes within the module, with significant terms shown along the bottom. The number of genes associated with each GO term is shown within each box along with the “purple” to “yellow” scale displayed. Fig.5B) The number of transcription factors (TFs) with enriched binding motifs within their respective module (TF and motifs are present in the same module) are shown. Fig.5C) Coefficient of variance (CV) of expression (TPM) was calculated for all genes within each module across all accessions and both conditions. CV was also calculated for a background set of genes based on an average module size (N = 193). The red dashed box outlines the “turquoise” module, which was found to be correlated with both lint yield and stress-response genes and was selected as a key module (highlighted with a red star) for further examination. Figures 6A-6D: Identifying HSFA6B and DREB2A targets with DAP-seq. Fig.6A) Identity of the seven key hub TFs within the lint yield associated module (Top). For each TF, the -log10 (Adj-P value) level of motif enrichment among co-expressed genes within the module is shown, along with the reported stress response. For the four abiotic stress-associated TFs, their Analysis of Motif Enrichment (AME) predicted binding to the genes with enriched gene ontology (GO) terms within the module, as well as five GWAS-identified lint yield genes, is shown (bottom). As GhHSFA6B-D and GhDREB2A-A were both stress response regulators and predicted to bind to at least one of the five lint yield associated genes, they were chosen for DAP-seq library creation (red dashed box). Fig.6B) Illustration of annotating peaks derived from GhHSFA6B-D and GhDREB2A-A DAP-seq data. The top panel highlights DAP-seq peaks within the 5-Kb upstream regions (distal promoters) and 5’UTR regions (proximal promoters) of annotated genes. The bottom panel depicts the two consensus motifs identified from 3,180 peaks (GhHSFA6B-D) and 5,218 peaks (GhDREB2A-A) that passed stringency filters. Fig.6C) Comparison of the log2 fold change of transcript abundance (WL relative to WW) for GhDREB2A-A and GhHSFA6B-D DAP-seq targets between genes in the lint yield associated module (192 GhHSFA6B-D targets and 76 GhDREB2A-A targets) and genes selected from transcriptome-wide through random sampling (N values shown below). Pairwise significance was performed using a paired Student’s t-Test. Fig.6D) Comparison of enriched GO terms among genes bound by GhHSFA6B-D and / or GhDREB2A-A within the lint yield module and targets from a transcriptome-wide group of expressed transcripts. The level of enrichment for each GO term is reflected by -log10-Qvalue, with higher levels of enrichment corresponding to larger values. The “star” logo highlights the “stress response” enriched GO term. Figures 7A-7C: Preparation of DAP-seq libraries. Fig.7A) The genomic DNA of the TM-1 cotton genome was extracted with the gel running before and after the adaptor ligation; Fig.7B) The protein expression of the HSFA6B and DREB2A to confirm the expected size; Fig. 7C) Examination of the protein quality in between steps of beads binding, beads washing, and protein purification. Figures 8A-8E: Electrophoretic Mobility Shift Assay (EMSA) validation of interaction between ABP and two TFs. Fig.8A) Sequence alignment of GhABP and an GhABP paralog: promoter and coding region of a lint yield candidate ABP and its paralog were aligned using MUSCLE (iteration = 200), two peaks bound by GhDREB2A-A and GhHSFA6B-D were identified around the 1.5 Kb upstream of promoter regions (marked by red pentagram). Annotation of the peak region revealed three tandem heat shock elements (HSE, orange bracket), which overlapped with the motif region detected by motif scan (FIMO) and analysis (MEME) (purple and yellow bracket respectively). SNPs were detected within one HSE (orange box) while the rest HSEs are identical from two paralogs. One EMSA DNA probe and competitor pair was designed to span two identical HSEs and the other probe-competitor pair to cover the HSE bearing SNPs (bottom panel). Similarly, the GhDREB2A-A peak region was annotated and marked (top panel). Fig.8B) Pearson correlation between TPM values of GhABP-A and its two putative regulating TFs, GhDREB2A-A and GhHSFA6B-D. Fig.8C) Pearson correlation between the expression (TPM) GhABP-A (red) and its homeolog, GhABP-D (blue), with lint yield. Fig. 8D) Interaction between GhDREB2A-A and GhABP-A was examined by EMSA. The reaction components for each lane are listed below, including the IRD700 labeled probe, unlabeled probe (200×), competitor sequences (200×, containing 2 SNPs flanking the HSE), and empty vector. Fig.8E) Interaction between GhHSFA6B-D and GhABP-A was examined by EMSA. The reaction components for each lane are listed below, including the IRD700 labeled probe, unlabeled probe (200×), competitor sequences (200×, containing 3 SNPs flanking the HSE), and empty vector. Figures 9A-9B: Distribution of GhDREB2A and GhHSFA6B DAPseq peaks. An Integrative Genomics Viewer (IGV) screen shot revealed the binding region (peak) across the distal promoter regions of ABP (bound by both GhDREB2A-A and GhHSFA6B-D), IPS (bound by GhHSFA6B-D), and GhDREB2A-A (bound by GhHSFA6B-D) across two replicates of DAP- seq libraries. Figures 10A-10B: Integrated network display of a module of transcripts associated with lint yield. The module membership (MM) scores and gene significance (GS) for genes within the lint-yield module were plotted with the linear model fitted line (Fig.10A). The dashed black lines indicate the two cutoffs used to identify trait-related genes (orange dots, GS > 0.7, MM > 0.8). The integrated network derived from 133 out of 866 selected genes within the lint- yield associated module were presented by layering multiple levels of information (Fig.10B): gene categories are shown in the bottom line, as corresponded to the color of highlighted dots in MM-GS plot, including genes associated with stress response (blue), TFs (red), heat response (light purple), spliceosome pathways, biotic response pathways, and five lint yield QTLs (green), whereas connection types are shown in the top right. The solid “gray” lines connected the co- expressed genes from transcriptional level, the dashed “red” lines indicated protein-protein interactions derived from STRING database, the “blue” dash-arrowed lines highlighted targets bound by DREB2A and HSFA6B, as supported by DAP-seq data. Figure 11: Correlation between TFs and target genes (TPM) Pearson correlation between TPM values of IPS genes and HSFA6B (right panel), TPM values of GhABP-A and GhABP-D genes and GhHSFA6B (bottom panel), and TPM values of GhABP-A and GhABP-B genes and GhDREB2A (left panel). Figures 12A-12C: Sequence alignment and expression profile of four IPS in cotton. Fig.12A-12B) Sequence alignment of Four IPSs identified in upland cotton genome: promoter (left panel) and coding region (right panel) were aligned using MUSCLE (iteration = 200), only one peak bound by GhHSFA6B-D was identified around the 3.1 Kb upstream of promoter regions (marked by red pentagram). Fig.12C) Correlation between TPM value of four GhIPSs and lint yield, tested by Pearson correlation among 21 accessions. Figures 13A-13E: EMSA validation of interaction between GhIPS1-A and GhHSFA6B-D. Fig.13A) Top: schematic representation of GhIPS1-A and its homeolog, GhIPS1-D depicting the gene structure as well as the distal promoter region where the HSFA6B binding site was identified. Filled boxes represent annotated exons. Grey shading connecting the two genes represents sequence similarity, with the distal promoter region showing high structural and sequence variation. Bottom: An alignment for the DAP-seq peak for HSFA6B in the promoter region of GhIPS1-A and its corresponding region from GhIPS1-D. Also shown are the DAP-seq and Cistrome identified GhHSFA6B-D binding site (HSE, purple and yellow boxes, respectively), as well as the labeled probe and competitor oligos used for EMSA. A green asterisk denotes the site of the C:T SNP observed within our panel. Fig.13B) Pearson correlation between GhIPS1-A (red), GhIPS1-D (blue) expression (TPM), and lint yield. Fig.13C) Pearson correlation between TPM values of GhIPS1-A and GhHSFA6B-D. Fig.13D) Variation in lint yield between accessions associated with “C / C” (blue) and “T / T” (brown) genotypes of a SNP adjacent to the HSE element in the distal promoter of GhIPS1-A (~3.1 Kb upstream of GhIPS1-A start). Accessions lacking sufficient coverage at this site are shown in grey. Accessions bearing the C / C or T / T genotypes were also divided based on their phylogenetic groupings (i.e., “U.S.”, “Foreign”, and “Mixed”) from Fig.1A. Fig.13E) Interaction between the GhHSFA6B-A protein and GhIPS1-A probe was examined by EMSA. The reaction components for each lane are listed below the gel image, including the IRD700 labeled probe, unlabeled probe (200×, same sequence as the labeled probe), competitor sequence (200×, 400×, and 600×, containing the GhIPS1-D disrupted HSE), empty vector, and the other competitor sequence (200×, containing one of the three SNP probes, A, T, or G). Figures 14A-14C: Identification of a HSE-adjacent SNP that impacts GhHSFA6B-D binding to GhIPS1-A. Fig.14A) A phylogenetic tree containing GhIPS1 and GhIPS2 homeologous genes from G. hirsutum (AADD), G. barbadense (AADD), G. arboreum (AA), and G. raimondii (DD) was constructed using Arabidopsis thaliana IPS homologs as an outgroup. The “orange” and “light blue” label denotes the presence of the HSE upstream of GhIPS1-A genes in all AA genome containing species, but not in the distal promoters of GhIPS1- D genes found in DD genome species, nor in any of the GhIPS2 paralogs in either AA or DD genomes. Fig.14B) Distribution of genotypes of the GhIPS1-A DAP-seq peak-associated SNP (C:T) in a published whole genome sequencing-based panel (1024 accessions Yuan et al., 2021). The distribution is summarized in pie-charts reflecting the accession frequencies of the wild, cultivated, mutants, and improved G. hirsutum accessions across the “C / C” genotype (reference genome genotype), “T / T” genotype, and “T / C” genotype. Fig.14C) The level of the GhHSFA6B-D to GhIPS1-A binding efficiency of different IPS-associated SNP variants was examined by competitions between the IDR700 dye-labeled reference oligos and non-labeled oligos with the reference nucleotide “C”, or the non-native SNP oligos “G”, and “A”, respectively. Left panel: binding efficiency of the reference “C” containing probe was tested by titrating in increasing amounts of the competitor probe (unlabeled reference “C”). Right panel: the impact of this nucleotide position on GhHSFA6B-D binding was tested by competing the labeled reference probe with increasing amounts of two non-native probes (“G” or “A”). Components in each reaction are shown below. Figure 15: Comparison of HSFA6B - IPS binding among three genotypes. The band area of “Protein + probe” and “probe” region were extracted by calculating the pixels of bands using ImageJ software. To illustrate the variations among three types of nucleotides (reference genotype “C” and two naturally non-existed genotypes “G” and “A”) across a concentration gradient of non-labeled probe ( 1×, 2×, 5×, 10×, 20× relative to labeled probe, X-axis in the line graph), we used the ratio = (probe + protein) / probe (Y-axis in the line graph) to represent the level of competition between labeled probe and non-labeled probe. The higher ratio indicates the lower chances of the binding of non-labeled probe. Figure 16: Comparison of global sub-genome expression bias. A total of 57,763 genes (average TPM among samples > 1) were retained to study global sub-genome dominance based on G.hirsutum gene annotation. The log2 transformed TPM score was used to indicate the transcription abundance for each gene. The pair-wise t-test was performed to examine the mean gene expression biases between the A and D sub-genomes (P < 0.05, represented by light blue and yellow boxes respectively). Each accession was tested under two watering conditions. Figure 17: Summary of 23 cotton accessions, sequencing and mapping rate based on G. hirsutum reference genome. Figure 18: Summary of core genes regulated by HSFA6B and DREB2A in cotton. Figure 19A-19D: GhHSFA6B – GhIps1-A promoter binding: Microscale thermophoresis was performed to detect in vitro binding between the transcription factor, GhHSFA6B, and the heat shock promoter element associated with the GhIps1-A locus. Fig. 19A) Binding between GhHSFA6B and the native promoter element, Fig.19B) binding between GhHSFA6B and an improved sequence, and between a scrambled Fig.19C) and telomeric Fig. 19D) oligos (no binding detected). Figure 20: Conservation of HSE in the IPS locus of various dicots. Detailed Description of the Invention The information presented herein can be leveraged for the development of new elite cotton cultivars with improved adaptation to the hotter and drier climatic conditions that are predicted for the near future. The substantial impact of drought stress on crop physiology results in the alteration of growth and productivity. Understanding the genetic and molecular crosstalk between stress responses and agronomically important traits (such as fiber yield) is particularly complicated in the allopolyploid species, upland cotton (Gossypium hirsutum), due to reduced sequence variability between A and D subgenomes. To better understand how drought stress impacts yield, transcriptome sequencing was performed for 22 upland cotton accessions grown in the Arizona low desert and exposed to both well-watered and water-limited conditions. Phenotypic and metabolomic data were integrated into a co-expression network analysis to identify genes associated with improved yield under water-limited conditions. DAP-sequencing was performed on two transcription factors, GhDREB2A-A and GhHSFA6B-D, shown to have the highest association with lint yield and stress response genes across the panel. Using transcriptomic, biochemical, and phylogenomic approaches, GhHSFA6B-D was shown to bind to and positively regulate the lint yield-associated gene, GhIPS1-A, and the associated GhHSFA6B-D regulatory element within GhIPS1-A is only present in the A sub-genome and A-genome diploid progenitors. We also identify a lint yield- associated single nucleotide polymorphism directly adjacent to this GhIPS1-A regulatory element that influences GhHSFA6B-D binding suggesting additional selection at this locus during domestication. Definitions For purposes of the present invention, “a” or “an” entity refers to one or more of that entity; for example, “a cDNA” refers to one or more cDNA or at least one cDNA. As such, the terms “a” or “an,” “one or more” and “at least one” can be used interchangeably herein. It is also noted that the terms “comprising,” “including,” and “having” can be used interchangeably. Furthermore, a compound “selected from the group consisting of” refers to one or more of the compounds in the list that follows, including mixtures (i.e., combinations) of two or more of the compounds. According to the present invention, an isolated, or biologically pure molecule is a compound that has been removed from its natural milieu. As such, “isolated” and “biologically pure” do not necessarily reflect the extent to which the compound has been purified. An isolated compound of the present invention can be obtained from its natural source, can be produced using laboratory synthetic techniques or can be produced by any such chemical synthetic route. The term "genetic alteration" as used herein refers to a change from the wild-type or reference sequence of one or more nucleic acid molecules. Genetic alterations include without limitation, base pair substitutions, additions and deletions of at least one nucleotide from a nucleic acid molecule of known sequence. A “single nucleotide polymorphism (SNP)” refers to a change in which a single base in the DNA differs from the usual base at that position. These single base changes are called SNPs or "snips." Millions of SNP's have been cataloged in the plant genomes. Some SNPs such as that which are responsible for disease. Other SNPs are normal variations in the genome. A “drought-associated SNP” or “drought-associated specific marker” or “drought- associated informational sequence molecule” is a SNP or marker sequence which is associated with an altered response to drought not observed in plants lacking this SNP or marker sequence. The presence of such markers in a plant convey a drought-resistant phenotype to the plant. These drought-resistant phenotypes include without limitation, improved lint yield, length uniformity (UI), scaled photochemical reflective index (sPRI), carotenoid reflectance index (CRI), and water index / normalized difference vegetation index (WI / NDVI) when compared to plants lacking said markers. Such markers may include but are not limited to nucleic acids, proteins encoded thereby, or other small molecules. Thus, the phrase “drought-associated SNP containing nucleic acid” is encompassed by the above description. In certain embodiments, the drought-associated SNP is present in IPS1-A, such as GhIPS1-A. In certain embodiments, the SNP is in the DREB2A-A or HSFA6B-D binding site of IPS1-A. In certain embodiments, the SNP is at least one of the SNPs identified in Figure 8 or Figure 13E. In certain embodiments, the SNP is present in SEQ ID NO: 33. In certain embodiments, SEQ ID NO: 33 is the DREB2A-A or HSFA6B-D binding site of IPS1-A. In certain embodiments, the SNP is at nucleotide 19 of SEQ ID NO: 33. In certain embodiments, the SNP is an A allele, a G allele, or a C allele. In certain embodiments, the SNP is a C allele and the presence of the C allele is associated with enhanced drought resistance when compared to plants lacking a C allele. In certain embodiments, the SNP is a C allele and the presence of the C allele is associated with enhanced drought resistance when compared to plants having a T allele. In certain embodiments, the SNP is an A allele and the presence of the A allele is associated with enhanced drought resistance when compared to plants lacking an A allele. In certain embodiments, the SNP is an A allele and the presence of the A allele is associated with enhanced drought resistance when compared to plants having a T allele. In certain embodiments, the SNP is a G allele and the presence of the G allele is associated with enhanced drought resistance when compared to plants lacking a G allele. In certain embodiments, the SNP is a G allele and the presence of the G allele is associated with enhanced drought resistance when compared to plants having a T allele. The IPS gene encodes for the myo‐inositol‐phosphate synthase protein (INO‐1), which catalyzes the rate‐limiting step in the synthesis of myo‐inositol‐6‐phosphate, a key source of phosphate in seed endosperm. In addition, myo‐inositol is a precursor of the osmoprotectants galactinol and raffinose and thus is critical in a number of abiotic and biotic stress responses. IPS1-A contains DREB2A and HSFA6B-D binding sites and is present in multiple plant genomes. Exemplary IPS1-A genes include G. hjirsutum (GhIPS1-A – SEQ ID NO: 34), G. max (GmIPS1-A – SEQ ID NO: 35), A. thaliana (MIPS1 – SEQ ID NO: 36 and MIPS2 – SEQ ID NO: 37). Figure 20 shows that the HSE is present at orthologous IPS loci in these plants. In certain embodiments, the SNP is present in the DREB3A and / or HSFA6B-D binding site of any one of SEQ ID NO: 34-37. In certain embodiments, the DREB3A and / or HSF5B-D binding site is in SEQ ID NO: 34 at nucleotides 2356 to 2375. In certain embodiments, the SNP is present in SEQ ID NO: 34 at nucleotide 2374. A “copy number variation (CNV)” refers to the number of copies of a particular gene or segment thereof in the genome of an individual plant. CNVs represent a major genetic component of phenotypic diversity. Susceptibility to genetic traits is known to be associated not only with single nucleotide polymorphisms (SNP), but also with structural and other genetic variations, including CNVs. A CNV represents a copy number change involving a DNA fragment that is 1 kilobases (kb) or larger (Feuk et al.2006a). CNVs described herein do not include those that arise from the insertion / deletion of transposable elements (e.g., 6-kb KpnI repeats) to minimize the complexity of future CNV analyses. The term CNV therefore encompasses previously introduced terms such as large-scale copy number variants (LCVs; Iafrate et al.2004), copy number polymorphisms (CNPs; Sebat et al.2004), and intermediate- sized variants (ISVs; Tuzun et al.2005), but not retroposon insertions. The terminology “duplication-containing CNV” is also used herein below consistent with the CNV definition provided. The term "solid matrix" as used herein refers to any format, such as beads, microparticles, a microarray, the surface of a microtitration well or a test tube, a dipstick or a filter. The material of the matrix may be polystyrene, cellulose, latex, nitrocellulose, nylon, polyacrylamide, dextran or agarose. Drought-associated nucleic acids may be affixed or immobilized to a solid matrix. Affixed or immobilized as used herein refers to a linkage that is stable in solution, such that the nucleic acids remain attached to the solid matrix under different processing or experimental conditions. The phrase “consisting essentially of” when referring to a particular nucleotide or amino acid means a sequence having the properties of a given SEQ ID NO:. For example, when used in reference to an amino acid sequence, the phrase includes the sequence per se and molecular modifications that would not affect the functional and novel characteristics of the sequence. The phrase “partial informative CNV” is used herein to refer to a nucleic acid that hybridizes to sequences comprising a duplication on a chromosome however, the partial informative CNV may not be identical to the duplication, rather, the CNV may correspond to only a portion of the duplication, but yet is still informative for the same. "Target nucleic acid" as used herein refers to a previously defined region of a nucleic acid present in a complex nucleic acid mixture wherein the defined wild-type region contains at least one known nucleotide variation which may or may not be associated with drought-resistance but is informative of an altered response to drought. The nucleic acid molecule may be isolated from a natural source by cDNA cloning or subtractive hybridization or synthesized manually. The nucleic acid molecule may be synthesized manually by the triester synthetic method or by using an automated DNA synthesizer. When cloning a target nucleic acid comprising a deletion, the skilled artisan is well aware of methods for selecting nucleic acids of a sufficient length flanking the affected region to facilitate cloning the region into a vector of choice. With regard to nucleic acids used in the invention, the term "isolated nucleic acid" is sometimes employed. This term, when applied to DNA, refers to a DNA molecule that is separated from sequences with which it is immediately contiguous (in the 5' and 3' directions) in the naturally occurring genome of the organism from which it was derived. For example, the “isolated nucleic acid” may comprise a DNA molecule inserted into a vector, such as a plasmid or virus vector, or integrated into the genomic DNA of a prokaryote or eukaryote. An "isolated nucleic acid molecule" may also comprise a cDNA molecule. An isolated nucleic acid molecule inserted into a vector is also sometimes referred to herein as a recombinant nucleic acid molecule. With respect to RNA molecules, the term “isolated nucleic acid” primarily refers to an RNA molecule encoded by an isolated DNA molecule as defined above. Alternatively, the term may refer to an RNA molecule that has been sufficiently separated from RNA molecules with which it would be associated in its natural state (i.e., in cells or tissues), such that it exists in a “substantially pure” form. By the use of the term “enriched” in reference to nucleic acid it is meant that the specific DNA or RNA sequence constitutes a significantly higher fraction (2-5 fold) of the total DNA or RNA present in the cells or solution of interest than in normal cells or in the cells from which the sequence was taken. This could be caused by a person by preferential reduction in the amount of other DNA or RNA present, or by a preferential increase in the amount of the specific DNA or RNA sequence, or by a combination of the two. However, it should be noted that “enriched” does not imply that there are no other DNA or RNA sequences present, just that the relative amount of the sequence of interest has been significantly increased. It is also advantageous for some purposes that a nucleotide sequence be in purified form. The term "purified" in reference to nucleic acid does not require absolute purity (such as a homogeneous preparation); instead, it represents an indication that the sequence is relatively purer than in the natural environment (compared to the natural level, this level should be at least 2-5 fold greater, e.g., in terms of mg / ml). Individual clones isolated from a cDNA library may be purified to electrophoretic homogeneity. The claimed DNA molecules obtained from these clones can be obtained directly from total DNA or from total RNA. The cDNA clones are not naturally occurring, but rather are preferably obtained via manipulation of a partially purified naturally occurring substance (messenger RNA). The construction of a cDNA library from mRNA involves the creation of a synthetic substance (cDNA) and pure individual cDNA clones can be isolated from the synthetic library by clonal selection of the cells carrying the cDNA library. Thus, the process which includes the construction of a cDNA library from mRNA and isolation of distinct cDNA clones yields an approximately 10-6-fold purification of the native message. Thus, purification of at least one order of magnitude, preferably two or three orders, and more preferably four or five orders of magnitude is expressly contemplated. Thus, the term "substantially pure" refers to a preparation comprising at least 50-60% by weight the compound of interest (e.g., nucleic acid, oligonucleotide, etc.). More preferably, the preparation comprises at least 75% by weight, and most preferably 90-99% by weight, the compound of interest. Purity is measured by methods appropriate for the compound of interest. The term "complementary" describes two nucleotides that can form multiple favorable interactions with one another. For example, adenine is complementary to thymine as they can form two hydrogen bonds. Similarly, guanine and cytosine are complementary since they can form three hydrogen bonds. Thus, if a nucleic acid sequence contains the following sequence of bases, thymine, adenine, guanine and cytosine, a "complement" of this nucleic acid molecule would be a molecule containing adenine in the place of thymine, thymine in the place of adenine, cytosine in the place of guanine, and guanine in the place of cytosine. Because the complement can contain a nucleic acid sequence that forms optimal interactions with the parent nucleic acid molecule, such a complement can bind with high affinity to its parent molecule. With respect to single stranded nucleic acids, particularly oligonucleotides, the term "specifically hybridizing" refers to the association between two single-stranded nucleotide molecules of sufficiently complementary sequence to permit such hybridization under pre- determined conditions generally used in the art (sometimes termed "substantially complementary"). In particular, the term refers to hybridization of an oligonucleotide with a substantially complementary sequence contained within a single-stranded DNA or RNA molecule of the invention, to the substantial exclusion of hybridization of the oligonucleotide with single-stranded nucleic acids of non-complementary sequence. For example, specific hybridization can refer to a sequence which hybridizes to any drought-associated specific marker gene or nucleic acid but does not hybridize to other nucleotides. Also, polynucleotide which “specifically hybridizes” may hybridize only to a drought-associated specific marker shown in the Tables contained herein. Appropriate conditions enabling specific hybridization of single stranded nucleic acid molecules of varying complementarity are well known in the art. For instance, one common formula for calculating the stringency conditions required to achieve hybridization between nucleic acid molecules of a specified sequence homology is set forth below (Sambrook et al., Molecular Cloning, Cold Spring Harbor Laboratory (1989): Tm = 81.5ºC + 16.6Log [Na+] + 0.41(% G+C) - 0.63 (% formamide) - 600 / #bp in duplex As an illustration of the above formula, using [Na+] = [0.368] and 50% formamide, with GC content of 42% and an average probe size of 200 bases, the Tm is 57ºC. The Tm of a DNA duplex decreases by 1 - 1.5ºC with every 1% decrease in homology. Thus, targets with greater than about 75% sequence identity would be observed using a hybridization temperature of 42ºC. The stringency of the hybridization and wash depend primarily on the salt concentration and temperature of the solutions. In general, to maximize the rate of annealing of the probe with its target, the hybridization is usually carried out at salt and temperature conditions that are 20- 25°C below the calculated Tm of the hybrid. Wash conditions should be as stringent as possible for the degree of identity of the probe for the target. In general, wash conditions are selected to be approximately 12-20°C below the Tm of the hybrid. In regard to the nucleic acids of the current invention, a moderate stringency hybridization is defined as hybridization in 6X SSC, 5X Denhardt’s solution, 0.5% SDS and 100 μg / ml denatured salmon sperm DNA at 42°C, and washed in 2X SSC and 0.5% SDS at 55°C for 15 minutes. A high stringency hybridization is defined as hybridization in 6X SSC, 5X Denhardt’s solution, 0.5% SDS and 100 μg / ml denatured salmon sperm DNA at 42°C, and washed in 1X SSC and 0.5% SDS at 65°C for 15 minutes. A very high stringency hybridization is defined as hybridization in 6X SSC, 5X Denhardt’s solution, 0.5% SDS and 100 μg / ml denatured salmon sperm DNA at 42°C, and washed in 0.1X SSC and 0.5% SDS at 65°C for 15 minutes. The term "oligonucleotide," as used herein is defined as a nucleic acid molecule comprised of two or more ribo- or deoxyribonucleotides, preferably more than three. The exact size of the oligonucleotide will depend on various factors and on the particular application and use of the oligonucleotide. Oligonucleotides, which include probes and primers, can be any length from 3 nucleotides to the full length of the nucleic acid molecule, and explicitly include every possible number of contiguous nucleic acids from 3 through the full length of the polynucleotide. Preferably, oligonucleotides are at least about 10 nucleotides in length, more preferably at least 15 nucleotides in length, more preferably at least about 20 nucleotides in length. Methods of screening for plants having improved characteristics The present invention provides methods of identifying plants, particularly cotton plants, having improved growth or tolerance characteristics, including without limitation, drought- tolerance or resistance, heat tolerance, increased yield, etc. The methods include the steps of providing a biological sample from the plant to be tested, measuring the amount of particular sets, or any or all of drought-associated markers present in the biological sample, and identifying the plant as having a greater or lesser resistance to drought based on the amount and / or type of drought-associated marker expression level(s) measured relative to those expression levels identified in known drought resistant or drought sensitive plants. Test plants having marker expression profiles associated with drought resistance observed in drought resistance reference plants, having greater drought resistance. Accordingly, the compositions and methods of the invention are useful for the identification of plants with drought resistance and plants that produce higher lint yield under drought and heat conditions relative to unimproved lines. The compositions and methods further identify new biomarkers useful for modulating the drought response. In certain embodiments, drought resistance refers to a plant having at least drought- resistant phenotype. Drought-resistant phenotypes include without limitation, improved lint yield, length uniformity (UI), scaled photochemical reflective index (sPRI), carotenoid reflectance index (CRI), and water index / normalized difference vegetation index (WI / NDVI) when compared to unimproved lines. In another aspect, the plant may have been previously genotyped and thus the genetic sequence and expression profile in the sample may be available. Accordingly, the method may entail storing reference drought-associated marker sequence information in a database, i.e., those SNPs statistically associated with a more favorable or less favorable phenotype as described herein, and performance of comparative genetic analysis on the computer, thereby identifying those plants having a resistance to drought. Drought-resistance SNP-containing nucleic acids, including but not limited to those listed below may be used for a variety of purposes in accordance with the present invention. Drought- associated SNP-containing DNA, RNA, or fragments thereof may be used as probes to detect the presence of and / or expression of drought-associated specific markers. Methods in which drought-associated specific marker nucleic acids may be utilized as probes for such assays include, but are not limited to: (1) in situ hybridization; (2) Southern hybridization (3) northern hybridization; and (4) assorted amplification reactions such as polymerase chain reactions (PCR). The term "probe" as used herein refers to an oligonucleotide, polynucleotide or nucleic acid, either RNA or DNA, whether occurring naturally as in a purified restriction enzyme digest or produced synthetically, which is capable of annealing with or specifically hybridizing to a nucleic acid with sequences complementary to the probe. A probe may be either single stranded or double stranded. The exact length of the probe will depend upon many factors, including temperature, source of probe and use of the method. For example, for diagnostic applications, depending on the complexity of the target sequence, the oligonucleotide probe typically contains 1525 or more nucleotides, although it may contain fewer nucleotides. The probes herein are selected to be complementary to different strands of a particular target nucleic acid sequence. This means that the probes must be sufficiently complementary so as to be able to "specifically hybridize" or anneal with their respective target strands under a set of pre-determined conditions. Therefore, the probe sequence need not reflect the exact complementary sequence of the target. For example, a non-complementary nucleotide fragment may be attached to the 5' or 3' end of the probe, with the remainder of the probe sequence being complementary to the target strand. Alternatively, non-complementary bases or longer sequences can be interspersed into the probe, provided that the probe sequence has sufficient complementarity with the sequence of the target nucleic acid to anneal therewith specifically. Further, assays for detecting drought-associated SNPs may be conducted on any type of biological sample. Clearly, drought-associated SNP-containing nucleic acids, vectors expressing the same, drought-associated SNP-containing marker proteins and anti-drought-associated specific marker antibodies of the invention can be used to detect drought-associated SNPs in the plant, cells, or fluid, and alter drought-associated SNP-containing marker protein expression for purposes of assessing the genetic and protein interactions involved in the development of drought resistance. In most embodiments for screening for drought-associated SNPs, the drought-associated SNP-containing nucleic acid in the sample will initially be amplified, e.g. using PCR, to increase the amount of the templates as compared to other sequences present in the sample. This allows the target sequences to be detected with a high degree of sensitivity if they are present in the sample. This initial step may be avoided by using highly sensitive array techniques that are important in the art. Alternatively, new detection technologies can overcome this limitation and enable analysis of small samples containing as little as 1μg of total RNA. Using Resonance Light Scattering (RLS) technology, as opposed to traditional fluorescence techniques, multiple reads can detect low quantities of mRNAs using biotin labeled hybridized targets and anti-biotin antibodies. Another alternative to PCR amplification involves planar wave guide technology (PWG) to increase signal-to-noise ratios and reduce background interference. Both techniques are commercially available from Qiagen Inc. (USA). Any of the aforementioned techniques may be used to detect or quantify drought- associated SNP marker expression and accordingly, identify drought-resistant plants. Vectors In many instances the nucleotide sequences of the drought-associated SNPs for use in the methods of the present invention, are provided in transcriptional units for transcription in the plant of interest. A transcriptional unit is comprised generally of a promoter and a nucleotide sequence operably linked in the 3′ direction of the promoter, optionally with a terminator. Selectable marker gene refers to a gene that upon expression confers a phenotype by which successfully transformed plants or cells or tissues carrying the transformed plastid can be identified. The term “expression” as used herein in the context of a gene product refers to the biosynthesis of that gene product, including the transcription and / or translation of the gene product. Inhibition or enhancement of expression or function of a target gene product (i.e., a gene product of interest) can be in the context of a comparison between any two plants, for example, expression or function of a target gene product in a genetically altered plant versus the expression or function of that target gene product in a corresponding wild-type plant. Alternatively, inhibition or enhancement of expression or function of the target gene product can be in the context of a comparison between plant cells, organelles, organs, tissues, or plant parts within the same plant or between plants and includes comparisons between developmental or temporal stages within the same plant or between plants. Any method or composition that down- or up-regulates expression of a target gene product, either at the level of transcription or translation, or down- or up-regulates functional activity of the target gene product can be used to achieve inhibition or enhancement of expression or function of the target gene product. A “specific binding pair” comprises a specific binding member (sbm) and a binding partner (bp) which have a particular specificity for each other and which in normal conditions bind to each other in preference to other molecules. Examples of specific binding pairs are antigens and antibodies, biotin and streptavidin, ligands and receptors and complementary nucleotide sequences. The skilled person is aware of many other examples. Further, the term “specific binding pair” is also applicable where either or both of the specific binding member and the binding partner comprise a part of a large molecule. In embodiments in which the specific binding pair comprises nucleic acid sequences, they will be of a length to hybridize to each other under conditions of the assay, preferably greater than 10 nucleotides long, more preferably greater than 15 or 20 nucleotides long. According to the present invention, an isolated or biologically pure molecule or cell is a compound that has been removed from its natural milieu. As such, “isolated” and “biologically pure” do not necessarily reflect the extent to which the compound has been purified. An isolated compound of the present invention can be obtained from its natural source, can be produced using laboratory synthetic techniques or can be produced by any such chemical synthetic route. The term “vector” relates to a single or double stranded circular nucleic acid molecule that can be infected, transfected or transformed into cells and replicate independently or within the host cell genome. A circular double stranded nucleic acid molecule can be cut and thereby linearized upon treatment with restriction enzymes. An assortment of vectors, restriction enzymes, and the knowledge of the nucleotide sequences that are targeted by restriction enzymes are readily available to those skilled in the art, and include any replicon, such as a plasmid, cosmid, bacmid, phage or virus, to which another genetic sequence or element (either DNA or RNA) may be attached so as to bring about the replication of the attached sequence or element. A nucleic acid molecule of the invention can be inserted into a vector by cutting the vector with restriction enzymes and ligating the two pieces together. The introduced nucleic acid may or may not be integrated (covalently linked) into nucleic acid of the recipient cell or organism. In plant cells, for example, the introduced nucleic acid may be maintained as an episomal element or independent replicon, such as a plasmid. Alternatively, the introduced nucleic acid may become integrated into the nucleic acid of the recipient cell or organism and be stably maintained in that cell or organism and further passed on or inherited to progeny cells or organisms of the recipient cell or organism. Finally, the introduced nucleic acid may exist in the recipient cell or host organism only transiently. In certain embodiments, the invention includes plant transformation vectors for the transformation of the plant. DNA constructs or vectors of the invention may be introduced into the genome of the desired plant host by a variety of conventional techniques. For example, the DNA construct may be introduced directly into the genomic DNA of the plant cell using techniques such as electroporation and microinjection of plant cell protoplasts, or the DNA constructs can be introduced directly to plant tissue using ballistic methods, such as DNA particle bombardment. Alternatively, the DNA constructs may be combined with suitable T- DNA flanking regions and introduced into a conventional Agrobacterium tumefaciens host vector. The virulence functions of the Agrobacterium tumefaciens host will direct the insertion of the construct and adjacent marker into the plant cell DNA when the cell is infected by the bacteria. The vectors of the present invention may comprise a coding region for a plant selectable marker gene, to select transformed plant cells with the corresponding selection agent. The plant selectable marker may provide resistance to a positive selection compound, for example, antibiotic resistance (e.g., streptomycin / spectinomycin, kanamycin, chloramphenicol, tobramycin, gentamycin, etc.), or herbicide resistance (e.g., including but not limited to: glyphosate, Dicamba, glufosinate, sulfonylureas, imidazolinones, bromoxynil, dalapon, cyclohexanedione, protoporphyrinogen oxidase inhibitors, and isoxaflutole herbicides). Polynucleotide molecules encoding proteins involved in herbicide tolerance are known in the art and can also be introduced into target plant cells using the compositions and methods of the invention. In addition to a plant selectable marker, in some embodiments it may be desirable to use a reporter gene. As used herein, the terms “reporter,” “reporter system”, “reporter gene,” or “reporter gene product” shall mean an operative genetic system in which a nucleic acid comprises a gene that encodes a product that when expressed produces a reporter signal that is a readily measurable, e.g., by biological assay, immunoassay, radio immunoassay, or by colorimetric, fluorogenic, chemiluminescent or other methods. GFP is exemplified herein. The nucleic acid may be either RNA or DNA, linear or circular, single or double stranded, and is operatively linked to the necessary control elements for the expression of the reporter gene product. The required control elements will vary according to the nature of the reporter system and whether the reporter gene is in the form of DNA or RNA, but may include, but not be limited to, such elements as promoters, enhancers, translational control sequences, poly A addition signals, transcriptional termination signals and the like. The term “primer” as used herein refers to an oligonucleotide, either RNA or DNA, either single-stranded or double-stranded, either derived from a biological system, generated by restriction enzyme digestion, or produced synthetically which, when placed in the proper environment, is able to functionally act as an initiator of template-dependent nucleic acid synthesis. When presented with an appropriate nucleic acid template, suitable nucleoside triphosphate precursors of nucleic acids, a polymerase enzyme, suitable cofactors and conditions such as a suitable temperature and pH, the primer may be extended at its 3′ terminus by the addition of nucleotides by the action of a polymerase or similar activity to yield a primer extension product. The primer may vary in length depending on the particular conditions and requirement of the application. For example, in diagnostic applications, the oligonucleotide primer is typically 15-25, 30, 50, 75 or more nucleotides in length. The primer must be of sufficient complementarity to the desired template to prime the synthesis of the desired extension product, that is, to be able anneal with the desired template strand in a manner sufficient to provide the 3′ hydroxyl moiety of the primer in appropriate juxtaposition for use in the initiation of synthesis by a polymerase or similar enzyme. It is not required that the primer sequence represent an exact complement of the desired template. For example, a non-complementary nucleotide sequence may be attached to the 5′ end of an otherwise complementary primer. Alternatively, non-complimentary bases may be interspersed within the oligonucleotide primer sequence, provided that the primer sequence has sufficient complementarity with the sequence of the desired template strand to functionally provide a template-primer complex for the synthesis of the extension product. As used herein "promoter" includes reference to a region of DNA involved in recognition and binding of RNA polymerase and other proteins to initiate transcription. Promoters useful in some embodiments of the present invention may be tissue-specific or cell-specific. The term “tissue-specific” as it applies to a promoter refers to a promoter that is capable of directing selective expression of a nucleotide sequence of interest to a specific type of tissue in the relative absence of expression of the same nucleotide sequence of interest in a different type of tissue (e.g., flower vs. root vs. leaf). As used herein "operably linked" includes reference to a functional linkage between two sequences in a nucleic acid construct, for example, a promoter and a second sequence, wherein the promoter sequence initiates and mediates transcription of the DNA sequence corresponding to the second sequence. Generally, operably linked means that the nucleic acid sequences being linked are contiguous and, where necessary to join two protein coding regions, contiguous and in the same reading frame. Generally, in the context of an over expression cassette, operably linked means that the nucleotide sequences being linked are contiguous and, where necessary to join two or more protein coding regions, contiguous and in the same reading frame. In the case where an expression cassette contains two or more protein coding regions joined in a contiguous manner in the same reading frame, the encoded polypeptide is herein defined as a “heterologous polypeptide” or a “chimeric polypeptide” or a “fusion polypeptide”. The cassette may additionally contain at least one additional coding sequence to be co-transformed into the organism. Alternatively, the additional coding sequence(s) can be provided on multiple expression cassettes. Methods of Generating Recombinant Plants Plants can be transformed to express nucleic acids harboring sequences of interest from a variety of cellular compartments, including the nucleus, mitochondria, seeds and plastids. Transformation refers to the stable integration of transforming DNA and constitutive expression into the target genome that is transmitted to the progeny of plants containing the transformed organelles. Transient expression of heterologous DNA in these different cellular compartments can also be employed. Transforming DNA refers to homologous DNA, or heterologous DNA flanked by homologous DNA, which when introduced into the organelle of interest becomes part of the genome by homologous recombination. The terms "transform", "transfect", "transduce", shall refer to any method or means by which a nucleic acid is introduced into a cell or host organism and may be used interchangeably to convey the same meaning. Such methods include, but are not limited to, transfection, electroporation, microinjection, PEG-fusion, biolistic bombardment and the like. The term “delivery” as used herein refers to the introduction of foreign molecule (i.e., miRNA encoding the polypeptide of interest) into cells. The term “administration” as used herein means the introduction of a foreign molecule into a cell. The term is intended to be synonymous with the term “delivery”. The invention discloses methods of expressing the newly discovered sequences associated with improved plant characteristics in any plant of interest. Another benefit of the invention is that the improved characteristic, e.g., drought resistance or high lint yield associated SNP can be directly delivered to a plant cell in conjunction with stable incorporation of a transgene. This allows for a broad range of utilities from increasing the efficiency of plant cell transformation to enhancing agroinfection. Mao et al. provide detailed guidance for use of the CRISPR / Cas system in higher plants in Molecular Plant, 6: 2008–2011 (2013). The article entitled “Application of the CRISPR–Cas System for Efficient Genome Engineering in Plants” and its supplemental material is incorporated herein by reference as though set forth in full. "Floral dip transformation" refers to Agrobacterium mediated DNA transfer, in which the flower is brought in contact with the Agrobacterium solution. Floral dip transformation has been described in Arabidopsis (Clough and Bent, 1998) and Brassica spp. (Verma et al., 2008; Tan et al., 2011). "T-DNA" refers to the transferred-region of the Ti (tumor-inducing) plasmid of Agrobacterium tumefaciens. Ti plasmids are natural gene transfer systems for the introduction of heterologous nucleic acids into the nucleus of higher plants. Binary Agrobacterium vectors such pBIN20 and pPZP222 (GenBank Accession Number U10463.1) are known in the art. Proteins delivered to plant cells using the Agrobacterium-mediated method will generally be localized to the nucleus of the plant cell. Additional modifications to the protein could result in the targeting to other specific subcellular locations within the plant cell. For example, "signal" or "transit" peptides are capable of targeting a polypeptide to the chloroplast, as known in the art. It is recognized that in the Agrobacterium-mediated method of protein delivery to plant cells, the signal peptide would be added to the polypeptide of interest in the bacterium. The DNA construct of the present invention may be introduced into the genome of a desired plant host by a suitable Agrobacterium-mediated plant transformation method. Several Agrobacterium species mediate the transfer of T-DNA, described above, that can be genetically engineered to carry any desired piece of DNA into many plant species. Recently, non- Agrobacterium systems are coming to the fore for plant transformation, specifically Ensifer adhaerens, Ochrobactrum haywardense and Rhizobium etli (Rathore DS, Mullins E , 2018). Other methods for introducing DNA into cells may also be used. These methods are well known to those of skill in the art and can include chemical methods and physical methods such as microinjection, electroporation, and micro-projectile. Transformed plant cells that are derived by any of the above transformation techniques can be cultured to regenerate a whole plant that possesses the transformed genotype and thus the desired phenotype. Such regeneration techniques rely on manipulation of certain phytohormones in a tissue culture growth medium, typically relying on a biocide and / or herbicide marker that has been introduced together with the desired nucleotide sequences. Plant regeneration from cultured protoplasts is described in Evans et al., Protoplasts Isolation and Culture, Handbook of Plant Cell Culture, pp.124-176, MacMillilan Publishing Company, New York, 1983; and Binding, Regeneration of Plants, Plant Protoplasts, pp.21-73, CRC Press, Boca Raton, 1985. Regeneration can also be obtained from plant callus, explants, organs, or parts thereof. Such regeneration techniques are described generally in Klee et al., Ann. Rev. of Plant Phys.38:467- 486 (1987). Plant cell regeneration techniques rely on manipulation of certain phytohormones in a tissue culture growth medium, also typically relying on a biocide and / or herbicide marker that has been introduced together with the desired nucleotide sequences. Choice of methodology with suitable protocols being available for hosts from Leguminosae (alfalfa, soybean, clover, etc.), Umbelliferae (carrot, celery, parsnip), Cruciferae (cabbage, radish, canola / rapeseed, etc.), Cucurbitaceae (melons and cucumber), Gramineae (wheat, barley, rice, maize, etc.), Solanaceae (potato, tobacco, tomato, peppers), various floral crops, such as sunflower, and nut-bearing trees, such as almonds, cashews, walnuts, and pecans. Such regeneration techniques are known to those of skill in the art. Methods and compositions for transforming plants by introducing a transgenic DNA construct into a plant genome in the practice of this invention can include any of the well-known and demonstrated methods. Plant transformants containing a desired genetic modification as a result of any of the above described methods and resulting in the expression of the drought-associated SNPs of the invention can be selected by various methods known in the art. These methods include, but are not limited to, methods such as SDS-PAGE analysis, immunoblotting using antibodies which bind to the genetic modification, single nucleotide polymorphism (SNP) analysis, or assaying for the products of a reporter or marker gene, and the like. As used herein, the term "plant" includes reference to whole plants, plant organs (e.g., leaves, stems, roots, etc.), seeds and plant cells and progeny of same. Plant cell, as used herein includes, without limitation, plant cells within or isolated from seeds, suspension cultures, embryos, meristematic regions, callus tissue, leaves, roots, shoots, gametophytes, sporophytes, pollen, and microspores. The plants disclosed herein include any plant species. Exemplary plant species include, without limitation, soybean (Glycine max), corn (Zea mays), Brassica sp. (e.g., B. napus, B. rapa, B. juncea), particularly those Brassica species useful as sources of seed oil, alfalfa (Medicago sativa), rice (Oryza sativa), rye (Secale cereale), sorghum (Sorghum bicolor, Sorghum vulgare), millet (e.g., pearl millet (Pennisetum glaucum), proso millet (Panicum miliaceum), foxtail millet (Setaria italica), finger millet (Eleusine coracana)), sunflower (Helianthus annuus), safflower (Carthamus tinctorius), wheat (Triticum aestivum), tobacco (Nicotiana tabacum), potato (Solanum tuberosum), peanuts (Arachis hypogaea), cotton (Gossypium barbadense, Gossypium hirsutum), sweet potato (Ipomoea batatus), cassava (Manihot esculenta), coffee (Coffea spp.), coconut (Cocos nucifera), pineapple (Ananas comosus), citrus trees (Citrus spp.), cocoa (Theobroma cacao), tea (Camellia sinensis), banana (Musa spp.), avocado (Persea americana), fig (Ficus casica), guava (Psidium guajava), mango (Mangifera indica), olive (Olea europaea), papaya (Carica papaya), cashew (Anacardium occidentale), macadamia (Macadamia integrifolia), almond (Prunus amygdalus), sugar beets (Beta vulgaris), sugarcane (Saccharum spp.), oats, barley, vegetables, ornamentals, and conifers. Vegetables include soybean (Glycine max), tomatoes (Lycopersicon esculentum), lettuce (e.g., Lactuca sativa), green beans (Phaseolus vulgaris), lima beans (Phaseolus limensis), peas (Lathyrus spp.), and members of the genus Cucumis such as cucumber (C. sativus), cantaloupe (C. cantalupensis), and musk melon (C. melo). Ornamentals include azalea (Rhododendron spp.), hydrangea (Macrophylla hydrangea), hibiscus (Hibiscus rosasanensis), roses (Rosa spp.), tulips (Tulipa spp.), daffodils (Narcissus spp.), petunias (Petunia hybrida), carnation (Dianthus caryophyllus), poinsettia (Euphorbia pulcherrima), and chrysanthemum. Conifers that may be employed in practicing the present invention include, for example, pines such as loblolly pine (Pinus taeda), slash pine (Pinus elliotii), ponderosa pine (Pinus ponderosa), lodgepole pine (Pinus contorta), and Monterey pine (Pin us radiata); Douglas-fir (Pseudotsuga menziesii); Western hemlock (Tsuga canadensis); Sitka spruce (Picea glauca); redwood (Sequoia sempervirens); true firs such as silver fir (Abies amabilis) and balsam fir (Abies balsamea); and cedars such as Western red cedar (Thuja plicata) and Alaska yellow-cedar (Chamaecyparis nootkatensis). Preferably, plants of the present invention are crop plants (for example, corn, alfalfa, sunflower, Brassica, soybean, cotton, safflower, peanut, sorghum, wheat, millet, tobacco, etc.)In certain embodiments the plant refers to a strain of arabidopsis or soybean. In preferred embodiments, the plant refers to a strain of cotton. In certain embodiments, the strain of cotton is selected from Gossypium hirsutum, Gossypium barbadense, Gossypium arboretum, or Gossypium raimondii. In preferred embodiments, the cotton strain is Gossypium hirsutum. A "clone" or "clonal cell population" is a population of cells derived from a single cell or common ancestor by mitosis. A "cell line" is a clone of a primary cell or cell population that is capable of stable growth in vitro for many generations. Methods of transmitting drought-associated SNP through Breeding One of skill will recognize that the naturally-occurring drought-associated SNP or a recombinant drought-associated SNP, it can be introduced into other plants by sexual crossing. Any of a number of standard breeding techniques can be used, depending upon the species to be crossed. As used herein, the terms “cross”, “crossing”, or “cross breeding” refer to and / or include the process by which the pollen of one flower on one plant is applied (artificially or naturally) to the ovule (stigma) of a flower on another plant. It will be appreciated that in breeding the plants of the present invention can also be in with other plants of the same species(i.e., self-breeding or cross breeding) such as with cultivated or wild plants, novel hybrid plants or plant lines which exhibit at least some of the characteristics of the plants of the present invention will be generated. Plants resultant from crossing any of these with another plant can be utilized in pedigree breeding, transformation and / or backcrossing to generate additional cultivars which exhibit the characteristics of the genomically multiplied plants of the present invention and any other desired traits. Screening techniques employing molecular or biochemical procedures well known in the art, including those discussed at length above, can be used to ensure that the important commercial characteristics sought after are preserved in each breeding generation. The term “backcrossing” refers to a process in which a breeder repeatedly crosses hybrid progeny, for example a first generation hybrid (Fi), back to one of the parents of the hybrid progeny. Backcrossing can be used to introduce one or more single locus conversions from one genetic background into another. The goal of backcrossing is to alter or substitute a single trait or characteristic in a recurrent parental line. To accomplish this, a single gene of the recurrent parental line is substituted or supplemented with the desired gene from the nonrecurrent line, while retaining essentially all of the rest of the desired genes, and therefore the desired physiological and morphological constitution of the original line. Although backcrossing methods are simplified when the characteristic being transferred is a dominant allele, a recessive allele may also be transferred. In this instance, it may be necessary to introduce a test of the progeny to determine if the desired characteristic has been successfully transferred. Likewise, transgenes can be introduced into the plant using any of a variety of established transformation methods well- known to persons skilled in the art, such as described above. The term “self-crossing” refers to a process, such as by cloning by self-pollination or clipping, which may reduce the amount of heterozygosity within the genetics. For example, experiments have shown that doubly inbred plants, exhibit less genetic variation as compared to non-inbred plants. If self-crossed plants are crossed such that the resulting plants segregate into distinct phenotypes, it indicates that the self-crossed parent plants were still heterozygous at the relevant loci. Kits and articles of manufacture Any of the aforementioned products can be incorporated into a kit which may contain a drought-associated SNP specific marker polynucleotide or one or more such markers immobilized on a Gene Chip, an oligonucleotide, a polypeptide, a peptide, an antibody, a detectable label, marker, reporter, a pharmaceutically acceptable carrier, a physiologically acceptable carrier, instructions for use, a container, a vessel for administration, an assay substrate, or any combination thereof. Immobilization on a solid support refers to methods for linking the nucleic acid molecules to the support such that they cannot be stripped from the support via washing. The materials and methods below are provided to facilitate the practice of the present invention. Plant materials and experimental design Plant growth conditions have been described in Melandri et al. (2021). These same cotton accessions were examined for their metabolite profiles in response to drought stress over a two- year field experiment. In brief, a panel of 22 upland cotton (Gossypium hirsutum L.) accessions (Figure 17) were grown at the University of Arizona Maricopa Agricultural Center (MAC) in Maricopa, AZ, United States during the summer of 2019. The accessions were arranged in a randomized incomplete block design with half of the plots experiencing normal irrigation (well- watered, WW) and the other half treated with water-limited (WL) condition, starting when 50% of the plots were at first flower, and consisting of approximately half of the WW irrigation amount. Leaf tissue was harvested ~50 days later at the boll development stage and used for metabolomics (details in Melandri et al., 2021) and transcriptomics analyses. For RNA extraction, two biological replicates, comprised of 10 leaf discs (0.64 cm in diameter) of the upper-most expanded leaf derived from five randomly selected plants (two leaf discs from each plant) in each plot were collected within a single day during a time window of 3 hours (11:00-14:00). Leaf discs were stored in a 2 mL Eppendorf tube containing 1.5 mL of RNAlater (Fisher #AM7021) and stored on ice before being transferred from field to lab where they were then stored in a –80°C freezer before further processing steps. Metabolite and phenotypic data measurement All details on the procedures used to collect / determine the phenotypic trait data and metabolite data used in this study can be found in Melandri et al., (2021). In brief, cotton fiber quantity and quality traits included lint yield (grams / plot), micronaire (Mic, units of air permeability), upper half mean length (UHM, inches), length uniformity (UI, percent), strength (Str, grams per tex) and elongation (Elo, percent). Reflectance-based vegetation indices (VIs) included normalized difference vegetation index (NDVI), carotenoid reflectance index (CRI), scaled photochemical reflectance index (sPRI), and the ratio between the water index (WI) and NDVI (WI / NDVI). The levels of 451 metabolites were determined by untargeted gas chromatography-mass spectrometry (GC-MS, 27 metabolites) and liquid chromatography-mass spectrometry (LC-MS, 424 metabolites). Best Linear Unbiased Estimators (BLUEs) of each phenotypic trait and metabolite value were generated for each cotton accession before being used for downstream statistical analyses. RNA-seq library construction Extraction of RNA from leaf disks was performed using methods from previous studies (Pang et al., 2011; Wu et al., 2002). Hot borate buffer was prepared containing 0.2 M sodium borate pH 9.0, 30 mM EGTA (ethylene glycol-bis(β-aminoethyl ether)-N,N,N′,N′-tetraacetic acid), 1% SDS (sodium lauryl sulfate), and 1% sodium deoxycholate. Just before use, Polyvinylpyrrolidone (PVP-40), NP-40, and Dithiothreitol (DTT) were added to the hot borate buffer to a final concentration of 2%, 1%, and 10 mM, respectively. To extract RNA, a total of approximately 0.75 ml hot borate buffer per 50 mg tissue (~5 cotton leaf disks) was used. Buffer was heated to 80 °C, and 250 hot buffer was added to 5 leaf disks. Tissue was ground in the hot buffer in a mortar and pestle, then 15 μl of 20 mg / ml Proteinase K (Roche #46950800 / Sigma #3115887001) per sample was added and the tissue was ground again. A final 500 ul hot buffer was added before grinding a final time. Lysate was added to a Qiashredder column (Qiagen #79656) and centrifuged at 13,000 x g for 1 minute. Flow through was added to 0.5 volumes of 100% ethanol. The mixture was used as input for RNA cleanup using the RNeasy kit (Qiagen #74104) according to the manufacturer’s instructions. Samples were eluted with RNase-free water. RNA-seq libraries were then generated from mRNA-enriched samples using the Amaryllis Nucleics Full Transcript library prep kit (YourSeq Duet, available on the world wide web at: / / amaryllisnucleics.com / kits / duet-rnaseq-library-kit). Transcriptome sequencing and data processing Each of the 22 G. hirsutum accessions treated under both the WW or WL conditions were sequenced using Illumina HiSeqTM 2500 (San Diego, CA) paired-end libraries. Trimmomatic was used to trim adapters and low-quality reads (Bolger et al., 2014). Further, reads were aligned to the G. hirsutum reference genome (Ghirsutum_527_v2.0, accession ID: VKGJ01000000, acquired from Phytozome (Chen et al., 2020)), using RMTA v2.6.3 pipeline (Peri et al., 2020) with default parameters and further quantified reads mapped back to each locus using FeatureCounts (parameter: Multimapping; reads: counted; Multi-overlapping reads: counted) and then transformed into length normalized Transcripts Per Kilobase Million (TPM) by custom R scripts (Liao et al., 2014). The raw feature counts for each transcript across 88 samples were normalized by DESeq2 (Love et al., 2014). To test the reproducibility of replicates, we used the Pearson correlation of normalized counts between each set of replicates. Replicates with R2< 0.8 and P > 0.05 between samples were removed, leading to only 21 accessions being further examined. Further, pair-wise comparison of TPM values under WL and WW conditions for each accession was performed to identify respective accession-specific DEGs (model: ~entry + treatment + entry : treatment). For each accession, significantly DEGs under WL condition were classified using thresholds adjusted p-value < 0.01, Log2FC > 1 (upregulated) or < 1 (downregulated). We examined the average expression of all genes from A and D subgenomes with expression > 1 TPM to determine whether there is a shift in global subgenome gene expression between WW and WL conditions. This analysis tests the hypothesis that subgenomes may be adapted to different environments such that their expression dominance may be affected by spatiotemporal contexts. Variants calling and phylogenetic analysis Variant calling was conducted with the GATK4 pipeline using the haplotypecaller function (Brouard et al., 2019). The average mapping depth (DP) was calculated across 84 samples as screening cutoff which is equal to 4.12. Then we filtered out variants with sites with DP < 3, depth by quality (DQ) < 2, genotype quality (GQ) < 20, and minor allele frequency < 0.025 by vcftools (Danecek et al., 2011). Phylogenetic relationships were inferred using the IQ- TREE pipeline with filtered SNPs (Nguyen et al., 2015). Briefly, variants calling files (VCFs) of all samples were transformed into phylip format by vcf2phylip (Ortiz, E.M.2019) tools, a maximum-likelihood tree was constructed (parameter: -nt AUTO -m MFP) and plotted by Figtree (available on the world wide web at: tree.bio.ed.ac.uk / software / figtree / ). The 21 accessions were subdivided into different groups based on phylogenetic, geographic, and historical breeding information obtained from USDA-GRIN (available on the world wide web at: ars-grin.gov). A publicly available deep-sequenced (~20×) WGS dataset containing 17 out of 21 accessions used for field experiment was obtained from the NCBI SRA (Figure 17) to identify genome-wide variants. The reads were trimmed by trimonmatic (default settings(Bolger et al., 2014)) and mapped to the reference genome using bwa-mem along with retention of unique mapped reads by picard (MarkDuplicates -remove = TRUE) for variants calling. The haplotypecaller calling was performed by GATK4 and filtered by VCFtools (--max-missing 0.9, --maf 0.05, minGQ 20, minQ 200, minDP 5). The filtered variants were annotated using VEP (default settings) to classify variations to coding regions, promoter, and inter-genic regions. Principal component analysis (PCA) To estimate the strength of treatment effects and the impact of phylogenetic relatedness on transcriptomic and metabolomic profiles, a principal component analysis (PCA) was performed using normalized read counts of 3,768 genes associated with the top 10% of transcriptomic variance or using plotPCA function of DESeq2 (Love et al., 2014), and normalized Z-score of 451 metabolites using “factoextra” R package (available on the world wide web at: cran.r-project.org / web / packages / factoextra). Contribution rates for each component were calculated using plotPCA. The first three components from transcriptome and metabolites were visualized in a three-dimensional PCA using Cubemaker (available on the world wide web at: tools.altiusinstitute.org / cubemaker / ). Weighted gene co-expression network analysis Gene co-expression network was constructed using the WGCNA R package to classify gene expression modules and explore module-trait relationships (Langfelder and Horvath, 2008). Genes with both a high median absolute deviation (MAD) score (top 10%) and genes with high expression (average TPM across all samples > 1) were retained for expression module classification. The pickSoftThreshold function was used to identify the optimal soft power threshold (= 7) at which R2surpassed 0.85 and no further improvements in mean connectivity (module size) were observed (See e.g., Zhu et al., 2018). Block-wise modules were constructed using the following parameters (power = 7, maxBlockSize = 5000, TOMType = "unsigned", minModuleSize = 30, reassignThreshold = 0, mergeCutHeight = 0.2). Trait-module relationships were derived using the “modTraitCor” function in WGCNA. Modules that displayed a high trait correlation (R2> 0.6, p < 0.05) were selected for further analysis. To investigate modules with high module-trait membership, only co-expressed genes with strong connectivity (weight score > 0.1) were retained for downstream analysis. Functional characterization of trait-correlated module Gene functional annotations of filtered genes (weight score > 0.1) in lint yield-correlated module were obtained from the cotton functional genome database (CottonFGD) to perform enrichment of gene ontology (GO) and KEGG pathways (available on the world wide web at cottonfgd.org / ; Zhu et al., 2017) using Fisher Exact test (P < 0.01). Transcription factors (TFs), transcription regulators (TRs), and protein kinases within this module were classified using iTAK (Zheng et al., 2016). Reported protein-protein interactions (PPIs) among genes within the same module were screened using the STRING database (Szklarczyk et al., 2017) (confidence level = 0.6, interaction source: database, experiments). To identify the upstream master TF(s) that may bind to genes within the lint yield module, the enrichment of consensus TF motifs was tested using the 2 Kbp upstream region of the 866 co-expressed genes within the module. The enrichment test was performed by the Analysis of Motif Enrichment (AME) pipeline (McLeay and Bailey, 2010) with the Arabidopsis DAP-seq profiles of (Malley et al., 2016) as the consensus motif database (parameters: --scoring avg --method fisher --hit-lo-fraction 0.25 --e- value --kmer report-threshold 10.0, cutoff: TP values > 3, P-value < 0.001). DNA affinity purification sequencing (DAP-Seq) library construction The four TFs were selected as hub genes to construct DAP-seq libraries, including two GhDREB2A (Gohir.A13G021700, Gohir.D13G022300) and two GhHSFA6B (Gohir.D08G072600, Gohir.A08G064100). DAP-Seq assay was carried out as described by published protocol (Bartlett et al., 2017). The NEB Next® DNA Library Prep Master Mix set for Illumina kit (NEB #E6040S) was implemented to prepare the DAP-Seq gDNA library. The pIX- HALO Vector (cat#G184A, Promega) was used to fuse the GhDREB2A and GhHSFA6B into the HaloTag. We further used the TNT SP6 High-Yield Wheat Germ Protein Expression System (L3260, Promega) to express the four TFs-HaloTag fusion protein. Magens HaloTag Beads (G7281, Promega) was used to purify the fusion protein. The fusion protein and 500 ng of library DNA were co-incubated in 40 μl PBS buffer for 1.5 hours shaking in a cold room. The beads were washed with 200 μl PBS + NP40 (0.005%) for 5 times. The supernatant was discarded and an aliquot of 25 μl of elution buffer was added. Finally, beads were incubated at 98°C for 10 min to elute DNA fragments. According to the fragment size of the library, the DAP-Seq library concentration for a given read count was measured. DAP-Seq sequencing and peak analysis DAP-seq libraries were sequenced by Illumina short reads platform (Single end: 150 bp) with a total of expected 40 million reads for each protein to be sequenced. These reads were trimmed by TrimGalore (parameter: --phred33) and then mapped to the reference genome (TM- 1: Ghirsutum_527_v2.0, Phytozome accession ID: VKGJ01000000) using bowtie2 and sorted by sambamba. Before peak calling, only the unique mapped reads were retained by sambamba (parameter: -F "[XS] == null and not unmapped and not duplicate") to prevent false positives due to multiple mapped hits. Furthermore, the MAC2s tool was applied to call peaks (parameter: -- keep-dupall -g 2.3-e9) and high confidence peaks were captured using IDR (Padj < 0.05). These high-quality peaks were annotated based on reference gene annotation using the ChIPseeker R package. Regulatory regions of a target gene were defined as the 5 Kb upstream sequences before the transcription start sites (TSSs) and the downstream sequences which bearing the longest 5’UTRs among respective isoforms. In particular, only DAP-seq peaks which fell into the regulatory regions were kept as TF targeted genes. The sequences associated with DAP-seq peaks of TFs were harvested as query sequences to perform motif discovery by MEME suite, to identify conserved consensus motif sequences among peak sequences (parameter: -mod zoops - nmotifs 3 -minw 6 -maxw 50 -objfun classic -revcomp -markov_order 0). Genomic analysis of TF binding regions Paralogous genes of the selected lint module genes (GhABP and GhIPS) were identified using Blastn (Ye et al., 2006), which aligns the CDS sequences against the reference gene annotation (cutoff: E-value < 10-3, identities > 90%). Those hits been identified were further checked by functional annotation. Further, the promoter and coding regions of these paralogous gene pairs were extracted and then aligned using multiple sequences alignment by Geneious (tool: MUSCLE, iteration: 500) to identify variants (SNPs and INDELs) between paralogous gene pairs, specifically the regions bearing the DAP-seq peak. The potential binding sites of two TFs over the peak regions were defined by motif scan (FIMO) using Arabidopsis DAP-seq derived consensus motif sequences (Motif ID: ERF48_col_a, HSFA6B_cal_a, HSFA6B_colamp_a). In-vitro protein expression and electrophoretic mobility shift assay (EMSA) Protein expression was carried out using the in vitro transcription and translation system (TnT™ T7 Quick for PCR DNA, Promega) and the EMSA reaction was carried out following manufacturer’s instruction (odyssey). The vector containing the gene of interest (with flag tag) was PCR amplified to obtain the DNA template. The reaction mixture contained the DNA template, reaction buffer, amino acid mix, and T7 RNA polymerase, following the manufacturer's instructions. The reaction was incubated at the recommended temperature for a specified duration, and the resulting protein product was subsequently analyzed by SDS-PAGE to confirm expression. Immunoblot analyses were performed by using standard wet transfer method. The proteins were blotted using nitrocellulose membrane (Sigma Aldrich). Rabbit monoclonal halo tag antibody was used in the (1:5000) dilution (v / v). Horse-radish peroxidase conjugated goat anti-rabbit antibodies were used as secondary antibodies. Blots were developed using Super-signal Chemiluminescent substrate following the manufacturer’s instruction. The 5 blots were imaged using the BioRAD gel imager following the linear curve of detection in the signal. The interaction between the expressed protein and its putative DNA-binding partner was analyzed using EMSA. A double-stranded DNA probe containing the target binding site was prepared by annealing complementary oligonucleotides labeled with a fluorophore. The binding 0 reaction mixture consisted of the expressed protein, labeled DNA probe, binding buffer, and appropriate cofactors. The reaction mixture was incubated at room temperature for 30 minutes to allow for protein-DNA complex formation. Subsequently, the samples were resolved on a non- denaturing polyacrylamide gel electrophoresis (PAGE) gel with a 6% acrylamide concentration. The gel was run at 120 volts for 3 hours in 1× TBE. Following electrophoresis, the gel was 5 visualized using LiCor Odyssey Gel Scanner to capture the mobility shifts indicating protein- DNA complex formation. For competition assays, unlabeled competitor DNA or unlabeled mutated competitor DNA was added to the binding reaction at increasing molar excesses. Probes and competitor sequences used in EMSA are provided below. Scale Name Sequence (5'-3') (SEQ ID NO:) Label (nmol) Purification Note e1 e1 e2 e2 e1 e1 Probe-2-F- TTTTGGGGTATTCTTGAAGCTTTCTGGAAAA Standard Control TTCAAGCTCGATTCATGT (11) NA 100 Desalting ABP-HSF-site2 Probe-2-R- ACATGAATCGAGCTTGAATTTTCCAGAAAG Standard e2 e1 e1 e2 e2 The following examples are provided to illustrate certain embodiments of the invention. They are not intended to limit the invention in any way. Example I: Genetic Profiling of Upland Cotton Panels 5 To better understand the molecular mechanisms underlying cotton’s performance under drought conditions, 22 diverse upland cotton accessions from the Gossypium Diversity Reference Set were evaluated under water-limited (WL) and well-watered (WW) conditions in the field. These accessions were selected for reported variation in heat and drought tolerance as well as fiber qualities (Figure 17). Leaf tissue from plants at the flowering and boll development stage was collected (two replicates each, ~ 100 DAS; Figure 1A). On average, ~27 million PE x 150 bp reads were generated for each replicate. Reads were then aligned using the RMTA pipeline (Peri et al., 2020) to the G. hirsutum v2.0 reference genome (Chen et al., 2020); (Figure 17). Gene level expression data were counted using G. hirsutum gene-level annotations as meta- features in FeatureCounts (Liao et al., 2014;). The Pearson correlation coefficient (PCC) of normalized read counts revealed high correlation among replicates (average PCC = 0.904), except for accession “Tipo Chaco” (PCC = 0.61, cutoff: 0.8) (Table 1). Thus, this accession was discarded from further analyses resulting in 21 accessions for further analyses. TABLE 1: Pearson Correlation Coefficient (PCC) of replicates for each sample 3 5 3 1 8 1 6 1 8 1 1 1 3 9 8 9 8 7 7 2 9 3 7 WL TAM_94WE_37S 0.86 WL VIR_7065_ALLEN_333 0.92 2 8 5 3 6 4 2 2 7 4 2 2 3 9 9 4 1 3 9 7 6 To determine the relatedness of each of these accessions, these RNA-seq data were used to perform variant calling relative to the reference genome (108,396 bi-allelic SNPs after filtering). These data were used to reconstruct a phylogeny of the 21 accessions, rooted by VIR_7153_D_10, which resulted in three major subgroups (Figure 1B). Based on information from USDA GRIN-Global, the first of these three subgroups is referred to as “Foreign”, as this group contains accessions developed primarily outside of the U.S. (Figure 1B and Figure 17). The second subgroup consists of “U.S.” accessions of breeding lines and commercial cultivars developed in the U.S. (Figure 1B and Figure 17). The last group, “Mixed”, consists of commercial cultivars from the U.S. and improved breeding lines generally developed outside the U.S. (Figure 1B and Figure 17). Thus, these 21 accessions serve as a genetically diverse framework to investigate the regulatory mechanisms controlling the physiological responses to drought in cotton. Example 2: Analysis of Variations in Physiological and Molecular Responses to Drought in Major Cotton Subgroups The phylogenetic groupings were used to determine the degree to which genotype and environment influenced physiological responses to drought in the assembled panel. The effects of genotype (G: Foreign, U.S., and Mixed), environment (E: WL and WW), and G*E interaction on six fiber quality and agronomic traits and four vegetation indices, were examined using two- way ANOVA (Table 2). Among these traits, three of the four vegetation indices were significantly affected by drought (Table 2; two-way ANOVA cut-off: p < 0.05, Figure 2). TABLE 2: Two-way ANOVA test of 10 traits TraitsPhylogenyEnvironment (P) (E)G*ELint yield 0.0286 * 0.061 0.889 Micronaire (Mic) 0.0335 * 0.0020 ** 0.855 Upper-half mean length (UHM) 0.00978 ** 0.768 0.817 Length uniformity (UI) 0.069 0.834 0.957 Strength (Str) 0.0155 * 0.161 0.822 Elongation (Elo) 0.308 0.769 0.985 NDVI 0.752 2.66e - 6 *** 0.298 sPRI 0.068 0.121 0.977 CRI 0.953 1.71e-11 *** 0.679 WI / NDVI 0.854 4.6e-05 *** 0.311 The treatment effect was particularly significant in the “Foreign” subgroup (student t-test cutoff: p < 0.05). By contrast, long-term drought stress brought limited impacts to most fiber quality traits apart from the significantly reduced micronaire (two-way ANOVA: p = 0.002) observed in “Mixed” and “Foreign” subgroups (t-test: p = 0.015 and 0.047, respectively, Figure 2, Table 2). In contrast to treatment effects, pronounced genotypic effects were observed between major subgroups (“Foreign, U.S. and Mixed) for five of the six fiber quality traits but not for the four vegetation indices (Table 2). However, intragroup comparisons of responses to drought revealed stark differences. For instance, each subgroup contained one or more accessions that significantly outperformed close relatives when using a ratio of phenotypes from WL relative to WW condition as an indicator, particularly in fiber quality traits (Figure 1C). The lack of phylogenetic congruence in observed traits indicates that lint yield and drought associated traits have been selected differently in even closely related accessions. Importantly, the presence of outperforming accessions indicates that this panel is useful in identifying the factors associated with lint yield and drought stress. We analyzed metabolite profiling as molecular phenotypes for stress response to better understand the effects of drought stress. Changes in the metabolome of the panel were measured 5 by re-examining 451 metabolites previously analyzed in Melandri et al., 2021 (27 GC-MS and 424 LC-MS / MS).). A global profiling of the 451 metabolites revealed that the strongest effects were due to drought treatment (principle component (PC) 1, 62%; Figure 1D). The second and third principal components explained a further 5.0% and 3.4%, respectively, but the phylogenetic relationships (i.e., genotype) were not clearly separated along either of these axes. A large0 treatment effect was also observed in the metabolite profiles using a two-way ANOVA (Table 3). Drought treatment significantly impacted more than 95% of metabolites, but only 5% of the metabolites exhibited genotypic effects; around 25% of metabolites exhibited changes that could be associated with G*E effects (p < 0.05, Table 3). Among the 432 metabolites associated with significant treatment effects, 273 were up-regulated and 159 down-regulated by drought. Among5 these metabolites are the known glycine and proline metabolites (Fang et al., 2015); Table 3). Thus, these data demonstrate that alteration of specific metabolites associated with drought tolerance (osmoprotectants) is a common response to water deficit stress conditions across our panel. TABLE 3: Two-way ANOVA test of 451 metabolites MetabolitesName Main Class Genotype Treatment GxEGCMS_1Sucrose Carbohydrates & conjugates 0.709197 0.016374 0.9654GCMS_2Citric acid Other 0.149809 0.000394 0.473574GCMS_4Allose Carbohydrates & conjugates 0.056824 0.000254 0.16182GCMS_5Galactose Carbohydrates & conjugates 0.127502 0.004976 0.178845GCMS_7Saccharic acid Carbohydrates & conjugates 0.314507 3.80E-06 0.77589GCMS_8Glutamic acid Amino acids & peptides 0.672043 2.95E-06 0.985794GCMS_9Epigallocatechin Other 0.476557 0.000408 0.769388GCMS_10Glyceric acid Carbohydrates & conjugates 0.025439 3.06E-06 0.806191GCMS_11Glycine Amino acids & peptides 0.344745 4.43E-10 0.904961GCMS_ Tryptamine, 5- 14hydroxy-Other 0.194888 0.01058 0.265354 GCMS_15Glucose Carbohydrates & conjugates 0.447909 0.068571 0.571096GCMS_16Phosphoric acid Other 0.006541 1.36E-08 0.167221GCMS_ Butanoic acid, 4- 17amino-Amino acids & peptides 0.389599 2.95E-08 0.302172GCMS_19Xylobiose, D- Carbohydrates & conjugates 0.014697 0.052411 0.810341GCMS_ Glutaric acid, 2- 20hydroxy-Other 0.420396 1.77E-07 0.861262GCMS_21Aspartic acid Amino acids & peptides 0.544208 1.88E-08 0.437268GCMS_22Threonine Amino acids & peptides 0.005093 2.43E-11 0.057174GCMS_ Glucose, 1,6- 23anhydro-, beta-Other 0.677815 1.83E-08 0.886711GCMS_24Serine Amino acids & peptides 0.221658 2.30E-06 0.484869GCMS_ Quinic acid, 3- 25caffeoyl-, cis-Other 0.043843 6.06E-08 0.076596GCMS_26Pyroglutamic acid Amino acids & peptides 0.040093 0.006879 0.654474GCMS_27Malic acid, 2-methyl- Other 0.614961 3.40E-06 0.345808GCMS_28Arabinose Carbohydrates & conjugates 0.014694 0.001255 0.931263GCMS_29Diisopropanolamine Other 0.199198 1.43E-14 0.382999GCMS_ 2- 30 Piperidinecarboxylic Amino acids & peptides 0.782839 1.47E-05 0.905711 acid GCMS_31Diethanolamine Other 0.46321 0.000156 0.649162GCMS_ Norbornane-2- 32 carboxylic acid, 2- Terpenoids 0.286399 0.130853 0.471647 amino- LCMS_1bietatriene Terpenoids 0.674542 8.54E-16 0.00541,2-Di-(9Z,12Z,15Z- LCMS_ octadecatrienoyl)- 2 3-(Galactosyl-alpha- Polar lipids 0.659394 1.11E-16 0.403904 1-6-Galactosyl-beta- 1)-glycerol LCMS_3TG(15:0 / 15:0 / 15:0) Neutral lipids 0.39937 1.48E-18 0.053144LCMS_ PG(18:2(9Z,12Z) / 18: 40)Polar lipids 0.788146 7.93E-15 0.110387LCMS_5N1-Acetylspermine Other 0.898887 3.54E-17 0.0175LCMS_6Arachisprenol 12 Neutral lipids 0.058481 2.33E-17 0.014167LCMS_7Arachisprenol 11 Neutral lipids 0.061308 1.24E-17 0.089362LCMS_ DG(20:0 / 20:1(11Z) / 0 8:0)Neutral lipids 0.188561 1.24E-17 0.037417 LCMS_9Oxytocin 1-8 Amino acids & peptides 0.192194 3.68E-13 0.272601Isoscoparin 2''-(6- LCMS_ (E)- 10 feruloylglucoside) 4'- Other 0.653028 8.80E-11 0.027876 glucoside LCMS_11Betulin Terpenoids 0.499464 2.54E-15 0.130106LCMS_12beta-Phellandrene Terpenoids 0.439272 5.18E-17 0.0407744alpha-Formyl- LCMS_ 4beta-methyl- 13 5alpha-cholesta-8- Terpenoids 0.15013 1.68E-15 0.057698 en-3beta-ol LCMS_ DG(14:1(9Z) / 17:2(9Z 14,12Z) / 0:0)[iso2]Neutral lipids 0.611431 2.63E-15 0.016619LCMS_16PG(18:0 / 18:0) Polar lipids 0.760398 8.44E-13 0.1315222-methoxy-3-methyl- LCMS_ 6-all-trans- 17 nonaprenylhydroquin Other 0.548604 9.16E-15 0.03222 one LCMS_18PA(20:1(11Z) / 15:0) Polar lipids 0.704978 2.63E-09 0.542848LCMS_ TG(16:0 / 16:1(9Z) / 18 19:0)Neutral lipids 0.835356 1.03E-14 0.095774LCMS_20Goyaglycoside c Other 0.63508 3.92E-15 0.004606LCMS_21PG(16:0 / 16:0) Polar lipids 0.91605 3.27E-15 0.007286LCMS_22Uric acid Other 0.082622 9.26E-10 0.094658LCMS_23Arachisprenol 10 Neutral lipids 0.320607 8.73E-24 0.050414LCMS_24Plastoquinone Other 0.285955 1.52E-14 0.000177LCMS_ Medroxyprogesteron 25eOther 0.751357 8.42E-11 0.012784LCMS_26PC(20:0 / 18:1(11Z)) Polar lipids 0.197007 6.57E-09 0.002555LCMS_ TG(15:0 / 18:1(9Z) / 20 27:4(5Z,8Z,11Z,14Z))Neutral lipids 0.270957 2.99E-14 0.003084LCMS_ PA(22:2(13Z,16Z) / 2 280:2(11Z,14Z))Polar lipids 0.482052 5.64E-12 0.019729LCMS_ 1-18:2-2-18:3- 29 digalactosyldiacylgly Polar lipids 0.992819 0.019469 0.69941 cerol LCMS_30Cornusiin E Other 0.181063 1.36E-09 0.01477LCMS_ TG(15:0 / 15:0 / 15:1(9 31Z))[iso3]Neutral lipids 0.336162 1.16E-14 0.035123LCMS_ 1-a,24R,25- 32 Trihydroxyvitamin Other 0.591789 1.85E-13 0.196674 D2 LCMS_ TG(12:0 / 14:1(9Z) / 19 33:0)[iso6]Neutral lipids 0.291472 2.19E-12 0.077581 LCMS_ PA(20:3(8Z,11Z,14Z 34) / 17:2(9Z,12Z))Polar lipids 0.99803 6.07E-05 0.706524LCMS_ PI(22:6(4Z,7Z,10Z,1 35 3Z,16Z,19Z) / 22:1(11 Polar lipids 0.492282 6.73E-09 0.100393 Z)) LCMS_ PI(P- 36 16:0 / 18:3(9Z,12Z,15 Polar lipids 0.155026 0.018592 0.010856 Z)) LCMS_ 1-18:3-2-18:2- 37 monogalactosyldiacy Polar lipids 0.988779 0.00105 0.304152 lglycerol LCMS_38Avocadyne Other 0.436877 8.29E-12 0.062404LCMS_ bacteriohopanetetrol 40cyclitol etherOther 0.94233 2.07E-13 0.095894LCMS_ PA(20:0 / 22:2(13Z,16 41Z))Polar lipids 0.77063 8.15E-14 0.048022LCMS_43Triricinolein Neutral lipids 0.700563 5.27E-13 0.835914LCMS_ TG(15:0 / 18:2(9Z,12 44Z) / 22:0)Neutral lipids 0.683164 4.81E-13 0.037867LCMS_ PE(P- 4520:0 / 22:2(13Z,16Z))Polar lipids 0.063001 7.96E-17 0.390119LCMS_46Glycinoprenol 11 Neutral lipids 0.814714 1.90E-20 0.081649LCMS_ DG(18:1(9Z) / 18:2(9Z 47,12Z) / 0:0)Neutral lipids 0.187802 1.70E-10 0.017506LCMS_ PI(20:3(8Z,11Z,14Z) / 4822:1(11Z))Polar lipids 0.368618 2.09E-09 0.002551-O-(alpha-D- LCMS_ galactopyranosyl)-N- 49 hexacosanoylsphing Polar lipids 0.388766 2.34E-10 0.079925 anine LCMS_ PE(15:0 / 20:3(8Z,11Z 50,14Z))Polar lipids 0.221644 3.92E-12 0.00296LCMS_ Type IV cyanolipid 5120:0 esterOther 0.435823 5.87E-18 0.058542LCMS_ PC(18:2(9Z,12Z) / 18: 522(9Z,12Z))Polar lipids 0.916391 3.97E-13 0.297109LCMS_ PC(16:0 / 18:3(9Z,12 53Z,15Z))Polar lipids 0.226754 0.023123 0.694922LCMS_ 5a-Androst-3-en-17- 54oneOther 0.913835 2.18E-13 0.056908LCMS_55Neoxanthin Terpenoids 0.145295 1.42E-08 0.895921LCMS_ TG(18:1(9Z) / 18:1(9Z 56) / 18:2(9Z,12Z))Neutral lipids 0.532543 2.49E-05 0.542281LCMS_ PE(18:0 / 20:2(11Z,14 57Z))Polar lipids 0.954336 8.37E-12 0.303962LCMS_ PC(16:0 / 18:2(9Z,12 58Z))Polar lipids 0.922617 6.39E-10 0.334933LCMS_ Tetrahydrodeoxycort 59icosteroneOther 0.914155 0.000757 0.285072LCMS_ DG(18:4(6Z,9Z,12Z, 60 15Z) / 18:3(9Z,12Z,15 Neutral lipids 0.888785 7.39E-06 0.621708 Z) / 0:0) LCMS_ Glycerol 1,2-di-(9Z- 61 octadecenoate) 3- Neutral lipids 0.980911 2.05E-13 0.080655 tetradecanoate LCMS_ TG(18:2(9Z,12Z) / 18: 62 2(9Z,12Z) / 18:3(6Z,9 Neutral lipids 0.663771 4.77E-05 0.66575 Z,12Z)) LCMS_ TG(16:0 / 15:0 / 14:1(9 63Z))Neutral lipids 0.489859 4.59E-14 0.026144LCMS_ TG(18:2(9Z,12Z) / 18: 64 2(9Z,12Z) / 18:2(9Z,1 Neutral lipids 0.413067 0.000288 0.729692 2Z)) LCMS_66PC(16:0 / 18:1(9Z)) Polar lipids 0.623175 2.56E-12 0.010066LCMS_ 1-18:3-2-18:3- 67 monogalactosyldiacy Polar lipids 0.609939 3.48E-10 0.505445 lglycerol LCMS_68Pregeijerene Other 0.025923 8.35E-14 0.040556LCMS_69SM(d18:1 / 25:0) Polar lipids 0.595093 1.83E-11 0.168622LCMS_ DG(18:3(6Z,9Z,12Z) / 7018:2(9Z,12Z) / 0:0)Neutral lipids 0.557449 5.77E-05 0.05215LCMS_71Pheophytin-a Other 0.813139 0.118184 0.102728LCMS_ alpha-Tocopherol 72succinateOther 0.974022 6.48E-10 0.102415LCMS_73PG(16:0 / 18:0) Polar lipids 0.412872 0.0767 0.067989LCMS_ 1-18:2-2-18:2- 74 digalactosyldiacylgly Polar lipids 0.445699 4.74E-10 0.914073 cerol LCMS_ PC(P- 7518:1(11Z) / 16:1(9Z))Polar lipids 0.584808 2.04E-14 0.030752LCMS_ 18-Nor- 76 4(19),8,11,13- Terpenoids 0.892848 1.97E-16 0.04386 abietatetraene LCMS_77PG(O-20:0 / 16:0) Polar lipids 0.778147 1.43E-12 0.003204LCMS_ DG(18:2(9Z,12Z) / 18: 782(9Z,12Z) / 0:0)Neutral lipids 0.287232 4.90E-11 0.015621LCMS_ DG(20:5(5Z,8Z,11Z, 79 14Z,17Z) / 20:5(5Z,8Z Neutral lipids 0.374632 1.17E-07 0.715732 ,11Z,14Z,17Z) / 0:0) LCMS_ PA(18:3(6Z,9Z,12Z) / 8022:0)Polar lipids 0.816676 8.11E-05 0.523984LCMS_ (20R)-24- 81 Hydroxygeminivitami Other 0.816996 2.12E-10 0.827526 n D3 LCMS_ TG(18:2(9Z,12Z) / 18: 82 3(6Z,9Z,12Z) / 18:3(6 Neutral lipids 0.75222 1.27E-15 0.281235 Z,9Z,12Z)) LCMS_ TG(18:1(9Z) / 18:2(9Z 83,12Z) / 18:2(9Z,12Z))Neutral lipids 0.364424 1.77E-08 0.192097LCMS_ Glucosylceramide 84(d18:1 / 22:0)Polar lipids 0.200607 6.27E-15 0.735604 LCMS_85SM(d18:2 / 20:1) Other 0.010333 5.22E-15 0.001654LCMS_ PE(20:1(11Z) / 14:1(9 86Z))Polar lipids 0.989619 1.53E-11 0.050911Glycerol 2-(9Z,12Z- octadecadienoate) 1-hexadecanoate 3- LCMS_ O-[alpha-D- 87 galactopyranosyl-(1- Polar lipids 0.782905 2.90E-06 0.857072 >6) -beta-D- galactopyranoside] LCMS_ 3-Hydroxy-cis-5- 88 tetradecenoylcarnitin Other 0.938779 2.16E-09 0.073658 e LCMS_89D8'-Merulinic acid A Other 0.533852 1.00E-08 0.237721LCMS_90Solanthrene Other 0.352728 3.35E-13 0.057578LCMS_ TG(16:0 / 16:1(9Z) / 20 91:4(5Z,8Z,11Z,14Z))Neutral lipids 0.91794 7.31E-07 0.366473LCMS_ TG(13:0 / 14:0 / 18:1(9 92Z))[iso6]Neutral lipids 0.851271 4.91E-12 0.265282LCMS_93menaquinol-12 Neutral lipids 0.439019 2.64E-14 0.033555LCMS_94DG(24:0 / 24:0 / 0:0) Neutral lipids 0.035398 7.87E-22 0.000853LCMS_ DG(16:0 / 18:2(9Z,12 95Z) / 0:0)Neutral lipids 0.39216 6.02E-15 0.0539291-tetradecanoyl-2- LCMS_ hexadecanoyl- 96 sn-glycero-3- Other 0.64443 3.51E-15 0.022816 phosphosulfocholine LCMS_97Tricetin Other 0.791381 7.80E-09 0.758606DG(20:5(5Z,8Z,11Z, LCMS_ 14Z,17Z) / 98 22:6(4Z,7Z,10Z,13Z, Neutral lipids 0.278832 4.67E-15 0.097244 16Z,19Z) / 0:0) LCMS_99MG(0:0 / 16:0 / 0:0) Other 0.399572 8.14E-22 0.521112LCMS_ PA(15:0 / 20:3(8Z,11Z 100,14Z))Polar lipids 0.567896 0.00168 0.931856LCMS_101PI(21:0 / 17:1(9Z)) Polar lipids 0.040215 0.014509 0.91150325-Acetyl-6,7- LCMS_ didehydrofevicordin 102 F 3-[glucosyl-(1->6)- Other 0.56936 5.46E-06 0.086415 glucoside] LCMS_104Inugalactolipid A Polar lipids 0.973896 7.18E-12 0.082772LCMS_ PG(22:6(4Z,7Z,10Z, 105 13Z,16Z,19Z) / 17:1(9 Polar lipids 0.977266 1.76E-05 0.626192 Z)) LCMS_ LysoPE(18:1(9Z) / 0:0 107)Polar lipids 0.563401 1.71E-06 0.194492 LCMS_ PA(13:0 / 20:3(8Z,11Z 108,14Z))Polar lipids 0.990339 4.88E-12 0.046268LCMS_ TG(16:1(9Z) / 18:1(9Z 109Neutral lipids 0.865289 0.009692 0.302026LCMS_110Terpenoids 0.875047 2.49E-16 0.007336LCMS_111) Polar lipids 0.285722 1.13E-12 0.094237LCMS_ 112PC(20:0 / 20:1(11Z)) Polar lipids 0.839882 2.15E-09 0.367691LCMS_ PI(20:2(11Z,14Z) / 18: 1132(9Z,12Z))Polar lipids 0.564004 1.16E-10 0.98957LCMS_ PC(18:1(9Z) / 18:1(9Z 114))Polar lipids 0.437665 2.04E-12 0.009723LCMS_ PC(18:1(9Z) / 18:2(9Z 115,12Z))Polar lipids 0.720158 5.98E-14 0.023581LCMS_ 1-18:1-2-16:0- 116 digalactosyldiacylgly Polar lipids 0.256814 4.60E-06 0.084266 cerol LCMS_ PA(20:0 / 20:4(5Z,8Z, 11711Z,14Z))Polar lipids 0.864661 3.83E-10 0.112911LCMS_ PC(18:2(9Z,12Z) / 18: 1183(9Z,12Z,15Z))Polar lipids 0.450674 2.40E-10 0.373941LCMS_ PC(o- 119 16:1(9Z) / 18:2(9Z,12 Polar lipids 0.358592 5.92E-13 0.035265 Z)) LCMS_120Secoisotetrandrine Other 0.656638 5.78E-08 0.658276LCMS_ (3E,7E)-4,8,12- 121 Trimethyl-1,3,7,11- Terpenoids 0.823401 8.64E-18 0.092963 tridecatetraene LCMS_122PA(22:1(11Z) / 19:0) Polar lipids 0.068566 3.28E-16 0.021187LCMS_123Armillarin Terpenoids 0.364744 4.83E-12 0.0390721-hexadecanoyl-2- [(5S,6E,8Z,11Z,14Z) LCMS_ -5 126 - Polar lipids 0.980641 2.87E-09 0.211389 hydroxyicosatetraen oyl]-sn-glycero-3- phosphocholine LCMS_ 5-Hydroxy-6- 127 methoxycoumarin 7- Other 0.929241 1.46E-07 0.00703 glucoside LCMS_ PE(18:0 / 18:2(9Z,12Z 128))Polar lipids 0.032135 7.09E-08 0.243158LCMS_129PI(21:0 / 18:1(9Z)) Polar lipids 0.743315 1.17E-13 0.027848LCMS_ N-palmitoyl 130glutamineAmino acids & peptides 0.91588 6.84E-12 0.209062LCMS_ PE(22:2(13Z,16Z) / 2 1310:3(8Z,11Z,14Z))Polar lipids 0.87211 2.71E-12 0.050167LCMS_132Capsorubin Terpenoids 0.192828 3.16E-09 0.968642 LCMS_ TG(15:0 / 20:1(11Z) / 2 1330:2n6)Neutral lipids 0.872719 1.45E-09 0.007382TG(14:0 / 18:2(9Z,12 LCMS_ Z) / 22:6 134 (4Z,7Z,10Z,13Z,16Z, Neutral lipids 0.868276 0.054414 0.776609 19Z)) LCMS_135PI-Cer(d18:1 / 22:0) Polar lipids 0.266866 5.34E-20 0.071164LCMS_ TG(18:2(9Z,12Z) / 19: 136 0 / Neutral lipids 0.072048 2.27E-15 0.034865 20:2(11Z,14Z))[iso6] LCMS_1375alpha-androstane Other 0.058372 2.25E-11 0.215377LCMS_ PA(20:2(11Z,14Z) / 2 1382:1(11Z))Polar lipids 0.139331 1.72E-08 0.016651TG(12:0 / 22:4(7Z,10 LCMS_ Z,13Z,16Z) / 139 22:6(4Z,7Z,10Z,13Z, Neutral lipids 0.242203 5.49E-10 0.075215 16Z,19Z))[iso6] LCMS_140Vinaginsenoside R3 Terpenoids 0.104508 5.85E-07 0.094546LCMS_141Capsidiol Terpenoids 0.699654 0.00463 0.985767LCMS_142PE(20:0 / 22:1(13Z)) Polar lipids 0.250763 0.052202 0.316164LCMS_ PA(22:2(13Z,16Z) / 1 1437:2(9Z,12Z))Polar lipids 0.721548 4.95E-07 0.472373LCMS_ PE(18:3(6Z,9Z,12Z) / 14418:2(9Z,12Z))Polar lipids 0.825225 7.31E-09 0.264431-(O-alpha-D- LCMS_ glucopyranosyl)- 145 (1,3R,25R)- Other 0.519492 4.59E-07 0.147637 hexacosanetriol LCMS_ 13'-Hydroxy-alpha- 146tocopherolOther 0.46816 0.004882 0.068461LCMS_147O-Methylsomniferine Other 0.415332 3.22E-09 0.491195LCMS_ PI(19:1(9Z) / 20:3(8Z, 14811Z,14Z))Polar lipids 0.724656 2.32E-15 0.24104DGDG(18:4(6Z,9Z,1 LCMS_ 2Z,15Z) / 150 18:4(6Z,9Z,12Z,15Z) Polar lipids 0.03403 2.11E-08 0.209501 ) LCMS_151SM(d18:0 / 12:0) Polar lipids 0.82246 2.30E-11 0.003083LCMS_152Kaempferol Other 0.361196 0.008219 0.619241LCMS_153PG(16:1(9Z) / 16:0) Polar lipids 0.17257 5.83E-15 0.247696LCMS_ Medicarpin 3-O-(6'- 154malonylglucoside)Other 0.406921 0.602065 0.4957152-methoxy-6-all- LCMS_ trans- 155 nonaprenylhydroquin Other 0.417351 2.19E-11 0.139466 one LCMS_156PA(21:0 / 17:1(9Z)) Polar lipids 0.032365 0.032956 0.650406LCMS_ PG(18:0 / 20:2(11Z,1 1574Z))Polar lipids 0.635156 8.63E-13 0.530528Galalpha1- 3(Fucalpha1- LCMS_ 2)Galbeta1-3 158 (Fucalpha1- Polar lipids 0.571659 8.77E-10 0.002129 4)GlcNAcbeta1-3 Galbeta1-4Glcbeta- Cer(d18:1 / 16:0) LCMS_ LysoPE(0:0 / 18:1(11 159Z))Polar lipids 0.602 4.17E-08 0.1342233alpha,6alpha,12alp LCMS_ ha- 160 Trihydroxy-7-oxo- Other 0.457633 1.71E-10 0.064898 5beta-cholan- 24-oic Acid LCMS_161Dieporeticenin Other 0.438454 0.001422 0.670307LCMS_ OH- 162 Diaponeurosporene Terpenoids 0.910684 0.00079 0.247066 glucoside ester LCMS_ PI(16:2(9Z,12Z) / 16:0 163)Polar lipids 0.208791 2.02E-06 0.251946LCMS_16418:2 Sitosteryl ester Terpenoids 0.518514 0.002196 0.663028TG(16:1(9Z) / 16:1(9Z LCMS_ ) / 165 20:4(5Z,8Z,11Z,14Z) Neutral lipids 0.346696 1.86E-05 0.258342 ) LCMS_ TG(18:0 / 22:0 / 22:6 166 (4Z,7Z,10Z,13Z,16Z, Neutral lipids 0.248994 8.49E-13 0.091457 19Z)) LCMS_167DG(21:0 / 22:0 / 0:0) Neutral lipids 0.120948 4.07E-16 0.294269LCMS_ PA(17:1(9Z) / 18:2(9Z 169,12Z))Polar lipids 0.5398 3.48E-07 0.2216042-methoxy-6-all LCMS_ trans-decaprenyl-2- 170 methoxy- Other 0.889017 1.01E-14 0.009875 1,4-benzoquinol LCMS_171beta1-Chaconine Other 0.943511 8.76E-09 0.206589LCMS_173Pubesenolide Other 0.540696 2.37E-14 0.218267LCMS_174Zeranol Other 0.751329 1.17E-07 0.130292LCMS_ PG(18:3(9Z,12Z,15Z 175) / 16:1(9Z))Polar lipids 0.538818 6.96E-14 0.20182LCMS_176PA(21:0 / 21:0) Polar lipids 0.665952 6.17E-10 0.36498LCMS_ PC(18:1(9Z) / 18:3(9Z 177,12Z,15Z))Polar lipids 0.949375 2.99E-09 0.16018LCMS_ TG(16:0 / 16:0 / 18:1(9 178Z))Neutral lipids 0.427637 2.54E-13 0.108382 LCMS_ 3-demethylubiquinol- 1809Neutral lipids 0.436226 1.28E-07 0.000146LCMS_1811,2-dihydrostilbene Other 0.22044 9.09E-08 0.356669LCMS_ Zeaxanthin 182diglucosideTerpenoids 0.881468 4.57E-11 0.048652LCMS_183Capsicoside C3 Other 0.604453 8.79E-09 0.994968DG(22:6(4Z,7Z,10Z, LCMS_ 13Z,16Z,19Z) / 184 22:6(4Z,7Z,10Z,13Z, Neutral lipids 0.503069 0.003694 0.975002 16Z,19Z) / 0:0) LCMS_185PE(18:0 / 18:1(9Z)) Polar lipids 0.135183 0.002018 0.030445LCMS_186Neochlorogenin Terpenoids 0.33709 3.09E-05 0.467515LCMS_ PE(18:3(9Z,12Z,15Z 187) / 16:0)Polar lipids 0.79896 0.000647 0.148619LCMS_ 188 Polar lipids 0.978611 0.000241 0.523403 LCMS_190Other 0.049532 2.15E-10 0.254062LCMS_ 191PC(20:0 / 16:1(9Z)) Polar lipids 0.479975 1.27E-09 0.004373LCMS_ DG(16:1(9Z) / 18:2(9Z 192,12Z) / 0:0)Neutral lipids 0.625861 5.68E-14 0.040091LCMS_193PC(18:1(9Z) / 24:0) Polar lipids 0.934859 4.51E-06 0.290314LCMS_195PA(22:0 / 19:0) Polar lipids 0.0093 9.19E-11 0.002156LCMS_ MG(18:2(9Z,12Z) / 0: 1960 / 0:0)Neutral lipids 0.340166 4.01E-07 0.120077LCMS_198Glycyrrhetinic acid Terpenoids 0.227059 2.49E-05 0.543734LCMS_199Glabrin C Amino acids & peptides 0.84808 0.000233 0.527556LCMS_200Dichotellate A Other 0.454958 6.12E-12 0.008026LCMS_201Conicasterol B Other 0.807437 1.76E-12 0.066761LCMS_ PG(17:0 / 18:2(9Z,12 202Z))Polar lipids 0.670527 0.000484 0.717014LCMS_203Azithromycin Carbohydrates & conjugates 0.549169 0.000139 0.377264LCMS_ 1- 204 Arachidonoylglycero Polar lipids 0.921361 0.029592 0.260972 phosphoinositol 5-Hydroxy-4- LCMS_ methoxy- 205 3-methyl-2,6- Other 0.933026 2.67E-07 0.151307 canthinedione LCMS_ PI(22:6(4Z,7Z,10Z,1 2073Z,16Z,19Z) / 21:0)Polar lipids 0.648942 1.21E-06 0.366911LCMS_208MG(18:0 / 0:0 / 0:0) Other 0.915002 3.35E-11 0.166432 LCMS_ PI(O- 20916:0 / 17:2(9Z,12Z))Polar lipids 0.824202 2.23E-05 0.50593LCMS_ PG(18:2(9Z,12Z) / 16: 2100)Polar lipids 0.225402 0.005036 0.202403LCMS_ PI(19:0 / 22:6(4Z,7Z,1 2110Z,13Z,16Z,19Z))Polar lipids 0.490178 4.29E-07 0.0164344'-O- LCMS_ Methylneobavaisofla 212 vone Other 0.681062 2.10E-08 0.369628 7-O-(2''-p- coumaroylglucoside) LCMS_ beta-Sitosterol 3-O- 213 beta- Other 0.970326 4.91E-09 0.007516 D-galactopyranoside LCMS_ Isolimonic acid 214glucosideOther 0.531188 0.000316 0.801971LCMS_215Scillirosidin Other 0.878981 1.28E-14 0.004553LCMS_216NA Other 0.656653 1.48E-11 0.094851LCMS_ PA(22:6 217 (4Z,7Z,10Z,13Z,16Z, Polar lipids 0.919419 6.54E-08 0.48906 19Z) / 17:0) LCMS_218ent-16-Kaurene Terpenoids 0.303242 3.06E-16 0.010031LCMS_ 18:2 Stigmasteryl 219esterTerpenoids 0.760647 4.12E-07 0.645551LCMS_ DG(19:1(9Z) / 22:0 / 0: 2200)[iso2]Neutral lipids 0.420119 7.26E-12 0.190906LCMS_ Cer(d16:2(4E,6E) / 20 221:1(11Z)(2OH))Other 0.964958 5.21E-10 0.961891LCMS_222Cohibin A Other 0.19185 2.70E-09 0.201296LCMS_223Lactucaxanthin Terpenoids 0.11169 3.39E-20 0.255918LCMS_224PC(22:0 / 18:1(9Z)) Polar lipids 0.998493 8.32E-13 0.03014LCMS_ DG(18:4(6Z,9Z,12Z, 22515Z) / 16:0 / 0:0)Neutral lipids 0.839789 3.39E-10 0.032675LCMS_226PC(o-18:1(9Z) / 22:0) Polar lipids 0.671307 2.14E-09 0.293641LCMS_227PA(22:0 / 20:1(11Z)) Polar lipids 0.499737 1.59E-09 0.330283LCMS_228Phytoene Terpenoids 0.559448 2.72E-18 0.001035LCMS_ TG(14:0 / 14:1(9Z) / 20 229Neutral lipids 0.957241 1.76E-14 0.215652LCMS_ 230Polar lipids 0.870842 4.86E-05 0.019039LCMS_ 232 Neutral lipids 0.24323 0.007724 0.007458 LCMS_ 233 Neutral lipids 0.593958 2.76E-12 0.19889 LCMS_234alpha-Cryptoxanthin Terpenoids 0.320985 9.03E-15 0.886042LCMS_235Nonanoyl-CoA Other 0.738897 7.32E-06 0.007882LCMS_ TG(16:0 / 16:1(9Z) / 23618:2(9Z,12Z))Neutral lipids 0.530681 2.02E-11 0.057813LCMS_237propanoyl phosphate Other 0.187981 9.75E-06 0.974382LCMS_238PGP(18:1(11Z) / 18:0) Polar lipids 0.804689 1.11E-14 0.180521LCMS_ Galactosylceramide 240(d18:1 / 14:0)Polar lipids 0.134486 0.002897 0.059966LCMS_241Apo-10'-violaxanthal Terpenoids 0.096055 2.23E-18 0.201123LCMS_ 2-(3'- 243 Methylthio)propylmal Other 0.488928 1.17E-09 0.383975 ic acid LCMS_ 4,5-epoxy-17R- 244HDHAOther 0.978986 2.34E-07 0.356296LCMS_245[7]-Gingerol Other 0.875253 0.228021 0.193894LCMS_246Octyl phenylacetate Other 0.976034 1.38E-12 0.016872LCMS_ TG(14:1(9Z) / 15:0 / 16 247:1(9Z))Neutral lipids 0.305413 9.31E-08 0.161078LCMS_ 248Polar lipids 0.457335 3.34E-13 0.027215LCMS_ 250Neutral lipids 0.686809 0.014565 0.025937LCMS_251Polar lipids 0.284359 2.53E-05 0.110237LCMS_ 252Polar lipids 0.578786 0.000126 0.088447LCMS_253Polar lipids 0.233647 9.45E-09 0.012252LCMS_255 Other 0.787054 4.28E-16 0.553345LCMS_256PE(15:0 / 20:1(11Z)) Polar lipids 0.250862 0.000186 0.656681LCMS_257Artobiloxanthone Other 0.766437 6.25E-08 0.165113LCMS_ PI(20:0 / 18:2(9Z,12Z) 258)Polar lipids 0.001397 1.30E-10 0.093736LCMS_ Stigmasteryl 259glucosideOther 0.544907 0.001138 0.738928LCMS_ sn-caldarchaeo-1- 261 phosphoethanolamin Other 0.292951 7.31E-12 0.525586 e LCMS_ Campesterol 262glucosideOther 0.311573 3.67E-10 0.864014LCMS_ Uridine diphosphate- 263 N- Other 0.507292 2.30E-08 0.924648 acetylglucosamine LCMS_264Pipericine Other 0.988373 0.055565 0.126448 LCMS_265phenolic phthiocerol Other 0.6599 8.03E-15 0.011957LCMS_ PI(20:3(8Z,11Z,14Z) / 26618:2(9Z,12Z))Polar lipids 0.327902 4.00E-07 0.00311LCMS_ 1-O-palmitoyl- 267Cer(d18:1 / 16:0)Other 0.408518 4.18E-09 0.000863LCMS_ 7,8- 268 Diaminopelargonic Other 0.272737 2.04E-14 0.120085 acid LCMS_269Montecristin Other 0.569843 7.10E-08 0.72607LCMS_270PC(21:0 / 20:1(11Z)) Polar lipids 0.711463 6.45E-07 0.161697LCMS_271Megalomicin B Carbohydrates & conjugates 0.618523 2.95E-06 0.492962LCMS_ PA(20:5(5Z,8Z,11Z, 27214Z,17Z) / 22:0)Polar lipids 0.967528 1.76E-08 0.106303LCMS_273Majoroside F1 Terpenoids 0.869356 2.08E-13 0.842942LCMS_275PC(21:0 / 22:0) Polar lipids 0.994023 1.23E-05 0.040657LCMS_276PE(22:0 / 24:1(15Z)) Polar lipids 0.988333 0.39247 0.141404LCMS_ PI(17:0 / 22:4(7Z,10Z, 27713Z,16Z))Polar lipids 0.852294 1.53E-10 0.408236LCMS_ DG(21:0 / 22:4 278 (7Z,10Z,13Z,16Z) / 0: Neutral lipids 0.118695 9.53E-18 0.065348 0)[iso2] LCMS_279PE(20:0 / 20:1(11Z)) Polar lipids 0.286307 1.35E-11 0.156238LCMS_281protochlorophyll a Other 0.327059 5.38E-13 0.024205LCMS_282PC(P-18:0 / 22:0) Polar lipids 0.884797 3.05E-12 0.071187LCMS_ hexadecasphinganin 283e(1+)Other 0.59704 4.50E-13 0.0066692-(3,4- dihydroxyphenyl)- 3,5- LCMS_ dihydroxy-7-{[3,4,5- 284 trihydroxy-6- Other 0.852899 1.34E-09 0.599957 (hydroxymethyl)oxan -2-yl]oxy}-4H- chromen-4-one LCMS_285Sphinganine Other 0.854405 8.15E-12 0.061752LCMS_ PG(P- 286 16:0 / 20:4(5Z,8Z,11Z Polar lipids 0.67882 4.07E-06 0.434777 ,14Z)) LCMS_287Oleamide Other 0.511968 0.002557 0.063544LCMS_ 2- 288ArachidonylglycerolOther 0.227239 1.94E-09 0.046311LCMS_ 5,6-Epoxyoctadeca- 2907,9-diynoic acidOther 0.612901 7.16E-12 0.281682 LCMS_ 2-Oxo-4- 291 methylthiobutanoic Other 0.90248 3.48E-10 0.077083 acid LCMS_292Amino acids & peptides 0.482583 1.28E-12 0.018586LCMS_ PC(14:0 / 18:2(9Z,12 293Z))Polar lipids 0.701906 2.27E-09 0.408475LCMS_ DG(15:0 / 20:3(8Z,11 294Z,14Z) / 0:0)Neutral lipids 0.370044 2.38E-07 0.351053LCMS_ PA(15:0 / 22:6(4Z,7Z, 29510Z,13Z,16Z,19Z))Polar lipids 0.773734 0.000151 0.700757LCMS_ PS(O- 297 20:0 / 22:6(4Z,7Z,10Z Polar lipids 0.325413 2.57E-09 0.083344 ,13Z,16Z,19Z)) campest-22E-en- LCMS_ 3beta, 298 4beta,5alpha,6alpha, Other 0.922802 0.019464 0.338303 8beta,14alpha, 15alpha,25,28-nonol LCMS_299Panaxydol linoleate Other 0.913431 1.37E-09 0.058913LCMS_300Calomelanol D-1 Neutral lipids 0.481294 0.001769 0.07176LCMS_ 3- 302HydroxyechinenoneTerpenoids 0.918901 0.147432 0.064414LCMS_ N-Arachidonoyl 303GABAAmino acids & peptides 0.880976 1.48E-11 0.234406LCMS_ PI(20:5(5Z,8Z,11Z,1 3044Z,17Z) / 21:0)Polar lipids 0.125994 0.032508 0.066153(3beta,22R,23R,24S LCMS_ )-3,22,23- 305 Trihydroxystigmasta Other 0.244306 0.001678 0.031974 n-6-one LCMS_ Protoporphyrinogen 306IXOther 0.392729 0.000243 0.029873LCMS_ MGDG(18:3(9Z,12Z, 307 15Z) / 18:4(6Z,9Z,12Z Polar lipids 0.809917 6.61E-08 0.513263 ,15Z)) LCMS_ 3-Hydroxy-8'-apo- 308epsilon-caroten-8'-alTerpenoids 0.628283 1.40E-11 0.708052(3S,3'S,5R,5'R,6R)- 3,6-Epoxy-5,6- LCMS_ dihydro-3',5,8'- 309 trihydroxy- Terpenoids 0.15314 0.00445 0.393825 beta,kappa-caroten- 6'-one LCMS_ PGF2alpha methyl 311etherOther 0.16586 4.99E-08 0.066409LCMS_ Dehydroisocopropor 312phyrinogenOther 0.816889 3.67E-09 0.664538LCMS_ PI(P- 313 16:0 / 18:4(6Z,9Z,12Z Polar lipids 0.680343 3.28E-08 0.561429 ,15Z)) LCMS_315Munchiwarin Other 0.048734 0.152359 0.319275 LCMS_316PI(14:1(9Z) / 19:0) Polar lipids 0.939601 7.34E-06 0.315982LCMS_ 1-18:2-2-18:3- 317 monogalactosyldiacy Polar lipids 0.511295 2.41E-15 0.898211 lglycerol LCMS_318Kiwiionoside Other 0.996201 0.005286 0.332366LCMS_ PE(18:4(6Z,9Z,12Z, 319 15Z) / 20:3(8Z,11Z,14 Polar lipids 0.459557 6.40E-05 0.235216 Z)) LCMS_ 4-O- 320MethylmelleolideTerpenoids 0.714315 6.13E-12 0.0329LCMS_ PI(P- 322 18:0 / 18:4(6Z,9Z,12Z Polar lipids 0.614178 1.41E-09 0.493999 ,15Z)) LCMS_323DG(17:0 / 22:0 / 0:0) Neutral lipids 0.804456 1.60E-15 0.019922LCMS_324Mosinone A Other 0.422776 0.134288 0.186238LCMS_325menaquinol-11 Neutral lipids 0.462313 1.60E-07 0.071362LCMS_326Pheophytin-b Other 0.046938 8.92E-12 0.00128LCMS_327PA(19:0 / 20:0) Polar lipids 0.28854 0.000112 0.098374LCMS_ TG(18:1(9Z) / 18:1(9Z 328) / 18:3(6Z,9Z,12Z))Neutral lipids 0.154169 3.16E-07 0.01173LCMS_329PG(16:0 / 22:1(11Z)) Polar lipids 0.183833 7.65E-13 0.268299LCMS_ Vinaginsenoside 330R16Terpenoids 0.760181 4.85E-14 0.0804544-[N-(p- LCMS_ Coumaroyl)serotonin 331 -4''-yl] Other 0.99422 4.30E-14 0.171454 -N-feruloylserotonin LCMS_ PC(o- 332 22:0 / 18:3(9Z,12Z,15 Polar lipids 0.983722 2.82E-12 0.47933 Z)) LCMS_ DG(15:0 / 22:5 333 (7Z,10Z,13Z,16Z,19 Neutral lipids 0.799074 9.76E-08 0.420732 Z) / 0:0) LCMS_334Cohibin C Other 0.975524 1.28E-11 0.198234LCMS_335Proline betaine Amino acids & peptides 0.93384 3.71E-14 0.023448N-(24- LCMS_ hydroxytetracosanoy 336 l) Other 0.859981 0.15656 0.188425 sphinganine LCMS_337DDM-838 Other 0.720113 0.000924 0.610142LCMS_ Gomphrenol 3- 338 methylether 4'- Other 0.222181 0.002243 0.20834 glucuronide LCMS_ Protopanaxadiol 3- 339glucoside 20-Terpenoids 0.458608 1.40E-12 0.366417 [arabinosyl-(1->2)- glucoside] LCMS_340PE(22:0 / 20:0) Polar lipids 0.633422 3.25E-13 0.036354LCMS_ 1- 341 deoxytetradecasphin Other 0.312163 5.47E-08 0.066221 ganine LCMS_342Saxitoxin Other 0.192592 2.43E-07 0.030444LCMS_343Luteolin 7-xyloside Other 0.500541 3.89E-05 0.453432LCMS_344ancymidol Other 0.548893 4.30E-05 0.290307LCMS_ De-O- 345 methylsterigmatocys Other 0.942583 5.46E-06 0.279537 tin LCMS_346Butenylcarnitine Other 0.584022 3.66E-11 0.611758LCMS_ Tetracosapentaenoic 347acid (24:5n-3)Neutral lipids 0.52333 1.26E-10 0.007031LCMS_ Hexadecanedioic 349acidOther 0.970126 1.25E-05 0.201801LCMS_350Atenolol Other 0.450644 0.008269 0.8462132-Hexaprenyl-3- LCMS_ methyl-6- 351 methoxy-1,4 Other 0.58527 6.90E-06 0.924276 benzoquinone LCMS_352Farfugin A Other 0.125396 8.49E-07 0.342199LCMS_ 355Polar lipids 0.035683 0.000156 0.037227LCMS_ 356Polar lipids 0.886219 7.59E-06 0.199788LCMS_ 357Polar lipids 0.953753 1.56E-06 0.717958LCMS_358Neutral lipids 0.649333 4.27E-08 0.800018LCMS_359 Other 0.427097 6.52E-09 0.000123LCMS_ (all-E)-1,7,9- 362 Heptadecatriene- Other 0.935737 0.001885 0.279707 11,13,15-triyne LCMS_363Isoannonacin A Other 0.634934 1.85E-08 0.471539LCMS_36513-Hydroxylupanine Other 0.700459 2.27E-07 0.015621LCMS_366Muzanzagenin Other 0.770861 1.33E-12 0.100507LCMS_ 3-Methyl-5-pentyl-2- 367furannonanoic acidOther 0.310557 5.78E-08 0.021639LCMS_ 11'-Carboxy-alpha- 368chromanolTerpenoids 0.213782 1.18E-07 0.205378LCMS_372NA Other 0.351514 0.001286 0.098409 LCMS_374Dioctyl hexanedioate Other 0.944626 1.90E-05 0.362284LCMS_375N-Nitrosotomatidine Other 0.40498 3.12E-07 0.008468LCMS_376Neolinderatin Other 0.373281 0.007021 0.046027LCMS_ 11b,17a,21- 377 Trihydroxypreg Other 0.978774 1.77E-05 0.70515 -nenolone LCMS_378Armillaripin Terpenoids 0.93627 1.97E-11 0.021989LCMS_379Clausarinol Other 0.608484 2.94E-09 0.103816LCMS_ GlcCer(d16:2 381(4E,6E) / 24:0(2OH))Other 0.376382 0.007774 0.574716LCMS_382PC(P-20:0 / 19:1(9Z)) Polar lipids 0.093199 1.42E-05 0.398514LCMS_383alpha-Ionene Other 0.588866 1.52E-16 0.002332LCMS_ (6E,8E)-4,6,8- 384MegastigmatrieneOther 0.947244 7.94E-12 0.094976(2R)-2-[(1R)-1- hydroxy-17- {(1R,2R)- LCMS_ 2-[(2R)-22-methyl- 385 21-oxotetracontan-2- Neutral lipids 0.836242 5.59E-07 0.002902 yl] cyclopropyl}heptade cyl]hexacosanoic acid LCMS_ TG(16:1(9Z) / 18:1(9Z 386 ) / 20:4(5Z,8Z,11Z,14 Neutral lipids 0.294346 2.40E-13 0.078093 Z)) LCMS_ TG(16:1(9Z) / 18:0 / 18 387:2(9Z,12Z))Neutral lipids 0.029564 6.09E-05 0.743623LCMS_ PGP(18:0 / 22:4(7Z,1 3880Z,13Z,16Z))Polar lipids 0.458751 0.09948 0.193427LCMS_ PA(22:6(4Z,7Z,10Z, 389 13Z,16Z,19Z) / 17:1(9 Polar lipids 0.961646 0.000259 0.455171 Z)) LCMS_ PS(22:6(4Z,7Z,10Z, 390 13Z,16Z,19Z) / Polar lipids 0.884598 6.34E-08 0.855468 20:3(8Z,11Z,14Z)) LCMS_391Bacteriorubixanthin Terpenoids 0.383165 7.70E-13 0.5671LCMS_ PS(20:3(8Z,11Z,14Z 392 ) / 22:4 Polar lipids 0.040333 0.002465 0.577132 (7Z,10Z,13Z,16Z)) LCMS_393Tetrahydropteridine Other 0.746919 9.18E-15 0.019901LCMS_394Floionolic acid Neutral lipids 0.842545 0.015445 0.813848LCMS_ 2-Decaprenyl-6- 395methoxyphenolNeutral lipids 0.051213 0.000707 0.217743LCMS_ PS(22:2(13Z,16Z) / 2 3972:6Polar lipids 0.456482 0.003356 0.45411 (4Z,7Z,10Z,13Z,16Z, 19Z)) LCMS_ PI(P-16:0 / 22:6 398 (4Z,7Z,10Z,13Z,16Z, Polar lipids 0.937358 2.83E-06 0.815643 19Z)) LCMS_399olivetol Other 0.467285 2.15E-11 0.198283LCMS_400Lucidenic acid A Terpenoids 0.387039 1.59E-14 0.847024LCMS_401Dalbergioidin Other 0.971447 8.54E-06 0.22977LCMS_402Herniarin Other 0.271888 1.30E-12 0.020198LCMS_ N-(11Z,14Z)- 404 eicosadienoylethanol Other 0.912372 0.020987 0.027126 amine LCMS_ N- 405 gondoylethanolamin Other 0.696575 0.000267 0.042722 e LCMS_ 9'-Carboxy-gamma- 406chromanolOther 0.608028 0.000316 0.028559LCMS_ PI- 407 Cer(t18:0 / 18:0(2OH) Polar lipids 0.966636 0.000268 0.73223 ) LCMS_ 16- 408 Dehydroprogesteron Other 0.547234 1.19E-08 0.005933 e LCMS_409C20 sphinganine(1+) Other 0.856298 1.04E-14 0.298526LCMS_410Resveratrol Other 0.969733 1.70E-09 0.70077LCMS_ 3,4-Dihydro-2H-1- 411benzopyran-2-oneOther 0.410648 0.00045 0.02884LCMS_ N- 412 (icosanoyl)ethanola Other 0.717473 5.56E-14 0.010904 mine LCMS_413Heteroflavanone C Other 0.73068 9.49E-10 0.074644LCMS_414PS(18:0 / 18:0) Polar lipids 0.076925 3.93E-08 0.795072LCMS_415Astaxanthin Terpenoids 0.091034 0.013159 0.7911954alpha-Carboxy- LCMS_ 4beta-methyl- 416 5alpha-cholesta-8- Terpenoids 0.449636 0.275382 0.483817 en-3beta-ol LCMS_417Solavetivone Terpenoids 0.700585 4.48E-07 0.047039LCMS_ TG(14:1(9Z) / 14:1(9Z 418 ) / 18:4 Neutral lipids 0.240164 0.000773 0.151033 (6Z,9Z,12Z,15Z)) LCMS_ PE(22:2(13Z,16Z) / 2 4190:1(11Z))Polar lipids 0.471568 2.18E-05 0.2991LCMS_ 4alpha-Formyl- 4204beta-Terpenoids 0.650665 2.86E-09 0.18503 methyl-5alpha- cholesta-8,24-dien- 3beta-ol LCMS_ DG(16:0 / 22:5 421 (7Z,10Z,13Z,16Z,19 Neutral lipids 0.685106 0.000536 0.145605 Z) / 0:0) LCMS_ trans-2-Hexenedioic 422acidOther 0.715962 0.002181 0.109838LCMS_423Pitheduloside I Terpenoids 0.754825 0.012472 0.632134LCMS_425Vitamin K1 Other 0.879921 9.54E-05 0.293639LCMS_427Ginsenoside Rg3 Terpenoids 0.063535 0.014471 0.003089LCMS_ 6-[5]-ladderane-1- 428hexanolOther 0.791844 0.008388 0.334098LCMS_429Sativic acid Neutral lipids 0.645702 0.01315 0.769245LCMS_ 2,2'- 430DiketospirilloxanthinTerpenoids 0.306131 2.76E-05 0.117463LCMS_ 2-trans,6-trans- 431FarnesalTerpenoids 0.67136 2.81E-05 0.30917LCMS_432Lactarazulene Terpenoids 0.541525 9.24E-05 0.411908LCMS_4332,5-Undecadienal Other 0.090351 1.94E-05 0.932559LCMS_437PS(18:0 / 20:0) Polar lipids 0.4457 0.001689 0.482701LCMS_438Phylloquinol Terpenoids 0.805511 7.57E-05 0.170988LCMS_ 3a,7a,12a- 439 Trihydroxy-5b- Other 0.964629 0.001591 0.124755 cholestan-26-al LCMS_ DG(14:0 / 20:5(5Z,8Z, 44111Z,14Z,17Z) / 0:0)Neutral lipids 0.757866 0.055339 0.020241LCMS_ PI(21:0 / 22:2(13Z,16 443Z))Polar lipids 0.315263 9.14E-07 0.023826LCMS_446Cer(d18:0 / 16:0) Other 0.499642 2.41E-09 0.548982LCMS_447Coproporphyrin III Other 0.338993 1.04E-09 0.909709LCMS_448Canesceol Other 0.189063 0.001859 0.050136LCMS_ PC(14:0 / 18:3(9Z,12 450Z,15Z))Polar lipids 0.642848 0.029378 0.251561LCMS_452d-Tocotrienol Other 0.264007 8.71E-06 0.309771LCMS_453Docosanamide Other 0.904822 0.024512 0.172847LCMS_ PC(16:1(9Z) / 22:6(4Z 455 ,7Z,10Z,13Z,16Z,19 Polar lipids 0.658716 0.011128 0.483625 Z)) LCMS_456Cer(d18:0 / 18:0) Other 0.758076 1.44E-07 0.071378LCMS_457Phenylethylamine Other 0.301614 5.78E-09 0.050521 LCMS_ Diadenosine 458heptaphosphateOther 0.702066 5.53E-05 0.002607LCMS_460Avrainvilloside Polar lipids 0.442306 2.63E-05 0.051714LCMS_ Didodecyl 463thiobispropanoateOther 0.862032 9.41E-06 0.160319LCMS_ alpha-Tocopherol 464acetateOther 0.658137 1.49E-08 0.836967LCMS_ PS(22:6(4Z,7Z,10Z, 465 13Z,16Z,19Z) / 20:2(1 Polar lipids 0.901617 4.67E-11 0.91869 1Z,14Z)) LCMS_466PS(O-20:0 / 22:0) Polar lipids 0.666304 1.17E-05 0.068865LCMS_ Cyclotricuspidoside 467COther 0.07347 9.49E-05 0.755785LCMS_ Cer(d16:2(4E,6E) / 18 469:1(9Z)(2OH)) 0.112233 8.06E-08 0.801432LCMS_ PI(P- 47016:0 / 22:2(13Z,16Z))Polar lipids 0.81906 0.000363 0.618745LCMS_ Cer(d16:2(4E,6E) / 20 471:0(2OH))Other 0.832419 4.78E-15 0.139415LCMS_ Glucosylceramide 472(d18:1 / 24:1(15Z))Polar lipids 0.31679 0.014744 0.067301LCMS_ P1,P4-Bis(5'- 475 xanthosyl) Other 0.991681 0.000367 0.087315 tetraphosphate LCMS_ Cyclohex-1,5-diene- 4761-carboxyl-CoAOther 0.830443 0.032419 0.039335LCMS_ Feruloyl C1- 478glucuronide LCMS_ PI(P- 479 18:0 / 22:6(4Z,7Z,10Z Polar lipids 0.099867 3.24E-07 0.085609 ,13Z,16Z,19Z)) LCMS_ TG(18:0 / 20:0 / 20:4(5 480Z,8Z,11Z,14Z))Neutral lipids 0.141749 3.50E-10 0.042362LCMS_481Phaeophorbide b Other 0.27198 0.069475 0.052618LCMS_ PI(17:1(9Z) / 20:2(11Z 482,14Z))Polar lipids 0.370972 0.000976 0.44772LCMS_ PC(P- 48418:0 / 20:1(11Z))Polar lipids 0.078264 5.62E-07 0.071955LCMS_ 6-methoxy-2- 485 octaprenylhydroquin Other 0.338956 0.001725 0.33888 one LCMS_ Galabiosylceramide 486(d18:1 / 16:0)Polar lipids 0.84516 7.02E-11 0.834176LCMS_ 1- 487 Piperidinecarboxalde Other 0.950049 5.68E-08 0.216399 hyde To test the basis for metabolic changes, we examined transcriptome profiles among these accessions in response to drought stress. Consistent with the observed metabolomic changes, 68% of the variation within the panel could be explained by the irrigation treatment (Figure 1E). In addition, neither PC 2 nor PC 3 could be attributed to genotypic differences. Pairwise comparisons of gene expression under the two conditions generated a range of differentially expressed genes (DEGs) for each accession (Table 4). We further compared transcriptomes to determine if DEGs were shared between all intragroup accessions, or between all accessions within multiple subgroups in response to drought. Out of the thousands of DEGs in each accession, there were few genes that were shared among all accessions within a subgroup, or between subgroups (Figure 3A, Table 4). The foreign accessions shared the least group-wide DEGs (up-regulated: 57 and downregulated: 86), whereas the mixed accessions featured the most (up-regulated: 194 and down-regulated: 177) (Figure 3B). In contrast, there were 300 DEGs (194 upregulated, 106 downregulated) shared among all examined accessions. Gene ontology (GO) enrichment of these shared up-regulated genes revealed an over-representation of the stress response and phosphorylation signal transduction genes while the shared down- regulated genes were involved in fatty acid biosynthetic processes and transport processes (Figure 3C and Figure 16). Thus, the data indicates that these shared DEGs represent a common set of genes involved in drought stress response. TABLE 4: Summary of DEGs derived from pair-wise comparison of 21 accessions Accession P-adjDown- Up- <0.05regulatedregulated*Other GroupAC_134_CB_4029 4719 3671 1007 CB_4012_KING_KARAJAZSY 3219 1692 1246 281 FELISTANA_UA_7_18 11050 5676 4296 1078 FUNTUA_FT_5 5123 2190 2512 421 Foreign VIR_6615_MCU_5 6796 2958 3071 767 VIR_7065_ALLEN_333 8928 4073 3918 937 WESTERN_STORMPROOF 6604 2665 3314 625 DP_393 6689 3234 2797 658 DP_493 10819 4372 5268 1179 FM_958 7071 3153 3157 761 LOCKETT_BXL 11558 5190 5180 1188 U.S. TAM_86III_26 6887 3461 2689 737 TAM_91C_34 6020 3151 2290 579 TAM_94WE_37S 5662 2606 2522 534 COKER_310 10175 4645 4500 1030 MEXICO_910 7135 3084 3434 617 PD_3 9485 4735 3780 970 Mixed PLAINS 7295 4212 2350 733 VIR_7094_COKER_310 12870 5946 5550 1374 VIR_7153_D_10 2948 1420 1283 245 VIR_7223 11365 6155 3947 1263 Upland cotton has a well reported subgenome expression bias towards the D (new world) subgenome when comparing homeologous gene pairs across a wide range of tissues (Chen et al., 2020). Subgenome expression dominance is believed to be influenced by changes in the environment (Bird et al., 2018), thus, we next examined how subgenome expression dominance was impacted at a global scale (comparing all expressed genes in A and D subgenomes) across our panel in response to water limiting conditions. After removal of genes with low expression (average TPM across samples < 1; 75,376 genes retained), we first compared expression dominance under well-watered and water limited conditions separately. Under well-watered conditions, 10 / 21 accessions displayed a significant bias in mean gene expression towards the A subgenome (pair-wise t-test, P < 0.05; Figure 16), with the remainder not showing any bias. In contrast, under water limiting conditions the bias towards the A subgenome became more pronounced (17 / 21 accessions). Interestingly, three accessions with a bias under well-watered conditions lost that bias under water limiting: Western Stormproof, DP 393, and Mexico 910. Each of these three accessions have reported tolerance to hot and dry conditions, although they are not the only accessions with this trait in our panel (Figure 17). Thus, this global analysis suggests that genes from the old-world subgenome are predominantly expressed under hot (well- watered) and hot / dry (water-limited) conditions. Example 3: Examination of Transcriptome-Trait Connections in Response to Drought To determine a relationship between transcriptomic and phenotypic variability, we performed a weighted gene co-expression analysis (WGCNA) to uncover the association of gene networks with phenotypes. For the WGCNA, we incorporated the top 10% most variable genes across the 21 accessions under WL conditions (determined by median absolute deviation [MAD], N = 4,552 genes). We obtained 22 modules that satisfied a scale-free topology (R2= 0.86; Figure 4A) using a soft threshold (Beta = 7) for network construction (Figure 4B). These 22 modules displayed clear separation, with only 45 genes (< 0.2 %) unclassified (Figure 4C). Using PCC, these 22 well-clustered modules were then correlated to the trait dataset, which consisted of 10 phenotypes and 451 metabolites. Among these traits being tested, five traits (lint percentage, referred to here as lint yield, length uniformity (UI), scaled photochemical reflective index (sPRI), carotenoid reflectance index (CRI), and water index / normalized difference vegetation index (WI / NDVI)) and nine metabolites exhibited significant correlation to 18 modules (Figure 5, P < 0.05). In particular, lint yield displayed the highest positive correlation to the turquoise module (r = 0.68, p = 7e-04), whereas terpenoid and carbohydrate metabolites displayed a positive correlation with the salmon module (P < 0.05; Figure 5). Example 4: Identification of Trait Associated Functional Modules Following module-trait correlation, we selected key module(s) for further functional analysis by incorporating the profile of GO enrichment, transcription factor (TF) binding motif enrichment, and gene expression variance (Figure 5A). Using Fisher test (P < 0.01) for GO enrichment, 14 of the 18 modules exhibited some degree of enrichment for genes involved in biological processes of interest. Notably, the lint yield-associated module contained a high degree of enriched stress response genes (Q-value: 1.6e-12) as well as protein folding genes (Q- value: 4.6e-8). Also, a large number of photosynthesis-related and redox-related genes (Q-value: 2.3e-10and Q-value: 8.4e-4, respectively) were identified within the blue module, negatively correlated with the abundance of multiple metabolites (Figure 5A). These data indicate that a co- expression network-based approach can uncover key genes integrating plant development and yield in the response to water limitation. To further investigate the upstream regulators of those genes associated with enriched GO terms, we took the 2-kb upstream promoter region sequences of all genes (WGCNA weight score > 0.1) in each of the 18 modules to perform motif enrichment analysis using AME (Analysis of Motif Enrichment; McLeay and Bailey, 2010). To narrow down the list of possible enriched TFs / motifs, we focused on TF / motif pairs where the TF was also present in the same module and thus displayed similar expression patterns across treatments and genotypes. This approach uncovered four modules with enriched motifs corresponding to 10 different transcription factors, with the lint-yield module (turquoise) containing the most enriched TFs (n = 6, AME TFs, Figure 5B). As the data suggested that significant trait-associated modules display a more pronounced response to treatment, we assessed expression variation for genes within each module between accessions and conditions by calculating the coefficient of variance (CV). While several modules showed a significant increase in coefficient of variance (CV) relative to background, the turquoise module contained numerous stress response genes with high CV (Figure 5C). In sum, these data indicate that the genes within the “turquoise” module (highlighted with a star in Figure 5A), and their associated TFs, may be critical for maintaining yield under water limiting conditions. Example 5: DAP-seq Analysis of HSFA6B and DREB2A Among six TFs in the turquoise module that are associated with yield, four were AME TFs that are predicted to be associated with heat or drought stress based on homology with functionally characterized TFs in Arabidopsis and cotton (Figure 6A). We further compared the gene from the enriched GO / KEGG terms (N = 119) in this module, as well as seven genes identified in previous GWAS, with lint yield for the presence of TF motifs (Table 4). These four TFs: HSF7, HSF6, HSFA6B, and DREB2A (Figure 6A, upper panel) are predicted to bind to a number of stress and heat response genes (Figure 6A, bottom panel). HSFA6B and DREB2A are also predicted to bind to the seven lint-yield associated genes in the module (Figure 6A (boxed)). This indicates that the cotton DREB2A and HSFA6B homologs are candidates for the observed association between water stress and fiber yield. Given their association with fiber yield and drought stress, these two TFs likely regulate downstream stress responsive genes. To test this, we performed DNA affinity purification sequencing (DAP-seq) using cotton DREB2A and HSFA6B and examined the DNA-binding ability of the homeologs from both subgenomes (e.g., GhDREB2A-A and GhDREB2A-D). More than 40 million (DREB2A) and 37 million (HSFA6B) single-end reads (SE: 150 bp) per replicate enabled the identification of more than 20,000 peaks with high reproducibility for each sample after peak calling. For each TF, despite both being expressed in an in vitro wheat germ system (Figure 7), only a single homeolog, namely, GhHSFA6B-D, derived from the D subgenome (Gohir.D08G072600) and GhDREB2A-A, derived from the A subgenome (Gohir.A13G021700), could bind to DNA above background. Using Irreproducibility Discovery Rate as a cutoff (IDR, P < 0.05), a total of 16,927 and 16,704 high confidence peaks were identified between replicates of GhDREB2A-A and GhHSFA6B-D (Figure 6B,). After filtering for GhHSFA6B-D and GhDREB2A-A peaks in the 5-kb upstream region of annotated coding genes (distal promoter), as well as peaks within the 5’ untranslated regions (UTRs) of coding genes (proximal promoter), a total of 5,229 and 3,178 genes were identified as likely regulatory targets of GhDREB2A-A and GhHSFA6B-D, respectively (Figure 6B, Table 5 and Figure 17). Genomic sequences associated with these peaks were further processed by motif analysis (MEME suite; Bailey et al., 2015) to identify binding sites and test the levels of conservation of consensus motifs. As expected, both target sequences of the two TFs revealed high-level conservation (E-value: 3.8e-855across 3180 peaks and 5.9e-603among 5218 peaks; Figure 6B), compared to core HSFA6B and DREB2A binding elements (HSEs and DREs) from other species. Interestingly, although our AME TF approach only predicted a GhHSFA6B-D binding motif in the lint yield-associated gene GhIPS1-A, (Myo- inositol-phosphate synthase protein), we observed DAP-seq peaks for GhHSFA6B-D binding to both GhIPS-A and GhABP-A (Figure 8). In addition, we observed GhHSFA6B-D peaks in the promoter region of both GhDREB2A-A and GhDREB2A-D homeologs (Figure 9) but did not observe reciprocal GhDREB2A-A peaks in the promoter of GhHSFA6B-D, indicating that GhHSFA6B-D acts as a regulator of GhDREB2A-A. Together, these data indicate that GhHSFA6B-D and GhDREB2A-A regulate expression of stress responsive pathway genes under drought in cotton. TABLE 5A: Summary of lint yield candidate genes identified from previous studies Publication Network- Gene -GeneID GeneIDNameDescription Publication IdentitiesGh_A02G1 Gohir.A02 Inositol-3-phosphate Su et al., 268G132300IPSsynthase 2016 100 Gh_A12G0 Gohir.A12ABP Auxin-binding protein T85Zhu et al. 043G006700(2020) 100 Gh_D01G0 DNA-binding DBC br Gao et al., 099Gohir.D01omodomain-containing G009700protein2021100Gh_D06G2 Gohir.D10APIApoptosis inhibitory protein Su et al. 153G02990055 (API5)(2020)90.86Lung seven Gh_D12G2 Gohir.A12 L7T transmembrane receptor Sun et al. G240700 family protein 98.246Gh_D08G0 Gohir.A08 HSP8 HEAT SHOCK PROTEIN Fang et al. 299 G025400 1.4 81.4 (2017) 98.524 Gh_D08G0 Gohir.D08 HSP8 HEAT SHOCK PROTEIN Fang et al. 300 G036000 1.4 81.4 (2017) 99.905 TABLE 5B: Summary of lint yielded candidate genes identified from previous studies Publication- eneIDLengQuerry Querry Ref- Ref- Gth SNP Gap-start -end startendE-valueGh_A02G12681533 0 0 1 1533 1 1533 0Gh_A12G0043573 0 0 1 573 1 573 0Gh_D01G00991683 0 0 1 1683 1 1683 0Gh_D06G21531652 145 3 1 1646 1 1652 0Gh_D12G2354 1311 23 0 1 1311 1 1311 0 Gh_D08G0299 2100 31 0 1 2100 1 2100 0 Gh_D08G0300 2100 2 0 1 2100 1 2100 0 To determine if GhHSFA6B-D and GhDREB2A-A might have a specific impact under stress on genes found within the lint yield module, we profiled transcriptome changes between treatments. A comparison of average log2FC (WL relative to WW) for in-module targets of GhHSFA6B-D or GhDREB2A-A revealed a significant up-regulation relative to global putative TF targets (P < 0.001, Figure 6C). The analysis of GO enrichment (P < 0.01) for these downstream target genes found that only the genes bound by GhHSFA6B-D in the module were enriched for stress response terms (-log10(Qvalue) > 8), whereas both GhHSFA6B-D and GhDREB2A-A exhibited in-module specificity to chaperone / protein folding related genes (Figure 6D). These data indicate that while both TFs display a high degree of regulatory connections between fiber yield and stress responses, GhHSFA6B-D may be a major regulator. To better visualize the interactions between these two TFs and other genes within the lint yield module, including each other, we synthesized our data into a regulatory network using Cytoscape (Figure 10A). To simplify this network, 866 genes in this module were reduced to a subset of 126 core genes by selecting genes with enriched GO terms, KEGG pathways, annotated as TFs or lint yield associated genes, and those highly correlated with lint yield (trait-correlated genes; Figure 10A). Trait-correlated genes are those with high module membership (MM), which is derived from the correlation between gene expression and the eigenvalue (first principal component) of the lint yield module, and gene significance (GS), a term reflecting the correlation between gene expression and lint yield (Figure 10B, orange circles). These target genes are those whose expression change across the 21 accessions is most likely to explain variation in lint yield in this panel. To further develop this network, we integrated the co-expression information derived from WGCNA, TF-target binding from our DAP-seq data, and protein-protein interactions (PPIs) from the STRING database (Figure 10, top panel). This network illustrates the high degree of connectivity between GhHSFA6B-D and the trait-correlated genes, both in terms of expression and direct connection (Figure 10A). Of the 24 trait-correlated genes, GhHSFA6B-D is predicted to bind 16 target genes based on DAP-seq analysis (Figure 10A ), whereas GhDREB2A-A is predicted to only bind two genes based on AME motif enrichment. As mentioned above, GhHSFA6B-D binds to both GhDREB2A-A and GhDREB2A-D homoeologs based on DAP-seq data. In addition, GhHSFA6B-D was found to bind to its own distal promoter region, indicating that it acts in an autoregulatory loop (Figure 10A ). Finally, despite AME predicting interactions between GhHSFA6B-D, GhDREB2A-A, and the five lint yield associated genes, only two DAP-seq derived peaks between GhHSFA6B-D and GhIPS1-A and GhABP-D were observed. As the initial AME TF predictions were made based on DAP-seq data from Arabidopsis (O’Malley et al., 2016), it is not unexpected that TF binding preferences will have shifted slightly for these TFs in cotton, highlighting the importance of our TF binding validation. The inferred gene regulatory network, developed from multiple lines of evidence, suggests a direct regulatory interaction between HSFA6B, DREB2A, and two of the five lint-yield associated genes, GhABP and GhIPS. GhABP-A (Gohir.A12G006700) is an auxin-binding protein involved in cell elongation and cell division that was previously identified as a lint yield QTL (Zhu et al., 2020). TF binding motif predictions based on the Arabidopsis Cistrome Database (O’Malley et al,.2016) identified DREB2A and HSFA6B motifs for both homeologous GhABP loci (GhABP-D and GhABP-A; yellow boxes, Figure 8A and Figure 11). GhABP-A, but not GhABP-D, (Gohir.D12G006100; GhABP-D), was bound by both TFs based on our DAP-seq data (GhHSFA6B-D peak center: -561 bp TSS and GhDREB2A-A peak center: -1,445bp TSS, Figure 8A). GhABP-D was also absent from the lint yield module, and a comparison of expression versus lint yield across the 21 accessions revealed that GhABP-D was expressed at lower levels than GhABP-A (Figure 9B). In addition, GhABP-A displayed a stronger correlation with the lint yield across the accessions. Despite the lower expression level of GhABP-D than GhABP-A, the abundance of both transcripts was positively correlated with both TFs (Figure 8). A sequence comparison of the identified TF binding regions between GhABP-A and GhABP-D revealed only minor changes near the GhDREB2A-A (two SNPs, Figure 8A, top) and within the HSFA6B motifs (four SNPs, Figure 8A, bottom). No changes were observed within the core dehydration responsive element (DRE) or heat shock element (HSE) that these two TFs are known to bind to in other systems (Figure 8A, red boxes). To determine if the observed SNPs are responsible for the altered transcript abundance 5 and DAP-seq peaks arising from these two homologous loci, we performed an electrophoretic mobility shift assay (EMSA, Table 6), using ~50-bp probes that spanned the respective DRE and HSE motifs (Figure 8A, pink bars). Both recombinant GhDREB2A-A and GhHSFA6B-D bound to their respective labeled probes, causing a gel shift that was not evident in the probe alone or probe + empty vector controls (Figures 8D and 8E). As expected, adding a large molar excess 10 (200 ×) unlabeled probe to the reaction was sufficient to compete for protein binding. Interestingly, the probes corresponding to the GhABP-D homeolog were also able to compete for binding with GhDREB2A-A and GhHSFA6B-D (Figures 8D and 8E). These data suggest that the SNPs observed between the GhABP-A and GhABP-D distal promoter regions are insufficient to explain the specificity observed between these paralogs. TABLE 6: Summary of DAP-seq data processing Sequenced Clean Mapped TF Reads Reads reads Mapping Unique Unique Mapping map Called IDR Number Number Number Rate ping reads rate peaks Peak # DREB2CA_A7,226,216 7,054,076 6,846,699 0.9706018 4,196,865 0.594956 16,607DREB2 4,470CA_B5,407,037 5,264,651 5,097,153 0.9681844 3,102,120 0.589235 9,661HSFA6bD_A7,229,447 7,010,821 6,808,153 0.9710921 4,195,744 0.598466 19,255HSFA6 9,493bD_B10,728,070 10,448,269 10,166,834 0.973064 6,357,220 0.608447 21,472DREB2CA_A40,728,909 39,286,098 37,234,736 0.947784 22,046,413 0.561175 24,398DREB2 16,927CA_B45,421,136 43,688,226 41,280,514 0.9448888 24,261,062 0.555322 24,213HSFA6bD_A49,585,378 47,423,554 45,078,256 0.9505457 26,897,952 0.567185 20,558HSFA6b 16,704D_B37,542,217 36,138,119 34,501,636 0.9547159 20,990,704 0.580846 26,40715 *Batch 1 is bold; Batch 2 is italicized Example 6: Analysis of Inositol-Phosphate Synthase (IPS) Modulation on Cotton Lint Yield The IPS gene, also known as MIPS in Arabidopsis, encodes for the Myo-inositol-phosphate synthase protein (INO-1), which catalyzes the rate limiting step in the synthesis of Myo-inositol- 6-phosphate, a key source of phosphate in seed endosperm (Mitsuhashi et al., 2008). In addition, myo-inositol is a precursor of the osmoprotectants galactinol and raffinose, and thus is critical in a number of abiotic and biotic stress responses (Vinson et al., 2020). In contrast to ABP, there are four IPS genes in upland cotton (Gohir.D03G043600 - GhIPS1-D, Gohir.A02G132300 - GhIPS1-A, Gohir.D11G224000 - GhIPS2-D, and Gohir.A11G199700 - GhIPS2-A, Figure 12), two of which (GhIPS1-D and GhIPS1-A) show a positive correlation between RNA abundance and lint yield across our panel (Figure 13B). These two homeologous IPS genes fall in syntenic regions of the “D” and “A” subgenomes and are phylogenetically distinct from GhIPS2-A and GhIPS2-D (Figure 14A). Despite a reported expression bias towards D subgenome homeologs (Chen et al., 2020), only GhIPS1-A was predicted to contain a GhHSFA6B-D binding site based on DAP-seq data and the Arabidopsis Cistrome Database (Figure 13A, purple and yellow boxes, respectively). A pairwise comparison of the distal promoter regions of GhIPS1-A and GhIPS1-D revealed substantial polymorphisms between the two regions that appear to have disrupted the core HSE in this region in GhIPS1-D, as well as between GhIPS1-D and GhIPS2-A / D (Figure 13A, bottom; Figure 12A-12B). An examination of IPS1 distal promoter elements in G. hirsutum, G. barbadensis, G. raimondii (a D-subgenome representative) and G. arboreum (old-world cotton and an A-subgenome representative) revealed conservation of the HSE within GhIPS1-A in G. hirsutum, G. barbadensis, and G. arboreum, but not in IPS1-D loci for any of these species (Figure 14A). In agreement with the DAP-seq data, we observe a strong positive correlation between GhHSFA6B-D and GhIPS1-A transcript abundance, but not between GhHSFA6B-D and any of the other GhIPS loci (Figure 13C; Figure 8). Interestingly, an HSE, and AtHSFA6B DAP-seq peak, were observed upstream of the Arabidopsis IPS1 and IPS2 paralogs in Arabidopsis Cistrome data (AT4G39800 and At2G22240, respectively; Figure 14A), suggesting that this regulatory mechanism is conserved between these two species. An examination of polymorphisms within our RNA-seq data relative to the TM-1 reference genome uncovered a lint-yield associated SNP directly adjacent to the predicted GhHSFA6B-A binding motif (Figure 13A). While this nucleotide is a cytosine (C) at this position in the reference genome and in G. barbadensis and G. arboreum, we observed two genotypes in our panel, either with a cytosine (C, n = 8) or a thymine (T, n = 9; Figure 13D). Interestingly, the “C” genotypes, which were either the US or “mixed” accessions, showed increased lint yield under water-limited conditions relative to the “T” genotypes (Figure 13D), suggesting this region might impact GhHSFA6B-D binding. To define the GhHSFA6B-D binding site more carefully in the GhIPS1-A promoter region, we designed a labeled probe centered on the core HSE (Figure 13A, pink box). A gel shift was observed when this probe was combined with in- vitro expressed GhHSFA6B-D protein (Figure 13E). As expected, this signal was abolished when a large excess of unlabeled probe was added. The addition of an unlabeled competitor, corresponding to the homologous GhIPS1-A promoter region with a disrupted HSE (Oligo 2, Figure 13E) was unable to abolish binding, even when adding a large molar excess (400× and 600×; Figure 13E). As the labeled probe contained the reference nucleotide (C) at the site of the observed lint yield-associated SNP, we next tested if altering this nucleotide, but leaving the rest of the HSE intact, had an impact on HSFA6B binding. Oligos with this site altered to an A, T, or G nucleotide (SNP probes 1-3) could compete for GhHSFA6B-D binding when added in excess (200×; Figure 13E). While the SNP oligos were competing with the reference oligo when added in excess, due to its location outside of the HSE it is not clear if the site of the SNP impacts GhHSFA6B-D binding. To precisely address this question, we performed a competition experiment using lower concentrations of two non-native SNP probes, SNP-A and SNP-G (SNP probes 1 and 3), that are not present in the genomes of our panel or a much larger sequenced panel comprised of 1024 accessions (Yuan et al., 2021). Surprisingly, equimolar amounts of either non-native competitor SNP oligos were capable of competing with the reference probe to a higher degree than the unlabeled reference oligo (Figure 13B, Figure 14C, and Figure 15). These data suggest that this site, which contains the only observed SNP within the DAP-seq peak region of GhIPS1-A in our panel, has a strong influence on GhHSFA6B protein binding, GhIPS1-A expression, and lint yield. Given the differentially accumulated frequency of SNPs at this site in extant accessions (wild, cultivated, improved, and mutant) (Figure 14B), this site has likely been under different levels of selection in wild and cultivated accessions due to its connection to improved yield. Understanding the regulatory crosstalk between agronomic traits of interest (e.g., yield) and heat and drought stress responses are critical for developing drought-tolerant cultivars with minimal impacts on yield (Alizadeh et al., 2020; Sinclair, 2011). Characterizing genotype- specific molecular responses under water-limiting conditions is challenging due to the genetic complexity underlying quantitative physiological traits (Welcker et al., 2011; Tardieu et al., 2011, 2014). In most crop species this genetic complexity is influenced at the pre- and post- transcriptional level by cis-regulatory control, gene structural variation, and post-translational modifications (Tardieu et al., 2018; Joshi et al., 2016). In cotton, this complexity is further compounded by a recent (~1.0-1.6 Ma) allopolyploidy and strong human selection within the last 8,000 years (Chen et al., 2020). Despite the strong selection existing in current elite G. hirsutum cultivars relative to other domesticated crops (Chen et al., 2020; Ma et al., 2018; Su et al., 2016), variation in gene expression was identified across the panel for a cohort of transcripts that could be associated with improved yield under water limiting conditions. Despite the developmental stage at which we sampled for transcriptomics (primarily the cell elongation stage within the developing bolls; Ma et al., 2021) there was already a strong transcriptomic signal in our data connecting abiotic stress factors and lint yield. This group of lint yield and abiotic stress-associated genes were largely regulated by the well-known transcription factors GhDREB2A-A and GhHSFA6B-D. Importantly, based on DAP-seq, motif enrichment, and co-expression, HSFA6B not only regulates GhDREB2A-A in this module, but also regulates itself and two lint yield-associated genes, GhABP and GhIPS1-A. While GhDREB2A-A was previously shown to be regulated by GhHSFA6B-D in Arabidopsis, this regulation was dependent on ABA in general, and specifically the ABA response element TF AREB1 (Sakuma et al., 2006, Huang et al., 2016). In contrast, in cotton this regulation does not appear to be ABA-dependent, as none of the typical ABA-responsive elements, nor ABA biosynthesis genes, were associated with this module, nor does HSFA6B contain a canonical AREB binding domain (ABRE) or share an expression profile with GhAREB1. However, this hypothesis requires further examination. This specialized regulatory mechanism appears to have emerged in cotton well before domestication. Indeed, the old-world variant of IPS1 arising from the A subgenome (GhIPS1-A) is the only IPS to be targeted by GhHSFA6B-D. Of importance to breeders, this regulatory mechanism appears to have undergone additional selection since speciation to maintain yield under water limited conditions. References Alizadeh, M.R., Adamowski, J., Nikoo, M.R., AghaKouchak, A., Dennison, P., and Sadegh, M. (2020). A century of observations reveals increasing likelihood of continental-scale compound dry-hot extremes. Sci. Adv.6: 1–12. Almoguera, C., Prieto-dapena, P., Díaz-martín, J., Espinosa, J.M., Carranco, R., and Jordano, J. (2009). The HaDREB2 transcription factor enhances basal thermotolerance and longevity of seeds through functional interaction with HaHSFA9.12: 1–12. Bae, H., Kim, S.H., Kim, M.S., Sicher, R.C., Lary, D., Strem, M.D., Natarajan, S., and Bailey, B.A. (2008). The drought response of Theobroma cacao (cacao) and the regulation of genes involved in polyamine biosynthesis by drought and other stresses. Plant Physiol. Biochem.46: 174–188. Bolger, A.M., Lohse, M., and Usadel, B. (2014). Trimmomatic: A flexible trimmer for Illumina sequence data. Bioinformatics 30: 2114–2120. Bradbury, P.J., Zhang, Z., Kroon, D.E., Casstevens, T.M., Ramdoss, Y., and Buckler, E.S. (2007). TASSEL: Software for association mapping of complex traits in diverse samples. Bioinformatics 23: 2633–2635. Brouard, J.S., Schenkel, F., Marete, A., and Bissonnette, N. (2019). The GATK joint genotyping workflow is appropriate for calling variants in RNA-seq experiments. J. Anim. Sci. Biotechnol.10: 1–6. Burke, J. J., and Wanjura, D. F. (2010). “Plant responses to temperature extremes,” in Physiology of Cotton, eds J. M. Stewart, D. M. Oosterhuis, J. J. Heitholt, and J. R. Mauney (Dordrecht: Springer), 123–128. doi: 10.1007 / 978-90-481-3195-2_12 Caturegli, L., Matteoli, S., Gaetani, M., Grossi, N., Magni, S., Minelli, A., Corsini, G., Remorini, D., and Volterrani, M. (2020). Effects of water stress on spectral reflectance of bermudagrass. Sci. Rep.10: 1–12. Chávez Montes RA, Haber A, Pardo J, Powell RF, Divisetty UK, Silva AT, Hernández- Hernández T, Silveira V, Tang H, Lyons E, Herrera Estrella LR, VanBuren R, Oliver MJ. A comparative genomics examination of desiccation tolerance and sensitivity in two sister grass species. Proc Natl Acad Sci U S A.2022 Feb 1;119(5):e2118886119. doi: 10.1073 / pnas.2118886119. Chen, Z.J. et al. (2020). Genomic diversifications of five Gossypium allopolyploid species and their impact on cotton improvement. Nat Genet 52: 525–533. Christensen, A., Svensson, K., Persson, S., Jung, J., Michalak, M., Widell, S., and Sommarin, M. (2008). Functional characterization of Arabidopsis calreticulin1a: A key alleviator of endoplasmic reticulum stress. Plant Cell Physiol.49: 912–924. Chu, X., Wang, C., Chen, X., Lu, W., Li, H., Wang, X., Hao, L., and Guo, X. (2015). The cotton WRKY gene GhWRKY41 positively regulates salt and drought stress tolerance in transgenic Nicotiana benthamiana. PLoS One 10: 1–21. Claeys, H. and Inzé, D. (2013). The Agony of Choice: How Plants Balance Growth and Survival under Water-Limiting Conditions. Plant Physiol.162: 1768–1779. Danecek, P., Auton, A., Abecasis, G., Albers, C.A., Banks, E., DePristo, M.A., Handsaker, R.E., Lunter, G., Marth, G.T., Sherry, S.T., McVean, G., and Durbin, R. (2011). The variant call format and VCFtools. Bioinformatics 27: 2156–2158. Dekkers, B.J.W., Schuurmans, J.A.M.J., and Smeekens, S.C.M. (2004). Glucose delays seed germination in Arabidopsis thaliana. Planta 218: 579–588. Delatte, T.L., Sedijani, P., Kondou, Y., Matsui, M., Jong, G.J. De, Somsen, G.W., Wiese- klinkenberg, A., Primavesi, L.F., Paul, M.J., and Schluepmann, H. (2011). Growth Arrest by Trehalose-6-Phosphate : An Astonishing Case of Primary Metabolite Control over Growth by Way of the SnRK1 Signaling Pathway 1 [ C ][ W ][ OA ].157: 160–174. Delorge I, Janiak M, Carpentier S, Van Dijck P. Fine tuning of trehalose biosynthesis and hydrolysis as novel tools for the generation of abiotic stress tolerant plants. Front Plant Sci. 2014 Apr 14;5:147. doi: 10.3389 / fpls.2014.00147. Dietrich, K., Weltmeier, F., Ehlert, A., Weiste, C., Stahl, M., Harter, K., and Dröge-Lasera, W. (2011). Heterodimers of the Arabidopsis transcription factors bZIP1 and bZIP53 reprogram amino acid metabolism during Low energy stress. Plant Cell 23: 381–395. Ding, L.N., Li, M., Wang, W.J., Cao, J., Wang, Z., Zhu, K.M., Yang, Y.H., Li, Y.L., and Tan, X.L. (2019). Advances in plant GDSL lipases: from sequences to functional mechanisms. Acta Physiol. Plant.41: 1–11. Do, P.T., Degenkolbe, T., Erban, A., Heyer, A.G., Kopka, J., Köhl, K.I., Hincha, D.K., and Zuther, E. (2013). Dissecting Rice Polyamine Metabolism under Controlled Long-Term Drought Stress. PLoS One 8. Fahad, S., Hussain, S., Saud, S., Khan, F., Hassan, S., Amanullah, Nasim, W., Arif, M., Wang, F., and Huang, J. (2016). Exogenously Applied Plant Growth Regulators Affect Heat-Stressed Rice Pollens. J. Agron. Crop Sci.202: 139–150. Fang, D.D., Jenkins, J.N., Deng, D.D., McCarty, J.C., Li, P., and Wu, J. (2014). Quantitative trait loci analysis of fiber quality traits using a random-mated recombinant inbred population in Upland cotton (Gossypium hirsutum L.). BMC Genomics 15. Fang, Y., Liao, K., Du, H., Xu, Y., Song, H., Li, X., and Xiong, L. (2015). A stress-responsive NAC transcription factor SNAC3 confers heat and drought tolerance through modulation of reactive oxygen species in rice. J. Exp. Bot.66: 6803–6817. Fernie, A.R., Stitt, M., On the Discordance of Metabolomics with Proteomics and Transcriptomics: Coping with Increasing Complexity in Logic, Chemistry, and Network Interactions Scientific Correspondence, Plant Physiology, Volume 158, Issue 3, March 2012, Pages 1139–1145, https: / / doi.org / 10.1104 / pp.112.193235 Goodstein, D.M., Shu, S., Howson, R., Neupane, R., Hayes, R.D., Fazo, J., Mitros, T., Dirks, W., Hellsten, U., Putnam, N., and Rokhsar, D.S. (2012). Phytozome: A comparative platform for green plant genomics. Nucleic Acids Res.40: 1178–1186. Guo, W.J. and Ho, T.H.D. (2008). An abscisic acid-induced protein, HVA22, inhibits gibberellin-mediated programmed cell death in cereal aleurone cells. Plant Physiol.147: 1710–1722. Gupta, A., Rico-Medina, A., and Caño-Delgado, A.I. (2020). The physiology of plant responses to drought. Science (80-. ).368: 266–269. Hayashi, H., De Bellis, L., Hayashi, Y., Nito, K., Kato, A., Hayashi, M., Hara-Nishimura, I., and Nishimura, M. (2002). Molecular characterization of an Arabidopsis acyl-coenzyme a synthetase localized on glyoxysomal membranes. Plant Physiol.130: 2019–2026. Hinze, L.L., Fang, D.D., Gore, M.A., Scheffler, B.E., Yu, J.Z., Frelichowski, J., and Percy, R.G. (2015). Molecular characterization of the Gossypium Diversity Reference Set of the US National Cotton Germplasm Collection. Theor. Appl. Genet.128: 313–327. Hinze, L.L., Gazave, E., Gore, M.A., Fang, D.D., Scheffler, B.E., Yu, J.Z., Jones, D.C., Frelichowski, J., and Percy, R.G. (2016). Genetic diversity of the two commercial tetraploid cotton species in the gossypium diversity reference set. J. Hered.107: 274–286. Huang, G.Q., Li, W., Zhou, W., Zhang, J.M., Li, D. Di, Gong, S.Y., and Li, X.B. (2013). Seven cotton genes encoding putative NAC domain proteins are preferentially expressed in roots and in responses to abiotic stress during root development. Plant Growth Regul.71: 101–112. Huang, J.G., Yang, M., Liu, P., Yang, G.D., Wu, C.A., and Zheng, C.C. (2009). GhDREB1 enhances abiotic stress tolerance, delays GA-mediated development and represses cytokinin signalling in transgenic Arabidopsis. Plant, Cell Environ.32: 1132–1145. Huang, Q., Wang, Y., Li, B., Chang, J., Chen, M., Li, K., Yang, G., and He, G. (2015). TaNAC29, a NAC transcription factor from wheat, enhances salt and drought tolerance in transgenic Arabidopsis. BMC Plant Biol.15: 1–15. Imkampe, J. et al. (2017). The arabidopsis leucine-rich repeat receptor kinase BIR3 negatively regulates BAK1 receptor complex formation and stabilizes BAK1. Plant Cell 29: 2285– 2303. Jaspers, P., Blomster, T., Brosché, M., Salojärvi, J., Ahlfors, R., Vainonen, J.P., Reddy, R.A., Immink, R., Angenent, G., Turck, F., Overmyer, K., and Kangasjärvi, J. (2009). Unequally redundant RCD1 and SRO1 mediate stress and developmental responses and interact with transcription factors. Plant J.60: 268–279. Joshi, R., Wani, S.H., Singh, B., Bohra, A., Dar, Z.A., Lone, A.A., Pareek, A., and Singla- Pareek, S.L. (2016). Transcription factors and plants response to drought stress: Current understanding and future directions. Front. Plant Sci.7: 1–15. Karim S, Aronsson H, Ericson H, Pirhonen M, Leyman B, Welin B, Mäntylä E, Palva ET, Van Dijck P, Holmström KO. Improved drought tolerance without undesired side effects in transgenic plants producing trehalose. Plant Mol Biol.2007 Jul;64(4):371-86. doi: 10.1007 / s11103-007-9159-6. Langfelder, P. and Horvath, S. (2008). WGCNA: An R package for weighted correlation network analysis. BMC Bioinformatics 9. Lei, L. et al. (2015). Ribosome profiling reveals dynamic translational landscape in maize seedlings under drought stress. Plant J.84: 1206–1208. Leyman, B., Van Dijck, P., and Thevelein, J.M. (2001). An unexpected plethora of trehalose biosynthesis genes in Arabidopsis thaliana. Trends Plant Sci.6: 510–513. Li HW, Zang BS, Deng XW, Wang XP. Overexpression of the trehalose-6-phosphate synthase gene OsTPS1 enhances abiotic stress tolerance in rice. Planta.2011 Nov;234(5):1007-18. doi: 10.1007 / s00425-011-1458-0. Liao, Y., Smyth, G.K., and Shi, W. (2014). featureCounts: an efficient general purpose program for assigning sequence reads to genomic features. Bioinformatics 30: 923–930. Lin Q, Yang J, Wang Q, Zhu H, Chen Z, Dao Y, Wang K. Overexpression of the trehalose-6- phosphate phosphatase family gene AtTPPF improves the drought tolerance of Arabidopsis thaliana. BMC Plant Biol.2019 Sep 2;19(1):381. doi: 10.1186 / s12870-019-1986-5. Love, M.I., Anders, S., and Huber, W. (2014). Differential analysis of count data - the DESeq2 package. Ma, L., Hu, L., Fan, J., Amombo, E., Khaldun, A.B.M., Zheng, Y., and Chen, L. (2017). Cotton GhERF38 gene is involved in plant response to salt / drought and ABA. Ecotoxicology 26: 841–854. Ma, Z. et al. (2018). Resequencing a core collection of upland cotton identifies genomic variation and loci influencing fiber quality and yield. Nat Genet 50: 803–813. MacIntyre AM, Meline V, Gorman Z, Augustine SP, Dye CJ, Hamilton CD, Iyer-Pascuzzi AS, Kolomiets MV, McCulloh KA, Allen C. Trehalose increases tomato drought tolerance, induces defenses, and increases resistance to bacterial wilt disease. PLoS One. 2022 Apr 27;17(4):e0266254. doi: 10.1371 / journal.pone.0266254. Mahmood, T., Khalid, S., Abdullah, M., Ahmed, Z., Shah, M.K.N., Ghafoor, A., and Du, X. (2019). Insights into Drought Stress Signaling in Plants and the Molecular Genetic Basis of Cotton Drought Tolerance. Cells 9. Malhotra, S. and Sowdhamini, R. (2014). Interactions among plant transcription factors regulating expression of stress-responsive genes. Bioinform. Biol. Insights 8: 193–198. Malley, R.C.O., Carol, S., Song, L., Galli, M., Ecker, J.R., Malley, R.C.O., Huang, S.C., Song, L., Lewsey, M.G., Bartlett, A., and Nery, J.R. (2016). Cistrome and Epicistrome Features Shape the Regulatory DNA Landscape Resource Cistrome and Epicistrome Features Shape the Regulatory DNA Landscape. Cell 165: 1280–1292. Miranda JA, Avonce N, Suárez R, Thevelein JM, Van Dijck P, Iturriaga G. A bifunctional TPS-TPP enzyme from yeast confers tolerance to multiple and extreme abiotic-stress conditions in transgenic Arabidopsis. Planta.2007 Nov;226(6):1411-21. doi: 10.1007 / s00425-007-0579-y. McLeay, R.C. and Bailey, T.L. (2010). Motif Enrichment Analysis: A unified framework and an evaluation on ChIP data. BMC Bioinformatics 11. Melandri, G., Thorp, K.R., Broeckling, C., Thompson, A.L., Hinze, L., and Pauli, D. (2021). Assessing Drought and Heat Stress-Induced Changes in the Cotton Leaf Metabolome and Their Relationship With Hyperspectral Reflectance. Front. Plant Sci.12: 1–19. Nguyen, L.T., Schmidt, H.A., Von Haeseler, A., and Minh, B.Q. (2015). IQ-TREE: A fast and effective stochastic algorithm for estimating maximum-likelihood phylogenies. Mol. Biol. Evol.32: 268–274. Obata, T. and Fernie, A.R. (2012). The use of metabolomics to dissect plant responses to abiotic stresses. Cell. Mol. Life Sci.69: 3225–3243. Pang, M., McD Stewart, J., and Zhang, J. (2011). A mini-scale hot borate method for the isolation of total RNA from a large number of cotton tissue samples. African J. Biotechnol. 10: 15430–15437. Patro, R., Duggal, G., Love, M.I., Irizarry, R.A., and Kingsford, C. (2017). Salmon provides fast and bias-aware quantification of transcript expression. Nat. Methods 14: 417–419. Paul Shannon, 1, Andrew Markiel, 1, Owen Ozier, 2 Nitin S. Baliga, 1 Jonathan T. Wang, 2 Daniel Ramage, 2, Nada Amin, 2, Benno Schwikowski, 1, 5 and Trey Ideker2, 3, 4, 5, (2003). Cytoscape: A Software Environment for Integrated Models. Genome Res.13: 426. Pego, J. V., Kortstee, A.J., Huijser, C., and Smeekens, S.C.M. (2000). Photosynthesis, sugars and the regulation of gene expression. J. Exp. Bot.51: 407–416. Peri, S., Roberts, S., Kreko, I.R., McHan, L.B., Naron, A., Ram, A., Murphy, R.L., Lyons, E., Gregory, B.D., Devisetty, U.K., and Nelson, A.D.L. (2020). Read Mapping and Transcript Assembly: A Scalable and High-Throughput Workflow for the Processing and Analysis of Ribonucleic Acid Sequencing Data. Front. Genet.10. Ponnu, J., Wahl, V., and Schmid, M. (2011). Trehalose-6-phosphate : connecting plant metabolism and development.2: 1–6. Prieto-dapena, P., Espinosa, M., and Jordano, J. (2005). Functional Interaction between Two Transcription Factors Involved in the Developmental Regulation of a Small Heat Stress Protein Gene Promoter 1 [ w ].139: 1483–1494. Rolland F, Baena-Gonzalez E, Sheen J. Sugar sensing and signaling in plants: conserved and novel mechanisms. Annu Rev Plant Biol.2006;57:675-709. doi: 10.1146 / annurev.arplant.57.032905.105441. Sakuma, Y., Maruyama, K., Qin, F., Osakabe, Y., Shinozaki, K., and Yamaguchi- Shinozaki, K. (2006). Dual function of an Arabidopsis transcription factor DREB2A in water-stress-responsive and heat-stress-responsive gene expression. Proc. Natl. Acad. Sci. U. S. A.103: 18822–18827. Schwacke, R., Ponce-Soto, G.Y., Krause, K., Bolger, A.M., Arsova, B., Hallab, A., Gruden, K., Stitt, M., Bolger, M.E., and Usadel, B. (2019). MapMan4: A Refined Protein Classification and Annotation Framework Applicable to Multi-Omics Data Analysis. Mol. Plant 12: 879–892. Shi, Y., Pang, X., Liu, W., Wang, R., Su, D., Gao, Y., Wu, M., Deng, W., Liu, Y., and Li, Z. (2021). SlZHD17 is involved in the control of chlorophyll and carotenoid metabolism in tomato fruit. Hortic. Res.8: 1–16. Shinozaki, K. and Yamaguchi-Shinozaki, K. (2007). Gene networks involved in drought stress response and tolerance. J. Exp. Bot.58: 221–227. Sinclair, T.R. (2011). Challenges in breeding for yield increase for drought. Trends Plant Sci. 16: 289–293. Singh, D. and Laxmi, A. (2015). Transcriptional regulation of drought response: A tortuous network of transcriptional factors. Front. Plant Sci.6: 1–11. Su, J., Fan, S., Li, L., Wei, H., Wang, C., and Wang, H. (2016). Detection of Favorable QTL Alleles and Candidate Genes for Lint Percentage by GWAS in Chinese Upland Cotton.7: 1–11. Sun, F., Chen, Q., Chen, Q., Jiang, M., Gao, W., and Qu, Y. (2021). Screening of Key Drought Tolerance Indices for Cotton at the Flowering and Boll Setting Stage Using the Dimension Reduction Method. Front. Plant Sci.12: 1–10. Szklarczyk, D., Morris, J.H., Cook, H., Kuhn, M., Wyder, S., Simonovic, M., Santos, A., Doncheva, N.T., Roth, A., Bork, P., Jensen, L.J., and Von Mering, C. (2017). The STRING database in 2017: Quality-controlled protein-protein association networks, made broadly accessible. Nucleic Acids Res.45: D362–D368. Takahashi, F., Kuromori, T., Urano, K., Yamaguchi-Shinozaki, K., and Shinozaki, K. (2020). Drought Stress Responses and Resistance in Plants: From Cellular Responses to Long-Distance Intercellular Communication. Front. Plant Sci.11: 1–14. Tang, J., Yang, X., Xiao, C., Li, J., Chen, Y., Li, R., Li, S., Shiyou, L., and Hu, H. (2020). GDSL lipase occluded stomatal pore 1 is required for wax biosynthesis and stomatal cuticular ledge formation. New Phytol.228: 1880–1896. Tardieu, F., Granier, C., and Muller, B. (2011). Water deficit and growth. Co-ordinating processes without an orchestrator? Curr. Opin. Plant Biol.14: 283–289. Tardieu, F., Parent, B., Caldeira, C.F., and Welcker, C. (2014). Genetic and physiological controls of growth under water deficit. Plant Physiol.164: 1628–1635. Tardieu, F., Simonneau, T., and Muller, B. (2018). The Physiological Basis of Drought Tolerance in Crop Plants: A Scenario-Dependent Probabilistic Approach. Annu. Rev. Plant Biol.69: 733–759. Thimm, O., Bläsing, O., Gibon, Y., Nagel, A., Meyer, S., Krüger, P., Selbig, J., Müller, L.A., Rhee, S.Y., and Stitt, M. (2004). MAPMAN: A user-driven tool to display genomics data sets onto diagrams of metabolic pathways and other biological processes. Plant J.37: 914– 939. Usadel, B., Poree, F., Nagel, A., Lohse, M., Czedik-Eysenberg, A., and Stitt, M. (2009). A guide to using MapMan to visualize and compare Omics data in plants: A case study in the crop species, Maize. Plant, Cell Environ.32: 1211–1229. Wang, F. et al. (2015). GmWRKY27 interacts with GmMYB174 to reduce expression of GmNAC29 for stress tolerance in soybean plants. Plant J.83: 224–236. Welcker, C., Sadok, W., Dignat, G., Renault, M., Salvi, S., Charcosset, A., and Tardieu, F. (2011). A common genetic determinism for sensitivities to soil water deficit and evaporative demand: Meta-analysis of quantitative trait loci and introgression lines of maize. Plant Physiol.157: 718–729. Weltmeier, F. et al. (2009). Expression patterns within the Arabidopsis C / S1 bZIP transcription factor network: Availability of heterodimerization partners controls gene expression during stress response and development. Plant Mol. Biol.69: 107–119. Wu, Y., Llewellyn, D.J., and Dennis, E.S. (2002). A quick and easy method for isolating good- quality RNA from cotton (Gossypium hirsutum L.) tissues. Plant Mol. Biol. Report.20: 213–218. Xiao, S., Jiang, L., Wang, C., and Ow, D.W. (2021). Arabidopsis OXS3 family proteins repress ABA signaling through interactions with AFP1 in the regulation of ABI4 expression. J. Exp. Bot.72: 5721–5734. Xie, Z., Nolan, T.M., Jiang, H., and Yin, Y. (2019). AP2 / ERF transcription factor regulatory networks in hormone and abiotic stress responses in Arabidopsis. Front. Plant Sci.10: 1–17. Zandalinas, S.I., Mittler, R., Balfagón, D., Arbona, V., and Gómez-Cadenas, A. (2018). Plant adaptations to the combination of drought and high temperatures. Physiol. Plant.162: 2–12. Zheng, Y. et al. (2016). iTAK: A Program for Genome-wide Prediction and Classification of Plant Transcription Factors, Transcriptional Regulators, and Protein Kinases. Mol. Plant 9: 1667–1670. Zhu, G. et al. (2018). Rewiring of the Fruit Metabolome in Tomato Breeding. Cell 172: 249- 261.e12. Zhu, T., Liang, C., Meng, Z., Sun, G., Meng, Z., Guo, S., and Zhang, R. (2017). CottonFGD: An integrated functional genomics database for cotton. BMC Plant Biol.17: 1–9. Example 7: Generation of Transgenic Plants The above examples have indicated that the transcription factor GhHSFA6B-D is tightly, and positively, correlated with expression of a lint-yield associated gene, GhIps1-A, under water limiting conditions. In our panel of 21 cotton accessions, those lines harboring higher levels of GhHsfa6b-D and GhIps1-A had higher lint yield. We further identified specific SNPs which are positively associated with lint yield under water limiting conditions. To identify additional HSE permutations that enhance GhHSFA6B binding to the regulatory region of GhIps1-A, microscale thermophoresis (MST) was used. MST captures the difference in DNA bound versus unbound proteins moving through a temperature gradient in a capillary tube. MST is able to quantify the DNA binding affinity of transcription factors at nanomolar scales. Using a MST machine, the binding affinity of in vitro transcribed, translated, and purified Halo-tagged GhHSFA6B to fluorescently labeled oligos containing variants of the GhIps1-A HSE was measured. The known SNP variants (mentioned above) that impact binding were tested and additional variants based on consensus motifs for HSFA6B in DAP-seq data that we have generated were designed. (FIG.19) The improved oligo shows a ~10x improvement in binding relative to the endogenous oligo. This construct is being used in the design of a transgenic construct for cotton transformation. In an effort to improve this increased yield, we will introduce a GhIps1-A with an altered HSE that incorporates the non-native nucleotides described above. Importantly, introducing a transgene that was still under the appropriate contextual regulatory control (GhHSFA6B-D during drought) will prevent pleotropic effects of over-expressing GhIps1-A or GhHSFA6B-D under non-stress conditions. Additionally, further adjustment to GhHSFA6B-D binding will be performed by testing a range of non-native GhIps1-A HSE permutations (referred to as HSEpm::GhIps1-A). This experiment will identify additional HSE permutations that enhance HSFA6B binding and generate transgenic cotton lines harboring different variations of these HSEpm::GhIps1-A constructs. These HSEpm::GhIps1-A constructs will show lint yield improvements under water limited conditions in a background that retains the endogenous GhIps1-A. Accordingly, these lines would be over-expressing HSEpm::GhIps1-A under very specific conditions, i.e. drought. The two GhIps1-A HSE permutations (HSEpm::GhIps1-A ) that display the highest affinity for GhHSFA6B will be transformed into Coker312. Additional permutations that bind with greater affinity than the two SNP permutations may be found. Following transformation, the transformants will be grown over multiple generations to homozygosity. T1 or T2 plants will be tested for performance under water limited conditions. Upland cotton (Gossypium hirsutum L.) variety Coker312 will be used for genetic transformation and whole plant somatic regeneration. A modified pCAMBIA plasmid harboring HSEpm::GhIps1-A will be transformed into A. tumefaciens strain EHA105. Cotton seeds will be surface sterilized and cultured in germination bottles on germination media and kept in the dark for 7 days as described in (5). The hypocotyls will be excised from the 7-days-old, aseptic, etiolated seedlings, cut into 5–7 mm pieces, and used as explants for Agrobacterium cocultivation. Agrobacterium suspension for cocultivation will be prepared as described in (6). A total of 100 independent explants will be cocultured in 30 ml of Agrobacterium suspension for 5 minutes and then transferred to sterilized filter paper to blot dry. After drying, explants are transferred to sterilized filter paper on the surface of co-culture medium for ~72 h. Transformed explants will be transferred to selection medium containing hygromycin (50 mg / L) (6). After 4 weeks on selection media, hygromycin resistant calli (transformed) be moved to new media every 30 days and monitored for polar structure formation with high-resolution imaging. Once embryogenic calli (EC) are observed, these will be transferred to differentiation media as described in (6) for shoot and root development. Rooted embryos will be transferred to magenta boxes containing rooting media with 0.5 mg / L IBA. Regenerated plantlets at the 2-3 leaves stage with good root formation will be transferred to soil in small pots and kept in mist chamber for hardening and then transferred to green house in 3 gallons pots. Fifteen to twenty plants from up to 10 independent transformation events can be generated this way. References: 1. Yuan, D., Grover, C.E., Hu, G., Pan, M., Miller, E.R., Conover, J.L., Hunt, S.P., Udall, J.A., and Wendel, J.F. (2021). Parallel and Intertwining Threads of Domestication in Allopolyploid Cotton. Adv. Sci.8: 1–17. 2. Mueller, A.M., Breitsprecher, D., Duhr, S., Baaske, P., Schubert, T., Längst, G. (2017). MicroScale Thermophoresis: A Rapid and Precise Method to Quantify Protein–Nucleic Acid Interactions in Solution. In: Kaufmann, M., Klinger, C., Savelsbergh, A. (eds) Functional Genomics. Methods in Molecular Biology, vol 1654. Humana Press, New York, NY. https: / / doi.org / 10.1007 / 978-1-4939-7231-9_10 3. J. Y. Li et al., Multi-omics analyses reveal epigenomics basis for cotton somatic embryogenesis through successive regeneration acclimation process. Plant Biotechnol J 17, 435-450 (2019). 4. F. A. Ran et al., Genome engineering using the CRISPR-Cas9 system. Nat Protoc 8, 2281- 2308 (2013). 5. L. Wen et al., Transcriptomic profiles of non-embryogenic and embryogenic callus cells in a highly regenerative upland cotton line (Gossypium hirsutum L.). BMC Dev Biol 20, 25 (2020). 6. S. X. Jin, G. Z. Liu, H. G. Zhu, X. Y. Yang, X. L. Zhang, Transformation of Upland Cotton (Gossypium hirsutum L.) with gfp Gene as a Visual Marker (vol 11, pg 910, 2012). J Integr Agr 11, 1433-1433 (2012). While certain features of the invention have been described herein, many modifications, substitutions, changes, and equivalents will now occur to those of ordinary skill in the art. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the scope of the invention.

Claims

What is claimed is:

1. A method for identifying a drought resistant plant comprising, a) obtaining a nucleic acid sample from said plant; and b) determining whether said sample contains at least one SNP containing nucleic acid in a DREB2A-A or HSFA6B-D binding site of IPS1-A, the presence of said SNP containing nucleic acid conferring at least one desirable plant characteristic selected from increased drought resistance, heat tolerance and lint yield in said plant.

2. The method of claim 1, wherein the SNP is present in nucleotide 19 of SEQ ID NO:

33.

3. The method of claim 2, wherein the SNP is a C allele, G allele, or A allele.

4. The method of claim 2 or 3 wherein the C allele is present and the presence of the C allele is associated with enhanced drought resistance when compared to plants lacking a C allele.

5. A method for identifying a drought resistant plant comprising, a) obtaining a nucleic acid sample from said plant; and b) determining whether said sample contains at least one SNP containing nucleic acid provided in Figure 8 or Figure 13E, the presence of said SNP containing nucleic acid conferring at least one desirable plant characteristic selected from increased drought resistance, heat tolerance and lint yield in said plant.

6. The method of any one of the preceding claims, wherein the target nucleic acid is amplified prior to detection.

7. The method of any one of the preceding claims, wherein the step of detecting the presence of said SNP is performed using a process selected from the group consisting of detection of specific hybridization, measurement of allele size, restriction fragment length polymorphism analysis, allele-specific hybridization analysis, single base primer extension reaction, and sequencing of an amplified polynucleotide.

8. The method as claimed in any of the preceding claims, wherein in the target nucleic acid is DNA.

9. The method of any of claims 5-8, wherein the SNP is in a DREB2A-A or HSFA6B-D binding site of IPS1-A.

10. The method of claim 9, wherein the SNP is present in nucleotide 19 of SEQ ID NO: 33 wherein the presence of the C allele is associated with enhanced drought resistance when compared to plants lacking a C allele.

11. The method of any one of the preceding claims, wherein the plant is a cotton plant selected from Gossypium hirsutum, Gossypium barbadense, Gossypium arboretum, and Gossypium raimondii.

12. The method of claim 11, wherein the cotton plant is Gossypium hirsutum.

13. A multiplex SNP panel comprising nucleic acids informative of the presence of increased drought resistance, wherein said panel contains nucleic acids comprising at least one SNP in a DREB2A-A or HSFA6B-D binding site of IPS1-A.

14. The multiplex SNP panel of claim 13, wherein the SNP is present in nucleotide 19 of SEQ ID NO:

33.

15. The multiplex SNP panel of claim 14, wherein the SNP is a C allele, G allele, or an A allele.

16. A multiplex SNP panel comprising nucleic acids informative of the presence of increased drought resistance, wherein said panel contains nucleic acids comprising SNPs as provided in Figure 8 or Figure 13E.

17. A vector comprising nucleic acids encoding a IPS1-Aprotein having at least one SNP in the DREB2A-A or HSFA6B binding site, the SNP comprising an allele that enhances drought resistance when compared to proteins without the allele.

18. The vector of claim 17, wherein the SNP is at least one SNP as provided in Figure 8, or 13E.

19. The vector of claim 18, wherein the SNP is present in nucleotide 19 of SEQ ID NO:

33.

20. The vector of claim 19, wherein the SNP is a C allele, G allele, or an A allele.

21. The vector of any one of claims 17-20, wherein the presence of the C allele is associated with enhanced drought resistance when compared to plants lacking the C allele.

22. A host cell comprising the vector of any one of claims 17-21.

23. A solid support comprising the drought resistant SNP containing nucleic acid of claim 17.

24. A recombinant plant comprising nucleic acids containing at least one SNP as provided in Figure 8, or Figure 13E.

25. A recombinant plant comprising nucleic acids containing at least one SNP in a DREB2A- A or HSFA6B-D binding site of IPS1-A.

26. The recombinant plant of claim 24 or 26, wherein the SNP is present in nucleotide 19 of SEQ ID NO: 33.

27. The recombinant plant of claim 26, wherein the SNP is a C allele, G allele, or an A allele.

28. The recombinant plant of claims 27, wherein the presence of the C allele is associated with enhanced drought resistance when compared to plants lacking the C allele.

29. The recombinant plant of anyone of claims 24-28, wherein the plant is a cotton plant selected from Gossypium hirsutum, Gossypium barbadense, Gossypium arboretum, and Gossypium raimondii.

30. The recombinant plant of claim 29, wherein the cotton plant is Gossypium hirsutum.

31. A method of increasing drought resistance in a cotton plant as compared to the wild type, the method comprising transforming a plant cell with the vector of any one of claims 17- 21.

32. A method of producing a transgenic plant with increased drought resistance as compared to the wild type, the method comprising transforming a plant cell with the vector of any one of claims 17-21 and regenerating a plant from the plant cell.

33. A method of producing a transgenic plant with an increased drought resistance expression profile as compared to the wild type, the method comprising crossing the recombinant cotton plant of any one of claims 24-30 with a wild type cotton plant.

34. A method of producing a plant with increased drought resistance as compared to the wild type, the method comprising crossing a first cotton plant with at least one SNP as provided in Figure 8, or Figure 13E with a second cotton plant.

35. A method of producing a plant with increased drought resistance as compared to the wild type, the method comprising crossing a first cotton plant with at least one SNP in a DREB2A-A or HSFA6B-D binding site of IPS1-A.

36. The method of claim 35, wherein the SNP is present in nucleotide 19 of SEQ ID NO:

33.

37. The method of claim 36, wherein the SNP is a C allele, G allele, or A allele.

38. The method of claim 37, wherein the C allele is present and the presence of the C allele is associated with enhanced drought resistance when compared to plants lacking a C allele.

39. The method of claim 33-38, further comprising detecting in the plant at least one of the SNPs.

40. The method of any one of claims 31-39, wherein the second plant is selected from the first cotton plant, a wild-type cotton plant, or a recombinant cotton plant.

41. The method of any one of claims 31-40, further comprising detecting at least one drought-resistant phenotype in said plant.

42. The method of claim 41, wherein the drought resistant phenotype is selected from increased lint yield, increased length uniformity (UI), scaled photochemical reflective index (sPRI), carotenoid reflectance index (CRI), and water index / normalized difference vegetation index (WI / NDVI).

43. A seed, plant, or plant cell produced by the method of any one of claims 31-41.

44. The seed, plant, or plant cell of claim 43, wherein the at least one SNP as provided in Figure 8, or Figure 13E is present in the seed, plant, or plant cell.

45. The seed, plant, or plant cell of claim 43, wherein the SNP is present in a DREB2A-A or HSFA6B-D binding site of IPS1-A.

46. The seed, plant, or plant cell of claim 45, wherein the at least one SNP is present in nucleotide 19 of SEQ ID NO:

33.

47. The seed, plant, or plant cell of claim 47, wherein the presence of the C allele is associated with enhanced drought resistance when compared to plants lacking the C allele.

Citation Information

Patent Citations

  • Cotton polymorphisms and methods of genotyping

    US20140255922A1