Genetic markers for plant height

Genetic markers for Cannabis plant height enable precise breeding, addressing variability issues and improving cannabinoid production and harvesting efficiency through marker-assisted selection.

WO2025184562A1PCT designated stage Publication Date: 2025-09-04PHYLOS BIOSCIENCE INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/017931
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-01
Filing Date
2025-02-28
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing Cannabis breeding methods lack efficient genetic markers for selecting plants with specific height traits, leading to variability in plant height and potential shading issues affecting cannabinoid production and mechanical harvesting efficiency.

Method used

Development of genetic markers, particularly SNPs, for vegetative height, harvest height, and stretch in Cannabis, allowing for marker-assisted selection (MAS) to achieve uniform plant height and desired growth characteristics.

Benefits of technology

Enables precise breeding for uniform plant height, enhancing cannabinoid production uniformity and mechanical harvesting efficiency by selecting plants with modified height traits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000054_0001
    Figure IMGF000054_0001
  • Figure IMGF000055_0001
    Figure IMGF000055_0001
  • Figure IMGF000066_0001
    Figure IMGF000066_0001
Patent Text Reader

Abstract

Provided are genetic markers associated with modified plant height in Cannabis. The genetic markers are useful, for example, for identifying, selecting, and / or breeding Cannabis plants having modified plant height. Also provided are genes conferring modified plant height in Cannabis.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] GENETIC MARKERS FOR PLANT HEIGHT

[0002] CROSS REFERENCE TO RELATED APPLICATIONS

[0003] This application claims priority to U.S. Provisional Application No. 63 / 560,285, filed March 1, 2024, which is incorporated by reference in its entirety.

[0004] FIELD

[0005] The present disclosure relates to genetic markers and genes associated with plant height, including vegetative height, harvest height, and / or stretch, and methods of use thereof.

[0006] INCORPORATION OF ELECTRONIC SEQUENCE LISTING

[0007] The Sequence Listing is submitted as an XML file named “Sequence.xml,” created on February 27, 2025, 258,319 bytes, which is incorporated by reference herein.

[0008] BACKGROUND

[0009] Cannabis sativa as a species varies considerably for plant height, from <50 cm for day-neutral (autoflowering) varieties to >200 cm for some fiber hemp cultivars. Plant height in Cannabis can be broken down into three components that are important for breeding for different production systems: (1 ) Vegetative Height (plant height at the end of the vegetative phase before lights are switched from 18 hours light to 12 hours light) (2) Harvest Height (plant height of mature plants), and (3) Stretch (the difference between Harvest Height and Vegetative Height). Different Cannabis production systems have different height requirements. Tiered growth room systems require short plants. Greenhouse systems require different heights depending on whether the production system makes use of trellis and / or if there is a desire for a more open top of the canopy, which is a result of stretch. Trellis-based greenhouse systems have a limit on the amount of stretch due to the need to anticipate trellis height. Outdoor systems require sturdy plants that do not fall over due to wind damage.

[0010] Genetic markers for plant height (e.g., vegetative height, harvest height, and stretch) allow breeders to perform marker assisted selection (MAS) to develop varieties for these various production systems. In addition, these markers allow breeders to select for uniformity in plant height. Uniformity of height is an important trait for mechanical harvesting as well as ensuring all plants get adequate light; the amount of light determines cannabinoid production and shorter plants could be shaded by taller plants, thus resulting in variable cannabinoid concentrations. SUMMARY

[0011] Disclosed are methods for selecting a Cannabis plant having modified plant height, including: (i) analyzing one or more genetic markers in a nucleic acid sample from the Cannabis plant or its germplasm; (ii) detecting one or more genetic markers that indicate modified plant height; and (iii) selecting the Cannabis plant, thereby selecting the Cannabis plant having modified plant height. Also disclosed are methods for producing one or more Cannabis plants having modified plant height, including: (i) analyzing one or more genetic markers in a nucleic acid sample from a Cannabis plant or its germplasm; (ii) detecting one or more genetic markers that indicate modified plant height, (iii) crossing the Cannabis plant comprising the one or more genetic markers indicating modified plant height, and (iv) obtaining one or more progeny plants comprising the one or more genetic markers indicating modified plant height, wherein the one or more progeny plants have modified plant height relative to a control.

[0012] Modified plant height refers to an increase or decrease in plant height, or increase in uniformity of plant height, relative to a control. Measurements of plant height can include, for example, vegetative height, harvest height, and stretch. In some aspects, the modified plant height is an increase in vegetative height, har vest height, and / or stretch relative to a control. In some aspects the control is a Cannabis plant without the one or more markers indicating modified plant height. In some aspects, the control is a parent or sibling of the Cannabis plant having modified plant height. In some aspects, the control is the Abacus Cannabis variety.

[0013] In some aspects, the one or more genetic markers include a polymorphism relative to a reference genome in a plant height haplotype. In some examples, the plant height haplotype includes a region on chromosome 4 between position 0.02 Mbp and 1.32 Mbp (when using the Abacus Cannabis genome version Csat_AbacusV2, NCBI assembly accession GCA_025232715.1 as a reference).

[0014] In some aspects, the genetic markers that indicate modified plant height include one or more genetic markers disclosed herein, for example, a marker disclosed in any of Tables 2-10. In some aspects, the genetic markers that indicate modified plant height include one or more genetic markers disclosed in Table 8. The SNP markers disclosed herein are described in several ways, including by reference to a marker name, by reference to a particular chromosomal position (in reference to the Abacus Cannabis reference genome Csat_AbacusV2, NCBI assembly accession GCA_025232715.1), or by reference to a corresponding SEQ ID NO. Cross reference of these descriptions in the disclosed Tables allows a person of ordinary skill in the art to readily and accurately identify a referenced SNP.

[0015] In some examples, the genetic markers that indicate modified plant height include a polymorphism at position 51 of one or more of: SEQ ID NOs: 1-205. In a non-limiting example, the genetic marker that indicates modified plant height includes a polymorphism at position 51 of SEQ ID NO: 44. In some examples, the methods include detecting one or more of SNP markers: 142603.12546539, 142603.12519907, Cannabis.vl_scf542.169826_101, 142603.12500071, 142603.12492164, Cannabis.vl_scf542.132451_99, Cannabis.vl_scf542.122780.100, 142603.12456339, 142603.12451790, 142603.12439239, 142603.12351777, 142603.12346103, 142603.12333141, 142603.12209629, 142603.12122639, 142603.12107817, 142603.12087936,

[0016] 142603.12081998, 142603.12064923, 142603.12038515, 142603.12001291, 142603.11997059, 142603.11989491, 142603.11951907, 142603.11459458, 142603.12546539, 142603.12519907, or 142603.11711278. In some examples, the methods include detecting one or more of SNP markers: 142603.12439239, 142603.12333141, 142603.12107817, 142708.323806, 142708.353959, or

[0017] 142603.12081998, In some examples, the methods include detecting one or more of SNP markers: 142603.12439239, 142603.12333141, 142603.12107817, 142708.323806, or 142708.353959, In a non-limiting example, the methods include detecting SNP marker 142603.12107817 and / or SNP marker

[0018] 142603.12081998, In a non-limiting example, the methods include detecting SNP marker 142603.12107817, In some examples, the genetic markers are genetically linked to a modified plant height trait locus.

[0019] The methods can include analyzing or detecting a plur ality of genetic markers, for example, analyzing or detecting at least two, at least three, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, or more, genetic markers. Analyzing can be performed using any suitable molecular biology technique for analyzing nucleic acids, for example, PCR, quantitative PCR (qPCR), and / or sequencing. Detection can be performed using any suitable method of detecting nucleic acids, including techniques that utilize oligonucleotide primers or probes.

[0020] Cannabis plants having modified plant height can be selected for further analysis, propagation, crossing, or to make a product. In some examples, a Cannabis plant having modified plant height is crossed, and one or more progeny plants having modified plant height are produced / obtained. Crossing includes, for example, selfing, sibling crossing, outcrossing, or backcrossing. In some examples, crossing includes marker-assisted selection (MAS), for example, for at least one, two, three, four, five, or more generations.

[0021] Further disclosed are alleles of WRKY21, LBD16, SAUR51, and / or TCS1 conferring modified plant height. Methods of introducing such alleles, or genetic mutations in WRKY21, LBD16, SAUR51, and / or TCS 1 , that confer a plant height phenotype in Cannabis are provided.

[0022] Also disclosed are Cannabis plants identified, selected, or produced by a method disclosed herein, including a seed, plant part, tissue culture, or protoplast thereof. Further disclosed are Cannabis products made from a Cannabis plant identified, selected, or produced by a method disclosed herein. The Cannabis product as used herein is a manufactured product, and is not or otherwise excludes, naturally occurring products. In some examples, the product is a kief, hashish, bubble hash, an edible product, solvent reduced oil, sludge, e-juice, or tincture.

[0023] The foregoing and other objects, features, and advantages of the disclosure will become more apparent from the following detailed description.

[0024] SEQUENCES

[0025] The nucleic and amino acid sequences listed in the accompanying sequence listing are shown using standard letter abbreviations for nucleotide bases, and single letter code for amino acids, as defined in 37 C.F.R. 1.822. Only one strand of each nucleic acid sequence is shown, but the complementary strand is understood as included by any reference to the displayed strand. In the accompanying sequence listing:

[0026] SEQ ID NOs: 1-205 are genomic DNA sequences flanking exemplary SNP markers.

[0027] SEQ ID NOs: 206-227 are exemplary oligonucleotide primers.

[0028] SEQ ID NOs: 228 and 229 are exemplary coding sequences of WRKY21.

[0029] SEQ ID NOs: 230 and 231 are exemplary amino acid sequences of WRKY21.

[0030] SEQ ID NO: 232 is genomic DNA sequence flanking a putative causative WRKY21 SNP.

[0031] SEQ ID NOs: 233-236 are exemplary coding sequences of LBD16.

[0032] SEQ ID NOs: 237-240 are exemplary amino acid sequences of LBD16.

[0033] SEQ ID NOs: 241-244 are genomic DNA sequences flanking putative causative LBD16 nucleotide polymorphisms.

[0034] SEQ ID NOs: 245-246 are genomic DNA sequences immediately upstream of the start codon of SAUR51.

[0035] SEQ ID NOs: 247-249 are genomic DNA sequences flanking putative causative SAUR51 SNPs.

[0036] SEQ ID NOs: 250-251 are exemplary coding sequences of TCS1.

[0037] SEQ ID NOs: 252-253 are exemplary amino acid sequences of TCS1 .

[0038] SEQ ID NOs: 254 is genomic DNA sequence flanking a putative causative TCS1 SNP.

[0039] SEQ ID NO: 255 is an exemplary coding sequence of SAUR51.

[0040] SEQ ID NO: 256 is an exemplary amino acid sequence of S AUR51.

[0041] DETAILED DESCRIPTION

[0042] I. Introduction

[0043] Cannabis has long been used for drug and industrial puiposes, fiber (hemp), for seed and seed oils, for medicinal purposes, and for recreational puiposes. Industrial hemp products arc made from Cannabis plants selected to produce an abundance of fiber. Some Cannabis varieties have been bred to produce minimal levels of THC, the principal psychoactive constituent responsible for the psychoactivity associated with marijuana. Marijuana has historically consisted of the dried flowers of Cannabis plants selectively bred to produce high levels of THC and other psychoactive cannabinoids. As a drug it usually comes in the form of dried flower buds (marijuana), resin (hashish), or various extracts collectively known as hashish oil.

[0044] Cannabis is an annual, dioecious, flowering herb. The leaves are palmately compound or digitate, with serrate leaflets. Cannabis normally has imperfect flowers, with staminate “male” and pistillate “female” flowers occurring on separate plants. It is not unusual, however, for individual plants to separately bear both male and female flowers (z.e., have monoecious plants). Although monoecious plants are often referred to as “hermaphrodites,” true hermaphrodites (which are less common in Cannabis) bear staminate and pistillate structures on individual flowers, whereas monoecious plants bear male and female flowers at different locations on the same plant.

[0045] The life cycle of Cannabis varies with each variety but can be generally summarized into germination, vegetative growth, and reproductive stages. Because of heavy breeding and selection by humans, most Cannabis seeds have lost dormancy mechanisms and do not require any pre-treatments or winterization to induce germination. Seed placed in viable growth conditions are expected to germinate in about 3 to 7 days. The first true leaves of a Cannabis plant contain a single leaflet, with subsequent leaves developing in opposite formation with increasing number of leaflets. Leaflets can be narrow or broad depending on the morphology of the plant grown. Cannabis plants are normally allowed to grow vegetatively for the first 4 to 8 weeks. During this period, the plant responds to increasing light with faster and faster growth. Under ideal conditions, Cannabis plants can grow up to 2.5 inches a day and are capable of reaching heights of up to 20 feet. Indoor growth pruning techniques tend to limit Cannabis size through careful pruning of apical or side shoots.

[0046] Cannabis is diploid, having a chromosome complement of 2n=20, although polyploid individuals have been artificially produced. The first genome sequence of Cannabis, which is estimated to be 820 Mb in size, was published in 2011 by a team of Canadian scientists (Bakel et al., “The draft genome and transcriptome of Cannabis sativa” Genome Biology 12:R102).

[0047] All known varieties of Cannabis are wind-pollinated and the fruit is an achene. Most varieties of Cannabis are short day plants, with the possible exception of C. sativa subsp. Sativa var. spontanea (=C. ruderalis), which is commonly described as “auto-flowering” and may be day-neutral.

[0048] The genus Cannabis was formerly placed in the Nettle (Urticaceae) or Mulberry (Moraceae) family, and later, along with the Humulus genus (hops), in a separate family, the Hemp family (Cannabaccac sensu stricto). Recent phylogenetic studies based on cpDNA restriction site analysis and gene sequencing strongly suggest that the Cannabaceae sensu stricto arose from within the former Celtidaceae family, and that the two families should be merged to form a single monophyletic family, the Cannabaceae sensu lato.

[0049] Cannabis plants produce a variety of secondary metabolites, including cannabinoids, terpenoids, and other compounds, which are often secreted by glandular- trichomes that occur most abundantly on the floral calyxes and bracts of female plants. Cannabinoids are the most studied group of secondary metabolites in Cannabis. Most exist in two forms, as acids and in neutral (decarboxylated) forms. The acid form is designated by an “A” at the end of its acronym (z.e. THCA). The phytocannabinoids are synthesized in the plant as acid forms, and while some decarboxylation does occur in the plant, it increases significantly post-harvest and the kinetics increase at high temperatures (Sanchez and Verpoorte 2008). The biologically active forms for human consumption are the neutral forms. Decarboxylation is usually achieved by thorough drying of the plant material followed by heating it, often by either combustion, vaporization, or heating or baking in an oven.

[0050] Cannabinoids found in Cannabis plants include, but are not limited to, A9-Tetrahydrocannabinol (A9-THC), A8-Tetrahydrocannabinol (A8-THC), Cannabichromene (CBC), Cannabicyclol (CBL), Cannabidiol (CBD), Cannabielsoin (CBE), Cannabigerol (CBG), Cannabinidiol (CBND), Cannabinol (CBN), Cannabitriol (CBT), and their- propyl homologs, including, but are not limited to cannabidivarin (CBDV), A9-Tetrahydrocannabivarin (THCV), cannabichromevarin (CBCV), and cannabigerovarin (CBGV). See, Holley et al. (Constituents of Cannabis sativa L. XI Cannabidiol and cannabichromene in samples of known geographical origin, J. Pharm. Sci. 64:892-894, 1975) and De Zeeuw et al. (Cannabinoids with a propyl side chain in Cannabis, Occurrence and chromatographic behavior. Science 175:778-779). Non-THC cannabinoids can be collectively referred to as “CBs”, wherein CBs can be one of THCV, CBDV, CBGV, CBCV, CBD, CBC, CBE, CBG, CBN, CBND, and CBT cannabinoids.

[0051] Terpenes are primarily produced in glandular trichomes of female inflorescences (Livingston et al., "Cannabis glandular trichomes alter morphology and metabolite content during flower maturation," The Plant Journal 101.1 (2020): 37-56). Besides affecting aroma and fragrance, terpenes may have a synergic effect with cannabinoids (Sommano et al., "The cannabis terpenes," Molecules 25.24 (2020): 5792), and have been attributed medicinal properties (Maggini et al., "An Optimized Terpene Profile for a New Medical Cannabis Oil," Pharmaceutics 14.2 (2022): 298). Two main groups of terpenes in Cannabis are the monoterpenes and sesquiterpenes, which are produced in the methylerythritol phosphate pathway (MEP) and mevalonic acid pathway (MEV), respectively (Booth et al., "Terpene synthases from Cannabis sativa," Pios one 12.3 (2017): e0173911). Monoterpenes have a ten-carbon isoprenoid precursor, geranyl diphosphate (GPP). Sesquiterpenes have a fifteen-carbon isoprenoid precursor, farncsyl diphosphate (FPP). GPP and FPP arc converted to different monoterpenes and sesquiterpenes, respectively, by terpene synthases (TPS; Booth et al., "Terpenes in Cannabis sativa-From plant genome to humans," Plant Science 284 (2019): 67-72).

[0052] Here, single nucleotide polymorphism (SNP) markers associated with Cannabis plant height are identified and described to allow selective breeding of Cannabis plants with appropriate height traits for target growing systems.

[0053] II. Summary of Terms

[0054] Unless otherwise noted, technical terms are used according to conventional usage. Definitions of many common terms in molecular- biology may be found in Krebs et al. (eds.), Lewin’s genes XII, published by Jones & Bartlett Learning, 2017. As used herein, the singular forms “a,” “an,” and “the,” refer to both the singular as well as plural, unless the context clearly indicates otherwise. For example, the term “a plant” includes singular or plural plants and can be considered equivalent to the phrase “at least one plant.” As used herein, the term “comprises” means “includes.” It is further to be understood that any and all base sizes or amino acid sizes, and all molecular weight or molecular mass values, given for nucleic acids or polypeptides are approximate, and are provided for descriptive purposes, unless otherwise indicated. Although many methods and materials similar- or equivalent to those described herein can be used, particular suitable methods and materials are described herein. In case of conflict, the present specification, including explanations of terms, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting. To facilitate review of this disclosure, the following explanations of terms are provided:

[0055] The term “Abacus” as used herein refers to the Cannabis sativa Abacus variety or reference genome known as the Abacus reference genome version Csat_AbacusV2 (NCBI assembly accession GCA_025232715.1, incorporated by reference herein), which is also sometimes also referred to as CsaAba2.

[0056] The term “about” refers to a range of 5% of the referenced value unless otherwise indicated. For example, about 100 refers to a range of 95 to 105.

[0057] The term “alternative nucleotide call” is a nucleotide polymorphism relative to a reference nucleotide for a SNP marker that is significantly associated with a desired phenotype (e.g., plant height).

[0058] The term “backcrossing” or “to backcross” refers to a process in which a breeder repeatedly crosses hybrid progeny, for example a first generation hybrid (Fl ), back to one of the parents of the hybrid progeny. Backcrossing can be used to introduce one or more single locus conversions from one genetic background into another.

[0059] The term “beneficial” as used herein refers to a genetic clement (e.g., gene, allele, or polymorphism) conferring or associated with a desired trait (e.g., plant height). In some examples, a “beneficial polymorphism” or “beneficial allele” refers to a polymorphism or allele associated with a desired trait (e.g., plant height). In some aspects, the desired trait is increased vegetative height, harvest height, or stretch. In some aspects, the desired trait is decreased vegetative height, harvest height, or stretch.

[0060] The term “Cannabis” refers to plants of the genus Cannabis, including Cannabis sativa, Cannabis indica, and Cannabis ruderalis.

[0061] The term “cell” refers to a prokaryotic or eukaryotic cell, including plant cells, capable of replicating DNA, transcribing RNA, translating polypeptides, and secreting proteins.

[0062] The term “coding sequence” (CDS) refers to a DNA sequence which codes for a specific amino acid sequence. “Regulatory sequences” refer to nucleotide sequences located upstream (5' non-coding sequences), within, or downstream (3’ non-coding sequences) of a coding sequence, and which influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences may include, but are not limited to, promoters, translation leader sequences, introns, and polyadenylation recognition sequences.

[0063] The term “control” refers to a reference standard. A control can be a negative or positive control. In some examples, the control is a historical control or a known reference value (or range of values). A practitioner can select a suitable control based on the teachings provided herein. Non-limiting examples of suitable controls include a Cannabis plant not including a marker or gene associated with plant height disclosed herein. In some aspects, the control is measurements from the Abacus Cannabis variety or a parent plant or sibling.

[0064] The term “cross” or “crossing” refer to the process by which the pollen of one flower on one plant is applied (artificially or naturally) to the ovule (stigma) of a flower on another plant (or the same plant when selfing). Exemplary types of crosses include selfing, backcrossing, outcrossing, and sibling crossing.

[0065] The term “cultivar” means a group of similar plants that by structural features and performance (e.g., morphological and physiological characteristics) can be identified from other varieties within the same species. Furthermore, the term “cultivar” variously refers to a variety, strain or race of plant that has been produced by horticultural or agronomic techniques and is not normally found in wild populations. The terms cultivar, variety, strain, plant and race are often used interchangeably by plant breeders, agronomists and farmers.

[0066] The term “detect” or “detecting” refers to any method for determining the presence of a nucleic acid. Methods of detecting nucleic acid polymorphisms, for example, have been described and can include amplification of a target polynucleotide (e.g., by PCR) and / or detection by a probe (e.g., hybridization assays). PCR uses a particular amplification primer pair that specifically hybridizes to a target polynucleotide and produces an amplification product (the amplicon). Primers can be designed such that the amplicon can contain a nucleic acid polymorphism of interest. Methods for designing PCR primers and PCR conditions have been described, for example, in Sambrook et al. (2014) Molecular Cloning: A Laboratory Manual (Fourth Edition, Cold Spring Harbor Laboratory Press, Plainview, N.Y.).

[0067] The term “expression” or “gene expression” relates to the process by which the coded information of a nucleic acid transcriptional unit (including, e.g., genomic DNA) is converted into an operational, non-operational, or structural part of a cell, often including the synthesis of a protein. Gene expression can be influenced by external signals; for example, exposure of a cell, tissue, or organism to an agent that increases or decreases gene expression. Expression of a gene can also be regulated anywhere in the pathway from DNA to RNA to protein. Regulation of gene expression occurs, for example, through controls acting on transcription, translation, RNA transport and processing, degradation of intermediary molecules such as mRNA, or through activation, inactivation, compartmentalization, or degradation of specific protein molecules after they have been made, or by combinations thereof. Gene expression can be measured at the RNA level or the protein level by any suitable, known method, including, without limitation, Northern blot, RT-PCR, Western blot, or in vitro, in situ, or in vivo protein activity assay(s). Elevated levels refer to higher than average levels of gene expression in comparison to a reference, e.g., a control plant.

[0068] The term "expression cassette" refers to a discrete nucleic acid fragment into which a nucleic acid sequence or fragment can be moved, typically for expression in a host cell. In some examples, an expression cassette is included on a vector, such as an expression vector.

[0069] The term “functional” as used herein refers to DNA or amino acid sequences which are of sufficient size and sequence to have the desired function (i.e. the ability to cause expression of a gene resulting in gene activity expected of the gene found in a reference genome, e.g., the Abacus reference genome).

[0070] The term “gene” or “allele” refer to a nucleic acid fragment that expresses a specific protein, including regulatory sequences preceding (5’ non-coding sequences) and following (3’ non-coding sequences) the coding sequence. “Native gene” refers to a gene as found in nature with its own regulatory sequences. “Chimeric gene” or “recombinant expression construct,” which are used interchangeably, are not naturally occurring molecules and refer to any artificial gene chimera, such as a cDNA sequence or gene operably linked to a non-native regulatory sequence. Accordingly, a chimeric gene may comprise regulatory sequences and coding sequences that are derived from different sources, or regulatory sequences and coding sequences derived from the same source, but arranged in a manner different than that found in nature. “Endogenous gene” refers to a native gene in its natural location in the genome of an organism. A “heterologous” gene or allele refers to a gene or allele that is not naturally found in the host, but is artificially introduced (e.g., by genetic engineering or selective plant breeding). Heterologous genes can comprise genes inserted into a non-native host, or chimeric genes.

[0071] The term "genetic modification” as used herein refers to a change from the wild-type or reference sequence of one or more nucleic acid molecules. Genetic modifications or alterations include without limitation, base pair substitutions, additions, or deletions of at least one nucleotide from a nucleic acid molecule of known sequence.

[0072] The term “genome” as it applies to plant cells encompasses not only chromosomal DNA found within the nucleus, but organelle DNA found within subcellular components (e.g.. mitochondrial, plastid) of the cell.

[0073] The term “genotype” refers to the genetic makeup of an individual cell, cell culture, tissue, organism (e.g., a plant), or group of organisms. A “beneficial genotype” refers to a desired genotype, such as a genotype associated with a desired height trait, for example, increased vegetative height, harvest height, or stretch. Conversely, a “detrimental genotype” is a genotype that is not associated with a desired height trait (e.g., not associated with increased vegetative height, harvest height, or stretch). A genotype may refer to a particular genetic marker (e.g., a polymorphism), such as a marker associated with plant height.

[0074] The term “germplasm” refers to genetic material of or from an individual (e.g., a plant), a group of individuals (e.g., a plant line, variety, or family), or a clone derived from a line, variety, species, or culture. The germplasm can be part of an organism or cell, or can be separate from the organism or cell. In general, germplasm provides genetic material with a specific molecular makeup that provides a physical foundation for some or all of the hereditary qualities of an organism or cell culture. As used herein, germplasm includes cells, seed or tissues from which new plants can be grown, as well as plant parts, such as leaves, stems, pollen, or cells that can be cultured into a whole plant.

[0075] The term “haplotype” refers to the genotype of a plant at a plurality of genetic loci, e.g., a combination of alleles or markers. Haplotype can refer to sequence of polymorphisms at a particular locus, such as a single marker locus, or sequence polymorphisms at multiple loci along a chromosomal segment in a given genome. As used herein, a haplotype can be a nucleic acid region spanning two markers.

[0076] A plant is "homozygous" if the individual has only one type of allele at a given locus (e.g., a diploid individual has a copy of the same allele at a locus for each of two homologous chromosomes). An individual is “heterozygous” if more than one allele type is present at a given locus (e.g., a diploid individual with one copy each of two different alleles). The term “homogeneity” indicates that members of a group have the same genotype at one or more specific loci. In contrast, the term “heterogeneity” is used to indicate that individuals within the group differ in genotype at one or more specific loci. The term “hybrid” refers to a variety or cultivar that is the result of a cross of plants of two different varieties. A hybrid, as described here, can refer to plants that are genetically different at any particular loci. A hybrid can further include a plant that is a variety that has been bred to have at least one different characteristic from the parent. A “Fl hybrid” refers to the first generation hybrid, “F2 hybrid” the second generation hybrid, “F3 hybrid” the third generation, and so on. A hybrid refers to any progeny that is either produced, or developed using research and development to create a new line having at least one distinct characteristic.

[0077] The terms "hybridizing specifically to", "specific hybridization", or “selectively hybridize to,” as used herein refer to the binding, duplexing, or hybridizing of a nucleic acid molecule preferentially to a particular nucleotide sequence under stringent conditions. The term “stringent conditions” refers to conditions under which a nucleic acid will hybridize preferentially to a target sequence, and to a lesser extent to, or not at all to, other off-target sequences. A “stringent hybridization” and “stringent hybridization wash conditions” in the context of nucleic acid hybridization (e.g., as in array, Southern or Northern hybridizations) are sequence dependent, and are different under different environmental parameters. An extensive guide to the hybridization of nucleic acids can be found in, e.g., Tijssen (1993) Laboratory Techniques in Biochemistry and Molecular Biology — Hybridization with Nucleic Acid Probes part I, Ch. 2, “Overview of principles of hybridization and the strategy of nucleic acid probe assays,” Elsevier, N.Y. (“Tijssen”). Generally, highly stringent hybridization and wash conditions are selected to be about 5 °C lower than the thermal melting point for the specific sequence at a defined ionic strength and pH. The Tmis the temperature (under defined ionic strength and pH) at which 50% of the target sequence hybridizes to a perfectly matched probe. Very stringent conditions are selected to be equal to the Tmfor a particular probe. An example of stringent hybridization conditions for hybridization of complementary nucleic acids which have more than 100 complementary residues on an array or on a filter in a Southern or northern blot is 42°C using standard hybridization solutions (see, e.g., Sambrook et al. (2014) Molecular Cloning: A Laboratory Manual (Fourth Edition, Cold Spring Harbor Laboratory Press, Plainview, N.Y.)).

[0078] As used herein, the term “inbreeding” refers to the production of offspring via the mating between relatives. The plants resulting from the inbreeding process are referred to herein as “inbred plants” or “inbreds.”

[0079] The terms “increase” or “decrease” refer to a positive (increase) or negative (decrease) difference relative to a reference value, such as a control. The difference can be quantitative. In some examples, the difference is statistically significant e.g., P-value less than 0.05 or 0.01). In some examples, the difference is an increase relative to a control of at least 5%, such as at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 150%, at least 200%, at least 250%, at least 300%, at least 350%, at least 400%, at least 500%, or greater than 500%. In some examples, the difference is a decrease relative to a control of at least 5%, such as at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, or 100%.

[0080] The term "introduced" refers to the incorporation of a particular nucleic acid sequence or protein into a cell, for example, via transformation. “Transformation” encompasses all techniques by which a nucleic acid molecule or protein might be introduced into such a cell, including chemical methods (e.g., calcium-phosphate transfection), physical methods (e.g., electroporation, microinjection, particle bombardment), fusion (e.g., liposomes), lipofection, nucleofection, receptor-mediated endocytosis (e.g., DNA-protein complexes, viral envelope / capsid-DNA complexes), agrobacterium-mediated transformation, biolistics (particle gun accelerator or gene gun), or other transduction and / or transfection methods. Transformation can include stable transformation, where a nucleic acid fragment is incorporated into the genome of a host cell (e.g., chromosome, plasmid, plastid or mitochondrial DNA), or transient transformation, e.g., transformation of an autonomous replicon or other transient molecule (e.g., transfected mRNA). In some examples, a genetic modification (e.g., a substitution, insertion, or deletion) is introduced through a gene editing technique, such as an RNAi, CRISPR / Cas9, ZFN, or TALEN based technique. In some examples, a gene (or vector carrying a gene) is introduced into a cell by transformation.

[0081] The terms “isolated” or “purified” in reference to biological components (such as nucleic acids, proteins, or cells) are components that have been substantially separated from other biological components in the environment in which the component occurs, e.g., separated from other chromosomal and extra-chromosomal DNA and RNA, proteins and / or cells. Nucleic acids and proteins that have been “isolated” include nucleic acids and proteins purified by standard purification methods. The term also embraces nucleic acids and proteins prepared by recombinant expression in a host cell as well as chemically synthesized nucleic acids.

[0082] Absolute purity or isolation is not required, it is intended as a relative term. Thus, for example, a purified / isolated protein, nucleic acid, or cell preparation is one in which the protein, nucleic acid, or cell is more enriched than the protein, nucleic acid, or cell is in its initial environment. In one example, a preparation is purified / isolated such that the protein, nucleic acid, or cell represents at least 50% of the total content of the preparation. A substantially purified protein or nucleic acid is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% pure. Thus, in one specific, non-limiting example, a substantially purified protein or nucleic acid is 90% free of other components.

[0083] The term “line” is used broadly to include, but is not limited to, a group of plants vegetatively propagated from a single parent plant, via tissue culture techniques or a group of inbred plants which are genetically very similar due to descent from a common parent(s). A plant is said to “belong” to a particular line if it (a) is a primary transformant (TO) plant regenerated from material of that line; (b) has a pedigree comprised of a TO plant of that line; or (c) is genetically very similar due to common ancestry (e.g., via inbreeding or selfing). In this context, the term “pedigree” denotes the lineage of a plant, e.g. in terms of the sexual crosses affected such that a gene or a combination of genes, in heterozygous (hemizygous) or homozygous condition, imparts a desired trait to the plant (e.g., modified plant height).

[0084] The term “locus” refers to a position on a genome that corresponds to a measurable property, e.g., a trait. Thus, a “modified plant height trait locus” as used herein is a position on the genome of a subject plant having genetic differences, in comparison to a reference genome that results in modified plant height in comparison to a reference plant.

[0085] The term “marker,” “genetic marker,” or “molecular marker,” refer to a nucleotide sequence or encoded product thereof (e.g., a protein) used as a point of reference for identifying a linked locus. A marker can be derived from genomic nucleotide sequence or from expressed nucleotide sequences (e.g., from a spliced RNA, a cDNA, etc.), or from an encoded polypeptide, and can be represented by one or more particular variant sequences, or by a consensus sequence. A “marker probe” is a nucleic acid molecule that can be used to identify the presence of a marker, e.g., a nucleic acid probe that is complementary to a marker locus sequence. Alternatively, in some aspects, a marker probe refers to a probe of any type that is able to distinguish (i.e., genotype) the particular allele that is present at a marker locus. A “marker locus” is a locus that can be used to track the presence of a second linked locus, e.g., a linked locus that encodes or contributes to expression of a phenotypic trait. For example, a marker locus can be used to monitor segregation of alleles at a locus, such as a QTL, that are genetically or physically linked to the marker locus. Thus, a “marker allele,” alternatively an “allele of a marker locus,” is one of a plurality of polymorphic nucleotide sequences found at a marker locus in a population that is polymorphic for the marker locus. Examples of markers include restriction fragment length polymorphism (RFLP) markers, amplified fragment length polymorphism (AFLP) markers, single nucleotide polymorphisms (SNPs), microsatellite markers (e.g. SSRs), sequence-characterized amplified region (SCAR) markers, cleaved amplified polymorphic sequence (CAPS) markers or isozyme markers or combinations of the markers described herein which defines a specific genetic and chromosomal location.

[0086] The term “marker assisted selection” refers to the process of identifying, optionally followed by selecting, a plant from a group of plants using the presence of a molecular marker as a selection criterion. The process usually involves detecting the presence of a certain nucleic acid sequence or polymorphism in the genome of a plant. The term “modified Cannabis plant,” “modified plant,” “engineered Cannabis plant,” or “engineered plant,” refer to an artificially produced plant that does not, or otherwise excludes, naturally occurring plants.

[0087] The term “modified plant height” refers to a Cannabis plant having a phenotypic change in plant height (e.g., taller, shorter, or more uniform) relative to a control (e.g., a parent or Abacus). In some aspects, plant height is a measure of vegetative height, harvest height, or stretch. In some examples, the control is a plant not having the same genotype for the markers indicating modified plant height. In some examples, the control plant is related to the plant having modified plant height, for example, a sibling or parent. Genetic markers that indicate modified plant height include markers associated with taller plant height (e.g., an average harvest height of over 95 cm) or shorter plant height (e.g., an average harvest height of less than 95 cm) depending on the desired plant height to be selected.

[0088] The term “nucleotide” refers to an organic molecule that serves as a monomeric unit of DNA and RNA. The nucleotide position is the position along a reference sequence wherein any particular monomeric unit of DNA or RNA is positioned relative to the other monomeric units of DNA or RNA.

[0089] The term “offspring” or “progeny” refer to a plant resulting as from a vegetative or sexual reproduction from one or more parent plants. For instance, an offspring / progeny plant may be obtained by cloning or selfing of a parent plant or by crossing two parent plants. An Fl is a first-generation offspring produced from parents at least one of which is used for the first time as donor of a trait, while offspring of second generation (F2) or subsequent generations (F3, F4, etc.) are specimens produced from selfings of Fl’s, F2’s etc. An Fl may thus be (and usually is) a hybrid resulting from a cross between two true breeding parents (true-breeding is homozygous for a trait), while an F2 may be (and usually is) an offspring resulting from self-pollination.

[0090] The term "operably linked" refers to the association of nucleic acid sequences on a single nucleic acid fragment such that the function of one is affected by the other. For example, a promoter is operably linked with a coding sequence when it is capable of inducing expression of that coding sequence (i.e., that the coding sequence is under the transcriptional control of the promoter). Coding sequences can be operably linked to regulatory sequences in sense or antisense orientation.

[0091] The term "plant" refers to a whole plant, cell, tissue, or other plant parts. Plant parts include any part(s) of a plant, including, for example and without limitation: seed (including mature seed and immature seed); a plant cutting; a plant cell; a plant cell culture; a plant organ (e.g., pollen, embryos, flowers, trichomes (e.g., glandular trichomes), fruits, shoots, leaves, roots, stems, and explants). Plant tissue refers to any tissue of a plant, including but not limited to, tissue from an embryo, shoot, root, stem, seed, stipule, leaf, trichome, petal, flower bud, flower, ovule, bract, branch, petiole, internode, bark, pubescence, tiller, rhizome, frond, blade, ovule, pollen, stamen. A plant tissue or plant organ may be a seed, protoplast, callus, or any other group of plant cells that is organized into a structural or functional unit. A plant cell or tissue culture may be capable of regenerating a plant having the physiological and morphological characteristics of the plant from which the cell or tissue was obtained, and of regenerating a plant having substantially the same genotype as the plant. In contrast, some plant cells are not capable of being regenerated to produce plants. Regenerable cells in a plant cell or tissue culture may be embryos, protoplasts, meristematic cells, callus, pollen, leaves, anthers, roots, root tips, silk, flowers, kernels, ears, cobs, husks, or stalks. Plant parts include harvestable parts and parts useful for propagation of progeny plants. Plant parts useful for propagation include, for example and without limitation: seed; fruit; a cutting; a seedling; a tuber; and a rootstock. A harvestable part of a plant may be any useful part of a plant, including, for example and without limitation: flower; pollen; seedling; tuber; leaf; stem; fruit; seed; and root.

[0092] A plant cell is the structural and physiological unit of the plant. A plant cell may be in the form of an isolated single cell, or an aggregate of cells (e.g., a friable callus and a cultured cell), and may be part of a higher organized unit (e.g., a plant tissue, plant organ, and plant). Thus, a plant cell may be a protoplast, a gamete producing cell, or a cell or collection of cells that can regenerate into a whole plant. As such, a seed, which comprises multiple plant cells and is capable of regenerating into a whole plant, is considered a “plant cell.” Described herein are plants in the genus of Cannabis and plants derived therefrom, which can be produced by asexual or sexual reproduction.

[0093] The term “plant height” refers to the height of a plant, and can include, for example, the vegetative height (e.g., height at the end of the vegetative phase before lights are switched from 18 hours light to 12 hours light, or height of a plant after topping), harvest height (height of mature plants when harvested), or stretch (the difference between harvest height and vegetative height).

[0094] The term “polymorphism” refers to a difference in the nucleotide or amino acid sequence of a given region as compared to a nucleotide or amino acid sequence in a homologous-region of another individual, in particular, a difference in the nucleotide of amino acid sequence of a given region which differs between individuals of the same species. A polymorphism is generally defined in relation to a reference sequence. Unless indicated otherwise, the reference sequence is the Cannabis Abacus reference genome (version Csat_AbacusV2, NCBI assembly accession GCA_025232715.1) or CDS derived from the Cannabis Abacus reference genome. Polymorphisms include single nucleotide differences, differences in sequence of more than one nucleotide, and single or multiple nucleotide insertions, inversions and deletions; as well as single amino acid differences, differences in sequence of more than one amino acid, and single or multiple amino acid insertions, inversions, and deletions.

[0095] The terms “polynucleotide,” “polynucleotide sequence,” “nucleotide sequence,” “nucleic acid sequence,” and “nucleic acid fragment,” are used interchangeably. These terms encompass polymers composed of nucleotide units (ribonucleotides, deoxyribonucleotides, related naturally occurring structural variants, and synthetic non-naturally occurring analogs thereof). The term “oligonucleotide” typically refers to short polynucleotides, generally no greater than 150 nucleotides. It will be understood that when a nucleic acid sequence is represented as a DNA sequence ( / '.<?., A, T, G, C), this also includes an RNA sequence (z.e., A, U, G, C) in which “U” replaces “T.” Nucleic acids can be single- or doublestranded. Exemplary nucleic acids include cDNA, genomic DNA, synthetic DNA, RNA, or mixtures thereof.

[0096] The term “polypeptide” or “protein” refers to a polymer of amino acid residues. The terms apply to amino acid polymers in which one or more amino acid residue is an artificial chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers. The term “amino acid residue” or “amino acid” includes reference to an amino acid that is incorporated into a protein, polypeptide, or peptide. The amino acid can be a naturally occurring amino acid and, unless otherwise limited, can encompass known analogs of natural amino acids that can function in a similar manner as naturally occurring amino acids. As used herein, “recombinant” includes reference to a protein produced using cells that do not have, in their native state, an endogenous copy of the DNA able to express the protein. The cells produce the recombinant protein because they have been genetically altered by the introduction of the appropriate isolated nucleic acid sequence. The term also includes reference to a cell, or nucleic acid, or vector, that has been modified by the introduction of a heterologous nucleic acid or the alteration of a native nucleic acid to a form not native to that cell, or that the cell is derived from a cell so modified.

[0097] The term "primer" as used herein refers to an oligonucleotide, either RNA or DNA, either singlestranded or double-stranded, either derived from a biological system, generated by restriction enzyme digestion, or produced synthetically which, when placed in the proper environment, is able to functionally act as an initiator of template-dependent nucleic acid synthesis. When presented with an appropriate nucleic acid template, suitable nucleoside triphosphate precursors of nucleic acids, a polymerase enzyme, suitable cofactors and conditions such as a suitable temperature and pH, the primer may be extended at its 3' terminus by the addition of nucleotides by the action of a polymerase or similar' activity to yield a primer extension product. The primer may vary in length depending on the particular conditions and requirements of the application. For example, in diagnostic applications, the oligonucleotide primer is typically 15-25 or more nucleotides in length. The primer must be of sufficient complementarity to the desired template to prime the synthesis of the desired extension product, that is, to be able anneal with the desired template strand in a manner sufficient to provide the 3’ hydroxyl moiety of the primer in appropriate juxtaposition for use in the initiation of synthesis by a polymerase or similar enzyme. It is not required that the primer sequence represent an exact complement of the desired template. For example, a non-complementary nucleotide sequence may be attached to the 5’ end of an otherwise complementary primer. Alternatively, non-complementary bases may be interspersed within the oligonucleotide primer sequence, provided that the primer sequence has sufficient complementarity with the sequence of the desired template strand to functionally provide a template-primer complex for the synthesis of the extension product.

[0098] The term “probe,” “nucleic acid probe,” or “oligonucleotide probe” as used herein, is one or more synthetic nucleic acid molecules that are complementary to a nucleic acid sequence of interest (target sequence), and hybridize to a sequence of interest when under hybridization conditions. Probes can be used to detect, analyze, and / or visualize the nucleic acid sequence of interest on a molecular level. Specific hybridization of a probe to a nucleic acid sequence of interest can be detected, for example, through a label on the probe. Probes have a length suitable to achieve a desired specificity to the target sequence, however, are generally at least 10 nucleotides long, for example, at least 15 nucleotides, at least 20 nucleotides, or at least 50 nucleotides long. Probes can be immobilized on a solid surface (e.g.. nitrocellulose, glass, quartz, fused silica slides), as in an array. The precise sequence of the particular probes described herein can be modified to a certain degree to produce probes that are “substantially identical” to the disclosed probes, but retain the ability to specifically bind to (i.e., hybridize specifically to) the same targets as the probe from which they were derived. Such modifications are specifically covered by reference to the individual probes described herein.

[0099] The term “product” as used in reference to a Cannabis product, is a composition including Cannabis (or an extract thereof). Products include, but are not limited to: a kief, hashish, bubble hash, an edible product, solvent reduced oil, sludge, e-juice, tincture, or other compositions including Cannabis (e.g., a Cannabis plant disclosed herein, or an extract thereof).

[0100] The term “promoter” refers to a nucleic acid control sequence that directs transcription of a nucleic acid. A promoter includes necessary nucleic acid sequences near the start site of transcription, and may include distal enhancer or repressor elements. A “constitutive promoter” is a promoter that is continuously active and is not subject to regulation by external signals or molecules. In contrast, the activity of an “inducible promoter” is regulated by an external signal or molecule (for example, a transcription factor). Exemplary promoters include pol III promoters (e.g., U6), pol II promoter, ubiquitin promoter, Cauliflower Mosaic Virus (CaMV) 35S promoter, or RUBISCO promoter. The terms “initiate transcription,” “initiate expression,” “drive transcription,” and “drive expression” are used interchangeably herein and all refer to the primary function of a promoter.

[0101] The term "quantitative trait loci" or "QTL" refers to the genetic elements controlling a quantitative trait. The term “recombinant” refers to a nucleic acid or protein that has a sequence made by an artificial combination of two otherwise separated segments of sequence (e.g., a “chimeric” sequence). This artificial combination can be accomplished by chemical synthesis or by manipulation of isolated segments of nucleic acids, for example, by standard molecular- biology techniques (e.g., cloning). A “recombinant expression construct” refers to an expression vector into which a nucleic acid sequence or fragment can be moved. Preferably, it is a plasmid vector, or a fragment thereof, comprising a promoter. The choice of plasmid vector is dependent upon the method that will be used to transform host plants. Similarly, genetic elements that must be present on the plasmid vector to successfully transform, select and propagate host cells containing the chimeric gene is dependent on the specific transformation method. Different independent transformation events typically result in different levels and patterns of expression and thus multiple events must be screened to obtain lines displaying the desired expression level and pattern. Such screening may be accomplished by PCR and Southern analysis of D A, RT-PCR and Northern analysis of mRNA expression, Western analysis of protein expression, or phenotypic analysis.

[0102] The term “reference plant” or “reference genome” refers to a reference sequence that genetic markers or sequences of a test sample can be compared to in order to detect a modification of the sequence in the test sample. In some examples, the reference plant or genome is Abacus (Csat_AbacusV2, NCBI assembly accession GCA_025232715.1).

[0103] The terms “sequence identity” or “percent identity” are used interchangeably to refer to a sequence comparison based on identical matches between correspondingly identical positions in two or more amino acid or nucleotide sequences that are being compared. The percent identity refers to the extent to which two optimally aligned polynucleotide or peptide sequences are invar iant throughout a window of alignment of components, e.g., nucleotides or amino acids. Hybridization experiments and mathematical algorithms known in the art may be used to determine percent identity. Many mathematical algorithms exist as sequence alignment computer programs known in the art that calculate percent identity. These programs may be categorized as either global sequence alignment programs or local sequence alignment programs.

[0104] The NCBI Basic Local Alignment Search Tool (BLAST) tool is often used and is available from several sources, including the National Center for Biotechnology Information (blast.ncbi.nlm.nih.gov / Blast.cgi). Various types of BLAST are available, for example, blastp, blastn, blastx, tblastn and tblastx. A description of how to determine sequence identity using this program is available on the NCBI website and other resources. In some examples, percent sequence identity is determined by using BLAST with default parameters.

[0105] The term “substantially similar” as used herein refers to nucleic acid fragments wherein changes in one or more nucleotide bases do not affect the ability of the nucleic acid fragment to mediate gene expression or produce a certain phenotype. These terms also refer to modifications of nucleic acid fragments, such as deletion or insertion of one or more nucleotides that do not substantially alter the functional properties of the resulting nucleic acid fragment relative to the initial, unmodified fragment. A “substantially homologous sequence” refers to variants of the disclosed sequences such as those that result from site-directed mutagenesis, as well as synthetically derived sequences. A substantially homologous sequence also refers to fragments of a particular promoter nucleotide sequence disclosed herein that operate to promote the constitutive expression of an operably linked heterologous nucleic acid fragment. These promoter fragments will include at least about 20 contiguous nucleotides, for example, at least 50 contiguous nucleotides, at least 75 contiguous nucleotides, or at least 100 contiguous nucleotides of the particular- promoter nucleotide sequence disclosed herein. The nucleotides of such fragments will usually comprise the TATA recognition sequence of the particular promoter sequence. Such fragments may be obtained by use of restriction enzymes to cleave the naturally occurring promoter nucleotide sequences disclosed herein; by synthesizing a nucleotide sequence from the naturally occurring promoter DNA sequence; or may be obtained through the use of PCR technology. Functional variants of these promoter fragments, such as those resulting from site-directed mutagenesis, are encompassed by the present disclosure.

[0106] The term “single nucleotide polymorphism (SNP)” refers to a change in which a single base in the DNA differs from the base at the corresponding position of a reference genome or sequence.

[0107] The term "target region" or "nucleic acid target" refers to a nucleotide sequence that resides at a specific chromosomal location. The "target region" or "nucleic acid target" can be specifically recognized by a probe.

[0108] The term “transformant” refers to a cell, tissue or organism that has undergone transformation. The original transformant is designated as “TO” or “TO.” Selfing the TO produces a first transformed generation designated as “Tl” or “Tl.”

[0109] The term “transgenic” refers to any cell, cell line, callus, tissue, plant part or plant, the genome of which has been altered by the presence of a heterologous nucleic acid, such as a recombinant DNA construct, including those initial transgenic events as well as those created by sexual crosses or asexual propagation from the initial transgenic event. The term “transgenic” as used herein does not encompass the alteration of the genome (chromosomal or extra-chromosomal) by conventional plant breeding methods or by naturally occurring events such as random cross-fertilization, non-recombinant viral infection, non-recombinant bacterial transformation, non-recombinant transposition, or spontaneous mutation. A “transgene” is a gene that has been introduced into the genome by a transformation procedure. The term “variety” as used herein has identical meaning to the corresponding definition in the International Convention for the Protection of New Varieties of Plants (UPOV treaty), of Dec. 2, 1961, as Revised at Geneva on Nov. 10, 1972, on Oct. 23, 1978, and on Mar. 19, 1991. Thus, “variety” means a plant grouping within a single botanical taxon of the lowest known rank, which grouping, irrespective of whether the conditions for the grant of a breeder’s right are fully met, can be i) defined by the expression of the characteristics resulting from a given genotype or combination of genotypes, ii) distinguished from any other plant grouping by the expression of at least one of the said characteristics and iii) considered as a unit with regard to its suitability for being propagated unchanged.

[0110] The term “vector” refers to a nucleic acid molecule that can be introduced into a host cell (for example, by transformation), thereby producing a transformed host cell. A vector can include nucleic acid sequences that permit it to replicate in a host cell, such as an origin of replication. Recombinant DNA vectors are vectors containing recombinant DNA. A vector can also include one or more selectable marker genes and other genetic elements. Often vectors are DNA plasmids, however, they can also be viral vectors (DNA or RNA), cosmids, or artificial chromosomes.

[0111] III. Genetic Markers and Detection Thereof

[0112] The present disclosure describes the discovery of novel markers associated with plant height. Cannabis sativa as a species varies considerably for plant height, from <50 cm for day-neutral (autoflowering) varieties to >200 cm for some fiber hemp cultivars. However, different Cannabis production systems have different height requirements. Shorter plants (e.g., plants under 95 cm, for example, 85-94.5 cm) are optimal for tiered growth room systems, however, some tiered systems can accommodate plants over 100 cm (e.g., up to 112 cm). Greenhouse systems require different heights depending on whether the production system makes use of trellis and / or if there is a desire for a more open top of the canopy, which is a result of stretch. Trellis-based greenhouse systems have a limit on the amount of stretch that can be accommodated due to the need to anticipate trellis height. Trellis systems can typically accommodate plants up to 155 cm tall, although plants 90-155 cm may be most suitable. A typical height range for outdoor systems is 150-185 cm, however, taller plants can be grown “orchard style” outdoors.

[0113] Genetic markers for vegetative height, harvest height, and / or stretch allow breeders to develop varieties for use in various production systems and / or develop varieties having improved height uniformity. Uniformity of height is an important trait for mechanical harvesting as well as ensuring all plants get adequate light; the amount of light determines cannabinoid production and shorter plants could be shaded by taller plants, thus resulting in variable cannabinoid concentrations. Disclosed herein are methods of selecting and / or producing Cannabis plants having modified plant height (e.g., taller, shorter, or more uniform relative to a parent variety or Abacus). The methods include analyzing and / or detecting one or more genetic markers in a nucleic acid sample from a Cannabis plant or its germplasm that indicate modified plant height. The genetic markers are associated with a height trait, for example, tall height (e.g., a harvest height of 95 cm or more) or short height (e.g., a harvest height of less than 95 cm). When selecting for uniformity, plants can be selected to have uniform marker genotypes and / or homozygous markers. A practitioner can select the appropriate marker and genotype depending on the target trait and application. In some aspects, methods disclosed herein select or produce Cannabis plants having increased plant height. In some aspects, methods disclosed herein select or produce Cannabis plants having decreased plant height. In some aspects, methods disclosed herein select or produce Cannabis plants having more uniform plant height.

[0114] In some aspects, the methods disclosed herein include a step of selecting a Cannabis plant comprising the one or more genetic markers that indicate modified plant height. The plants can be selected, for example, for further analysis, propagation, breeding, and / or to make a product. In some aspects, the plants are selected, for further analysis, propagation, and / or to make a product. In some aspects, the one or more genetic markers that indicate modified plant height include one or more genetic markers described in any one of Tables 2-10 (e.g., SNPs).

[0115] In some aspects, the methods include analyzing one or more genetic markers in a nucleic acid sample from a Cannabis plant (or its germplasm) and / or detecting one or more genetic markers that indicate modified plant height. In some aspects, the methods include a first step of obtaining nucleic acids from a sample plant or its germplasm. In some examples, the sample plant is a progeny plant obtained from a cross between a first plant having modified plant height (e.g., a desired height) and a second plant not having modified plant height.

[0116] In some aspects, markers that indicate modified plant height indicate a tall plant height. In some aspects, tall Cannabis plants have a greater average vegetative height, harvest height, and / or stretch relative to the Abacus Cannabis variety. In some examples, a tall Cannabis plant is a plant having an average harvest height of at least 95 cm, for example, at least 100 cm, at least 105 cm, 110 cm, 115 cm, 120 cm, 125 cm, 130 cm, 135 cm, 140 cm, etc. In some aspects, a tall Cannabis plant is plant having an average harvest height between 95-550 cm, for example, 100-550 cm, 105-550 cm, 110-550 cm, 115-550 cm, 120-550 cm, 125-550 cm, 130-550 cm, 135-550 cm, 140-550 cm, 145-550 cm, 150-550 cm, 200-550 cm, 250-550 cm, 300-550 cm, 350-550 cm, 400-550 cm, 95-400 cm, 100-400 cm, 105-400 cm, 110-400 cm, 115-400 cm, 120-400 cm, 125-400 cm, 130-400 cm, 135-400 cm, 140-400 cm, 145-400 cm, 150-400 cm, 200-400 cm, 250-400 cm, 300-400 cm, 350-400 cm, 95-300 cm, 100-300 cm, 105-300 cm, 110-300 cm, 115-300 cm, 120-300 cm, 125-300 cm, 130-300 cm, 135-300 cm, 140-300 cm, 145-300 cm, 150-300 cm, 200-300 cm, 250-300 cm, 95-210 cm, 100-210 cm, 105-210 cm, 110-210 cm, 115-210 cm, 120-210 cm, 125-210 cm, 130-210 cm, 135-210 cm, 140-210 cm, 145-210 cm, 150-210 cm, 100-210 cm, 105-210 cm, 110-210 cm, 115-210 cm, 120-210 cm, 125-210 cm, 130-210 cm, 135-210 cm, 140-210 cm, 145-210 cm, 150-210 cm, 95-200 cm, 95-200, 100-200 cm, 105-200 cm, 110-200 cm, 115-200 cm, 120-200 cm, 125-200 cm, 130-200 cm, 135-200 cm, 140-200 cm, 145-200 cm, 150-200 cm, 95-180 cm, 100-180 cm, 105-180 cm, 110-180 cm, 115-180 cm, 120-180 cm, 125-180 cm, 130-180 cm, 135-180 cm, 140-180 cm, 145-180 cm, 150-180 cm, 95-160 cm, 100-160 cm, 105-160 cm, 110-160 cm, 115-160 cm, 120-160 cm, 125-160 cm, 130-160 cm, 135-160 cm, 140-160 cm, 145-160 cm, 150-160 cm, 95-141 cm, 100-141 cm, 105-141 cm, 110-141 cm, 115-141 cm, 120-141 cm, 125-141 cm, 130-141 cm, or 135-141 cm. In some aspects, a tall Cannabis plant is plant having an average harvest height between 95-210 cm. In some aspects, a tall Cannabis plant is plant having an average harvest height between 95-310 cm.

[0117] In some aspects, markers that indicate modified plant height indicate a short plant height. In some aspects, short Cannabis plants are Cannabis plants having an average vegetative height, harvest height, and / or stretch that is less than, or not statistically different from, the Abacus Cannabis variety. In some examples, a short Cannabis plant is a plant having an average harvest height of less than 95 cm, for example, less than 90 cm, 85 cm, 80 cm, 75 cm, 70 cm, 65 cm, 60 cm etc. In some examples, a short Cannabis plant is a plant having an average harvest height between 20-95 cm, for example, 40-95 cm, 45- 95 cm, 50-95 cm, 55-95 cm, 60-95 cm, 65-95 cm, 70-95 cm, 75-95 cm, 80-95 cm, 85-95 cm, 90-95 cm, 20-90 cm, 40-90 cm, 45-90 cm, 50-90 cm, 55-90 cm, 60-90 cm, 65-90 cm, 70-90 cm, 75-90 cm, 80-90 cm, 85-90 cm, 20-80 cm, 40-80 cm, 45-80 cm, 50-80 cm, 55-80 cm, 60-80 cm, 65-80 cm, 70-80 cm, 75- 80 cm, 20-70 cm, 40-70 cm, 45-70 cm, 50-70 cm, 55-70 cm, 60-70 cm, 65-70 cm, 20-65 cm, 40-65 cm, 45-65 cm, 50-65 cm, 55-65 cm, or 60-65 cm. In some examples, a short Cannabis plant is a plant having an average harvest height between 90-95 cm.

[0118] In some aspects, markers that indicate modified plant height are used to select for greater height uniformity. In such examples, a practitioner screens one or more plant height markers, and selects for a uniform genotype among the one or more genetic markers.

[0119] The markers of the present disclosure were discovered as described herein, and comprise polymorphisms relative to the Abacus Cannabis reference genome Csat_AbacusV2; NCBI assembly accession GCA_025232715.1 (CsaAba2). Table 8 describes markers and sequence identifiers, and the positioning on their respective chromosomes. Additional tables in the specification include the reference call of the nucleotide at the respective position within the reference genome, as well as the alternate call describing the polymorphism in plants having modified plant height. The markers disclosed herein are useful, for example, for selecting Cannabis plants that grow to a desired height, for example, to breed new varieties suitable for a particular growth format, and / or for producing Cannabis plants having modified plant height relative to a parent plant (e.g., taller or shorter relative to a parent).

[0120] In some aspects, the methods include analyzing and / or detecting one or more genetic markers (e.g., SNPs) described in any one of Tables 2-10. In some aspects, the methods include analyzing and / or detecting one or more genetic markers (e.g., SNPs) described in Table 2. In some aspects, the methods include analyzing and / or detecting one or more genetic markers (e.g., SNPs) described in Table 3. In some aspects, the methods include analyzing and / or detecting one or more genetic markers (e.g.. SNPs) described in Table 4. In some aspects, the methods include analyzing and / or detecting one or more genetic markers (e.g., SNPs) described in Table 5. In some aspects, the methods include analyzing and / or detecting one or more genetic markers (e.g., SNPs) described in Table 6. In some aspects, the methods include analyzing and / or detecting one or more genetic markers (e.g., SNPs) described in Table 7. In some aspects, the methods include analyzing and / or detecting one or more genetic markers (e.g., SNPs) described in Table 8. In some aspects, the methods include analyzing and / or detecting one or more genetic markers (e.g., SNPs) described in Table 9. In some aspects, the methods include analyzing and / or detecting one or more genetic markers (e.g., SNPs) described in Table 10.

[0121] In some aspects, analyzing and / or detecting one or more genetic markers includes analyzing and / or detecting one or more of SNP markers: 142603_12546539, 142603_12519907, Cannabis.vl_scf542.169826.101, 142603.12500071, 142603_12492164, Cannabis.vl_scf542.132451.99, Cannabis.vl_scf542.122780_100, 142603.12456339, 142603.12451790, 142603.12439239, 142603.12351777, 142603.12346103, 142603.12333141, 142603.12209629, 142603.12122639, 142603.12107817, 142603.12087936, 142603.12081998, 142603.12064923, 142603.12038515, 142603.12001291, 142603.11997059, 142603.11989491, 142603.11951907, 142603.11459458, 142603.12546539, 142603.12519907, and 142603.11711278. In some aspects, the methods include detecting and / or analyzing one or more of SNP markers: 142603.12439239, 142603.12333141 , 142603.12107817, 142708.323806, 142708.353959, and 142603.12081998. In some aspects, the methods include detecting and / or analyzing one or more of SNP markers: 142603.12439239, 142603.12333141, 142603.12107817, 142708.323806, and 142708.353959. In some aspects, the methods include detecting and / or analyzing one or more of SNP markers: 142603.12439239, 142603.12107817, 142708.323806, 142708.353959, and 142708.453318. In a non-limiting example, analyzing and / or detecting one or more genetic markers includes analyzing and / or detecting SNP marker 142603.12107817 and / or 142603.12081998. In some examples, the one or more genetic markers are associated with or indicate modified plant height (e.g., associated with tall or short plant height). Additional information for the SNP markers arc provided, for example, in Tables 2-7. In some aspects, the SNP markers refer to a reference genome that is Abacus Cannabis genome version Csat_AbacusV2, NCBI assembly accession GCA_025232715.1.

[0122] In a non-limiting example, the methods disclosed herein include detecting and / or analyzing one or more of SNP markers: 142603_12107817, 142603.12439239, 142603.12107817, 142603.12439239, and 142708.323806. In some aspects, the SNP markers refer to a reference genome that is Abacus Cannabis genome version Csat. Abacus V2, NCBI assembly accession GCA.025232715.1.

[0123] In some aspects, at least two genetic markers described herein are analyzed in a nucleic acid sample from the Cannabis plant or its germplasm, for example, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, or 17 genetic markers are analyzed. In some aspects, the methods disclosed herein include analyzing at least 3 genetic markers. In some aspects, the methods disclosed herein include analyzing at least 4 genetic markers. In some aspects, the methods disclosed herein include analyzing at least 5 genetic markers. In some aspects, the methods disclosed herein include analyzing 1 to 205 genetic markers, for example, analyzing 1 to 2, 1 to 3, 1 to 4, 1 to 5, 1 to 6, 1 to 7, 1 to 8, 1 to 9, 1 to 10, 1 to 11, 1 to 12, 1 to 13, 1 to 14, 1 to 15, 1 to 20, 1 to 30, 1 to 50, 1 to 100, 1 to 200, 2 to 3, 2 to 4, 2 to 5, 2 to 6, 2 to 7, 2 to 8, 2 to 9, 2 to 10, 2 to 11, 2 to 12, 2 to 13, 2 to 14, 2 to 15, 2 to 20, 2 to 30, 2 to 50, 2 to 100, 2 to 200, 2 to 205, 3 to 4, 3 to 5, 3 to 6, 3 to 7, 3 to 8, 3 to 9, 3 to 10, 3 to 11, 3 to 12, 3 to 13, 3 to 14, 3 to 15, 3 to 20, 3 to 30, 3 to 50, 3 to 100, 3 to 200, 3 to 205, 4 to 5, 4 to 6, 4 to 7, 4 to 8, 4 to 9, 4 to 10, 4 to 11, 4 to 12, 4 to

[0124] 13, 4 to 14, 4 to 15, 4 to 20, 4 to 30, 4 to 50, 4 to 100, 4 to 200, 4 to 205, 5 to 6, 5 to 7, 5 to 8, 5 to 9, 5 to

[0125] 10, 5 to 11, 5 to 12, 5 to 13, 5 to 14, 5 to 15, 5 to 20, 5 to 30, 5 to 50, 5 to 100, 5 to 200, 5 to 205, 6 to 7, 6 to 8, 6 to 9, 6 to 10, 6 to 11, 6 to 12, 6 to 13, 6 to 14, 6 to 15, 6 to 20, 6 to 30, 6 to 50, 6 to 100, 6 to 200, 6 to 205, 10 to 20, 10 to 30, 10 to 50, 10 to 100, 10 to 200, or 10 to 205 genetic markers. In some aspects, the methods disclosed herein include analyzing 2-10 genetic markers. In some aspects, the methods disclosed herein include analyzing 4-10 genetic markers. In some aspects, the methods disclosed herein include analyzing 3-6 genetic markers. In some aspects, the methods disclosed herein include analyzing 4-6 genetic markers.

[0126] In some examples, the method includes detecting at least one genetic marker (e.g., SNP) indicating modified plant height as disclosed herein (see, e.g., Tables 2-10). In some examples, at least two genetic markers that indicate modified plant height are detected, for example, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, or at least 17 genetic markers that indicate modified plant height are detected. In some aspects, the methods include detecting a number of genetic markers that indicate modified plant height, for example, 1 to 2, 1 to 3, 1 to 4, 1 to 5, 1 to 6, 1 to 7, 1 to 8, 1 to 9, 1 to 10, 1 to 11, 1 to 12, 1 to 13, 1 to 14, 1 to 15, 1 to 16, 1 to 17, 2 to 17, 3 to 17, 4 to 17, 5 to 17, 6 to 17, 7 to 17, 8 to 17, 9 to 17, 10 to 17, 11 to 17, 12 to 17, 13 to 17, 14 to 17. 15 to 17, 16 to 17, 2 to 16, 2 to 15, 2 to 14, 2 to 13, 2 to 12, 2 to 11, 2 to 10, 2 to 9, 2 to 8, 2 to 7, 2 to 6, 2 to 5, 2 to 4, 2 to 3, 3 to 16, 3 to 15, 3 to 14, 3 to 13, 3 to 12, 3 to 11, 3 to 10, 3 to 9, 3 to 8, 3 to 7, 3 to 6, 3 to 5, 3 to 4, 4 to 16, 4 to 15, 4 to 14, 4 to 13, 4 to 12, 4 to 11, 4 to 10, 4 to 9, 4 to 8, 4 to 7, 4 to 6, 4 to 5, 5 to 16, 5 to 15, 5 to 14, 5 to 13, 5 to 12, 5 to 11, 5 to 10, 5 to

[0127] 9, 5 to 8, 5 to 7, 5 to 6, 6 to 15, 6 to 14, 6 to 13, 6 to 12, 6 to 11, 6 to 10, 6 to 9, 6 to 8, 6 to 7, 7 to 15, 7 to

[0128] 14, 7 to 13, 7 to 12, 7 to 11, 7 to 10, 7 to 9, 7 to 8, 8 to 16, 8 to 15, 8 to 14, 8 to 13, 8 to 12, 8 to 11, 8 to

[0129] 10, 8 to 9, 9 to 16, 9 to 15, 9 to 14, 9 to 13, 9 to 12, 9 to 11, 9 to 10, 10 to 16, 10 to 15, 10 to 14, 10 to 13, 10 to 12, 10 to 11, 11 to 16, 11 to 15, 11 to 14, 11 to 13, 11 to 12, 12 to 16, 12 to 15, 12 to 14, 12 to 13, 13 to 16, 13 to 15, or 13 to 14 genetic markers that indicate modified plant height. In some aspects, the methods disclosed herein include detecting 2-10 genetic markers that indicate modified plant height. In some aspects, the methods disclosed herein include detecting 4-10 genetic markers that indicate modified plant height. In some aspects, the methods disclosed herein include detecting 3-6 genetic markers that indicate modified plant height. In some aspects, the methods disclosed herein include detecting 4-6 genetic markers that indicate modified plant height. In some aspects, the methods disclosed herein include detecting 2-4 genetic markers that indicate modified plant height. Detecting one or more such genetic markers indicates that the Cannabis plant will have a modified height (e.g., taller or shorter), or increases the likelihood of the Cannabis plant having a modified height.

[0130] The markers of the present disclosure are described in numerous fashions. To illustrate, for nonlimiting exemplary purposes, marker 142603_12107817, which is described, for example, in Tables 2, 3, 4, 8, and 9, is described as being positioned at base pair (bp) position 484,010 on chromosome 4 of the CsaAba2 reference genome. Likewise, marker 142603_12107817 is also described as nucleotide 51 of SEQ ID NO: 44.

[0131] The present disclosure further describes the discovery of novel haplotype markers for plants, including Cannabis. Haplotypes refer to the genotype of a plant at a plurality of genetic loci, e.g., a combination of alleles or markers. Haplotypes can refer to sequence polymorphisms at a particular locus, such as a single marker locus, or sequence polymorphisms at multiple loci along a chromosomal segment in a given genome. Markers of the present disclosure and within the haplotypes described are significantly associated with plant height, and thus can be used to screen plants exhibiting a modified plant height phenotype. Table 8 describes markers within a haplotype that identify polymorphisms associated with plant height. In particular, Table 8 describes the haplotype both with respect to the left and right flanking markers, and with respect to the left and right flanking positioning on their respective chromosomes. To illustrate, for non-limiting exemplary puiposes, marker 142603_12107817 is within a haplotype defined as being between left flanking marker 142603_12113682 at position 478,145 on chromosome 4 and right flanking marker 142603_12067549 at position 524,335 on chromosome 4 of the CsaAba2 reference genome. In some aspects, the methods include detecting one or more genetic markers (e.g.. genetic markers that indicate modified plant height) within a haplotype (e.g.. a haplotype that indicates modified plant height), such as one or more haplotypes described in Table 8.

[0132] The present disclosure further describes the discovery of novel trait loci for plant height. A chromosomal interval designating a contiguous linear span of genomic DNA that resides usually on a single chromosome is a genetic locus. Thus, a plant height trait locus is a chromosome interval linked to the modified plant height phenotype. A chromosome interval may comprise a quantitative trait locus (“QTL”) linked with a genetic trait and the QTL may comprise a single gene or multiple genes associated with the genetic trait. The boundaries of a chromosome interval comprising a QTL are drawn such that a marker that lies within the chromosome interval can be used as a marker for the genetic trait, as well as markers genetically linked thereto. Each interval comprising a QTL comprises at least one gene conferring a given trait, however knowledge of how many genes are in a particular interval is not necessary to make or practice the disclosure, as such an interval will segregate at meiosis as a linkage block. In accordance with the present disclosure, a chromosomal interval comprising a QTL may therefore be readily introgressed and tracked in a given genetic background using the methods and compositions provided herein.

[0133] Identification of chromosomal intervals and QTL is therefore beneficial for detecting and tracking a genetic trait, such as plant height, in plant populations. In some aspects, this is accomplished by identification of markers linked to a particular QTL. The principles of QTL analysis and statistical methods for calculating linkage between markers and useful QTL include regression analysis, single point marker analysis, complex pedigree analysis, Bayesian MCMC, identity-by-descent analysis, interval mapping, composite interval mapping (CIM), and Haseman-Elston regression. QTL analyses may be performed with the help of a computer and specialized software available from a variety of public and commercial sources.

[0134] Markers of the present disclosure that are within a plant height trait locus are significantly correlated to modified plant height and genetically linked, and thus can be used to screen plants exhibiting a modified plant height phenotype.

[0135] The present disclosure includes analyzing and / or detecting genetic markers associated with plant height, such as nucleic acid polymorphisms (e.g., SNPs). In some aspects, the genetic markers indicate modified plant height e.g., taller or shorter relative to a parent or Abacus). Methods of analyzing or detecting nucleic acid polymorphisms can include, but are not limited to, amplification techniques (e.g., by PCR), sequencing techniques, or techniques for the detection of a label (e.g., labeled probe or other indicator). PCR uses a particular amplification primer pair that specifically hybridize to a target polynucleotide and produce an amplification product (the amplicon). Primers can be designed such that the amplicon can contain a nucleic acid polymorphism of interest. Methods for designing PCR primers and PCR conditions have been described, for example, in Sambrook et al. (2014) Molecular- Cloning: A Laboratory Manual (Fourth Edition, Cold Spring Harbor Laboratory Press, Plainview, N.Y.). A number of parameters in a specific PCR protocol may need to be adjusted to specific laboratory conditions and may be slightly modified, and yet allow for the collection of similar results. The primers can be radiolabeled, or labeled by any suitable means (e.g., using a non-radioactive fluorescent tag), to allow for rapid visualization of the different size amplicons following an amplification reaction without any additional labeling step or visualization step.

[0136] Other examples of nucleic acid amplification methods include, but are not limited to, reversetranscription PCR (RT-PCR), quantitative real-time PCR (qPCR), quantitative real-time reverse transcriptase PCR (qRT-PCR) (see, e.g., Adams, A beginner’s guide to RT-PCR, qPCR and RT-qPCR, Biochemist (Lond) (2020) 42(3): 48-53), isothermal amplification methods (see, e.g., Zanoli et al., Biosensors (2013) 3(1): 18-43), nucleic acid sequence-based amplification (NASBA) (see, e.g., Deiman and Sillekens, Mol Biotechnol (2002) 20(2):163-79), loop-mediated isothermal amplification (LAMP) (see, e.g., Notomi et al., (2000) Nucleic Acids Res. 28(12): e63), helicase-dependent amplification (HD A) (see, e.g., Cao et al.. Helicase-dependent amplification of nucleic acids, Curr Protoc Mol Biol, 104:15.11.1-15.11.12, 2013), rolling circle amplification (RCA) (see, e.g, Yao et al. Nature Protocols (2021) 16, 5460-5483), multiple displacement amplification (MDA) (see, e.g, Spits et al. Nature Protocols (2006) 1: 1965-1970), recombinase polymerase amplification (RPA) (see, e.g., Lobato et al., Trends Analyt Chem (2018) 98: 19-35), ligase chain reaction (LCR) (see e.g., Gibriel and Adel, Mutat Res Rev Mutat Res. (2017) 773: 66-90), transcription amplification (see e.g., Kwoh et al. (1989) Proc. Natl. Acad. Sci. USA 86: 1173), self-sustained sequence replication (see e.g., Guatelli et al. (1990) Proc. Natl. Acad. Sci. USA 87: 1874), dot PCR, and linker adapter PCR. Additional information and amplification methods can be found, for example, in Sambrook et al. (2014) Molecular Cloning: A Laboratory Manual (Fourth Edition, Cold Spring Harbor Laboratory Press, Plainview, N.Y.).

[0137] In some aspects, amplification produces an amplicon that is at least 20 nucleotides in length, for example, at least 50 nucleotides in length, or alternatively, at least 100 nucleotides in length, at least 200 nucleotides in length, at least 300, at least 400, at least 500, at least 1000, at least 2000, etc., nucleotides in length. In some examples, the amplicon is no longer than 10,000 nucleotides in length, for example, no longer than 3,000, no longer than 5,000, no longer than 7,000, or no longer than 9,000 nucleotides in length. In some examples, marker amplification produces an amplicon that is 20 to 10,000 nucleotides in length, for example, 50 to 9000 nucleotides, 50 to 8000 nucleotides, 50 to 7000 nucleotides, 50 to 6000 nucleotides, 50 to 5000 nucleotides, 50 to 4000 nucleotides, 50 to 3000 nucleotides, 50 to 2000 nucleotides, 50 to 1000 nucleotides, 50 to 500 nucleotides, 50 to 400 nucleotides, 50 to 300 nucleotides, 50 to 200 nucleotides, 50 to 150 nucleotides, 50 to 100 nucleotides, 100 to 9000 nucleotides, 100 to 8000 nucleotides, 100 to 7000 nucleotides, 100 to 6000 nucleotides, 100 to 5000 nucleotides, 100 to 4000 nucleotides, 100 to 3000 nucleotides, 100 to 2000 nucleotides, 100 to 1000 nucleotides, 100 to 500 nucleotides, 100 to 400 nucleotides, 100 to 300 nucleotides, 100 to 200 nucleotides, 100 to 150 nucleotides, 250 to 9000 nucleotides, 250 to 8000 nucleotides, 250 to 7000 nucleotides, 250 to 6000 nucleotides, 250 to 5000 nucleotides, 250 to 4000 nucleotides, 250 to 3000 nucleotides, 250 to 2000 nucleotides, 250 to 1000 nucleotides, 250 to 500 nucleotides, 250 to 400 nucleotides, or 250 to 300 nucleotides in length. In some examples, the amplicon is 50 to 5000 nucleotides in length. In some examples, the amplicon is 50 to 3000 nucleotides in length. In some examples, the amplicon is 50 to 1000 nucleotides in length. In some examples, the amplicon is at least 50 nucleotides in length. In some examples, the amplicon is at least 100 nucleotides in length.

[0138] The presence of a nucleic acid polymorphism in an amplicon can be determined, for example, by directly sequencing the amplicon, performing a restriction enzyme digest (e.g, restriction fragment length polymorphism (RFLP)), or by using a detection probe. In some aspects, detection includes a PCR method (e.g., PCR, quantitative PCR (qPCR), reverse-transcription PCR (RT-PCR), quantitative real-time reverse transcriptase PCR (qRT-PCR)), and / or sequencing. In some examples, detection includes quantitative PCR (qPCR), and / or sequencing.

[0139] PCR detection and quantification using dual-labeled fluorogenic oligonucleotide probes, commonly referred to as “TaqMan™” probes, can be performed according to the present disclosure. These probes are composed of short (e.g., 20-25 base) oligodeoxynucleotides that are labeled with two different fluorescent dyes. On the 5' terminus of each probe is a reporter dye, and on the 3' terminus of each probe a quenching dye is found. The oligonucleotide probe sequence is complementary to an internal target sequence present in a PCR amplicon. When the probe is intact, energy transfer occurs between the two fluorophores and emission from the reporter is quenched by the quencher by FRET. During the extension phase of PCR, the probe is cleaved by 5' nuclease activity of the polymerase used in the reaction, thereby releasing the reporter from the oligonucleotide-quencher and producing an increase in reporter emission intensity. TaqMan™ probes are oligonucleotides that have a label and a quencher, where the label is released during amplification by the exonuclease action of the polymerase used in amplification, providing a real time measure of amplification during synthesis. A variety of TaqMan™ reagents are commercially available, e.g., from Applied Biosystems as well as from a variety of specialty vendors such as Bioscarch Technologies. In some aspects, detection of nucleic acid polymorphisms includes use of an oligonucleotide primer or probe. In general, synthetic methods for making oligonucleotides, including probes or primers are known. For example, oligonucleotides can be synthesized chemically according to the solid phase phosphoramidite triester method described. Oligonucleotides, including modified oligonucleotides, can also be ordered from a variety of commercial sources. Nucleic acid probes to the marker loci can be cloned and / or synthesized. Any suitable label can be used with a probe. Detectable labels suitable for use with nucleic acid probes include, for example, any composition detectable by spectroscopic, radioisotopic, photochemical, biochemical, immunochemical, electrical, optical or chemical means. Useful labels include biotin for staining with labeled streptavidin conjugate, magnetic beads, fluorescent dyes, radio labels, enzymes, and colorimetric labels. Other labels include ligands which bind to antibodies labeled with fluorophores, chemiluminescent agents, and enzymes. A probe can also constitute radio labeled PCR primers that are used to generate a radio labeled amplicon. It is not intended that the nucleic acid probes be limited to any particular size, however, nucleic acid probes are typically 20-100 base pairs.

[0140] Amplification is not always required for detection of a nucleic acid polymorphism (e.g. Southern blotting or RFLP detection). Separ ate detection probes can also be omitted in amplification / detection methods, e.g., by performing a real time amplification reaction that detects product formation by modification of the relevant amplification primer upon incorporation into a product, incorporation of labeled nucleotides into an amplicon, or by monitoring changes in molecular rotation properties of amplicons as compared to unamplified precursors (e.g., by fluorescence polarization).

[0141] In some aspects, a genetic marker (e.g., nucleic acid polymorphism) is detected by sequencing a nucleic acid fragment comprising a target sequence of interest (e.g., a particular haplotype) or by whole genome sequencing (or whole transcriptome sequencing). Non-limiting examples of suitable sequencing methods include capillary electrophoresis (e.g., Sanger sequencing) and high-throughput sequencing (e.g., Illumina® or 454 Sequencing®). High-throughput sequencing includes short read or long read techniques. In some aspects, sequencing includes whole genome sequencing (e.g., sequencing the genome of a Cannabis plant). In some examples, sequencing includes targeted sequencing (sequencing of a particular nucleic acid or amplicon of interest). In some examples, sequencing includes sequencing a transcriptome (RNA-Seq) (e.g., sequencing the transcriptome of a Cannabis plant of interest). In some aspects, sequencing does not include sequencing of RNA. In some aspects, the genome is sequenced.

[0142] In general, synthetic methods for making oligonucleotides, including probes, primers, molecular beacons, PNAs, LNAs (locked nucleic acids), etc., are known. For example, oligonucleotides can be synthesized chemically according to the solid phase phosphoramidite tricstcr method described. Oligonucleotides, including modified oligonucleotides, can also be ordered from a variety of commercial sources.

[0143] IV. Genes Conferring Modified Plant Height

[0144] Also disclosed are genes and genetic modifications associated with plant height. Such genes are useful, for example, to produce (e.g., introduce a genetic modification or selectively breed) Cannabis plants having modified plant height (e.g.. taller or shorter than a parent plant). In some aspects, disclosed are engineered Cannabis plants comprising a genetic modification in a WRKY21, LBD16, SAUR51, and / or TCS1 gene, or comprising a heterologous WRKY21, LBD16, SAUR51, and / or TCS1 allele that is associated with modified plant height.

[0145] In some aspects, the genetic modification in WRKY21 that confers modified height is an amino acid substitution at position 146 with reference to the Abacus WRKY21 coding sequence. In some aspects, the amino acid substitution at position 146 of the WRKY21 coding sequence comprises a threonine substitution. In some aspects, the amino acid substitution at position 146 of the WRKY21 coding sequence comprises a serine substitution. In some aspects, the amino acid substitution at position 146 of the WRKY21 coding sequence comprises a serine to threonine substitution (S146T).

[0146] In some aspects, the genetic modification in the LBD16 gene that confers modified height is an amino acid substitution at position 9 with reference to the Abacus LB 16 coding sequence. In some aspects, the amino acid substitution at position 9 of the LBD16 coding sequence is a glycine substitution. In some aspects, the amino acid substitution at position 9 of the LBD16 coding sequence is a serine substitution. In some aspects, the amino acid substitution at position 9 of the LBD16 coding sequence is a serine to glycine substitution (S9G).

[0147] In some aspects, the genetic modification in the LBD16 gene that confers modified height is an amino acid substitution at position 130 with reference to the Abacus LBD16 coding sequence. In some aspects, the amino acid substitution at position 1 0 of the LBD16 coding sequence is a histidine substitution. In some aspects, the amino acid substitution at position 130 of the LBD16 coding sequence is a glutamine substitution. In some aspects, the amino acid substitution at position 130 of the LBD16 coding sequence is a glutamine to histidine substitution (Q130H).

[0148] In some aspects, the genetic modification in the LBD16 gene that confers modified height is an amino acid substitution at position 131 with reference to the Abacus LBD16 coding sequence. In some aspects, the substitution at position 131 of the LBD16 coding sequence introduces a preliminary stop codon (e.g., Y132stop). In some aspects, the substitution at position 131 of the LBD16 coding sequence removes a preliminary stop codon (e.g., replaces a stop codon with a codon for tyrosine, e.g., stop!32Y). In some aspects, the genetic modification in the LBD16 gene that confers modified height is an indel (insertion or deletion) between amino acid positions 141 - 144 with reference to the Abacus LBD16 coding sequence. In some aspects, the indel between amino acid positions 141 - 144 of the LBD16 coding sequence is a two amino acid deletion relative to the Abacus LBD16 coding sequence. In some aspects, the amino acid deletion or insertion between amino acid positions 141 - 144 of the LBD 16 coding sequence is a two amino acid insertion (e.g., to match the sequence of Abacus LBD16).

[0149] In some aspects, the genetic modification in SAUR51 that confers modified height is a nucleotide substitution at position -4 upstream the start codon with reference to the Abacus SAUR51 gene. In some aspects, the nucleic acid substitution is an adenine substitution. In some aspects, the nucleic acid substitution is a guanine substitution. In some aspects, the nucleic acid substitution at position -4 of the SAUR51 gene comprises a G to A substitution. In some aspects, the nucleic acid substitution at position - 4 of the SAUR51 gene comprises an A to G substitution.

[0150] In some aspects, the genetic modification in SAUR51 that confers modified height is a nucleotide substitution at position -37 upstream the start codon with reference to the Abacus SAUR51 gene. In some aspects, the nucleic acid substitution is an adenine substitution. In some aspects, the nucleic acid substitution is a thymine substitution. In some aspects, the nucleic acid substitution at position -37 of the SAUR51 gene is a T to A substitution. In some aspects, the nucleic acid substitution at position -37 of the SAUR51 gene is an A to T substitution

[0151] In some aspects, the genetic modification in SAUR51 that confers modified height is a nucleotide substitution at position -94 upstream the start codon with reference to the Abacus SAUR51 gene. In some aspects, the substitution is a guanine substitution. In some aspects, the substitution is a cytosine substitution. In some aspects, the nucleic acid substitution at position -94 of the SAUR51 gene is a C to G substitution. In some aspects, the nucleic acid substitution at position -94 of the SAUR51 gene is a G to C substitution.

[0152] In some aspects, the genetic modification in TCS1 that confers modified height is an amino acid substitution at position 802 with reference to the Abacus TCS1 coding sequence. In some aspects, the amino acid substitution at position 802 of the TCS1 coding sequence is an asparagine substitution. In some aspects, the amino acid substitution at position 802 of the TCS1 coding sequence is a serine substitution. In some aspects, the amino acid substitution at position 802 of the TCS1 coding sequence is a S to N substitution (S802N).

[0153] Further contemplated by this disclosure are heterologous WRKY21, LBD16, SAUR51, and / or TCS1 genes comprising a polymorphism described herein that confers modified plant height. A heterologous WRKY21, LBD 16, SAUR51, and / or TCS1 gene can be introduced, for example, into a Cannabis plant, thereby producing an engineered Cannabis plant having modified plant height. In some aspects, the WRKY21, LBD16, SAUR51, and / or TCS1 gene includes at least 70% sequence identity, for example, at least 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% identity to a WRKY21, LBD16, SAUR51, and / or TCS1 gene sequence disclosed herein (see, e.g., SEQ ID NOs: 228, 229, 233-236, 246-249, 250, and 251) and further comprises a polymorphism disclosed herein conferring modified plant height. In some aspects, the WRKY21, LBD16, SAUR51, and / or TCS1 gene includes at least 90% sequence identity to a WRKY21, LBD16, SAUR51, and / or TCS1 gene disclosed herein see, e.g., SEQ ID NOs: 228, 229, 233-236, 246-249, 250, and 251) and further comprises a polymorphism disclosed herein conferring modified plant height. In some aspects, the WRKY21, LBD16, SAUR51, and / or TCS1 gene includes at least 95% sequence identity to a WRKY21, LBD16, SAUR51, and / or TCS1 gene sequence disclosed herein (see, e.g., SEQ ID NOs: 228, 229, 233-236, 246-249, 250, and 251) and further comprises a polymorphism disclosed herein conferring modified plant height. In some aspects, the WRKY21, LBD16, SAUR51, and / or TCS1 gene includes at least 98% sequence identity to a WRKY21, LBD16, SAUR51, and / or TCS1 gene sequence disclosed herein (see, e.g., SEQ ID NOs: 228, 229, 233-236, 246-249, 250, and 251) and further comprises a polymorphism disclosed herein conferring modified plant height.

[0154] In some aspects, the WRKY21, LBD16, and / or TCS1 gene encode a protein product having at least 70% sequence identity, for example, at least 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%. 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%. 95%. 96%. 97%. 98%. 99%, or 100% sequence identity, to a WRKY21, LBD16, and / or TCS1 protein sequence disclosed herein (see, e.g., SEQ ID NOs: 230, 231, 237-240, 252, or 253) and further comprise a polymorphism conferring modified plant height. In some aspects, the WRKY21, LBD16, and / or TCS1 gene encode a protein product having at least 90% sequence identity to a WRKY21, LBD16, and / or TCS1 protein sequence disclosed herein (see, e.g., SEQ ID NOs: 230, 231, 237-240, 252, or 253). In some aspects, the WRKY21 , LBD16, and / or TCS1 gene encode a protein product having at least 95% sequence identity to a WRKY21, LBD16, and / or TCS1 protein sequence disclosed herein (see, e.g., SEQ ID NOs: 230, 231, 237-240, 252, or 253). In some aspects, the WRKY21, LBD16, and / or TCS1 gene encode a protein product having at least 98% sequence identity to a WRKY21, LBD16, and / or TCS1 protein sequence disclosed herein (see, e.g., SEQ ID NOs: 230, 231, 237-240, 252, or 253).

[0155] Local sequence alignment programs such as BLAST® can be used to compare specific regions of two sequences. A BLAST® comparison of two sequences results in an E-value, or expectation value, that represents the number of different alignments with scores equivalent to or better than the raw alignment score, S, that arc expected to occur in a database search by chance. The lower the E value, the more significant the match. Because database size is an element in E-value calculations, E-values obtained by BLASTing against public databases, such as GENBANK, have generally increased over time for any given query / entry match. In setting criteria for confidence of polypeptide function prediction, a "high" BLAST® match is considered herein as having an E-value for the top BLAST® hit of less than IE-30; a medium BLASTX E-value is IE-30 to IE-8; and a low BLASTX E-value is greater than IE-8. The protein function assignment in the present disclosure is determined using combinations of E- values, percent identity, query coverage and hit coverage. Query coverage refers to the percent of the query sequence that is represented in the BLAST® alignment. Hit coverage refers to the percent of the database entry that is represented in the BLAST® alignment. In one aspect of the present disclosure, function of a query polypeptide is inferred from function of a protein homolog where either (1) hit_p<le-30 or % identity >35% AND query _coverage >50% AND hit_coverage >50%, or (2) hit_p<le-8 AND query_coverage >70% AND hit_coverage >70%. The following abbreviations are produced during a BLAST® analysis of a sequence. SEQ_NUM provides the SEQ ID NO for the listed recombinant polynucleotide sequences. CONTIG_ID provides an arbitrary sequence name taken from the name of the clone from which the cDNA sequence was obtained. PROTEIN_NUM provides the SEQ ID NO for the recombinant polypeptide sequence NCBI_GI provides the GenBank ID number for the top BLAST® hit for the sequence. The top BLAST® hit is indicated by the National Center for Biotechnology Information GenBank Identifier number. NCBI_GI_DESCRIPTION refers to the description of the GenBank top BLAST® hit for sequence. E_VALUE provides the expectation value for the top BLAST® match. MATCH_LENGTH provides the length of the sequence which is aligned in the top BLAST® match TOP_HIT_PCT_IDENT refers to the percentage of identically matched nucleotides (or residues) that exist along the length of that portion of the sequences which is aligned in the top BLAST® match. CAT_TYPE indicates the classification scheme used to classify the sequence. GO_BP=Gene Ontology Consortium-biological process; GO_CC=Gene Ontology Consortium— cellular component;

[0156] GO_MF=Gene Ontology Consortium molecular function; KEGG=KEGG functional hierarchy (KEGG=Kyoto Encyclopedia of Genes and Genomes); EC=Enzyme Classification from ENZYME data bank release 25.0; POI=Pathways of Interest. CAT_DESC provides the classification scheme subcategory to which the query sequence was assigned. PRODUCT_CAT_DESC provides the FunCAT annotation category to which the query sequence was assigned. PRODUCT_HIT_DESC provides the description of the BLAST® hit which resulted in assignment of the sequence to the function category provided in the cat_desc column. HIT_E provides the E value for the BLAST® hit in the hit_desc column. PCT_IDENT refers to the percentage of identically matched nucleotides (or residues) that exist along the length of that portion of the sequences which is aligned in the BLAST® match provided in hit_dcsc. QRY_RANGE lists the range of the query sequence aligned with the hit. HIT_RANGE lists the range of the hit sequence aligned with the query, provides the percent of query sequence length that matches QRY_CVRG provides the percent of query sequence length that matches to the hit (NCBI) sequence in the BLAST® match (% qry cvrg=(match length / query total length)xlOO). HIT_CVRG provides the percent of hit sequence length that matches to the query sequence in the match generated using BLAST® (% hit cvrg=(match lengthy hit total length)xlOO).

[0157] Alternative programs that can be used to determine percent identity between two polynucleotides or amino acid sequences include, but are not limited to, AlignX™ alignment program of the Vector NTI™ suite (Invitrogen™, Carlsbad, Calif.) and MegAlign™ program of the LASERGENE™ bioinformatics computing suite (DNASTAR™. Madison, Wis.). AlignX™ and MegAlign™ alignment programs include global sequence alignment tools for polynucleotides or proteins.

[0158] Also disclosed are methods of producing an engineered Cannabis plant having modified plant height, including introducing an allele of WRKY21, LBD16, SAUR51, and / or TCS1 that is associated with or confers modified plant height. An allele can be introduced, for example, by genetically modifying an endogenous WRKY21, LBD16, SAUR51, and / or TCS1 gene to a sequence that is associated with plant height, or by introducing a heterologous WRKY21, LBD16, SAUR51, and / or TCS1 allele comprising a polymorphism described herein that is associated with or confers modified plant height.

[0159] In some aspects, the methods include introducing an allele of WRKY21, LBD16, SAUR51, and / or TCS 1 into a Cannabis plant conferring tall plant height. In some aspects, the modification includes introducing an allele of WRKY21, LBD16, SAUR51, and / or TCS1 conferring short plant height. Introducing an allele includes introducing a full-length gene (e.g., expressing an exogenous gene) as well as introducing a mutation in an endogenous gene, such that the sequence is changed to a desired sequence (e.g., mutagenesis or gene editing). Sequences of exemplary alleles associated with or conferring modified plant height are provided herein, for example, see Example 2. In some examples, introducing an allele of WRKY21, LBD16, SAUR51, and / or TCS1 disclosed herein modifies height of the engineered Cannabis plant, for example, producing engineered Cannabis plants that are shorter or taller than a parent variety. In some aspects, an allele of WRKY21 , LBD16, SAUR51 , and / or TCS 1 disclosed herein is introduced through selective breeding (e.g., MAS). In some aspects, an allele of WRKY21, LBD16, SAUR51 , and / or TCS 1 disclosed herein is introduced through genetic modification.

[0160] Introducing a genetic modification into a WRKY21, LBD16, SAUR51, and / or TCS1 gene in a Cannabis plant includes introducing a beneficial allele into a plant cell, plant tissue, or plant part. In some aspects, a genetic modification in WRKY21, LBD16, SAUR51, and / or TCS1 is introduced into a Cannabis plant cell or tissue, and a Cannabis plant is regenerated from the cell or tissue.

[0161] In some aspects, genetic modification of WRKY21 causes a change in protein structure of WRKY21, resulting in increased levels of indolc-3-acctic acid (IAA) and / or gibberellin 3 (GA3) and increased plant height. In some aspects, genetic modification of LBD16 disrupts LBD16 function, thereby resulting in increased plant height. In some aspects, genetic modification of LBD16 improves LBD16 function, thereby resulting in decreased plant height. In some aspects, genetic modification of SAUR51 decreases SAUR51 expression, thereby increasing plant height. In some aspects, genetic modification of SAUR51 increases SAUR51 expression, thereby decreasing plant height. In some aspects, genetic modification of TCS1 improves TCS1 function, thereby increasing plant height. In some aspects, genetic modification of TCS1 decreases TCS1 function, thereby decreasing plant height.

[0162] In some aspects, the methods include reducing expression or activity (e.g., knocking out or knocking down) of one or more of WRKY21, LBD16, SAUR51, and / or TCS1, relative to Cannabis Abacus or an unmodified parent plant. In some aspects, expression is reduced by at least 5%, for example, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 98%, or 100%.

[0163] In some aspects, the methods include increasing expression or activity (e.g., overexpressing or upregulating) of one or more of WRKY21, LBD16, SAUR51, and / or TCS1, relative to Cannabis Abacus or an unmodified parent plant. In some aspects, expression or activity is increased by at least 5%, for example, at least 10%, at least 25%, at least 50%, at least 75%, at least 80%, at least 90%, at least 100%, at least 150%, at least 200%, at least 400%, or more.

[0164] A Cannabis gene can be modified, for example, by introducing a nucleic acid substitution, insertion, or deletion into the gene. The modification can be homozygous or heterozygous in the modified plant. In some examples, the modification is heterozygous in the modified plant. In some examples, the modification is homozygous in the modified plant. In some aspects, the genetic modification is introduced by mutagenesis or a gene editing technique (e.g., RNAi, CRISPR / Cas9, ZFN, or TALEN based systems).

[0165] Many methods of gene editing have been described and can be used with the present disclosure. Such methods can use various genetic or biochemical tools to over-express or suppress one or more genes or regulatory elements thereof, for example and without limitation, guide RNAs, nucleases, antisense DNA, antisense RNA, ribozymes, DNAzymes, locked nucleic acids (LNA), and aptamers.

[0166] Clustered regularly interspaced short palindromic repeats (CRISPR) / CRISPR associated protein (Cas) systems include genome engineering tools based on the bacterial CRISPR / Cas prokaryotic adaptive immune system. CRISPR technology allows targeted cleavage of genomic DNA guided by a customizable small noncoding RNA, resulting in gene modifications by both non-homologous end joining (NHEJ) and homology-directed repair (HDR) mechanisms (Belhaj K. et al., 2013. Plant Methods 2013, 9:39). By including a DNA repair template, a target gene can be edited to include a specific sequence of interest. In some aspects, a CRISPR system is used to introduce a genetic modification in WRKY21, LBD16, SAUR51, and / or TCS1. In some aspects, the CRISPR system is a CRISPR / Cas9 system. CRISPR-based gene editing systems need not be limited to Cas9 systems, as other analogous editing enzymes exist, e.g., MAD7.

[0167] RNA interference (RNAi) is an exemplary method for reducing gene function in plants. RNAi is mediated by RNA-induced silencing complex (RISC), a sequence-specific, multicomponent nuclease that destroys messenger RNAs homologous to the silencing trigger. RISC is known to contain short RNAs (approximately 22 nucleotides) derived from the double-stranded RNA trigger. The short-nucleotide RNA sequences are homologous to the target gene that is being suppressed. Thus, the short-nucleotide sequences appear to serve as guide sequences to instruct a multicomponent nuclease, RISC, to destroy the specific mRNAs. The dsRNA used to initiate RNAi, may be isolated from a native source or produced by known means, e.g., transcribed from DNA. Plasmids and vectors for generating RNAi molecules against target sequence are now readily available from commercial sources.

[0168] DNAzyme molecules, enzymatic oligonucleotides, and methods of mutagenesis are also commonly used to introduce genetic modifications or modulate gene function. Any available mutagenesis procedure can be used, including but not limited to, site-directed point mutagenesis, random point mutagenesis, in vitro or in vivo homologous recombination (DNA shuffling), uracil-containing templates, oligonucleotide-directed mutagenesis, phosphorothioate-modified DNA mutagenesis, mutagenesis using gapped duplex DNA, point mismatch repair, repair-deficient host strains, restrictionselection and restriction-purification, deletion mutagenesis, total gene synthesis, double-strand break repair, zinc-finger nucleases (ZFN), transcription activator-like effector nucleases (TALEN), any other mutagenesis procedure known in the art.

[0169] In some aspects, modifying a Cannabis gene includes a step of introducing a nucleic acid into a Cannabis host cell (e.g., transforming a WRKY21, LBD16, SAUR51, and / or TCS1 gene including a genetic modification associated with plant height). In some aspects, the nucleic acid is included on a vector (e.g., an expression cassette). Methods for transformation of plant cells are well known in the art, and selection of an appropriate transformation technique can be determined by the practitioner. Suitable methods include, but are not limited to, electroporation, liposome-mediated transformation, polyethylene glycol (PEG) mediated transformation, virus mediated transformation, micro-injection, particle bombardment, and Agrobacterium tumefaciens mediated transformation. Transformation includes both stable or transient introduction of a nucleic acid.

[0170] In planta transformation techniques (e.g., vacuum-infiltration, floral spray, or floral dip procedures) are also useful to introduce nucleic acids (e.g., expression cassettes typically in an Agrobactcrium vector) into meristematic or gcrmlinc cells of a whole plant. Such methods provide a simple and reliable method of obtaining transformants at high efficiency while avoiding the use of tissue culture (see, e.g., Bechtold et at. 1993 C. R. Acad. Sci. 316:1194-1199; Chung et at. 2000 Transgenic Res. 9:471-476; Clough et al. 1998 Plant J. 16:735-743; and Desfeux et at. 2000 Plant Physiol 123:895- 904). In such examples, seed produced by the plant comprise the transformed nucleic acids. Introduced nucleic acids can further encode a selectable marker, allowing the seed to be selected based on the ability to germinate under conditions that inhibit germination of the untransformed seed.

[0171] If transformation techniques require use of tissue culture, transformed cells may be regenerated into plants in accordance with known techniques. Regenerated plants may then be grown and crossed with the same or different plant varieties using traditional breeding techniques to produce seed, which are then selected under the appropriate conditions.

[0172] A genome editing protein itself may be introduced into the plant cell. For example, an expression cassette can be used to express nucleic acids or proteins necessary for gene editing. In some aspects, the expression vector is integrated into the genome of the plant cells, in which case subsequent generations will express the genome editing proteins unless the expression vector is bred out. Alternatively, in some aspects the expression cassette is not integrated into the genome of the plant’s cell, in which case the genome editing protein is transiently expressed in the transformed cells and is not expressed in subsequent generations.

[0173] In some aspects, the genome editing protein is provided in sufficient quantity to modify the cell but does not persist after a contemplated period of time has passed or after one or more cell divisions. In such aspects, the genome editing proteins are provided transiently, for example, the genome editing protein is prepared in vitro prior to introduction to a plant cell using well known recombinant expression systems (bacterial expression, in vitro translation, yeast cells, insect cells and the like). After expression, the protein is isolated, refolded if needed, purified and optionally treated to remove any purification tags, such as a His-tag. Once crude, partially purified, or more completely purified genome editing proteins are obtained, they are introduced to a plant cell via a transformation method (e.g., electroporation, microinjection, particle bombardment, or chemical transfection).

[0174] The genome editing protein can also be expressed in Agrobacterium as a fusion protein, fused to an appropriate domain of a virulence protein that is translocated into plants (e.g., VirD2, VirE2, VirE2 and VirF). The Vi r protein fused with the genome editing protein travels to the plant cell's nucleus, where the genome editing protein would produce the desired double stranded break in the genome of the cell. (see Vergunst et al. 2000 Science 290:979-82).

[0175] In some aspects gene editing is used with plant breeding to develop plants having modified plant height. In a non-limiting example, an endogenous nucleic acid sequence is replaced with an exogenous nucleic acid sequence conferring modified plant height. The plant is then crossed or selfed, thereby producing a plurality of progeny. One or more progeny plants comprising the nucleic acid sequence conferring modified plant height are selected, thereby selecting plants having conferring modified plant height.

[0176] In some aspects, the plant of any of the methods disclosed herein is a Cannabis plant. In some examples, the Cannabis plant is Cannabis sativa, Cannabis indica, or Cannabis ruderalis. In a nonlimiting example, the plant of any of the methods disclosed herein is Cannabis sativa.

[0177] N. Cannabis Breeding

[0178] Cannabis is an important and valuable crop. Thus, a continuing goal of Cannabis plant breeders is to develop stable, high yielding Cannabis cultivars that are agronomically sound. To accomplish this goal, the Cannabis breeder preferably selects and develops Cannabis plants with traits that result in superior cultivars. The plants described herein can be used to produce new plant varieties. In some aspects, the plants are used to develop new, unique, and superior varieties or hybrids with desired phenotypes (e.g.. modified height).

[0179] The development of commercial Cannabis cultivars requires the development of Cannabis varieties, the crossing of these varieties, and the evaluation of the crosses. Pedigree breeding and recurrent selection breeding methods may be used to develop cultivars from breeding populations. Breeding programs may combine desirable traits from two or more varieties or various broad-based sources into breeding pools from which cultivars are developed by selfing and selection of desired phenotypes. The new cultivars may be crossed with other varieties and the hybrids from these crosses are evaluated to determine which have commercial potential.

[0180] Details of existing Cannabis plants varieties and breeding methods are described in Potter et al. (2011, World Wide Weed: Global Trends in Cannabis Cultivation and Its Control), Holland (2010, The Pot Book: A Complete Guide to Cannabis, Inner Traditions / Bear & Co, ISBN1594778981, 9781594778988), Green I (2009, The Cannabis Grow Bible: The Definitive Guide to Growing Marijuana for Recreational and Medical Use, Green Candy Press, 2009, ISBN 1931 160589, 9781931 160582), Green II (2005, The Cannabis Breeder's Bible: The Definitive Guide to Marijuana Genetics, Cannabis Botany and Creating Strains for the Seed Market, Green Candy Press, 1931160279, 9781931160278), Starks (1990, Marijuana Chemistry: Genetics, Processing & Potency, ISBN 0914171399, 9780914171393), Clarke (1981, Marijuana Botany, an Advanced Study: The Propagation and Breeding of Distinctive Cannabis, Ronin Publishing, ISBN 091417178X, 9780914171782), Short (2004, Cultivating Exceptional Cannabis: An Expert Breeder Shares His Secrets, ISBN 1936807122, 9781936807123), Cervantes (2004, Marijuana Horticulture: The Indoor / Outdoor Medical Grower’s Bible, Van Patten Publishing, ISBN 187882323X, 9781878823236), Franck ct al. (1990, Marijuana Grower's Guide, Red Eye Press, ISBN 0929349016, 9780929349015), Grotenhermen and Russo (2002, Cannabis and Cannabinoids: Pharmacology, Toxicology, and Therapeutic Potential, Psychology Press, ISBN 0789015080, 9780789015082), Rosenthal (2007, The Big Book of Buds: More Marijuana Varieties from the World's Great Seed Breeders, ISBN 1936807068, 9781936807062), Clarke, RC (Cannabis Evolution and Ethnobotany 2013 (In press)), King, J (Cannabible Vols 1-3, 2001-2006), and four volumes of Rosenthal's Big Book of Buds series (2001, 2004, 2007, and 2011).

[0181] Pedigree selection, where both single plant selection and MAS selection practices are employed, may be used for the generating varieties as described herein. Pedigree selection, also known as the “Vilmorin system of selection,” is described in Fehr, Walter; Principles of Cultivar Development, Volume I, Macmillan Publishing Co. Pedigree breeding is used commonly for the improvement of selfpollinating crops or inbred lines of cross-pollinating crops. Two par ents which possess favorable, complementary traits are crossed to produce an Fl. An F2 population is produced by selfing one or several Fl's or by intercrossing two Fl's (sib mating). Selection of the best individuals usually begins in the F2 population; then, beginning in the F3, the best individuals in the best families are usually selected. Replicated testing of families, or hybrid combinations involving individuals of these families, often follows in the F4 generation to improve the effectiveness of selection for traits with low heritability. At an advanced stage of inbreeding (e.g., F6 and F7), the best lines or mixtures of phenotypically similar lines are tested for potential release as new cultivars.

[0182] Choice of breeding or selection methods depends on the mode of plant reproduction, the heritability of the trait(s) being improved, and the type of cultivar used commercially (e.g.. Fl hybrid cultivar, pureline cultivar, etc.). For highly heritable traits, a choice of superior individual plants evaluated at a single location will be effective, whereas for traits with low heritability, selection should be based on mean values obtained from replicated evaluations of families of related plants. Popular selection methods commonly include pedigree selection, modified pedigree selection, mass selection, and recurrent selection.

[0183] Mass and recurrent selections can be used to improve populations of either self- or crosspollinating crops. A genetically variable population of heterozygous individuals may be identified or created by intercrossing several different parents. The best plants may be selected based on individual superiority, outstanding progeny, or excellent combining ability. Preferably, the selected plants are intercrossed to produce a new population in which further cycles of selection are continued.

[0184] B ackcross breeding has been used to transfer genes for a simply inherited, highly heritable trait into a desirable homozygous cultivar or line that is the recurrent parent. The source of the trait to be transferred is called the donor parent. The resulting plant is expected to have the attributes of the recurrent parent (e.g., cultivar) and the desirable trait transferred from the donor parent. After the initial cross, individuals possessing the phenotype of the donor parent may be selected and repeatedly crossed (backcrossed) to the recurrent parent. The resulting plant is expected to have the attributes of the recurrent parent (e.g.. cultivar) and the desirable trait transferred from the donor parent.

[0185] A single-seed descent procedure refers to planting a segregating population, harvesting a sample of one seed per plant, and using the one-seed sample to plant the next generation. When the population has advanced from the F2 to the desired level of inbreeding, the plants from which lines are derived will each trace to different F2 individuals. The number of plants in a population declines each generation due to failure of some seeds to germinate or some plants to produce at least one seed. As a result, not all of the F2 plants originally sampled in the population will be represented by a progeny when generation advance is completed.

[0186] Mutation breeding is another method of introducing new traits into Cannabis varieties. Mutations that occur spontaneously or are artificially induced can be useful sources of variability for a plant breeder. The goal of artificial mutagenesis is to increase the rate of mutation for a desired characteristic. Mutation rates can be increased by many different means including temperature, long-term seed storage, tissue culture conditions, radiation (such as X-rays, Gamma rays, neutrons. Beta radiation, or ultraviolet radiation), chemical mutagens (such as base analogs like 5 -bromo-uracil), antibiotics, alkylating agents (such as sulfur mustards, nitrogen mustards, epoxides, ethyleneamines, sulfates, sulfonates, sulfones, or lactones), azide, hydroxylamine, nitrous acid or acridines. Once a desired trait is observed through mutagenesis the trait may then be incorporated into existing germplasm by traditional breeding techniques. Details of mutation breeding can be found in Principles of Cultivar Development by Fehr, Macmillan Publishing Company, 1993.

[0187] The complexity of inheritance also influences the choice of the breeding method. Backcross breeding may be used to transfer one or a few favorable genes for a highly heritable trait into a desirable cultivar. This approach has been used extensively for breeding disease-resistant cultivars. Various recurrent selection techniques are used to improve quantitatively inherited traits controlled by numerous genes. The use of recurrent selection in self-pollinating crops depends on the ease of pollination, the frequency of successful hybrids from each pollination, and the number of hybrid offspring from each successful cross.

[0188] Additional breeding methods have been described, e.g., methods discussed in Chahal and Gosal (Principles and procedures of plant breeding: biotechnological and conventional approaches, CRC Press, 2002, ISBN 084931321X, 9780849313219), Taji et al. (In vitro plant breeding, Routledge, 2002, ISBN 156022908X, 9781560229087), Richards (Plant breeding systems, Taylor & Francis US, 1997, ISBN 0412574500, 9780412574504), Hayes (Methods of Plant Breeding, Publisher: READ BOOKS, 2007, ISBN1406737062, 9781406737066). Cannabis genome has been sequenced (Bakcl et al., The draft genome and transcriptome of Cannabis sativa, Genome Biology, 12(10):R102, 2011). Molecular markers for Cannabis plants are described in Datwyler et al. (Genetic variation in hemp and marijuana Cannabis sativa L.) according to amplified fragment length polymorphisms, J Forensic Sci. 2006 March; 51 (2):371 -5), Pinarkara et al., (RAPD analysis of seized marijuana (Cannabis sativa L.) in Turkey, Electronic Journal of Biotechnology, 12(1), 2009), Hakki et al., (Inter simple sequence repeats separate efficiently hemp from marijuana (Cannabis sativa L.), Electronic Journal of Biotechnology, 10(4), 2007), Datwyler et al., (Genetic Variation in Hemp and Marijuana (Cannabis sativa L.) According to Amplified Fragment Length Polymorphisms, J Forensic Sci, March 2006, 51(2):371-375), Gilmore et al. (Isolation of microsatellite markers in Cannabis sativa L. (marijuana), Molecular Ecology Notes, 3(1): 105-107, March 2003), Pacifico et al., (Genetics and marker-assisted selection of chemotype in Cannabis sativa L.), Molecular Breeding (2006) 17:257-268), and Mendoza et al., (Genetic individualization of Cannabis sativa by a short tandem repeat multiplex system, Anal Bioanal Chem (2009) 393:719-726).

[0189] The production of double haploids can also be used for the development of homozygous varieties in a breeding program. Double haploids are produced by first making haploid lines via one of several methods (e.g., a haploid inducer line, androgenesis, or gynogenesis) and then the doubling of a set of chromosomes to produce a completely homozygous individual. For example, see, Wan et al., Theor. Appl. Genet., 77:889-892, 1989.

[0190] Marker assisted selection (MAS) can be used in any of the methods disclosed herein to produce plants with desired traits. MAS is a powerful shortcut to selecting for desired phenotypes and for introgressing desired traits into cultivars (e.g., introgressing desired traits into elite lines). MAS is easily adapted to high throughput molecular analysis methods that can quickly screen large numbers of plant or germplasm genetic material for the markers of interest and is much more cost effective than raising and observing plants for visible traits.

[0191] Introgression refers to the transmission of a desired allele of a genetic locus from one genetic background to another, which is significantly assisted through MAS. For example, introgression of a desired allele at a specified locus can be transmitted to at least one progeny via a sexual cross between two parents of the same species, where at least one of the parents has the desired allele in its genome. Alternatively, for example, transmission of an allele can occur by recombination between two donor genomes, e.g., in a fused protoplast, where at least one of the donor protoplasts has the desired allele in its genome. The desired allele can be, e.g., a selected allele of a marker, a QTL, a transgene, or the like.

[0192] The introgression of one or more desired loci from a donor line into another is achieved via repeated backcrossing to a recurrent parent accompanied by selection to retain one or more loci from the donor parent. Markers associated with plant height may be assayed in progeny and those progeny with one or more desired markers arc selected for advancement. In another aspect, one or more markers can be assayed in the progeny to select for plants with the genotype of the agronomically elite parent. This disclosure anticipates that trait introgression will require more than one generation, wherein progeny are crossed to the recurrent (agronomically elite) parent or selfed. Selections are made based on the presence of one or more plant height markers and can also be made based on the recurrent parent genotype, wherein screening is performed on a genetic marker and / or phenotype basis. In another aspect, markers of this disclosure can be used in conjunction with other markers, for example, at least one on each chromosome of the Cannabis genome.

[0193] Genetic markers are used to identify plants that contain a desired genotype at one or more loci, and that are expected to transfer the desired genotype, along with a desired phenotype to their progeny. Genetic markers can be used to identify plants containing a desired genotype at one locus, or at several unlinked or linked loci (e.g., a haplotype), and that would be expected to transfer the desired genotype, along with a desired phenotype to their progeny. The present disclosure provides the means to identify plants that exhibit desired plant height traits by identifying plants having plant height markers.

[0194] In general, MAS uses polymorphic markers that have been identified as having a significant likelihood of co-segregation with a desired trait. Such markers are presumed to map near a gene or genes that give the plant its desired phenotype, and are considered indicators for the desired trait, and are termed QTL markers. Plants are tested for the presence or absence of a desired allele in the QTL marker.

[0195] Genomic selection is another form of marker-assisted selection in which a very large number of genetic markers covering the whole genome are used. With genomic selection, all SNPs are included, each with a different level of effect, in a model to explain the variation of the trait. Genomic selection is based on the analysis of many SNPs, for example tens of thousands or even millions of SNPs. This high number of SNP markers is used as input in a genomic prediction formula that predicts the desired phenotype for MAS.

[0196] Identification of plants or germplasm that include a marker locus or marker loci linked to a desired trait or traits provides a basis for performing MAS. Plants that comprise favorable markers or favorable alleles are selected for, while plants that comprise markers or alleles that are negatively correlated with the desired trait can be selected against. Desired markers and / or alleles can be introgressed into plants having a desired (e.g., elite or exotic) genetic background to produce an introgressed plant or germplasm having the desired trait. In some aspects, it is contemplated that a plurality of markers for desired traits are sequentially or simultaneously selected and / or introgressed. The combinations of markers that are selected for in a single plant are not limited, and can include any combination of markers disclosed herein or any marker linked to the markers disclosed herein, or any markers located within the QTL intervals defined herein.

[0197] In some aspects, a first Cannabis plant or germplasm exhibiting a desired trait (the donor) can be crossed with a second Cannabis plant or germplasm (the recipient, e.g., an elite or exotic Cannabis, depending on characteristics that are desired in the progeny) to create an introgressed Cannabis plant or germplasm as part of a breeding program. In some aspects, the recipient plant can also contain one or more loci associated with one or more desired traits, which can be qualitative or quantitative trait loci. In another aspect, the recipient plant can contain a transgene.

[0198] MAS, as described herein, using additional markers flanking either side of the DNA locus provide further efficiency because an unlikely double recombination event would be needed to simultaneously break linkage between the locus and both markers. Moreover, using markers tightly flanking a locus, a practitioner can reduce linkage drag by more accurately selecting individuals that have less of the potentially deleterious donor parent DNA. Any marker linked to or among the chromosome intervals described herein can thus find use within the scope of this disclosure.

[0199] Similarly, by identifying plants that do not have a desired plant height can be identified and eliminated from subsequent crosses. These marker loci can be introgressed into any desired genomic background, germplasm, plant, line, variety, etc., as part of an overall MAS breeding program designed to increase or decrease plant height. The present disclosure also provides chromosome QTL intervals and identified trait loci that can be used in MAS to select plants that demonstrate modified plant height. The QTL intervals and trait loci can also be used to counter-select plants that do not have a desired height.

[0200] Thus, the present disclosure permits a practitioner to detect the presence of modified height genotypes in the genomes of Cannabis plants as part of a MAS program, as described herein. In one aspect, a breeder ascertains the genotype at one or more markers for a parent having a desired height trait, and the genotype at one or more markers for a parent without a desired height trait. A breeder can then reliably track the inheritance of the desired phenotype through subsequent populations derived from crosses between the two parents by genotyping offspring with the markers used on the parents and comparing the genotypes at those markers with those of the parents. Depending on how tightly linked markers are with the trait, progeny that share genotypes can be reliably predicted to have the desirable phenotype (or conversely the undesirable phenotype). Thus, the laborious, inefficient, and potentially inaccurate process of manually phenotyping the progeny for height traits is avoided.

[0201] Closely linked markers flanking the locus of interest that have alleles in linkage disequilibrium, thus markers described herein, such as those listed in Table 8, as well as other markers genetically linked to the same chromosome interval, may be used to select for Cannabis plants having modified height. Often, a set of these markers will be used, (c.g., 2 or more, 3 or more, 4 or more, 5 or more) in the flanking regions of a locus. Optionally, as described above, a marker flanking or within the actual locus may also be used. The parents and their progeny may be screened for these sets of markers, and the markers that arc polymorphic between the two parents used for selection. In an introgression program, this allows for selection of the gene or locus genotype at the more proximal polymorphic markers and selection for the recurrent parent genotype at the more distal polymorphic markers.

[0202] In a non-limiting example, MAS is used to select and / or breed one or more Cannabis plants comprising modified plant height, the method comprising: (i) obtaining nucleic acids from a sample plant or its germplasm; (ii) detecting one or more markers that indicate modified plant height (eg., indicate tall or short plant height), and (iii) selecting the one or more plants comprising the one or more markers that indicate modified plant height. Markers that indicate modified plant height and suitable methods of detecting such markers, are described herein. Modified plant height includes, for example, an increase or decrease in vegetative height, harvest height, and / or stretch relative to a control. In some aspects, modified plant height is an increase in vegetative height, harvest height, and / or stretch relative to a control. In some aspects, modified plant height is an increase in harvest height relative to a control. In some aspects, the control is Abacus or a parent plant. In some aspects, the one or more markers that indicate modified plant height are associated with a tall plant height phenotype (eg.. a harvest height of 95 cm or more). In some aspects, the one or more markers that indicate modified plant height are associated with a short plant height phenotype (eg., a harvest height of less than 95 cm). In some aspects, Cannabis plants including the one or more markers that indicate modified plant height are selected for further analysis, propagation, and / or to make a product.

[0203] In some aspects, the Cannabis plant including the one or more markers that indicate modified plant height is further crossed with another Cannabis plant to obtain one or more progeny plants comprising the one or more genetic markers indicating modified plant height. Crossing includes, for example, selfing, sibling crossing, outcrossing, and backcrossing. In some aspects, the progeny plants have a tall height phenotype. In some aspects, the progeny plants have a short height phenotype. In some aspects, the one or more progeny plants have modified plant height relative to a control (eg., Abacus or a parent plant). In some aspects, the progeny plants are Fl progeny.

[0204] VI. Plants and Products

[0205] Disclosed are engineered Cannabis plants comprising a genetic modification conferring a desired plant height trait or an exogenous gene conferring a desired height trait. In some aspects, the engineered Cannabis plants include a genetic modification in WRKY21, LBD16, SAUR51, and / or TCS1 that is associated with or confers a particular plant height trait, or include an allele of WRKY21, LBD16, SAUR51, and / or TCS1 that is associated with or confers a desired height trait (eg., tall or short height trait). Exemplary alleles of WRKY21, LBD16, SAUR51, and / or TCS1 that confer height traits (eg., tall or short height trait) arc disclosed herein. Such genetic modifications or alleles can be introduced into Cannabis plants using genetic engineering and / or breeding techniques described herein. Further disclosed are Cannabis plants made by any of the methods disclosed herein. Material derived from the Cannabis plants disclosed herein, including seed, flowers, trichomes, or other tissues or cells (including protoplasts); are encompassed by the disclosure. In some aspects, the Cannabis plant is Cannabis sativa, Cannabis indica, or Cannabis ruderalis. In a non-limiting example, the Cannabis plant is Cannabis sativa.

[0206] Further disclosed are products including or derived from any of the Cannabis plants disclosed herein, including plants made by any of the methods disclosed herein. The product may be any product known in the Cannabis arts, and can include, but is not limited to, a kief, hashish, bubble hash, an edible product, extract, solvent reduced oil, sludge, e-juice, or tincture. Kief refers to a composition of concentrated Cannabis trichomes, which are accumulated by being sifted from Cannabis flowers or buds using a mesh screen or sieve. Hashish (or hash) refers to a compressed or purified preparation from Cannabis tissue containing trichomes (e.g., flowers). Bubble hash refers to a solid concentration of Cannabis trichomes made from a solventless extraction method. As used herein. Cannabis sludges are solvent-free Cannabis extracts made via multigas extraction including the refrigerant 134A, butane, isobutane and propane in a ratio that delivers a very complete and balanced extraction of cannabinoids and essential oils. E-juice (vape juice) refers to a liquid composition for use in an e-cigarette. A tincture refers to an alcohol-based extract, for example, an extract of Cannabis tissue dissolved in an alcohol. A practitioner can readily determine suitable known methods for producing any products disclosed herein.

[0207] Also disclosed are compositions (e.g., pharmaceutical, nutraceutical, or supplement) including or derived from any of the Cannabis plants disclosed herein. The composition can be formulated for any suitable administration route, for example, pulmonary, oral, or topical administration, or administration by injection. The compositions can be prepared according to conventional methods.

[0208] In some aspects, the composition is for pulmonary administration. The compositions include, but are not limited to, dry powder compositions consisting of the powder of a Cannabis oil described herein, and the powder of a suitable carrier and / or lubricant. The compositions for pulmonary administration can be inhaled from any suitable dry powder inhaler device. In certain instances, the compositions may be conveniently delivered in the form of an aerosol spray from pressurized packs or a nebulizer, with the use of a suitable propellant, for example, dichlorodifluoromethane, trichlorofluoromethane, dichlorotetrafluoroethane, carbon dioxide, or other suitable gas. In the case of a pressurized aerosol, the dosage unit can be determined by providing a valve to deliver a metered amount. Capsules and cartridges of, for example, gelatin for use in an inhaler or insufflator can be formulated containing a powder mix of the compound(s) and a suitable powder base, for example, lactose or starch.

[0209] In some aspects, the composition is formulated for oral administration. For oral administration, a composition (e.g., a pharmaceutical composition) can take the form of, e.g., a tablet or a capsule prepared by conventional means with a pharmaceutically acceptable excipient. Preferred are tablets and gelatin capsules comprising the active ingredient(s), together with (a) diluents or fillers, e.g., lactose, dextrose, sucrose, mannitol, maltodextrin, lecithin, agarose, xanthan gum, guar gum, sorbitol, cellulose (e.g., ethyl cellulose, microcrystalline cellulose), glycine, pectin, polyacrylates and / or calcium hydrogen phosphate, calcium sulfate, (b) lubricants; e.g., silica, anhydrous colloidal silica, talcum, stearic acid, its magnesium or calcium salt (e.g., magnesium stearate or calcium stearate), metallic stearates, colloidal silicon dioxide, hydrogenated vegetable oil, corn starch, sodium benzoate, sodium acetate and / or polyethyleneglycol; for tablets also (c) binders, e.g., magnesium aluminum silicate, starch paste, gelatin, tragacanth, methylcellulose, sodium carboxymethylcellulose, polyvinylpyrrolidone and / or hydroxypropyl methylcellulose; if desired (d) disinte grants, e.g., starches (e.g., potato starch or sodium starch), glycolate, agar, alginic acid or its sodium or potassium salt, or effervescent mixtures; (e) wetting agents, e.g., sodium lauryl sulfate, and / or (f) absorbents, colorants, flavors and sweeteners. Tablets can be either uncoated or coated according to methods known in the art. The excipients described herein can also be used for preparation of buccal dosage forms and sublingual dosage forms (e.g., films and lozenges) as described, for example, in U.S. Pat. Nos. 5,981,552 and 8,475,832. Formulation in chewing gums as described, for example, in U.S. Pat. No. 8,722,022, is also contemplated.

[0210] Further preparations for oral administration can take the form of, for example, solutions, syrups, suspensions, and toothpastes. Liquid preparations for oral administration can be prepared by conventional means with pharmaceutically acceptable additives, for example, suspending agents, for example, sorbitol syrup, cellulose derivatives, or hydrogenated edible fats; emulsifying agents, for example, lecithin, xanthan gum, or acacia; non-aqueous vehicles, for example, almond oil, sesame oil, hemp seed oil, fish oil, oily esters, ethyl alcohol, or fractionated vegetable oils; and preservatives, for example, methyl or propyl-p-hydroxybenzoates or sorbic acid. The preparations can also contain buffer salts, flavoring, coloring, and / or sweetening agents as appropriate.

[0211] The compositions disclosed herein can be formulated for topical administration. Typical formulations for topical administration include creams, ointments, sprays, lotions, hydrocolloid dressings, and patches, as well as eye drops, ear drops, and deodorants. Cannabis oils can be administered via transdermal patches as described, for example, in U.S. Pat. Appl. Pub. No. 2015 / 0126595 and U.S. Pat. No. 8,449,908. Formulation for rectal or vaginal administration is also contemplated. The Cannabis oils can be formulated, for example, as suppositories containing conventional suppository bases such as cocoa butter and other glycerides as described in U.S. Pat. Nos. 5,508,037 and 4,933,363. Suppositories are preferably prepared from fatty emulsions or suspensions. Compositions can contain other solidifying agents such as shea butter, beeswax, kokum butter, mango butter, illipc butter, tamanu butter, carnauba wax, emulsifying wax, soy wax, castor wax, rice bran wax, and candelilla wax. Compositions can further include clays (e.g., Bentonite, French green clays, Fuller's earth, Rhassoul clay, white kaolin clay) and salts (e.g., sea salt, Himalayan pink salt, and magnesium salts such as Epsom salt).

[0212] The compositions disclosed herein can be formulated for administration by injection, for example, by bolus injection or continuous infusion. Formulations for injection can be presented in unit dosage form, for example, in ampoules or in multi-dose containers, optionally with an added preservative. Injectable compositions are preferably aqueous isotonic solutions or suspensions. The compositions may be sterilized and / or contain pharmaceutically acceptable agents, such as preserving, stabilizing, wetting or emulsifying agents, solution promoters, salts for regulating the osmotic pressure, buffers, and / or other ingredients. Alternatively, injectable compositions can be in powder form for reconstitution with a suitable vehicle, for example, a carrier oil, before use. The compositions may further include therapeutic agents or substances.

[0213] In some aspects, the compositions include 1 to about 75%, for example, 1 to about 50%, of a Cannabis product, for example, an oil or extract. In general, subjects receiving a Cannabis oil composition orally are administered doses ranging from about 1 to about 2000 mg of Cannabis oil. A small dose ranging from about 1 to about 20 mg can typically be administered orally when treatment is initiated, and the dose can be increased (e.g., doubled) over a period of days or weeks until a maximum or effective dose is reached.

[0214] VII. Kits for Use in Diagnostic Applications

[0215] Kits for use in diagnostic, research, and prognostic applications are also provided. Such kits may include, for example: primers and / or probes to detect one or more markers disclosed herein, nucleic acid sequences comprising one or more genes associated with plant height as disclosed herein, nucleic acid sequences for targeting / gene editing of WRKY21, LBD16, SAUR51, and / or TCS1 (e.g., siRNA or gRNA targeting WRKY21, LBD16, SAUR51, and / or TCS1), enzymes (e.g., a nuclease (e.g., Cas9), polymerase, restriction enzymes), assay reagents (e.g., qPCR reagents), and buffers. Kit components can be included in one or more containers. The kits may include instructional materials containing directions (e.g., protocols) for the practice of the methods disclosed herein. While the instructional materials typically include written or printed materials, they are not limited to such. Any medium capable of storing such instructions and communicating them to an end user is contemplated. Such media include, but are not limited to electronic storage media (e.g., magnetic discs, tapes, cartridges, chips), optical media (e.g., CD ROM), cloud-based media, and the like. Such media may include addresses to internet sites that provide such instructional materials.

[0216] Clauses I Clause 1. A method for producing one or more Cannabis plants having modified plant height, comprising: (i) analyzing one or more genetic markers in a nucleic acid sample from a Cannabis plant or its germplasm; (ii) detecting one or more genetic markers that indicate modified plant height, (iii) crossing the Cannabis plant comprising the one or more genetic markers indicating modified plant height, and (iv) obtaining one or more progeny plants comprising the one or more genetic markers indicating modified plant height, and wherein the one or more progeny plants have modified plant height relative to a control.

[0217] Clause 2. A method for selecting a Cannabis plant having modified plant height, comprising: (i) analyzing one or more genetic markers in a nucleic acid sample from the Cannabis plant or its germplasm; (ii) detecting one or more genetic markers that indicate modified plant height; and (iii) selecting the Cannabis plant, thereby selecting the Cannabis plant having modified plant height. Clause 3. The method of clause 1 or clause 2, wherein the Cannabis plant having modified plant height is selected for further analysis, propagation, crossing, or to make a product.

[0218] Clause 4. The method of any one of the prior clauses, further comprising crossing the Cannabis plant having modified plant height and producing one or more progeny plants having modified plant height. Clause 5. The method of any one of the prior clauses, wherein the modified plant height is an increase or decrease in vegetative height, harvest height, and / or stretch relative to a control.

[0219] Clause 6. The method of any one of the prior clauses, wherein the modified plant height is an increase in vegetative height, harvest height, and / or stretch relative to a control.

[0220] Clause 7. The method of any one of the prior clauses, wherein: analyzing comprises using PCR, quantitative PCR (qPCR), and / or sequencing; and / or detecting comprises using an oligonucleotide primer set or probe.

[0221] Clause 8. The method of any one of the prior clauses, wherein crossing comprises selfing, sibling crossing, outcrossing, or backcrossing.

[0222] Clause 9. The method of one of the prior clauses, wherein the selfing, sibling crossing, outcrossing, or backcrossing comprises marker-assisted selection for at least two generations.

[0223] Clause 10. The method of one of the prior clauses, wherein detecting one or more genetic markers in the nucleic acid sample comprises detecting one or more genetic markers described in Table 2, 3, 4, 5, 6, or 7. Clause 11. The method of one of the prior clauses, wherein detecting one or more genetic markers comprises detecting one or more of SNP markers: 142603_l 2546539. 142603.12519907, Cannabis.vl_scf542.169826_101 , 142603J 2500071. 142603.12492164, Cannabis.vl_scf542.132451_.99. Cannabis.vl_scf542.122780.J00, 142603.12456339, 142603.12451790, 142603. J 2439239, 142603.12351777, 142603.12346103, 142603.12333141, 142603.12209629, 142603.12122639, 142603.12107817, 142603.12087936, 142603.12081998, 142603_l 2064923. 142603_12038515, 142603_12001291, 142603 J 1997059, 142603_l 1989491, 142603..11951907, 142603 1459458, 142603. J 2546539, 142603..12519907, or 142603.11711278. Clause 12. The method of one of the prior clauses, wherein detecting one or more genetic markers comprises detecting SNP marker 142603_12107817 and / or SNP marker 142603_12081998.

[0224] Clause 13. The method of any one of the prior clauses, wherein the one or more genetic markers that indicate modified plant height comprise a polymorphism at position 51 of one or more of: SEQ ID NOs: 1-200.

[0225] Clause 14. The method of any one of the prior clauses, wherein the one or more genetic markers that indicate modified plant height comprise a polymorphism at position 51 of SEQ ID NO: 44 and / or position 51 of SEQ ID NO: 46.

[0226] Clause 15. The method of any one of the prior clauses, wherein the one or more genetic markers comprise a polymorphism relative to a reference genome in a plant height haplotype, wherein the plant height haplotype comprises the region on chromosome 4 between position 0.02 Mbp and 1.32 Mbp, and wherein the reference genome is Abacus Cannabis genome version Csat_AbacusV2, NCBI assembly accession GCA_025232715.1.

[0227] Clause 16. The method of any one of the prior clauses, wherein analyzing one or more genetic markers comprises analyzing at least 2 genetic markers.

[0228] Clause 17. The method of any one of the prior clauses, wherein detecting one or more genetic markers that indicate modified height comprises detecting at least 2 genetic markers.

[0229] Clause 18. The method of any one of the prior clauses, wherein the one or more genetic markers are genetically linked to a modified plant height trait locus.

[0230] Clause 19. The method of any one of the prior clauses, wherein the control is a Cannabis plant without the one or more markers indicating modified plant height.

[0231] Clause 20. A Cannabis plant produced by the method of any one of the prior clauses.

[0232] Clause 21 . A seed, plant part, tissue culture, or protoplast of the plant of clause 20.

[0233] Clause 22. A method of Cannabis breeding, comprising crossing the Cannabis plant of clause 20.

[0234] Clause 23. The method of clause 22, wherein crossing comprises selfing, sibling crossing, outcrossing, or backcrossing.

[0235] Clause 24. A Cannabis product produced from the plant of clause 20, or the seed, plant part, tissue culture, or protoplast of clause 21.

[0236] Clause 25. The Cannabis product of clause 24, wherein the product is a kief, hashish, bubble hash, an edible product, solvent reduced oil, sludge, e-juice, or tincture.

[0237] Clauses II Clause 1. A method for selecting a Cannabis plant having modified plant height, comprising: (i) analyzing one or more genetic markers in a nucleic acid sample from the Cannabis plant or its germplasm; (ii) detecting one or more genetic markers that indicate modified plant height; and (iii) selecting the Cannabis plant comprising the one or more genetic markers that indicate modified plant height.

[0238] Clause 2. The method of clause 1 , wherein the Cannabis plant having modified plant height is selected for further analysis, propagation, or to make a product.

[0239] Clause 3. The method of clause 1 or clause 2, further comprising: (iv) crossing the selected Cannabis plant and obtaining one or more progeny plants comprising the one or more genetic markers indicating modified plant height.

[0240] Clause 4. The method of any one of the prior clauses, wherein the one or more progeny plants have modified plant height relative to a control.

[0241] Clause 5. The method of any one of the prior clauses, wherein crossing comprises selfing, sibling crossing, outcrossing, or backcrossing.

[0242] Clause 6. The method of any one of the prior clauses, wherein the selfing, sibling crossing, outcrossing, or backcrossing comprises marker-assisted selection for at least two generations.

[0243] Clause 7. The method of any one of the prior clauses, wherein the modified plant height is an increase or decrease in vegetative height, harvest height, and / or stretch relative to a control.

[0244] Clause 8. The method of any one of the prior clauses, wherein the modified plant height is an increase in vegetative height, harvest height, and / or stretch relative to a control.

[0245] Clause 9. The method of any one of the prior clauses, wherein: analyzing comprises using PCR, quantitative PCR (qPCR), and / or sequencing; and / or detecting comprises using an oligonucleotide primer set or probe.

[0246] Clause 10. The method of any one of the prior clauses, wherein detecting one or more genetic markers comprises detecting SNP marker 142603_l 2107817.

[0247] Clause 11. The method of any one of the prior clauses, wherein detecting one or more genetic markers comprises detecting one or more of SNP markers: 142603_12107817, 1426O3_12439239, i 42603 i 2333141, 142708.. 323806, and 142708.. 353959.

[0248] Clause 12. The method of one of the prior clauses, wherein detecting one or more genetic markers in the nucleic acid sample comprises detecting one or more genetic markers described in Table 2, 3, 4, 5, 6, or 7. Clause 13. The method of any one of the prior clauses, wherein the one or more genetic markers that indicate modified plant height comprise a polymorphism at position 51 of one or more of: SEQ ID NOs: 1-205. Clause 14. The method of any one of the prior clauses, wherein the one or more genetic markers that indicate modified plant height comprise a polymorphism at position 51 of SEQ ID NO: 44, SEQ ID NO: 27, SEQ ID NO: 31, SEQ ID NO: 203, or SEQ ID NO: 193.

[0249] Clause 15. The method of any one of the prior clauses, wherein the one or more genetic markers comprise a polymorphism relative to a reference genome in a plant height haplotype, wherein the plant height haplotype comprises the region on chromosome 4 between position 0.02 Mbp and 1.32 Mbp, and wherein the reference genome is Abacus Cannabis genome version Csat_AbacusV2, NCBI assembly accession GCA_025232715.1.

[0250] Clause 16. The method of any one of the prior clauses, wherein analyzing one or more genetic markers comprises analyzing at least 2 genetic markers.

[0251] Clause 17. The method of any one of the prior clauses, wherein detecting one or more genetic markers that indicate modified height comprises detecting at least 2 genetic markers.

[0252] Clause 18. The method of any one of the prior clauses, wherein the one or more genetic markers are genetically linked to a plant height trait locus.

[0253] Clause 19. The method of any one of the prior clauses, wherein the control is Abacus, or a parent plant without the one or more markers indicating modified plant height.

[0254] Clause 20. An engineered Cannabis plant comprising a genetic modification in a WRKY21, LBD16, SAUR51, and / or TCS1 gene that is associated with plant height.

[0255] Clause 21. A method of producing an engineered Cannabis plant having modified plant height, comprising introducing a genetic modification into a WRKY21, LBD16, SAUR51, and / or TCS1 gene that is associated with plant height.

[0256] Clause 22. The engineered Cannabis plant of clause 20 or the method of clause 21, wherein: a. the genetic modification in the WRKY21 gene comprises an amino acid substitution at position 146 of the WRKY21 coding sequence; b. the genetic modification in the LBD16 gene comprises an amino acid substitution at position 9, 130, or 131 , or a two amino acid deletion or insertion between amino acid positions 141 - 144 of the LBD16 coding sequence; c. the genetic modification in the SAUR51 gene comprises a nucleic acid substitution at nucleotide position -4, -37, or -94 bp of the SAUR51 start codon; or d. the genetic modification in the TCS1 gene comprises an amino acid substitution at position 802 in the TCS1 coding sequence; wherein the amino acid or nucleic acid positions are in reference to Abacus Cannabis genome version Csat_AbacusV2, NCBI assembly accession GCA_025232715.1, or a coding sequence obtained therefrom.

[0257] Clause 23. The method of any one of clauses 20-22, wherein: a. the amino acid substitution at position 146 of the WRKY21 coding sequence is a serine or threonine substitution; b. the amino acid substitution at position 9 of the LBD16 coding sequence is a serine or glycine substitution, the amino acid substitution at position 130 of the LBD16 coding sequence is a glutamine or histidine substitution, the substitution at position 131 of the LBD16 coding sequence introduces a preliminary stop codon or removes a preliminary stop codon, or the amino acid deletion or insertion between amino acid positions 141 - 144 of the LBD16 coding sequence is a two amino acid insertion or deletion; c. the nucleic acid substitution at position -4 of the SAUR51 gene comprises a G or A substitution; the nucleic acid substitution at position -37 of the SAUR51 gene is a T or A substitution, or the nucleic acid substitution at position -94 of the SAUR51 gene is a C or G substitution; or d. the amino acid substitution at position 802 of the TCS1 coding sequence is a S or N substitution; wherein the amino acid or nucleic acid positions are in reference to Abacus Cannabis genome version Csat_AbacusV2, NCBI assembly accession GCA_025232715.1 or a coding sequence obtained therefrom.

[0258] Clause 24. A method of selecting a Cannabis plant having modified plant height, comprising: (i) analyzing a WRKY21, LBD16, SAUR51, and / or TCS1 gene in a nucleic acid sample from the Cannabis plant or its germplasm; (ii) detecting one or more polymorphisms that indicate modified plant height; and (iii) selecting a Cannabis plant having the one or more nucleic acid or amino acid substitutions that indicate modified plant height, thereby selecting the Cannabis plant having modified plant height.

[0259] Clause 25. The engineered Cannabis plant of clause 20, 22, or 23, or the method of clause 24, wherein the Cannabis plant having modified plant height is selected for further analysis, propagation, crossing, or to make a product.

[0260] Clause 26. A Cannabis plant produced by the method of any one of the prior clauses.

[0261] Clause 27. A seed, plant part, tissue culture, or protoplast of the plant of clause 26.

[0262] Clause 28. A method of Cannabis breeding, comprising crossing the Cannabis plant of clause 26.

[0263] Clause 29. The method of clause 28, wherein crossing comprises selfing, sibling crossing, outcrossing, or backcrossing.

[0264] Clause 30. A Cannabis product produced from the plant of clause 26, or the seed, plant part, tissue culture, or protoplast of clause 27.

[0265] Clause 31. The Cannabis product of clause 30, wherein the product is a kief, hashish, bubble hash, an edible product, solvent reduced oil, sludge, e-juice, or tincture.

[0266] VIII. EXAMPLES

[0267] The following examples are provided to illustrate particular features of certain aspects of the disclosure, but the scope of the claims should not be limited to those features exemplified.

[0268] Example 1 Identifying and Validating Genetic Markers Plant material and height phenotyping

[0269] In total, four diversity panels were used for height mapping: two consisted of photosensitive (=PS) cannabis seed lots, one consisted of day-neutral seed lots (containing the Autoflowerl, AF1, locus; Dowling, Caroline A., et al. "A FLOWERING LOCUS T ortholog is associated with photoperiodinsensitive flowering in hemp {Cannabis sativa L.)." bioRxiv (2023): 2023-04.) grown under 18 hours light, and one consisted of F2 populations with hemp background that initiate flowers under long day length (>12 hours; containing the Autoflower2, AF2, locus; Dowling et al. "i FLOWERING LOCUS T ortholog is associated with photoperiod-insensitive flowering in hemp (Cannabis sativa L.)." bioRxiv (2023): 2023-04) grown under 18 hours light.

[0270] The first photosensitive diversity panel consisted of 61 seed lots varying in size between 2 - 131 accessions which were evaluated during fall of 2019 and winter of 2020 (n=450; 47 days vegetative time under 18 hours light, followed by 12 hours light until plants reached maturity; Table 1). All plants were clones with the majority of plants grown as a single clonal replicate, one F2 population (n=131; Table 1) was grown as up to three clonal replicates per accession where the average across clonal replicates was used for mapping. The second photosensitive diversity panel consisted of 49 seed lots varying in size between 1 - 16 accessions which were evaluated during spring / summer / fall of 2020 (n=327; Table 1; 35 days vegetative time under 18 hours light, followed by 12 hours light until plants reached maturity). All plants in the second photosensitive diversity panel were grown from seed and were topped. Topping takes place between day 9-10 after transplant, which is 4-5 days before switching to the 12 hour light: 12 hour dark light cycle. The AF1 day-neutral diversity panel consisted of 45 seed lots ranging in size between 3 - 35 plants, which were evaluated from seed during 2021-2023 (n=416). Since day-neutral production plants are grown from sow under 18 hours light only Har vest Height data were collected. The AF2 diversity panel consisted of 3 F2 populations ranging in size between 49 - 50 plants, derived from selfing three sister FIs, which were evaluated from seed during summer / fall of 2022 (n=148; Table 1 ). Height of the AF2 diversity panel plants was measured when they reached 6+ pistil stage under 18 / 6 and was therefore referred to as Vegetative Height.

[0271] Height was measured as Vegetative Height, Harvest Height and Stretch. For plants that are not topped, Vegetative Height is the height of a plant when light is flipped from 18 hours to 12 hours. For topped plants, Vegetative Height is the height of a plant after topping. Harvest Height is the height of a plant at full maturity when it is harvested. Stretch is the difference between Harvest Height and Vegetative Height.

[0272] Table 1. Germplasm used for height mapping. First column: project ID(s); Second column: flowering type. PS=photosensitive. AFl=autoflowering / day neutral, AF2=flowering under long day length; Third column: number of seed lots evaluated; Fourth column: number of accessions evaluated; Fifth column: Vegetative Plant Height (cm), *AF2 plants were grown until the stage where inflorescence clusters containing 6+ pistils are starting to appear; Sixth column: Stretch (cm), not collected for AF1 and AF2; Seventh column: Harvest Height (cm), not collected for AF1 and AF2; Eighth column: day length under which plants were evaluated: 18 / 6=18 hours light, followed by 6 hours dark during a 24 hour cycle, 12 / 12=12 hours light, followed by 12 hours dark during a 24 hour cycle.

[0273] Association mapping of height and stretch

[0274] All diversity panels (Table 1) were genotyped with an Illumina bead array. After initial SNP QC further filtering steps were performed to filter out known low quality SNPs, followed by filtering for missing data (<10%; ) and minor allele frequency (> 1%) using vcftools (Danecek et al. "The variant call format and VCFtools." Bioinformatics 27.15 (201 1): 2156-2158). Missing data were subsequently imputed (R package NAM “snpQC” option; Xavier et al. "NAM: association studies in multiple populations." Bioinformatics 31.23 (2015): 3862-3864).

[0275] Subsequently, nested association mapping (NAM) was performed for Vegetative Height, Harvest Height, and Stretch for the two photosensitive diversity panels (35,576 and 34,807 SNPs, respectively), Harvest Height for the AF1 diversity panel, and Vegetative Height for the AF2 diversity panel (11,790 and 28,492 SNPs, respectively). In addition, NAM was performed for Vegetative Height, Harvest Height, and Stretch for the 19GAR2 data (an F2 mapping population consisting of 142 accessions; 11,183 SNPs). NAM was performed using the R package NAM after missing data imputation as described above using seed lots as family structure and a kinship matrix to control for relatedness (GWAS2 function).

[0276] In total, 39 SNP markers between 21,770 and 1,320,570 bp on chromosome 4 of the Abacus reference genome (version Csat_AbacusV2; NCBI assembly accession GCA_025232715.1) were significantly associated with Harvest Height in 19GAR1 / 2 / 3 (Bonferroni multi-test threshold p=1.41E-06; Table 2). Within this set of 39 SNPs, ten were significantly associated with Vegetative Height and 23 SNPs were significantly associated with Stretch in 19GAR1 / 2 / 3 (Table 2).

[0277] In total, 88 SNP markers located on chromosomes 1, 4, 7, and 8 were significantly associated with Harvest Height in 20TP1B / C / D (Bonferroni multi-test threshold p=1.44E-06; Table 3). The majority of 82 markers were located on chromosome 4, 25 of those markers overlapped with the Harvest Height markers discovered in 19GAR1 / 2 / 3 (142603.12546539, 142603_12519907, Cannabis.vl_scf542.169826_101 , 142603_l 2500071 , 142603_12492164, Cannabis.vl_scf542.132451_99, Cannabis. vl_scf542.122780_100, 142603_12456339, 142603.12451790, 142603_12439239, 142603_12351777, 142603_12346103, 142603.12333141, 142603.12209629, 142603_12122639, 142603_12107817, 142603.12087936, 142603.12081998, 142603.12064923, 142603.12038515, 142603.12001291, 142603.11997059, 142603.11989491, 142603.11951907, 142603.11459458). There was no significant association with Vegetative Height, whereas 170 SNP markers on chromosomes 1, 2, 3, 4, 6, 7, 8, 9, and X were significantly associated with Stretch (Table 4), 75 of those were also significantly associated with Harvest Height in 20TP1B / C / D, 17 of those were also significantly associated with Stretch in 19GAR1 / 2 / 3 (142603.12546539, 142603.12519907, Cannabis.vl_scf542.132451.99, 142603.12456339, 142603.12451790, 142603.12439239, 142603.12346103, 142603.12333141, 142603.12209629, 142603.12122639, 142603.12107817, 142603.12087936, 142603.12081998, 142603.12064923, 142603.11989491, 142603.11711278, 142603.11459458). Since Harvest Height and Stretch were measured on plants that were topped, these SNP markers are associated with stretch after topping.

[0278] In total, 7 SNP markers located on chr omosome X were significantly associated with Harvest Height in the AF1 diversity panel (Bonferroni multi-test threshold p=2.14E-06; Table 5).

[0279] In total, three SNP markers located on chromosomes 4 and 6 were significantly associated with Vegetative Height in the AF2 diversity panel (Bonferroni multi-test threshold p=l .75E-06; Table 6). The marker on chromosome 4, 142603.12333141, was also associated with Harvest Height and Stretch in the two PS diversity panels.

[0280] In total, 13 SNP markers located on chromosome 4 between positions 199,720 - 561,849 bp were significantly associated with Harvest Height and Stretch in the 19GAR2 F2 mapping population (Bonferroni multi-test threshold p=4.47E-06). Five of those 13 SNP markers were also significantly associated with Vegetative Height in 19GAR2 (Table 7).

[0281] Height and stretch SNP marker validation

[0282] SNP markers were validated by comparing beneficial genotypes for the mapped markers. Beneficial genotypes were determined as the genotype with the highest average value for Harvest Height, Vegetative Height, and Stretch per genotype (homozygous reference allele, heterozygous, or homozygous alternate allele). The heterozygous genotype was added to the homozygous beneficial genotype if the heterozygous genotype had an intermediate or similar' average value of the phenotypic score between homozygous reference allele and homozygous alternate allele, which indicates an additive or dominant effect, respectively. Validation was performed when missing data did not exceed 10%. Validation was performed using diversity panels used for mapping (Table 1) as well as an AF1 validation set from 14 seed lots containing 3 - 4 accessions per seed lot (n=54) which were grown in a greenhouse under 18 hours light during spring of 2024 (24ATX1). This additional set of AF1 accessions was added for the validation of height and stretch markers mapped in the two PS and the AF2 diversity panels since the initial AF1 diversity panel displayed in general reduced levels of variation for some of these markers.

[0283] Validation of the 39 SNP markers associated with Harvest Height discovered in the first PS diversity panel resulted in 37 of those SNP markers validated in the second PS diversity panel (SNP markers 142603_12209629 and 142603_12122639 were not validated); all of the 39 SNP markers were validated in the AF2 diversity panel and 18 SNP markers (Cannabis.vl_scf542-132451_99, 142603.12456339, 142603.12439239, 142603.12333141, 142603.12315001, 128506.1124, 142603.12258791, 142603.12224772, 142603.12116693, 142603.12107817, 142603.12087936, 142603.12081998, 142603.12064923, 142603.12034397, 142603.11745385, 142603.11693304, 142603.11459458, Cannabis. vl_scfl557-58243_100) were validated in the 24ATX1 AF1 validation set.

[0284] Validation of the 88 SNP markers associated with Harvest Height discovered in the second PS diversity panel resulted in 81 of those markers validated in the first PS diversity panel (SNP markers 142603.12209629, 142603.12122639, 142603.11444265, 142603.11267672, 142603.8477734, 142704.5112285, and 142193.3819651 were not validated); all 88 SNP markers were validated in the AF2 diversity panel and 23 SNP markers (Cannabis. vl_scf542-132451_99, 142603.12456339, 142603.12439239, 142603.12333141 , 128506.14057, 142603.12209629, 142603.12173963, 142603.12122639, 142603.12107817, 142603.12087936, 142603.12081998, 142603.12067549, 142603.12064923, 142603.12052981, 142603.12047821, 142603.12024545, 142603.12018251, 142603.11994824, 142603.11818595, 142603.11735270, 142603.11459458, 142603.11444265, 157.2560640) were validated in the 24ATX1 AF1 validation set.

[0285] Validation of the seven SNP markers associated with Harvest Height discovered in the AF1 diversity panel was performed with the AF1 accessions from 14 seed lots containing 3 - 4 accessions per seed lot (n=54) which were grown in a greenhouse under 18 hours light:6 hours dark between April - July 2024 (24ATX1). As a result, of the seven mapped markers three SNP markers were validated: 142708.323806, 142708.353959, and 142708.453318. Four of the seven SNP markers were validated in the first PS diversity panel (142708.186978, 142708.209874, 142708.349228, 142708.353959 and 142708.453318) and five SNP markers were validated in the second PS diversity panel (142708.186978, 142708.209874, 142708.323806, 142708.349228, 142708.399189).

[0286] Validation of the three SNP markers associated with Harvest Height discovered in the AF2 diversity panel resulted in two of those markers validated in the fust PS diversity panel (SNP marker 139181.38701 was not validated); all three SNP markers were validated in the second PS diversity panel as well as the 24ATX1 AF1 validation set.

[0287] Validation of the 13 SNP markers associated with Har vest height discovered in the 19GAR2 F2 mapping population resulted in eight of these markers validated in the second PS diversity panel (142603.12333141, 134305.1882, 142603.12258791, 142603.12209629, 142603.12122639, 142603.12087936, 142603.12081998, 142603.12034397); a different set of eight SNP markers was validated in the AF2 diversity panel (142603.12369183, 142603.12333141, 134305.1882, 142603.12209629, 142603.12087936, 142603.12081998, 142603.12034397, 142603.12030168); a set of 10 SNP markers was validated in the AF1 diversity panel (142603.12369183, 142603.12333141, 134305.1882, 142603.12258791, 142603.12224772, 142603.12122639, 142603.12087936, 142603.12081998, 142603.12034397, 142603.12030168).

[0288] Validation of the 23 SNP markers associated with Stretch discovered in the first PS diversity panel resulted in 21 of those SNPs validated in the second PS diversity panel (SNP markers 142603.12209629 and 142603.12122639 were not validated).

[0289] Validation of the 170 SNP markers associated with Stretch discovered in the second PS diversity panel resulted in 156 of those markers validated in the first PS diversity panel (SNP markers 142603.11267672, 142603.8477734, 142704.4441230, 142250.1252684, 141136.547048, 141588.963838, 142582.866445, 142078.2688971, 142078.3416091, 142193.2428295, 142193.2666513, 142193.2920302, 142193.3819651, 157345.7773, and 90.2781977 were not validated). The 14 not validated SNP markers had lower significance (p-values running between 9.55E-08 and 1.0E-05).

[0290] Validation of the five SNP markers associated with Stretch in the 19GAR2 F2 mapping population resulted in two of those markers validated in the second PS diversity panel (142603.12333141, 134305.1882).

[0291] The main Harvest Height, Vegetative Height, and Stretch trait locus, for both topped and untopped plants, mapped in the two PS populations and the AF2 population spans 1.3 Mbp between 0.02 - 1.32 Mbp on chromosome 4 of the Abacus reference genome (version Csat_AbacusV2; NCBI assembly accession GCA.025232715.1). The main Harvest Height trait locus mapped in the AF1 population spans 0.28 Mbp between 1.13 - 1.41 Mbp on chromosome X of the Abacus reference genome (version Csat_AbacusV2; NCBI assembly accession GCA_025232715.1).

[0292] A highly associated SNP marker (142603_12107817) mapped based on NAM of Harvest Height and Stretch in the two PS diversity panels, is located at position 484,010 bp on chromosome 4 of the Abacus reference genome (version Csat_AbacusV2; NCBI assembly accession GCA_025232715.1); it is located 10.3 Kbp upstream of candidate gene LBD16 (Lateral Organ Boundaries 16; AT2G42420; based on homology with Arabidopsis thaliana) and 30.9 Kbp downstream of candidate gene LBD29 (Lateral Organ Boundaries 29; AT3G58190; based on homology with Arabidopsis thaliana). The CBDRx (version cslO) homolog of LBD16 is LOC115714754 (annotated as LOB domain-containing protein 16- like; over-expression of LBD16 in wheat resulted in shorter plants; Wang, Huifang, et al. "Overexpression of TaLBD16-4D alters plant architecture and heading date in transgenic wheat." Frontiers in Plant Science 13 (2022): 911993.) and is located between 91,377,423-91,379,032 bp on chromosome 4 of the CBDRx reference genome. The CBDRx (version cslO) homolog of LBD29 is LOCI 15714879 (annotated as LOB domain-containing protein 29; over-expression of LBD29 in Eucalyptus grandis resulted in shorter plants; Lu, Qiang, et al. "Genomewide analysis of the lateral organ boundaries domain gene family in Eucalyptus grandis reveals members that differentially impact secondary growth." Plant biotechnology journal 16.1 (2018): 124-136.) and is located between 91,393,087 - 91,391,755 bp on chromosome 4 of the CBDRx reference genome.

[0293] Marker combinations

[0294] Combinations of the most significant mapped markers in various mapping panels containing both topped and untopped plants (142603_12439239, 142603_12107817, 142708_323806, 142708_353959, and 142708_453318) were validated in the mapping panels, the 24ATX1 AF1 panel (14 seed lots, n=54, not topped), and a panel of 77 diverse photosensitive seed lots (n=366; topped). In general, a B (=homozygous alternative allele) or X (=heterozygous) genotype for 142603_I21078I7 in combination with a B or X genotype for 142603_l 2439239 was best at predicting highest values for Harvest Height, Vegetative Height, and Stretch, for both topped and untopped plants. However, when 142603_12107817 has genotype A (=homozygous reference genome allele), then highest values for Harvest Height, Vegetative Height, and Stretch are obtained by selecting for 142603_12439239 genotype B or X. When 142603_12439239 has genotype A, then highest values for Harvest Height, Vegetative Height, and Stretch are obtained by selecting for 142603 12107817 genotype B or X. In some genetic backgrounds (mainly AF1) where 142603_12107817 has genotype A and 142603_l 2439239 has genotype B highest values for Harvest Height are obtained by selecting for an A genotype for 142708_323806, 142708_353959, or 142708_453318. For successful prediction of shortest plants an A genotype for 142603_12107817 in combination with an A genotype for 142603_12439239, and a B or X genotype for 142708.323806, 142708.353959, or 142708.453318.

[0295] In the 19GAR2 F2 mapping population 142603.12439239, 142603.12107817, 142708.323806, 142708.353959, and 142708.453318 are all fixed and 142603.12333141 is the main SNP marker associated with Harvest Height, Vegetative Height and Stretch. This marker also had the largest effect on Harvest Height in 22ALV1. Therefore, genetic backgrounds where 142603.12439239, 142603.12107817, 142708.323806, 142708.353959, and 142708.453318 are all fixed or have a minor effect on height and stretch SNP marker 142603.12333141 genotype of B should be used to predict the tallest plants and genotype A should be used to predict the shortest plants.

[0296] Table 8 shows the mapped SNP markers and their 50 bp flanking sequences in the Abacus reference genome (version Csat_AbacusV2; NCBI assembly accession GCA.025232715.1).

[0297] Example 2 Discovery of Gene Sequences Associated with Plant Height

[0298] Four genes, WRKY21, LBD16, SAUR51, and TCS1, located within the haplotypes of SNP markers 142603.12439239, 142603.12107817, 142708.323806, and 142708.353959, respectively, were explored for putative causative SNPs for the mapped plant height traits.

[0299] SNP marker 142603.12439239, which was mapped based on NAM of Harvest Height and Stretch in the two PS diversity panels (validated in both the PS diversity panels, the AF2 and 24ATX1 AF1 diversity panels), is located at position 128,827 bp on chromosome 4 of the Abacus reference genome (version Csat_AbacusV2; NCBI assembly accession GCA.025232715.1); it is located 5.9 Kbp upstream of candidate gene WRKY21 (WRKY DNA-binding protein 21; AT2G30590; based on homology with Arabidopsis thaliana). The CBDRx (version cslO) homolog of LBD16 is LOC115713071, which is located between positions 91,717,509 - 91,720,124 bp on chromosome 4 of the CBDRx (version cslO) reference genome and annotated as probable WRKY transcription factor 21 . Over-expression of WRKY21 in rice results in a semi-dwarf phenotype with short internodes and reduced levels of indole-3-acetic acid (IAA) and gibberellin 3 (GA3; Wei, Xiaoshuang, et al. "Genome-wide association study in rice revealed a novel gene in determining plant height and stem development, by encoding a WRKY transcription factor." International Journal of Molecular Sciences 22.15 (2021): 8192.).

[0300] SNP marker 142603.12107817, which was mapped based on NAM of Harvest Height and Stretch in the two PS diversity panels (validated in both the PS diversity panels, the AF2 and 24ATX1 AF1 diversity panels), is located at position 484,010 bp on chromosome 4 of the Abacus reference genome (version Csat_AbacusV2; NCBI assembly accession GCA.025232715.1); it is located 10.3 Kbp upstream of candidate gene LBD16 (Lateral Organ Boundaries 16; AT2G42420; based on homology with Arabidopsis thaliana). The CBDRx (version cslO) homolog of LBD16 is LOC115714754 (annotated as LOB domain-containing protein 16-like), located between 91,393,087 - 91,391,755 bp on chromosome 4 of the CBDRx reference genome. Over-expression of LBD16 in wheat resulted in shorter plants; Wang, Huifang, et al. "Overexpression of TaLBD16-4D alters plant architecture and heading date in transgenic wheat." Frontiers in Plant Science 13 (2022): 911993.). LBD16 is activated by auxin via AUXIN RESPONSE FACTOR 7 (ARF7) and ARF19 (ARF7 / 19; Okushima. Yoko, et al. "ARF7 and ARF19 regulate lateral root formation via direct activation of LBD / ASL genes in Arabidopsis." The Plant Cell 19.1 (2007): 118-130.).

[0301] SNP marker 142708_353959 was mapped based on NAM of Harvest Height in the AF1 diversity panel (validated in the 24ATX1 AF1 set and the first PS diversity panel), is located at position 1,300,352 bp on chromosome X of the Abacus reference genome (version Csat_AbacusV2; NCB1 assembly accession GCA_025232715.1); it is located 5.3 Kbp upstream of candidate gene TCS1 (Trichome Cell Shape 1; AT1G75590; based on homology with Arabidopsis thaliana). The CBDRx (version cslO) homolog of TCS1 is LOCI 15701735 (annotated as filament-like plant protein 4 isoform XI), located between 103,680,341 - 103,686,491 bp on chromosome X of the CBDRx reference genome. TCS1 promotes microtubule assembly and indirectly plant height as shown by shorter hypocotyl length of oryzalin-treated Arabidopsis T-DNA mutants; Chen, Liangliang, et al. "TCS1, a microtubule-binding protein, interacts with KCBP / ZWICHEL to regulate trichome cell shape in Arabidopsis thaliana." PLoS Genetics 12.10 (2016): el006266.).

[0302] SNP marker 142708_323806 was mapped based on NAM of Harvest Height in the AF1 diversity panel (validated in the 24 ATX 1 AF1 set and the second and first PS diversity panel, respectively) and is located at position 1,270,322 bp on chromosome X of the Abacus reference genome (version Csat_AbacusV2; NCBI assembly accession GCA_025232715.1); it is located 1.6 Kbp upstream of candidate gene SAUR51 (SMALL AUXIN UPREGULATED RNA 51 ; AT1G75580; based on homology with Arabidopsis thaliana). The CBDRx (version cslO) homolog of SAUR51 is LOCI 15701761 (annotated as auxin-responsive protein SAUR50), located between 103,704,381 - 103,705,454 bp on chromosome X of the CBDRx reference genome. Over-expression of SAUR58 in tomato results in reduced plant height (Liu, Yue, et al. "Cytokinin-inducible response regulator S1RR6 controls plant height through gibberellin and auxin pathways in tomato." Journal of Experimental Botany 74.15 (2023): 4471- 4488.)).

[0303] WRKY21, LBD16, SAUR51, and TCS1 were evaluated for coding sequence variation in Cannabis accessions varying for Harvest Height and the mapped markers as described above (Tables 2, 3, 4, and 5). Sequences were compared with CBDRx (NCBI assembly accession GCA_900626175.1) reference genome sequence. CBDRx reference genome marker genotypes were determined after performing a BLASTN search based on 50-100 bp flanking sequences surrounding each SNP marker in the Abacus reference genome (Table 9).

[0304] RNA was extracted from leaf tissue from 12 accessions (Table 9; Nucleospin RNA Plant and Fungi kit, Macherey-Nagel). After concentration adjustment and treatment with DNAse, the RNA was used directly for RT-PCR (OneTaq® One-Step RT-PCR Kit, New England Biolabs). Sanger sequencing of coding sequence (CDS) was performed based on RT-PCR product. Primers for amplification and sequencing of WRKY21, LBD16, SAUR51, and TCS1 can be found in Table 10 (SEQ ID NOs: 206 - 227). For both Abacus and Finola reference genome sequences (Abacus: NCBI assembly accession GCA_025232715.1; Finola: NCBI assembly accession GCA_003417725.2) were used in combination with RT-PCR sequencing data to identify correct splice sites. For CBDRx reference genome sequence (NCBI assembly accession GCA_900626175.1) was used and splice sites were inferred based on alignment with RT-PCR sequences from other accessions.

[0305] Alignment of Sanger sequenced fragments was performed per accession for all four genes. The resulting consensus sequences were subsequently aligned per gene. Functional CDS were translated to protein sequences, which were compared among tall and short plants differing for SNP marker genotypes. WRKY21

[0306] Alignment of WRKY21 CDS and protein sequences of Abacus / 20LCMP-l:l, 23TRC1-15-1, 23PT2-101-1, Finola / 23TRC1-16-1, 23PT2-86-30, 23PT2-106-22, 23PT2-108-1, and 21TP1B-1-1 revealed one amino acid substitution distinguishing between short and tall accessions differing for SNP marker 142603_12439239 (Table 9). This amino acid substitution is a S (Serine, observed in Abacus / 20LCMP-l:l, 23TRC1-15-1, and 23PT2-101-1) to T (Threonine, observed in Finola / 23TRC16-1, 23PT2-86-30, 23PT2-106-22, 23PT2-108-1, and 21TP1B-1-1) at amino acid position 146 in Abacus (S146T; SEQ ID NO: 230) and Finola / 23TRC1-16-1 (SEQ ID NO: 231), caused by a T to A nucleotide substitution (Abacus / 20LCMP-1 :1 , 23TRC1 -15-1 , and 23PT2-101-1 : T; Finola / 23TRC16-1 , 23PT2-86- 30, 23PT2-106-22, 23PT2-108-1, and 21TP1B-1-1: A) at CDS position 436 bp of Finola / 23TRC1-16-1 (SEQ ID NO: 229), and Abacus (SEQ ID NO: 228; located at position 51 bp for SEQ ID NO: 232; Table 10).

[0307] NCBI conserved domain search (Marchler-Bauer, Aron, et al. "CDD / SPARCLE: functional classification of proteins via subfamily domain architectures." Nucleic acids research 45.D1 (2017): D200-D203.) identified the location of the zinc binding domain between amino acid positions 267 - 302 and the WRKY domain between amino acid positions 305 - 362. Alignment with Arabidopsis thaliana WRKY21 protein sequence revealed that the S146T amino acid substitution observed in Cannabis sativa plants which have the beneficial genotype for SNP marker 142603_12439239 is located in a region with polar residues and therefore an amino acid substitution could lead to changes in protein structure and conformation (Uniprot 004336). Without being bound by any particular theory, it is possible that the S146T amino acid substitution causes a change in protein structure of WRKY21 resulting in increased levels of indole-3-acetic acid (IAA) and gibberellin 3 (GA3) causing increased plant height (Wei, Xiaoshuang, et al. "Genome-wide association study in rice revealed a novel gene in determining plant height and stem development, by encoding a WRKY transcription factor." International Journal of Molecular Sciences 22.15 (2021): 8192.).

[0308] LBD16

[0309] Alignment of LBD16 CDS and protein sequences of Abacus / 20LCMP-l:l, 23TRC1-15-1, 23PT2-101-1, Finola / 23TRC 1-16-1, 23PT2-86-30, 23PT2-106-22, 23PT2-108-1, and 21TP1B-1-1 revealed five amino acid substitutions / indels in total, where a two amino acid indel distinguishes between short and tall accessions differing for SNP marker 142603_12107817 (Table 9), an amino acid substitution early in the sequence distinguishes a tall accession from a different genetic background (21TP1B-1-1) and an amino acid substitution followed by a nucleotide substitution causing a preliminary stop codon distinguishes a tall accession from another genetic background (Finola / 23TRC1-16-1).

[0310] The first amino acid substitution is an S (Serine, observed in Abacus / 20LCMP-l:l, 23TRC1-15- 1, 23PT2-101-1, 23TRC1-16-1, 23PT2-86-30, 23PT2-106-22, and 23PT2-108-1) to G (Glycine, observed in 21TP1B-1-1) at amino acid position 9 in Abacus (S9G; SEQ ID NO: 237) and 21TP1B-1-1 (SED ID NO: 238) caused by an A to G nucleotide substitution (Abacus / 20LCMP-l:l, 23TRC1-15-1, 23PT2-101- 1, Finola / 23TRC1-16-1, 23PT2-86-30, 23PT2-106-22, and 23PT2-108-1: A; 21TP1B-1-1: G) at CDS position 25 bp of 21TP1B-1-1 (SEQ ID NO: 234) and Abacus (SEQ ID NO: 233; located at position 51 bp for SEQ ID NO: 241; Table 10).

[0311] The second amino acid substitution is a Q (Glutamine, observed in Abacus / 20LCMP-l:l, 23TRC1-15-1, 23PT2-101-1, 23PT2-86-30, 23PT2-106-22, 23PT2-108-1, and 21TP1B-1-1) to an H (Histidine, observed in Finola / 23TRC1 -16-1 ) at amino acid position 130 in Abacus (Q130H; SEQ ID NO: 237) and 23TRC1-16-1 (SEQ ID NO: 239) caused by a G to C nucleotide substitution (Abacus / 20LCMP- 1:1, 23TRC1-15-1, 23PT2-101-1, 23PT2-86-30, 23PT2-106-22, 23PT2-108-1, and 21TP1B-1-1: G; 23TRC1-16-1: C) at CDS position 390 bp of 23TRC1-16-1 (SEQ ID NO: 235) and Abacus (SEQ ID NO: 233; located at position 51 bp for SEQ ID NO: 242; Table 10).

[0312] The preliminary stop codon observed at amino acid position 131 for 23TRC1-16-1 (Y132stop; SEQ ID NO: 239) is a Y (Tyrosine, observed in Abacus / 20LCMP-l:l, 23TRC1-15-1, 23PT2-101-1, 23PT2-86-30, 23PT2-106-22, 23PT2-108-1, and 21TP1B-1-1). This is caused by a T to G nucleotide substitution (Abacus / 20LCMP-l:l, 23TRC1-15-1, 23PT2-101-1, 23PT2-86-30, 23PT2-106-22, 23PT2- 108-1, and 21TP1B-1-1: T; 23TRC1-16-1: G) at CDS position 396 bp of 23TRC1-16-1 (SED ID NO: 235) and Abacus (SEQ ID NO: 233; located at position 51 bp for SEQ ID NO: 243; Table 10).

[0313] The indel is a deletion of 6 nucleotides AATCAT (insertion present in Abacus) or ACTCAT (insertion present in 23TRC1-15-1 and 23PT2-101-1) between CDS position 423 - 424 bp of 23PT2-86- 30 (SEQ ID NO: 236; as well as 23PT2-106-22, 23PT2-108-1) causing a deletion of NH (Asparagine and Histidine insertion present in Abacus) or TH (Threonine and Histidine insertion present in 23TRC1-15-1 and 23PT2-10I-I) in 23PT2-86-30 (SEQ ID NO: 240; as well as 23PT2-106-22, 23PT2-108-1). In the Abacus protein sequence the NH insertion is present between amino acid positions 141 - 144 (N142del, H143del ; SEQ ID NO: 237), corresponding with nucleotide positions 424 - 429 bp in Abacus CDS (SEQ ID NO: 233; located at position 51 bp for SEQ ID NO: 244, Table 10).

[0314] NCBI conserved domain search (Marchler-Bauer, Aron, et al. "CDD / SPARCLE: functional classification of proteins via subfamily domain architectures." Nucleic acids research 45.D1 (2017): D200-D203.) identified the location of the LOB domain between amino acid positions 118 - 117. The S9G amino acid substitution observed in 21TP1B-1-1 as well as the N142del, H143del deletion observed in 23PT2-86-30, 23PT2-106-22, and 23PT2-108-1 are located outside the LOB domain, however, it is possible that this amino acid substitution as well as the deletion still affect the function of LBD16. The truncated LBD16 protein observed in 23TRC1 / Finola may result in the opposite phenotype (taller plants) of over-expressed LBD16, which resulted in shorter plants (Wang, Huifang, et al. "Overexpression of TaLBD16-4D alters plant architecture and heading date in transgenic wheat." Frontiers in Plant Science 13 (2022): 911993) possibly as a result of reduced number of lateral roots as observed in Arabidopsis thaliana T-DNA mutants (Okushima, Yoko, et al. "ARF7 and ARF19 regulate lateral root formation via direct activation of LBD / ASL genes in Arabidopsis." The Plant Cell 19.1 (2007): 118-130.).

[0315] SAVR51

[0316] Alignment of SAUR51 CDS and protein sequences of Abacus / 20LCMP-l:l, 23TRC1-15-1, 23PT2-108-1, 24ALV1-1-4, 24ATX1-70-3, 24ATXI-71-3, and 24ATX1-73-2 did not reveal any amino acid substitutions, however within 94 bp upstream of the start codon (Abacus: SEQ ID NO: 245; 24ALV1-4: SEQ ID NO: 246) there were three SNPs distinguishing between short and tall accessions differing for SNP marker 142708_323806 (Table 9).

[0317] The first SNP is a G (observed in Abacus / 20LCMP-l:l, 24ATX1-70-3, and 24ATX1-71-3) to A (24ALV1-L4, 24ATXL73-2) substitution at nucleotide position -4 bp upstream from the start codon (located at position 51 bp of SEQ ID NO: 247). Accessions 23TRC1-15-1, 23PT2-108-1, which were heterozygous for SNP marker 142708_323806, were also heterozygous for this SNP.

[0318] The second SNP is a T (observed in Abacus / 20LCMP-l:l, 24ATX1-70-3, and 24ATX1-71-3) to A (24ALV1-1-4, 24ATX1-73-2) substitution at nucleotide position -37 bp upstream from the start codon (located at position 51 bp of SEQ ID NO: 248). Accessions 23TRC1-15-1, 23PT2-108-1, which were heterozygous for SNP marker 142708_323806, were also heterozygous for this SNP.

[0319] The third SNP is a C (observed in Abacus / 20LCMP-l:l, 24ATX1-70-3, and 24ATX1-71-3) to G (24ALV1-1-4, 24ATX1-73-2) substitution at nucleotide position -94 bp upstream from the start codon (located at position 51 bp of SEQ ID NO: 249). Accessions 23TRC1-15-1, 23PT2-108-1, which were heterozygous for SNP marker 142708_323806, were also heterozygous for this SNP.

[0320] Since over-expression of SAUR58 in tomato results in reduced plant height (Liu, Yue, et al. "Cytokinin-inducible response regulator S1RR6 controls plant height through gibberellin and auxin pathways in tomato." Journal of Experimental Botany 74.15 (2023): 4471-4488) an A at position -4 bp, an A at position -37 bp and a G at position -94 bp upstream of the start codon of S AUR51 may positively affect its transcriptional activation resulting in increased expression of SAUR51 and reduced plant height. Alternatively, G at position -4 bp, an T at position -37 bp and a C at position -94 bp upstream of the start codon of SAUR51 negatively affect its transcription activation resulting in decreased expression of S AUR51 and increased plant height.

[0321] TCS1

[0322] Alignment of TCS1 CDS and protein sequences of Abacus / 20LCMP-l:l, 23TRC1-15-1, 23PT2- 101-1, 23PT2-86-30, 23PT2-108-1, and 24ALV1-1-4 revealed one amino acid substitution distinguishing between short and tall accessions differing for SNP marker 142708_353959 (Table 9). This amino acid substitution is a S (Serine, observed in Abacus / 20LCMP-l:l, 23PT2-86-30, and 23PT2-101-1) to N (Asparagine, observed in, 23PT2-108-1, 23TRC1-15-1, and 24ALV1-1-4) at amino acid position 802 in Abacus (S802N; SEQ ID NO: 252) and 24ALV1-1-4 (SEQ ID NO: 253), caused by a G to A nucleotide substitution (Abacus / 20LCMP-l:l, 23PT2-86-30, and 23PT2-101-1: G: 23PT2-108-1, 23TRC1-15-1, and 24ALV1-1-4: A) at CDS position 2405 bp of 24ALV1-1-4 (SEQ ID NO: 251), and Abacus (SEQ ID NO: 250; located at position 51 bp for SEQ ID NO: 254; Table 10).

[0323] NCBI conserved domain search (Marchler-Bauer, Aron, et al. "CDD / SPARCLE: functional classification of proteins via subfamily domain architectures." Nucleic acids research 45. DI (2017): D200-D203) identified the location of the FPP domain between amino acid positions 81 - 955. Based on an alignment with Arabidopsis thaliana, the S802N substitution appears to be located in the coiled coil region inside this domain. Without being bound by any particular theory, it is possible that this amino acid substitution affects TCS 1 gene function such that it improves microtubule assembly and therefore plant height (Chen, Liangliang, et al. "TCS1, a microtubule-binding protein, interacts with KCBP / ZWICHEL to regulate trichome cell shape in Arabidopsis thaliana." PLoS Genetics 12.10 (2016): el006266.).

[0324]

[0325]

[0326] Table 2. Significant SNP markers based on NAM of Harvest Height, Vegetative Height, and Stretch of GAR1 / 2 / 3 (n=450). First column, SNP marker number; Second column, SNP marker name,A=SNP marker is also significantly associated with Vegetative Height, *=SNP marker is also significantly associated with Stretch; Third column, NAM p-value for Harvest Height; Fourth column, beneficial genotype for increased

[0327] 5 height / stretch (A=homozygous for reference allele, B=homozygous for alternative allele, X=heterozygous); Fifth column, reference allele call;

[0328] Sixth column, alternative allele call; Seventh column, chromosome; Eighth column, Abacus reference genome position in bp; Ninth column, left flanking SNP of haplotype surrounding SNP marker; Tenth column, right flanking SNP of haplotype surrounding SNP marker; Eleventh column, Abacus reference genome position in bp for left flanking SNP of haplotype surrounding SNP marker; Twelfth column, Abacus reference genome position in bp for right flanking SNP of haplotype surrounding SNP marker.

[0329]

[0330]

[0331]

[0332]

[0333]

[0334]

[0335] Table 3. Significant SNP markers based on NAM of Harvest Height in 20TP1B / C / D (n=327). First column, SNP marker number; Second column, SNP marker name; Third column, NAM p-value for Harvest Height; Fourth column, beneficial genotype for increased Harvest Height (A=homozygous for reference allele, B=homozygous for alternative allele, X=heterozygous), *=B inferred based on segregation patterns, **=A

[0336] 5 inferred based on segregation patterns; Fifth column, reference allele call; Sixth column, alternative allele call; Seventh column, chromosome; Eighth column, Abacus reference genome position in bp; Ninth column, left flanking SNP of haplotype surrounding SNP marker; Tenth column, right flanking SNP of haplotype surrounding SNP marker; Eleventh column, Abacus reference genome position in bp for left flanking SNP of haplotype surrounding SNP marker; Twelfth column, Abacus reference genome position in bp for right flanking SNP of haplotype surrounding SNP marker.

[0337]

[0338]

[0339]

[0340]

[0341]

[0342]

[0343]

[0344]

[0345]

[0346]

[0347]

[0348]

[0349] Table 4. Significant SNP markers based on NAM of Stretch in 20TP1B / C / D (n=327). First column, SNP marker number; Second column, SNP

[0350] marker name; Third column, NAM p-value for Stretch; Fourth column, beneficial genotype for increased Stretch (A=homozygous for reference allele, B=homozygous for alternative allele, X=heterozygous), *=B inferred based on segregation patterns, **=A inferred based on segregation patterns; Fifth column, reference allele call; Sixth column, alternative allele call; Seventh column, chromosome; Eighth column, Abacus reference genome position in bp; Ninth column, left flanking SNP of haplotype surrounding SNP marker; Tenth column, right flanking SNP of haplotype

[0351] 5 surrounding SNP marker; Eleventh column, Abacus reference genome position in bp for left flanking SNP of haplotype surrounding SNP marker; Twelfth column, Abacus reference genome position in bp for right flanking SNP of haplotype surrounding SNP marker.

[0352] Table 5. Significant SNP markers based on NAM of Harvest Height in the AF1 diversity panel (n=416). First column, SNP marker number;

[0353] 10 Second column, SNP marker name; Third column, NAM p-value for Harvest Height; Fourth column, beneficial genotype for increased Harvest

[0354] Height (A=homozygous for reference allele, B=homozygous for alternative allele, X=heterozygous); Fifth column, reference allele call; Sixth column, alternative allele call; Seventh column, chromosome; Eighth column, Abacus reference genome position in bp; Ninth column, left flanking SNP of haplotype surrounding SNP marker; Tenth column, right flanking SNP of haplotype surrounding SNP marker; Eleventh column, Abacus reference genome position in bp for left flanking SNP of haplotype surrounding SNP marker; Twelfth column, Abacus reference genome

[0355] 5 position in bp for right flanking SNP of haplotype surrounding SNP marker.

[0356] Table 6. Significant SNP markers based on NAM of Harvest Height in the AF2 diversity panel (n=148). First column, SNP marker number;

[0357] Second column, SNP marker name; Third column, NAM p-value for Harvest Height; Fourth column, beneficial genotype for increased Harvest

[0358] 10 Height (A=homozygous for reference allele, B=homozygous for alternative allele, X=heterozygous); Fifth column, reference allele call; Sixth column, alternative allele call; Seventh column, chromosome; Eighth column, Abacus reference genome position in bp; Ninth column, left flanking SNP of haplotype surrounding SNP marker; Tenth column, right flanking SNP of haplotype surrounding SNP marker; Eleventh column, Abacus reference genome position in bp for left flanking SNP of haplotype surrounding SNP marker; Twelfth column, Abacus reference genome position in bp for right flanking SNP of haplotype surrounding SNP marker.

[0359]

[0360] Table 7. Significant SNP markers based on NAM of Harvest Height and Stretch in the 19GAR2 F2 mapping population (n=142). First column, SNP marker number; Second column, SNP marker name, *SNP marker is also significantly associated with Vegetative Height; Third column, NAM p-value for Harvest Height; Fourth column, beneficial genotype for increased Harvest Height (A=homozygous for reference allele, B=homozygous for alternative allele, X=heterozygous); Fifth column, reference allele call; Sixth column, alternative allele call; Seventh column,

[0361] 5 chromosome; Eighth column, Abacus reference genome position in bp; Ninth column, left flanking SNP of haplotype surrounding SNP marker; Tenth column, right flanking SNP of haplotype surrounding SNP marker; Eleventh column, Abacus reference genome position in bp for left flanking SNP of haplotype surrounding SNP marker; Twelfth column, Abacus reference genome position in bp for right flanking SNP of haplotype surrounding SNP marker.

[0362]

[0363]

[0364]

[0365]

[0366]

[0367]

[0368]

[0369]

[0370]

[0371]

[0372]

[0373] Table 8. 50 bp flanking sequences surrounding SNP markers. First column: SNP marker number; second column: SNP marker name; third column: 101 bp sequence with the SNP marker at position 51 bp, sequence is from the Abacus reference genome (version Csat_AbacusV2, NCBI assembly accession GCA_025232715).

[0374]

[0375] Table 9. Accessions used for gene sequencing and their SNP marker genotypes. First column: accession name, *beneficial genotype of mapped SNP markers near candidate genes (A=homozygous reference allele, X=heterozygous, B homozygous alternate allele), #Abacus and Finola splice variants of the CDS were determined based on Sanger sequencing of RNA, **selfed progeny of 20GAQ-1229, ##Autoflowering; Second column: plant height measured at full maturity when plants are harvested, except for Finola and the autoflowering plants all plants were topped 9-10 days

[0376] 5 after transplant, which is 4-5 days before flip to 12 hours light; Third - sixth columns: genotypes for SNP markers (A=homozygous reference

[0377] allele, X=heterozygous, B homozygous alternate allele). NA=data not available.

[0378]

[0379]

[0380]

[0381]

[0382]

[0383]

[0384]

[0385]

[0386]

[0387]

[0388]

[0389] Table 10. provides additional sequence information. First column: corresponding SEQ ID NO; Second column: sequence description, incomplete CDS and protein sequence, coordinates in brackets show the start and end of the sequence as compared to the Abacus reference genome homologous CDS and protein sequence, respectively, **truncated protein sequence as a result of a preliminary stop codon; Third column:

[0390] 5 sequences (genomic DNA, CDS, or protein sequences as indicated in the second column description of the sequences). N=missing nucleotide sequence, X=missing protein sequence.

[0391] It will be apparent that the precise details of the methods or compositions described may be varied or modified without departing from the spirit of the described aspects of the disclosure. We claim all such modifications and variations that fall within the scope and spirit of the claims below.

Claims

We claim:

1. A method for selecting a Cannabis plant having modified plant height, comprising:(i) analyzing one or more genetic markers in a nucleic acid sample from the Cannabis plant or its germplasm;(ii) detecting one or more genetic markers that indicate modified plant height; and(iii) selecting the Cannabis plant comprising the one or more genetic markers that indicate modified plant height.

2. The method of claim 1 , wherein the Cannabis plant having modified plant height is selected for further analysis, propagation, or to make a product.

3. The method of claim 1, further comprising:(iv) crossing the selected Cannabis plant and obtaining one or more progeny plants comprising the one or more genetic markers indicating modified plant height.

4. The method of claim 3, wherein the one or more progeny plants have modified plant height relative to a control.

5. The method of claim 3, wherein crossing comprises selfing, sibling crossing, outcrossing, or backcrossing.

6. The method of claim 5, wherein the selfing, sibling crossing, outcrossing, or backcrossing comprises marker-assisted selection for at least two generations.

7. The method of claim 1, wherein the modified plant height is an increase or decrease in vegetative height, harvest height, and / or stretch relative to a control.

8. The method of claim 7, wherein the modified plant height is an increase in vegetative height, harvest height, and / or stretch relative to a control.

9. The method of claim 1, wherein: analyzing comprises using PCR, quantitative PCR (qPCR), and / or sequencing; and / or detecting comprises using an oligonucleotide primer set or probe.

10. The method of claim 1, wherein detecting one or more genetic markers comprises detecting SNP marker 142603_12107817.

11. The method of claim 1, wherein detecting one or more genetic markers compri ses detecting one or more of SNP markers: 142603 2107817, 142603J 2439239, 1426O3_12333141, 142708_323806. and 142708_353959.

12. The method of one of the prior claims, wherein detecting one or more genetic markers in the nucleic acid sample comprises detecting one or more genetic markers described in Table 2, 3, 4, 5, 6, or 7.

13. The method of any one of the prior claims, wherein the one or more genetic markers that indicate modified plant height comprise a polymorphism at position 51 of one or more of: SEQ ID NOs: 1-205.

14. The method of any one of the prior claims, wherein the one or more genetic markers that indicate modified plant height comprise a polymorphism at position 51 of SEQ ID NO: 44, SEQ ID NO: 27, SEQ ID NO: 31, SEQ ID NO: 203, or SEQ ID NO: 193.

15. The method of any one of the prior claims, wherein the one or more genetic markers comprise a polymorphism relative to a reference genome in a plant height haplotype, wherein the plant height haplotype comprises the region on chromosome 4 between position 0.02 Mbp and 1.32 Mbp, and wherein the reference genome is Abacus Cannabis genome version Csat_AbacusV2, NCBI assembly accession GCA_025232715.1.

16. The method of any one of the prior claims, wherein analyzing one or more genetic markers comprises analyzing at least 2 genetic markers.

17. The method of any one of the prior claims, wherein detecting one or more genetic markers that indicate modified height comprises detecting at least 2 genetic markers.

18. The method of any one of the prior claims, wherein the one or more genetic markers are genetically linked to a plant height trait locus.

19. The method of any one of the prior claims, wherein the control is Abacus, or a parent plant without the one or more markers indicating modified plant height.

20. An engineered Cannabis plant comprising a genetic modification in a WRKY21, LBD16, S AUR51 , and / or TCS 1 gene that is associated with plant height.

21. A method of producing an engineered Cannabis plant having modified plant height, comprising introducing a genetic modification into a WRKY21, LBD16, SAUR51, and / or TCS1 gene that is associated with plant height.

22. The engineered Cannabis plant of claim 20 or the method of claim 21, wherein: a. the genetic modification in the WRKY21 gene comprises an amino acid substitution at position 146 of the WRKY21 coding sequence; b. the genetic modification in the LB 16 gene comprises an amino acid substitution at position 9, 130, or 131, or a two amino acid deletion or insertion between amino acid positions 141 - 144 of the LBD16 coding sequence; c. the genetic modification in the SAUR51 gene comprises a nucleic acid substitution at nucleotide position -4, -37, or -94 bp of the SAUR51 start codon; or d. the genetic modification in the TCS1 gene comprises an amino acid substitution at position 802 in the TCS1 coding sequence; wherein the amino acid or nucleic acid positions are in reference to Abacus Cannabis genome version Csat_AbacusV2, NCBI assembly accession GCA_025232715.1, or a coding sequence obtained therefrom.

23. The method of claim 22, wherein: a. the amino acid substitution at position 146 of the WRKY21 coding sequence is a serine or threonine substitution; b. the amino acid substitution at position 9 of the LBD16 coding sequence is a serine or glycine substitution, the amino acid substitution at position 130 of the LBD16 coding sequence is a glutamine or histidine substitution, the substitution at position 131 of the LBD16 coding sequence introduces a preliminary stop codon or removes a preliminary stop codon, or the amino acid deletion or insertion between amino acid positions 141 - 144 of the LBD16 coding sequence is a two amino acid insertion or deletion; c. the nucleic acid substitution at position -4 of the SAUR51 gene comprises a G or A substitution; the nucleic acid substitution at position -37 of the SAUR51 gene is a T or A substitution, or the nucleic acid substitution at position -94 of the SAUR51 gene is a C or G substitution; ord. the amino acid substitution at position 802 of the TCS 1 coding sequence is a S or N substitution; wherein the amino acid or nucleic acid positions are in reference to Abacus Cannabis genome version Csat_AbacusV2, NCBI assembly accession GCA_025232715.1 or a coding sequence obtained therefrom.

24. A method of selecting a Cannabis plant having modified plant height, comprising:(i) analyzing a WRKY21, LBD16, SAUR51, and / or TCS1 gene in a nucleic acid sample from the Cannabis plant or its germplasm;(ii) detecting one or more polymorphisms that indicate modified plant height; and(iii) selecting a Cannabis plant having the one or more nucleic acid or amino acid substitutions that indicate modified plant height, thereby selecting the Cannabis plant having modified plant height.

25. The engineered Cannabis plant of claim 20 or the method of claim 24, wherein the Cannabis plant having modified plant height is selected for further analysis, propagation, crossing, or to make a product.

26. A Cannabis plant produced by the method of any one of the prior claims.

27. A seed, plant part, tissue culture, or protoplast of the plant of claim 26.

28. A method of Cannabis breeding, comprising crossing the Cannabis plant of claim 26.

29. The method of claim 28, wherein crossing comprises selfing, sibling crossing, outcrossing, or backcrossing.

30. A Cannabis product produced from the plant of claim 26, or the seed, plant part, tissue culture, or protoplast of claim 27.

31. The Cannabis product of claim 30, wherein the product is a kief, hashish, bubble hash, an edible product, solvent reduced oil, sludge, e-juice, or tincture.

Citation Information

Patent Citations

  • Cannabis genomes and uses thereof

    US20160177404A1