Process for regulating plant gene expression
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- PHYTOFORM LABS LTD
- Filing Date
- 2023-05-23
- Publication Date
- 2026-06-01
AI Technical Summary
Conventional methods for designing plant gene regulatory sequences are largely trial-and-error and biased towards known solutions, lacking the ability to identify subtle or non-intuitive regulatory changes, and are not precise enough for targeted genetic modifications.
An informatics-led approach using in silico analysis of the plant genome to identify promoterome sequences and transcriptome data, combined with machine learning algorithms, to predict and design novel expression control sequences that can regulate gene expression more precisely.
Enables the generation of novel gene regulatory sequences that can be verified in vivo, providing desired expression profiles with high precision and efficiency, allowing for targeted genetic modifications in plants.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of in silico design involving in vivo validation of any novel plant gene regulatory sequences.
Background Art
[0002] Plants exhibit a surprising ability to phenotypic plasticity resulting from genomic changes brought about by natural mutations, DNA translocations, polyploidization, or mating with distant relatives. Since the beginning of agriculture, farmers have utilized this plasticity to develop new plant varieties that enhance crop yields and favor traits such as disease resistance. By repeatedly selecting for desirable traits, such as increased seed or fruit size, plant species have changed phenotypically over time, providing the wide range of varieties and cultivars used in modern agriculture. Since the latter half of the previous century, knowledge of the plant genome has increasingly guided efforts to create new crop varieties. However, modern breeding still largely relies on the genetic diversity that occurs spontaneously in nature.
[0003] During the past few decades, recombinant genetic engineering methods have emerged that enable the creation of targeted genetic mutations, eliminating the need to rely on naturally occurring mutations that happen by chance. One such method is genetic recombination, which involves adding new genes conferring desirable traits, such as herbicide tolerance, to the plant genome. However, with few exceptions, genetic recombination has only been introduced into a very small number of high-value crops and is considered controversial in some countries. Mutations can also be created by altering the genome of specific crops using mutagenic chemicals such as ethyl methanesulfonate or high-energy radiation. The drawback of mutation breeding is that it is largely uncontrollable and requires generating and screening large mutation populations to identify the very few valuable genetic mutations. Newer, more targeted mutagenesis methods, such as / CAS-genome editing, offer higher precision, but it is not possible to clearly understand what the valuable results of mutagenesis are.
[0004] Conventional approaches have attempted to perform so-called rational design of promoter elements in plant cells using an intuitive researcher-led approach (see Yang et al. (2021) Plant Biotech.J. 19: 1364-1369). This methodology is highly trial-and-error and is heavily biased towards a narrow range of existing solutions known to researchers from the literature and existing natural mutations. These approaches also typically rely on alignment methods for de novo motif discovery, such as motif-based sequence analysis tools like the meme-suite (https: / / mem-suite.org), thereby rationally constructing a narrow range of solutions. Thus, these techniques greatly reduce the degrees of freedom that may exist to identify more subtle or non-intuitive positive or negative regulatory changes.
[0005] Therefore, it is desirable to provide a new method for developing methods for improving specific traits in plants. In particular, it is desirable to improve such traits through the development of novel plant gene regulatory sequences beyond natural or mutagenic mutations that control plant gene expression to provide such desirable traits.
[0006] These and other uses, features, and advantages of the present invention should be apparent to those skilled in the art from the teachings provided herein.
SUMMARY OF THE INVENTION
[0007] The inventors have provided a new method that utilizes a novel informatics-led approach to drive the evolution of novel expression control sequences in plants. These novel sequences may be further verified using in vivo screening methods.
[0008] Accordingly, a first aspect of the present invention provides a method for generating a nucleic acid library comprising a plurality of expression control sequences configured to regulate the expression of a coding sequence operably linked in a plant cell, the method comprising Performing in silico analysis of the genome of a plant, the in silico analysis including identifying a plurality of sequences that collectively define the promoterome of the plant, each sequence within the promoterome consisting of an expression control region extending from the start codon of an open reading frame to about 100 kilobases 5' and / or 3' of the start codon, and Obtaining a transcriptome in the form of mRNA expression data of the plant, and Performing an analysis using an array-based modeling algorithm to provide a predicted value of the expression level of each protein-coding gene contained within the transcriptome, and linking that value to the sequence of the corresponding expression control region contained within the promoterome, and Generating a plurality of non-wild-type sequence designs of expression control sequences that are most likely to provide a desired expression profile to an operably linked coding sequence, the plurality of non-wild-type sequence designs being informed by the predicted values, and Synthesizing a plurality of non-wild-type expression control sequences corresponding to the plurality of non-wild-type sequence designs, and Generating a nucleic acid sequence library comprising the plurality of non-wild-type expression control sequences.
[0009] A second aspect of the invention provides a nucleic acid library comprising a plurality of non-wild-type expression control sequences configured to regulate the expression of an operably linked coding sequence within a plant cell. Suitably, the plurality of non-wild-type expression control sequences are generated by the methods defined herein.
[0010] Plant cells, protoplasts, plant tissues, calli, seedlings, seeds, whole plants, and other biological materials comprising the nucleic acid library or its component sequences are also provided in various aspects and embodiments of the invention.
[0011] A third aspect of the invention provides a method that can be implemented, wholly or in part, on a computer for performing an analysis of the genome of a plant, the analysis comprising Identification of a plurality of arrays that collectively define a promoterome dataset for a plant, wherein each array within the promoterome dataset consists of array data corresponding to an expression control region extending from the start codon of an open reading frame to approximately 100 kilobases 5' and / or 3' of the start codon, the identification, and Obtaining a transcriptome dataset in the form of mRNA expression data for a plant, and Performing an analysis using an array-based machine learning modeling algorithm trained using all or part of the promoterome dataset and all or part of the transcriptome dataset to provide a predicted value of the expression level of a gene placed under the operative control of any given expression control sequence contained within the promoterome. Querying an array-based machine learning modeling algorithm using one or more query sequence designs, wherein the array-based machine learning modeling algorithm applies the predicted value to the one or more query sequence designs to provide a prediction of gene expression within the plant when the one or more query sequence designs are introduced into the gene's expression control region.
[0012] A fourth aspect of the present invention provides a method for generating a modified plant cell, comprising performing a genomic analysis of a plant cell as described herein, identifying at least one sequence design having a desired prediction of gene expression within the plant, and modifying an expression control region to conform to the at least one sequence design within the plant cell to obtain a modified plant cell having the desired expression of the gene.
[0013] A fifth aspect of the present invention provides a modified plant cell obtained by, or obtainable by, the method described herein.
[0014] A sixth aspect of the present invention provides a modified plant derived from the plant cell described herein.
[0015] Within the scope of the present application, it is expressly intended that the various aspects, embodiments, examples and alternatives described in the previous paragraphs, the claims, and / or the following description and drawings, in particular their individual features, can be used alone or in any combination. That is, all embodiments and / or features of any embodiment can be combined in any way and / or combination as long as such features are not incompatible.
[0016] One or more embodiments of the present invention will now be described by way of example only, with reference to the accompanying drawings.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3
Figure 4(a)
Figure 4(b)
Figure 4(c)
Figure 5
Figure 6
Figure 7
Figure 8(a)
Figure 8(b)
Figure 8(c)
Figure 9
Figure 10
Figure 11(a)
Figure 11(b)
Figure 12(a)
Figure 12(b)
Figure 13(a)
Figure 13(b)
Figure 14(a)
Figure 14(b)
BEST MODE FOR CARRYING OUT THE INVENTION
[0018] Before explaining the present invention, many definitions are provided to assist in the understanding of the present invention.
[0019] Unless otherwise indicated, the practice of the present invention employs conventional techniques of chemistry, molecular biology, microbiology, recombinant DNA technology, and chemical methods within the capabilities of those of ordinary skill in the art. Such techniques are described in the literature, for example, M.R. Green, J. Sambrook, 2012, Molecular Cloning: A Laboratory Manual, 4th Edition, Volumes 1-3, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, Ausubel, F.M. et al. (Current Protocols in Molecular Biology, John Wiley & Sons, Online ISSN: 1934-3647), B. Roe, J. Crabtree and A. Kahn, 1996, DNA Isolation and Sequencing: Essential Techniques, John Wiley & Sons, J.M. Polak and James O’D. McGee, 1990, In Situ Hybridisation: Principles and Practice, Oxford University Press, M.J. Gait (Editor), 1984, Oligonucleotide Synthesis: A Practical Approach, IRL Press, and D.M.J. Lilley and J.E.Dahlberg, 1992, Methods of Enzymology: DNA Structure Part A: Synthesis and Physical Analysis of DNA Methods in Enzymology, Academic Press; Synthetic Biology, Part A, Methods in Enzymology, edited by Chris Voigt, Volume 497, pages 2 - 662 (2011); Synthetic Biology, Part B, Computer Aided Design and DNA Assembly, Methods in Enzymology, edited by Christopher Voigt, Volume 498, pages 2 - 500 (2011); RNA Interference, Methods in Enzymology, David R. Engelke and John J. Rossi, Volume 392, pages 1 - 454 (2005) are also described. All references cited herein are incorporated by reference in their entirety. Unless otherwise expressly stated, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs..
[0020] As used herein, the term "comprising" means that any of the recited elements are necessarily included and other elements may optionally be included as well. "Consisting essentially of" means that the recited elements are necessarily included and elements that materially affect the basic and novel characteristics of the recited elements are excluded, and other elements may optionally be included. "Consisting of" means that all elements other than those recited are excluded. Embodiments defined by each of these terms are within the scope of the present invention.
[0021] The term "expression vector" is used to refer to a linear or circular DNA molecule that can incorporate another DNA sequence fragment of appropriate size. Such DNA fragment(s) can include additional segments that provide for the transcription of the gene(s) encoded by the DNA sequence fragment. The additional segments can include, but are not limited to, a promoter, a transcription terminator, an enhancer, an internal ribosome entry site, an untranslated region, a polyadenylation signal, a selectable marker, an origin of replication, etc. In embodiments of the present invention, the DNA sequence fragment can include a plant expression control sequence that includes a novel non-wild-type plant promoter, enhancer or silencer element. Expression vectors are typically derived from plasmids, cosmids, viral vectors, yeast artificial chromosomes. The vector is suitably a recombinant molecule that includes DNA sequences from several sources. Nucleic acid expression and reporter libraries of the type described herein may be included within the expression vector.
[0022] The term "operably linked", when applied to a DNA sequence in, for example, an expression vector or a genetically engineered gene construct, indicates that the sequences are arranged so as to function in concert to achieve an intended purpose, e.g., the start of transcription through a linked coding sequence to a termination sequence is enabled by an expression control sequence such as a promoter sequence.
[0023] A "polynucleotide" is a single-stranded or double-stranded covalently bonded nucleotide sequence in which the 3' and 5' ends of each nucleotide are linked by phosphodiester bonds. A polynucleotide can be composed of deoxyribonucleotide bases or ribonucleotide bases. Polynucleotides include DNA and RNA and can be synthetically produced in vitro or isolated from natural sources. The size of a polynucleotide is typically represented by the number of base pairs (bp) of a double-stranded polynucleotide or, in the case of a single-stranded polynucleotide, the number of nucleotides (nt). 1000 bp or nt corresponds to a kilobase (kb). Polynucleotides less than about 40 nucleotides in length are usually referred to as "oligonucleotides".
[0024] A "polypeptide" is a polymer of amino acid residues linked by peptide bonds, whether produced naturally or in vitro by synthetic means. Polypeptides that are less than about 12 amino acid residues in length are usually referred to as "peptides". As used herein, the term "polypeptide" means a naturally occurring polypeptide, a precursor form, or the product of a protein. Polypeptides can undergo maturation or post-translational modification processes, including but not limited to glycosylation, proteolytic cleavage, lipidation, signal peptide cleavage, propeptide cleavage, phosphorylation, etc. A "protein" is a macromolecule that contains one or more polypeptide chains.
[0025] As used herein, the term "promoter" refers to a genetic regulatory element within a DNA sequence to which RNA polymerase binds to initiate transcription of DNA. The promoter plays an important role in gene expression by providing a binding site for RNA polymerase. When RNA polymerase binds to the promoter region, it initiates the transcription process. Promoters are usually present in the 5' non-coding region of a gene, but not always. The 5' region refers to the upstream region of the gene, which typically means it is before the actual coding sequence of the gene, often indicated by the ATG start codon (e.g., before the first exon). The non-coding region is a segment of DNA that does not directly contribute to the formation of a polypeptide or other gene product. These regions may contain various regulatory elements such as promoters. The main function of the promoter sequence is to provide a recognition site for RNA polymerase and other transcription regulatory proteins, enabling them to interact with DNA and initiate the transcription process. The binding of RNA polymerase to the promoter region indicates the starting point for the assembly of the transcription machinery, ultimately leading to the synthesis of an RNA molecule known as the primary transcript or pre-mRNA. Thus, promoters are highly diverse in terms of their sequence and structure. They contain specific DNA motifs and sequences recognized by transcription factors that further regulate gene expression. Transcription factors can enhance or inhibit the binding of RNA polymerase to the promoter, thereby often affecting the level of gene transcription in a cell-type or tissue-specific manner.
[0026] As used herein, the term "enhancer" refers to a genetic regulatory element within a DNA sequence that, when bound by one or more transcription factors, enhances the transcription of the associated gene. Enhancers play a critically important role in gene expression by regulating the transcription of associated genes or gene sets within a locus. When an enhancer binds to one or more transcription factors, it increases the rate of transcription. Enhancers are typically located at various distances from the gene(s) they regulate. They can be present either upstream (upstream enhancer) or downstream (downstream enhancer) of the gene(s), and in some cases, can also be present within the introns of the gene itself. Unlike promoters, enhancers are not necessarily direction-specific and can function regardless of their orientation relative to the gene. An important function of enhancers is to provide binding sites for transcription factors and regulatory complexes. When a specific transcription factor recognizes and binds to an enhancer, it can facilitate the assembly of the transcriptional machinery at the promoter region of the associated gene. This recruitment and interaction of transcription factors at the enhancer and promoter regions enable the efficient initiation and regulation of gene transcription. Enhancers exhibit great flexibility and can act over long distances. They can interact with the promoter region of the target gene through three-dimensional loops of DNA, bringing regulatory elements into proximity. This spatial arrangement allows transcription factors bound to the enhancer to directly interact with the transcriptional machinery of the promoter, leading to increased transcriptional activity. Enhancers may also have cell-type-specific or developmental-stage-specific activity. This means that an enhancer may be activated only in a specific cell type or at a specific stage of development, which contributes to the precise regulation of gene expression. The specificity and activity of enhancers are governed by the combination of transcription factors that bind to them, forming a complex regulatory network that determines the timing, level, and specificity of gene expression. Furthermore, enhancers can act synergistically in combination with other enhancers or regulatory elements. This cooperation between multiple enhancers allows for fine-tuning of gene expression patterns, enabling cells to respond to various environmental cues and signaling pathways.The combinatorial effect of enhancers provides a powerful and dynamic mechanism for gene regulation, ensuring the proper function and adaptation of cells in various situations, particularly when conferring tissue specificity in the form of phenotypic gene expression.
[0027] As used herein, the term "silencer" refers to genetic regulatory elements within a DNA sequence that reduce transcription from the associated promoter. Typically, they are the inhibitory counterparts of enhancers. Silencers play an important role in reducing or suppressing the transcriptional activity of an associated or adjacent promoter and are useful for the fine-tuning of gene expression. Silencers are usually located near the promoter region of the gene(s) they regulate. They can be found upstream (upstream silencers), downstream (downstream silencers), or even within the introns of a gene. Similar to enhancers, silencers are not necessarily direction-specific and can function regardless of their orientation relative to the gene. The main function of a silencer is to provide a binding site for transcription factors that have an inhibitory effect on gene transcription. When specific transcription factors recognize and bind to a silencer, they can recruit corepressor proteins or inhibit the binding of activator proteins to the promoter region. This interference suppresses the transcriptional activity from the associated promoter. Silencers can exert their inhibitory effect in various ways. They can interact directly with the transcriptional machinery at the promoter region and prevent the assembly of components necessary for transcription initiation. Silencers can also induce chromatin modifications such as the addition of methyl groups to DNA or the removal of acetyl groups from histones. These modifications change the chromatin structure, making the DNA less accessible to the transcriptional machinery and inhibiting gene expression. Similar to enhancers, silencers can exhibit cell-type-specific or developmental-stage-specific activity. This means that a silencer may be activated only in a specific cell type or at a specific stage of development, adding another level of complexity to gene regulation. The specific combination of transcription factors that bind to a silencer determines its activity and the inhibitory effect on gene transcription. Silencers can also interact and function cooperatively with other regulatory elements such as other silencers or enhancers to regulate gene expression. These elements work together to fine-tune transcriptional activity and establish precise gene expression patterns in response to various signals and environmental cues.Thus, through the recruitment of repressive transcription factors and chromatin modifications, silencers function as dampeners of transcriptional activity, enabling cells to precisely regulate gene expression levels. Due to their cell-type specific and cooperative nature, gene regulatory networks become complex, ensuring appropriate gene expression patterns during development and in response to various cellular contexts. In certain situations, silencers can also be bifunctional regulatory elements that function as enhancers, again depending on the cellular context.
[0028] As used herein, the terms "3′" ("3 prime") and "5′" ("5 prime") shall have their ordinary meaning in the art, namely, to distinguish the ends or directionality within a polynucleotide sequence. A polynucleotide has a 5' end and a 3' end, and a polynucleotide sequence is typically written in the 5' to 3' direction. The 5' end is appropriately considered to be upstream of the 3' end of the polynucleotide sequence. Thus, a sequence referred to as upstream of a specific reference point within a gene, such as the transcription start codon of an open reading frame (ORF), is a sequence that is 5' relative to the reference point. Similarly, a sequence designated as downstream is 3' of the reference point.
[0029] According to the present invention, the homology to any nucleic acid sequence such as the expression control sequences described herein is not limited to simply 100%, 99%, 98%, 97%, 95%, or even 90% sequence identity. Despite having clearly low sequence identity, many nucleic acid sequences can exhibit significantly biochemical homology to each other. In the present invention, homologous nucleic acid sequences are considered to be sequences that hybridize to each other under low stringency conditions (Sambrook J. et al., Molecular Cloning: a Laboratory Manual, Cold Spring Harbor Press, Cold Spring Harbor, NY). However, in some cases, it may be desirable to distinguish between two sequences that can hybridize to each other but contain some mismatches ("inaccurate matches", "incomplete matches", or "inaccurate complementarities") and those that have no mismatches, i.e., two sequences that can hybridize to each other and contain "accurate matches", "perfect matches", or "accurate complementarities". Further, the degree of possible mismatches is also considered.
[0030] Accordingly, the present invention provides a novel discovery platform for plant expression control sequences, including novel plant promoter, enhancer and / or silencer elements obtained beyond natural diversity. In embodiments, the present invention utilizes a combination of artificial intelligence (AI) techniques and an approach that provides additional options for in vivo high-throughput library screening for precise control of plant gene expression. In certain embodiments of the present invention, a novel nucleic acid sequence library of gene expression control elements is provided that can be screened for desired gene expression elements. Plant cells, plant tissues, and plants (species, varieties, or cultivars) containing one or more novel gene expression elements can be generated, and desired traits can be enhanced.
[0031] Embodiments of the present invention utilize a bioinformatics approach using AI to direct the evolution of novel gene expression control sequences. Exemplary embodiments of the present invention are described in more detail below.
[0032] Unless already assembled, de novo DNA genome assembly is performed on a target plant, such as a plant species or variety, and related varieties. The working data for genome assembly may include a mixture of Illumina short reads and long reads, or long reads and short reads (e.g., using PacBio or Oxford Nanopore sequencing), and may include bioinformatics assembly techniques that take into account the ploidy of the plant and the distribution of alleles. For example, when using a Bayesian statistical framework with a haplotype-based variant detector such as FreeBayes, see Garrison E, Marth G. Haplotype-based variant detection from short-read sequencing.arXiv preprint arXiv:1207.3907 [g-bio.GN]2012.
[0033] Sequences are identified for the regions around / above each annotated coding gene identified within the assembled genome. Coding genes can be identified using reference genome data, supplementation of published literature, or homology analysis of the constructed genome. These genomic DNA sequences are fixed around a conceptual translation start position (e.g., start codon) or transcription start position (e.g., TATA box) and define individual data strings that extend up to approximately 100,000 base pairs upstream in their 5'. The promoterome may further or alternatively include a 3' downstream of a similar distance, and / or may similarly include the coding gene. This captures multiple DNA sequences that may collectively define the "promoterome" of the target plant. In another embodiment of the invention, the promoterome may include only sequences located 5' or upstream of the open reading frame. In another embodiment, the promoterome may include only sequences located 3' or downstream of the open reading frame. The promoterome may include one or more exons, one or more introns, one or more transposable elements, and / or one or more heterologous insertion elements that may have been introduced by previous gene editing or recombination techniques.
[0034] In certain embodiments of the present invention, the promoterome may include a plurality of nucleic acid sequences extending up to 100,000 base pairs (bps), 90,000 bps, 80,000 bps, 70,000 bps, 60,000 bps, 50,000 bps, 40,000 bps, 30,000 bps, 20,000 bps, 10,000 bps, and 5,000 bps upstream and / or downstream of the first translation start position. In embodiments of the present invention, the promoterome is defined as including a plurality of sequences that extend a distance longer than the distance extending downstream from a reference point within the gene, suitably at least 1.5-fold, at least 2-fold, at least 3-fold, and up to at least 4-fold upstream. In another alternative embodiment of the present invention, the promoterome is defined as including a plurality of sequences that extend a distance shorter than the distance extending downstream from a reference point within the gene, suitably at least 1.5-fold, at least 2-fold, at least 3-fold, and up to at least 4-fold upstream. As described above, the reference point can be the transcription or translation start site, and suitably can include any one of the possible transcription start sites when alternative variants or transcripts may exist for a given gene.
[0035] Transcriptome information may be provided in the form of expression data for a plant subject. The transcriptome can be obtained for a specific plant cell, plant cell type, plant tissue type (such as leaves or fruits). The transcriptome can be obtained from bulk RNA sequencing technology (RNA-seq), and UMI (unique molecular identifier or "barcode")-based single-cell RNA-seq, such as, but not limited to, drop-RNAseq, and other RNA-seq technologies that utilize transient transfection such as the STARR-seq method and the SuRE-seq method. Also, in certain embodiments of the present invention, the data input can be supplemented with an add-on dataset to evaluate chromatin accessibility and transcription factor binding information.
[0036] Suitable chromatin-based sequencing methodologies include the following. -ATAC-seq (Assay for Transposase-Accessible Chromatin using High-Throughput Sequencing): ATAC-Seq is a method for investigating chromatin accessibility within a sample. In this assay, the genome is treated with a transposase (enzyme) called Tn5. Tn5 marks open chromatin regions by cutting and inserting adapter sequences, which can then be detected by subsequent sequencing. ATAC-seq has shown utility in assessing changes in genome-wide chromatin accessibility following editing events (Buenrostro et al., Curr Protoc Mol Biol. (2015); 2015:21.29.1-21.29.9). -Using chromosome conformation capture (3C) methodologies such as Hi-C analysis, chromatin accessibility, the 3D organization of the genome, and interconnectivity can be evaluated, and changes to chromatin accessibility and the 3D structure of the genome, including local chromosomal neighborhoods and / or transcription factories, can be identified (Lieberman-Aiden et al., Science. October 9, 2009; 326(5950):289-293) -Chromatin immunoprecipitation followed by sequencing (ChIP-seq): ChIP-seq is a central method in epigenomic research. Genome-wide analysis of histone modifications, such as enhancer analysis and genome-wide chromatin state annotation, allows for a systematic analysis of how the epigenomic landscape contributes to cell identity, development, lineage specification, and disease. In particular, it is necessary to assess the post-editing effects on histone modifications, including but not limited to regulatory elements (H3K27Ac, H3K4Me1), promoter accessibility (H3K4Me3), heterochromatin formation (H3K9Me3), gene bodies (H3K36Me3, H3K27Me3), etc. (see Furey (2012) Nat. Rev. Genet., 13 (12), pp. 840-852). -Methyl-seq: This approach assesses the impact of editing on the DNA methylation profile within the genome, thereby inferring changes to chromatin accessibility. Methyl-seq can be performed using either a chemical approach (bisulfite sequencing) or an enzymatic approach (EM-Seq) (see Vaisvila et al. (2021) Genome Res. Jul; 31(7): 1280-1289). -DNase accessibility data: Deoxyribonuclease (DNase for short) is an enzyme that catalyzes the hydrolytic cleavage of phosphodiester bonds in the DNA backbone to degrade DNA. Deoxyribonuclease is a type of nuclease, a general term for enzymes that can hydrolyze the phosphodiester bonds connecting nucleotides. DNase activity is one way to evaluate chromatin accessibility and define the importance of cell-type specific regions within the target plant tissue / cells.
[0037] Analysis of data obtained from one or more of the above sequencing techniques enables the evaluation of promoterromes at the epigenetic and transcriptional levels. This evaluation includes, for example, the determination of the activity and chromosomal accessibility of transcription factor binding sites, transcription factor coding regions, gene regulatory elements (such as enhancers, silencers, repressors, etc.), DNA methylation, histone modifications, transcription factories, etc.
[0038] Analysis of the input data for promoterromes and transcriptomes is achieved by using a sequence-based modeling algorithm configured to provide predicted values of the expression levels of each protein-coding gene contained within the transcriptome and linking those predicted values to the corresponding sequences of the expression control regions contained within the promoterrome.
[0039] Bioinformatics processing of expression data may include, but is not limited to, the following techniques. -Mapping of next-generation sequencing reads to create read counts - for example, HISAT2 Normalization and expression analysis - for example, DESeq2 Removal of batch effects - for example, combat
[0040] Array-based modeling algorithms may utilize artificial intelligence (AI) and machine learning (ML) approaches. The model can be trained using the input data related to the aforementioned transcriptomics and proteomics to provide predictions called predicted values of gene expression contained within the genome assembly. The predicted values can be assigned to genes that enable the investigation of all genes identified within the database of said genes. The database can be investigated based on a number of criteria and metadata assigned to each entry, such as gene sequences, coding sequences, protein sequences, expression levels, associated UMIs, and predicted values. The gene sequence may include other parts of the gene, such as regulatory regions (promoters and non-coding parts), 5' and 3' untranslated regions (UTRs), coding sequences (CDSs), introns, exons, and other regions. Furthermore, the model can be used to investigate the expression values of new sequences outside of the gene database and, if necessary, subsequently perform further transcriptional and in vivo validation.
[0041] As used herein, the term "predicted value" may refer to a numerical value or score that characterizes the expression of a gene. The predicted value can be a relative value such as an expression level (positive or negative) compared to a conceptual benchmark such as the expression of a housekeeping or other appropriate reference gene. Thus, the predicted value can be a numerical value, a logarithmic value, or other non-dimensional value, for example, where the predicted value is a color or alphanumeric character code within a specific threshold band. In alternative embodiments, the predicted value may be an absolute value corresponding to a predicted quantification of gene expression typically determined via one or more techniques including, but not limited to, RNA sequencing (RNA-seq), microarray analysis, quantitative RT-PCR (qPCR), Northern or Western blotting, in situ hybridization, and immunohistochemistry. Thus, according to certain embodiments of the present invention, the predicted value is represented as an RNA-seq count for a given sequence design. The RNA-seq count can be provided as a unit count such as FPKM (fragments per kilobase of transcript per million) or TPM (transcripts per million).
[0042] Array-based modeling architectures and array prediction algorithms that include commonalities in natural language processing can be used in embodiments of the present invention to provide predicted values of the expression levels of each protein-coding gene included in a transcriptome and to link the predicted values to the corresponding sequences of expression control regions included in a promotorome. Suitable array-based modeling algorithms can include, for example, artificial neural networks (ANN) such as convolutional neural networks (CNN), recurrent neural networks (RNN) including bidirectional RNN, transformers, and masked language models. These algorithms specifically implement a series of machine learning models called neural networks, and the "neurons" of the network are arranged in an architecture suitable for modeling continuous data such as genomic sequences. They are used in various biological applications known to those skilled in the art. An explanation of the algorithms and their successful applications to biology is provided, for example, by Greener et al., Nature Reviews Molecular Cell Biology, Volume 23: 40-55 (2022).
[0043] According to one embodiment of the present invention, multiple non-wild-type sequence designs are generated for novel expression control sequences that are most likely to provide a desired expression profile to an operably linked coding sequence. The multiple non-wild-type sequence designs are informed by the predicted values generated by the array-based modeling steps described herein. Thus, if it is desired to increase the expression of an operably linked coding sequence, sequences showing predicted high-level expression values are searched for. Conversely, if it is desired to decrease the expression of an operably linked coding sequence, sequences showing predicted low-level expression values are searched for.
[0044] Multiple array designs for novel expression control arrays are generated via an in-silico mutagenesis approach. Such an approach may include rule-based array design and / or generated array design. In certain embodiments of the invention, a "minimal change" specification is employed that prioritizes the smallest amount of base pair changes, insertions, or deletions compared to a template / starting wild-type (WT) promoter. The crop-specific requirement of "minimal change" can address concerns regarding the use of novel breeding technologies (NBTs) such as CRISPR-CAS-guided endonuclease-based genome editing to create non-transgenic and targeted changes in the crop genome. In some jurisdictions, such an approach is classified as a non-genetically modified organism (non-GMO).
[0045] The minimal change approach is also a significant advantage compared to other known synthetic AI-based design approaches used in industrial biotechnology or medicine, e.g., "fitness" to local or global maxima, regardless of the number of changes compared to a particular starting sequence. In contrast, the minimal change approach enables optimization of attributes, e.g., fitness to pass quality control (QC) thresholds for in-plant performance, taking into account (a) the original function of the promoter / enhancer / silencer, and (b) the regulatory aspects of real-world legislation and DNA changes in plants intended for use as commercial crops.
[0046] Thus, in embodiments of the present invention, a minimum change threshold is established for non-wild-type sequence designs in relation to the number and / or type of modifications that are allowed to be made to the expression control sequence from the wild type. This minimum change threshold can be incorporated into a rule-based sequence and / or generative sequence design algorithm that employs sequence perturbation to design mutant sequences having a desired gene expression profile. In embodiments of the present invention, concepts from the field of explainable AI can be used to estimate the location and structure of changes in the wild-type sequence that are most likely to achieve a desired change in the predicted expression value, such as by assigning importance of features or strength of interactions (see Molnar, C., "Interpretable Machine Learning: A Guide For Making Black Box Models Explainable", 2nd Edition (2022), ISBN: 979-8411463330).
[0047] The novel non-wild-type sequence design may be scored using in-silico high-throughput techniques, such as one or more AI-based trained models, to determine the top sequence designs that are most likely to provide an optimal desired expression profile.
[0048] In vivo validation of the novel non-wild-type sequence design can be performed via the design and construction of a library of novel gene expression control sequences created by novel DNA synthesis. Equivalent wild-type gene expression control sequences can be newly synthesized or amplified from plant source materials by standard recombinant techniques such as polymerase chain reaction (PCR).
[0049] The method of the present invention can be used to generate a variety of gene control elements that can be introduced into plant cells and / or tissues. Suitably, the plant cells can be selected from gametes, germ cells, vegetative cells and / or meristematic cells.
[0050] In certain embodiments, the plant cell is in the form of a protoplast. As used herein, the term "plant protoplast" (also simply referred to as "protoplast" throughout the present disclosure) refers to a plant cell from which the cell wall has been completely or partially removed. Removal of the cell wall can be accomplished by mechanical, chemical, or enzymatic means. In embodiments, the protoplasts are obtained from suitable plant material using cell wall-digesting enzymes. For example, enzymes such as cellulase, macerozyme, pectinase, hemicellulase, pectolyase, driselase, xylanase, and combinations thereof may be suitable for use in the context of the present invention. In embodiments, cellulase can be used at a concentration of 1 w% to 1.5 w%. In embodiments, macerozyme can be used at a concentration of 0.2 w% to 0.4 w%. In embodiments, hemicellulase can be used at a concentration of 2 w% to 5 w%. In embodiments, pectolyase can be used at a concentration of 0.01 w% to 0.5 w%. In embodiments, driselase can be used at a concentration of 0.5 w% to 2 w%. Protocols for obtaining protoplasts from plant tissue are known in the art and will not be described further herein.
[0051] In embodiments, suitable plant tissue is selected from leaves, stems, roots, tubers, seeds, branches, trichomes, root nodules, leaf axils, flowers, pollen, stamens, pistils, petals, pedicels, petioles, stigmas, styles, bracts, fruits, trunks, carpels, scales, anthers, ovules, pedicels, needles, cones, rhizomes, stolons, shoots, pericarp, endosperm, placenta, berries, stamens, or leaf sheaths. In certain embodiments, the plant cell material can include root tissue, mesophyll, and / or cultured callus. In embodiments, the protoplasts are obtained using the protocol described in Yoo, Cho, & Sheen (2007) Nature Protocols, Volume 2, pages 1565 - 1572, which is incorporated herein by reference.
[0052] In embodiments, in vivo verification may include a method of detecting and optionally selecting encapsulated protoplasts or plant cells based on one or more desired characteristics. The system of the present invention includes an optical detection system such as a fluorescence detection system. By combining this with a sorting mechanism, it becomes possible to select protoplasts or cells having a fluorescence marker or a simple expression reporter indicating the desired characteristics. In embodiments, the method of the present invention includes introducing a coding sequence of a fluorescent protein into a protoplast or a plant cell culture. In such embodiments, the method of the present invention may include using a fluorescence detection system such as a fluorescence-activated cell sorting (FACS) system to detect the expression of the fluorescent protein in the protoplasts or plant cells to be assayed. Suitable screening and propagation methodologies are described and exemplified, for example, in WO-A-2020 / 212713.
[0053] Examples of intracellular reporter proteins that can function as fluorescence markers for gene expression include green fluorescent protein (GFP) and its homologs or derivatives, such as enhanced GFP (eGFP), blue fluorescent protein (BFP, Azurite, mKalama1), cyan fluorescent protein (CFP, CyPet), yellow fluorescent protein (TFP, Citrine), and mCherry. This allows transfected protoplasts to be easily identified using conventional cell sorting techniques.
[0054] In embodiments, the fluorescence detection system may be configured to detect both a signal indicating the presence of chlorophyll and a signal indicating the presence of a fluorescence marker indicating the characteristic of interest. In such embodiments, the detection system may be coupled to a sorting mechanism, and the system may be configured such that the sorting mechanism separates microcapsules between two different channels based on the presence of a combination of a signal indicating the presence of chlorophyll and a signal indicating the presence of a fluorescence marker.
[0055] In an embodiment, a chemiluminescence detection system can also be utilized. By combining this with a sorting mechanism, it becomes possible to select protoplasts having a chemiluminescent marker indicating the target characteristics. In an embodiment, the method of the present invention includes introducing a coding sequence of one or more chemiluminescent proteins (for example, luciferin, aequorin, etc.) into a protoplast culture. In such an embodiment, the method of the present invention can include detecting the expression of the chemiluminescent protein in the encapsulated protoplasts using a luminescence detection system. As described above, the protoplasts can be manually or automatically placed on a tissue culture system, for example, on an agar gel plate containing a plant growth medium, or on a microwell plate containing a gel and a liquid containing a plant growth promoting medium, in the form of a single encapsulated protoplast.
[0056] In an embodiment, a callus induction medium can be used. For example, media such as Gamborg B5, Murashige Skoog, etc. may be used. As will be understood by those skilled in the art, various salts, vitamins, auxins, cytokinins, and / or other hormones that promote the growth of a single protoplast into callus may be included in the medium or the tissue culture plate.
[0057] In an embodiment, callus that has reached a predetermined size can be transferred to a plant growth medium having different auxin and cytokinin ratios to induce shoot formation. In an embodiment, the callus that has undergone shoot formation can be further transferred to another plant growth medium to induce root formation. Preferably, all manipulations of the callus can be performed under aseptic conditions.
[0058] The small seedlings obtained from the above process can be used in conventional micropropagation techniques.
[0059] According to embodiments of the present invention, methods and apparatuses can be provided for efficiently manipulating and recovering plants or propagable plant materials (e.g., seeds) that contain a desired non-wild-type gene expression pattern or trait. The ability to utilize rapid phenotypic analysis of single cells in a novel library and the high recovery efficiency of whole plants afforded by the present invention are particularly advantageous in the context of plant genetic engineering.
[0060] The target plant for which the analysis using the method of the present invention is shown or desired is referred to as a "plant" or "plant target". The target species and genera of plants for which the present invention is assumed to be particularly useful include, but are not limited to, Solanum plants (e.g., S. lycopersicum, S. tuberosum, S. melongena, S. muricatum, S. betaceum), Brassica plants (e.g., B. oleracea, B. napobrassica, B. napus, B. cretica, B. rupestris and B. rapa), Capsicum plants (e.g., C. annum, C. baccatum, C. chinense, C. frutescens, C. pubescens), Lupinus plants (e.g., L. angustifolius, L. albus, L. mutabilis and L. luteus), Phaseolus plants (e.g., P. acutifolius, P. coccineus, P. lunatus, P. vulgaris and P. dumosus), Vigna species (e.g., V. aconitifolia, V. angularis, V. mungo, V. radiata, V. subterranea, and V. unguiculata), Vicia faba, Cotyledonary bean, Pea, Lathyrus plants (e.g., L. sativus and L. tuberosus), Lens plants (e.g., L. culinaris and L. esculenta), Soybean, Psophocarpus, Cajanus cajan, Arachis hypogaea, Latuca plants (e.g., L. sativa, L. serriola, L. saligna, L. virosa, and L. tatarica), Asparagus officinalis, Apium graveolens, Allium plants (e.g., A. cepa, A. oschaninii, A. ampeloprasum, A. wakegi, A. porrum, A. sativum and A. schoenoprasum), Beta vulgaris, Chicory, Globe artichoke, Eruca plants (e.g., E. vesicaria and E. sativa), Cucurbita plants (e.g., C. argyosperma, C. digitata, C. pepo, C. moschata, C. ecuadorensis, C. ficifolia, C. foetidissima, C. galeottii, C. lundelliana, C.maxima, C. moshata, C. pedatifolia, C. radicans), Spinacia oleracea, Crambe abyssinica, Cucumis species (e.g., C. sativus, C. melo, C. hystrix, C. picrolobus and C. anguria), Olea europaea, Daucus carota, Ipomoea batatas, Ipomoea eriocarpa, Manihot esculenta, Zingiber officinale, Armoracia rusticana, Helianthus species (e.g., H. annuus and H. tuberosus), Cannabis species (e.g., C. sativa and C. indica), Pastinaca sativa, Raphanus sativus, Curcuma longa, Dioscorea species (e.g., D. rotundata, D. alata, D. polystachya, D. bulbifera, D. esculenta, D. dumetorum, D. trifida and D. cayennensis), Piper species (e.g., P. aduncum, P. guineense, and P. nigrum), Zea species (e.g., Z. mays and Z. diploperennis), Hordeum species (e.g., H. vulgare, H. pusillum, H. murinum, H. marinum, H. jubatum and H. intercedens), Gossypium species (e.g., G. hirsutum, G. barbadense, G. arboreum and G. herbaceum), Triticum species (e.g., T. aestivum and T. timopheevii), Vitis vinifera, Prunus species (e.g., P. avium, P. armeniaca, P. cerasifera, P. cerasus, P. domestica, P. persica, and P. dulcis), Malus domestica, Pyrus species (e.g., P. communis, P. cordata, and P. pyrifolia), Fragaria vesca and Fragaria × ananassa, Rubus idaeus, Saccharum officinarum, Sorghum bicolor, Musa balbisiana and Musa x paradisiaca, Oryza sativa, Nicotiana tabacum, Arabidopsis thaliana, Citrus species (e.g., C. × aurantiifolia, C. × aurantium, C. × latifolia, C. × limon, C. × limonia, C. × paradise, C. × sinensis and C.× Tangierina), Populus plants (e.g., P. tremula, P. balsamifera, and P. tomentosa), Tulipa gesneriana, Medicago sativa, Abies balsamea, Avena orientalis, Bromus mango, Chrysanthemum, Costmary, Dianthus caryophyllus, Eucalyptus plants (e.g., E. leucoxylon, E. maculata, E. polybractea, E. sargentii), Impatiens biflora, Linum usitatissimum, Lycopersicon esculentum, Mangifera indica, Nelumbo plants (e.g., N. nucifera and N. pentapatala), Gramineae plants, Secale cereale, Tagetes erecta, and Tagetes minuta are included. Plants and plant cells of any of the aforementioned species having a modified sequence contained within the genome, particularly within one or more expression control sequences, can be generated by the methods described herein.
[0061] Novel gene expression control sequences, such as those identified by the methods described herein, can be used in techniques involving targeted gene editing in plant cells or plants. Those skilled in the art can utilize various approaches to targeted gene editing, including techniques that rely on sequence-guided endonucleases such as CRISPR / Cas-based genome editing systems. Accordingly, embodiments of the present invention can provide methods for generating genome-edited plants or plant propagation materials (e.g., seeds). CRISPR / Cas-based genome editing systems can target specific nucleic acid sequences. For example, the guide RNA of CRISPR / Cas is designed to bind to a nucleic acid molecule such that the Cas endonuclease can recognize a protospacer adjacent motif (PAM) sequence within the nucleic acid molecule and cleave (or nick) the nucleic acid molecule. As is known to those skilled in the art, the guide RNA (gRNA) in the CRISPR / Cas genome editing system targets specific positions within the genome of the plant of interest that are adjacent or proximal to an appropriate nucleotide protospacer adjacent motif (PAM) sequence. The gRNA can be used to target specific positions within the plant genome that are contained within or near the wild-type expression control elements. Cleavage of sequences within or near the gene expression control elements may enable disruption of endogenous sequences or insertion of new gene control sequences into the gene. Suitable novel gene control sequences will include one or more of the verified non-wild-type sequence designs identified by the present method. Highly specific targeting enables maximizing phenotypic effects while minimizing genome changes by CRISPR / Cas-based genome editing technologies. Thus, such gene editing technologies are considered to hold great promise for plant genome engineering due to their simplicity and efficiency.
[0062] In certain embodiments of the present invention, a process is provided for utilizing the novel array designs identified by the methods of the present invention described herein for the modification of plant cells. The novel array designs are used to inform the selection of target genes or genomic regions, and then a guide RNA (gRNA) is designed to direct a Cas protein (e.g., Cas9 or Cas12a) to the target site, and the construction of a CRISPR / Cas system comprising the Cas protein and the gRNA, and the delivery of the CRISPR / Cas system into plant cells are performed via established methods (such as Agrobacterium-mediated transformation). DNA cleavage occurs at the desired genomic location by the Cas protein, and then the DNA repair machinery of the plant cell is activated, and subsequent screening and selection of the edited plant cells or tissues are performed, for example, via non-homologous end joining (NHEJ) and homology-directed repair (HDR). Thus, this method enables precise modifications that introduce new array designs into the plant genome, for example, by disrupting target genes via HDR or NHEJ. The edited plant cells or tissues can be regenerated into whole plants, enabling further characterization and evaluation of the genomic modification and the resulting phenotypic effects. The disclosed CRISPR / Cas genome editing method holds great promise for applications in plant research, crop improvement, disease resistance enhancement, and the development of new plant varieties.
[0063] In certain embodiments of the present invention, a system is provided comprising at least one processor configured to operate a DNA sequence model based on a convolutional neural network and a transformer-based architecture. This model is trained to provide a predicted value representing the level of gene expression in numerical output format from DNA sequence information alone. In one embodiment, the input gene is defined as a sequence of approximately 3000 bp centered on the TSS (transcription start site) of the gene. The predicted value is represented as a predicted RNA abundance count, which is a standard way of representing gene expression, and is determined by measuring the amount of RNA molecules corresponding to an individual gene at a specific time point and within a specific tissue. The abundance of RNA is typically obtained experimentally by performing an RNA-seq assay. In this way, RNA is extracted from plant tissue and a sequencing methodology is performed to count the number of RNA transcripts bound to the genes present. The numerical values are typically represented in units such as FPKM (fragments per kilobase of transcript per million reads) or TPM (transcripts per million). Using 80% of the genes of a particular species as input genes for model training allows determination of the RNA-seq count of FPKM or TPM under specific conditions. For example, the RNA-seq count can be determined for a given plant for a specific tissue type such as a leaf, or for an environmental stress such as drought, and according to the method of the present invention, it is possible to train a model to predict the RNA-seq count based only on the input gene sequence, i.e., there is no need to measure the RNA-seq count experimentally. This surprisingly powerful ability to predict gene expression output from only the query DNA sequence enables the generation of new expression control sequences by iterative modification, etc., and the predicted value can be changed positively (upregulating gene expression) or negatively (downregulating gene expression).
[0064] Embodiments of the present invention enable, for example, in silico DNA sequence generation / promoter mutagenesis and subsequent model prediction (setting of predicted values) by assigning RNA-seq counts to natural or synthetic query sequences. This approach eliminates the need to verify each round of sequence changes in the laboratory and allows for the rapid regulation and / or improvement of gene expression in plants by using an iterative loop to predict gene expression of promoters and gene sequences. In practice, sequence designs that have been generated and passed multiple levels of evaluation based on predicted values from the model (such as RNA-seq counts) can be further verified in the wet lab to further characterize the behavior of plant cells and whole plants. The data generated by wet lab verification can also be used to provide additional information to the model.
[0065] Here, the present invention is further illustrated by the following non-limiting examples.
Example
[0066] Example 1 - In silico - assisted design of the modified PPO2 (PGSC0003DMG400018916) promoter in potato (cultivar Georgina) Potato bruising is a common agricultural problem, causing significant losses every year. One of the main genes affecting potato bruising is polyphenol oxidase 2 (PPO2). Controlling the expression of PPO2 in potatoes may result in bruise-free potatoes.
[0067] The objective of this example is to evolve the regulatory region of the PPO2 gene in potato (cultivar Georgina) to downregulate the expression of PPO2 in potatoes and minimize the discoloration associated with potato bruising.
[0068] Start of data input: The whole genome sequence of potato (vr.Georgina, HiFi PacBio+Illumina PE 150 reads).
[0069] A 5,000 - base pair (bp) sequence region upstream of the 5’UTR of all annotated coding - sequence genes known in the genome of jagaimo (vr.Georgina) was selected.
[0070] Differential expression analysis of the transcriptome of jagaimo (vr.Georgina) after 36 - hour exposure to bruise injury and the non - bruised transcriptome (0 hours).
[0071] Results: Analysis of the jagaimo transcriptome identified genes that were up - regulated in response to tuber bruising (see Figure 1). Among these, the polyphenol oxidase 2 (PPO2) gene was found to be in the top 3% of the up - regulated genes. PPO2 was selected as a candidate for further analysis.
[0072] 19,000,000 promoter mutants of 5,000 bp derived from the 5’UTR PPO2 promoter region were computationally mutated and screened across 917 DNA sites in the jagaimo genome. The maximum amount of changed base pairs was 150 bp, or 3% of the maximum of the selected 5,000 base pairs.
[0073] Using inference with a model trained to classify the expression of the jagaimo promoterome in tubers among 19,000,000 promoter mutants, 74,000 in - silico evolved promoter mutants were predicted to significantly down - regulate PPO2 with a classified differential expression of - 11log2 - fold change between jagaimo (vr.Georgina) exposed to 36 - hour bruise injury and non - exposed jagaimo (see Figure 2). An exemplary mutant promoter sequence called mutant_50981[SEQ ID NO:1] is shown in Figure 9 alongside the wild - type PPO2[SEQ ID NO:1].
[0074] Example 2 - In - silico assisted evolution of the modified MYB113 (AT1G66370) promoter in Arabidopsis thaliana (ecotype col - 0) Arabidopsis thaliana is a well-characterized plant science model organism and was used here to demonstrate the approach.
[0075] Anthocyanin-based purple discoloration is a change that occurs when Arabidopsis thaliana is stressed. This is controlled by the transcription factor gene MYB113 and is very tightly regulated. Under non-stress conditions, it is not expressed at physiologically relevant levels. By evolving the regulatory region of MYB113 that can drive the expression of MYB113 in the absence of stress, plants that grow normally but have purple leaves can be generated.
[0076] The purpose of this experiment was to evolve the regulatory region of the MYB113 gene of Arabidopsis thaliana (ecotype Col-0) using current in-silico assisted evolution algorithms to upregulate the expression of leaf MYB113 in the absence of stress, thereby generating plants with purple leaves.
[0077] Start of data input: The genome of Arabidopsis thaliana (ecotype Col-0, TAIR10 assembly) accessed from Ensembl Plants (see Yates et al., Nucleic Acids Research, Volume 50, Issue D1, January 7, 2022, pages D996 - D1003).
[0078] The variable base pair sequence region upstream of the 5’UTR of all annotated coding sequence genes within the Arabidopsis thaliana (ecotype Col-0) genome.
[0079] The transcriptome expression dataset of Arabidopsis thaliana (ecotype Col-0) at the RNA level in various plant organs at various developmental ages obtained from the University of Toronto (Klepikova Atlas) (Klepikova et al., (2016), The Plant J., Volume 88, Issue 6; 1058 - 1070).
[0080] Results: 9,000,000 promoter variants of the 754 bp 5’UTR MYB113 promoter region were screened by computer. The maximum amount of changed base pairs (bp) was 35 bp, or a maximum of 4.6% of the selected 754 base pairs.
[0081] This model evaluated 16,000 promoter variants out of 9,000,000 promoter variants that were predicted to have significantly upregulated MYB113 expression in Arabidopsis leaves compared to the wild-type MYB113 sequence.
[0082] To verify the model prediction, 1250 sequences (nMYB113) were arbitrarily selected considering the ease of DNA library synthesis.
[0083] A plasmid pool containing 1250 mutations in the 754 bp 5’ regulatory region of the MYB113 gene immediately upstream of the 5’ MYB113 UTR was assembled using conventional techniques known to those skilled in the art. Other components of the plasmid include the following. 1. 5’ MYB113 UTR derived from the Arabidopsis TAIR10 genome, 2. A coding gene encoding green fluorescent protein (EGFP), 3. A 4-codon variable region at the beginning of the GFP gene that functions as a unique molecular identifier (UMI), 4. 3’ MYB113 UTR derived from the Arabidopsis TAIR10 genome, 5. A terminator derived from 35S 6. A 35S expression cassette containing a 35S promoter, 5’UTR, an mCherry gene encoding a fluorescent reporter gene, 3’UTR, and a terminator.
[0084] The vector map of the plasmid is shown in Figure 3.
[0085] By observing a population of transiently transfected Arabidopsis protoplast leaf cells, it was shown that within the nMYB113 library pool, there are promoters that can drive the expression of GFP, which causes an increase in green fluorescence compared to the wild type, and there are mutants of the MYB113 promoter sequence that can drive protein expression with minimal changes (see Figure 4).
[0086] Using a transcriptional activity assay with unique molecular identifiers (UMIs), individual promoter mutants involved in increased gene expression were elucidated. Specifically, for the transcriptional activity assay, a library pool of 1250 mutations in the 754 bp 5' regulatory region of the MYB113 gene, immediately upstream of the 5' MYB113 UTR, was cloned into a uniquely partner-supplied SuRE plasmid vector (Arensbergen et al., (2017) Nat Biotechnol. February; 35(2): 145 - 153).
[0087] Transcriptional activity assays based on the use of SuRE plasmids with the nMYB113 promoter element pool quantified that a portion of the in - silico - designed library can up - regulate transcription. In combination with UMI barcodes, 23 non - wild - type promoter mutants with higher transcriptional activity than wild - type MYB113 in protoplast cells were identified. This is an important validation before full - plant physiological testing. It also surprisingly shows the advantage of robust high - throughput biological validation of AI - designed libraries, as only a small portion of the in - silico - designed library was executed as expected in plant cells (see Figures 5 and 6).
[0088] Figure 7 shows the expression distribution of one of the designed nMYB113 elements called OP625_short_225. The distribution was performed using multiple different UMIs. The x - axis shows the log2 fold - change in expression on an axis normalized based on the WT MYB113 promoter distribution. The y - axis shows the density of the various UMIs for each element.
[0089] Individual designed promoter elements can be further verified when taken out of the context of a library pool and tested at low throughput by fluorescence microscopy. One such element in this experiment, called A2.1DEL, was observed to drive protein expression and generate a strong fluorescence signal in protoplast cells transiently transfected in the absence of stress (see Figure 8(b)). For comparison, GFP expression driven by the wild-type MYB113 promoter was negligible fluorescence (Figure 8(a)). An exemplary mutant promoter sequence called OP625_225 [SEQ ID NO:3] is shown in Figure 10 alongside the wild-type MYB113 [SEQ ID NO:4].
[0090] Therefore, following high-throughput and low-throughput confirmation of increased gene expression in Arabidopsis protoplasts, a small subset of the nMYB113 promoter library can be further selected to verify changes in plant physiology in whole plants implemented by either stable transformation or genome editing.
[0091] Materials and Methods: Protoplast generation. Arabidopsis protoplasts were obtained using the protocol described in Yoo, Cho, and Sheen, (2007) (see above) and washed twice with buffer to remove free calcium ions. The washed protoplasts were then resuspended in the first solution to a cell concentration of 200 cells / μl. The first solution contained 100 mM CaCl2, 100 mM EDTA, 2% w / v sodium alginate, and 0.5 M mannitol, sterilized by autoclave.
[0092] Transformation: Isolated Arabidopsis protoplasts were transfected with either the nMYB113 promoter library or a control plasmid using the PEG-mediated transfection method described by Yoo, Cho, and Sheen (2007). The control plasmid carried either GFP or mCHERRY fluorescent gene under the 35s expression cassette. The nMYB113 promoter library was characterized by the presence of MYB113 promoter mutants upstream of the GFP gene and the presence of the 35s promoter upstream of the mCHERRY gene used as an internal control. The transfected protoplasts were incubated overnight at room temperature in a buffer containing sodium chloride (154 mM), calcium chloride (135 mM), potassium chloride (5 mM), glucose (5 mM), and MES (1.5 mM) to ensure the expression of the fluorescent protein downstream of the promoter.
[0093] FACS: Fluorescence-activated cell sorting (FACS) was used to sort the transfected cells based on the fluorescence properties and measurements using a method previously described by Gronlund et al. in 2012. Control protoplasts fluoresced due to the expression of GFP (green) or mCHERRY (red) downstream of the 35s promoter. On the other hand, protoplasts carrying the nMYB113 library were sorted based on GFP expression driven by the activity of these novel promoters and mCHERRY expression driven by the 35s promoter used as an internal control. Sorted cells containing the active MYB113 promoter were collected for DNA / RNA extraction.
[0094] DNA extraction and PCR for array determination: DNA was extracted from the sorted protoplasts using the Zymo Quick-DNA Plant / Seed Miniprep Kit (https: / / files.zymoresearch.com / protocols / _d6020_quick-dna_plant-seed_miniprep_kit.pdf). Next, this DNA was used as a template to amplify the MYB113 promoter variants present in the sorted cells and sent for next-generation sequencing (NGS) for identification.
[0095] RNA extraction, cDNA synthesis and PCR for array determination: Alternatively, total RNA was extracted from the sorted protoplasts using the Sigma Spectrum Plant Total RNA Extraction Kit (https: / / www.sigmaaldrich.com / deepweb / assets / sigmaaldrich / product / documents / 885 / 605 / strn10bul.pdf). The extracted RNA was used as a template for cDNA synthesis using the NEB ProtoScript® First Strand cDNA Synthesis Kit (https: / / international.neb.com / protocols / 0001 / 01 / 01 / first-strand-cdna-synthesis-e6300). The MYB113 promoter variants present in the sorted cells were amplified from the purified cDNA and sent for NGS for identification.
[0096] Transcriptional activity assay: In the transcriptional activity assay, the nMYB113 library was cloned into the SuRE plasmid using the unique molecular identifiers (UMIs) of each promoter variant. Next, this plasmid library was transfected into Arabidopsis protoplasts using the method described above (Yoo et al., 2007). Next, total RNA was extracted from the transfected protoplasts using the Sigma Spectrum Plant Total RNA Kit and sent to the provider for further transcriptional analysis.
[0097] Example 3 - The in-silico model based on convolutional neural network and transformer architecture can predict RNA-seq counts in FPKM or TPM for multiple genes of specific plant tissues and environmental conditions. The machine learning (ML) model (CRE.AI.TIVE™ v3.5) is trained to predict RNA-seq count values of genes under various conditions based on a dataset composed of multiple different genomes of plant species. In the accuracy metric, the variation of biological replicates is considered. For example, in the case of gene A sampled in tissue A (e.g., seeds), the RNA-seq count varies slightly among individual seeds. When considering the accuracy of the model's prediction, it is desirable to check the number of predictions made by the model that cannot be distinguished from natural biological replicate variations. The gene inputs used to create the model predictions are held as a test set separate from the model's training, and the model's weights were not affected by the test set.
[0098] Wheat Genome input: DNA input from the annotated reference genome Triticum aestivum (IWSC.v51) - Ensembl Plants Transcriptome input: RNA-seq dataset (within FPKM) from the Plant Public RNA-seq Database (https: / / doi.org / 10.1111 / pbi.13798) 25 RNA-seq datasets 11,311 genes in the natural variation dataset 2,273,511 data points in the natural variation dataset 90,487 genes in the training set 11,311 genes in the test set A total of 2,544,950 data points (training and validation)
[0099] Results: Natural variation in gene expression and model-predicted gene expression are shown in Figure 11. In 25 RNA-seq experiments, 11,311 genes were tested for both natural variation and model prediction ability. Figure 11(a) shows the natural variation of biological replicates, and the RNA-seq counts (FPKM) measurements of 94.71% of the genes in biological replicates are within + / - 5 FPKM of the mean value. Figure 11(b) shows the model that predicted the gene expression of genes in the test set, and 83.38% of the genes had predicted RNA-seq count values within + / - 5 FPKM of the mean experimental value.
[0100] Natural variation in gene expression and model-predicted gene expression are shown in Figure 12. In one RNA-seq experiment in root tissue, 11,311 genes were tested for both natural variation and model prediction ability. Figure 12(a) shows the average gene expression of root tissue for 11,311 genes (3 replicates per gene) plotted from experimental RNA-seq data. The x-axis shows the individual gene expression values of the measured genes and their minimum / maximum replicate values in order of gene expression intensity, showing the lowest to highest expression levels from left to right. The y-axis shows the strength of gene expression in FPKM. Figure 12(b) shows the predictions made by the trained model plotted for each individual gene. The dark gray dots are within 5 FPKM of the mean value observed in the experiment. The light gray dots are outside the 5 FPKM boundary of the mean value observed experimentally. Thus, it is clear that the model's predictions are in close agreement with the in vivo expression data.
[0101] Maize Genome input: DNA input from annotated reference genome B73 - B73.4 - Ensembl Plants Transcriptome input: RNA-seq dataset (within FPKM) from the Plant Public RNA-seq Database (https: / / doi.org / 10.1111 / pbi.13798) 116 RNA-seq datasets 4,554 genes in the natural variation dataset 1,962,774 data points within the natural variation dataset 36,435 genes within the training set 4,554 genes within the test set A total of 4,754,724 data points (training and validation)
[0102] Results: The natural variation of gene expression and the model-predicted gene expression were shown in Figure 13. In 116 RNA-seq experiments, 4,554 genes were tested for both natural variation and model prediction ability. Figure 13(a) shows the natural variation of biological replicates, and the RNA-seq count (FPKM) measurements of 91.9% of the genes in biological replicates are within + / - 5 FPKM of the mean value. Figure 13(b) shows the model that predicted the gene expression of genes in the test set, and 64.95% of the genes had predicted RNA-seq count values within + / - 5 FPKM of the mean experimental value.
[0103] Soybean Genomic input: DNA input from the annotated reference genome Williams82 (Wm82.a2.v1) - Genome assembly from Soybase, annotation from Ensembl Plants Transcriptome input: RNA-seq dataset (within FPKM) from the Plant Public RNA-seq Database (https: / / doi.org / 10.1111 / pbi.13798) 152 RNA-seq datasets 5,559 genes within the natural variation dataset 2,295,867 data points within the natural variation dataset 44,835 genes within the training set 5,559 genes within the test set A total of 7,659,888 data points (training and validation)
[0104] Results: Natural variation in gene expression and model-predicted gene expression are shown in Figure 14. In 152 RNA-seq experiments, 5,559 genes were tested for both natural variation and model prediction ability. Figure 14(a) shows the natural variation of biological replicates, and the RNA-seq count (FPKM) measurements of biological replicate for 96.48% of the genes are within + / - 5 FPKM of the mean value. Figure 14(b) shows that the model predicted the gene expression of genes in the test set, and 68.64% of the genes had predicted RNA-seq count values within + / - 5 FPKM of the mean experimental value.
[0105] Conclusion This result demonstrates a robust gene expression prediction model that can generate highly predictive RNA-seq count data purely from gene sequence data. This model has demonstrated this predictive ability across various plant species, various tissues, and developmental stages.
[0106] Although specific embodiments of the present invention have been disclosed in detail herein, this is for illustrative purposes only and is done by way of example. The foregoing embodiments are not intended to limit the scope of the following appended claims. The selection of the nucleic acid starting material, the clone of interest, or the type of library to use would be considered routine matters for those skilled in the art with knowledge of the embodiments described herein. The inventors believe that various substitutions, changes, and modifications can be made to the present invention without departing from the spirit and scope of the invention as defined by the claims.
[0107] [Non-Patent Document 1] Yang et al. (2021) Plant Biotech. J. 19: 1364 - 1369
Claims
1. A method for generating a nucleic acid library comprising a plurality of expression regulatory sequences configured to regulate the expression of coding sequences operably linked in plant cells, wherein the method is: The method involves performing in silico analysis of a plant genome, wherein the in silico analysis includes the identification of multiple sequences that collectively define a promoterome for the plant, and each sequence within the promoterome consists of an expression regulatory region extending from the start codon of the open reading frame to approximately 100 kilobases of 5' and / or 3' of the start codon. To obtain a transcriptome in the form of mRNA expression data for the aforementioned plant, To provide predicted expression levels for each protein-coding gene contained within the transcriptome, a sequence-based modeling algorithm is used for analysis, and these values are concatenated to the sequences of the corresponding expression regulatory regions contained within the promoterome. To generate multiple non-wild-type sequence designs of expression regulatory sequences that are most likely to provide a desired expression profile to a operably concatenated coding sequence, wherein the multiple non-wild-type sequence designs are known by the predicted values, Synthesizing multiple non-wild-type expression regulatory sequences corresponding to the aforementioned multiple non-wild-type sequence designs, A method comprising generating a nucleic acid sequence library containing the aforementioned plurality of non-wild-type expression control sequences.
2. The method according to claim 1, wherein the expression control region comprises one or more sequences selected from the group consisting of promoters, enhancers, silencers, transcription factor binding sites, introns, transgenic sequences, and transposons.
3. The method according to claim 1, wherein the promotrome comprises all or part of the protein-coding region.
4. The method according to claim 1, wherein the promotrome does not include a protein-coding region.
5. The method according to claim 1, wherein each sequence within the promotrome comprises an expression control region extending from the start codon of the open reading frame to less than 100 kilobases of the 5' and / or 3' of the start codon, preferably less than 50 kilobases of the 5' and / or 3' of the start codon.
6. The method according to claim 1, wherein the analysis using an array-based modeling algorithm includes training an artificial intelligence (AI) or machine learning (ML) model.
7. The method according to claim 6, wherein the AI or ML model includes an array prediction algorithm.
8. The method according to claim 7, wherein the sequence prediction algorithm is selected from artificial neural network (ANN) algorithms such that the group consists of convolutional neural networks (CNNs), recurrent neural networks (RNNs) (including bidirectional RNNs, masked language models, and transformer networks).
9. The desired expression profile prioritizes expression regulatory sequences that provide an expression level higher than the predicted value, or prioritizes expression regulatory sequences that provide an expression level lower than the predicted value, or The aforementioned multiple non-wild-type sequence designs are constrained by the minimum change requirement. The method according to claim 1.
10. The method according to claim 1, further comprising validating one or more of the plurality of non-wild-type expression regulatory sequences in vivo.
11. The method according to claim 10, wherein the in vivo validation comprises cloning one or more of the plurality of non-wild-type expression regulatory sequences into one or more plasmids containing a reporter gene whose expression is operably linked to the corresponding non-wild-type expression regulatory sequence.
12. The method according to claim 11, wherein the in vivo validation includes performing a massively parallel reporter assay (MPRA).
13. The method according to claim 12, wherein the MPRA includes fluorescence-activated cell sorting (FACS) analysis.
14. The method according to claim 11, wherein each of the one or more plasmids containing the reporter gene further comprises a unique molecular identifier (UMI) sequence.
15. The method according to claim 14, wherein the in vivo validation includes measuring the amount of the generated reporter gene mRNA and performing a transcriptional activity assay that correlates the amount with the corresponding non-wild-type expression regulatory sequence via the UMI.
16. The aforementioned in vivo verification, Inside plant protoplasts, inside plant cells, Inside a plant, or Within a part of a plant, The method according to claim 10, as implemented in [location].
17. The method according to claim 1, wherein the plant subject includes a plant species or a plant variety.
18. The aforementioned plant species or plant variety may include plants of the Solanum genus (e.g., S. lycopersicum, S. tuberosum, S. melongena, S. muricatum, S. betaceum), plants of the Brassica genus (e.g., B. oleracea, B. napobrassica, B. napus, B. cretica, B. rubestris and B. rapa), plants of the Capsicum genus (e.g., C. annuum, C. baccatum, C. chinense, C. fruitescens, C. pubescens), and plants of the Lupinus genus (e.g., L. angu). Plants of the genus P. stifolius, L. albus, L. mutabilis and L. luteus), plants of the genus P. acutifolius, P. coccineus, P. lunatus, P. vulgaris and P. dumosus, species of the genus V. (e.g. V. aconitifolia, V. angularis, V. mung, V. radiata, V. subterranea and V. unguiculata), Bithynia faba, chickpea, pea, plants of the genus Lathyrus (e.g. L. sativus and L. tuberosus), plants of the genus Lenticularis ( For example, L. culinaris and L. esculenta), soybeans, Psophocarpus, Cajanus cayan, Arachis hypogaea, Latica plants (for example, L. sativa, L. seriola, L. salinha, L. villosa, and L. taterica), Asparagus officinalis, Apium graveolen, Allium plants (for example, A. cepa, A. oschaninii, A. ampeloprasum, A. wakegi, A. porrum, A. sativum and A. schoenoprasum), Beta vulgaris, chicory, Common dandelion, plants of the genus Eluca (e.g., E. vesicalia and E. sativa), plants of the genus Cucurbita (e.g., C. argyosperma, C. digitata, C. pepo, C. moschata, C. ecuadolensis, C. ficifolia, C. foetidissima, C. galeottii, C. lundelliana, C. maxima, C. moschata, C. pedatifolia, C. radicans), Spinacia oresea, nasturtium, plants of the genus Cucumis (e.g., C. sativus,C. melo, C. histrix, C. picrocarpas and C. angulia), Olea aeropaea, Daucus carota, Ipomoea batatas, Ipomoea aeriocarpa, Manihot esculenta, ginger, Almorasia rusticana, sunflowers (e.g., H. annuus and H. tuberosus), hemp (e.g., C. sativa and C. indica), Pastinaca sativa, radish, Curcuma longa, Dioscorea plants (e.g., D. rotundata, D. alata, D. polystachya, D. bulbifera, D. esculenta, D. dumetorum, D. trifida and D. cayennensis), plants of the genus Piper (e.g., P. aduncum, P. guineense and P. nigrum), plants of the genus Zea (e.g., Z. mays and Z. diploperennis), Plants of the genus Barley (e.g., H. vulgare, H. pusillum, H. murinum, H. marinum, H. jubatum and H. intercedens), plants of the genus Gossipium (e.g., G. hirsutum, G. barbadense, G. arboreum and G. herbaceum), plants of the genus Wheat (e.g., T. aestivum and T. timopheevii), Viti Prunus vinifera, Prunus genus plants (e.g., P. avium, P. armeniaca, P. cerasifera, P. cerasus, P. domestica, P. persica, and P. dulcis), Mars domestica, Pilas genus plants (e.g., P. communis, P. cordata, and P. pyrifolia), Fragaria vesca and Fragaria × ananassa, Rubus Idaeus, Saccharum officinarum, Sorghum saccharum, Musa balbisiana and Musa x paradisiaca, Rice, Nicotiana tabacum, Arabidopsis thaliana, Citrus plants (e.g., C. x auranthifolia, C. x aurantium, C. x latifolia, C. x limon, C. x limonia, C. x paradise, C. x sinensis and C. x tangerina), Populus plants (e.g., P. tremula, P. balsamifera and P. tomentosa), Tulipa gesneriana, Medicago sativa,The method according to claim 17, selected from the group consisting of Abies balsamea, Avena orientalis, Bromus mango, Calendula, Costmary, Caryophyllus, Eucalyptus plants (e.g., E. leucoxylon, E. maculata, E. polybractea, E. sargentii), Impatiens biflora, Linum oxitachysimum, Lycopersicon esculentum, Mangifera indica, Nerumbo plants (e.g., N. nucifera and N. pentapatala), Grape species, Secare cereale, Tagetes erecta and Tagetes minuta.
19. A nucleic acid library comprising a plurality of non-wild-type expression control sequences configured to regulate the expression of a coding sequence operablely linked within a plant cell, wherein the plurality of non-wild-type expression control sequences are generated by the method of claim 1.