Method and system for producing regulatory elements
Synthetic regulatory elements are generated through computer-implemented methods, and nucleotide sequences are optimized using scoring functions and iterative processes, which solves the problem of low efficiency in gene expression regulation in the prior art, and achieves more refined and effective control of gene expression in transgenic cells and organisms.
Patent Information
- Application Number
- CN202380054590.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-07-01
- Filing Date
- 2023-06-28
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art has problems of inefficiency and difficulty in regulating gene expression in transgenic cells and organisms, especially in the lack of effective methods in the temporal and spatial expression of heterologous genes.
Synthetic regulatory elements are generated through computer-implemented methods, and nucleotide sequences are optimized using scoring functions and iterative processes to enhance gene expression in transgenic cells and organisms. The method includes identifying the starting sequence, scoring and optimizing the sequence based on the scoring function until a regulatory element that meets the gene expression target is generated.
More refined and effective control of gene expression in transgenic cells and organisms is achieved, and new regulatory elements are provided to avoid gene silencing and improve expression levels.
Smart Images

Figure CN119998821A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 358,085, filed on July 1, 2022. The entire disclosure of the above application is incorporated herein by reference. Technical Field
[0003] The present disclosure relates generally to methods and systems for plant biotechnology, and specifically to methods and systems for generating regulatory elements, and more specifically to methods and systems for synthesizing regulatory elements (e.g., promoters, introns and related polynucleotides, transgenic cells, and transgenic organisms). Background Art
[0004] This section provides background information related to the present disclosure which is not necessarily prior art.
[0005] Molecular biologists often generate transgenic cells and organisms by incorporating heterologous genes. Methods for incorporating isolated nucleotide sequences into expression cassettes, generating transformation vectors, and transforming various types of cells and organisms are well known. However, regulation or control of gene expression is essential for the development of transgenic cells and organisms for commercial use. For example, in transgenic plants containing heterologous genes that confer herbicide tolerance, it may be beneficial to express the heterologous genes in a temporal and spatial manner, for example, corresponding to when the plant is exposed to the herbicide and on which parts of the plant the herbicide normally exerts certain effects. Summary of the invention
[0006] This section provides a general summary of the disclosure, and is not a comprehensive disclosure of its full scope or all of its features.
[0007] Example embodiments of the present disclosure generally relate to computer-implemented methods for generating regulatory elements (e.g., synthetic regulatory elements). In an example embodiment, such a method generally includes: identifying an input sequence associated with a regulatory element as a starting sequence by a computing device; calculating a score of the starting sequence based on a scoring function by a computing device; initializing N iterations at at least one parameter, wherein N is an integer; for each iteration in the N iterations: (i) changing at least one nucleotide in the input sequence of the iteration; (ii) calculating a score of the changed sequence based on the scoring function; (iii) advancing the changed sequence to the next iteration based on the score, probability function, and / or predefined threshold (e.g., in response to accepting a probability function and / or the calculated score of the changed sequence meets a threshold and / or is greater than the calculated score of the starting sequence and / or the iteration is less than N and advancing the changed sequence to the next iteration); and (iv) when the calculated score indicates that the input sequence is enhanced and the iteration is equal to N, identifying the changed sequence as the output sequence of the N iterations; and after the N iterations, directing the output sequence to a verification phase, thereby synthesizing a sequence defining a regulatory element.
[0008] Example embodiments of the present disclosure also generally relate to systems for generating regulatory elements (e.g., synthetic regulatory elements). In an example embodiment, such a system generally includes a computing device configured to: identify an input sequence associated with a regulatory element as a starting sequence; calculate a score for the starting sequence based on a scoring function; initialize N iterations at at least one parameter, where N is an integer; for each iteration in the N iterations: (i) change at least one nucleotide in the input sequence of the iteration; (ii) calculate a score for the changed sequence based on a scoring function; (iii) advance the changed sequence to the next iteration based on a score, a probability function, and / or a predefined threshold (e.g., in response to accepting a probability function and / or the calculated score of the changed sequence meets a threshold and / or is greater than the calculated score of the starting sequence and / or the iteration is less than N and the changed sequence is advanced to the next iteration); and (iv) when the calculated score indicates that the input sequence is enhanced and the iteration is equal to N, identify the changed sequence as the output sequence of the N iterations; and after the N iterations, guide the output sequence to a verification phase, thereby synthesizing a sequence defining a regulatory element.
[0009] Example embodiments of the present disclosure also generally relate to a non-transitory computer-readable storage medium comprising computer-executable instructions that, when executed by at least one processor, cause the at least one processor to generate a regulatory element (eg, a synthetic regulatory element). In an example embodiment, such a non-transitory computer-readable storage medium includes instructions that, when executed by at least one processor, cause the at least one processor to: identify an input sequence associated with a regulatory element as a starting sequence; calculate a score for the starting sequence based on a scoring function; initialize N iterations at at least one parameter, where N is an integer; for each of the N iterations: (i) change at least one nucleotide in the input sequence of the iteration; (ii) calculate a score for the changed sequence based on the scoring function; (iii) advance the changed sequence to the next iteration based on the score, probability function, and / or a predefined threshold (e.g., in response to accepting the probability function and / or the calculated score of the changed sequence meets the threshold and / or is greater than the calculated score of the starting sequence and / or the iteration is less than N, advance the changed sequence to the next iteration); and (iv) when the calculated score indicates that the input sequence is enhanced and the iteration is equal to N, identify the changed sequence as the output sequence of the N iterations; and after the N iterations, direct the output sequence to a validation phase to synthesize a sequence defining a regulatory element.
[0010] Further scope of applicability will become apparent from the description provided herein.The description and specific examples in this summary are for purposes of illustration only and are not intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings described herein are for illustrative purposes only of selected embodiments and not all possible implementations, and are not intended to limit the scope of the present disclosure.
[0012] Figure 1 An example system of the present disclosure is shown, which is suitable for generating synthetic regulatory elements based on probabilities associated with gene expression;
[0013] Figure 2 shows an example feature map related to feature extraction, which can be combined with Figure 1 system to achieve this;
[0014] Figure 3 is available in Figure 1 A block diagram of an example computing device used in the system; and
[0015] Figure 4 Shows that Figure 1 The system is combined with an example method for generating one or more synthetic regulatory elements.
[0016] Corresponding reference numerals indicate corresponding parts throughout the several views of the drawings. DETAILED DESCRIPTION
[0017] Example embodiments will now be described more fully with reference to the accompanying drawings.The description and specific examples included herein are intended for purposes of illustration only and are not intended to limit the scope of the present disclosure.
[0018] The ability to control transgene expression through certain regulatory elements is limited. Regulatory elements generally include naturally occurring regulatory elements and synthetic regulatory elements made through complex biological experiments, editing and / or expression testing.
[0019] Uniquely, the systems and methods herein allow for the definition of specific, novel or synthetic regulatory elements and for use in enhancing the expression of neutral traits such as transgenic cells, organisms (including viruses and viral vectors), and polynucleotides. Synthetic regulatory sequences are defined as meeting gene expression goals, including the ability to stack multiple heterologous genes (e.g., "gene stacking" or "stacking") for expression in a single cell while avoiding gene silencing or reduced expression levels. In this regard, the systems and methods herein allow for a biological understanding of certain factors involved in specific gene expression patterns and rely on analysis of genomic and / or phenotypic data. In this manner, various advantages are provided, including, for example, but not limited to, (1) providing a source of unique synthetic regulatory elements; (2) providing expression patterns, regulation, and properties that are not available from naturally occurring regulatory elements; (3) limiting and / or alleviating gene silencing problems; and / or (4) providing additional compact synthetic regulatory sequences, etc.
[0020] This application includes subject matter that may be related to subject matter contained in applicant's following patent applications: U.S. Provisional Patent Application No. 61 / 529,001, filed on August 30, 2011; U.S. Provisional Patent Application No. 61 / 535,109, filed on September 15, 2011; U.S. Provisional Patent Application No. 61 / 535,117, filed on September 15, 2011; U.S. Patent Application No. 13 / 599,254, filed on August 30, 2012; U.S. Patent Application No. 13 / 599,255, filed on August 30, 2012; U.S. Patent Application No. 15 / 408,402, filed on January 17, 2017; and U.S. Patent Application No. 17 / 165,734, filed on February 2, 2021. The entire disclosure of each of the above applications is incorporated herein by reference. In incorporating these applications, Applicant does not waive the confidentiality provisions of 35 U.S.C. §122.
[0021] That being said, Figure 1An example system 100 for generating synthetic regulatory elements is shown in which one or more aspects of the present disclosure may be implemented. Although in the described embodiments, the various parts of the system 100 are presented in one arrangement, other embodiments may include the same or different parts arranged in other ways, depending on, for example, the accessibility of specific data, the distribution of operations associated with the present disclosure, etc.
[0022] like Figure 1 As shown, system 100 generally includes a computing device 102 and a database 104, wherein database 104 is communicatively coupled to computing device 102. It should be understood that database 104 may be separate (physically or logically) from computing device 102, such as Figure 1 As shown, or in other embodiments, the database is included in whole or in part in the computing device 102.
[0023] The computing device 102 is configured to start with one or more start sequences 106, such as Figure 1 As shown.Computing device 102 is configured to evaluate starting sequence 106 based on, for example, a scoring function, and then changes starting sequence 106 (as described in more detail below), to generate one or more output sequences 108.Computing device 102 is configured to evaluate the sequence (or output sequence 108) of change subsequently, and continues to change and evaluate the sequence.Finally, computing device 102 is configured to identify one or more output sequences 108 based on the relative evaluation of starting / changing sequence.Output sequence 108 represents synthetic regulatory elements, which may be applicable to, for example, plants, animals, algae, fungi, bacteria or viruses, etc.Once output sequence 108 is identified, the sequence 108 of synthetic regulatory elements is exposed to the verification stage 110 of system 100.In this embodiment, verification stage 110 includes, for example, converting the sequence 108 representing the synthetic regulatory elements into plasmids, and converting the gene of interest in the plasmids, to utilize the expression regulation of the synthetic regulatory elements generated, etc.
[0024] As used herein, the term "regulatory element" refers to a nucleotide sequence that participates in controlling the expression of genes in an organism of interest. Genetic regulatory elements include, for example, promoters, leader sequences (also referred to as 5'UTRs), enhancers, introns, transcription termination regions (or 3'UTRs), polyadenylation signals and chromatin control elements, or other sequences that affect RNA transcription, mRNA processing, RNA turnover or abundance, or RNA translation, etc. The regulatory element may include a 5'-untranslated region (5'UTR) or a portion thereof, a 3'-untranslated region (3'UTR) or a portion thereof, or an intron sequence, etc. It should be recognized that the regulatory element may include one or more additional regulatory elements, such as enhancers, etc. It should be further recognized that the regulatory element may act in concert with other regulatory elements to control the regulation of an operably connected gene of interest. In addition, it is recognized that enhancers may sometimes be separated from the transcription region of a gene of interest by 1, 2, 3 or more kilobase pairs of DNA.
[0025] In this example embodiment, database 104 includes hundreds, thousands, tens of thousands or hundreds of thousands or more or less genes of a variety of different organisms. Database 104 includes the genetic sequence of the gene. In addition, database 104 identifies the specific regulatory elements of the control gene expression contained in the gene. For each gene in database 104, the database may include expression data, which indicates that the regulatory elements contained in the gene, for example, effectively express the gene in the organism. Therefore, subsequently, database 104 can separate genes and / or regulatory elements according to the expression of the genes and / or regulatory elements.
[0026] In this way, for a particular organism and a particular gene to be expressed, database 104 may include a set of known regulatory elements that have one or more selected gene expression characteristics, and a set of known regulatory elements that do not have one or more selected gene expression characteristics. These sets of known regulatory elements may be understood as training sets for the purposes described herein, etc. A training set of regulatory elements may also include one or more species or genera (or virus families).
[0027] It should be understood that a group of regulatory elements with one or more selected gene expression characteristics may include all known sequences (as described above) from one or more selected species or genera (or viroides), and it is known that it exhibits one or more selected characteristics (or includes features associated with one or more gene expression characteristics). The computing device 102 herein may be configured to be carried out based on a regulatory element group or a subgroup of these sequences. The regulatory element group may include at least about 10 regulatory elements until about 10,000 or more (but not limited thereto). Preferably, in one or more example embodiments, the regulatory element group includes about 25 to about 300 elements. In certain embodiments, the regulatory element group with one or more selected gene expression characteristics may include at least about 25 elements, at least about 30 elements, at least about 35 elements, or at least about 40 elements, or at least about 100 elements. In other embodiments, the computing device 102 may be configured to operate at least about 300, at least about 350, or at least about 400 such regulatory elements. These existing sequence groups can be obtained from various publicly available genomes.
[0028] It is further recognized that the number of genes will vary depending on many factors including, for example, the choice of target organism, genetic regulatory elements, and word length or oligomer window length. In general, a sufficient number of sequences should be used to provide adequate statistical power.
[0029] In addition to genes and / or regulatory elements, database 104 may also include other data related to genes and / or regulatory elements, such as characteristics of genes and / or regulatory elements based on prior knowledge of genes and / or regulatory elements (e.g., GC content, deleterious motifs, etc.), etc. Other data is provided in more detail in the description of computing device 102 below.
[0030] In this example embodiment, system 100 employs one or more feature extraction techniques to identify certain features of a sequence that are associated with a selected or desired gene expression characteristic (e.g., a list of motifs and sequence features for templating within (or outside) a target element, sequence features associated with a target of interest (or a target element of interest), etc.).
[0031] In one example feature extraction technique, computing device 102 may be configured to determine, for a set of regulatory elements, values for a list of features of regulatory elements from database 104. The features may include, for example, different base pair patterns (e.g., GC content, etc.), which may be position-dependent or position-independent; k-mer frequencies (e.g., the frequency with which a given oligomer window (e.g., word) appears in a nucleotide or protein sequence, etc.); RNA secondary structure; codon frequency; other specialized features (e.g., codon stability coefficient, codon adaptation index, presence of intronic motifs, etc.); and the like.
[0032] Specifically, in this example, a user associated with the system 100 (e.g., a scientist, a breeder, other users, etc.) selects at least one organism (e.g., a target plant, etc.) and one or more specific gene expression characteristics (e.g., one or more selected gene expression characteristics, etc.). Examples of gene expression characteristics include, but are not limited to, expression levels in the selected organism (e.g., strong expression, weak expression, etc.), inducible expression (e.g., stress-induced expression, hormone or chemical-induced expression, etc.), temporal control of gene expression, and spatial control of gene expression. In some embodiments, the selected gene expression characteristics may include constitutive expression (e.g., high or low constitutive expression, etc.), cell-specific expression, tissue-specific expression, or organ-specific expression. In some embodiments, the selected gene expression characteristics may be expressed in response to biotic stress (e.g., fungal, bacterial and viral pathogens, insects, herbivores, etc.) and / or abiotic stress (e.g., injury, drought, cold, heat, high nutrient levels, low nutrient levels, metals, light, herbicides, pesticides, other synthetic chemicals, etc.). In other embodiments, the selected characteristics of gene expression may be developmentally controlled in one or more of plant stems, leaves, roots, and seeds. In one embodiment, the selected expression pattern can be constitutive expression, such as constitutive expression in the plant root, constitutive expression in all tissues of the root, constitutive expression in meristem, etc.
[0033] In connection therewith, computing device 102 is configured to access database 104, and in particular, to access a first set of regulatory elements (e.g., nucleotides or amino acids, etc.) of a selected organism that includes a selected gene expression characteristic, and a second set of regulatory elements of the selected organism that does not include the selected gene expression characteristic. Computing device 102 is configured to access a list of features of interest (e.g., defined by a user, etc.) in database 104, and to evaluate the first and second sets of regulatory elements for a particular feature of interest.
[0034] For example, some features may include the frequency of a specific motif with a length of k. Therefore, computing device 102 is configured to extract the frequency that all motifs or k-polymers with a length of k are contained in the regulatory element. In general, according to the number of motifs searched in the element, features may include, for example, 4k different features. Obviously, the selection of k value may be related to the performance capability of computing device 102, or it may be irrelevant. As other potential features, computing device 102 may be configured to determine the occurrence of a specific motif in the regulatory element. For example, computing device 102 may be configured to access a set of harmful motif lists (e.g., defined by a user, database 104, etc.), and determine the occurrence of harmful motifs in each regulatory element in one or more scenarios (such as a list of μRNA targeting sites). Specific motifs may have different lengths or arbitrary lengths to suit specific implementations and / or implementation schemes, etc.
[0035] Other features may include one or more sequence metrics, structural metrics and / or gene-specific metrics. For example, computing device 102 may be configured to determine the GC content of each regulatory element. Other sequence metrics may include, for example, codon stability coefficient, codon adaptation index, codon usage index, GC3 content (e.g., GC content of every three nucleotides, etc.), free energy of predicted RNA secondary structure, the number of unpaired bases in predicted RNA secondary structure, etc. It should be understood that other sequence metrics may define features in other system implementations, etc. In another example, computing device 102 may be configured to determine structural metrics, which may include metrics related to features of secondary structures of RNA transcripts that tend to simulate or approximate the secondary structures produced by sequences. And, in various examples, computing device 102 may be configured to determine gene-specific metrics, which may include frequencies of different codons, codon adaptation index (e.g., correlation metrics between codons in input sequence and their use in reference groups of highly expressed genes, etc.), codon stability coefficients, etc.
[0036] It will be appreciated that each of the above features can be targeted to specific regions within a gene and / or regulatory element (eg, near the 5' or 3' end, adjacent to the transcription start site, etc.).
[0037] Based on the above, the computing device 102 may also be configured to generate a feature matrix, which is generally defined as an n×m matrix, where n is the number of regulatory elements and m is the number of features included in the list accessed in the database 104. Therefore, the feature matrix may include one or more rows including a list of features for a given regulatory element, such as sequence GC content, sequence predicted secondary structure-based energy, frequency of a given k-mer, etc. That being said, in some implementations of the present disclosure, a feature matrix may not be required or generated.
[0038] Thus, the feature matrix (when generated) provides a certain dimension, where m>>n, such that feature reduction may be desired in various example embodiments (although such feature reduction is not required in all example embodiments of the present disclosure). In embodiments where feature reduction is desired, computing device 102 may optionally be configured to provide feature reduction through various techniques.
[0039] In one feature reduction example, the computing device 102 is configured to rely on information content-based reduction. In this regard, the computing device 102 is configured to measure the variance of each individual feature when isolated. For example, a feature with only 0 in the matrix will be considered to have a variance of 0, while a column with many different values will be considered to have a high variance for the regulatory elements contained in the matrix. Therefore, the computing device 102 is typically configured to identify high variance columns as being more likely to contain information useful to the machine learning process. Therefore, in this example, the computing device 102 may be configured to select the top k columns (features) with the largest variance.
[0040] Additionally or alternatively, in another feature reduction example, the computing device 102 is configured to rely on model-based reduction. In connection therewith, the computing device 102 is configured to rely on the ranking of features through some learning algorithm to weight certain features when defining an internal model. The computing device 102 is then configured to exploit the ranking by selecting features for a single model (e.g., random forest, lasso-based selection, etc.) or by polling multiple simple models and selecting a consistent set of features.
[0041] Regardless of the feature reduction technique (if any) employed, the computing device 102 is configured to then perform feature analysis in combination with the features included in, for example, a feature matrix, etc. Specifically, in this example, the computing device 102 is configured to select the strongest features that correlate or anti-correlate gene expression with the organism. Specifically, each feature is analyzed so that feature contributions can be explained.
[0042] An example technique employed by computing device 102 may include applying a Shapley Additive Explanation (or SHAP) graph, e.g., Figure 2 As shown in 200 of . For example, consistent with SHAP, an explanatory model is generated to intuitively provide feature contributions. Therefore, the SHAP plot provides an indicator where each feature is assigned a SHAP value representing the expected change in prediction in the prediction model when adjusting the feature. Figure 2 , the shading indicates whether a given feature is high or low, and the position of the shaded indicator toward the left or right side of the vertical axis 202 (e.g., the target value, etc.) indicates whether the feature value is positively or negatively correlated with the target value. Figure 2 In some examples, a ranking system can be used in addition to (or in lieu of) the visual analysis to assess the prevalence of different features using non-linear correlations of features with the target class (e.g., similar to the correlations described above, etc.) so as to discard unimportant features (e.g., if the determined correlations are sorted, ranked, etc.).
[0043] In view of the above, computing device 102 is configured to identify a set of motifs and / or metrics and related features to include in the scoring of potential regulatory elements (as described below). Thus, the algorithm developed by computing device 102 can effectively and accurately define specific features from existing expression data as having a specific contribution to the expression of genes in an organism. In doing so, the various features identified herein can then be considered in a given algorithm as needed (e.g., depending on the design of a given algorithm, etc.), and therefore directly considered in the relevant scores.
[0044] It should be understood that the identified motifs and / or feature groups described above are stored in the database 104 for use in the operations described below. Therefore, the computing device 102 can be configured to select the in-group motif group (e.g., Figure 1 ) and the outgroup motif group (as Figure 1 Then, the computing device 102 is configured to define an ingroup and an outgroup, wherein the ingroup motifs may include a set of motifs associated with the expression of selected genes in certain organisms, and the outgroup motifs may include all other motifs, or motifs that are particularly associated with different gene expressions (e.g., lack of gene expression, low gene expression, high gene expression, inducible gene expression, constitutive gene expression, tissue-specific gene expression, etc.), or positive effects on the sequence, or negative effects on the sequence (e.g., deleterious motifs, etc.), etc.
[0045] In another example, computing device 102 may employ a position sensitive word group (POWRS) algorithm associated with feature analysis / extraction. Specifically, POWRS can be used to identify sequence features associated with a selected or desired gene expression characteristic. For example, POWRS uses position-specific enrichment of regulatory elements near the transcription start site to increase sensitivity while also providing information about the preferred positioning of these elements.
[0046] In another example, the computing device 102 may employ one or more additional steps or processes associated with feature analysis / extraction. For example, the computing device 102 may be configured to determine the frequency of short oligomer windows of a predetermined length in a known sequence. As used herein, an "oligomer window" refers to a short nucleotide sequence. In addition, "frequency" may refer to the number of times each such oligomer window occurs; or to the fraction or percentage of all oligomer windows contained in such a count; or to the ratio of such fractions between two sets of known sequences, and thus reflect the frequency "enrichment" of oligomer windows in one set relative to another set.
[0047] In certain embodiments, when determining the position-dependent or position-independent enrichment of an oligomer window, the computing device 102 may be configured to determine the enrichment relative to a group of background elements (or "second group") that do not have (or are predicted not to have) the selected characteristic. In general, the second group of regulatory elements includes all or part of the regulatory element category in the organism. In some embodiments, the second group may include about 20,000 to about 60,000 regulatory elements, but in other embodiments, the second group may only include a subgroup of these regulatory elements from the target organism. Typically, the second group includes at least about 100 regulatory elements. Having said that, it should be understood that in certain other embodiments, a "simulated background" process can be used, so that the second group of elements can be omitted. For example, the simulated background method can be used for the design of viral promoters. In short, the simulated background scheme involves determining the position-dependent enrichment of oligomer windows in the first group of regulatory elements with gene expression characteristics (or multiple characteristics) relative to the total number of occurrences of oligomer windows in the regulatory element group.
[0048] It should be appreciated that various other techniques may be employed to associate a particular metric with a characteristic representation in a component group.
[0049] That is, as disclosed in detail below, scoring function can be realized by computing device 102, to calculate the position dependence or position-independent enrichment of each oligomer window (or " word ") of selected size in the sequence in the regulatory element group with selected gene expression characteristic (or multiple characteristics). That is, select oligomer window size (such as 4-mer, 5-mer, 6-mer, 7-mer, 8-mer, 9-mer, 10-mer, 11-mer, 12-mer, 13-mer, 14-mer, 15-mer, 16-mer, 17-mer, 18-mer, 19-mer, 20-mer, etc.), and analyze each oligomer window in the sequence (or a part of the sequence), to obtain position dependence and / or position-independent enrichment in the regulatory element group with selected characteristic. Then can determine total score, it represents the probability that sequence has selected gene expression characteristic (or multiple characteristics) in species of interest. Can adopt known algorithm to predict the possibility that nucleotide sequence has selected characteristic, for example, adopt Bayes rule in some embodiments.
[0050] In certain embodiments, computing device 102 is configured to construct genetic control elements, such as introns, that may appear more than once in a gene of interest. In such embodiments, a first group of genetic control elements may include all introns that appear at a specified position (e.g., the first or last intron in a gene, etc.), and a second group of genetic control elements may include all introns that are located outside a specified position in the genome of an organism. In one embodiment, a first group of genetic control elements includes the first intron from a highly expressed constituent gene, which appears in a 5'UTR or coding region and is located within 500 base pairs (bp) of a transcription start site (TSS), and a second group of nucleotide sequences includes all non-first introns of all genes in a target organism.
[0051] The computing device 102 is configured to compare the regulatory element group around one or more conserved sequences or "landmark" sequences to perform position-dependent analysis of enriched sequences. The conserved sequence or landmark can be, for example, a transcription start site (TSS), a TATA box, a transcription termination signal, a polyadenylation signal, a splice acceptor site, a splice donor site, or a branch site. In certain embodiments, the conserved sequence is a TSS or a TATA box. In some embodiments, the landmark sequence includes the 5' and / or 3' ends of the element, or other conserved motifs or subelements within the genetic element. However, any method for comparing sequences known in the art can be used. For example, when the gene regulatory element is an intron, the computing device 102 can be configured to compare the intron sequences on both the 5' and 3' splice sites, and copy or truncate the intermediate sequence as needed to provide an appropriate length. In addition, negative motifs (e.g., motifs excluded from the final sequence, etc.) can also be identified.
[0052] It should be understood that computing device 102 can be configured to select an oligomer window or word length for comparing sequences, wherein the oligomer window or word length includes the number of consecutive nucleotides in the oligomer window. For a given implementation, word length can be fixed. Word length can be about 4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,19,20 etc. For each word length x, there are 4x possible words, because each nucleotide position may have A, G, C or T, although not all words may be represented in the nucleotide sequence of a group of genetic control elements.
[0053] Generally, as described (e.g., as part of feature extraction, etc.), different elements in the in-group and out-group may be aligned around a given landmark (because they may have length differences and therefore need to be centered around something). Then, different features may be extracted, as described above.
[0054] As disclosed in various embodiments herein, computing device 102 is configured to calculate position dependence and / or position independence scores of multiple oligomer windows based on a scoring function, and determine the probability that the sequence started or changed has the selected characteristic based on the sum or factor of the position dependence and / or position independence scores. Position dependence enrichment of oligomer windows in the regulatory sequence group with the selected characteristic means that the oligomer sequence is enriched at the same position or defined as within ±200 nucleotides, or within ±100 nucleotides, or within ±30 nucleotides in some embodiments. In some embodiments, position dependence enrichment is limited to within ±20 nucleotides or within ±10 nucleotides.
[0055] In various embodiments, only a part of nucleotide sequence is analyzed to obtain the position-dependent enrichment of oligomer window, because the predicted importance of positioning may depend on the type of element or change in element. In other words, the different parts of the training group herein can be processed differently according to its position relative to each landmark. For example, when the synthetic regulating element is a promoter, the position-dependent enrichment of oligomer window may be less important in the region away from TSS or TATA box. Therefore, in some embodiments, the position-dependent enrichment of oligomer window can be determined in the regulating element group with selected characteristics in at least 20bp zone in TSS or TATA box upstream and / or downstream. For example, relative to TSS, the position-dependent enrichment of oligomer window in the region comprising -50 to +20, or -100 to +20, or -200 to +20, or -50 to +50, or -100 to +50, or -200 to +50 can be analyzed. In other embodiments, determine the position-dependent enrichment of at least about 50 bases in TSS or TATA box upstream, or at least about 100 bases. Other oligomer windows outside of these regions can be analyzed in a position-dependent or position-independent manner.
[0056] In some embodiments, the process maintains a certain level of sequence complexity or weighted local sequence complexity by inserting repeated base pairs, for example, so that the synthetic regulatory element approximates the sequence complexity of the regulatory element group having the desired gene expression property (or multiple properties) (including locally in some embodiments). Sequence complexity can be defined by GC or AT content, or by dinucleotide content (e.g., AA, AT, AC, AG, TT, TA, TC, TG, CC, CG, CT, CA, GG, GC, GA, and GT, etc.), or by A, T, G, and / or C scores. A separate local sequence complexity score can be determined for each fragment of a polynucleotide. The length of such a fragment can be at least 30 bp, and in some embodiments, the length is at least 50 bp, or at least 100 bp, or at least 125 bp. In such an embodiment, the computing device 102 can use an algorithm (e.g., as part of the entropy element in the score herein (e.g., Z4 (S) herein, etc.) etc.) to calculate the local sequence complexity, thereby constraining the local sequence complexity to approximate the local sequence complexity of the element with the selected property.
[0057] When modifying the sequence iteratively or non-iteratively, the computing device 102 may be configured to employ any suitable technique to modify the sequence. In some embodiments, for example, the computing device 102 may be configured to employ simulated annealing, while in other embodiments, the computing device 102 may be configured to employ other types of algorithms, including but not limited to genetic algorithms, tabu search, simplex algorithm, steepest descent, conjugate gradient, and dynamic programming.
[0058] After the signature expression and identification of the signature and / or motif, the computing device 102 is configured to generate a synthetic regulatory element in dependence thereon. First, the computing device 102 is configured to select an input sequence 106 for which a synthetic regulatory element is to be generated. A variety of different techniques may be employed to select the input sequence.
[0059] In various embodiments, synthetic regulatory elements are generated based on iterative modification of sequences, and then scored by one or more scoring functions, rather than by combining sequences from defined subsequence groups. As described herein, scoring functions can be probabilistic in nature, and can be used to design sequences similar to members of a group of naturally occurring regulatory elements included in the first group of regulatory elements (e.g., having desired one or more gene expression characteristics, etc.), but the designed sequences have limited extended sequence homology to the naturally occurring sequences. In this way, the scoring function does not require prior knowledge of functional motifs, cis elements, transcription factor binding sites, etc. Due to these features, the configurations and processes described herein are widely applicable to promoter and non-promoter regulatory elements, including, for example, introns and untranslated regions (UTRs), where little or no functional motif information is available.
[0060] In this regard, as described above, computing device 102 can be configured to obtain at least a first set of sequences of genetic control elements or parts thereof, wherein the first set of sequences are from a selected organism (e.g., a target organism), and each gene in the first set of genes is known or expected to be expressed in a desired manner in the target organism. Computing device 102 can also be configured to determine the frequency of each word of a predetermined word length for the first set of sequences. The position dependence or position-independent enrichment of each word can be determined as described herein. Computing device 102 can be further configured to define an output sequence 108 of a synthetic genetic control element or part thereof by starting from input sequence 106 and generating at least one modified sequence, and then using a scoring function to score at least one modified sequence, in pursuit of an improved score.
[0061] In this exemplary embodiment, in view of the above, the input sequence 106 can be, for example, a sequence from the first group of nucleotide sequences described above, a known sequence associated with the characteristic, or a sequence generated using the scoring function described below.
[0062] The score of sequence is derived from such scoring function as described herein, which for example shows at least in part the similarity of this sequence to the first group of regulatory elements. The score is derived from the frequency of the word in the first group of regulatory elements. Usually, the expected score is the score of the score of the nucleotide sequence of about 1%, 5% or 10% higher than the first group of regulatory elements. In some embodiments, the expected score is the score of the score of the gene expression element of about 20%, 25% or 30% higher than the first group. In other embodiments, the expected score is the score of the score of the nucleotide sequence of about 40%, 50%, 60% or more higher than the first group of nucleotide sequences. It should be understood that other scoring thresholds can be adopted in other embodiments. It should also be understood that computing device 102 can be configured to continue to generate additional related sequences and score them, until generating related sequences containing expected scores (for example, relative to scoring thresholds, etc.).
[0063] Thus, as described in further detail below, the computing device 102 may be configured to determine: (i) the frequency of each word in the first set of genetic regulatory elements; (ii) the enrichment of each word in the genetic regulatory element relative to the occurrence of each word in the second set of genetic regulatory elements or relative to the frequency of the word at all positions in the first set of genetic regulatory elements (e.g., without using the second set of genetic regulatory elements, etc.); (iii) the sequence entropy of the genetic regulatory element associated with the scoring function of the present invention; and / or (iv) the local and / or global similarity of the sequence to the regulatory elements included in the regulatory element group, etc.
[0064] Specifically, computing device 102 can be configured to compare nucleotide sequences from a first group (A) with nucleotide sequences from a second background group (B), and determine which features of the genetic regulatory elements of A may contribute to the unique expression patterns of these genes or elements. For example, the genetic regulatory element of interest may be a promoter. For example, promoters from A and B are aligned relative to their TSS and can be compared in a position-specific manner, for example, as a function of the distance from the TSS. As a variant, the sequence can be aligned around conservative elements near the TSS, such as the TATA box. Specifically, at each position, determine whether a word or oligomer window sequence (also referred to herein as a "k-polymer", such as 4 to 10 consecutive bases) is overexpressed in the gene of interest.
[0065] In this manner, computing device 102 is configured to rely on features that are present in the first group but not in the second group as indicators of expression.
[0066] In this example embodiment, computing device 102 is configured to select input sequence 106 or input control element from a group of control elements known to show the desired expression characteristics of the selected gene of the selected organism. For example, it should be understood that some existing, naturally occurring control elements with selected gene expression characteristics (or multiple characteristics) from source species or organisms can be generally understood from genome data. For example, microarray or RNA sequence analysis can be used to quantify the transcripts in the cells and tissues of interest, and expression patterns are associated with homologous genetic control elements. And, the target species can be one or more plants, and various types and species of target plants are described elsewhere herein. Genetic data from these target species can be used to prepare synthetic control elements.
[0067] However, in other embodiments (as generally noted above), computing device 102 may be configured to generate an input sequence or input regulatory elements (eg, synthetic regulatory elements) based on a set of regulatory elements having selected gene expression characteristics.
[0068] Regardless of the way in which the input sequence 106 is identified, selected or compiled, the computing device 102 is configured to score the input sequence 106 based on a scoring function (as described below), and change or modify the sequence (e.g., provide instructions for doing so, etc.). The computing device 102 is configured to then score the modified sequence, and repeat in an iterative or non-iterative manner as required. In this way, the computing device 102 is configured to define a sequence characterized by a suitable score or a statistically significant score, whereby the sequence may have one or more selected gene expression characteristics. In this context, the term "statistically significant" means that the sequence contains a position-dependent or position-independent enrichment of an oligomer window sequence found in a regulatory sequence group with one or more selected gene expression characteristics, and the enrichment level is unlikely to occur by chance. For example, a statistically significant score may have a p-value of about 0.05 or less, or a p-value of about 0.005 or less, etc.
[0069] In conjunction with the above, in one or more example embodiments, the computing device 102 is configured to generate a nucleotide sequence S that approximately maximizes the probability of the expression pattern E in the following equations (1)-(4), for example, (approximately) maximizes P(E|S). For convenience, k is used to represent the length of the short sequence (usually 4-10 bp) and the sequence itself (e.g., GCCCA, etc.). And, G represents the merge of the sequence groups A and B. For each position i relative to the TSS and each k-mer k, G k,i Include those sequences in G that contain k at position i. The k-mer at i and the k-mer at i+1 overlap each other by k-1 bases. In addition, G i Include the sequence containing position i in G (due to the different lengths of regulatory elements, some G may not include position i). Given the above, in this example, the scoring function includes P(E|k,i) as the probability that the sequence with k at position i will display expression pattern E:
[0070]
[0071] The probability P(E|S) that a sequence S gives expression pattern E can be estimated by assuming that the positional probabilities are independent and multiplying them. This procedure is similar to a naive Bayes classifier. These probabilities can be normalized by the base probability of expression pattern E and log-transformed to yield a score Z1(S) that is greater than zero if the sequence S is more likely than the average to show pattern E and less than zero if S is less likely than the average to show pattern E, as shown in example equation (5).
[0072]
[0073] where k is understood as k S,i, i.e., the k-mer at position i of sequence S. Therefore, the term within the logarithm is simply the fold enrichment of k in the gene of interest compared to the entire genome.
[0074] In various embodiments, additional terms may be desirable. First, it is understood that the oligomer window length of a k-mer may indicate the information content of the k-mer, where, for example, longer k-mers are generally more informative, but it is understood that for certain transcription factors, for example, the length of a given nucleotide feature may ideally coincide with the window length of the k-mer to provide sufficient information without additional spurious information. That is, there are generally many more possible k-mers than genes of interest, which means that ||A k,i || is rarely greater than 1, and is often zero. For example, there are 4096 possible 6-mers, and 65,536 possible 8-mers. Second, some k-mers are intrinsically uncommon in the genome, so a limited number of occurrences in A would result in an apparent high enrichment.
[0075] In conjunction with the above, the scoring function can be modified by counting the number of occurrences of k on the local oligomer window, not just at position i. The counting is done in the form of kernel density estimation using a cosine kernel with a half-width and half-height of w (w = 10 bps, or 5 bps, or 15 bps or other, etc.), as shown in example equation (6).
[0076]
[0077] Those skilled in the art will recognize that other kernels (eg, Gaussian, triangular, square, etc.) or methods (eg, standard, smoothed, or mean shifted histogram, etc.) may be used in other embodiments.
[0078] In addition, the scoring function can be modified by adding a pseudo count p to the actual observations; this is equivalent to assuming a uniform distribution as a Bayesian prior. For most embodiments disclosed herein, p is 20, or may include values of 10 to 50. Each of these modifications is provided below in the example equation (7) for Z2(S).
[0079]
[0080] In spite of the above situation, optionally, the score function can be modified to limit the contribution of the k-aggregate counts from a single gene, while still making the count on the local oligomer window smooth. This modification can be used to explain the gene containing the same k-aggregate repeatedly in a small region, which may occur in the case of (k-1) bases in the overlapping k bases of the k-aggregate and the previous k-aggregate. In these cases, only a long repeated gene is needed to make the k-aggregate as "GGGGGG" obviously enriched. Modification includes the following example equation (8).
[0081]
[0082] In equation (8), if gene “a” contains k-mer k at position j, then ||a k,j || = 1, otherwise 0. Through the above, the scoring function is modified to the scoring function Z3(S), as shown in the example equation (9).
[0083]
[0084] It will be appreciated from the above that altered sequences that provide improved scores by Z3(S) should have the potential to drive gene expression following pattern E.
[0085] However, simply improving the Z3(S) score does not guarantee that the sequence is promoter-like: all promoters may have certain common features or properties that cannot be detected by Z3(S). In fact, sequences with high Z3(S) scores are almost entirely composed of k-mers that are actually observed at significant frequencies in natural promoters. However, it has been observed that for some species (e.g., rice, etc.), sequences designed to improve the Z3(S) score may continuously and repeatedly exhibit similar motifs, resulting in unnaturally low complexity.
[0086] To vary the complexity of the sequence, the local sequence entropy may be constrained at positions along the design sequence (e.g., at each position along the sequence, etc.). The local sequence entropy may be determined by the computing device 102 using single nucleotides, dinucleotides, trinucleotides, etc. In a particular embodiment, the computing device 102 may be configured to calculate the entropy in an oligomer window of 2ω bases (2ω=128 bp) using dinucleotide composition, as shown in example equation (10).
[0087]
[0088] In equation (10), ||S n (i-ω,i+ω)|| is the number of occurrences of the dinucleotide n between positions i-ω and i+ω in sequence S. For comparison, the average local entropy H0 and its variance σ for all sequences and all positions in A can be calculated 2 H0 ( and ).
[0089] In some embodiments, the promoter is designed based on a viral promoter from the same family as 35S (Cauliflower mosaic virus family). In this case, there is no obvious outgroup (B) to compare the sequence to. In this case, a "simulated" background can be calculated, comparing the frequency of the motif at a specific position in A to its average frequency over all positions in A, and is defined as follows in example equation (11).
[0090]
[0091] As described above, Z(S) may be calculated using equation (11) above instead of Z3(S). In certain embodiments of the present disclosure, even in the presence of an obvious outgroup B, a "simulated background" approach is applied.
[0092] A score Z4(S) may then be defined in example equation (12), which imposes a penalty on S for having a local entropy that is too high or too low.
[0093]
[0094] Furthermore, those skilled in the art will recognize that other measures of sequence complexity may be substituted for entropy with similar results.
[0095] As described above, in certain embodiments, it may be desirable to include certain motifs that are common in A, rather than motifs that are particularly enriched relative to G. Empirically, this may provide the ability to avoid unnaturally low complexity, particularly in the case of introns, where some motifs are strongly enriched in a relative position-independent manner. The motif frequency score may then be defined by example equation (13).
[0096]
[0097] In equation (13), p = 1 for all work done so far. This score assumes that all 4 k The a priori likelihood of each of the possible k-mers is equal, for example, the expected frequency of any given motif at any given position is 4 -k ; Therefore, for a random sequence, Z5(S) is expected to be zero. In some cases, this may distort the designed sequence, resulting in any imbalance between A / T. G / C content is present in naturally occurring sequences. In this case, the expected frequency can be determined separately for each k-mer based on the ratio of A, C, G and T bases in the naturally occurring sequence.
[0098] In addition to the above, similarity can be further used to modify the scoring function. For example, Z6 is provided below, which includes local similarity (as S with S0 and S p convolution) and global similarity (as Hamming distance similarity with the original sequence S0) and an optional sequence group S p , which is a set of sequences that are referenced to avoid similarity between the sequences and the original sequence S0. The convolution g between w and f is defined in example equation (14).
[0099]
[0100] Referring also to the example equation (15) below, local similarity starts with the convolution g and minimizes an oligo window size K, which is equal to the number of base pairs to avoid plus 1. For example, to avoid 22 bp, then K=23.
[0101] Z6(S,S o , S P )=∑g(S,y)+max(H(S,y)) (15)
[0102] Finally, in various embodiments, the position-dependent k-mer enrichment score can be combined with the entropy constraint and the frequency score to obtain the final position-dependent scoring function Z(S), as generally indicated in example equation (16). The components are weighted by empirically determined coefficients that balance the k-mer composition with the sequence complexity (in most embodiments disclosed herein, And∈ z = 0.07, although for some embodiments where the genetic regulatory element is an intron, And∈ z =150 may be preferred).
[0103]
[0104] A promoter sequence S with a high Z(S) value is expected to confer a desired expression pattern on any gene of interest to which it is coupled.
[0105] Those skilled in the art will recognize that many different techniques may be used to generate a sequence S with a high Z(S) value. These techniques may include, but are not limited to, function optimization methods such as simulated annealing, genetic algorithms, tabu search, simplex algorithm, steepest descent, conjugate gradient, and dynamic programming. Such methods may or may not contain probabilistic, random, or stochastic elements; and may or may not involve an iterative process.
[0106] In certain embodiments of the present disclosure, the computing device 102 is configured to employ simulated annealing to iteratively change a sequence that may be improved as defined by a score of the sequence based on a scoring function.
[0107] In particular, in embodiments, system 100 comprises database 104, which may include but is not limited to motif data structure and gene distribution data structure. The motif data structure may include a motif list, which should be included in the regulating and controlling element, or may not be included in the regulating and controlling element. In this way, motifs may be considered to include motifs and exclude motifs. Similarly, gene distribution data structure may include a general list of regulating and controlling elements, which are positive examples and counterexamples. Therefore, regulating and controlling elements are intended to be more similar to positive examples, and different from counterexamples.
[0108] The computing device 102 can then be configured to access the database 104, including the motif data structure and the gene distribution data structure, and generate a configuration file for generating a synthetic regulatory element. The configuration file includes, but is not limited to, details related to the scoring of the synthetic regulatory element (e.g., weights, etc.) and the number of iterations, parameter schedules (e.g., temperature schedules, etc.) related to the generation of the synthetic regulatory element.
[0109] In this example embodiment, computing device 102 is configured to then initialize the scoring function in accordance with the configuration file, as described above (or as described below), and then enter the first iteration of the simulated annealing operation with an initial or starting sequence. Any sequence can be used as a starting sequence. For example, a member of group A of known expression can be used, wherein the motif defined in data structure 102 is included and / or excluded. It should be understood that in this embodiment, the iterative process is based on the change of required parameters (e.g., temperature in this example, etc.), which is a representative of stability in this context. In general, an annealing schedule (also referred to as a cooling schedule, also may be referred to as a temperature schedule) is provided, which progresses from high temperature to lower temperatures, thereby imposing low stability to high stability (e.g., lower temperatures may limit the number of equivalent codons, etc.) to the iterative process. The temperature schedule in this example embodiment may include, for example, {2.0, 1.0, 0.5, 0.2, 0.1, 0.01}.
[0110] Specifically, the computing device 102 is configured to iteratively modify the original sequence (or the starting sequence of the iteration) according to the temperature of the annealing schedule. In this way, the computing device 102 is configured to modify the sequence in multiple iterations, where the temperature is adopted with probability so that the sequence advances to the next iteration. For example, for each iteration of the temperature, the computing device 102 is configured to score the sequence based on the scoring function as described above and compare the score with the score of the previous iteration (or the score of the initial sequence of the first iteration). Z(S, S o , S P ) When the score of the modified sequence increases, the computing device 102 is configured to pass the modified sequence to the next iteration as the starting sequence of the next iteration.
[0111] In this example embodiment, when the score of the modified sequence is not improved (based on the score), the computing device 102 is configured to use the following example equation (17), under the given parameter (T) herein, the score of the iteration (e.g., E') is compared with the score (e.g., E) of the original coding sequence or the previous new coding sequence. The computing device 102 is configured to subsequently generate a random threshold (e.g., between 0 and 1, etc.) and compare the probability (p) of the new coding sequence with the random threshold. When the probability meets the random threshold (e.g., higher than the random threshold, etc.), the computing device 102 is configured to feed the new coding sequence as the starting coding sequence to the next iteration. When the probability does not meet the random threshold, the computing device 102 is configured to discard the new coding sequence and feed the previous coding sequence as the starting coding sequence to the next iteration.
[0112]
[0113] Although the scores obtained by the above equations are directly compared to the specific probability functions of the above probability equations (for determining which coding sequence to implement in the next iteration), it should be understood that the scores can be compared in other ways to make the same decision in other example embodiments. In addition, in some example embodiments, the threshold value can be a static threshold value.
[0114] After each iteration at a particular temperature is completed (or stopped), the computing device 102 is configured to advance to the next temperature in the annealing schedule and repeat the iterative modification and scoring of the sequence. It should be understood that in some embodiments, when doing so, the computing device 102 can be configured to use parallel processing to simultaneously modify the coding sequence and score it as part of multiple different iterations, whereby the iterations vary, for example, as defined by the annealing schedule. In this way, the computing device 102 can provide a parallel tempering algorithm (e.g., in combination with replica exchange Markov chain Monte Carlo (MCMC) sampling, etc.), thereby achieving multi-target enhancement of genes of interest (as generally described in conjunction with method 400).
[0115] Next, in the system 100, after the iteration of each temperature in the annealing schedule is completed, the computing device 102 is configured to identify one or more sequences as sequences of one or more synthetic regulatory elements based on the score as described above. The computing device 102 is configured to perform one or more checks on the output sequence 108 before advancing to verification. For example, the computing device 102 is configured to confirm that the output sequence 108 includes a threshold similarity with the input sequence 106. In this regard, the computing device 102 can confirm that the output sequence has a similarity of up to about 75% with the input sequence 106, which may be sufficient to avoid including the input sequence 106 and the output sequence 108 in the same stack based on silence. In addition to the overall similarity, the computing device 102 is also configured to confirm the local similarity threshold. For example, the computing device 102 can be configured to compare the input sequence 106 and the output sequence 108 step by step to ensure, confirm, etc. that there is no common extension of X base pairs, where X is an integer greater than 5 (e.g., X is 22bps, etc.), etc. It should be understood that other checks may be performed after the iteration, whether associated with local or global similarity, or other checks, etc.
[0116] Thereafter, the sequence-defined synthetic regulatory element is provided to a validation stage 110, which is configured to implement and / or test the sequence. The sequence is then converted to RNA by a manual or automated process, as is well known to those skilled in the art.
[0117] And, in some embodiments, in the verification stage 110, the DNA encoding RNA is transformed into plants or plant cells by conventional methods known in the art by artificial or process. For example, the transgenic (i.e., gene of interest) encoding RNA of interest can be inserted into Agrobacterium, and then Agrobacterium introduces genetic material into plant tissue. Then the transformed plant tissue is cultured in a suitable culture medium to promote the formation of roots. After bud formation, the transgenic plant is transferred to a suitable soil to verify the desired phenotype or gene of interest. In another example, the gene of interest operably connected to the synthetic regulatory element defined by the sequence can be inserted into Agrobacterium, and then Agrobacterium introduces genetic material into plant tissue. Then the transformed plant tissue is cultured in a suitable culture medium to promote the formation of roots. After bud formation, the transgenic plant is transferred to a suitable soil to verify the ability of the generated synthetic regulatory element to regulate or drive the expression of the gene of interest. Alternatively, genetic material can be introduced into plant cells by gene gun, electroporation, zinc finger nuclease (ZFN), transcription activator-like effector nuclease (TALEN), CRISPR / Cas9 system or CRISPR / Cpf1 system, etc. In other embodiments, during the validation phase 110, the genomic DNA is edited to encode an output coding sequence.
[0118] Notwithstanding the foregoing, it should be appreciated that computing device 102 may be configured in other ways to modify and score sequences in different ways, as described below.
[0119] In one example embodiment, in connection with parallel annealing, the computing device 102 is configured to initialize the input code sequence 110 into a plurality of different chains, wherein each chain is associated with a different temperature in a given annealing schedule (or randomly selected). The computing device 102 is then configured to initialize each chain in parallel for a desired number of iterations.
[0120] Specifically, for each chain and iteration thereof, the computing device 102 is configured to change the sequence as constrained or indicated by a given temperature selected for the chain. The computing device 102 is configured to then score the sequences, as described above. For this example embodiment, the computing device 102 is configured to then compare the score of the original coding sequence (or a previously evaluated coding sequence) with the score of the modified coding sequence, and advance the coding sequence with the higher score (or the coding sequence with the lower score, randomly selected based on the above-described scoring function, etc., as described above) to the next iteration. Z(S, S o , S P ) When multiple iterations are completed, the encoding sequence of the last iteration is identified and advanced. Next, the computing device 102 is configured to swap the sequences of the chain with sequences from other chains, and then continue with further iterations. The swapping causes certain sequences to change under different temperature constraints, which may be randomly selected (or related to a given annealing schedule).
[0121] After a defined number of iterations and exchanges between chains, computing device 102 is configured to identify one or more sequences as output sequences based on the scoring as described above. Output sequences 108 are then provided to validation stage 110, where output sequences 108 are synthesized and tested as described above.
[0122] Despite the above, due to the form of the scoring function, it may be appropriate to use a weighted combination of such scoring functions (minimum, maximum, sum, etc.). The component functions can be trained on different k-polymer lengths or gap structures, or can be trained on different data sets. For example, a scoring function derived from a gene that is relatively highly expressed in the root can be combined with a function derived from a gene that is relatively highly expressed in the shoot, thereby producing a design that should be relatively highly expressed in both the root and the shoot.
[0123] In certain embodiments, a plurality of scoring functions are combined to retain the information portion of each function. For each k-aggregate and position, either use the value of the most significant scoring function, or if there is no significant scoring function, then all scoring functions are averaged. In certain embodiments, the position-independent method can be used to design a synthetic genetic control element or its part. In other embodiments, a hybrid approach can be used, wherein adopting the above-mentioned position-dependent method to design the first part of a synthetic control element sequence, and adopting the position-independent method to design the second part of a synthetic control element.
[0124] Position-independent method is based on the observation of promoter.However, description in this article is not limited to promoter, but can be used together with any genetic control element.For promoter, it is observed that the most significant position-specific enrichment of k-aggregates in promoter may occur at about 200 bases before TSS.In the more upstream of TSS, enrichment signal is usually weaker and may be unreliable.This is consistent with the understanding of this area, i.e., there is a "core promoter" element of high position sensitivity near TSS, and there is an enhancing or regulating element with lower position specificity at a distance from TSS.Therefore, a hybrid synthetic promoter is designed, which optimizes the Z (S) (about -200 to +50) in the core promoter region and the alternative score (about -500 to -200) in the upstream regulatory region.According to the size of the naturally occurring Arabidopsis (Arabidopsis) promoter, the regulatory region of 300bp is selected for experimental testing, but longer or shorter regions may play similar functions.
[0125] In certain embodiments, including embodiments involving genetic regulatory elements of viral promoters, the TSS may be unknown. In embodiments where the TSS is unknown or even in embodiments where the TSS is known, the promoter may be aligned on its TATA box instead. For example, for viral promoters, some signals (e.g., TATA boxes, etc.) are much stronger than other signals, so that it is difficult to select a suitable bandwidth w for the kernel density estimation step: too little smoothing makes it difficult to detect more dispersed signals, but too much smoothing can lead to tandem repeats of strong motifs (such as TATA boxes). Therefore, if necessary, standard kernel density estimation can be replaced by an adaptive variant. Depending on the local density, the bandwidth of each motif and each position is different: weak signals are smoothed more and strong signals are smoothed less. For large background groups, the computational cost is very high, so it is particularly suitable for the "simulated background" method, in which only a small group of sequences need to be processed. Alternatively, adaptive KDE can be used for the inner group, and fixed bandwidth KDE can be used for the outer group, because the outer group is highly heterogeneous, and therefore sharp peaks are not expected to appear (except for the possible TATA box).
[0126] Figure 3An example computing device 300 is shown that may be used in system 100. Computing device 300 may include, for example, one or more servers, workstations, personal computers, laptops, tablet computers, distributed computing systems, embedded systems, stand-alone electronic devices, cloud-based platforms, mobile devices, network devices, etc. Additionally, computing device 300 may include a single computing device, or it may include multiple computing devices located in close proximity or distributed over a geographic area, so long as the computing devices are specifically configured to function as described herein. Figure 1 In the system 100 of the present invention, the computing device 102 and the database 104 may include a computing device consistent with the computing device 300, or may be implemented in a computing device consistent with the computing device 300. That is, the system 100 or portions thereof should not be understood as limited to the computing device 300, as other computing devices may be employed in other system embodiments. In addition, different components and / or component arrangements may be used in other computing devices.
[0127] Furthermore, although the computing device 102 is illustrated as a physical machine, the computing device 102 may be implemented in one or more virtual machines, or through a cloud service, to implement the operations and / or methods therein.
[0128] See also Figure 3 , the example computing device 300 includes a processor 302 and a memory 304 coupled to (and in communication with) the processor 302. The processor 302 may include one or more processing units (e.g., a multi-core configuration, etc.). For example, the processor 302 may include, but is not limited to, a central processing unit (CPU), a microcontroller, a reduced instruction set computer (RISC) processor, an application specific integrated circuit (ASIC), a programmable logic device (PLD), a gate array, and / or any other circuit or processor capable of implementing the functionality described herein.
[0129] As described herein, memory 304 is one or more devices that allow data, instructions, etc. to be stored and retrieved therefrom. Memory 304 may include one or more computer-readable storage media, such as, but not limited to, dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), solid-state device, flash drive, CD-ROM, thumb drive, floppy disk, tape, hard disk and / or any other type of volatile or non-volatile physical or tangible computer-readable storage medium. Memory 304 may be configured to store (but not limited to) temperature schedules, harmful motifs, miRNA target sites, initialization parameters, scoring functions, sequences, and / or other types of data (and / or data structures) as needed and / or suitable for the purposes described herein. In addition, in various embodiments, computer executable instructions may be stored in memory 304 for processor 302 to perform, so that processor 302 performs one or more functions described herein, so that memory 304 is a physical, tangible and non-temporary computer-readable storage medium. Such instructions often improve the efficiency and / or performance of the processor 302 in performing one or more of the various operations herein (e.g., one or more of the operations of method 400, etc.), whereby the computing device 300 can be converted into a special-purpose computing device. It should be understood that the memory 304 may include a variety of different memories, each of which is implemented in one or more operations or processes described herein.
[0130] In an example embodiment, computing device 300 includes an output device 306 (or presentation unit) coupled to (and in communication with) processor 302 (however, it should be understood that computing device 300 may include output devices other than output device 306, etc.). Output device 306 outputs information (e.g., output coded sequences, scores, etc.) to a user of computing device 300 (e.g., a user associated with conversion engine 130, a user associated with verification stage 110, etc.) in a visual or audible manner. Various interfaces (e.g., interfaces defined by a web-based application, etc.) may be displayed at computing device 300, and in particular at output device 306, to display such information. Output device 306 may include, but is not limited to, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic LED (OLED) display, an "electronic ink" display, a speaker, etc. In some embodiments, output device 306 includes multiple devices.
[0131] The computing device 300 also includes an input device 308 that receives input from a user (e.g., user input, etc.), such as a selection of a gene of interest, etc., or input from another computing device. The input device 308 is coupled to (and communicates with) the processor 302 and may include, for example, a keyboard, a pointing device, a touch-sensitive panel (e.g., a touchpad or touch screen, etc.), another computing device, and / or an audio input device. In addition, in various example embodiments, a touch screen, such as a touch screen included in a tablet computer, a smart phone, or a similar device, acts as both an output device 306 and an input device 308.
[0132] In addition, the illustrated computing device 300 includes a network interface 310 coupled to (and in communication with) the processor 302 and the memory 304. The network interface 310 may include, but is not limited to, a wired network adapter, a wireless network adapter, a mobile network adapter, or other devices capable of communicating to / with one or more different networks. Furthermore, in some example embodiments, the computing device 300 includes a processor 302 and one or more network interfaces incorporated into or with the processor 302.
[0133] Consistent with the above, it should be understood that the present disclosure may be related to a computer system or a computer-implemented method. In general, in such an embodiment, the system 100 may, for example, include a data source in a database 104 (e.g., one or more databases or data structures generated or made, or linked to an external database, etc.), such as nucleotide sequence and / or gene expression data. The computing device 102 may, for example, then include a computer executable program or routine to process data (e.g., software, firmware, hardware, or any combination thereof, etc.), and it configures the computing device 102 to perform operations described herein. The program provides output (e.g., in the form of stored data, etc.) for a memory or output device, or provides output for the verification phase 110, thereby implementing and testing a regulatory element.
[0134] Figure 4 An example method 400 for generating or identifying synthetic regulatory elements based on gene expression probabilities is shown. Method 400 is described with reference to computing device 102, system 100, and computing device 300. That being said, it should be understood that the methods herein are not limited to system 100 or computing device 300, as other architectures, devices, etc. may be employed. Likewise, it should also be understood that the systems and / or devices herein should not be construed as limited to method 400, as other suitable methods may be employed in the systems and / or devices herein.
[0135] At the outset of method 400, in this example, it should be understood that one or more feature extraction techniques (as generally described above in system 100) may be employed, whereby certain features of a sequence associated with a selected or desired gene expression characteristic may be identified (e.g., a list of motifs and sequence features for templating within (or outside) a target element, sequence features associated with a target of interest (or a target element of interest), etc.).
[0136] It should also be understood that method 400 is provided as a simulated annealing process. Generally speaking, simulated annealing is a stochastic optimization technique that includes a parameter exploration phase, in which a large number of different parameter configurations can be evaluated, followed by a parameter exploitation phase, in which a small set of well-performing parameter configurations are identified (in order to further improve the identified configurations). By using the exploration and exploitation phases, the simulated annealing implemented in method 400 can enhance the ability of the algorithm to effectively converge to a desired solution.
[0137] Nevertheless, as described above, parallel tempering can provide an alternative method to utilize such exploration and exploitation phases by executing the phases simultaneously in parallel threads and allowing the phases to exchange information about the parameter landscapes they have discovered, for example, at set time intervals. For example, parallel tempering can provide N copies (or chains) of a sequence, randomly initialized at different temperatures. Each of these sequence copies is called a chain or replica. Moreover, parallel tempering can make a sequence modified at a higher temperature available for modification at a lower temperature, and vice versa, whereby parallel tempering can simultaneously use a high temperature chain to explore low-performance regions and use a low temperature chain to exploit high-performance regions (e.g., as a way to refine regions, etc.). Therefore, parallel tempering can improve the performance of the algorithm by increasing the amount and quality of information seen and / or available to each chain.
[0138] Specific reference Figure 4 , method 400 is shown as an experiment, in which a starting sequence of a regulatory element is selected, and then method 400 is performed to obtain one or more output sequences of the regulatory element. The experiment can then be repeated as needed with the same or different starting sequences, parameters, etc. A combination of methods such as Figure 1 The verification stage shown (eg, verification stage 110, etc.) is used to determine success.
[0139] Specifically, combined Figure 4 At 402 , the computing device 102 compiles a configuration file of the regulatory element and parameters related to the regulatory element to be defined.
[0140] Specifically, as shown in the figure, computing device 102 accesses data from motif data structure 440 from database 104, and described data structure comprises inner group motif and outer group motif.As mentioned above, inner group of motif comprises the motif that should be included in the regulating element sequence.Similarly, outer group of motif comprises the motif that should not be included in the regulating element sequence.Described group can comprise several motifs, tens of motifs or may comprise hundreds of or more or less motifs etc.In addition, computing device 102 accesses distribution data structure 442 from database 104, and it comprises one or more gene distributions of sequence.Specifically, in this example, database 104 comprises distribution data structure 442, and it is usually specific to the type of regulating element.Therefore, for example, when the regulating element to be generated is a promoter, gene distribution data structure 442 comprises previously tested or known promoter sequence, and described sequence is desirable or undesirable. In this way, the gene distribution data structure 442 includes another set of ingroup regulatory elements and a set of outgroup regulatory elements, whereby the method 400 evaluates the similarity of the potential sequence to sequences in the ingroup and outgroup, and promotes the similarity to the ingroup and reduces (or penalizes) the similarity to the outgroup, etc.
[0141] Consistent with the above, the ingroup / outgroup data structure 440 of motifs provides specific inclusion / exclusion of subsets of sequences, while the gene distribution data structure 442 provides more general guidance regarding sequence content.
[0142] The configuration file based on the above may include specific motifs, gene distributions, and different parameters of the simulated annealing process, such as the number of iterations, temperature schedules, the way to modify the sequence (e.g., random, etc.), etc. It should be understood that some parameters may be set by a user (e.g., a technician, etc.) associated with the computing device 102, and some parameters may be automatically selected (e.g., default values, etc.), which may or may not be changed by the user. Other parameters (e.g., weight values of different items of the scoring function, etc.) may be determined by analysis based on user preferences, historical success mining, machine learning methods, random chance, and / or any other appropriate method.
[0143] Thereafter, at 404, the computing device 102 initializes the scoring function for the experiment and applies the profile parameters and the starting sequence. At the start of the experiment, the computing device 102 calculates a score for the starting sequence at 406. The score is based on one or more equations described above. In this example embodiment, the computing device 102 employs equation (16) and thus generates a score indicating the predicted performance of one sequence relative to one or more other sequences (during a simulated annealing process, etc.).
[0144] Next, computing device 102 initiates an iterative process at 408 that includes n iterations, where n may be any integer greater than 1. In this example, n is 50, 100, 1,000, 10,000, 100,000, or more or less, etc.
[0145] In addition, according to the temperature schedule 410 in the data structure (e.g., memory 204, etc.), the iteration process is set to a specific temperature (T), where the temperature may include 10, 5, 3, 1, 0.1, 0.01 (as a dimensionless unit) or others. For example, when the temperature value is 10 (which affects the degree of change of the sequence), the computing device 102 starts multiple iterations (e.g., iteration (i) = 1 to 100, etc.) at 408 to modify the sequence of the given temperature selection. Specifically, to start, the computing device 102 first generates a new sequence at 412, which is based on the starting sequence and the selected Nth temperature.
[0146] Specifically, in this example embodiment, the starting sequence is changed by random changes in the sequence, such as one nucleotide per iteration. Random changes may include, for example, random position changes of a point, in which one nucleotide is randomly exchanged with another (e.g., point mutations at random positions, etc.). In general, the change in each iteration cycle will involve one bp, and the position of which bp is modified is random. Therefore, the temperature generally indicates the probability of accepting a new sequence that is worse than the old sequence in score (e.g., a high temperature generally indicates that the algorithm explores more, and as the system cools, it becomes less risk-averse, but more and more inclined to improve the current best sequence; etc.).
[0147] Continue to refer Figure 4 After changing the sequence, the computing device 102 calculates the sequence score of the changed sequence at 414. To this end, the computing device 102 uses the above-mentioned scoring function to calculate the score. Thereafter, the computing device 102 determines whether the changed sequence is enhanced relative to the input sequence based on the score of the iterative changed sequence and the score of the input sequence (from step 406) at 416. For example, when the score of the changed sequence is higher than the score of the input sequence, the changed sequence is accepted at 418, saved in a memory (e.g., memory 204, etc.) and returns to step 408 as the input sequence for the next iteration (e.g., i=2, etc.).
[0148] Conversely, when the score of the changed sequence is less than the score of the input sequence, computing device 102 may reject the changed sequence at 420 and then return to step 408, where the previous input sequence is again relied upon as the input sequence in the next iteration (e.g., i=2, etc.).
[0149] Additionally or alternatively, when the score of the changed sequence is less than the score of the input sequence, the computing device 102 may apply equation (17) above to determine the probability between the scores based on the temperature, and then compare the probability to a randomly generated threshold (e.g., a threshold between 0 and 1, etc.). It should be understood that one or more other equations (in addition to equation (17)) may be used to perform this analysis in other embodiments. When the probability does meet the randomly generated threshold, it is accepted and returns to step 408 as the input sequence for the next iteration (e.g., i=2, etc.). Conversely, when the probability does not meet the random threshold, the new encoding sequence is rejected at 420 and the previous input encoding sequence is reused as the input encoding sequence for the next iteration (e.g., i=2, etc.). In this way, despite the lower score, the changed sequence with the lower score still has the opportunity to advance as the starting sequence for the next iteration, thereby potentially applying flexibility in method 400.
[0150] When each iteration is completed (e.g., when i=100 in the illustrated embodiment, etc.), computing device 102 identifies the output sequence and stores it in memory (e.g., memory 204, etc.). Then, as Figure 4 As shown, the computing device 102 returns to the iteration in step 408 and continues to process the next temperature (e.g., the next Nth temperature, etc.) in the temperature data structure (e.g., 5, 3, 1, 0.1, 0.01, etc.), where the output sequence from the previous temperature is the starting sequence for the new temperature, or a different starting sequence may be used, etc. The computing device then proceeds as described above with respect to steps 412-420.
[0151] Repeat and complete steps 408-420 for each temperature in the temperature schedule, then computing device 102 identifies the output sequence and stores it in a memory (e.g., memory 204, etc.). Then the output sequence is subjected to one or more post-inspections. Specifically, in this example embodiment, computing device determines whether to meet the similarity post-inspection at 422 places. In this example, the post-inspection is global similarity (e.g., whether the output sequence is sufficiently different from the starting sequence (e.g., similarity is less than 75%, etc.) and local similarity, i.e., whether the specific section / block of the sequence is different (e.g., there is no common extension section of 22 nucleotides, etc.). For example, as mentioned above, the scoring function of this paper can punish sequences with redundancy, but will not completely forbid them to be considered (e.g., in order to avoid not considering the temporary conversion between two valid sequences, etc.). Post-inspection can be used to evaluate this type of retained sequence to ensure that it should actually be retained. When computing device 102 determines that the sequence does not meet the post-inspection at 424 places, new encoding sequences are discarded at 426 places.
[0152] In contrast, when computing device 102 determines at 424 that the sequence satisfies the post-check, the sequence is accepted at 428 , whereupon the sequence is stored in memory (eg, memory 304 , etc.), and then synthesized according to the description herein.
[0153] Synthetic regulatory elements are not limited to any particular size, but in some embodiments, the sequence generated or operably linked to the gene of interest is at least 25 nucleotides, at least about 30 nucleotides, at least about 40 nucleotides, at least about 50 nucleotides, at least about 60 nucleotides, at least about 70 nucleotides, at least about 80 nucleotides, at least about 90 nucleotides, at least about 100 nucleotides, at least about 150 nucleotides, at least about 200 nucleotides, at least about 250 nucleotides, at least about 300 nucleotides, at least about 350 nucleotides, at least about 400 nucleotides, at least about 450 nucleotides, at least about 500 nucleotides, at least about 550 nucleotides, at least about 600 nucleotides, or at least about 1 kb in length.
[0154] In certain embodiments, the method also includes synthesizing a nucleic acid molecule comprising a synthetic nucleotide sequence and / or testing the nucleic acid molecule as a synthetic control element to determine whether the synthetic genetic control element can regulate the gene expression of the gene of interest operably connected in a desired manner and / or in a desired cell or organism. As used herein, the term "operably connected" refers to the association of a nucleic acid sequence so that the function of a nucleic acid sequence is regulated by another nucleic acid sequence. For example, when a promoter can regulate the expression of a coding sequence (for example, the coding sequence is under the transcriptional control of a promoter, etc.), the promoter is operably connected to the coding sequence. The coding sequence can be operably connected to a control sequence in a sense or antisense direction. In another example, the complementary RNA region can be directly or indirectly at the 5' of the target mRNA, or at the 3' of the target mRNA, or operably connected in the target mRNA, or the first complementary region is located at the 5' of the target mRNA and its complement is located at 3'.
[0155] Generally, the function of genetic control elements is determined by transforming organisms or at least one cell thereof with a polynucleotide construct comprising the genetic control elements operably connected to a gene of interest. If desired or necessary to express a gene of interest in an organism or at least one cell thereof, the polynucleotide construct may further comprise additional naturally occurring or synthetic genetic control elements. It will be appreciated by those skilled in the art that whether the genetic control elements of determination synthesis can regulate the expression of a gene operably connected in a target organism or any other organism of interest in a desired manner may depend on many factors, including the type of genetic control elements such as produced by the methods disclosed herein, the presence of additional genetic elements in the expression construct, a gene of interest to be expressed, an organism or its part or cell measured therein, expression assay, detection method (e.g., GFP visible fluorescence, GFP RNA detected by qPCR), environmental conditions during determination, etc.
[0156] For example, in certain embodiments, where the synthetic genetic control element is a promoter, and the expression of the gene of interest is evaluated by the expression of the encoded protein, in the absence of an enhancing intron in the polynucleotide construct, about 5-15% of the genetic control elements produced by the method herein can show expression that can be detected by GFP fluorescence confocal imaging in T1 generation Arabidopsis. However, when the polynucleotide construct further comprises an enhancing intron, when measured in Arabidopsis by the method disclosed herein, about 60% of the synthetic genetic control elements show detectable expression by GFP fluorescence confocal imaging in T1 generation. Similarly, when promoter activity is determined at the nucleic acid level, such as by sensitive qPCR detection, about 60% of the genetic control elements can show detectable promoter activity without adding enhancing introns. These results show that most of the synthetic promoters produced by the method disclosed herein have biological promoter activity in plants.
[0157] When determining whether the synthetic genetic control element can regulate the expression of the operably connected gene in a desired manner, a reporter gene can be used. As used herein, "reporter" or "reporter gene" refers to a nucleic acid molecule encoding a detectable label. Preferred reporter genes include, for example, luciferase (e.g., firefly luciferase or Renilla luciferase, etc.), beta-galactosidase, chloramphenicol acetyltransferase (CAT) and fluorescent proteins (e.g., green fluorescent protein (GFP), red fluorescent protein (DsRed), yellow fluorescent protein, blue fluorescent protein, cyan fluorescent protein, or its variants, including enhanced variants, such as enhanced GFP (eGFP) etc.). Reporter genes can be detected by reporter assays. Reporter assays can measure the level of reporter gene expression or activity in a variety of ways, including, for example, measuring the level of reporter mRNA, the level of reporter protein or the amount of reporter protein activity. Reporter assays are known in the art or otherwise disclosed herein.
[0158] The synthetic genetic control elements produced as described herein are not limited to use in the target organisms derived from one or more groups of genes as described herein. In one example, the genetic control elements produced as described herein using the first group of nucleotide sequences from the genetic control elements of Arabidopsis thaliana can be used for regulating and controlling the expression of the gene of interest operably connected in Arabidopsis thaliana, soybean plant and / or one or more other interested dicotyledons. In another example, the synthetic genetic control elements produced as described herein using the first group of nucleotide sequences from the synthetic genetic control elements of rice can be used for regulating and controlling the expression of the gene of interest operably connected in rice plant, corn plant and / or one or more other interested monocotyledons. In another example, the synthetic genetic control elements produced as described herein using the first group of nucleotide sequences from the synthetic genetic control elements of Cauliflower mosaic virus family can be used for regulating and controlling the expression of the gene of interest operably connected in Arabidopsis thaliana, soybean plant, rice plant, corn plant and / or one or more other interested monocotyledons and / or dicotyledons.
[0159] In some embodiments, the synthetic regulatory element is a promoter. "Promoter" refers to a nucleic acid that can control the expression of an operably connected coding sequence or other sequences of the RNA that the coding is not necessarily translated into protein. The promoter sequence may include proximal and more distal upstream elements, the latter element is generally referred to as an enhancer. "Enhancer" is a DNA sequence that can stimulate promoter activity, and can be an intrinsic element of the promoter or a heterologous element inserted to enhance the level of the promoter or tissue specificity. It will be appreciated by those skilled in the art that different promoters can instruct genes to be expressed in different tissues or cell types, or at different developmental stages, or in response to different environmental conditions. It is further recognized that, since in most cases, the exact boundaries of the regulatory sequence have not yet been fully determined, some nucleic acid fragments of variations may have the same promoter activity.
[0160] In some embodiments, the promoter is a plant promoter. "Plant promoter" is a promoter that can initiate transcription in a plant cell, whether or not its source is a plant cell. For example, it is well known that the Agrobacterium promoter has a function in a plant cell. Therefore, plant promoters include promoter DNA obtained from plants, plant viruses and bacteria such as Agrobacterium and Bradyrhizobium, and synthetic promoters that can initiate transcription in a plant cell. Plant promoters can be constitutive promoters, non-constitutive promoters, inducible promoters, repressible promoters, tissue-specific promoters (e.g., root-specific promoters, stem-specific promoters, leaf-specific promoters, ovary-specific promoters, anther-specific promoters, tassel-specific promoters, etc.), tissue-preferred promoters (e.g., root-preferred promoters, stem-preferred promoters, leaf-preferred promoters, ovary-preferred promoters, anther-preferred promoters, tassel-preferred promoters, etc.), cell type-specific or preferred promoters (e.g., meristem cell-specific / preferred promoters, egg cell-specific / preferred promoters, pollen-specific / preferred promoters, etc.) or any other type.
[0161] A constitutive promoter is a promoter that is active under most conditions and / or during most developmental stages. The use of a constitutive promoter in an expression vector for plant biotechnology may have certain advantages, such as: high-level production of proteins for selection of transgenic cells or plants; high-level expression of reporter proteins or scorable markers, easy to detect and quantify; high-level production of transcription factors as part of a regulatory transcription system; production of compounds that require ubiquitous activity in plants; and production of compounds required at all stages of plant development. For illustration, constitutive promoters include the CaMV 35S promoter, the coronary promoter, the ubiquitin promoter, the actin promoter, the alcohol dehydrogenase promoter, and the like. In some embodiments, a synthetic promoter prepared as described herein is used to drive expression of a heterologous sequence, while the CaMV35S promoter is used to drive expression of a second sequence.
[0162] A non-constitutive promoter is a promoter that is active under certain conditions, in certain types of cells, and / or during certain developmental stages. For example, tissue-specific or preferred, cell-type-specific or preferred, inducible promoters, and promoters under developmental control are non-constitutive promoters. Examples of promoters under developmental control include promoters that preferentially initiate transcription in certain tissues (e.g., stems, leaves, roots, embryos, or seeds).
[0163] An "inducible" or "repressible" promoter is a promoter that is controlled by a chemical or environmental factor. Examples of environmental conditions that may affect transcription from an inducible promoter include cold, heat, drought, light, or certain chemicals.
[0164] A "tissue-specific" promoter is a promoter that initiates transcription only in certain tissues. Unlike constitutive expression of a gene, tissue-specific expression is the result of several interacting levels of gene regulation. Therefore, it is sometimes preferred to use a promoter from a homologous or closely related plant species in order to achieve efficient and reliable expression of a transgene in a specific tissue. This is one of the main reasons why a large number of tissue-specific promoters isolated from specific plants and tissues are found in the scientific and patent literature. Non-limiting tissue-specific promoters include the β-amylase gene or hordein gene promoters (for seed gene expression), the tomato pz7 and pz130 gene promoters (for ovary gene expression), the tobacco RD2 gene promoter (for root gene expression), the banana TRX promoter and the melon actin promoter (for fruit gene expression), as well as embryo-specific promoters, such as those associated with the amino acid permease gene (AAP1), the oleate 12-hydroxylase:desaturase gene (LFAH12) from Lesquerella fendleri, the 2S2 albumin gene (2S2), the fatty acid elongase gene (FAE1), or the leafy cotyledon gene (LEC2). For example, a "root-specific" promoter is a promoter that initiates transcription only in root tissue.
[0165] A "tissue-preferred" promoter is one that initiates transcription primarily, but not necessarily exclusively or exclusively, in certain tissues. For example, a "root-preferred" promoter is one that initiates transcription primarily, but not necessarily exclusively or exclusively, in root tissue.
[0166] A "cell type-specific" promoter is one that drives expression primarily in certain cell types in one or more organs (eg, roots, vascular cells in leaves, stem cells, and stem cells).
[0167] A "cell type preferred" promoter is one that drives expression primarily, but not necessarily exclusively or exclusively, in certain cell types in one or more organs (eg, roots, vascular cells in leaves, stem cells, or stem cells).
[0168] In some embodiments, the synthetic regulatory element is an expression enhancing intron.An "expression enhancing intron" or "enhancer intron" is an intron that is capable of increasing the expression of a gene to which it is operably linked.
[0169] In some embodiments, the synthetic regulatory element is a transcription termination region (or 3'UTR). As used herein, the term "3' transcription termination molecule", "3' untranslated region" or "3'UTR" refers to a DNA molecule used during transcription of the untranslated region of the 3' portion of an mRNA molecule. The 3' untranslated region (also referred to as the polyA tail) of an mRNA molecule can be generated by specific cleavage and 3' polyadenylation. The 3'UTR can be operably connected to a transcribable DNA molecule (e.g., a gene of interest, etc.) and located downstream thereof, and can include polyadenylation signals and other regulatory signals that can affect transcription, mRNA processing or gene expression. The polyA tail is believed to play a role in mRNA stability and translation initiation. Examples of 3' transcription termination molecules in the art are nopaline synthase 3' region, wheat hsp17 3' region, pea ribulose diphosphate carboxylase (rubisco) small subunit 3' region, cotton E6 3' region, and coixin 3'UTR. The 3'UTR is generally beneficial to the recombinant expression of a specific DNA molecule.
[0170] In other aspects, the disclosure herein provides preparation of expression vectors, transgenic cells or non-human transgenic organisms, as described herein, for producing synthetic regulatory elements. The disclosure relates to synthetic regulatory elements produced as described herein operably linked to a gene of interest, thereby producing an expression construct. Such genes of interest will depend on the desired results and may include nucleotide sequences encoding proteins and / or RNA of interest. Nucleic acid molecules may then be synthesized or produced using a variety of methods known in the art. These methods include chemical synthesis and recombinant techniques. The disclosure herein may also relate to transforming at least one cell with a polynucleotide construct. The disclosure may also relate to propagating cells or regenerating transgenic organisms from transformed cells.
[0171] As used herein, the phrases "recombinant construct", "expression construct", "chimeric construct", "construct" and "recombinant DNA construct" are used interchangeably. Recombinant constructs include artificial combinations of nucleic acid fragments, such as regulatory sequences and coding sequences that do not exist simultaneously in nature. For example, a chimeric construct may include regulatory sequences and coding sequences derived from different sources, or regulatory sequences and coding sequences derived from the same source but arranged in a manner different from that found in nature. Such constructs may be used alone or in combination with a vector. If a vector is used, the choice of vector depends on the method for transforming host cells well known to those skilled in the art. For example, a plasmid vector may be used. The skilled person is fully aware of the genetic elements that must be present on the vector in order to successfully transform, select and propagate host cells containing any isolated nucleic acid fragment of the present disclosure. Screening for transformants can be accomplished by Southern analysis of DNA, Northern analysis of mRNA expression, immunoblot analysis of protein expression, or phenotypic analysis, etc. The vector may be a plasmid, a virus, a phage, a provirus, a phagemid, a transposon, an artificial chromosome, etc., which can replicate autonomously or integrate into the chromosome of the host cell. The vector may also be a naked RNA polynucleotide, a naked DNA polynucleotide, a polynucleotide composed of both DNA and RNA in the same chain, polylysine-bound DNA or RNA, peptide-bound DNA or RNA, liposome-bound DNA, etc., which do not replicate autonomously.
[0172] The cassette may additionally contain at least one additional gene to be co-transformed into the organism. Alternatively, additional genes may be provided on multiple expression cassettes. Such expression cassettes have multiple restriction sites and / or recombination sites for inserting polynucleotides to be regulated by transcription of the regulatory region. Where appropriate, the gene of interest may be optimized to increase its expression in the transformed plant. That is, polynucleotides may be synthesized using plant preferred codons to improve expression. For a discussion of host preferred codon usage, see, for example, Campbell and Gowri (1990) Plant Physiol. 92: 1-11. Methods for synthesizing plant preferred genes are available in the art. See, for example, U.S. Pat. Nos. 5,380,831 and 5,436,391 and Murray et al. (1989) Nucleic Acids Res. 17: 477-498, which are incorporated herein by reference.
[0173] The expression cassette may also include a selectable marker gene for selecting transformed cells. The selectable marker gene is used to select transformed cells or tissues. The marker gene includes a gene encoding antibiotic resistance, such as a gene encoding neomycin phosphotransferase II (NEO) and hygromycin phosphotransferase (HPT), and a gene that imparts resistance to herbicidal compounds, such as glufosinate, bromoxynil, imidazolinone, sulfonylurea, glyphosate, glufosinate, L-phosphinothricin, triazine, benzonitrile and 2,4-dichlorophenoxyacetic acid (2,4-D). Other selectable markers include phenotypic markers, such as β-galactosidase and fluorescent proteins, such as green fluorescent protein (GFP), cyan fluorescent protein (CYP), yellow fluorescent protein, etc.
[0174] Those of ordinary skill in the art will recognize that the systems and methods for generating synthetic regulatory elements are applicable to organisms other than plants. In some aspects, the disclosure herein provides preparing transgenic cells or non-human organisms by incorporating synthetic regulatory elements operably associated with coding sequences or other transcriptional genes into one or more cells, wherein the synthetic regulatory elements have a statistically significant score using a scoring function described herein. The cells are bred to produce transgenic cells or non-human organisms. It should be recognized that the genetic regulatory elements of the present disclosure and the expression cassettes comprising one or more such genetic regulatory elements can be used for expression in humans and non-human host cells. In some aspects, the disclosure herein provides preparing pluripotent stem cells by incorporating synthetic regulatory elements operably associated with coding sequences or other transcriptional genes into one or more pluripotent stem cells, wherein the synthetic regulatory elements have a statistically significant score using a scoring function described herein. In one embodiment of the present disclosure, the host cell is a human host cell or a host cell line that cannot be differentiated into humans.
[0175] The present disclosure may also involve introducing a polynucleotide construct into a plant. The term "introducing" means presenting the polynucleotide construct to a plant in a manner that enables the polynucleotide construct to enter the interior of a plant cell. The present disclosure does not rely on a particular method for introducing a polynucleotide construct into a plant, but only relies on the polynucleotide construct being able to enter the interior of at least one cell of a plant. This transformation may be stable or transient.
[0176] "Stable transformation" means that a polynucleotide construct introduced into a plant is integrated into the genome of the plant and can be inherited by its progeny. "Transient transformation" means that a polynucleotide construct introduced into a plant is not integrated into the genome of the plant.
[0177] Suitable methods for introducing nucleotide sequences into plant cells and subsequent insertion into the plant genome may include, for example, microinjection such as electroporation, direct gene transfer, and ballistic particle acceleration.
[0178] The polynucleotides herein can be introduced into plants by contacting plants with viruses or viral nucleic acids. In general, such methods involve integrating polynucleotide constructs into viral DNA or RNA molecules. In addition, it should be appreciated that the promoters produced as described herein also encompass promoters for transcription by viral RNA polymerases.
[0179] The transformed cells can be grown into plants according to conventional techniques. These plants can then be grown and pollinated with the same transformed strain or a different strain, and the resulting hybrids with constitutive expression of the desired phenotypic characteristic identified. Two or more generations can be grown to ensure that the expression of the desired phenotypic characteristic is stably maintained and inherited, and then the seeds are harvested to ensure that the expression of the desired phenotypic characteristic has been achieved.
[0180] As used herein, the term plant includes plant cells, plant protoplasts, plant cell tissue cultures from which plants can be regenerated, plant calli, plant pieces, and intact plant cells in plants or plant parts such as embryos, pollen, ovules, seeds, leaves, flowers, branches, fruits, roots, root tips, anthers, etc. Progeny, variants, and mutants of regenerated plants are also included within the scope of the present disclosure, as long as these parts contain the introduced polynucleotides (e.g., containing synthetic regulatory elements, etc.).
[0181] Particularly for plants, the genes of interest controlled by synthetic regulatory elements reflect the market and / or interests of those involved in crop development. General categories of genes of interest include, for example, genes related to information, such as zinc fingers; genes related to communication, such as kinases; and genes related to housekeeping, such as heat shock proteins. For example, more specific transgenic categories include genes encoding important traits of agronomy, insect resistance, disease resistance, herbicide resistance, sterility, grain characteristics, yield, abiotic stress tolerance and commercial products. Genes of interest generally include genes related to oil, starch, carbohydrate or nutrient metabolism. In addition, genes of interest include genes encoding enzymes and other proteins from plants and other sources (including prokaryotes and other eukaryotes).
[0182] In certain embodiments, the disclosure relates to transgenic plants and methods for preparing the same. As used herein, the term "plant" refers to any organism belonging to the plant kingdom (e.g., any genus / species in the plant kingdom, etc.). In some embodiments, the plant is a tree, herb, shrub, grass, liana, pteridophyte, moss, or green algae. The plant can be a monocot or a dicot. The example of a specific plant includes, but is not limited to, Arabidopsis thaliana, short-stalked grass, switchgrass, corn, potato, rose, apple tree, sunflower, wheat, rice, banana, tomato, gourd, pumpkin, zucchini, lettuce, cabbage, oak, impatiens, geranium, hibiscus, clematis, poinsettia, sugarcane, taro, duckweed, pine, Kentucky bluegrass, zoysia, coconut palm, cauliflower, cabbage (cavalo), kale, kale, kohlrabi, mustard, rapeseed and other cruciferous plants. Leafy vegetable crops, bulb vegetables (e.g., garlic, leeks, onions (shallots, green onions and scallions), shallots, etc.), citrus fruits (e.g., grapefruit, lemons, limes, oranges, tangerines, citrus hybrids, pomelos and other citrus fruit crops, etc.), cucurbit vegetables (e.g., cucumbers, pomelos, edible gourds, gherkins, melons (including cucumber hybrids and / or cultivars), watermelons, cantaloupes, etc.), fruiting vegetables (including eggplant, sour berries, peppers, tomatoes, melons, etc.), physalis (such as tarragon), grapes, leafy vegetables (such as romaine lettuce, etc.), root / tuber and bulb vegetables (such as potatoes, etc.), and tree nuts (almonds, pecans, pistachios, and walnuts), berries (such as tomatoes, barberry, gooseberry, elderberry, currant, honeysuckle, American clover, barberry, Oregon grape, buckthorn, hackberry, bearberry, huckleberry, strawberry, sea grape, blackberry, cloudberry, blackberry, raspberry, sweetberry, raspberry, cranberry, etc.), cereal crops (such as corn, rice, wheat, barley, sorghum, millet, oats, rye, triticale, buckwheat, fonio, quinoa, oil palm, etc.), cruciferous plants and legumes, pome fruits (e.g., apples, pears, etc.), stone fruits (e.g., coffee, dates, mangoes, olives, coconuts, oil palms, pistachios, almonds, apricots, cherries, damsons, nectarines, peaches and plums, etc.), vines (e.g., table grapes, wine grapes, etc.), fiber crops (e.g., hemp, cotton, etc.), ornamental plants, etc.
[0183] In some embodiments, the nucleic acid molecule is introduced into a plant by cloning the nucleic acid molecule into a binary vector suitable for plant-specific transformation. For example, to introduce the nucleic acid molecule into Brassica species, the nucleic acid molecule is cloned into a binary vector suitable for Brassica species transformation.
[0184] In certain embodiments, the plant is a cultivar. As used herein, the term "cultivar" refers to a plant variety, strain, or race that is produced by horticultural or agronomic techniques and is not normally found in wild populations.
[0185] In some aspects, the disclosure includes plant parts derived from transgenic plants described herein. As used herein, the term "plant part" refers to any part of a plant, including but not limited to buds, roots, stems, petioles, trunks, tillers, seeds, endosperms, pedicels, tubers, rhizomes, stipules, runners, nodules, leaves or sheaths, needles, cones, petals, flowers, ovules, fruits, berries, stigmas, bracts, pedicels, branches, styles, carpels, pericarps, petioles, internodes, bark, pubescence, pollen, stamens, pistils, sepals, anthers, placentas, etc. The two main parts of a plant grown in a certain medium (e.g., soil) are commonly referred to as the "aboveground" part (also often referred to as "buds") and the "underground" part (also often referred to as "roots").
[0186] In some embodiments, the present disclosure provides for preparing transgenic plants having a gene of interest controlled by a synthetic promoter, which is operable, for example, in rice. The transgenic plant may or may not be a rice plant. Furthermore, the synthetic promoter may be a highly constitutive promoter.
[0187] In some embodiments, the present disclosure provides a method for preparing a transgenic plant having a gene of interest controlled by a synthetic promoter. Also, the synthetic promoter can be a constitutive promoter.
[0188] In some embodiments, the present disclosure provides methods for preparing transgenic plants having a gene of interest under the control of a synthetic promoter. Also, the synthetic promoter can be a highly constitutive promoter.
[0189] In some embodiments, the present disclosure provides methods for preparing a transgenic plant having a gene of interest operably associated with a synthetic intron. Also, the synthetic intron can be an expression enhancing intron.
[0190] In some embodiments, transgenic plants have a gene of interest under the control of the synthetic promoter and synthetic intron described above.
[0191] The disclosed content can be used in combination with basic plant breeding techniques. For example, transgenic plants can be inbred lines or monoallelic conversion plants. As used herein, the term "inbred line" or "inbred plant" includes any single gene conversion of the inbred line. The phrase "single allele conversion plant" refers to a plant developed by a plant breeding technique called backcrossing, wherein in addition to the monoallelic gene transferred to the inbred line by backcrossing technology, all the required morphological and physiological characteristics of the inbred line are restored. In some embodiments, offspring plants can be obtained by cloning or selfing of parental plants or by hybridization of two parental plants, and include selfing and F1 or F2 or more distant generations. F1 is the first generation offspring produced by at least one parent as a trait donor for the first time, and the offspring of the second generation (F2) or subsequent generations (F3, F4, etc.) are samples produced by selfing of F1, F2, etc. Thus, F1 may be (and usually is) a hybrid produced by crossing two purebred parents (purebred is homozygous for a trait), and F2 may be (and usually is) a progeny produced by self-pollination of the F1 hybrid. Developing transgenic plants may also include crossing. As used herein, the terms "cross", "crossing", "cross pollination" or "cross breeding" refer to the process of applying pollen (artificial or natural) from a flower on one plant to the ovules (stigma) of a flower on another plant.
[0192] In certain embodiments, the disclosure relates to cell transformation. As used herein, the term "transformant" refers to a cell, tissue or organism that has undergone transformation. The original transformant may be designated as "T0" or "T0". T0 selfing produces the first transformation generation designated as "T1" or "T1".
[0193] In some embodiments, the transgenic cell or organism is hemizygous for a gene of interest that is controlled by a synthetic regulatory element. As used herein, the term "hemizygous" refers to a cell, tissue, or organism in which a gene is present only once in the genotype, such as a gene in a haploid cell or organism, a sex-linked gene in a heterogametic sex, or a gene in a chromosome segment in which the partner segment has been deleted in a diploid cell or organism.
[0194] In some embodiments, cell or organism is heterozygous for the gene of interest controlled by synthetic control elements.As used herein, the term "heterozygote" refers to a diploid or polyploid single cell or plant having different alleles (form of a given gene) on at least one locus.Similarly, the term "heterozygote" refers to the presence of different alleles (form of a given gene) at a specific gene locus.In other embodiments, cell or organism is a homozygote of the gene of interest controlled by synthetic elements.As used herein, the term "homozygote" refers to a single cell or plant having identical alleles at one or more loci.Therefore, the term "homozygote" refers to the presence of identical alleles at one or more loci in homologous chromosome fragments.
[0195] Any transgenic plant comprising one or more synthetic promoters and / or synthetic introns can be used as a donor to produce more transgenic plants by plant breeding methods well known to those skilled in the art. The overall goal is to develop new, unique and superior varieties and hybrids. In some embodiments, selection methods (e.g., marker-assisted selection) can be combined with breeding methods to accelerate the process.
[0196] In some embodiments, an exemplary method may include (i) crossing any plant of the present disclosure comprising one or more synthetic promoters and / or synthetic introns as a donor with a recipient plant line to produce an F1 population; (ii) evaluating transgene expression in progeny derived from the F1 population; and (iii) selecting progeny having functional transgene expression under the control of the synthetic promoter and / or synthetic intron.
[0197] In some embodiments, the entire chromosome of the donor plant is transferred. For example, a transgenic plant with a synthetic promoter and / or a synthetic intron can be used as a male or female parent in a cross-pollination to produce a progeny plant, wherein the progeny plant obtains the synthetic promoter and / or the synthetic intron by receiving the transgene from the donor plant. In some embodiments, only the genomic fragment containing the transgene (e.g., having a synthetic promoter and / or a synthetic intron, etc.) is integrated into the recipient plant.
[0198] The present disclosure further provides the use of plant breeding techniques in plant breeding programs to cultivate plants, including recurrent selection, backcrossing, pedigree breeding, molecular marker (isoenzyme electrophoresis, restriction fragment length polymorphism (RFLP), random amplified polymorphic DNA (RAPD), arbitrary primer polymerase chain reaction (AP-PCR), DNA amplification fingerprinting (DAF), sequence characteristic amplified region (SCAR), amplified fragment length polymorphism (AFLP) and simple sequence repeat (SSR), also known as microsatellite, etc.) enhanced selection, genetic marker enhanced selection and transformation. Seeds, plants and parts thereof produced by such breeding methods are also part of the present disclosure.
[0199] In this connection, several embodiments relate to the recombinant nucleic acid comprising one or more output coding sequences.As used herein, "recombinant nucleic acid" refers to a nucleic acid molecule (DNA or RNA) having a coding and / or non-coding sequence that can be distinguished from the endogenous nucleic acid found in the natural system.In some respects, the recombinant nucleic acid provided herein is used in any composition, system or method provided herein.In some respects, the recombinant nucleic acid provided herein is a transgenic.In some respects, the recombinant nucleic acid can encode any protein and can be used in any composition, system or method provided herein.In some embodiments, the recombinant nucleic acid comprises one or more output coding sequences that are operably connected to a heterologous promoter.In one aspect, the recombinant nucleic acid provided herein comprises one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more or ten or more heterologous promoters, and the heterologous promoter is operably connected to one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more or ten or more output coding sequences.
[0200] In some aspects, the recombinant nucleic acid may be included in a vector. As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid connected thereto. Vectors include, but are not limited to, single-stranded, double-stranded or partially double-stranded nucleic acid molecules; nucleic acid molecules comprising one or more free ends, no free ends (e.g., circular, etc.); nucleic acid molecules comprising DNA, RNA or both; and polynucleotides of other varieties known in the art. One type of vector is Agrobacterium T-DNA. Another type of vector is a viral vector, in which a virally derived DNA or RNA sequence is present in a vector for packaging into a virus (e.g., retrovirus, replication-defective retrovirus, tobacco mosaic virus (TMV), potato virus X (PVX) and cowpea mosaic virus (CPMV), tobacco mosaic virus, geminivirus, adenovirus, replication-defective adenovirus and adeno-associated virus, etc.). Viral vectors also include polynucleotides carried by viruses for transfection into host cells. In some embodiments, viral vectors can be delivered to plants using Agrobacterium. Certain vectors can replicate autonomously in host cells introduced into them. Other vectors are integrated into the genome of the host cell after being introduced into the host cell, and thus replicated with the host genome. In addition, some vectors can guide the expression of genes operably connected thereto. Such vectors are referred to herein as "expression vectors". It will be appreciated by those skilled in the art that the design of expression vectors may depend on factors such as the selection of host cells to be transformed, the required expression level, etc. The vector may be introduced into the host cell, thereby producing transcripts, proteins or peptides encoded by nucleic acids as described herein, including fusion proteins or peptides. In some embodiments, the expression vector may include one or more output coding sequences in a form suitable for expressing the output coding sequence in a plant cell, which means that the expression vector includes one or more regulatory elements operably connected to the output coding sequence to be expressed. Regulatory elements may include enhancers, terminator sequences, introns, etc.
[0201] In one aspect, the vector provided herein comprises any recombinant nucleic acid, and the recombinant nucleic acid comprises the output coding sequence provided herein. On the other hand, the plant cell provided herein comprises the recombinant nucleic acid provided herein. On the other hand, the plant cell provided herein comprises the vector provided herein. The plant cell can be a monocot or a dicot. In some embodiments, the plant cell can be from or belong to a crop or a cereal plant, such as cassava, corn, sorghum, alfalfa, cotton, soybean, rape, wheat, oats or rice. The plant cell can also be an algae, a tree, or a production plant, fruit or vegetable (e.g., a tree, such as a citrus tree (e.g., an orange, grapefruit or lemon tree; a peach or nectarine tree; an apple or pear tree, etc.); a nut tree, such as an almond or walnut or pistachio tree; a plant of the Solanaceae family; a plant of the genus Brassica; a plant of the genus Lactuca; a plant of the genus Spinach; a plant of the genus Capsicum; cotton, tobacco, asparagus, avocado, papaya, cassava, carrot, cabbage, broccoli, cauliflower, tomato, eggplant, pepper, lettuce, spinach, strawberry, potato, pumpkin, melon, blueberry, raspberry, blackberry, grape, coffee, cocoa, etc.).
[0202] That is, and as generally described above, a variety of methods are known in the art for transforming chromosomes or plastids in plant cells with recombinant DNA molecules, which can be used in accordance with the systems and methods of the present application to produce plant cells and plants comprising one or more output coding sequences (e.g., in validation stage 110, etc.).
[0203] In plants, recombinant nucleic acids can be delivered using particle bombardment or gene gun delivery. Particle bombardment is suitable for transforming plants with DNA, RNA, protein, or any combination thereof. Methods for transforming plants using gene gun delivery of DNA are described in PCT / US2019 / 033984 and are incorporated herein by reference in their entirety.
[0204] In plants, Agrobacterium-mediated transformation is a suitable method for delivering recombinant nucleic acids to one or more T-DNAs. Agrobacterium-mediated transformation is widely used in monocotyledons and dicotyledons. In one embodiment, an expression cassette comprising one or more output coding sequences can be provided as a double tumor induction (Ti) plasmid border construct, the construct having the right border (RB or AGRtu.RB) and left border (LB or AGRtu.LB) regions of the Ti plasmid isolated from the Agrobacterium tumefaciens comprising T-DNA, the T-DNA together with the transfer molecule provided by the Agrobacterium tumefaciens cell allowing the T-DNA to be integrated into the genome of the plant cell (see, e.g., U.S. Patent No. 6,603,061, which is incorporated herein by reference in its entirety). The construct may also contain a plasmid backbone DNA segment that provides replication function and antibiotic selection in bacterial cells, such as an Escherichia coli replication origin, such as ori322; a broad host range replication origin, such as oriV or oriRi; and a coding region for a selectable marker, such as Spec / Strp encoding Tn7 aminoglycoside adenylyltransferase (aadA) that confers resistance to spectinomycin or streptomycin or gentamicin (Gm, Gent) selectable marker genes. In some embodiments, one or more expression cassettes containing one or more output coding sequences are provided in a T-DNA binary vector (e.g., an OriRi vector backbone) with a low copy replication origin. For plant transformation, the host bacterial strain is often Agrobacterium tumefaciens ABI, C58, or LBA4404; however, other strains known to those skilled in the art of plant transformation may work in the present disclosure. In some embodiments, an Agrobacterium strain lacking certain DNA recombination functions (e.g., RecA) is used to deliver expression vectors encoding CAST system components to plant cells.
[0205] In some embodiments, an expression cassette comprising one or more output coding sequences described herein is provided on a single T-DNA. In some embodiments, an expression cassette encoding one or more output coding sequences described herein is provided on multiple separate T-DNAs and delivered to a plant cell in a single transformation process or in separate continuous transformation processes.
[0206] Several embodiments relate to plants comprising an output coding sequence as described herein in its genome. In certain embodiments, genome editing methods are used to modify or replace existing genome coding sequences in plant genomes with sequences encoding output coding sequences provided by the methods described herein, such as coding sequences of proteins that confer herbicide tolerance. In some embodiments, natural genome coding sequences are modified to include one or more targeted nucleotide changes, additions, deletions or other modifications to introduce output coding sequences provided by the methods described herein. Several embodiments relate to using known genome editing methods and site-specific genome modification enzymes, such as zinc finger nucleases, engineered or natural meganucleases, TALE-endonucleases or RNA-guided endonucleases (e.g., clustered regularly spaced short palindromic repeats (CRISPR) / Cas9 systems, CRISPR / Cpf1 systems, CRISPR / CasX systems, CRISPR / CasY systems, CRISPR / Cascade systems) to modify or replace existing coding sequences in plant genomes. Several embodiments relate to providing a site-specific genome modification enzyme that is capable of recognizing a specific gene of interest within a plant genome, thereby allowing the native sequence to be altered by non-templated editing or by template editing to introduce an output coding sequence provided by the methods described herein. Several embodiments relate to providing a base modifier (e.g., a deaminase) linked to a site-specific genome modification enzyme that is capable of recognizing a specific nucleotide sequence of interest within a plant genome, thereby allowing the base modifier to alter the native sequence to provide an output coding sequence.
[0207] Again, and as previously described, it should be understood that in some embodiments, the functions described herein may be described by computer executable instructions stored on a computer readable medium and executable by one or more processors. A computer readable medium is a non-transitory computer readable storage medium. For example, and not limitation, such computer readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage device or any other medium that can be used to carry or store a desired instruction or data structure form of program code and can be accessed by a computer. The above combination should also be included in the scope of the computer readable medium.
[0208] It should also be appreciated that one or more aspects of the present disclosure transform a general-purpose computing device into a special-purpose computing device when configured to perform the functions, methods and / or processes described herein.
[0209] Based on the foregoing description, it should be understood that the above embodiments of the present disclosure can be implemented using computer programming or engineering techniques including computer software, firmware, hardware, or any combination or subset thereof, wherein the technical effect can be implemented by performing at least one operation described herein and / or described in the corresponding technical scheme below. For example, the technical effect can be achieved in the following manner: (a) identifying an input sequence associated with a regulatory element as a starting sequence; (b) calculating a score for the starting sequence based on a scoring function; (c) initializing N iterations at at least one parameter (e.g., temperature, etc.), wherein N is an integer; (d) for each iteration in the N iterations: (i) changing at least one nucleotide in the input sequence for iteration; (ii) calculating a score for the changed sequence based on a scoring function; (iii) advancing the changed sequence to the next iteration based on at least the calculated score and a threshold of the changed sequence; and (iv) when the calculated score indicates an enhancement relative to the input coding sequence and the iteration is equal to N, identifying the changed sequence as the output sequence of the N iterations; and (e) after the N iterations, directing the output sequence to a verification stage, thereby synthesizing a sequence defining a regulatory element.
[0210] Example embodiments are provided so that the present disclosure will be exhaustive and will fully convey the scope to those skilled in the art. Many specific details, such as examples of specific components, devices, and methods, are set forth to provide a thorough understanding of the embodiments of the present disclosure. It is apparent to those skilled in the art that the example embodiments may be implemented in many different forms without employing specific details and that any one embodiment should not be construed as limiting the scope of the present disclosure. In some example embodiments, well-known processes, well-known device structures, and well-known techniques are not described in detail.
[0211] The specific values disclosed herein are essentially examples and do not limit the scope of the present disclosure. The disclosure of the specific values and specific value ranges of given parameters herein does not exclude other values and value ranges that can be used in one or more examples disclosed herein. In addition, it is envisioned that any two specific values of the specific parameters described herein can define the endpoints of the value range that can also be applicable to the given parameter (for example, the disclosure of the first value and the second value of the given parameter can be interpreted as disclosing any value between the first value and the second value and can also be used for the given parameter). For example, if parameter X is exemplified herein as having value A and is also exemplified as having value Z, it is envisioned that parameter X can have a value range from about A to about Z. Similarly, it is envisioned that the disclosure of two or more ranges of the value of the parameter (whether such ranges are nested, overlapping or different) covers all possible range combinations of values that may be claimed using the endpoints of the disclosed range. For example, if parameter X is illustrated herein as having a value within the range of 1-10, or 2-9, or 3-8, it is also contemplated that parameter X may have other ranges of values, including 1-9, 1-8, 1-3, 1-2, 2-10, 2-8, 2-3, 3-10, and 3-9.
[0212] The terms used herein are only for the purpose of describing specific example embodiments and are not intended to be limiting. As used herein, the singular forms "a, an" and "the" may be intended to include the plural forms as well, unless the context clearly indicates otherwise. The terms "comprises", "comprising", "includes" and "including" are inclusive, thus indicating the presence of the features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof. The method steps, processes and operations described herein should not be understood as necessarily requiring them to be performed in the specific order discussed or shown, unless specifically indicated as the order of performance. It should also be understood that additional or alternative steps may be used.
[0213] When a feature is referred to as being "on," "engaged to," "connected to," "coupled to," "associated with," "communicating with," or "included in" another element or layer, it may be directly on, engaged, connected or coupled to, associated with, communicating with, or included in another feature, or intervening features may be present. As used herein, the term "and / or" and the phrase "at least one of" include any and all combinations of one or more of the associated listed items.
[0214] Although the terms first, second, third, etc. may be used herein to describe various features, these features should not be limited by these terms. These terms may only be used to distinguish one feature from another. For example, terms such as "first", "second", and other numerical terms do not imply an order or sequence when used herein unless the context clearly indicates. Therefore, without departing from the teachings of the example embodiments, the first feature described herein may be referred to as the second feature.
[0215] No element recited in a claim is intended to be a means-plus-function element within the meaning of 35 U.S.C. §112(f) unless the element is expressly recited using the phrase "means for..." or in the case of a method claim using the phrase "operation of..." or "step of..."
[0216] The above description of the embodiments has been provided for the purpose of illustration and description. The description is not intended to be exhaustive or to limit the present disclosure. The individual elements or features of a particular embodiment are generally not limited to the particular embodiment, but are interchangeable where appropriate and may be used in selected embodiments, even if not specifically shown or described. Also, variations may be made in many aspects. Such variations should not be considered as departing from the present disclosure, and all such modifications are intended to be included within the scope of the present disclosure.
Claims
1. A computer-implemented method for identifying a regulatory element, the computer-implemented method comprising: identifying, by a computing device, an input sequence associated with a regulatory element as a starting sequence; Calculating, by the computing device, a score of the starting sequence based on a scoring function; Initialize N iterations at at least one parameter, where N is an integer; For each of the N iterations: changing at least one nucleotide in an input sequence to perform said iteration; Calculating a score for the altered sequence based on the scoring function; advancing the altered sequence to a next iteration based at least on the calculated score of the altered sequence and a threshold; as well as When the calculated score of the altered sequence indicates an enhancement relative to the input sequence and the iterations are equal to N, identifying the altered sequence as the output sequence of the N iterations; And then After the N iterations, the output sequence is directed to a validation phase to synthesize a sequence defining the regulatory element. 2 . The computer-implemented method of claim 1 , wherein advancing the altered sequence to the next iteration is based on the calculated score of the altered sequence being greater than the calculated score of the starting sequence. 3 . The computer-implemented method of claim 1 , wherein advancing the altered sequence to a next iteration is further based on the iteration being less than N. 4 .
4. The computer-implemented method of any one of claims 1 to 3, wherein advancing the altered sequence to the next iteration is further based on a probability function.
5. The computer-implemented method of any one of claims 1 to 4, further comprising discarding the altered sequence in response to the calculated score of the altered sequence being less than the calculated score of the starting sequence.
6. The computer-implemented method of any one of claims 1 to 5, wherein the regulatory element comprises one or more of a promoter, an intron, and / or an untranslated region (UTR).
7. The computer-implemented method of any one of claims 1 to 6, wherein calculating the score for the altered sequence comprises calculating the score based on:
8. The computer-implemented method of claim 1 , wherein based on a probability function satisfying the threshold and the iteration being less than N, advancing the altered sequence to the next iteration; and The threshold comprises a static threshold and a threshold randomly generated in each iteration.
9. The computer-implemented method of claim 8, wherein the probability function comprises: as well as wherein E is the calculated score of the starting sequence, E' is the calculated score of the altered sequence, and T is the at least one parameter.
10. The computer-implemented method of any one of claims 1 to 9, wherein N is less than 100.
11. The computer-implemented method of any one of claims 1 to 10, wherein the at least one parameter is temperature.
12. The computer-implemented method of any one of claims 1 to 11, further comprising synthesizing the regulatory element defined by the output sequence.
13. A non-transitory computer-readable storage medium comprising executable instructions for identifying a regulatory element, the instructions, when executed by at least one processor, causing the at least one processor to: identifying an input sequence associated with a regulatory element as a starting sequence; Calculating a score for the starting sequence based on a scoring function; Initialize N iterations at at least one parameter, where N is an integer; For each of the N iterations: changing at least one nucleotide in an input sequence to perform said iteration; Calculating a score for the altered sequence based on the scoring function; advancing the altered sequence to a next iteration based at least on a calculated score and / or a threshold of the altered sequence; as well as When the calculated score of the altered sequence indicates an enhancement relative to the input sequence and the iterations are equal to N, identifying the altered sequence as the output sequence of the N iterations; And then After the N iterations, the output sequence is directed to a validation phase to synthesize an output sequence defining the regulatory element.
14. The non-transitory computer-readable storage medium of claim 13, wherein the executable instructions, when executed by the at least one processor to advance the changed sequence to a next iteration, cause the at least one processor to advance the changed sequence to the next iteration based on the calculated score of the changed sequence being greater than the calculated score of the starting sequence and the iteration being less than N.
15. The non-transitory computer-readable storage medium of claim 13 or claim 14, wherein the regulatory element comprises one or more of a promoter, an intron, and / or an untranslated region (UTR).
16. The non-transitory computer-readable storage medium of claim 15, wherein the executable instructions, when executed by the at least one processor to calculate the score for the altered sequence, cause the at least one processor to calculate the score based on:
17. The non-transitory computer-readable storage medium of claim 13, wherein the executable instructions, when executed by the at least one processor to advance the altered sequence to the next iteration, cause the at least one processor to further advance the altered sequence to the next iteration based on a probability function satisfying the threshold and the iteration being less than N; The threshold comprises a static threshold and a threshold randomly generated in each iteration; The probability function includes: as well as wherein E is the calculated score of the starting sequence, E' is the calculated score of the altered sequence, and T is the at least one parameter.
18. The non-transitory computer-readable storage medium of claim 17, wherein the executable instructions, when executed by the at least one processor to calculate the score for the altered sequence, cause the at least one processor to calculate the score based on:
19. The non-transitory computer-readable storage medium of claim 17 or claim 18, wherein N is less than 100.
20. The non-transitory computer readable storage medium of claim 19, wherein the at least one parameter is temperature.
21. A system for identifying a regulatory element, the computer-implemented system comprising: At least one computing device configured by executable instructions to: identifying an input sequence associated with a regulatory element as a starting sequence; Calculating a score for the starting sequence based on a scoring function; Initialize N iterations at at least one parameter, where N is an integer; For each of the N iterations: changing at least one nucleotide in an input sequence to perform said iteration; Calculating a score for the altered sequence based on the scoring function; advancing the altered sequence to a next iteration based at least on the calculated score of the altered sequence and a threshold; as well as When the calculated score of the altered sequence indicates an enhancement relative to the input sequence and the iterations are equal to N, identifying the altered sequence as the output sequence of the N iterations; as well as A validation phase is configured to synthesize the regulatory element defined by the output sequence after the N iterations.
22. A system as claimed in claim 21, wherein the computing device is configured to advance the changed sequence to a next iteration based on the calculated score of the changed sequence being greater than the calculated score of the starting sequence and the iteration being less than N.
23. The system of claim 22, wherein to calculate the score of the altered sequence, the computing device is configured to calculate the score based on:
24. The system of claim 23, wherein the computing device is configured to advance the altered sequence to the next iteration further based on a probability function of satisfying the threshold in order to advance the altered sequence to the next iteration; The threshold comprises a static threshold and a threshold randomly generated in each iteration; The probability function includes: as well as wherein E is the calculated score of the starting sequence, E' is the calculated score of the altered sequence, and T is the at least one parameter.
25. The system of any one of claims 21 to 24, wherein the regulatory element comprises one or more of a promoter, an intron and / or an untranslated region (UTR); wherein N is less than 100; and / or Wherein the at least one parameter is temperature.
Citation Information
Patent Citations
Genetic regulatory elements
US11377663B1
Genetic regulatory elements
US11932862B1
Methods for making genetic regulatory elements
US12188028B1
Synthetic insecticidal crystal protein gene
US5380831A
Synthetic insecticidal gene, plants of the genus oryza transformed with the gene, and production thereof
US5436391A