Method for predicting promoter activity and method for modifying promoter based on prediction result
A machine learning-based method using DNABERT and CRISPR/Cas for promoter sequence modification addresses the challenge of precise gene expression control in plants, enhancing or suppressing gene function by predicting and editing promoter activity.
Patent Information
- Application Number
- JP2024001583
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-10
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2043-12-28
AI Technical Summary
Current methods for genome editing in plants primarily focus on introducing mutations in the coding region, leading to loss-of-function, while methods to enhance or suppress gene expression levels by modifying promoter regions are limited due to ambiguous sequence-function correspondence, making precise control of gene expression challenging.
A machine learning-based approach using DNABERT to predict promoter activity, followed by genome editing techniques like CRISPR/Cas, to generate and select modified promoter sequences achieving desired expression levels, enabling precise control of gene expression.
The method allows for accurate prediction and modification of promoter sequences to enhance or suppress gene expression levels, providing a precise and efficient means to regulate gene function in plants.
Smart Images

Figure 2025105366000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the prediction of promoter activity and the modification of promoters based on the prediction results.
Background Art
[0002] Currently, when imparting traits to plants by genome editing, the loss-of-function method of introducing mutations into the coding region (CDS) and causing gene function to be lost due to frameshift is mainly used.
[0003] On the other hand, methods such as enhancing function by increasing the expression level of the target gene or suppressing the side effects of knockout by leaving a small amount (rather than reducing the expression to zero) can be considered. Since research on enhancing function by introducing transgenes and findings from functional analysis by RNAi can be applied, such methods are expected to greatly expand the variations of traits that can be imparted by genome editing and are important.
Prior Art Documents
Non-Patent Documents
[0004]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] One possible method is to edit the promoter region instead of the CDS. Since regions such as promoters and enhancers determine the gene expression intensity, it is considered that by modifying the nucleotide sequence of this region, the gene function can be kept intact while the expression level can be increased or decreased.
[0006] In fact, the range (dynamic range) of gene expression levels at the mRNA level measured by the RNA-seq method is very large. There are mRNAs that are contained only in a few copies in one cell, while there are also mRNAs that are contained in about 10 5 copies. Such differences in gene expression levels are considered to be mainly caused by promoters and enhancers, indicating the potential of the method of modifying the nucleotide sequence of the promoter region. Therefore, providing a technique for precisely controlling gene expression levels by genome editing is one of the objectives of the present disclosure.
Means for Solving the Problems
[0007] Regions such as promoters and enhancers have an ambiguous sequence-function correspondence of "this part has this function" compared to the CDS. Some Cis regulatory elements with highly conserved consensus sequences, such as the TATA box, have been discovered and are in a database. For example, software such as PromoterCAD of the RIKEN and NEW PLACE of the National Agriculture and Food Research Organization is a system that searches for Cis regulatory elements based on such findings and infers whether a transcription factor binds to a promoter, but its accuracy is insufficient for the purpose of predicting and designing expression levels.
[0008] Therefore, the present inventors developed a method of learning the relationship between nucleotide sequences and expression levels by machine learning instead of the database method of registering individual elements. As a result, the coefficient of determination showing the correlation between the measured value and the predicted value is R 2A high value of 0.75 was obtained. Then, the prediction system was used to design an array, and experiments using plants were conducted to verify the prediction accuracy. If any base sequence can be given to a computer to predict transcriptional activity, the expression level can be "designed". As shown in the present disclosure, the inventors have demonstrated that based on computer prediction, by editing the promoter of a gene, the function (expression level) of the gene can be increased or decreased.
[0009] The present disclosure is based on these findings and includes the following aspects: [Aspect 1] A method for obtaining a promoter sequence modified to have a desired activity, comprising: Preparing an original promoter sequence to be modified; Generating a set of a plurality of modified promoter sequences that can be created by genome editing technology based on the promoter sequence; Predicting the activity of each individual modified promoter sequence included in the generated set of modified promoter sequences using a machine learning model; Selecting a modified promoter sequence predicted to have the desired activity . [Aspect 2] The method according to Aspect 1, wherein the desired activity is a gene expression induction activity higher than that of the original promoter sequence or a gene expression induction activity lower than that of the original promoter sequence. [Aspect 3] The method according to Aspect 1, wherein the promoter sequence is a promoter sequence of a plant cell. [Aspect 4] The method according to Aspect 1, wherein the original promoter sequence includes a core promoter sequence and the sequence upstream thereof. [Aspect 5] The method according to Aspect 1, wherein the genome editing technology is a genome editing technology using the CRISPR / Cas system. [Aspect 6] The method according to embodiment 5, wherein a set of a plurality of modified promoter sequences is generated by a sequence deletion caused by cleavage induced by a combination of guide RNA sequences designed based on two PAM recognition sequences. [Embodiment 7] The method according to embodiment 1, wherein the set of modified promoter sequences includes at least 1000 different sequences. [Embodiment 8] The method according to embodiment 1, wherein the machine learning model is a regression model trained to predict gene expression induction activity from a promoter sequence, using gene expression induction activity data of a plurality of promoter sequences in plant cells as teacher data. [Embodiment 9] The method according to embodiment 1, further comprising a step of visualizing the activity of individual modified promoter sequences included in the generated set of modified promoter sequences on a computer display. [Embodiment 10] An information processing apparatus for predicting a promoter sequence modified to have a desired activity, comprising: a sequence input unit that receives an input of an original promoter sequence to be modified; a modified sequence generation unit that generates a set of a plurality of modified promoter sequences that can be created by a genome editing technique based on the promoter sequence; an activity prediction unit that predicts the activity of individual modified promoter sequences included in the generated set of modified promoter sequences by a machine learning model; a sequence selection unit that selects a modified promoter sequence predicted to have a desired activity An information processing apparatus comprising: [Embodiment 11] A non-transitory computer-readable medium storing instructions that, when executed by a processor, perform the following steps: receiving an input of an original promoter sequence to be modified; generating a set of a plurality of modified promoter sequences that can be created by a genome editing technique based on the promoter sequence; Predicting the activity of each modified promoter array included in the set of the generated modified promoter arrays by a machine learning model, Selecting a modified promoter array predicted to have a desired activity A computer-readable medium capable of executing the above. [Aspect 12] For a computer, A function of receiving an input of an original promoter array to be modified, A function of generating a set of a plurality of modified promoter arrays that can be created by a genome editing technique based on the promoter array, A function of predicting the activity of each modified promoter array included in the set of the generated modified promoter arrays by a machine learning model, A function of selecting a modified promoter array predicted to have a desired activity A program for realizing the above. [Aspect 13] A method for genome editing of cells for regulating the expression level of a desired gene, comprising: Obtaining a modified promoter array having a desired activity by the method according to Aspect 1, Preparing cells to be genome-edited, Editing the genome of the cells so as to generate the modified promoter array A method comprising the above. [Aspect 14] A method for producing a genome-edited plant in which the expression level of a desired gene is regulated, comprising: Editing the genome of a desired plant cell by the method according to Aspect 13, Obtaining a plant individual derived from the genome-edited cells A method comprising the above. [Aspect 15] A machine learning model for predicting gene expression induction activity in plant cells from a promoter array, which is a regression model trained to predict gene expression induction activity from a promoter array using gene expression induction activity data of a plurality of promoter arrays in plant cells as teacher data. [Aspect 16] A method for generating a machine learning model for predicting gene expression induction activity in plant cells from a promoter sequence, the method comprising training a model using gene expression induction activity data of a plurality of promoter sequences in plant cells as training data. [Aspect 17] The method according to Aspect 16, wherein a pre-trained base model based on a transformer is used as the model to be trained.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
DETAILED DESCRIPTION OF THE INVENTION
[0011] In the process of crop breeding, breeders have the desire to enhance or suppress the expression of specific genes. As gene functions are being analyzed, the relationships between genes and traits are being gradually clarified. For example, taking the case of semi-dwarfness in rice, the trait of semi-dwarfness has been a breeding goal in rice, and the semi-dwarf variety "Reimei" was developed in 1966. In semi-dwarf rice, even when fertilizers are applied, the plant height hardly increases, and the yield does not decrease. When fertilizers are applied, normal rice usually grows vertically, but as a result, it is easily lodged by typhoons or strong winds, and the harvest may not be possible or the quality may deteriorate. Therefore, fertilization above a certain level leads to lodging and does not result in an increase in the harvest. In semi-dwarf rice, since it is difficult for the plant height to increase, it is resistant to lodging even when a large amount of fertilizer is applied, and the yield increases. It is now clear that the mutation that gave Reimei semi-dwarfness exists in the gene encoding the G20 oxidase in the gibberellin biosynthesis system. Semi-dwarf crops were a great invention that brought a huge advance to the expansion of food production.
[0012] In this way, the relationships between genes and traits have been gradually revealed one after another, and knowledge has been accumulated. As a result, in order to obtain the target trait, the target gene is often set. For example, when trying to introduce semi-dwarfism into rice, mutants of the above gene (sd1 mutants) can be selected. It is possible to search for naturally occurring mutants, but mutations can also be introduced randomly by chemical or radiation irradiation. Although the loss-of-function allele of sd1 is considered recessive, individuals in which both alleles become the loss-of-function type sd1 can be bred by planned mating.
[0013] With the advent of genetic recombination technology, it has become possible to artificially design genotypes. That is, by introducing an arbitrary gene sequence into the target crop, functions and traits can be imparted. A typical example is corn with herbicide tolerance (Roundup Ready (registered trademark)). Roundup (registered trademark) (glyphosate) is a herbicide and is actually an amino acid analog. In some plants and bacteria, when glyphosate is taken up, amino acid synthesis is inhibited and they die. By introducing a gene that detoxifies glyphosate from bacteria, plants are given tolerance to glyphosate, which was a great invention that greatly reduced the cost of weeding. On the other hand, the production and use of genetically modified plants that introduce foreign genes are limited, and not all crops on Earth will be replaced by genetic modification. Especially in Europe, there has been strong opposition, and even now, the cultivation of genetically modified crops in Europe is greatly restricted, and only one variety (corn in Spain) is being cultivated.
[0014] Since the invention of genome editing, the situation has changed further. It has become possible to directly introduce mutations into target genes without genetic recombination. This can be regarded as a great invention in history. For example, when considering introducing semi-dwarfness into Koshihikari, a small-scale indel of about 1 to 10 bases can be introduced at the SD1 locus to induce a frameshift and cause loss of function. By this method, it can be developed earlier than mutation induction by chemicals or radiation (followed by about 5 to 7 backcrosses), and the expression of unintended traits due to linked drag can be avoided, which has a great advantage in terms of cost. Examples of foods produced by this technique and sold in Japan include, for example, red sea bream sold by Regional Fish Co., Ltd., which lacks the myostatin gene.
[0015] In genome editing, in addition to small-scale indels of about 1 to 10 bases, medium-scale deletions up to several thousand bases can also be induced. In this case, when two guide RNAs are simultaneously delivered into cells and two cleavage sites that cut the target sequences of DNA at two locations are physically close, such as being linked, the region between the two cleavage sites may be deleted and repaired. A specific example is the waxy corn of Corteva Agriscience, which has a deletion of approximately 4 kb, including the entire coding sequence of the Wx locus.
[0016] The mutations in the varieties developed through genome editing, as mentioned above, are of the types that can also occur in nature, and the resulting crops are essentially no different from those bred through conventional mating. In contrast, it is also possible to induce homologous recombination using genome editing, and the target position in the genome can be replaced with any DNA sequence. By applying this, various mutations such as point mutations, deletions, and insertions can be introduced. Furthermore, by replacing with foreign genes, gene recombination can also be induced. Thus, even though it is called genome editing in a general sense, mutants of essentially different categories such as mutations that can occur in nature and gene recombination that cannot occur at all under natural conditions can be created. To avoid confusion, varieties created by genome editing are classified into three categories: SDN1 to SDN3. The above examples of Regional Fish Co., Ltd. and Corteva Agriscience are both classified as SDN1. For genome-edited crops, procedures such as notification to administrative agencies and prior consultation are required in many countries including Japan with the creation of new varieties. In SDN1, since the newly introduced mutations are essentially equivalent to those that occur naturally and there are also accumulated precedents, it is expected that procedures will proceed quickly in many countries, which has industrial advantages.
[0017] By the way, as mentioned above, genome editing can induce frameshift, which can easily create functional knockouts of specific genes. However, for breeders, functional knockouts of genes are not the only useful ones. Mutations accompanied by increased gene expression or mutants accompanied by decreased gene expression have been selected innumerable times, either intentionally or unintentionally. Also, in the history of plant science, a huge number of gene silencing (RNAi, KD) experiments have been reported. The phenotypes of knockout (KO) and knockdown (KD) mutants may differ. There may be cases where slight residual expression avoids harmful phenotypes. Thus, a significant decrease in gene expression levels, mimicking KD mutants, is also often considered useful. The method according to the present disclosure has the advantage of enabling the design of such increases or decreases in gene expression levels using genome editing of SDN1, which can be commercialized only through the process of notification to relevant ministries and agencies in Japan.
[0018] As a method for enhancing function, one possible approach is to introduce mutations into genes. For example, it is thought that the activity can be enhanced by disrupting the autoinhibitory domain of an enzyme. Taking crops sold in Japan as an example, in the high-GABA accumulating tomato of Sanatech Seed Co., Ltd., the autoinhibitory domain present at the C-terminus of the GABA-synthesizing enzyme GAD is disrupted by frameshift, thereby enhancing GABA synthesis activity. This method is very effective, but it can only be used when the target protein (gene) has an autoinhibitory domain. Since proteins with autoinhibitory domains are relatively few in number, the scope of application is narrow. As other methods, there are considerations such as disrupting the protein's localization signal to always keep it in the active form, substituting the amino acid that undergoes phosphorylation with glutamic acid, or introducing mutations into the active site of the enzyme to improve reactivity. However, such techniques are generally not methods that can be used for any gene.
[0019] As a method for enhancing functions, one generally applicable to many genes is to introduce mutations into regions that regulate gene expression levels. That is, the idea is to increase the expression level by editing the promoter, thereby enhancing the function. However, this idea has drawbacks. The drawback is that the correspondence between the sequence and the function is ambiguous even for the promoter. In genes, the correspondence between codons and amino acid sequences is strict, and the functional domains formed by specific amino acid sequences also have high conservation, so the function can be inferred from the sequence. Therefore, it is possible to disrupt the target domain with high accuracy. On the contrary, in the promoter (non-ORF), it is difficult to establish a hypothesis such as "if this part is done in this way, such and such will happen." Of course, there are also relationships between sequences and functions. Cis-regulatory elements typified by the TATA box are understood as binding sites for basal transcription factors and transcription factors, and there have been cases where transcription factors have been inferred from sequences and binding sites have been deduced from transcription factors. For example, web applications such as PromoterCAD of the RIKEN and NEW PLACE of the National Agriculture and Food Research Organization are tools for searching for such elements. While such tools have been proposed, there seem to be limited examples of successful breeding using them. Song et al., 2022 reported an improvement in gene expression by genome editing of the promoter (Song, X., Meng, X., Guo, H. et al. Targeting a gene regulatory element enhances rice grain yield by decoupling panicle number and size. Nat Biotechnol 40, 1403-1411 (2022). https: / / doi.org / 10.1038 / s41587-022-01281-7). This report states that an increase in yield can be obtained by editing the promoter or 5’UTR of the rice IPA1 gene. It follows a procedure of creating lines with various deletion patterns by genome editing, evaluating expression patterns and traits, and finding important regulatory elements. This shows how difficult it is to design functions from regulatory elements. The variability of regulatory elements (typically allowing a few-base mismatches from the consensus sequence) may be one reason for the difficulty. Recently, there have been reports of inferring Cis-regulatory elements by applying deep learning, but there are absolutely no examples of it being linked to actual crop traits.
[0020] Rather than using conventional methods to infer functions from Cis regulatory elements, the inventors adopted a strategy of directly inferring expression levels from DNA sequences. For this purpose, the target was limited to promoters. Although the relationship between sequences and expression levels is a black box, methods using deep learning often exhibit high performance even when the principles and theories are unclear. For example, in the field of machine translation, until the 1990s, methods were explored to construct a dictionary by collecting vocabulary and associating it with functions (semantic information) to convert languages (e.g., from English to German), but high performance was not achieved. In contrast, around 2011, when WATSON, which applied neural networks (NNs), demonstrated high performance in natural language processing, translation with high accuracy was realized without inputting the meaning of individual sentences by learning a vast amount of text data, rather than having humans create a dictionary and make the computer understand the meaning. Today's Google Translate and DeepL are typical examples of machine translation using NNs, both of which provide extremely high utility. Thus, in machine learning using NNs, high performance can sometimes be achieved without humans understanding and organizing the principles and theories.
[0021] DNABERT is a model pre-trained to handle DNA sequences, referring to the natural language model BERT created for natural language processing. DNABERT was developed for the purpose of analyzing non-coding regions based on DNA sequences and has shown high performance in tasks such as the discovery of promoters and splice sites. Therefore, it was decided to use it for the problem of predicting the transcriptional activity of the core promoter in the present invention. However, there were two problems when applying DNABERT to the present invention.
[0022] First, the human genome has been used as pre-trained data, and it was not clear whether the plant genome could be handled by DNABERT. In plants, as core elements that make up the core promoter, in addition to factors common to animals such as the TATA box, Initiator, and Kozak, there may be plant-specific elements such as the Y patch, CA, and GA. Also, in the promoters of genes in animals and plants, transcription activity may be controlled by DNA methylation. In animals, cytosine in the CG (CpG) sequence is mostly methylated. In contrast, in plants, cytosine in non-CG sequences is also methylated. Thus, although the basic mechanism of transcription is common between animals and plants, the components such as core elements and methylation sites are different. Therefore, it was unclear whether DNABERT could perform well in plants with promoters that are significantly different in structure from those of humans.
[0023] Second, although DNABERT has shown high performance in tasks such as the discovery of transcription start points and splice sites, these are tasks of determining whether an element of interest "exists" or "does not exist" in a given sequence, which are binary classification problems. In contrast, the task of the present inventors was to predict the gene expression intensity downstream of the core promoter sequence, and it was required to output continuous values. It was completely unclear whether DNABERT was effective or could perform well in such a regression problem. When dealing with continuous values related to the DNA sequence, it was common to use a CNN. Based on reports that the natural language processing model BERT can be applied to handle regression problems, the present inventors modified the code of DNABERT. Specifically, classes were added to enable handling of continuous values.
[0024] Despite having such problems, we succeeded in fine-tuning the pre-trained model of DNA BERT according to the separately shown procedure and predicting the strength of the promoter with high accuracy. Furthermore, the inventors have also succeeded in developing a program that incorporates this learning model and changes the function of the promoter by genome editing.
[0025] In one aspect, the present disclosure relates to a method for obtaining a promoter sequence modified to have a desired activity. In some embodiments, the method according to the present disclosure includes preparing an original promoter sequence to be modified, generating a set of a plurality of modified promoter sequences that can be created by genome editing technology based on the promoter sequence, predicting the activity of each modified promoter sequence included in the generated set of modified promoter sequences using a machine learning model, and selecting a modified promoter sequence predicted to have the desired activity.
[0026] A promoter is a DNA sequence located (usually) upstream of a gene. It is not transcribed itself but has the function of determining the transcription level (expression level) of the gene and, in multicellular organisms, the site and tissue specificity of expression. In this specification, the region near the gene that interacts with the basal transcription factors is particularly referred to as the "core promoter". The core promoter can be defined as the position from -200 to +50 when the position of the base at the transcription start point is represented as +1. More preferably, it can be defined as the position from -180 to +25. Even more preferably, it can be defined as the position from -170 to +17. Most preferably, the position from -165 to +5 can be defined as the core promoter. The core promoter is considered to have the function of determining the basic transcription level.
[0027] In some embodiments, the promoter sequence is a promoter sequence of a plant cell. It is known that the types of core elements included in the core promoter are different between animals and plants. In this specification, "plant" is not particularly limited. For example, it can include a wide range of plants such as bryophytes, ferns, gymnosperms, magnolias of angiosperms, monocotyledons, and eudicots (Rosidae I, Rosidae II, Asteridae I, Asteridae II, and their outgroups). More specific examples of plants include solanaceous plants such as tomato, pepper, chili pepper, eggplant, tobacco, and torvum; cucurbitaceous plants such as cucumber, pumpkin, melon, and watermelon; cruciferous plants such as cabbage, broccoli, Chinese cabbage, and kale; leafy and aromatic plants such as perilla, celery, parsley, and lettuce; alliaceous plants such as leek, onion, and garlic; other fruit vegetables such as strawberry and melon; taproot plants such as radish, turnip, carrot, and burdock; tuber plants such as taro, cassava, potato, sweet potato, and yam; cereals such as rice, corn, wheat, sorghum, barley, rye, wild rice, and buckwheat; leguminous plants such as soybean, adzuki bean, mung bean, kidney bean, lima bean, peanut, pea, and broad bean; leafy vegetables such as asparagus, spinach, and clover; flower crops such as turkey lily, rose, stock, carnation, and chrysanthemum; turfgrasses such as bentgrass and Korean lawn grass; oil crops such as rapeseed, camelina, oilseed rape, Jatropha curcas, sesame, and perilla; fiber crops such as cotton, rush, and hemp; forage crops such as clover, dent corn, and taro palm; deciduous fruit trees such as apple, pear, grape, and peach; citrus fruits such as satsuma mandarin, orange, lemon, and grapefruit; woody plants such as azalea, rhododendron, cedar, poplar, and gutta-percha tree, etc. Also, the promoter may be either tissue-specific or non-tissue-specific. A tissue-specific promoter can, for example, control specific expression in leaves or roots. In some embodiments, the desired activity required for the modified promoter can be a gene expression induction activity higher than that of the original promoter sequence or a gene expression induction activity lower than that of the original promoter sequence. The gene expression induction activity higher than that of the original promoter sequence can be, for example, any of 1.1-fold to 1000-fold of the original activity, such as 1.1-fold, 1.2-fold, 1.5-fold, 2-fold, 3-fold, 4-fold, 5-fold, 10-fold, 20-fold, 50-fold, 100-fold, 200-fold, 300-fold, 400-fold, 500-fold, 600-fold, 700-fold, 800-fold, 900-fold, or 1000-fold. The gene expression induction activity lower than that of the original promoter sequence can be, for example, any of 90% to 0.01% of the original activity, such as 90%, 80%, 50%, 10%, 5%, 1%, 0.5%, 0.1%, 0.05%, 0.04%, 0.03%, 0.02%, or 0.01%.
[0028] In some embodiments, the desired activity required for the modified promoter may be determined by comparison with the activity of another reference promoter sequence, rather than the original promoter sequence.
[0029] Alternatively, the desired activity required for the modified promoter may be determined by stratification such as high expression, medium expression, and low expression. In some embodiments, any number or range of stratifications can be performed.
[0030] In some embodiments, the modification of the promoter sequence can be due to deletion, substitution, or insertion of a part of the sequence. In some embodiments, the modification of the promoter sequence can be a deletion of a part of the sequence by cleavage at two positions in the promoter sequence.
[0031] In some embodiments, the original promoter sequence to be modified can be prepared by obtaining it from a database such as GenBank. As the promoter sequence, for example, any DNA region within the DNA region of -9900 to +100, preferably -4950 to +50 or -2975 to +25, more preferably -1995 to +5 around the transcription start point of the target gene, including the transcription start point, can be mentioned. Alternatively, any DNA region between the transcription start point of the target gene and the transcription termination point of the adjacent gene can be mentioned. That is, as the promoter sequence, a region including the core promoter and the sequence upstream thereof can be used. The length of the sequence upstream of the core promoter can be, for example, at least 200 bp, 400 bp, 600 bp, 800 bp, 1000 bp, 1200 bp, 1400 bp, 1600 bp, 1800 bp, 2000 bp, 3000 bp, 4000 bp, 5000 bp, 6000 bp, 7000 bp, 8000 bp, or 9000 bp. When there are multiple transcription start points, the multiple transcription start points can be selected. Also, for example, any range of 100 to 1800 bases included in -1995 to +5 around the transcription start point may be used. The length of the promoter (base pair number: bp) can be, for example, at least 200 bp, 400 bp, 600 bp, 800 bp, 1000 bp, 1200 bp, 1500 bp, 2000 bp, 3000 bp, 4000 bp, 5000 bp, 6000 bp, 7000 bp, 8000 bp, or 9000 bp, and the range can be a length in the range of 100 to 10000 bp, preferably 200 to 5000 bp, more preferably 500 to 3000 bp.
[0032] In some embodiments, to select a modified promoter sequence having a desired activity, a set of multiple modified promoter sequences that can be created by genome editing technology based on the promoter sequence is generated.
[0033] As genome editing technologies, ZFNs, TALENs, and CRISPR-Cas systems can be used. The CRISPR-Cas system includes Class 1 Type I CRISPR-Cas3, Class 2 Type II Cas9, Class 2 Type V Cas12a, Class 2 Type V Cas12f (Cas14a), Class 2 Type VI Cas13a, etc., regardless of their functional mechanisms and classifications. In addition, Cas proteins derived from various bacteria can be used. For example, for Cas9, SpCas9 derived from Streptococcus pyogenes, SaCas9 derived from Staphylococcus aureus, FnCas9 derived from Francisella novicida, CjCas9 derived from Campylobacter jejuni; for Cas12a, AsCas12a (AsCpf1) derived from Acidaminococcus sp., LbCas12a (LbCpf1) derived from Lachnospiraceae bacterium, ErCas12a derived from Eubacterium rectale, etc. are included, without limiting their origins. Also, those with modified nucleotide sequences encoding Cas proteins or amino acid sequences of Cas proteins, those fused with other proteins or functional domains or peptides or amino acid sequences, and those modified with compounds can also be used.
[0034] In the CRISPR-Cas system, Cas nuclease and guide RNA (gRNA) are used for cleavage of target sequences. Guide RNA consists of two elements, crRNA and tracrRNA, which may exist as individual RNAs or be ligated to form single-stranded RNA. Single-stranded guide RNA is also called sgRNA. Nucleotide sequences of 1 to 10 bases, 10 to 50 bases, 50 to 100 bases, or 100 to 500 bases may be added to the 5'-end and 3'-end of the guide RNA. The guide RNA binds complementarily to a specific DNA sequence and guides the Cas nuclease to a specific position in the genome. Thereby, the Cas nuclease cleaves the specific DNA sequence, enabling genome editing. For example, when using CRISPR, the PAM sequence in the promoter sequence is searched to list a list of cleavable positions. The PAM sequence depends on the Cas protein used. For example, in the case of SpCas9, it is NGG; in the case of SaCas9, it is NGRRT or NGNRRN; in the case of NmeCas9, it is NNNNGATT; in the case of CjCas9, it is NNNNRYAC; in the case of LbCas12a (Cpf1), it is TTTV; in the case of AsCas12a (Cpf1), it is TTTV; in the case of AacCas12a (Cpf1), it is TTN; in the case of BhCas12b v4, it is ATTN or TTTN or GTTN, etc. However, if the sequence can be recognized by the Cas protein, these PAM sequences may contain mismatches of 1 to 8 bases. The presence of the PAM sequence is important for enhancing the accuracy and specificity of genome editing of the CRISPR system. In the absence of the PAM sequence, the Cas nuclease cannot bind to the DNA sequence specified by the guide RNA. Due to this property, it becomes possible to accurately target a specific gene region for genome editing. Also, as described above, by using different types of Cas nucleases, regions with different PAM sequences can be targeted. For example, typically, for Cas9, the cleavage position is at 3 to 4 bases from the PAM sequence, and for Cas12a, the cleavage position is at 18 to 23 bases from the PAM sequence. A person skilled in the art can design an appropriate guide RNA based on the PAM sequence in the target genome.Cleavage sites are typically found about 50 times among 2000 bases. For example, among n cleavage sites, the number of combinations specifying 2 sites is nC2. That is, for 50 cleavage sites, there can be 1250 combinations. Thus, in some embodiments, to select a modified promoter sequence having a desired activity, for the nC2 combinations of PAM sequences in the promoter sequence to be modified, a set of base sequences where the bases between two cleavage sites are truncated (deleted) can be generated. In this way, in some embodiments, a set of multiple modified promoter sequences is generated by sequence deletions caused by cleavage induced by a combination of guide RNA sequences designed based on two PAM recognition sequences.
[0035] Also, as other promoter editing methods, a method of deleting between two nicks generated by Cas9 nickase, a method using base editing to substitute a targeted base on the genome with another base, and a method using prime editor capable of deleting, inserting, or substituting any base sequence consisting of 1 to 20 bases or 20 to 50 bases or 50 to 100 bases into the genome can also be used.
[0036] The number of different sequences included in the set of modified promoter sequences can be at least 5, 10, 20, 50, 100, 150, 200, 300, 500, 700, 1000, 1200, 1500, 2000, 3000, 4000, or 5000. For sufficient exploration of sequences affecting activity, the number of different sequences included in the set of modified promoter sequences is preferably 1000 or more. Also, to ensure a sufficient number of different sequences included in the set of modified promoter sequences, the length of the original promoter sequence is preferably, for example, 1800 bp or more.
[0037] After a set of multiple modified promoter arrays is generated, in some embodiments, the activity of each individual modified promoter array included in the generated set of modified promoter arrays is predicted by a machine learning model. For example, for each generated nucleotide sequence, 170 nucleotides on the 3'-end side are obtained and input into the machine learning model. In some embodiments, 100 to 250 nucleotides on the 3'-end side may be obtained and input into the machine learning model. Also, in some embodiments, some sequences included in the 170 nucleotides on the 3'-end side, for example, 100 to 169 nucleotides, may be input. A properly trained machine learning model can output the strength (transcriptional activity) of a given nucleotide sequence as a core promoter.
[0038] In some embodiments, the activity of each individual modified promoter array included in the generated set of modified promoter arrays is visualized on a computer display. The visualization can be performed, for example, as shown in FIG. 8 or FIG. 9.
[0039] Then, by selecting a modified promoter array predicted to have a desired activity from the prediction results of the machine learning model, a promoter array modified to have the desired activity can be obtained.
[0040] In some embodiments, the machine learning model can be a regression model trained to predict gene expression induction activity from a promoter sequence, using gene expression induction activity data of a plurality of promoter sequences in plant cells as training data. In some embodiments, the data of Jores et al. (2021) can be used for training (Jores, T., Tonnies, J., Wrightsman, T. et al. Synthetic promoter designs enabled by a comprehensive analysis of plant core promoters. Nat. Plants 7, 842-855 (2021). https: / / doi.org / 10.1038 / s41477-021-00932-y). Thus, one aspect according to the present disclosure is a method for generating a machine learning model for predicting gene expression induction activity in plant cells from a promoter sequence, the method including training the model using gene expression induction activity data of a plurality of promoter sequences in plant cells as training data. In some embodiments, a pre-trained base model based on a transformer can be used as the model to be trained.
[0041] In some embodiments, in constructing a machine learning model, deep learning models such as transformers, especially models based on BERT, can be used. BERT (Bidirectional Encoder Representations from Transformers) is one of the machine learning models widely used in natural language processing (NLP). It was developed by Google in 2018 and has shown remarkable performance in text understanding and generation. The main features of BERT include: 1) bidirectional context understanding, 2) utilization of the Transformer architecture, 3) pre-training and fine-tuning, 4) diverse application possibilities, etc. First, BERT is a "bidirectional" model that understands words within a given text in both left and right contexts. While conventional models only consider context in one direction (left to right or vice versa), BERT can comprehensively understand the entire text. Also, BERT is built based on a neural network architecture called Transformer. Transformer uses an "attention mechanism" to capture the relationships between each word in the input text. This enables more complex and refined text understanding. Furthermore, BERT is "pre-trained" with a large amount of text data and has acquired general language understanding. Subsequently, it can be optimized for specific applications by performing "fine-tuning" for specific tasks (such as sentiment analysis and question answering). BERT is applicable to various NLP tasks and is used in a wide range of fields, such as text classification, question answering, sentiment analysis, machine translation, etc.
[0042] Two main tasks are mainly used for training the BERT model. The MLM task is to randomly "mask" words from the input text and let BERT predict the masked words. The main purpose of this task is to let BERT understand the meaning of words by using context. To consider bidirectional context, the model utilizes the information of the entire sentence to predict the masked words. The NSP task is to let BERT determine whether two sentences are consecutive. The model is given a sentence (A) and another sentence (B), and it is required to predict whether B comes immediately after A. At this time, with a probability of half, B is the actual sentence following A, and the other half is an unrelated sentence randomly selected. The purpose of the NSP task is to let BERT understand the relationship between sentences. This ability is particularly useful in tasks where it is important to understand the connection and semantic flow of sentences, such as question answering and text summarization. Through these tasks, BERT cultivates the ability to understand not only at the word level but also the relationships between entire sentences and multiple sentences. As a result, BERT can achieve high performance in various natural language processing tasks.
[0043] In some embodiments, for machine learning, DNABERT, a pre-trained base model based on transformers, can be utilized (Yanrong Ji, Zhihan Zhou, Han Liu, Ramana V Davuluri, DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome, Bioinformatics, Volume 37, Issue 15, August 2021, Pages 2112-2120). DNABERT is a pre-trained bidirectional encoder representation that can capture a global understanding of genomic DNA sequences based on the context of upstream and downstream base sequences. In DNABERT, the NSP task used in BERT is not performed, and learning is carried out only with the MLM task. A certain ratio in the DNA sequence is masked, and the k-mer tokens at the masked sites are predicted. The training data used for the pre-training of DNABERT are DNA sequences sampled from the human genome. The inventors have demonstrated that by fine-tuning a pre-trained base model such as DNABERT, a machine learning model can be generated to predict the gene expression induction activity in plant cells from promoter sequences. Note that the machine learning model according to the present disclosure can handle not only binary classification problems but also regression problems, and as a result, the level of gene expression induction activity can be predicted.
[0044] One aspect of the present disclosure relates to a method for genome editing of cells for regulating the expression level of a desired gene. In some embodiments, the method can be a method comprising obtaining a modified promoter sequence having a desired activity by a method for obtaining a promoter sequence modified to have a desired activity according to the present disclosure, preparing a cell to be the subject of genome editing, and editing the genome of the cell to generate the modified promoter sequence.
[0045] In some embodiments, the cells can include plant cells, such as cells from mosses, ferns, gymnosperms, angiosperms including magnolias, monocots, and eudicots (Rosidae I, Rosidae II, Asteridae I, Asteridae II, and their outgroups). More specific examples of plants can include nightshades such as tomato, pepper, chili pepper, eggplant, tobacco, and Solanum torvum; cucurbits such as cucumber, pumpkin, melon, and watermelon; brassicas such as cabbage, broccoli, Chinese cabbage, and kale; leafy greens and herbs such as perilla, celery, parsley, and lettuce; alliums such as leek, onion, and garlic; other fruit vegetables such as strawberry and melon; root vegetables such as daikon radish, turnip, carrot, and burdock; tubers such as taro, cassava, potato (Solanum tuberosum), sweet potato, and yam; cereals such as rice, maize, wheat, sorghum, barley, rye, wild rice, and buckwheat; legumes such as soybean, adzuki bean, mung bean, cowpea, kidney bean, peanut, pea, and broad bean; leafy vegetables such as asparagus, spinach, and Japanese butterbur; flowers such as turkey mullein, rose, stock, carnation, and chrysanthemum; turfgrasses such as Kentucky bluegrass and Japanese lawngrass; oil crops such as rapeseed, camelina, oilseed rape, Jatropha curcas, sesame, and perilla; fiber crops such as cotton, hemp, and flax; forage crops such as clover, sorghum-sudangrass, and oil palm; deciduous fruit trees such as apple, pear, grape, and peach; citrus fruits such as satsuma mandarin, orange, lemon, and grapefruit; woody plants such as azalea, rhododendron, cedar, poplar, and gutta-percha tree, etc.
[0046] One aspect of the present disclosure relates to a polynucleotide having promoter activity, having a sequence obtained by the method according to the present disclosure. Such a polynucleotide can have, for example, the sequence of SEQ ID NO: 2 or 4.
[0047] One aspect of the present disclosure relates to a method for producing a genome-edited plant in which the expression level of a desired gene is regulated. In some embodiments, the method may include editing the genome of a desired plant cell by a method for genome editing of a cell for regulating the expression level of a desired gene according to the present disclosure, and obtaining a plant individual derived from the genome-edited cell. Further, one aspect of the present disclosure relates to a genome-edited plant produced by the production method according to the present disclosure.
[0048] When obtaining a genome-edited cell, a genome editing enzyme can be introduced into a plant tissue or cell in any form of DNA, RNP, or protein. Examples of the introduction site of the genome editing enzyme in a plant include a flower (egg cell, pollen, petal, etc.), a stem (cambium, pith, cortex, etc.), a leaf (including leaf primordium), a root, a shoot apex, a lateral bud, a flower bud, a root tip, a protoplast, and the like.
[0049] The introduction method is not particularly limited and can be appropriately selected according to the plant species to be introduced and the target cell / tissue to be introduced. Examples of the introduction method include the Agrobacterium method, the particle gun method, the whisker method, the nanopipette method, virus-mediated nucleic acid delivery, and the like.
[0050] In some embodiments, when obtaining a plant individual derived from a genome-edited cell, for example, a step of culturing the genome-edited cell or a tissue containing the edited cell under appropriate culture conditions (a), a step of growing the cells to form a callus (undifferentiated cell mass) (b), a step of inducing a shoot (new bud) from the callus (c1), a step of directly forming a shoot from a tissue containing the edited cell (c2), a step of inducing roots on the shoot to regenerate a plant body (d), a step of obtaining seeds of the next generation from the regenerated plant body (e1), or a step of propagating a part of the regenerated plant body by cutting (e2), etc. may be included.
[0051] One aspect of the present disclosure relates to an information processing apparatus for predicting a promoter sequence modified to have a desired activity (see FIG. 11).
[0052] In some embodiments, the apparatus 010 may be an information processing device including an array input unit 011 that receives an input of an original promoter sequence to be modified, a modified sequence generation unit 012 that generates a set of a plurality of modified promoter sequences that can be created by a genome editing technique based on the promoter sequence, an activity prediction unit 013 that predicts the activity of each modified promoter sequence included in the generated set of modified promoter sequences by a machine learning model, and an array selection unit 014 that selects a modified promoter sequence predicted to have a desired activity. The array input unit 011 may be connected to, for example, a communication device with an external network or an input device such as a keyboard. The array selection unit 014 may be connected to, for example, an output device such as a display or a printer, a communication device with an external network, or the like. Those skilled in the art will be able to understand the configuration necessary for the apparatus to implement the array acquisition method according to the present disclosure in light of the disclosure of this specification.
[0053] One aspect of the present disclosure relates to a non-transitory computer-readable medium storing instructions or a program.
[0054] In some embodiments, the computer-readable medium may be a computer-readable medium storing instructions or a program that, when executed by a processor, can perform the following steps: a step S100 of receiving an input of an original promoter sequence to be modified, a step S110 of generating a set of a plurality of modified promoter sequences that can be created by a genome editing technique based on the promoter sequence, a step S120 of predicting the activity of each modified promoter sequence included in the generated set of modified promoter sequences by a machine learning model, and a step S130 of selecting a modified promoter sequence predicted to have a desired activity. FIG. 12 shows an exemplary flowchart of such a program.
[0055] Step S110 of generating a set of a plurality of modified promoter sequences that can be created by genome editing technology may further include step S112 of comprehensively searching for cleavable positions (PAM sequences) by Cas9 from the input sequences, step S114 of combinatorially selecting all two of the searched cleavable positions (PAM sequences), and step S116 of generating a set of modified promoter sequences by deleting the sequences between the two selected cleavage sites for all combinations. Further, step S120 of predicting the activity of individual modified promoter sequences by a machine learning model may be performed using only the core promoter portion (a 150-200 bp sequence, for example, about 170 bp sequence on the 3'-end side) of the individual modified promoter sequences.
[0056] One aspect of the present disclosure relates to a computer program for predicting a promoter sequence modified to have a desired activity.
[0057] In some embodiments, the program can be a program that enables a computer to implement functions of receiving an input of an original promoter sequence to be modified, generating a set of a plurality of modified promoter sequences that can be created by genome editing technology based on the promoter sequence, predicting the activity of individual modified promoter sequences included in the generated set of modified promoter sequences by a machine learning model, and selecting a modified promoter sequence predicted to have a desired activity.
[0058] FIG. 13 is a schematic diagram showing an exemplary aspect of the device of the present disclosure. In FIG. 13, reference numeral 100 denotes a computer, which includes a control unit 101, a storage unit 102, a peripheral device I / F unit 103, an input unit 104, a display unit 105, and a communication unit 106, and these are connected by a bus 110. Note that this configuration is an example, and various configurations can be adopted as appropriate.
[0059] The control unit 101 is composed of a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), etc. The CPU calls a program stored in the storage unit 102, the ROM, a recording medium, etc. to the work memory area on the RAM and executes it, drives and controls each device connected via the bus 110, and realizes the processing performed by the computer. The ROM is a non-volatile memory and holds programs such as the boot program of the computer 100 and BIOS, and data. The RAM is a volatile memory, temporarily holds programs, data, etc. loaded from the storage unit 102, the ROM, a recording medium, etc., and has a work area used when the control unit 101 performs various processes. The storage unit 102 is, for example, an HDD (Hard Disk Drive) and stores programs executed by the control unit 101 and other various data.
[0060] The peripheral device I / F (Interface) unit 103 is a port for connecting the computer 100 and peripheral devices. The peripheral device I / F unit 103 is composed of USB, IEEE1394, RS-232C, etc. Note that the connection form with peripheral devices can be either wired or wireless. The input unit 104 has input devices such as a keyboard, a pointing device such as a mouse, and a numeric keypad, and gives operation instructions, operation commands, data input, etc. to the computer 100. The display unit 105 is a logic circuit or device driver for performing display of video, images, etc. on a display device such as a liquid crystal panel. The input unit 104 and the display unit 105 can also be integrally configured as a touch display.
[0061] The communication unit 106 has a communication control device, a communication port, etc., and is a wired or wireless communication interface that mediates communication with the network 120. The bus 110 is a communication path that mediates the exchange of control signals, data signals, etc. between each device. The network 120 may be further connected to an external server 130 or a network storage 140.
[0062] For example, the device in FIG. 13 can be made to function as an information processing device by reading a program recorded on a computer-readable medium according to the present disclosure, and including an array input unit that receives an input of the original promoter sequence to be modified, a modified sequence generation unit that generates a set of a plurality of modified promoter sequences that can be created by a genome editing technique based on the promoter sequence, an activity prediction unit that predicts the activity of each individual modified promoter sequence included in the generated set of modified promoter sequences by a machine learning model, and a sequence selection unit that selects a modified promoter sequence predicted to have a desired activity.
[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present invention, some possible, preferred methods and materials will now be described. All publications mentioned herein are incorporated herein by reference, and the methods and / or materials related to the cited publications are disclosed and described. It is understood that the present disclosure takes precedence over the disclosure of incorporated publications in case of conflict.
[0064] When a range of values is recited, it is to be understood that, unless the context clearly dictates otherwise, each intervening value between the upper and lower limits of that range, including the lower limit unit incremented by one-tenth thereof, is also specifically disclosed. Each smaller range between any of the recited values or intervening values in the recited range and any other recited value or intervening value in that recited range is also contemplated by the present disclosure. The upper and lower limits of these smaller ranges may independently be included in or excluded from the range, and each range including any one, both, or neither of the limits of the smaller range is also contemplated by the present invention, provided that any specifically excluded limit values in the recited range are excluded. When the recited range includes one or both of the limit values, ranges excluding either or both of the included limit values are also included in the present invention. The term "about" as used in connection with a numerical value means within 5% thereof.
[0065] The embodiments described in this specification are intended to be merely exemplary, and those skilled in the art will be able to make numerous variations and modifications without departing from the spirit of the invention. Also, certain variations and modifications, although they may not result in the optimal result, may still yield satisfactory results. All such variations and modifications are intended to be within the scope of the invention as defined by the appended claims. Also, any combination of the components disclosed in this specification, and any conversion of the expressions of this disclosure between methods, apparatuses, systems, computer programs, data structures, recording media, etc., are also valid as aspects related to this disclosure. Therefore, the details described with respect to the method of this disclosure can be applied to systems, computer programs, data structures, recording media, etc.
Examples
[0066] Example 1: Construction of a Machine Learning Model A machine learning model for predicting gene expression induction activity in plant cells from promoter sequences was constructed as follows. First, as the data for learning, the data shown in the paper by Jores et al. was obtained (Jores, T., Tonnies, J., Wrightsman, T. et al. Synthetic promoter designs enabled by a comprehensive analysis of plant core promoters. Nat. Plants 7, 842-855 (2021). https: / / doi.org / 10.1038 / s41477-021-00932-y). The overview of the dataset was a dataset of about 70,000, which was a combination of a 170-base DNA sequence and expression intensity.
[0067] Also, for machine learning, DNABERT, a pre-trained base model based on transformers, was used. Regarding DNABERT, it is reported, for example, in the paper by Ji et al. mentioned above. DNABERT mainly focuses on handling binary classification problems, and it has been shown that, for example, it can make a determination such as "whether a given base sequence has a splicing site or not". A binary classification problem is a problem of classifying into two categories, such as 1 if it has a splicing site and 0 if it does not. In contrast, a problem such as "predicting the downstream gene expression level for a given base sequence" is called a regression problem as opposed to a classification problem and needs to output a continuous value. Therefore, in order to handle continuous values with DNABERT, a class was added to the transformer and the model was trained.
[0068] Example 2: Prediction and Classification of Promoter Strength by a Trained Model Promoter sequences are often present upstream of genes and control the transcription level (expression level) of the gene directly below. It is considered that the expression level of a gene is controlled by the sequence of the promoter. The expression level of the downstream gene is defined as the strength of the promoter. It is considered that each promoter sequence shows a unique strength.
[0069] As described in Example 1, a machine learning model for predicting transcriptional activity was constructed based on the nucleotide sequence. Using this model, a list was created to evaluate the transcriptional activities of all genes in Arabidopsis thaliana.
[0070] All promoters were classified into the following three groups based on the predicted transcriptional activity: 1. "High-expression group" predicted to show high promoter strength; 2. "Low-expression group" predicted to have very low or almost no promoter strength; and 3. "Medium group" showing intermediate promoter strength.
[0071] Example 3: Comparison between Model Prediction and Actual Gene Expression Seven promoters were selected from the high-expression group, eight promoters were selected from the medium group, and four promoters were selected from the low-expression group, and synthetic genes connected to the luciferase (LUC) gene were created.
[0072] The structure of the gene is as shown in FIG. 1. The promoters have 19 different sequence patterns. The sequences of the downstream LUC gene and the upstream cauliflower mosaic virus (CaMV) 35S enhancer are common to each artificial gene.
[0073] The plasmid vector carrying this artificial gene was transfected into protoplasts collected from Arabidopsis thaliana leaves by the polyethylene glycol method. The protoplasts expressed the LUC gene at various expression levels. To evaluate the LUC expression level, the luminescence intensity was measured using a plate reader. The relationship between the measurement results and the predicted promoter strength is shown in FIG. 2. The relationship between the measured value of the expression level of the LUC gene and the predicted value of the promoter strength is shown as a scatter plot. The vertical axis represents the predicted transcriptional activity. The higher the value, the greater the expected expression level of the downstream gene. The horizontal axis represents the measured value of the logarithm-transformed LUC expression intensity, normalized by the expression level of the CaMV promoter, which is a positive control (P control).
[0074] From the results shown in Figure 2, it can be seen that the measured values are divided into three groups, consistent with the groups based on the predicted values. Thus, a clear correlation was confirmed between the predicted values and the measured values by the constructed machine learning model. This figure shows high predictability. In this figure, when converging linearly, it means that the predicted values and the measured values are in agreement, indicating high prediction performance. When evaluating this prediction performance with the correlation coefficient R, R = 0.9296673 was obtained.
[0075] Next, based on the correlation shown in the above scatter plot, the LUC expression level corresponding to the predicted transcriptional activity was calculated by linear regression prediction. A comparison of these is shown in the following Figure 3.
[0076] The luminescence intensity of LUC was plotted on the vertical axis. However, it is shown as a relative value with the luminescence intensity of the P control set to 1. As the P control, a cauliflower mosaic virus (CaMV) promoter with known high transcriptional activity was used. The numbers of 19 types of promoters were shown on the horizontal axis. Promoter numbers 3 - 10 were classified into the "high expression" group with high predicted expression intensity, promoter numbers 13 - 21 into the same "medium" group, and promoter numbers 22 - 25 into the same "low expression" group. The black bar graph indicates the measured value (expression level) of the LUC assay. The white bar graph indicates the predicted value of the expression intensity by the machine learning model. In Figure 3, focusing on the black bar graph, the "high expression" group showed a relatively high expression level, and the "low expression" group showed a low expression level. The "medium" group showed intermediate values.
[0077] Next, for each individual promoter, it was confirmed how accurately the expression level was predicted. Comparing the black and white bars of each promoter, it can be seen that promoters 3, 6, 10, and 18 were predicted with a high degree of consistency. On the other hand, for promoter 4, the measured value > predicted value. For promoter 5, conversely, the predicted value > measured value.
[0078] In summary, the constructed model was able to predict general trends such as high expression and low expression. For low-expression promoters, predictions were generally made with high accuracy. In the high-expression group, expression could be predicted within the range of 50% to 200% of the measured values.
[0079] Thus, the machine learning model according to the present disclosure can predict gene expression levels based on the base sequence of the promoter.
[0080] Example 4: Prediction and Demonstration of a Promoter with Reduced Activity Next, a method for reducing the transcriptional activity of the promoter was devised. Many Cas9 PAM recognition sequences (NGG) are present in each promoter. These sequences were searched, and approximately several thousand possible editing patterns were listed.
[0081] The more specific procedure is as follows: 1. Comprehensively search for positions (PAM sequences) that can be cleaved by Cas9 from the given base sequence; 2. Select two of them; 3. Cut at the two cleavage sites and simulate the new core promoter sequence that appears after repair; 4. Infer the strength as a core promoter for the simulated sequence; 5. Execute the above steps for all possible pairs of cleavage sites; 6. Select the one that shows the lowest score (the highest score when selecting a promoter with increased transcriptional activity) among all the simulated core promoters.
[0082] As described above, when it is desired to enhance / suppress the expression level, it is possible to determine which cleavage site should be selected. In this way, by inferring the transcriptional activity of the promoter for the nucleotide sequence of each editing pattern, an editing pattern capable of reducing the transcriptional activity of the promoter was searched for. The results are shown in Fig. 4. For promoter numbers 3, 4, 5, 6, 9, 10, and 21, an editing pattern capable of reducing the promoter strength was searched for. The predicted value of the score of the LUC assay, calculated based on the predicted new promoter strength, is shown by the black triangular plots. As a result, it was predicted that the transcriptional activity would be such that the LUC expression level would be about 14% to 1% compared to the original promoter.
[0083] Next, genes having these theoretical promoter sequences were artificially synthesized, transfected into protoplasts in the same manner, and the expression level was measured with a plate reader. The results are shown in Fig. 5. Overlaid on Fig. 4, the measured values of the LUC expression level of the promoters of the new sequences are shown as hatched bar graphs. Since the expression level was extremely low, the vertical axis is shown enlarged. For promoter number 5, there was a missing value. For the 6 promoters for which measured values were obtained, a significant decrease in the expression level was observed in all cases, which was consistent with the predicted values. For example, the sequence of promoter number 3 is CGGAAACTTGTCACTTCCTTTACATTTGAGTTTCCAACACCTAATCACGACAACAATCATATAGCTCTCGCATACAAACAAACATATGCATGTATTCTTACACGTGAACTCCATGCAAGTCTCTTTTCTCACCTATAAATACCAACCACACCTTCACCACATTCTTCACT (SEQ ID NO: 1) and the predicted value of the transcriptional activity is 5.43, and the LUC expression level predicted by linear regression from this value was 111.5% compared to the P control. Also, the measured value of the LUC expression level was 68.6%. In contrast, the sequence of promoter number 3 after editing is GAAACTGATTAGCTCCTATCAGTTCAGCAAACCACAAGCTGAAGAATCCAAGACTTGAGAAACAAATTTACAAAAGCCCATGTTCCAATCAAAACTGTTACCAAACATCTGAAATAGATCTAAATGAGCGTTGGTATAATTGAAACTTACCGAAGGCCCACATTCTTCAC (SEQ ID NO: 2) It occurs when 845 bases between -852 and -7 are truncated. The predicted value of transcriptional activity is 0.393, and the LUC expression level predicted by linear regression from this value was 7.98% relative to the P control. Also, the measured value of the LUC expression level was 3.41% relative to the P control.
[0084] Thus, by editing the base sequence of the promoter based on the prediction of the machine learning model according to the present disclosure, the expression can be decreased.
[0085] Example 5: Prediction and Demonstration of a Promoter with Increased Activity Finally, a method for increasing promoter activity was devised. Similar to the case of Example 4, among the base sequences of the promoters of conceivable editing patterns, those with an increased predicted value of promoter strength after editing were searched for. The results are shown in FIG. 6. For promoter numbers 13, 17, 22, 23, 24, and 25, those with increased promoter strength were selected, and the predicted expression levels are shown as white circles. It was predicted that the LUC expression level would be 7.4 to 125 times that of the original promoter.
[0086] Next, genes having these theoretical promoter sequences were artificially synthesized, transfected into protoplasts in the same manner, and the expression levels were measured with a plate reader. The results are shown in the following FIG. 7. Among the 6 promoters, the expression levels increased in 5 compared to the original promoter, but only promoter No. 13 had an expression level increase exceeding the predicted value. For example, the sequence of promoter No. 13 is TCAAGCAATCATTATCGACTACGGTCGTTCGTTAAAGATCATGCATGTGCTTAGTGGCAATACCCTACGCATCTTGATTCGTTACTGCGGCACGTGTCATGACCATGCACATGAATGATGATTAATGTTTAGTACATATAATGTTCACGCAAACGCATAGTGTTAGGAAA(SEQ ID NO: 3) It was as follows, the predicted value of transcriptional activity was 2.00, and the LUC expression level predicted by linear regression from this value was 6.94% compared to the P control. Also, the actually measured value of the LUC expression level was 6.60% compared to the P control. In contrast, for promoter No. 13, the edited sequence was GAAACTTGAAAATCAAATCAGTGAGTCGCAAGTAAGACTTTGTGGTTGTTGTATCAGATTTCGCCGTGCGCATCTTGATTCGTTACTGCGGCACGTGTCATGACCATGCACATGAATGATGATTAATGTTTAGTACATATAATGTTCACGCAAACGCATAGTGTTAGGAA(SEQ ID NO: 4) It was as follows, the predicted value of transcriptional activity was 4.18, and the LUC expression level predicted by linear regression from this value was 36.7% compared to the P control. Also, the actually measured value of the LUC expression level was 69.5% compared to the P control.
[0087] Thus, by editing the base sequence of the promoter based on the prediction of the machine learning model according to the present disclosure, the expression can be increased.
[0088] Example 6: Program for Changing the Function of a Promoter by Genome Editing The inventors developed a program used to change the function of a promoter by genome editing by the following procedure.
[0089] 1. Input base sequence X. For example, it is 2000 bases from -1995 to +5 around the transcription start point of the target gene. 2. Search for the PAM sequence in base sequence X. (Sequence of Cas9: NGG or Cas12a: TTTV) 3. Based on the PAM sequence, store the list of cleavage positions in an array. Here, for Cas9, the cleavage position is at the 4 - base position downstream of the PAM sequence, and for Cas12a, the cleavage position is at the 18 - base position downstream of the PAM sequence. The cleavage positions are found, for example, about 50 places within 2000 bases. 4. Among the n cleavage positions, the number of combinations of specifying 2 positions is \(_{n}C_{2}\). For example, for 50 cleavage positions, there can be 1250 combinations. For each of the \(_{n}C_{2}\) combinations, generate the base sequence with the bases between the 2 cleavage positions trimmed (deleted). 5. For each of the generated base sequences, obtain the 170 bases on the 3'-end side and input them into the learning model. 6. The learning model outputs the strength (transcription activity) of the given base sequence as a core promoter. 7. Overall, the following three types of information can be obtained. A. The transcription activity of base sequence X B. The set of cleavage positions where the promoter activity of the sequence formed by trimming the bases between the 2 cleavage positions is greater than A. C. The set of cleavage positions where the promoter activity of the sequence formed by trimming the bases between the 2 cleavage positions is smaller than A. 8. For example, when increasing the expression of a target gene, two types of guide RNAs targeting the cleavage positions obtained in 7.B are simultaneously introduced into cells, and genome editing by CRISPR - Cas9 or - Cas12a is performed, so that individuals with a deletion between the two cleavage sites are expected to be obtained.
[0090] In particular, it should be noted that using two different guide RNAs simultaneously is intended to induce small - to - medium - scale deletions within the scope of category SDN1.
[0091] Example 7: Visualization of Predicted Promoter Activity Values Finally, from the program described in Example 6, a set of the positions of the two guide RNAs and the predicted value of the new promoter activity of the sequence after deletion is obtained, which is difficult to intuitively understand. Therefore, the inventors have also devised a method to visualize these.
[0092] As the first method of visualization, a method for analyzing the overall image of the input base sequence X was devised. The vertical axis represents the predicted promoter strength, and the horizontal axis represents the position of the base sequence. Each plot shows the value when the DNA sequence at a certain position is cut out with a 170-base window and the transcriptional activity is predicted. For example, the point at position 0 on the horizontal axis is obtained by taking out 170 bases from the 5'-end of the input base sequence X and predicting the transcriptional activity. As a result, the transcriptional activity was 1.24 (since this 170-base base sequence is not in contact with the gene, it is unlikely to function as a core promoter, but if a gene is connected directly below this 170-base sequence, it is predicted to transcribe with a strength of 1.24). Therefore, it was plotted at the position of the point (0, 1.24). Next, the window was slid by 1 base, and the base sequence from the 2nd base to the 171st base of the base sequence X was taken out and predicted in the same way. As a result, the transcriptional activity was estimated to be 1.37. Therefore, it was plotted at the position of the point (1, 1.37). This operation was repeated, and a total of 1830 points were plotted (Figure 8).
[0093] In Figure 8, the rightmost point (1829, 1.47) is the evaluation of the 170 bases immediately upstream of the gene. If the sequence between the point with a value higher than this value and the 3'-end can be removed by genome editing, an increase in gene expression is expected. Similarly, it is considered possible to decrease gene expression. By using this figure, the overall image of sequence X can be overviewed, and it becomes clear which part should be designed with a guide RNA that hybridizes to obtain the desired expression level.
[0094] The above visualization method may be effective for roughly grasping the overall image of the target sequence. On the other hand, the positional relationship between the promoter strength and the positions of the two guide RNAs to be designed is not shown. To overcome this point, visualization by another method was devised. Figure 9 shown below is the result of visualization by another method.
[0095] A guide RNA designed at a position close to the target gene is called a proximal guide, and a guide RNA designed at a distant position is called a distal guide. Visualization was performed based on the position where the guide RNA can be designed (the position of the PAM sequence). However, only the range of 1000 bases directly upstream of the gene is shown.
[0096] The horizontal axis indicates the position of the base hybridized by the distal guide in sequence X. That is, it shows the distance between the transcription start point of the gene and the distal guide. The vertical axis shows the distance between the proximal guide and the transcription start point of the gene in the same way. Here, since the proximal guide is not designed further away than the distal guide, half of the plots in Figure 9 are hidden. At the points on the line y = x, both the proximal guide and the distal guide are located at approximately the same distance from the gene, indicating that they are close to each other. In this case, the number of bases to be deleted is small, resulting in minimal editing, which is desirable. When deleting the sequence between the proximal guide and the distal guide, 170 bases were taken from the 3'-end of the new sequence that could be formed, and the transcriptional activity was predicted. Based on the predicted values, the color of the plots was changed. In the color chart, light gray indicates a predicted high transcriptional activity, and black indicates a predicted low transcriptional activity.
[0097] For example, the plot at the point (849, 17) will be used as an example to explain. The position of the distal guide at this point is 849 bases away from the gene. Also, the position of the proximal guide is 17 bases away from the gene. A base sequence X' was created by deleting 832 bases between these two points from the base sequence X. When 170 bases were taken from the 3'-end of the base sequence X' and the transcriptional activity was predicted, it was 4.02. According to the color chart, a light gray plot was marked. By performing the above operations for all combinations of distal guides and proximal guides, the graph was created.
[0098] When it is desired to increase the expression level of this gene, select points greater than the original transcriptional activity of 2.34, and read the positions of the guide RNAs to be designed from the positions of their distal and proximal guides. For example, when selecting the point (849, 17) or the point (841, 29), high transcriptional activity is expected. Also, when selecting the point (432, 110) or the point (271, 77), low transcriptional activity is expected. Visualizing in this way may simplify the design of guide RNAs that achieve the structure of the target promoter.
[0099] After obtaining the position of the target sequence of the guide RNA in this way, the nucleotide sequence at that position was obtained, and 20 or 23 nucleotides were obtained to design the guide RNA. The sequence of this guide RNA was evaluated and verified using tools such as CRISPOR, and the high specificity and off-targets were evaluated. When successful in designing an appropriate guide RNA, a vector carrying the guide RNA was created by applying artificial DNA synthesis and the PCR method. Thereafter, genome editing can be carried out according to established methods.
[0100] Example 8: Example of a Soybean Gene Promoter The training data of the machine learning model used the promoter sequences of Arabidopsis thaliana, sorghum, and maize. Therefore, it was unclear whether the prediction system according to the present invention would function as expected for promoter sequences of organisms other than these three species. Therefore, a study was conducted to increase the transcriptional activity for a specific soybean gene. The results are shown in FIG. 10.
[0101] When predicting the transcriptional activity of the 170-base sequence of the core promoter of a certain soybean gene, it was predicted to be 0.927494. Based on this value, when predicting the LUC expression level, it was 1.26% of the positive control. This promoter sequence was artificially synthesized and introduced into protoplasts isolated from Arabidopsis thaliana leaves, and a LUC assay was performed. As a result, it was 4.15%.
[0102] Based on this promoter, two types of promoter sequences after genome editing were created to increase gene expression. Let them be edit1 and edit2 respectively. The transcriptional activity of edit1 was predicted to be 2.950119. Also, the LUC expression level was predicted to be 13.2%. The measured value of the LUC assay for this promoter was 19.0%.
[0103] The transcriptional activity of edit2 was predicted to be 2.432977. Also, the LUC expression level was predicted to be 7.27%. The measured value of the LUC assay for this promoter was 41.6%.
Explanation of symbols
[0104] 010 ··· Information processing device 011 ··· Sequence input part 012 ··· Modified sequence generation part 013 ··· Activity prediction part 014 ··· Sequence selection part 100 ··· Computer 101 ··· Control part 102 ··· Memory part 103 ··· Peripheral device I / F part 104 ··· Input part 105 ··· Display part 106 ··· Communication part 110 ··· Bus 120 ··· Network 130 ··· External server 140 ··· Database
Claims
1. A method for obtaining a promoter sequence modified to have a desired activity, comprising: preparing an original promoter sequence to be modified; generating a set of a plurality of modified promoter sequences that can be created by a genome editing technique based on the promoter sequence; predicting the activity of each modified promoter sequence included in the generated set of modified promoter sequences by a machine learning model; selecting a modified promoter sequence predicted to have the desired activity .
2. The method according to claim 1, wherein the desired activity is a gene expression induction activity higher than that of the original promoter sequence or a gene expression induction activity lower than that of the original promoter sequence.
3. The method according to claim 1, wherein the promoter sequence is a promoter sequence of a plant cell.
4. The method according to claim 1, wherein the original promoter sequence includes a core promoter sequence and the sequence upstream thereof.
5. The method according to claim 1, wherein the genome editing technique is a genome editing technique using the CRISPR / Cas system.
6. The method according to claim 5, wherein the set of a plurality of modified promoter sequences is generated by a sequence deletion caused by cleavage induced by a combination of guide RNA sequences designed based on two PAM recognition sequences.
7. The method according to claim 1, wherein the set of modified promoter sequences includes at least 1000 different sequences.
8. The method according to claim 1, wherein the machine learning model is a regression model trained to predict gene expression induction activity from a promoter sequence using gene expression induction activity data of a plurality of promoter sequences in a plant cell as teacher data.
9. The method according to claim 1, further comprising a step of visualizing the activity of each modified promoter sequence included in the generated set of modified promoter sequences on a computer display.
10. An information processing apparatus for predicting a promoter sequence modified to have a desired activity, comprising: a sequence input unit that receives an input of an original promoter sequence to be modified; a modified sequence generation unit that generates a set of a plurality of modified promoter sequences that can be created by a genome editing technique based on the promoter sequence; an activity prediction unit that predicts the activity of each modified promoter sequence included in the generated set of modified promoter sequences by a machine learning model; An array selection unit that selects a modified promoter array predicted to have a desired activity An information processing apparatus including the same.
11. A non-transitory computer-readable medium storing instructions that, when executed by a processor, perform the following steps: Receiving an input of an original promoter array to be modified; Generating a set of a plurality of modified promoter arrays that can be created by a genome editing technique based on the promoter array; Predicting the activity of each modified promoter array included in the generated set of modified promoter arrays using a machine learning model; Selecting a modified promoter array predicted to have a desired activity A computer-readable medium capable of performing the above steps.
12. A program for causing a computer to have a function of receiving an input of an original promoter array to be modified; have a function of generating a set of a plurality of modified promoter arrays that can be created by a genome editing technique based on the promoter array; have a function of predicting the activity of each modified promoter array included in the generated set of modified promoter arrays using a machine learning model; have a function of selecting a modified promoter array predicted to have a desired activity to be realized.
13. A method for genome editing of cells for regulating the expression level of a desired gene, comprising: obtaining a modified promoter array having a desired activity by the method according to claim 1; preparing cells to be subjected to genome editing; editing the genome of the cells so as to generate the modified promoter array A method including the above steps.
14. A method for producing a genome-edited plant in which the expression level of a desired gene is regulated, comprising: editing the genome of a desired plant cell by the method according to claim 13; obtaining a plant individual derived from the genome-edited cell A method including the above steps.
15. A machine learning model for predicting gene expression induction activity in plant cells from a promoter array, which is a regression model trained to predict gene expression induction activity from a promoter array using gene expression induction activity data of a plurality of promoter arrays in plant cells as teacher data.
16. A method for generating a machine learning model that predicts gene expression induction activity in plant cells from a promoter sequence, the method comprising training the model using gene expression induction activity data of a plurality of promoter sequences in plant cells as teacher data.
17. The method according to claim 16, wherein a transformer-based pre-trained base model is used as the model to be trained.
Citation Information
Patent Citations
Methods for predicting pathogenicity of gene sequence variants
JP2018527647A
Engineered CRISPR / Cas9 Systems for Simultaneous Long-term Regulation of Multiple Targets
US20220290132A1
Systems and methods for synthetic regulatory sequence design or production
US20230056396A1