Method for predicting promoter activity and method for modifying promoter based on results of predicting promoter activity

WO2025141959A1PCT designated stage expired Publication Date: 2025-07-03GRA&GREEN INC

Patent Information

Application Number
PCT/JP2024/030592
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-08-28
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

The prior art is difficult to accurately control gene expression levels. Traditional methods such as RNAi and transgenic methods have limitations, and the regulatory effect on enhanced or weakened gene function is poor.

Method used

Machine learning models such as DNABERT are used to predict gene expression levels, and gene promoter regions are edited in combination with the CRISPR/Cas system to generate a variety of modified promoter sequences, and select sequences that meet the expected activity.

Benefits of technology

Accurate regulation of gene expression levels is achieved, which can significantly enhance or weaken gene function and improve the accuracy and efficiency of gene editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JPOXMLDOC01-APPB-T000001
    Figure JPOXMLDOC01-APPB-T000001
  • Figure JPOXMLDOC01-APPB-T000002
    Figure JPOXMLDOC01-APPB-T000002
  • Figure 00000051_0000
    Figure 00000051_0000
Patent Text Reader

Abstract

The present disclosure provides a method for obtaining a promoter sequence modified so as to have a desired transcriptional activity. More specifically, provided is a method for obtaining a promoter sequence modified so as to have a desired transcription activity, the method comprising: preparing an original promoter sequence to be modified; generating a set of a plurality of modified promoter sequences that can be created by genome editing techniques on the basis of the promoter sequence; predicting, by a machine learning model, the transcription activity of each modified promoter sequence included in the set of modified promoter sequences generated; and selecting a modified promoter sequence predicted to have a desired active transcription performance.
Need to check novelty before this filing date? Find Prior Art

Description

Method for predicting promoter activity and method for modifying promoters based on the results of the prediction

[0001] The present disclosure relates to prediction of promoter activity and modification of promoters based on the results of the prediction.

[0002] Currently, when gene editing is used to confer traits to plants, the main method is loss-of-function, which involves introducing a mutation into the coding region (CDS) and causing a frameshift that results in the loss of gene function.

[0003] On the other hand, methods can be considered to enhance function by increasing the expression level of the target gene, or to suppress the side effects of knockout by leaving a small amount of expression (rather than reducing expression to zero). Since research on functional enhancement by introducing transgenes and knowledge gained from functional analysis using RNAi can be applied, such methods are expected to greatly expand the variety of traits that can be conferred by genome editing, and are important.

[0004] Zhang, Yi, et al. "Applications and potential of genome editing in crop improvement." Genome biology 19 (2018): 1-11.

[0005] One such method is to edit the promoter region instead of CDS. Because gene expression is determined by regions such as promoters and enhancers, it is thought that by modifying the base sequence of these regions, it is possible to increase or decrease the expression level while maintaining the gene's function.

[0006] In fact, the dynamic range of gene expression levels measured by RNA-seq is extremely large. While some mRNAs contain only a few copies in a single cell, there are many mRNAs that contain up to 10 copies. 5Some mRNAs contain only a few copies. Such differences in gene expression levels are thought to be mainly caused by promoters and enhancers, and this suggests the potential of methods that modify the base sequence of promoter regions. Therefore, one of the objectives of the present disclosure is to provide a technology for precisely controlling gene expression levels through genome editing.

[0007] In regions such as promoters and enhancers, the sequence-function relationship of "this part has this function" is less clear than in CDS. Several Cis regulatory elements with highly conserved consensus sequences, such as the TATA box, have been discovered and are now organized into databases. For example, software such as PromoterCAD from the RIKEN Institute and NEW PLACE from the National Agriculture and Food Research Organization use this knowledge to search for Cis regulatory elements and predict whether transcription factors will bind to promoters. However, the accuracy of these systems is insufficient for the purpose of predicting and designing expression levels.

[0008] Therefore, instead of using a database system to register individual elements, the inventors developed a system that uses machine learning to learn the relationship between base sequences and expression levels. As a result, the coefficient of determination (R) was calculated, which indicates the correlation between the actual and predicted values. 2 A high value of 0.75 was obtained. The prediction system was then used to design sequences, and experiments using plants were conducted to demonstrate the accuracy of the prediction. If transcriptional activity can be predicted by providing any base sequence to a computer, it becomes possible to "design" expression levels. As shown in this disclosure, the inventors have demonstrated that gene function (expression level) can be increased or decreased by editing the gene promoter based on computer predictions.

[0009] The present disclosure is based on these findings and includes the following aspects: [Aspect 1] A method for obtaining a promoter sequence modified to have a desired activity, comprising: preparing an original promoter sequence to be modified; generating a set of multiple modified promoter sequences that can be created by genome editing technology based on the promoter sequence; predicting the activity of each modified promoter sequence included in the generated set of modified promoter sequences using a machine learning model; and selecting a modified promoter sequence predicted to have the desired activity. [Aspect 2] The method of Aspect 1, wherein the desired activity is a gene expression induction activity that is higher than that of the original promoter sequence, or lower than that of the original promoter sequence. [Aspect 3] The method of Aspect 1, wherein the promoter sequence is a promoter sequence for a plant cell. [Aspect 4] The method of Aspect 1, wherein the original promoter sequence comprises a core promoter sequence and a sequence upstream of it. [Aspect 5] The method of Aspect 1, wherein the genome editing technology is genome editing technology using a CRISPR / Cas system. [Aspect 6] The method of Aspect 5, wherein the set of multiple modified promoter sequences is generated by sequence deletion caused by cleavage induced by a combination of guide RNA sequences designed based on two PAM recognition sequences. [Aspect 7] The method of Aspect 1, wherein the set of modified promoter sequences comprises at least 1,000 different sequences. [Aspect 8] The method of Aspect 1, wherein the machine learning model is a regression model trained to predict gene expression induction activity from a promoter sequence using training data on the gene expression induction activity of a plurality of promoter sequences in plant cells. [Aspect 9] The method of Aspect 1, further comprising a step of visualizing on a computer display the activity of each modified promoter sequence included in the generated set of modified promoter sequences.[Aspect 10] An information processing device that predicts a promoter sequence modified to have a desired activity, comprising: a sequence input unit that accepts input of an original promoter sequence to be modified, a modified sequence generation unit that generates a set of multiple modified promoter sequences that can be produced by genome editing technology based on the promoter sequence, an activity prediction unit that predicts the activity of each modified promoter sequence included in the generated set of modified promoter sequences using a machine learning model, and a sequence selection unit that selects the modified promoter sequence predicted to have the desired activity. [Aspect 11] A non-transitory computer-readable medium that stores instructions, which, when executed by a processor, can perform the following steps: accepting input of an original promoter sequence to be modified, generating a set of multiple modified promoter sequences that can be produced by genome editing technology based on the promoter sequence, using a machine learning model to predict the activity of each modified promoter sequence included in the generated set of modified promoter sequences, and selecting the modified promoter sequence predicted to have the desired activity. [Aspect 12] A program that causes a computer to perform the following functions: accepting input of an original promoter sequence to be modified, generating a set of multiple modified promoter sequences that can be created by genome editing technology based on the promoter sequence, predicting the activity of each modified promoter sequence included in the generated set of modified promoter sequences using a machine learning model, and selecting a modified promoter sequence predicted to have a desired activity. [Aspect 13] A method for genome editing of a cell to regulate the expression level of a desired gene, the method comprising: obtaining a modified promoter sequence having a desired activity by the method of Aspect 1; preparing a cell to be subjected to genome editing; and editing the genome of the cell to produce the modified promoter sequence.[Aspect 14] A method for producing a genome-edited plant in which the expression level of a desired gene is regulated, the method comprising: editing the genome of a desired plant cell by the method of Aspect 13; and obtaining an individual plant derived from the genome-edited cell. [Aspect 15] A machine learning model for predicting gene expression induction activity in a plant cell from a promoter sequence, the machine learning model being a regression model trained to predict gene expression induction activity from the promoter sequence using training data on the gene expression induction activity of a plurality of promoter sequences in the plant cell as training data. [Aspect 16] A method for generating a machine learning model for predicting gene expression induction activity in a plant cell from a promoter sequence, the method comprising training the model using training data on the gene expression induction activity of a plurality of promoter sequences in the plant cell as training data. [Aspect 17] The method of Aspect 16, in which a Transformer-based pre-trained basic model is used as the model to be trained. [Aspect 18] A method for obtaining a promoter sequence modified to have a desired activity, comprising: preparing an original promoter sequence to be modified; generating a set of multiple modified promoter sequences that can be produced by genome editing technology based on the promoter sequence; predicting the activity of each modified promoter sequence included in the generated set of modified promoter sequences using a machine learning model; and selecting a modified promoter sequence predicted to have the desired activity; wherein the promoter sequence is a promoter sequence of a plant cell; the original promoter sequence includes a core promoter and a sequence upstream of it, wherein the core promoter is a sequence included in positions −200 to +50 when the base position of the transcription start site is represented as +1; the genome editing technology is a genome editing technology using a CRISPR / Cas system; and the set of multiple modified promoter sequences is generated by sequence deletion caused by cleavage induced by a combination of guide RNA sequences designed based on two PAM recognition sequences, wherein at least one of the two PAM recognition sequences is located within the core promoter; and the length of the original promoter is at least 2,000 bp.

[0010] Figure 1 is a schematic diagram showing the structure of the artificial synthetic gene. Seven promoters from the high-expression group, eight promoters from the medium-expression group, and four promoters from the low-expression group were selected and connected to the luciferase (LUC) gene to create artificial synthetic genes. The promoters had 19 different sequence patterns. The downstream LUC gene and the upstream cauliflower mosaic virus (CaMV) 35S enhancer sequences were common to each artificial gene. Figure 2 is a scatter plot showing the relationship between the measured LUC gene expression levels and the predicted promoter strengths. The vertical axis represents the predicted promoter strength. Higher values ​​indicate higher expected expression levels of downstream genes. The horizontal axis represents the measured LUC expression levels, normalized by the transcriptional activity of the positive control CaMV promoter. Figure 3 shows the results of calculating and comparing the LUC expression levels corresponding to the predicted transcriptional activity. The vertical axis represents the LUC expression levels. The values ​​are expressed relative to the luminescence intensity of the positive control, which is set to 1. The cauliflower mosaic virus (CaMV) 35S promoter, known to exhibit high transcriptional activity, was used as a positive control. The horizontal axis shows the numbers of the 19 promoters. Promoter numbers 3 to 10 were classified into the "high expression" group, which predicted high expression intensity; promoter numbers 13 to 21 into the "moderate" group; and promoter numbers 22 to 25 into the "low expression" group. Black bars represent the actual measured values ​​of the LUC assay. White bars represent the predicted transcription activity values ​​using the prediction system of the present disclosure. Figure 4 shows the results of a search for editing patterns that can reduce transcription activity. Editing patterns that can reduce transcription activity were searched for promoter numbers 3, 4, 5, 6, 9, 10, and 21. The black triangle plot shows the predicted LUC assay score (expression level) calculated based on the predicted new transcription activity. As a result, transcription activity was predicted to be approximately 14% to 1% of that of the original promoter. Figure 5 shows the actual measured LUC expression levels of the new promoters in the shaded bars superimposed on Figure 4. Due to the very low expression levels, the vertical axis has been expanded. Promoter number 5 was a missing value.For all six promoters for which measurements were obtained, significant decreases in expression levels were observed, consistent with the predicted values. Figure 6 shows the results of a search for promoters whose predicted transcriptional activity increased after editing. Promoter numbers 13, 17, 22, 23, 24, and 25 were selected for increased transcriptional activity, and their predicted expression levels are indicated by white circles. LUC expression levels were predicted to be 7.4 to 125 times higher than the original promoter. Figure 7 shows the results of artificially synthesizing genes with theoretical promoter sequences, transfecting them into protoplasts in the same manner, and measuring their expression levels using a plate reader. Of the six promoters, five showed increased expression levels compared to the original promoter, but only promoter number 13 exceeded the predicted level. Figure 8 shows an example of a visualization method for analyzing the overall picture of input base sequence X. The vertical axis corresponds to predicted transcriptional activity, and the horizontal axis corresponds to the position in the base sequence. Figure 9 shows an example of visualization using a different method. The vertical axis corresponds to the position of the guide RNA (proximal guide) designed closer to the target gene, and the horizontal axis corresponds to the position of the guide RNA (distal guide) designed further away. The color of each plot changes based on the predicted transcriptional activity. For example, on a color chart, light gray indicates high transcriptional activity and black indicates low transcriptional activity. Figure 10 shows the results of an experiment to enhance the transcriptional activity of a specific soybean gene. Based on the promoter of a soybean gene, two post-genome editing promoter sequences (edit1 and edit2) were created to increase gene expression. The transcriptional activity of edit1 was predicted to be 2.950119. Furthermore, linear regression predicted LUC expression to be 13.2% of the P control. The actual LUC assay value of this promoter was 19.0% of the P control. The transcriptional activity of edit2 was predicted to be 2.432977. Furthermore, LUC expression was predicted to be 7.27% of the P control. The actual measured value of the LUC assay for this promoter was 41.6% of the P control. Figure 11 shows an exemplary configuration of a data processing device for predicting promoter sequences modified to have desired activity.Figure 12 shows an exemplary flowchart of a program that allows a computer to receive input of an original promoter sequence to be modified, generate a set of multiple modified promoter sequences that can be created using genome editing technology based on the promoter sequence, predict the activity of each modified promoter sequence included in the generated set of modified promoter sequences using a machine learning model, and select modified promoter sequences predicted to have the desired activity. Figure 13 shows a schematic configuration of a computer that can be used to implement aspects of the present disclosure. Figure 14 shows the results of in silico genome editing of the promoter sequence of the Arabidopsis thaliana AT4G39130 gene with one target (one guide RNA). Figure 15 shows the clear difference in changes in gene expression levels between two-target deletion and one-target deletion in the promoter region.

[0011] During the crop breeding process, breeders hope to enhance or suppress the expression of specific genes. As gene function is increasingly analyzed, the relationship between genes and traits is becoming increasingly clear. For example, in the case of semi-dwarf rice, semi-dwarfness was a breeding goal, and the semi-dwarf variety "Reimei" was developed in 1966. Semi-dwarf rice is less likely to grow taller even when fertilizer is applied, and yields do not decrease. Normally, rice plants grow vertically when fertilized, but this makes them prone to lodging in typhoons and strong winds, resulting in harvest failure or reduced quality. Therefore, applying fertilizer beyond a certain level leads to lodging and does not lead to increased yields. Because semi-dwarf rice is less likely to grow taller, it is more resistant to lodging, even when large amounts of fertilizer are applied, and yields increased. It is now clear that the mutation that conferred semi-dwarfness to Reimei is located in a gene encoding the G20 oxidase enzyme in the gibberellin biosynthesis pathway. Semi-dwarf crops were a major invention that brought about enormous advances in expanding food production.

[0012] As such, the relationship between genes and traits is being revealed one after another, and knowledge is accumulating. As a result, target genes are often selected to obtain desired traits. For example, if one wishes to introduce semi-dwarfism into rice, one can simply select a mutant of the above gene (the sd1 mutant). While it is possible to search for naturally occurring mutants, it is also possible to randomly introduce mutations using chemicals or radiation. Although the loss-of-function allele of sd1 is thought to be recessive, it is possible to breed individuals in which both alleles are loss-of-function sd1 through planned breeding.

[0013] The advent of genetic engineering made it possible to artificially engineer genotypes. This meant that introducing any gene sequence into a target crop could confer functions and traits. A prime example is herbicide-resistant corn (Roundup Ready®). Roundup® (glyphosate) is a herbicide and, in essence, an amino acid analog. When glyphosate is absorbed into some plants and bacteria, amino acid synthesis is inhibited, causing death. Introducing a glyphosate-neutralizing gene from bacteria gave plants resistance to glyphosate, a major breakthrough that significantly reduced the cost of weed control. However, the production and use of genetically modified plants, which incorporate foreign genes, was limited, preventing the replacement of all crops on Earth with genetically modified plants. There was particularly strong opposition in Europe, where the cultivation of genetically modified crops remains largely restricted, with only one variety (corn in Spain) currently in cultivation.

[0014] The situation has changed even further since the invention of genome editing. It now becomes possible to directly introduce mutations into target genes without genetic modification. This is a historic breakthrough. For example, if one wanted to introduce semi-dwarfism into Koshihikari rice, one could simply introduce a small indel (1–10 bases) into the SD1 locus, inducing a frameshift and causing a loss of function. This method allows for faster development than chemical or radiation-mediated mutagenesis (followed by 5–7 backcrosses), and it avoids the expression of unintended traits due to linkage drugs, offering significant cost advantages. An example of a food product developed using this method and sold in Japan is the red sea bream (Daphnia gracilis) lacking the myostatin gene, sold by Regional Fish Co.

[0015] Genome editing can induce not only small indels of 1-10 base pairs but also medium-sized deletions of up to several thousand base pairs. In this case, two guide RNAs are simultaneously delivered into cells, and if the two cut sites that cut the two target DNA sequences are physically close to each other, the gap between the two cut sites can be removed and repaired. A specific example is Corteva Agriscience's waxy corn, which contains a deletion of approximately 4 kb, including the entire coding sequence of the Wx locus.

[0016] The mutations in the varieties developed through genome editing mentioned above are of the kind that could occur naturally, and the resulting crops are essentially indistinguishable from those bred through traditional crossbreeding. Genome editing, on the other hand, can also be used to induce homologous recombination, replacing a target position in the genome with any DNA sequence. This can be used to introduce various mutations, such as point mutations, deletions, and insertions. Furthermore, replacing a target position with an exogenous gene can induce genetic recombination. Thus, while genome editing is a broad term, it can produce essentially different categories of mutants, including mutations that occur naturally and genetic modifications that would never occur in nature. To avoid confusion, varieties created through genome editing are classified into three categories: SDN1, SDN2, and SDN3. The above examples from Regional Fish and Corteva are both classified as SDN1. In many countries, including Japan, the creation of new genome-edited crop varieties requires procedures such as notification and prior consultation with government agencies. In SDN1, the newly introduced mutations are essentially equivalent to naturally occurring mutations, and there is a long history of precedent, so it is expected that the process will proceed quickly in many countries, which will be beneficial for industry.

[0017] As mentioned above, genome editing can induce frameshifts, which can easily create knockout mutants of specific genes. However, knockout mutants are not the only useful gene mutations for breeders. Countless mutants with increased or decreased gene expression have been intentionally or unintentionally selected. Furthermore, a vast number of gene silencing (RNAi, KD) experiments have been reported throughout the history of plant science. Phenotypes can differ between knockout (KO) and knockdown (KD) mutants. In some cases, even slight residual expression may prevent deleterious phenotypes. Thus, significant reductions in gene expression, such as those mimicking KD mutants, are often considered useful. The method disclosed herein has the advantage of enabling the design of increased or decreased gene expression using genome editing of SDN1, which can be commercialized in Japan simply by submitting a notification to the relevant government agencies.

[0018] One method for enhancing function is to introduce mutations into genes. For example, disrupting the autoinhibitory domain of an enzyme may increase activity. For example, in the case of Sanatec Seed's high-GABA-accumulating tomato, GABA synthesis activity is enhanced by disrupting the autoinhibitory domain at the C-terminus of the GABA synthase GAD through frameshifting. While this method is highly effective, it can only be used when the target protein (gene) contains an autoinhibitory domain. Because relatively few proteins possess autoinhibitory domains, its range of application is limited. Other methods include disrupting the protein's localization signal to ensure it is always active, substituting the amino acid subject to phosphorylation with glutamic acid, or introducing mutations into the enzyme's active site to improve reactivity. However, these methods are not generally applicable to all genes.

[0019] Another method for enhancing function that can be applied to many genes is to introduce mutations into the region that regulates gene expression. In other words, the idea is to increase expression by editing the promoter, thereby enhancing function. However, this idea has a drawback: even for promoters, the correspondence between sequence and function is unclear. In genes, the correspondence between codons and amino acid sequences is strict, and the functional domains created by specific amino acid sequences are highly conserved, so function can be inferred from the sequence. This allows for highly accurate disruption of target domains. Conversely, with promoters (non-ORFs), it is difficult to hypothesize, "If I do this part, this will happen." Of course, there is also a relationship between sequence and function. Cis regulatory elements, such as the TATA box, are understood as binding sites for basal transcription factors and transcription factors. There have been successful examples of inferring transcription factors based on their sequences and inferring binding sites from transcription factors. For example, web applications such as PromoterCAD from the RIKEN Institute and NEWPLACE from the National Agriculture and Food Research Organization are tools for searching for such elements. While these tools have been proposed, there appear to be limited examples of successful breeding using them. Song et al., 2022, reported improved gene expression through promoter genome editing (Song, X., Meng, X., Guo, H. et al. Targeting a gene regulatory element enhances rice grain yield by decoupling panicle number and size. Nat Biotechnol 40, 1403-1411 (2022). https: / / doi.org / 10.1038 / s41587-022-01281-7). In this report, they demonstrated that yield could be increased by editing the promoter or 5'UTR of the rice IPA1 gene. The procedure involved using genome editing to create lines with various deletion patterns, evaluating expression patterns and traits, and identifying important regulatory elements.This shows how difficult it is to design function from regulatory elements. One reason for this difficulty is likely the fluctuation of regulatory elements (which typically allow for a mismatch of several bases from the consensus sequence). Recently, there have been reports of applying deep learning to infer Cis regulatory elements, but there have been no examples of this being linked to actual crop traits.

[0020] Instead of the conventional method of inferring function from cis regulatory elements, the inventors adopted a strategy of directly inferring expression levels from DNA sequences. To achieve this, they limited their target to promoters. While the relationship between sequence and expression level remains a black box, deep learning methods often demonstrate high performance even when the underlying principles and theory are unknown. For example, in the field of machine translation, until the 1990s, attempts were made to convert language by collecting vocabulary and building dictionaries linking them to function (semantic information) (e.g., English to German), but these methods failed to achieve high performance. In contrast, around 2011, WATSON, which applied neural networks (NNs), demonstrated high performance in natural language processing. Rather than relying on humans to create dictionaries and train computers to understand meaning, learning from vast amounts of text data enabled highly accurate translation without the need to input the meaning of individual sentences. Today, Google Translate and DeepL are representative examples of NN-based machine translation, both of which provide significant benefits. Thus, NN-based machine learning can achieve high performance without humans having to understand or organize the underlying principles and theories.

[0021] DNABERT is a model pre-trained to handle DNA sequences, based on the natural language model BERT, which was created for the purpose of natural language processing. DNABERT was developed for the purpose of analyzing non-coding regions based on DNA sequences and has demonstrated high performance in tasks such as the discovery of promoters and splice sites. For this reason, we decided to use it for the present invention, which is to predict the transcriptional activity of core promoters. However, there were two problems with applying DNABERT to the present invention.

[0022] First, the human genome was used as pre-training data, and it was unclear whether DNABERT could handle plant genomes. In plants, core promoter elements include elements shared with animals, such as the TATA box, initiator, and Kozak, as well as plant-specific elements such as the Y patch, CA, and GA. Furthermore, transcriptional activity in animal and plant gene promoters can be controlled by DNA methylation. In animals, cytosines in CG (CpG) sequences are often methylated. In contrast, cytosines in non-CG sequences are also methylated in plants. Thus, although the basic mechanisms of transcription are common between animals and plants, the components, such as core elements and methylation sites, differ. Thus, it was unclear whether DNABERT would perform well in plants, whose promoters have structures significantly different from those of humans.

[0023] Second, DNABERT has demonstrated high performance in tasks such as identifying transcription start sites and splice sites. However, these tasks involve determining whether a given sequence contains a specific element of interest—a binary classification problem. In contrast, our task involved predicting gene expression levels downstream of a core promoter sequence, which required outputting a continuous value. It was unclear whether DNABERT would be effective or perform well on such regression problems. CNNs have traditionally been used to handle continuous values ​​related to DNA sequences. Based on reports that the BERT natural language processing model can also handle regression problems, we modified the DNABERT code. Specifically, we added classes to enable continuous values.

[0024] Despite these problems, we successfully predicted promoter strength with high accuracy by fine-tuning the DNABERT pre-training model using a procedure described separately. Furthermore, we have successfully developed a program that incorporates this training model and changes promoter function through genome editing.

[0025] In one aspect, the present disclosure relates to a method for obtaining a promoter sequence modified to have a desired activity. In some embodiments, the method according to the present disclosure may include: preparing an original promoter sequence to be modified; generating a set of multiple modified promoter sequences that can be created by genome editing technology based on the promoter sequence; predicting the activity of each modified promoter sequence included in the generated set of modified promoter sequences using a machine learning model; and selecting the modified promoter sequence predicted to have the desired activity.

[0026] A promoter is a DNA sequence (usually located upstream of a gene) that is not itself transcribed but determines the amount of gene transcription (expression level) and the site and tissue specificity of expression in multicellular organisms. Herein, the region that interacts with general transcription factors near the gene is referred to as a "core promoter." A core promoter can be defined as a region between -200 and +50, where +1 is the base position of the transcription start site. More preferably, it can be defined as a region between -180 and +25. Even more preferably, it can be defined as a region between -170 and +17. Most preferably, it can be defined as a region between -165 and +5. The position of a core promoter can also be defined as a range between -200, -195, -190, -185, -180, -175, -170, or -165 and +5, +8, +12, +17, +25, +30, +40, or +50. The core promoter is thought to have the function of determining the basal transcription level. The length of the core promoter can be, for example, about 250 bp, about 225 bp, or about 200 bp.

[0027] In some embodiments, the promoter sequence is a promoter sequence for plant cells. It is known that the types of core elements contained in core promoters differ between animals and plants. As used herein, "plant" is not particularly limited. Examples include a wide range of plants, including bryophytes, ferns, gymnosperms, and the angiosperms Magnolia, monocotyledons, and eudicotyledons (Rosaceae I, Rosaceae II, Chrysanthemum I, Chrysanthemum II, and their outgroups). More specific examples of plants include nightshades such as tomatoes, bell peppers, chili peppers, eggplants, tobacco, and torvum; gourds such as cucumbers, pumpkins, melons, and watermelons; vegetables such as cabbage, broccoli, Chinese cabbage, and kale; fresh and spicy vegetables such as shiso, celery, parsley, and lettuce; alliums such as leeks, onions, and garlic; other fruit vegetables such as strawberries and melons; taproots such as radishes, turnips, carrots, and burdock; tubers such as taro, cassava, potato, sweet potato, and Chinese yam; grains such as rice, corn, wheat, sorghum, barley, rye, wheatgrass, and buckwheat; soybeans, adzuki beans, mung beans, cowpeas, and yams. Examples of suitable promoters include legumes such as columbine, peanut, pea, and broad bean; soft vegetables such as asparagus, spinach, and mitsuba; flowering plants such as lisianthus, rose, stock, carnation, and chrysanthemum; turfgrass such as bentgrass and Zoysiagrass; oilseed crops such as rapeseed, camelina, rapeseed, Jatropha, sesame, and perilla; fiber crops such as cotton, rush, and hemp; forage crops such as clover, dent corn, and alfalfa; deciduous fruit trees such as apple, pear, grape, and peach; citrus fruits such as Satsuma mandarin, orange, lemon, and grapefruit; and woody plants such as azalea, azalea, cedar, poplar, and rubber tree. The promoter may be site-specific or non-site-specific. A site-specific promoter may, for example, control expression specifically in leaves or roots.

[0028] In some embodiments, the desired activity of the modified promoter may be a higher or lower gene expression induction activity than the original promoter sequence. The gene expression induction activity higher than the original promoter sequence may be, for example, 1.1 to 1,000 times the original activity, for example, 1.1, 1.2, 1.5, 2, 3, 4, 5, 10, 20, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1,000 times. The gene expression induction activity lower than that of the original promoter sequence can be, for example, anywhere from 90% to 0.01% of the original activity, such as 90%, 80%, 50%, 10%, 5%, 1%, 0.5%, 0.1%, 0.05%, 0.04%, 0.03%, 0.02%, or 0.01%.

[0029] In some embodiments, the desired activity of the modified promoter may be determined by comparison with the activity of another reference promoter sequence, rather than the original promoter sequence.

[0030] Alternatively, the desired activity of the modified promoter may be determined by stratifying high expression, medium expression, low expression, etc. In some embodiments, any number or range of stratifications may be used.

[0031] In some embodiments, the modification of the promoter sequence may be by deletion, substitution, or insertion of a portion of the sequence, hi some embodiments, the modification of the promoter sequence may be by deletion of a portion of the sequence by making two truncations in the promoter sequence.

[0032] In some embodiments, the original promoter sequence to be modified can be obtained from a database such as GenBank. Examples of promoter sequences include any DNA region containing the transcription start point within the DNA region from -9,900 to +100, preferably -4,950 to +50 or -2975 to +25, and more preferably -1,995 to +5, around the transcription start point of the target gene. Alternatively, any DNA region between the transcription start point of the target gene and the transcription termination point of the adjacent gene can be used. In other words, a region containing the core promoter and its upstream sequence can be used as the promoter sequence. The length of the sequence upstream of the core promoter may be, for example, at least 200 bp, 400 bp, 600 bp, 800 bp, 1,000 bp, 1,200 bp, 1,400 bp, 1,600 bp, 1,800 bp, 2,000 bp, 3,000 bp, 4,000 bp, 5,000 bp, 6,000 bp, 7,000 bp, 8,000 bp, or 9,000 bp. The length of the sequence upstream of the core promoter may be, for example, 15,000 bp or less, 12,500 bp or less, or 10,000 bp or less. If multiple transcription start sites are present, these multiple transcription start sites can be selected. Alternatively, any range of 100 to 1,800 bases within the range of -1,995 to +5 bases around the transcription start site may be used. The length (number of base pairs: bp) of the promoter may be, for example, at least 200 bp, 400 bp, 600 bp, 800 bp, 1,000 bp, 1,200 bp, 1,500 bp, 2,000 bp, 3,000 bp, 4,000 bp, 5,000 bp, 6,000 bp, 7,000 bp, 8,000 bp, or 9,000 bp, and may range from 100 to 10,000 bp, preferably from 200 to 5,000 bp, and more preferably from 500 to 3,000 bp.

[0033] In some embodiments, a set of multiple modified promoter sequences that can be created by genome editing techniques based on the promoter sequence is generated to select a modified promoter sequence with the desired activity.

[0034] Genome editing technologies that can be used include ZFNs, TALENs, and CRISPR-Cas systems. CRISPR-Cas systems include Class 1 Type I CRISPR-Cas3, Class 2 Type II Cas9, Class 2 Type V Cas12a, Class 2 Type V Cas12f (Cas14a), and Class 2 Type VI Cas13a, regardless of their functional mechanism or classification. Cas proteins derived from various bacteria can also be used. Examples of Cas9 include SpCas9 derived from Streptococcus pyogenes, SaCas9 derived from Staphylococcus aureus, FnCas9 derived from Francisella novicida, and CjCas9 derived from Campylobacter jejuni. Examples of Cas12a include AsCas12a (AsCpf1) derived from Acidaminococcus sp., LbCas12a (LbCpf1) derived from Lachnospiraceae bacterium, and ErCas12a derived from Eubacterium rectale, but the origin is not limited thereto. Cas proteins may also be modified in their nucleotide sequences or amino acid sequences, fused with other proteins, functional domains, peptides, or amino acid sequences, or modified with compounds.

[0035] The CRISPR-Cas system uses Cas nuclease and guide RNA (gRNA) to cleave target sequences. Guide RNAs consist of two elements, crRNA and tracrRNA, which can exist as individual RNAs or linked together to form a single-stranded RNA. Single-stranded guide RNAs are also called sgRNAs. The 5' and 3' ends of the guide RNA may contain sequences of 1 to 10 bases, 10 to 50 bases, 50 to 100 bases, or 100 to 500 bases. The guide RNA binds to a specific DNA sequence complementary to the target sequence, guiding the Cas nuclease to a specific location in the genome. This allows the Cas nuclease to cleave the specific DNA sequence, enabling genome editing. For example, when using CRISPR, a PAM sequence in the promoter sequence is searched for to generate a list of possible cleavage sites. The PAM sequence varies depending on the Cas protein used; for example, NGG for SpCas9, NGRRT or NGNRRN for SaCas9, NNNNGATT for NmeCas9, NNNNRYAC for CjCas9, TTTV for LbCas12a (Cpf1), TTTV for AsCas12a (Cpf1), TTN for AacCas12a (Cpf1), and ATTN, TTTN, or GTTN for BhCas12b v4. However, these PAM sequences can contain mismatches of one, two, three, four, five, six, seven, or eight bases as long as the sequence is recognized by the Cas protein. The presence of a PAM sequence is important for enhancing the accuracy and specificity of CRISPR-based genome editing. Without a PAM sequence, the Cas nuclease cannot bind to the DNA sequence specified by the guide RNA. This characteristic allows precise targeting of specific gene regions for genome editing. As mentioned above, different types of Cas nucleases can be used to target regions with different PAM sequences. For example, Cas9 typically cleaves 3-4 bases from the PAM sequence, while Cas12a typically cleaves 18-23 bases from the PAM sequence. Those skilled in the art can design appropriate guide RNAs based on the PAM sequence in the target genome.Typically, about 50 cleavage sites are found in a 2,000-base sequence. For example, there are nC2 combinations of specifying two of the n cleavage sites. In other words, there are 1,250 possible combinations for the 50 cleavage sites. Therefore, in some embodiments, to select a modified promoter sequence with the desired activity, a set of base sequences can be generated by truncating (deleting) the two cleavage sites for each of the nC2 combinations of PAM sequences in the promoter sequence to be modified. Thus, in some embodiments, a set of multiple modified promoter sequences can be generated by sequence deletion caused by cleavage induced by a combination of guide RNA sequences designed based on two PAM recognition sequences. In some embodiments, at least one of the two PAM recognition sequences can be located within a core promoter. In some embodiments, both of the two PAM recognition sequences can be located within a core promoter. In some embodiments, only one of the two PAM recognition sequences can be located within a core promoter. In some embodiments, the length of the sequence deletion resulting from cleavage induced by a combination of guide RNA sequences designed based on two PAM recognition sequences is not particularly limited, as long as it is a sequence deletion (large deletion) larger than the deletion size (deletion of 1 to 5 base sequences) caused by one target (one guide RNA), and can be, for example, in the range of 50 bp to 15,000 bp, preferably in the range of 50 bp to 5,000 bp.Thus, in some embodiments, the length of the sequence deletion is at least 50 bp, 75 bp, 100 bp, 150 bp, 200 bp, 150 bp, 200 bp, 250 bp, 300 bp, 350 bp, 400 bp, 450 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1,100 bp, 1,200 bp, 1,300 bp, 1,400 bp, 1,500 bp, 1,600 bp, 1,700 bp, It can be 1,800bp, 1,900bp, 2,000bp, 2,100bp, 2,200bp, 2,300bp, 2,400bp, 2,500bp, 3,000bp, 4,000bp, 5,000bp, 6,000bp, 7,000bp, 8,000bp, 9,000bp, 10,000bp, 11,000bp, 12,000bp, 13,000bp, 14,000bp, or 15,000bp, preferably at least 400bp. In addition, in some embodiments, the length of the sequence deletion resulting from cleavage induced by a combination of guide RNA sequences designed based on two PAM recognition sequences may be 15,000 bp or less, 14,000 bp or less, 13,000 bp or less, 12,000 bp or less, 11,000 bp or less, 10,000 bp or less, 9,000 bp or less, 8,000 bp or less, 7,000 bp or less, 6,000 bp or less, 5,000 bp or less, 4,000 bp or less, 3,000 bp or less, 2,500 bp or less, 2,250 bp or less, 2,000 bp or less, 1,000 bp or less, or 500 bp or less. In addition, in some embodiments, the length of the sequence deletion resulting from cleavage induced by a combination of guide RNA sequences designed based on two PAM recognition sequences can be any length between 400 bp and 3,000 bp, for example, 400 bp, 450 bp, 500 bp, 600 bp, 700 bp, 800 bp, 900 bp, 1,000 bp, 1,100 bp, 1,200 bp, 1,300 bp, 1,400 bp, 1,500 bp, 1,600 bp, 1,700 bp, 1,800 bp, 1,900 bp, 2,000 bp, 2,100 bp, 2,200 bp, 2,300 bp, 2,400 bp, or 2,500 bp.Furthermore, in some embodiments, at least one of the two PAM recognition sequences may be located within 200 bases of the base position of the transcription start site, e.g., within 190 bases, 180 bases, 170 bases, 160 bases, 150 bases, 140 bases, 130 bases, 120 bases, 110 bases, 100 bases, 90 bases, 80 bases, 70 bases, 60 bases, 50 bases, 40 bases, 30 bases, 20 bases, 10 bases, or 5 bases. Furthermore, in some embodiments, at least one of the two PAM recognition sequences may be located at a position 200 or more bases away from the base position of the transcription start site, for example, 250 bases, 300 bases, 400 bases, 500 bases, 750 bases, 1,000 bases, 1,500 bases, 2,000 bases, 3,000 bases, 4,000 bases, 5,000 bases, 7,500 bases, 10,000 bases, 12,500 bases, or 15,000 bases or more.

[0036] Other promoter editing methods that can be used include deleting the space between two nicks created by Cas9 nickase, using base editing to replace a targeted base in the genome with another base, and using prime editor, which can delete, insert, or replace any base sequence consisting of 1 to 20 bases, 20 to 50 bases, or 50 to 100 bases in the genome.

[0037] The number of different sequences included in the set of modified promoter sequences can be at least 5, 10, 20, 50, 100, 150, 200, 300, 500, 700, 1,000, 1,200, 1,500, 2,000, 3,000, 4,000, or 5,000. To fully identify sequences that affect activity, the number of different sequences included in the set of modified promoter sequences is preferably 1,000 or more. To ensure a sufficient number of different sequences included in the set of modified promoter sequences, the length of the original promoter sequence is preferably, for example, 1,500 bp or more, 1,600 bp or more, 1,700 bp or more, 1,800 bp or more, 1,900 bp or more, 2,000 bp or more, 2,100 bp or more, or 2,200 bp or more.

[0038] After generating a set of multiple modified promoter sequences, in some embodiments, the activity of each modified promoter sequence included in the generated set of modified promoter sequences is predicted using a machine learning model. For example, for each generated base sequence, the 3'-terminal 170 bases are obtained and input into the machine learning model. In some embodiments, 100 to 250 bases from the 3'-terminal, e.g., 100 bp, 110 bp, 120 bp, 130 bp, 140 bp, 150 bp, 160 bp, 170 bp, 180 bp, 190 bp, 200 bp, 210 bp, 220 bp, 230 bp, 240 bp, or 250 bp, may be obtained and input into the machine learning model. In addition, in some embodiments, a portion of the sequence included in the 3'-terminal 170 bases, e.g., 100 to 169 bases, may be input. A properly trained machine learning model can output the strength (transcriptional activity) of a given base sequence as a core promoter.

[0039] In some embodiments, the activity of each modified promoter sequence in the generated set of modified promoter sequences is visualized on a computer display, for example, as shown in Figure 8 or Figure 9.

[0040] Then, by selecting a modified promoter sequence predicted to have the desired activity from the prediction results of the machine learning model, a promoter sequence modified to have the desired activity can be obtained.

[0041] In some embodiments, the machine learning model may be a regression model trained to predict gene expression induction activity from a promoter sequence using training data on gene expression induction activity of multiple promoter sequences in plant cells. In some embodiments, data from Jores et al. (2021) may be used for training (Jores, T., Tonnies, J., Wrightsman, T. et al. Synthetic promoter designs enabled by a comprehensive analysis of plant core promoters. Nat. Plants 7, 842-855 (2021). https: / / doi.org / 10.1038 / s41477-021-00932-y). Thus, one aspect of the present disclosure relates to a method for generating a machine learning model that predicts gene expression induction activity in plant cells from a promoter sequence, the method comprising training the model using training data on gene expression induction activity of multiple promoter sequences in plant cells. In some embodiments, the model to be trained may be a pre-trained Transformer-based basic model.

[0042] In some embodiments, deep learning models such as Transformers, particularly models based on BERT, may be used to build machine learning models. BERT (Bidirectional Encoder Representations from Transformers) is a widely used machine learning model in natural language processing (NLP). Developed by Google in 2018, it has demonstrated remarkable performance in text understanding and generation. BERT's key features include 1) bidirectional context understanding, 2) the use of a Transformer architecture, 3) pre-training and fine-tuning, and 4) diverse application potential. First, BERT is a "bidirectional" model that understands words in a given text in both left- and right-hand context. While previous models only considered context in one direction (from left to right or vice versa), BERT can comprehensively understand the entire text. Furthermore, BERT is built on a neural network architecture called the Transformer. Transformers use an "attention mechanism" to capture the relationships between words in the input text. This enables more complex and sophisticated text understanding. Furthermore, BERT is "pre-trained" on large amounts of text data to acquire a general understanding of language. It can then be "fine-tuned" for specific tasks (e.g., sentiment analysis or question answering) to optimize it for specific uses. BERT is applicable to a wide range of NLP tasks, including text classification, question answering, sentiment analysis, and machine translation.

[0043] Two main tasks are used to train the BERT model. The MLM task involves randomly "masking" words from input text and having BERT predict the masked words. The main goal of this task is to train BERT to understand the meaning of words using context. Because it takes bidirectional context into account, the model utilizes information from the entire sentence to predict the masked word. The NSP task asks BERT to determine whether two sentences are consecutive. The model is given one sentence (A) and another sentence (B) and is asked to predict whether B immediately follows A. Half the time, B is the actual sentence that follows A, and the other half is a randomly selected, unrelated sentence. The goal of the NSP task is to train BERT to understand the relationships between sentences. This ability is particularly useful in tasks where understanding the connection and semantic flow of sentences is important, such as question answering and text summarization. These tasks train BERT to understand not only the word level but also the relationships between entire sentences and multiple sentences. As a result, BERT can achieve high performance in a variety of natural language processing tasks.

[0044] In some embodiments, machine learning may utilize DNABERT, a pre-trained Transformer-based base model (Yanrong Ji, Zhihan Zhou, Han Liu, Ramana V Davuluri, "DNABERT: Pre-Trained Bidirectional Encoder Representations from Transformers Model for DNA-Language in Genome," Bioinformatics, Volume 37, Issue 15, August 2021, Pages 2112-2120). DNABERT is a pre-trained bidirectional encoder representation that can capture a global understanding of genomic DNA sequences based on the context of upstream and downstream base sequences. DNABERT does not perform the NSP task used in BERT, but instead trains only on the MLM task. A certain percentage of the DNA sequence is masked, and k-mer tokens at the masked sites are predicted. The training data used for pre-training DNABERT are DNA sequences sampled from the human genome. The present inventors have demonstrated that fine-tuning a pre-trained basic model such as DNABERT can generate a machine learning model that predicts gene expression induction activity in plant cells from promoter sequences. The machine learning model disclosed herein can handle not only binary classification problems but also regression problems, and as a result, can predict the level of gene expression induction activity.

[0045] One aspect of the present disclosure relates to a method for genome editing of a cell to regulate the expression level of a desired gene. In some embodiments, the method may include obtaining a modified promoter sequence having a desired activity by a method for obtaining a promoter sequence modified to have a desired activity according to the present disclosure, providing a cell to be subjected to genome editing, and editing the genome of the cell to produce the modified promoter sequence.

[0046] In some embodiments, the cells are plant cells, including cells from a wide range of plants, including bryophytes, ferns, gymnosperms, and the angiosperms Magnolias, monocots, and eudicots (Rosaceae I, Rosaceae II, Chrysanthemum I, Chrysanthemum II, and their outgroups). More specific examples of plants include nightshades such as tomatoes, bell peppers, chili peppers, eggplants, tobacco, and torvum; gourds such as cucumbers, pumpkins, melons, and watermelons; vegetables such as cabbage, broccoli, Chinese cabbage, and kale; fresh and spicy vegetables such as shiso, celery, parsley, and lettuce; alliums such as leeks, onions, and garlic; other fruit vegetables such as strawberries and melons; taproots such as radishes, turnips, carrots, and burdock; tubers such as taro, cassava, potato, sweet potato, and Chinese yam; grains such as rice, corn, wheat, sorghum, barley, rye, wheatgrass, and buckwheat; soybeans, adzuki beans, mung beans, cowpeas, and yams. Examples of suitable crops include legumes such as columbine, peanut, pea, and broad bean; soft vegetables such as asparagus, spinach, and mitsuba; floral plants such as lisianthus, rose, stock, carnation, and chrysanthemum; turf grasses such as bentgrass and Korean lawn grass; oilseed crops such as rapeseed, camelina, rapeseed, jatropha, sesame, and perilla; fiber crops such as cotton, rush, and hemp; forage crops such as clover, dent corn, and alfalfa; deciduous fruit trees such as apples, pears, grapes, and peaches; citrus fruits such as Satsuma mandarins, oranges, lemons, and grapefruit; and woody plants such as azalea, azalea, cedar, poplar, and rubber tree.

[0047] One aspect of the present disclosure relates to a polynucleotide having promoter activity, the polynucleotide having a sequence obtained by the method of the present disclosure. Such a polynucleotide may have the sequence of SEQ ID NO: 2 or 4, for example.

[0048] One aspect of the present disclosure relates to a method for producing a genome-edited plant in which the expression level of a desired gene is regulated. In some embodiments, the method may include editing the genome of a desired plant cell using a cellular genome editing method for regulating the expression level of a desired gene according to the present disclosure, and obtaining a plant derived from the genome-edited cell. Furthermore, one aspect of the present disclosure relates to a genome-edited plant produced by the production method according to the present disclosure.

[0049] To obtain genome-edited cells, the genome-editing enzyme can be introduced into plant tissues or cells in the form of DNA, RNP, or protein. Examples of the site of introduction of the genome-editing enzyme in plants include flowers (egg cells, pollen, petals, etc.), stems (cambium, pith, cortex, etc.), leaves (including leaf primordia), roots, shoot tips, lateral buds, flower buds, root tips, protoplasts, etc.

[0050] The introduction method is not particularly limited and can be appropriately selected depending on the plant species and the target cells / tissues for introduction. Examples of the introduction method include the Agrobacterium method, particle gun method, whisker method, nanopipette method, and virus-mediated nucleic acid delivery.

[0051] In some embodiments, obtaining an individual plant derived from a genome-edited cell may include, for example, steps such as (a) culturing the genome-edited cell or tissue containing the edited cell under appropriate culture conditions, (b) proliferating the cells to form a callus (an undifferentiated cell mass), (c1) inducing a shoot (a new sprout) from the callus, (c2) directly forming a shoot from the tissue containing the edited cell, (d) inducing roots in the shoot to regenerate a plant, (e1) obtaining seeds of the next generation from the regenerated plant, or (e2) propagating a part of the regenerated plant by cutting.

[0052] One aspect of the present disclosure relates to an information processing device that predicts promoter sequences that have been modified to have a desired activity (see FIG. 11).

[0053] In some embodiments, the present apparatus 010 may be an information processing device including a sequence input unit 011 that accepts input of an original promoter sequence to be modified; a modified sequence generation unit 012 that generates a set of multiple modified promoter sequences that can be created using genome editing technology based on the promoter sequence; an activity prediction unit 013 that predicts the activity of each modified promoter sequence included in the generated set of modified promoter sequences using a machine learning model; and a sequence selection unit 014 that selects modified promoter sequences predicted to have the desired activity. The sequence input unit 011 may be connected to an input device such as a communication device for an external network or a keyboard. The sequence selection unit 014 may be connected to an output device such as a display or printer, or a communication device for an external network. Those skilled in the art will be able to understand the configuration necessary for the present apparatus to perform the sequence acquisition method disclosed herein in light of the disclosure herein.

[0054] One aspect of the present disclosure relates to a non-transitory computer-readable medium having instructions or programs stored thereon.

[0055] In some embodiments, the computer-readable medium may be a computer-readable medium having stored thereon instructions or a program that, when executed by a processor, can execute the following steps: step S100 of accepting input of an original promoter sequence to be modified, step S110 of generating a set of multiple modified promoter sequences that can be created using genome editing technology based on the promoter sequence, step S120 of predicting the activity of each modified promoter sequence included in the generated set of modified promoter sequences using a machine learning model, and step S130 of selecting the modified promoter sequence predicted to have the desired activity. Figure 12 shows an exemplary flowchart of such a program.

[0056] Step S110 of generating a set of multiple modified promoter sequences that can be created using genome editing technology may further include step S112 of comprehensively searching for Cas9-cleavable sites (PAM sequences) from the input sequences, step S114 of combinatorially selecting all two of the searched cleavable sites (PAM sequences), and step S116 of deleting the sequence between the two cleavage sites selected for all combinations to generate a set of modified promoter sequences. Furthermore, step S120 of predicting the activity of each modified promoter sequence using a machine learning model may be performed using only the core promoter portion (150-200 bp, e.g., approximately 170 bp, from the 3' end) of each modified promoter sequence.

[0057] One aspect of the present disclosure relates to a computer program that predicts promoter sequences that have been modified to have a desired activity.

[0058] In some embodiments, the program may be a program that causes a computer to perform the functions of accepting input of an original promoter sequence to be modified, generating a set of multiple modified promoter sequences that can be created using genome editing technology based on the promoter sequence, predicting the activity of each modified promoter sequence included in the generated set of modified promoter sequences using a machine learning model, and selecting a modified promoter sequence predicted to have the desired activity.

[0059] Fig. 13 is a schematic diagram showing an exemplary embodiment of the device of the present disclosure. In Fig. 13, reference numeral 100 denotes a computer, which includes a control unit 101, a storage unit 102, a peripheral device I / F unit 103, an input unit 104, a display unit 105, and a communication unit 106, all of which are connected by a bus 110. Note that this configuration is merely an example, and various other configurations may be adopted as appropriate.

[0060] The control unit 101 is composed of a CPU (Central Processing Unit), ROM (Read Only Memory), RAM (Random Access Memory), etc. The CPU loads programs stored in the memory unit 102, ROM, recording medium, etc. into a work memory area on the RAM and executes them, driving and controlling each device connected via the bus 110 and realizing the processing performed by the computer. The ROM is a non-volatile memory that stores programs, data, etc., such as the boot program and BIOS of the computer 100. The RAM is a volatile memory that temporarily stores programs, data, etc. loaded from the memory unit 102, ROM, recording medium, etc., and also provides a work area used by the control unit 101 when performing various processing. The memory unit 102 is, for example, a hard disk drive (HDD) and stores the programs executed by the control unit 101 and various other data.

[0061] The peripheral device I / F (interface) unit 103 is a port for connecting the computer 100 to peripheral devices. The peripheral device I / F unit 103 is configured with a USB, IEEE1394, RS-232C, or the like. The connection with the peripheral devices may be wired or wireless. The input unit 104 has input devices such as a keyboard, a pointing device such as a mouse, and a numeric keypad, and issues operation instructions, operational instructions, data input, and the like to the computer 100. The display unit 105 is a logic circuit or device driver for displaying videos, images, and the like on a display device such as a liquid crystal panel. The input unit 104 and the display unit 105 can also be configured integrally as a touch display.

[0062] The communication unit 106 has a communication control device, a communication port, etc., and is a wired or wireless communication interface that mediates communication with the network 120. The bus 110 is a communication path that mediates the exchange of control signals, data signals, etc. between the devices. The network 120 may further be connected to an external server 130 and a net storage 140.

[0063] For example, a program recorded on a computer-readable medium according to the present disclosure can be loaded into the device of Figure 13, and the computer can function as an information processing device equipped with a sequence input unit that accepts input of an original promoter sequence to be modified, a modified sequence generation unit that generates a set of multiple modified promoter sequences that can be created using genome editing technology based on the promoter sequence, an activity prediction unit that predicts the activity of each modified promoter sequence included in the generated set of modified promoter sequences using a machine learning model, and a sequence selection unit that selects modified promoter sequences predicted to have the desired activity.

[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can be used to practice or test the present invention, some potentially preferred methods and materials are now described. All publications mentioned herein are incorporated by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. It is understood that the present disclosure supersedes the disclosure of the incorporated publication in the event of a conflict.

[0065] Where a range of values ​​is described, unless the context clearly dictates otherwise, it is understood that each intervening value, to the tenth of the unit of the lower limit, between the upper and lower limits of that range is also specifically disclosed. Each smaller range between any stated or intervening value in a stated range and any other stated or intervening value within that stated range is also encompassed within the disclosure. The upper and lower limits of these smaller ranges may independently be included or excluded, and each range that includes either, either, or both limits in the smaller ranges is also encompassed within the invention, subject to the specifically excluded limits in the stated range. When a stated range includes one or both limits, ranges excluding either or both of those included limits are also included within the invention. The term "about" with respect to a numerical value means within 5%.

[0066] The embodiments described herein are intended to be merely illustrative, and those skilled in the art will be able to make numerous variations and modifications without departing from the spirit of the present invention. Furthermore, certain variations and modifications may produce less than optimal results, but still provide satisfactory results. All such variations and modifications are intended to be within the scope of the present invention as defined by the appended claims. Furthermore, any combination of the components disclosed herein, or any transformation of the disclosed expression into a method, apparatus, system, computer program, data structure, recording medium, or the like, is also valid as an aspect of the present disclosure. Therefore, details described regarding the method of the present disclosure may also be applied to the system, computer program, data structure, recording medium, or the like.

[0067] Example 1: Construction of a machine learning model A machine learning model that predicts gene expression induction activity in plant cells from promoter sequences was constructed as follows. First, the data shown in a paper by Jores et al. was obtained as the learning source data (Jores, T., Tonnies, J., Wrightsman, T. et al. Synthetic promoter designs enabled by a comprehensive analysis of plant core promoters. Nat. Plants 7, 842-855 (2021). https: / / doi.org / 10.1038 / s41477-021-00932-y). The dataset was an overview of approximately 70,000 datasets, each containing a 170-base DNA sequence and an expression intensity.

[0068] For machine learning, we used DNABERT, a pre-trained Transformer-based model. DNABERT is reported, for example, in the aforementioned paper by Ji et al. DNABERT primarily focuses on binary classification problems, and has been shown to be able to determine, for example, whether a given base sequence contains a splicing site. A binary classification problem is one in which a sequence is classified into two categories, such as 1 if it contains a splicing site and 0 if it does not. In contrast, a problem such as "predicting downstream gene expression levels for a given base sequence" is called a regression problem, and requires the output of a continuous value. Therefore, to handle continuous values ​​in DNABERT, we added classes to the Transformer and trained the model.

[0069] Example 2: Prediction and classification of promoter strength using a trained model Promoter sequences are often located upstream of genes and control the transcription level (expression level) of genes directly downstream. The expression level of a gene is thought to be controlled by the promoter sequence. The expression level of a downstream gene is defined as promoter strength. Each promoter sequence is thought to exhibit a unique strength.

[0070] As explained in Example 1, a machine learning model for predicting transcription activity was constructed based on the base sequence. Using this model, a list was created in which the transcription activity of all genes in Arabidopsis thaliana was evaluated.

[0071] All promoters were classified into three groups based on their predicted transcriptional activity: 1. a "high expression group" predicted to have high promoter strength; 2. a "low expression group" predicted to have very low or almost no promoter strength; and 3. a "moderate expression group" showing promoter strength between the above.

[0072] Example 3: Comparison of model predictions and actual gene expression Seven promoters were selected from the high expression group, eight promoters from the medium expression group, and four promoters from the low expression group, and artificial synthetic genes were created by connecting them to the luciferase (LUC) gene.

[0073] The gene structures are shown in Figure 1. The promoters have 19 different sequence patterns. The downstream LUC gene and the upstream cauliflower mosaic virus (CaMV) 35S enhancer sequences are common to all artificial genes.

[0074] The plasmid vector carrying this artificial gene was transfected into protoplasts extracted from Arabidopsis leaves using the polyethylene glycol method. The protoplasts expressed the LUC gene at various levels. To evaluate the LUC expression level, luminescence intensity was measured using a plate reader. Figure 2 shows the relationship between the measurement results and the predicted promoter strength. A scatter plot is shown of the relationship between the actual measured LUC gene expression level and the predicted promoter strength. The vertical axis indicates the predicted transcription activity. A higher value is expected to indicate greater expression of downstream genes. The horizontal axis indicates the actual measured LUC expression intensity, normalized to the expression level of the CaMV promoter, a positive control (P control).

[0075] The results shown in Figure 2 show that the actual measured values ​​are divided into three groups, consistent with the groups based on the predicted values. As such, a clear correlation was confirmed between the predicted values ​​by the constructed machine learning model and the actual measured values. This figure demonstrates the high level of predictability. In this figure, linear convergence means that the predicted values ​​and the actual measured values ​​are in agreement, demonstrating high predictive performance. This predictive performance was evaluated using the correlation coefficient R, which was R = 0.9296673.

[0076] Next, based on the correlation shown in the scatter plot, we calculated the LUC expression levels corresponding to the predicted transcriptional activity using linear regression prediction. The comparison is shown in Figure 3.

[0077] The vertical axis plots LUC luminescence intensity. The luminescence intensity of the P control is expressed relative to 1. The cauliflower mosaic virus (CaMV) promoter, known for its high transcriptional activity, was used as the P control. The horizontal axis shows the numbers of 19 promoters. Promoter numbers 3–10 were classified as the "high expression" group, predicted to have high expression intensity; promoter numbers 13–21 as the "moderate" group; and promoter numbers 22–25 as the "low expression" group. The black bars represent the actual values ​​(expression levels) measured by the LUC assay. The white bars represent the predicted expression levels using the machine learning model. Focusing on the black bars in Figure 3, the "high expression" group showed relatively high expression levels, while the "low expression" group showed low expression levels. The "moderate" group showed intermediate values.

[0078] Next, we checked how accurately the expression levels of individual promoters were predicted. Comparing the black and white bars for each promoter reveals that promoters 3, 4, 6, 7, 9, 10, 13, 15, 16, 20, and 24 were predicted with a high degree of agreement. On the other hand, for promoters 21 and 22, the actual values ​​were significantly higher than the predicted values. Conversely, for promoters 5, 17, and 19, the predicted values ​​were significantly higher than the actual values.

[0079] In summary, the constructed model was able to predict general trends such as high, medium, and low expression. For promoters with low expression, the tendency for low expression was predicted with high accuracy overall. For the high expression group, expression was predicted within a range of 50% to 280% of the actual measured value.

[0080] In this way, the machine learning model according to the present disclosure can predict gene expression levels based on promoter base sequences.

[0081] Example 4: Prediction and Demonstration of Reduced-Activity Promoters Next, we devised a method to reduce the transcriptional activity of promoters. Each promoter contains many Cas9 PAM recognition sequences (NGG). We searched these sequences and compiled a list of approximately several thousand possible editing patterns.

[0082] More specifically, the procedure is as follows: 1. Comprehensively search for positions (PAM sequences) that can be cleaved by Cas9 within the given base sequence; 2. Select two of these; 3. Cut at the two cleavage sites and simulate the new core promoter sequence that appears after repair; 4. Estimate the strength of the simulated sequence as a core promoter; 5. Repeat the above steps for all possible pairs of cleavage sites; 6. Select the one with the lowest score (or the highest score if selecting a promoter with increased transcriptional activity) from all simulated core promoters.

[0083] This allows us to determine which cleavage site to select when enhancing or suppressing expression levels. In this way, we estimated the promoter transcriptional activity for each editing pattern's base sequence and searched for editing patterns that could reduce promoter transcriptional activity. The results are shown in Figure 4. We searched for editing patterns that could reduce promoter strength for promoters 3, 4, 5, 6, 9, 10, and 21. The black triangle plots show the predicted LUC assay scores calculated based on the predicted new promoter strength. As a result, we predicted that LUC expression levels would be approximately 14% to 1% of the transcriptional activity of the original promoter. The deletion lengths for each promoter are shown in the table below.

[0084]

[0085] Next, genes with these theoretical promoter sequences were artificially synthesized, transfected into protoplasts in the same manner, and expression levels were measured using a plate reader. The results are shown in Figure 5. The shaded bar graphs overlaid on Figure 4 show the actual measured LUC expression levels for the new promoter sequences. The vertical axis has been expanded due to the extremely low expression levels. Promoter 5 was left blank. For all six promoters for which measurements were obtained, significant decreases in expression levels were observed, consistent with the predicted values. For example, promoter 3, whose sequence is CGGAAACTTGTCACTTCCTTTACATTTGAGTTTCCAACACCTAATCACGACAACAATCATATAGCTCTCGCATACAAACAAACATATGCATGTATTCTTACACGTGAACTCCATGCAAGTCTCTTTTCTCACCTATAAATACCAACCACACCTTCACCACATTCTTCACT (SEQ ID NO: 1), had a predicted transcriptional activity of 5.43, and the LUC expression level predicted by linear regression from this value was 111.5% of the P control. The actual LUC expression level was 68.6%. In contrast, the edited sequence of promoter 3 was GAAACTGATTAGCTCCTATCAGTTCAGCAAACCACAAGCTGAAGAATCCAAGACTTGAGAAACAAATTTACAAAAGCCCATGTTCCAATCAAAACTGTTACCAAACATCTGAAATAGATCTAAATGAGCGTTGGTATAATTGAAACTTACCGAAGGCCCACATTCTTCAC (SEQ ID NO: 2), which results from truncating 1,824 bases between -1,833 and -9. The predicted transcriptional activity was 0.393, and the predicted LUC expression level by linear regression using this value was 7.98% of the P control. The actual LUC expression level was 3.41% of the P control.

[0086] In this way, expression can be reduced by editing the promoter base sequence based on the predictions of the machine learning model according to the present disclosure.

[0087] Example 5: Prediction and Demonstration of Promoters with Enhanced Activity Finally, a method for increasing promoter activity was devised. As in Example 4, promoter sequences with possible editing patterns were searched for those whose predicted promoter strength increased after editing. The results are shown in Figure 6. For promoters 13, 17, 22, 23, 24, and 25, promoters with increased promoter strength were selected, and the predicted expression levels are indicated by white circles. LUC expression levels were predicted to be 7.4 to 125 times higher than the original promoter. The deletion lengths for each promoter are shown in the table below.

[0088]

[0089] Next, genes containing these theoretical promoter sequences were artificially synthesized and transfected into protoplasts in the same manner. Expression levels were measured using a plate reader. The results are shown in Figure 7. Of the six promoters, five showed increased expression levels compared to the original promoters, but only promoter 13 exceeded the predicted level. For example, promoter 13's sequence is TCAAGCAATCATTATCGACTACGGTCGTTCGTTAAAGATCATGCATGTGCTTAGTGGCAATACCCTACGCATCTTGATTCGTTACTGCGGCACGTGTCATGACCATGCACATGAATGATGATTAATGTTTAGTACATATAATGTTCACGCAAACGCATAGTGTTAGGAAA (SEQ ID NO: 3). The predicted transcriptional activity was 2.00, and the LUC expression level predicted by linear regression from this value was 6.94% compared to the P control. The actual LUC expression level was 6.60% compared to the P control. In contrast, the edited sequence of promoter 13 was GAAACTTGAAAATCAAATCAGTGAGTCGCAAGTAAGACTTTGTGGTTGTTGTATCAGATTTCGCCGTGCGCATCTTGATTCGTTACTGCGGCACGTGTCATGACCATGCACATGAATGATGATTAATGTTTAGTACATATAATGTTCACGCAAACGCATAGTGTTAGGAA (SEQ ID NO: 4), the predicted transcription activity was 4.18, and the LUC expression level predicted by linear regression from this value was 36.7% compared to the P control. The actual measured LUC expression level was 69.5% compared to the P control.

[0090] In this way, expression can be increased by editing the promoter base sequence based on the predictions of the machine learning model of the present disclosure.

[0091] Example 6: Program for changing promoter function by genome editing The present inventors developed a program used to change promoter function by genome editing using the following procedure.

[0092] 1. Input base sequence X. For example, 2,000 bases from -1,995 to +5 around the transcription start point of the target gene. 2. Search for the PAM sequence within base sequence X. (Cas9: NGG or Cas12a: TTTV sequence) 3. Based on the PAM sequence, a list of cleavage positions is stored in the sequence. Here, for Cas9, the cleavage position is 4 bases downstream of the PAM sequence, and for Cas12a, the cleavage position is 18 bases downstream of the PAM sequence. For example, approximately 50 cleavage positions are discovered within 2,000 bases. 4. Of n cleavage positions, there are nC2 combinations for specifying two. For example, with 50 cleavage positions, there are 1,250 possible combinations. For each of the nC2 combinations, generate a truncated (deleted) base sequence between the two cleavage positions. 5. For each generated base sequence, obtain 170 bases from the 3' end and input them into the learning model. 6. The learning model outputs the strength (transcriptional activity) of the given base sequence as a core promoter. 7. Overall, the following three types of information are obtained: A. Transcriptional activity of base sequence X. B. A set of cleavage sites where the promoter activity of the sequence obtained by truncating between the two cleavage sites is greater than that of A. C. A set of cleavage sites where the promoter activity of the sequence obtained by truncating between the two cleavage sites is less than that of A. 8. For example, to increase the expression of a target gene, two types of guide RNAs targeting the cleavage sites obtained in 7.B. are simultaneously introduced into cells, and genome editing using CRISPR-Cas9 or -Cas12a is performed, which is expected to produce individuals with a deletion between the two cleavage sites.

[0093] In particular, it should be noted that the simultaneous use of two different guide RNAs is intended to induce small to medium-sized deletions within the category SDN1.

[0094] Example 7: Visualization of predicted promoter activity Finally, the program described in Example 6 provides a set of predicted promoter activity values ​​for the positions of the two guide RNAs and the deleted sequence, but these values ​​are difficult to understand intuitively. Therefore, the present inventors devised a method to visualize these values.

[0095] As a first visualization method, we devised a method for analyzing the overall structure of input base sequence X. The vertical axis represents the predicted promoter strength, and the horizontal axis represents the position of the base sequence. Each plot shows the predicted transcription activity value for a 170-base window of DNA sequence at a certain position. For example, the point at position 0 on the horizontal axis represents the predicted transcription activity for 170 bases from the 5' end of the input base sequence X. The resulting transcription activity was 1.24 (this 170-base sequence is not adjacent to any gene, so it is unlikely to function as a core promoter. However, if a gene were connected directly below this 170-base sequence, the predicted transcription strength would be 1.24). Therefore, the point (0, 1.24) was plotted. Next, the window was slid by one base, and the base sequence from bases 2 to 171 of base sequence X was extracted and predicted in the same way. The transcription activity was estimated to be 1.37. Therefore, it was plotted at the point (1, 1.37). This process was repeated until a total of 1,830 points were plotted (Figure 8).

[0096] In Figure 8, the point on the far right (1829, 1.47) represents the 170 bases immediately upstream of the gene. If the sequence between any point with a value higher than this and the 3' end can be removed by genome editing, gene expression is expected to increase. Similarly, it is also possible to decrease gene expression. Using this diagram provides an overview of sequence X and clarifies which part of the guide RNA should hybridize to achieve the desired expression level.

[0097] The visualization method described above may be effective in providing a rough overview of the target sequence. However, it does not show the promoter strength or the relative positions of the two guide RNAs that need to be designed. To overcome this issue, we devised a different visualization method. Figure 9 below shows a visualization using this method.

[0098] Guide RNAs designed close to the target gene are called proximal guides, while guide RNAs designed farther away are called distal guides. The visualization was based on the positions of the available guide RNAs (position of the PAM sequence). Note that only the first 1,000 bases immediately upstream of the gene are shown.

[0099] The horizontal axis shows the position of the hybridizing base of the distal guide in sequence X. In other words, it shows the distance between the gene's transcription start point and the distal guide. Similarly, the vertical axis shows the distance between the proximal guide and the gene's transcription start point. Here, the proximal guide is never designed to be farther away than the distal guide, so half of the plot in Figure 9 is hidden. Points on the line y = x indicate that the proximal and distal guides are both located at approximately the same distance from the gene, indicating that they are close to each other. In this case, the deleted bases are small, resulting in minimal editing and is desirable. 170 bases were extracted from the 3' end of the new sequence created by deleting the sequence between the proximal and distal guides, and transcriptional activity was predicted. The plot color is changed based on the predicted value. In the color chart, light gray indicates high predicted transcriptional activity, and black indicates low predicted transcriptional activity.

[0100] For example, let's take the plot of point (849, 17). The distal guide position for this point is 849 bases from the gene. The proximal guide position is 17 bases from the gene. The 832 bases between these two points are deleted from base sequence X to create base sequence X'. 170 bases are removed from the 3' end of base sequence X', and the predicted transcription activity is 4.02. A light gray plot is drawn according to the color chart. The diagram is created by performing the above steps for all combinations of distal and proximal guides.

[0101] If we wanted to increase the expression level of this gene, we could simply select a point with a transcriptional activity greater than the original 2.34 and use the distal and proximal guide positions to determine the location of the guide RNA we should design. For example, selecting points (849, 17) or (841, 29) would result in high transcriptional activity. Selecting points (432, 110) or (271, 77) would result in low transcriptional activity. Visualizing the promoter structure in this way may simplify the design of guide RNAs that achieve the desired promoter structure.

[0102] After obtaining the location of the guide RNA's target sequence in this way, the base sequence of that location was obtained, and a guide RNA was designed using 20 or 23 bases. This guide RNA sequence was evaluated and verified using tools such as CRISPOR to assess specificity and off-target effects. If an appropriate guide RNA was successfully designed, a vector carrying the guide RNA was created using artificial DNA synthesis and PCR. From this point, genome editing can be performed according to standard methods.

[0103] Example 8: Example of a soybean gene promoter The training data for the machine learning model were promoter sequences from Arabidopsis thaliana, sorghum, and maize. Therefore, it was unclear whether the prediction system of the present invention would function as expected for promoter sequences from organisms other than these three species. Therefore, we investigated the enhancement of transcriptional activity for specific soybean genes. The results are shown in Figure 10.

[0104] The transcriptional activity of a 170-base sequence in the core promoter of a soybean gene was predicted to be 0.927494. Based on this value, the predicted LUC expression level was 1.26% of the positive control. This promoter sequence was artificially synthesized and introduced into protoplasts isolated from Arabidopsis leaves, and an LUC assay was performed. The result was 4.15%.

[0105] Based on this promoter, two types of promoter sequences were created after genome editing to increase gene expression. These are called edit1 and edit2. The transcriptional activity of edit1 was predicted to be 2.950119. The LUC expression level was predicted to be 13.2%. The actual LUC assay value for this promoter was 19.0%.

[0106] The predicted transcriptional activity of edit2 was 2.432977, and the predicted LUC expression level was 7.27%. The actual LUC level measured by the LUC assay for this promoter was 41.6%.

[0107] Construction of Cas9 Expression Vectors. We constructed a guide RNA expression cassette for the edit1 / edit2 promoter sequence, a Cas9 gene expression cassette, and a hygromycin resistance gene expression cassette, which were confirmed to be highly expressed in Arabidopsis protoplasts. Standard molecular biology techniques were used to construct the expression vectors. PCR amplification was performed using PrimeSTAR Max (Takara Bio Inc.). PCR products were purified using NucleoSpin Gel and PCR Cleanup (Machley-Nagel). Fragments excised from the plasmid DNA were assembled using the NEB Golden Gate Assembly Kit (BsaI-HF v2) (New England BioLabs Japan). Finally, the Cas9 gene expression cassette, guide RNA expression cassette, and hygromycin resistance gene expression cassette were inserted into the pTTK352 binary vector to obtain the genome editing tool expression vector.

[0108] Cultivation of soybean plants The soybean cultivar "Enrei" was used as plant material. Seeding was carried out in the following order: seed disinfection, humidity control, and sowing. For seed disinfection, the seeds were placed in a sealed container with a beaker containing 20 mL of bleach (Kao) and 1 mL of 12N hydrochloric acid for 16 hours. After seed disinfection, the seeds were placed in a sealed container with a beaker containing 100 mL of deionized water for humidity control at 4°C for 1-2 weeks. After humidity control, the seeds were placed on MS medium for sowing. For cultivation, the sown plants were grown in a climate control room set at 25°C with a 16-hour light / 8-hour dark period.

[0109] Soybean plants approximately one day after the start of cultivation were used for explants. They were cut in half under a stereomicroscope, and the hypocotyl was excised with a scalpel, then cut again vertically in half. The explants were scratched 10 times with a microbrush or scalpel from the base of the hypocotyl to the tip. Isolation was carried out in a clean bench. The isolated explants were placed on solid coculture medium prepared as previously described (Tabei, 2012, Transformation Protocol [Plant Edition], Kagaku Dojin).

[0110] Gene introduction into explants: The Agrobacterium strain EHA105, transformed with a genome editing tool expression vector, was used. Agrobacterium infection and co-cultivation were performed in the following order: Agrobacterium pre-culture, preparation of an Agrobacterium suspension for infection, tissue immersion in the suspension, and co-cultivation. The medium used was prepared as previously described (Tabei, 2012, Transformation Protocol [Plant Edition], Kagaku Dojin). For Agrobacterium pre-culture, a glycerol stock of Agrobacterium was streaked onto LB medium and cultured at 28°C in the dark for 2 days. For the Agrobacterium suspension for infection, the pre-cultured Agrobacterium was resuspended in liquid co-cultivation medium and adjusted to an OD600 of 0.8. For tissue immersion in the suspension, explants were infiltrated with the Agrobacterium suspension for infection for 20 minutes. For co-cultivation, sterilized filter paper was placed on top of the solid co-cultivation medium, and the explants that had been immersed were placed on top of it and left to stand at 25°C with a 16-hour light / 8-hour dark period for 5 days.

[0111] The medium used for shoot regeneration was prepared as previously described (Tabei, 2012, Transformation Protocol [Plant Edition], Kagaku Dojin). After co-cultivation, the explants were washed with liquid adventitious bud induction medium and then inserted into solid adventitious bud induction medium. The explants were cultured at 25°C under a 16-hour light / 8-hour dark cycle for 2 weeks. The elongated shoots (not recombinant) were then excised and transferred to solid adventitious bud induction medium supplemented with 20 mg / L hygromycin for an additional 2 weeks. The browned areas were then removed with a scalpel, and the explants were transferred to adventitious bud elongation medium supplemented with 20 mg / L hygromycin. Hygromycin-resistant shoots were obtained by transferring the regenerated shoots to rooting medium. After rooting, the explants were potted and grown in a closed greenhouse for seed harvesting.

[0112] Expression analysis in the next generation (T1 generation): After sowing the seeds, DNA was extracted from the seedlings, and individuals with edit1 or edit2 editing in the next generation (T1) were identified and selected by sequencing analysis. RNA was then extracted from the selected individuals and reverse-transcribed, and expression analysis was performed using real-time PCR on soybean target genes containing the edit1 or edit2 promoter sequence.

[0113] Example 9: Genome Editing with a Single Target (Single Guide RNA) The graph in Figure 14 shows the results of in silico genome editing of the promoter sequence of the Arabidopsis thaliana AT4G39130 gene with a single target (single guide RNA). Sequences were created by adding 1- to 5-bp deletions, typically introduced by Cas9 cleavage, to six PAM sequences, and gene expression activity was predicted. Results showed that the activity did not change significantly from the intact sequence (without genome editing). This result demonstrates that simply applying genome editing technology to a single-target promoter region does not produce the desired modifications. Thus, our extensive research has revealed that single-target deletion (which typically results in a 1- to 5-bp deletion) does not significantly alter expression intensity. This was an unexpected result, demonstrating that the single-target deletion we initially anticipated would not produce the results we expected.

[0114] On the other hand, as described in Examples 4 and 5, when two deletion targets were performed, gene expression levels were predicted to increase by more than six-fold or decrease to 1 / 8 or less compared to when a single deletion target was performed (see Figure 15). The inventors speculated that this is because two deletion targets can cause longer deletions (also called large deletions) compared to when a single deletion target was performed. While it is academically known that deletion occurs when two gRNAs are used, it is inefficient, has few applications, and is not suitable for exploratory tasks (there are inhibiting factors). Given this, this was a very surprising discovery and resulted in unexpectedly favorable effects.

[0115] DESCRIPTION OF THE SYMBOLS 010: Information processing device 011: Sequence input unit 012: Modified sequence generation unit 013: Activity prediction unit 014: Sequence selection unit 100: Computer 101: Control unit 102: Memory unit 103: Peripheral device I / F unit 104: Input unit 105: Display unit 106: Communication unit 110: Bus 120: Network 130: External server 140: Database

Claims

1. A method for obtaining a promoter sequence modified to have a desired activity, comprising: preparing an original promoter sequence to be modified; generating a set of a plurality of modified promoter sequences that can be created by genome editing technology based on the promoter sequence; predicting the activity of each modified promoter sequence included in the generated set of modified promoter sequences using a machine learning model; and selecting a modified promoter sequence predicted to have the desired activity.

2. The method according to claim 1, wherein the desired activity is a gene expression induction activity higher than that of the original promoter sequence or a gene expression induction activity lower than that of the original promoter sequence.

3. The method according to claim 1, wherein the promoter sequence is a promoter sequence of a plant cell.

4. The method according to claim 1, wherein the original promoter sequence includes a core promoter sequence and a sequence upstream thereof.

5. The method according to claim 1, wherein the genome editing technology is a genome editing technology using the CRISPR / Cas system.

6. The method according to claim 5, wherein the set of a plurality of modified promoter sequences is generated by a sequence deletion caused by cleavage induced by a combination of guide RNA sequences designed based on two PAM recognition sequences.

7. The method according to claim 1, wherein the set of modified promoter sequences includes at least 1,000 different sequences.

8. The method according to claim 1, wherein the machine learning model is a regression model trained to predict gene expression induction activity from a promoter sequence using gene expression induction activity data of a plurality of promoter sequences in plant cells as teacher data.

9. The method according to claim 1, further comprising a step of visualizing the activity of each modified promoter sequence included in the generated set of modified promoter sequences on a computer display.

10. An information processing apparatus for predicting a promoter sequence modified to have a desired activity, comprising: a sequence input unit that receives an input of an original promoter sequence to be modified; a modified sequence generation unit that generates a set of a plurality of modified promoter sequences that can be created by a genome editing technique based on the promoter sequence; an activity prediction unit that predicts the activity of each modified promoter sequence included in the generated set of modified promoter sequences using a machine learning model; and a sequence selection unit that selects a modified promoter sequence predicted to have a desired activity.

11. A non-transitory computer-readable medium storing instructions that, when executed by a processor, perform the following steps: receiving an input of an original promoter sequence to be modified; generating a set of a plurality of modified promoter sequences that can be created by a genome editing technique based on the promoter sequence; predicting the activity of each modified promoter sequence included in the generated set of modified promoter sequences using a machine learning model; and selecting a modified promoter sequence predicted to have a desired activity.

12. A program for causing a computer to realize a function of receiving an input of an original promoter sequence to be modified, a function of generating a set of a plurality of modified promoter sequences that can be created by a genome editing technique based on the promoter sequence, a function of predicting the activity of each modified promoter sequence included in the generated set of modified promoter sequences using a machine learning model, and a function of selecting a modified promoter sequence predicted to have a desired activity.

13. A method for genome editing of a cell for regulating the expression level of a desired gene, comprising: obtaining a modified promoter sequence having a desired activity by the method according to claim 1; preparing a cell to be subjected to genome editing; and editing the genome of the cell so as to generate the modified promoter sequence.

14. A method for producing a genome-edited plant in which the expression level of a desired gene is regulated, the method comprising editing the genome of a desired plant cell by the method according to claim 13, and obtaining a plant individual derived from the genome-edited cell.

15. A machine learning model for predicting the gene expression induction activity in a plant cell from a promoter sequence, which is a regression model trained to predict the gene expression induction activity from a promoter sequence using gene expression induction activity data of a plurality of promoter sequences in a plant cell as teacher data.

16. A method for generating a machine learning model for predicting the gene expression induction activity in a plant cell from a promoter sequence, the method comprising training the model using gene expression induction activity data of a plurality of promoter sequences in a plant cell as teacher data.

17. The method according to claim 16, wherein a transformer-based pre-trained base model is used as the model to be trained.

Citation Information

Patent Citations

  • Methods for predicting pathogenicity of gene sequence variants

    JP2018527647A

Cited By

  • Application of recombinant promoter in improvement of bovine muscle performance

    CN121687205A