High-throughput evaluation method for in-vitro activity of crispr / cas system, activity prediction model of crispr / cas system, and utilization thereof
A high-throughput method for assessing CRISPR/Cas system activity in vitro involves transfecting cells with target-guide pairs, immobilizing them, and sequencing reactions, overcoming existing limitations and enabling efficient prediction and selection of active guide nucleic acids.
Patent Information
- Application Number
- PCT/KR2024/018654
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-08
- Filing Date
- 2024-11-22
- Publication Date
- 2025-05-30
AI Technical Summary
Current methods for assessing the in vitro activity of CRISPR/Cas systems are limited by the lack of high-throughput techniques, making it difficult to efficiently evaluate the nucleic acid cleavage activity of CRISPR/Cas systems in an in vitro environment.
A high-throughput method is developed by transfecting cells with vectors containing target-guide pairs, immobilizing cells to simulate an in vitro environment, independently inducing CRISPR/Cas system activation in each cell, and simultaneously sequencing the reactions to assess activity.
This method enables the high-throughput evaluation of CRISPR/Cas system activity, allowing for the prediction of target nucleic acid cleavage activity and the selection of guide nucleic acids with high selective cleavage activity, thereby enhancing the efficiency of nucleic acid enrichment and detection.
Smart Images

Figure KR2024018654_30052025_PF_FP_ABST
Abstract
Description
A high-throughput evaluation method for the in vitro activity of the CRISPR / Cas system, a prediction model for the activity of the CRISPR / Cas system, and its applications.
[0001] This specification discloses inventions related to the technical field of CRISPR / Cas systems. More specifically, the inventions disclosed in this specification relate to methods for assessing the in vitro target nucleic acid cleavage activity of a CRISPR / Cas system, a machine learning model for predicting the target nucleic acid cleavage activity of a CRISPR / Cas system, and methods for enriching rare nucleic acids using a CRISPR / Cas system.
[0002]
[0003] The CRISPR / Cas system is a type of programmable nuclease known to possess target-specific nucleic acid (DNA or RNA) cleavage activity. The CRISPR / Cas system can be utilized not only for in vivo gene editing but also for in vitro nucleic acid enrichment and detection. The technology to utilize the CRISPR / Cas system in vitro is also a critical research topic.
[0004] While methods for studying the function of the CRISPR / Cas system in vivo are well established, methods for studying its activity in vitro are relatively limited. In particular, high-throughput methods are essential for studying the outcomes of numerous reactions under slightly different reaction conditions. However, high-throughput methods for use in vitro are rare.
[0005] If large amounts of data are obtained using high-throughput methods, machine learning can be used to create predictive models. Indeed, models have been developed to predict the efficiency of indel introduction in vivo. However, models predicting the activity of the CRISPR / Cas system in vitro remain elusive due to the lack of suitable high-throughput methods.
[0006]
[0007] High-throughput assessment method for in vitro activity of CRISPR / Cas systems
[0008] This specification sets out the technical challenge of developing a high-throughput method for assessing the in vitro activity of the CRISPR / Cas system.
[0009] The CRISPR / Cas system can be utilized not only for in vivo gene editing but also for in vitro nucleic acid detection. In particular, specific applications such as nucleic acid detection, sample enrichment, and diagnostics utilize the CRISPR / Cas system in vitro, assuming it cleaves target nucleic acids. In these fields, developing a CRISPR / Cas system that cleaves target nucleic acids as intended in vitro is a critical challenge.
[0010] To discover a CRISPR / Cas system that exhibits nucleic acid cleavage activity suitable for its intended use, it is necessary to select an appropriate target nucleic acid and identify a guide nucleic acid that effectively targets it. To efficiently conduct this search on a large scale, an appropriate high-throughput method must be used. However, to the best of the inventors' knowledge, no method for assessing nucleic acid cleavage activity in a high-throughput in vitro environment has been properly developed. For a high-throughput method to be used, independent reactions must be processed in quantities of at least thousands, tens of thousands, or even hundreds of thousands or even millions. However, it is difficult to create independent reaction chambers for each target nucleic acid-guide nucleic acid pair while maintaining an in vitro environment. Moreover, since the CRISPR / Cas system cleaves a target nucleic acid, it is even more difficult to incorporate multiple components into each reaction chamber.
[0011] Conventional techniques for creating libraries of vectors encoding target nucleic acid-guide nucleic acid pairs, transfecting cells with these vectors, inducing a CRISPR / Cas system response, and then simultaneously sequencing them to identify CRISPR / Cas systems that function well in vivo are well-established (cited paper). However, because these conventional techniques are performed in vivo, they assess the CRISPR / Cas system's ability to induce indels in target nucleic acids, not its ability to cleave nucleic acids. Even when targeting the same target nucleic acid with a CRISPR / Cas system, there is a significant difference between its ability to induce indels in vivo and its ability to cleave nucleic acids in vitro (see the experimental examples in this application). Therefore, high-throughput activity assessment methods designed based on in vivo environments are difficult to apply to assessing activity in vitro.
[0012] There are important unmet technical requirements in this area.
[0013] Activity prediction model for the CRISPR / Cas system
[0014] This specification aims to develop an activity prediction model for the CRISPR / Cas system. Furthermore, the goal is to develop a method for selecting guide nucleic acids that selectively cleave specific nucleic acids by applying this activity prediction model.
[0015] In previous studies, models predicting the activity of the CRISPR / Cas system were constructed and used for the purpose of predicting 1) how well the guide nucleic acid would cleave the target nucleic acid, or 2) whether there would be any adverse effects of cleaving nucleic acids with a sequence nearly identical to the target nucleic acid (non-target nucleic acids). In other words, existing CRISPR / Cas system activity prediction models distinguished between on-target activity, which cleaves the target, and off-target activity, which cleaves non-targets, and were designed for that purpose.
[0016] In contrast, this specification aims to develop a model that predicts how well a specific guide nucleic acid will cleave a specific target nucleic acid, without distinguishing between on-target and off-target activity. Unlike conventional models, this model does not assume a perfect match between the guide and target nucleic acids, and must encompass cases where the mismatch between the guide and target nucleic acids is significant, allowing for diverse patterns of interaction.
[0017] We need a way to construct a new active prediction model and train it.
[0018] Rare nucleic acid sample enrichment method
[0019] This specification aims to develop a method for enriching rare nucleic acid (RNA) samples using the CRISPR / Cas system. Specifically, the goal is to develop a method for enriching rare nucleic acids in situations where the original sample contains trace amounts of rare nucleic acids and a large amount of background nucleic acids with similar sequences to the rare nucleic acids.
[0020] There are two main methods for enriching rare nucleic acids using the CRISPR / Cas system. One method involves cleaving nucleic acids other than the rare nucleic acid contained in the original sample (i.e., background nucleic acids) with the CRISPR / Cas system, and then amplifying and enriching the uncleaved rare nucleic acids. The basic concept and implementation of this method are disclosed in Korean Publication No. 2015-0138074 A. The other method involves blocking the ends of the nucleic acids in the original sample, cleaving the rare nucleic acids with the CRISPR / Cas system to expose the unblocked ends, linking an adapter to the unblocked ends, and amplifying and enriching the rare nucleic acid fragments linked to the adapters. The basic concept and implementation of this method are disclosed in U.S. Publication No. 2019-0300935 A1.
[0021] What these two methods have in common is that they use a CRISPR / Cas system that targets a rare nucleic acid or a background nucleic acid, that is, a CRISPR / Cas system that includes a guide nucleic acid that perfectly matches the nucleic acid to be cut. However, it is known in the prior art that the CRISPR / Cas system can cleave not only the target nucleic acid but also a non-target nucleic acid that is very similar to the target nucleic acid. This is the so-called off-target cleavage activity problem. Accordingly, when the sequences of the rare nucleic acid and the background nucleic acid are very similar, the CRISPR / Cas system cleaves not only the nucleic acid to be cut but also the nucleic acid that should not be cut, making it difficult to achieve the intended enrichment effect. In addition, the prior art has a limitation in that it can only be used when the sequences of the rare nucleic acid and the background nucleic acid are very similar and when the different bases are located in the protospacer adjacent motif (PAM).
[0022] Therefore, there is a need to develop a method that can be applied even when the rare nucleic acid sequence and the background nucleic acid sequence are very similar.
[0023]
[0024] High-throughput assessment method for in vitro activity of CRISPR / Cas systems
[0025] The inventors of this application have implemented the invention of this specification based on the idea that an in vitro environment can be simulated by immobilizing cells to suppress their biological activity, while partially utilizing existing well-established high-throughput assessment methods for the activity of the in vivo environment. Specifically, the technical problem has been solved by developing a method for 1) transfecting cells with vectors containing various target-guide pairs, 2) immobilizing such cells so that each cell simulates an in vitro environment, 3) independently inducing an activation reaction of the CRISPR / Cas system in each of these cells, and 4) analyzing the effects by simultaneous sequencing after the reaction.
[0026] Activity prediction model for the CRISPR / Cas system
[0027] The inventors of this application solved the technical problem by 1) collecting data on cleavage activity using various combinations of target-guide vectors, 2) deriving two or more alignment patterns to augment the data when the target nucleic acid and the guide nucleic acid are mismatched, and 3) developing a method for training a model using the aligned target nucleic acid sequence and the aligned guide nucleic acid sequence as key input variables.
[0028] Furthermore, we verified whether the model learned in this way could predict the nucleic acid cleavage activity of the CRISPR / Cas system.
[0029] In addition, the above inventors developed and demonstrated 1) a method for predicting target nucleic acid cleavage selectivity and 2) a method for deriving a guide nucleic acid with high selective cleavage activity for a specific target among multiple targets by utilizing the above prediction model.
[0030] Rare nucleic acid sample enrichment method
[0031] As described above, the inventors of this application invented an activity prediction model for the CRISPR / Cas system, and based on this, developed a method for deriving a guide nucleic acid with high selective cleavage activity for a specific target among multiple targets. Using the above method, the inventors selected a guide nucleic acid (optimized guide nucleic acid) that selectively cleaves a target nucleic acid (rare nucleic acid or background nucleic acid) compared to a nucleic acid with a similar sequence (background nucleic acid or rare nucleic acid), and developed a method for enriching rare nucleic acids using the optimized guide nucleic acid.
[0032] The inventors of this application demonstrated that the above rare nucleic acid enrichment method can be used to enrich circulating tumor nucleic acids (CTUs) present in trace amounts in a sample. Furthermore, they discovered optimized guide nucleic acids capable of enriching known CTUs. Furthermore, the inventors demonstrated that the use of these optimized guide nucleic acids enables much more effective enrichment of rare nucleic acids than the use of perfectly corresponding guide nucleic acids in the prior art. Furthermore, they implemented a high-throughput method capable of simultaneously enriching various types of rare nucleic acids.
[0033]
[0034] Using the high-throughput evaluation method for in vitro activity of the CRISPR / Cas system disclosed in this specification, the activity of the CRISPR / Cas system to cleave a target nucleic acid depending on the guide nucleic acid can be evaluated in high-throughput by varying the target-guide combination.
[0035] Using the activity prediction model for the CRISPR / Cas system disclosed in this specification and its application method, target-guide activity can be predicted with high reliability and accuracy, and guide nucleic acids that selectively cleave specific target nucleic acids according to the purpose can be easily derived.
[0036] Using the rare nucleic acid sample enrichment method disclosed in this specification, rare nucleic acids contained in trace amounts in a sample can be effectively enriched and detected.
[0037]
[0038] Figure 1 schematically illustrates exemplary implementations of target-guide vectors included in a target-guide library used to measure sgRNA activity in vitro. Each vector in the library comprises a protospacer adjacent motif (PAM), a target nucleic acid, a unique molecular identifier (UMI), a T7 promoter, a T7 terminator (T7), and an EcoRI restriction site.
[0039] Figure 2 schematically illustrates the Cut-seq1 method, which is an exemplary implementation example of a method for evaluating the activity of the CRISPR / Cas system in high-throughput in Chapter 1.
[0040] Figure 3 shows the correlation between the truncation indices of two replicates. Pearson's correlation coefficient (r) and Spearman's correlation coefficient (R) are shown. The number of sgRNA-target pairs (n) = 109,266.
[0041] Figure 4 shows the effect of Cut-seq1 on cell line strains. Both endonuclease A-positive (EndA+) and endonuclease A-negative (EndA-) strains were tested. The asterisk (*) indicates the expected PCR band size after adapter ligation to the target cleaved by SpCas9 in T7_2k.
[0042] Figure 5 shows the effect of anchoring methods on Cut-seq1. Asterisks (*) indicate the expected PCR band sizes after adapter ligation to the target cleaved by Cas9.
[0043] Figure 6 presents experimental results demonstrating that each cleavage reaction was compartmentalized. Cut-seq1 was performed using two different target sequences, called EMX1 and FANCF, and their corresponding sgRNA sequences, called sgEMX1 and sgFANCF, respectively, which either matched or did not match the target sequence. Each vector had the structure disclosed in Figure 2 . The asterisk (*) indicates the expected PCR band size after ligation of the adapter to the target cleaved by Cas9.
[0044] Figure 7 illustrates the structure of a target-guide vector used to measure target-guide activity in cultured cells, according to an exemplary embodiment of a method for high-throughput evaluation of the activity of a CRISPR / Cas system in Chapter 1. The vector comprises a long terminal repeat (LTR), a U6 promoter (U6), a poly(T) sequence (polyT), a unique molecular identifier (UMI), and a protospacer adjacent motif (PAM).
[0045] Figure 8 shows the correlation between Cas9 activity (indel incorporation rate) for sgRNA and target nucleic acid matching pairs in cultured cells. The correlation between Cas9 activities in the same cell type shows the correlation between two biological replicates. Each replicate was individually transduced with a lentiviral library of sgRNA and target sequence pairs. Pearson (r) and Spearman (R) correlation coefficients are shown. The number of sgRNA and target pairs analyzed (n) was 33,309 for HEK293T, U6_120k, 354 for HeLa, U6_6k, 352 for Hepa 1-6, U6_6k, and 348 for B16-F10, U6_6k. Figure 9 shows the correlation between Cas9 activity (cleavage activity vs. indel incorporation rate) for sgRNA and target nucleic acid matching pairs in cultured cells and in vitro. Pearson (r) and Spearman (R) correlation coefficients are presented. The number of sgRNA and target pairs analyzed (n) is 33,309 for HEK293T, U6_120k in vitro, 33,309 for T7_120k in vitro, 295 for HeLa, U6_6k, 294 for Hepa 1-6, U6_6k, and 290 for B16-F10, U6_6k. Correlation between Cas9 activity (cleavage activity vs. indel introduction rate) for sgRNA and target nucleic acid matching pairs in cultured cells and in vitro is presented. Pearson (r) and Spearman (R) correlation coefficients are presented. The number of sgRNA and target pairs analyzed (n) was 33,309 for HEK293T, U6_120k, 33,309 for in vitro, T7_120k, 295 for HeLa, U6_6k, 294 for Hepa 1-6, U6_6k, and 290 for B16-F10, U6_6k.
[0046] Figure 10 shows heatmaps showing Pearson and Spearman correlation coefficients (left and middle heatmaps) and Cas9 activity ratios (right) representing Cas9 activity in vitro and in cultured cells (HEK293T, HeLa, Hepa 1-6, or B16-F10). Cleavage index and indel frequency were used to measure Cas9 activity in vitro and in cultured cells, respectively.
[0047] Figure 11 shows the sequence preference for the target at each position. Efficiently cleaved target sequences, including the top 20% of target sequences, and inefficiently cleaved target sequences, including the bottom 20% of target sequences, were compared with the matching sgRNA in vitro and in cultured cells (HEK293T, HeLa, Hepa 1-6, and B16-F10). The y-axis represents the log-2 odds ratio of nucleotide frequencies between efficiently cleaved and inefficiently cleaved target sequences. A positive log-2 odds ratio indicates a preference for a nucleotide in an efficiently cleaved target at that position, while a negative value indicates a preference for a nucleotide in an inefficiently cleaved target at that position. The number of target sequences (n) was 33,309 in vitro, 33,309 in HEK293T, 354 in HeLa, 352 in Hepa 1-6, and 348 in B16-F10 cells.
[0048] Figure 12 shows the correlation between the average in vitro cleavage index for 4-nt PAM sequences and the average indel frequency in HEK293T cells. Each dot represents the average value of the cleavage index or indel frequency for all possible PAM sequences. Pearson (r) and Spearman (R) correlation coefficients are shown. The number of 4-nt PAM sequences (n) = 256 (= 4 4 )am.
[0049] Figure 13 shows a heatmap showing the average in vitro target-guided cleavage activity (left) and the average indel induction frequency (right) in HEK293T cells for each 4-nt PAM sequence. The number of sgRNA-target sequence pairs (n) per 4-nt PAM sequence is 27 in vitro and 30 in HEK293T cells.
[0050] Figure 14 shows the correlation between the cleavage selectivity index (SI) measured in vitro using Cut1_30k and the indel SI measured in HEK293T cells using U6_30k containing the same target and sgRNA pairs. Pearson correlation coefficients (r) and Spearman correlation coefficients (R) are shown. The number of sgRNA-target pairs (n) = 13,965.
[0051] Figure 15 shows the effects of SpCas9 variants, SpCas9-HF1 (HF1), SuperFi-Cas9 (SuperFi), and evoCas9 (evo), on the cleavage selectivity index (SI). SI was measured using Cut1_30k. The boxes in FIG. 5B represent the 25th, 50th, and 75th percentiles, and the whiskers represent the 10th and 90th percentiles. The number of sgRNA and target pairs (n) for each SpCas9 variant was 13,391.
[0052] Figure 16 shows a heatmap showing the average cleavage selectivity index (SI) for sgRNAs with perfect match, 1-nt mismatch, 2-nt mismatch, 1-nt DNA bulge, and 1-nt RNA bulge used with SpCas9 variants (HF1, SuperFi, and evo). The number of sgRNA and target pairs (n) for each SpCas9 variant is 13,391.
[0053] Figure 17 shows the ratio of the type of sgRNA (i.e., optimized sgRNA) with the highest cleavage selectivity index (SI) for 91 target mutations for SpCas9, SpCas9-HF1 (HF1), SuperFi-Cas9 (SuperFi), and evoCas9 (evo). The number of optimized sgRNAs (n) = 91.
[0054] Figure 18 shows the impact of SpCas9 variants on the cleavage selectivity index (SI) of the sgRNA with the highest cleavage selectivity index (SI) for each target mutation. SI was measured using Cut1_30k. Boxes represent the 25th, 50th, and 75th percentiles, and whiskers represent the 10th and 90th percentiles. The optimized number of sgRNAs (n) for each SpCas9 variant was 91.
[0055] Figure 19 schematically illustrates the Cut-seq2 method, which is an exemplary implementation of a method for high-throughput evaluation of the activity of the CRISPR / Cas system in Chapter 1. The target-guide vector used in the Cut-seq2 method includes, in addition to the target nucleic acid, the guide nucleic acid encoding sequence (sgRNA), and the scaffold, a constant sequence for identifying reads not cleaved by Cas9, a T7 promoter (T7), and an EcoRV restriction site.
[0056] Figure 20 shows a heatmap showing the mean (left) and median (right) cleavage ratio index differences for perfectly matched sgRNAs and sgRNAs with 1-nt mismatches, 2-nt mismatches, 1-nt DNA bulges, and 1-nt RNA bulges for SpCas9 variants (SpCas9-HF1 (HF1), SpCas9-NRRH-HF1 (NRRH-HF1), and SpCas9-NRCH-HF1 (NRCH-HF1)). The number of sgRNA and target pairs (n) for each Cas9 variant is 18,155.
[0057] Figure 21 shows the ratio of the sgRNA type (i.e., optimized sgRNA) with the highest cleavage ratio index difference for target nucleic acids relative to noise nucleic acids of 77, 68, and 68 for SpCas9-HF1, SpCas9-NRRH-HF1, and SpCas9-NRCH-HF1, respectively. The number of optimized sgRNAs (n) is as follows: HF1, n = 77; NRRH-HF1, n = 68; NRCH-HF1, n = 68.
[0058] Figure 22 shows the cleavage ratio indices of optimal and perfect matched sgRNAs for WT (wild-type) and MT (mutant) targets determined using SpCas9 variants (HF1, NRRH-HF1, or NRCH-HF1). The numbers (n) of optimal and perfect matched sgRNAs are 77 and 71 for HF1, 69 and 66 for NRRH-HF1, and 68 and 60 for NRCH-HF1, respectively.
[0059] Figure 23 shows a heatmap showing the difference in the average cleavage rate index induced by SpCas9 variants (HF1, NRRH-HF1, or NRCH-HF1) for each 4-nt candidate PAM sequence. If the SpCas9 variant has the highest average cleavage rate index difference among the three variants evaluated for a particular 4-nt PAM sequence, that PAM is indicated by a solid border and a 'v' in the box.
[0060] Figure 24 schematically illustrates the structures of vectors used in Cut-seq1 and Cut-seq2, which are exemplary implementations of a method for high-throughput evaluation of the activity of the CRISPR / Cas system in Chapter 1. EcoRV, EcoRV restriction enzyme site; Constant, constant sequence for identifying reads not cleaved by Cas9; UMI, unique molecular identifier; T7, T7 promoter; T7-term, T7 terminator; EcoR1, EcoR1 restriction enzyme site.
[0061] Figure 25 schematically illustrates the Cut-seq2 method, which is an exemplary implementation example of a method for evaluating the activity of the CRISPR / Cas system in high-throughput in Chapter 1.
[0062] Figure 26 shows the correlation between the cleavage ratio indices of the two replicates. Pearson's correlation coefficient (r) and Spearman's correlation coefficient (R) are shown. The number of sgRNA-target pairs (n) is 25,644 for HF1, 27,260 for NRRH-HF1, and 25,553 for NRCH-HF1.
[0063] Figure 27 shows a heatmap showing the impact of the number of mismatched nucleotides, DNA bulges, or RNA bulges in the sgRNA relative to the target on the average cleavage ratio index of Cas9 variants (SpCas9-HF1 (HF1), SpCas9-NRRH-HF1 (NRRH-HF1), and SpCas9-NRCH-HF1 (NRCH-HF1)) measured using Cut2_30k and Cut2_42k.
[0064] Figure 28 shows a schematic diagram of the deep learning algorithm used to develop DeepCut, an implementation example of a CRISPR / Cas system cleavage activity prediction model in Chapter 2. M, mismatch; D, DNA bulge; R, RNA bulge; Tm, melting temperature; MFE, minimum free energy; ΔGH, change in sgRNA-DNA hybridization free energy.
[0065] Figure 29 shows the results of using the truncation rate index dataset, which was not used for training, to evaluate DeepCut-HF1. A scatterplot of the measured and predicted truncation rate indices is shown. The Pearson correlation coefficient (r) and Spearman correlation coefficient (R) are also shown. The number of sgRNA and target pairs (n) in the test set is 5,831 for HF1.
[0066] Figure 30 shows the results of using the cleavage rate index dataset, which was not used for training, to evaluate DeepCut-NRRH-HF1. A scatterplot of the measured and predicted cleavage rate indices is shown. The Pearson correlation coefficient (r) and Spearman correlation coefficient (R) are also shown. The number of sgRNA and target pairs (n) in the test set is for NRRH-HF1.
[0067] Figure 31 shows the cleavage rate index dataset, which was not used for training, used for evaluating DeepCut-NRCH-HF1. A scatterplot of the measured and predicted cleavage rate indices is shown. The Pearson correlation coefficient (r) and Spearman correlation coefficient (R) are also shown. The number of sgRNA and target pairs (n) in the test set is 5,570 for NRCH-HF1.
[0068] Figure 32 shows the evaluation results of the HF1 (d), NRRH-HF1 (e), and NRCH-HF1 (f) models. The models used to evaluate the above cleavage rate index dataset were trained without additional features. The dataset used for evaluation was not used for training. A scatterplot of the measured and predicted cleavage rate indices is shown. Pearson's correlation coefficient (r) and Spearman's correlation coefficient (R) are also shown. The number of sgRNA and target pairs (n) in the test set is 5,831 for HF1, 4,402 for NRRH-HF1, and 5,570 for NRCH-HF1.
[0069] Figure 33 shows the models predicting cleavage activity for SpCas9-HF1 using various algorithms, and the Pearson (top row) or Spearman (bottom row) correlation coefficients between the cleavage rate index measured through 5-fold cross-validation and the predicted cleavage rate index for each model. The left row shows the results targeting wild-type (WT) nucleic acids, and the right row shows the results targeting mutant (MT) nucleic acids. The number of correlation coefficients (n) = 5 (Fold 0 to Fold 4). Statistical comparisons between the top two algorithms were evaluated using a two-tailed Steiger test. The bars and error bars represent the mean and standard deviation of the correlation coefficients, respectively. CNN, convolutional neural network; DNN, deep neural network; SVM, support vector machine; GRU, gated recurrent unit; XGBoost, extreme gradient boosting; LightGBM, lightweight gradient boosting machine; CatBoost, categorical boosting; RNN, recurrent neural network; LSTM, long short-term memory; RF, random forest; Ridge, ridge regression; Linear, linear regression; GB, gradient boosting; Huber, Huber regression; LightSVM, light support vector machine.
[0070] Figure 34 shows the models predicting cleavage activity for SpCas9-NRRH-HF1 using various algorithms, and the Pearson (top row) or Spearman (bottom row) correlation coefficients between the cleavage rate index measured through 5-fold cross-validation and the predicted cleavage rate index for each model. The left row shows the results targeting wild-type (WT) nucleic acids, and the right row shows the results targeting mutant (MT) nucleic acids. The number of correlation coefficients (n) = 5 (Fold 0 to Fold 4). Statistical comparisons between the top two algorithms were evaluated using a two-tailed Steiger test. The bars and error bars represent the mean and standard deviation of the correlation coefficients, respectively. CNN, convolutional neural network; DNN, deep neural network; SVM, support vector machine; GRU, gated recurrent unit; XGBoost, extreme gradient boosting; LightGBM, lightweight gradient boosting machine; CatBoost, categorical boosting; RNN, recurrent neural network; LSTM, long short-term memory; RF, random forest; Ridge, ridge regression; Linear, linear regression; GB, gradient boosting; Huber, Huber regression; LightSVM, light support vector machine.
[0071] Figure 35 shows the models predicting cleavage activity for SpCas9-NRCH-HF1 using various algorithms, and the Pearson (top row) or Spearman (bottom row) correlation coefficients between the cleavage rate index measured through 5-fold cross-validation and the predicted cleavage rate index for each model. The left row shows the results targeting wild-type (WT) nucleic acids, and the right row shows the results targeting mutant (MT) nucleic acids. The number of correlation coefficients (n) = 5 (Fold 0 to Fold 4). Statistical comparisons between the top two algorithms were evaluated using a two-tailed Steiger test. The bars and error bars represent the mean and standard deviation of the correlation coefficients, respectively. CNN, convolutional neural network; DNN, deep neural network; SVM, support vector machine; GRU, gated recurrent unit; XGBoost, extreme gradient boosting; LightGBM, lightweight gradient boosting machine; CatBoost, categorical boosting; RNN, recurrent neural network; LSTM, long short-term memory; RF, random forest; Ridge, ridge regression; Linear, linear regression; GB, gradient boosting; Huber, Huber regression; LightSVM, light support vector machine.
[0072] Figure 36 shows the key features associated with the cleavage rate index differences of the models for SpCas9-HF1. The 15 most important features associated with the cleavage rate index differences were determined by Tree SHAP using a correlation coefficient threshold of 0.7 after accounting for multicollinearity. A high SHAP value indicates that the feature is associated with a high cleavage rate index difference. The darker dots indicate high and low values of related features, such as melting temperature (Tm) and whether the sgRNA spacer starts with G (GN19). The number of sgRNA-target sequence pairs (n) = 62,397.
[0073] Figure 37 shows the key features associated with the cleavage rate index differences of the models for SpCas9-NRRH-HF1. The 15 most important features associated with the cleavage rate index differences were determined by Tree SHAP using a correlation coefficient threshold of 0.7 after accounting for multicollinearity. A high SHAP value indicates that the feature is associated with a high cleavage rate index difference. The darker dots indicate high and low values of related features, such as melting temperature (Tm) and whether the sgRNA spacer starts with G (GN19). The number of sgRNA-target sequence pairs (n) = 64,641.
[0074] Figure 38 shows the key features associated with the cleavage rate index differences of the models for SpCas9-NRCH-HF1. The 15 most important features associated with the cleavage rate index differences were determined by Tree SHAP using a correlation coefficient threshold of 0.7 after accounting for multicollinearity. A high SHAP value indicates that the feature is associated with a high cleavage rate index difference. The darker dots indicate high and low values of related features, such as melting temperature (Tm) and whether the sgRNA spacer starts with G (GN19). The number of sgRNA-target sequence pairs (n) = 63,599.
[0075] Figure 39 schematically illustrates a method for enriching rare nucleic acids by cleaving background nucleic acids with a CRISPR / Cas system.
[0076] Figure 40 shows the number of cancer-associated mutations in each of the 40 most frequent genes in a gene library containing 2,612 mutations. These mutations were included in a panel of 2,612 target mutations (i.e., corresponding to rare nucleotides).
[0077] Figure 41 shows a schematic diagram of cfDNA mimics. In the mutant cfDNA mimic (cf_MT), dark squares indicate the locations of cancer-associated mutations in the mutant target sequence. In the wild-type cfDNA mimic (cf_WT), black squares indicate the locations of cancer-associated mutations.
[0078] Figure 42 shows the structure of a cfDNA mimic surrounded by an adapter. The cfDNA mimic also contains a barcode and a unique molecular identifier (UMI).
[0079] Figure 43 shows the ratio of optimized sgRNA types that showed the highest cleavage ratio index difference among sgRNAs with a cleavage ratio index less than 0 at the mutant target for 2,612 target mutations for SpCas9-HF1 (HF1) and SpCas9-NRRH-HF1 (NRRH-HF1). The number of optimized sgRNAs (n) = 2,612 for HF1 and 2,612 for NRRH-HF1.
[0080] Figure 44 shows the ratio of mutation positions (positions within the protospacer (i.e., selectivity by the sgRNA) versus positions within the PAM (i.e., selectivity by the PAM sequence)) of the sgRNAs that showed the highest cleavage ratio index difference for 2,612 target mutations for HF1 (top) and NRRH-HF1 (bottom).
[0081] Figure 45 shows the effect of the SpCas9-HF1:sgRNA molar ratio on the enrichment fold (VAF fold increase) during selective enrichment of rare nucleic acid (mutant nucleic acid, MT) sequences. A single cleavage reaction cycle was performed using the optimal_HF1_2k sgRNA library and a 1:1,000 mixture of cf_WT and cf_MT. The number of target mutations (n) for each condition was 2,612.
[0082] Figure 46 shows the effect of SpCas9-HF1 concentration on the enrichment fold (VAF fold increase) during selective enrichment of rare nucleic acid (mutant nucleic acid, MT) sequences. A single cleavage reaction cycle was performed using the optimal_HF1_2k sgRNA library and a 1:1,000 mixture of cf_WT and cf_MT. The number of target mutations (n) for each condition was 2,612.
[0083] Figure 47 shows the effect of SpCas9-HF1 cleavage reaction time on the enrichment fold (VAF fold increase) during selective enrichment of rare nucleic acid (mutant nucleic acid, MT) sequences. A single cleavage reaction cycle was performed using the optimal_HF1_2k sgRNA library and a 1:1,000 mixture of cf_WT and cf_MT. The number of target mutations (n) for each condition was 2,612.
[0084] Figure 48 shows the fold change (left, A) and VAF (middle, B) of the variant allele frequency (VAF) before and after cleavage with either optimal_HF1_2k or perfect match_2k using HF1-guided cleavage, depending on the number of PCR cycles. cf_WT and cf_MT were mixed at a ratio of 1,000:1 and used as cfDNA mimics. The dots represent the mean fold change (A) and mean VAF (B). The boxes represent the 25th, 50th, and 75th percentiles, and the whiskers represent the 10th and 90th percentiles. 'V' indicates that the median is 0, which cannot be displayed on the graph. The number of target mutations (n) = 2,612. C shows the proportion of target mutations with a VAF exceeding the limit of reliable detection (VAF > 0.2%) after cleavage and PCR cycles using HF1 and optimal_HF1_2k or perfect_match_2k. The number of target mutations (n) = 2,612.
[0085] Figure 49 shows the fold increase in variant allele frequency (VAF) for each of the 40 most frequently mutated genes among all genes containing 2,612 mutations. Each dot represents a cancer-associated mutation. The VAF fold increase was measured after five rounds of cleavage and PCR using a 1,000:1 mixture of cf_WT and cf_MT with an optimized sgRNA for HF1. The number of target mutations (n) was 346 for TP53, 78 for PIK3CA, 67 for FAT4, 58 for KMT2C, 55 for APC, 49 for KMT2D, 44 for RNF213, 43 for ARID1A, 43 for PTEN, 36 for KRAS, 36 for NFE2L2, 35 for FAT1, 33 for ERBB4, 29 for CDKN2A, 28 for FBXW7, 28 for CTNNB1, 27 for NF1, 26 for BCL7A, 24 for ATM, 24 for PTPRT, 23 for ZFHX3, 22 for SPEN, 20 for NTRK3, 20 for SF3B1, and 20 for MET. 19 for EGFR, 18 for NRAS, 17 for TRRAP, 17 for RUNX1, 17 for TSPOAP1-AS1, 17 for KMT2A, 17 for MTOR, 16 for PDE4DIP, 16 for GNAS, 15 for BRAF, 13 for KIT, 13 for HRAS, 13 for ERBB2, 12 for DNMT3A, and 12 for SMARCA4.
[0086] Figure 50 shows the results of performing one cycle of cleavage and PCR using SpCas9-HF1 and an optimized sgRNA (left) or a perfectly matched sgRNA (right). The graphs are plotted by VAF fold change range. The percentage above each bar represents the proportion of sgRNAs within the corresponding VAF fold change range. The number of target nucleic acids (n) is 2,612 for the optimized sgRNA and 2,612 for the perfectly matched sgRNA.
[0087] Figure 51 shows the VAF fold increase according to the difference in the cleavage ratio index predicted by DeepCut-HF1. The VAF fold increase was measured after 1, 3, or 5 iterations of cleavage and PCR using HF1 and Optimal_HF1_2k. The black dot represents the mean. The boxes represent the 25th, 50th, and 75th percentiles, and the whiskers represent the 10th and 90th percentiles. The number of target mutations (n) = 2,612.
[0088] Figure 52 shows the fold change (left, A) and VAF (middle, B) of the variant allele frequency (VAF) before and after cleavage with either optimal_NRRH_2k or perfect match_2k using NRRH-HF1-guided cleavage, depending on the number of PCR cycles. cf_WT and cf_MT were mixed at a ratio of 1,000:1 and used as cfDNA mimics. The dots represent the mean fold change (A) and mean VAF (B). The boxes represent the 25th, 50th, and 75th percentiles, and the whiskers represent the 10th and 90th percentiles. 'V' indicates that the median is 0, which cannot be displayed on the graph. The number of target mutations (n) = 2,612. C shows the proportion of target mutations with a VAF exceeding the limit of reliable detection (VAF > 0.2%) after cleavage and PCR cycles using NRRH-HF1 and optimal_NRRH_2k or perfect_match_2k. The number of target mutations (n) = 2,612.
[0089] Figure 53 shows the fold increase in variant allele frequency (VAF) for each of the 40 most frequently mutated genes among all genes containing 2,612 mutations. Each dot represents a cancer-associated mutation. The VAF fold increase was measured after five rounds of cleavage and PCR with the optimal sgRNA for NRRH-HF1 using a 1,000:1 mixture of cf_WT and cf_MT. The number of target mutations (n) was 346 for TP53, 78 for PIK3CA, 67 for FAT4, 58 for KMT2C, 55 for APC, 49 for KMT2D, 44 for RNF213, 43 for ARID1A, 43 for PTEN, 36 for KRAS, 36 for NFE2L2, 35 for FAT1, 33 for ERBB4, 29 for CDKN2A, 28 for FBXW7, 28 for CTNNB1, 27 for NF1, 26 for BCL7A, 24 for ATM, 24 for PTPRT, 23 for ZFHX3, 22 for SPEN, 20 for NTRK3, 20 for SF3B1, and 20 for MET. 19 for EGFR, 18 for NRAS, 17 for TRRAP, 17 for RUNX1, 17 for TSPOAP1-AS1, 17 for KMT2A, 17 for MTOR, 16 for PDE4DIP, 16 for GNAS, 15 for BRAF, 13 for KIT, 13 for HRAS, 13 for ERBB2, 12 for DNMT3A, and 12 for SMARCA4.
[0090] Figure 54 shows the results of performing one cycle of cleavage and PCR using SpCas9-NRRH-HF1 and an optimized sgRNA (left) or a perfectly matched sgRNA (right). The graph is plotted by VAF fold change range. The percentage above each bar represents the proportion of sgRNAs within the corresponding VAF fold change range. The number of target nucleic acids (n) is 2,612 for the optimized sgRNA and 2,612 for the perfectly matched sgRNA.
[0091] Figure 55 shows the VAF fold increase according to the difference in the cleavage ratio index predicted by DeepCut-NRRH-HF1. The VAF fold increase was measured after 1, 3, or 5 repetitions of cleavage and PCR using NRRH-HF1 and Optimal_NRRH_2k. The black dots represent the mean. The boxes represent the 25th, 50th, and 75th percentiles, and the whiskers represent the 10th and 90th percentiles. The number of target mutations (n) = 2,612.
[0092] Figure 56 shows the results of enrichment through PCR after cleavage by treating SpCas9-HF and optimized sgRNA for each of the cases where the sequence of the rare nucleic acid (mutant nucleic acid) is 1) mutation near the NGG PAM, 2) mutation near the NAG or NGA PAM, 3) mutation near the non-NGG, NAG, or NGA PAM, and 4) mutation within the PAM when cleaving the background nucleic acid to enrich the rare nucleic acid. The graph on the left shows the difference in the cleavage ratio index predicted by the cleavage activity prediction model. The graph on the right shows the experimentally measured VAF fold change according to the PCR cycle.
[0093] Figure 57 shows the results of enrichment through PCR after cleavage of rare nucleic acids (mutant nucleic acids) compared to the background nucleic acid (wild-type nucleic acid) sequence when cleaving the background nucleic acid to enrich rare nucleic acids, for each of the following cases: 1) mutations near the NRRH PAM, 2) mutations near the non-NRRH PAM, and 3) mutations within the PAM. The graph on the left shows the difference in cleavage ratio index predicted by the cleavage activity prediction model. The graph on the right shows the experimentally measured VAF fold change according to PCR cycles.
[0094] Figure 58 shows the effect of deep sequencing read depth (i.e., approximately 2,000x versus approximately 10,000x) on enrichment of rare nucleic acids (i.e., mutant nucleic acids). Each box represents the 25th, 50th, and 75th percentiles, and the whiskers represent the 10th and 90th percentiles. The number of target mutations (n) = 2,612.
[0095] Figure 59 schematically illustrates an implementation example of a method of enriching rare nucleic acids by cutting them with a CRISPR / Cas system, among the rare nucleic acid enrichment methods of Chapter 3.
[0096] Figure 60 shows the fold change in variant allele frequency (VAF) using the indicated sgRNA libraries and PCR after cleavage of rare nucleic acids using SpCas9-HF1. 600_WT and 600_MT were mixed at a 1,000:1 ratio. The black dots represent the average fold change in VAF. The graph is presented in order of decreasing VAF fold change. Subsets of sgRNA libraries that did not show statistically significant differences in VAF fold change (P > 0.05, determined by analysis of variance using the Mann-Whitney U test followed by a Bonferroni correction post hoc test) are indicated by letters a, b, c, etc. Boxes represent the 25th, 50th, and 75th percentiles, and whiskers represent the 10th and 90th percentiles. The number of target mutations (n) = 200.
[0097] Figure 61 shows the variant allele frequency (VAF) after HF1-induced cleavage using PCR with uncut (left), optimized (middle), or perfectly matched (right) sgRNA.
[0098] Figure 62 shows a schematic of a strategy to remove non-cancer-related targets (i.e., background nucleic acids) using dCas9 and sgRNA binding to cf_MT.
[0099] Figure 63 shows the ratio of mutant nucleic acids to rare nucleic acids (i.e., mutant nucleic acids) and adjacent regions, depending on the number of cycles of treatment with dCas9 and sgRNA. The control group was set at a ratio of 1,000:1 between cf_non-cancer and cf_MT, which did not receive dCas9 or sgRNA treatment.
[0100] Figure 64 shows the variant allele frequency (VAF) of the MT sequence before and after dCas9-mediated capture and PCR cycles using a 1,000,000:1,000:1 mixture of cf_non-cancer:cf_WT:cf_MT. The number of target mutations (n) = 2,612.
[0101] Figure 65 shows the ratio of cf_MT, cf_WT, and cf_non-cancer before and after dCas9-mediated capture and PCR cycles using a 1,000,000:1,000:1 mixture of cf_non-cancer:cf_WT:cf_MT.
[0102] Figure 66 shows a flowchart for a high-throughput evaluation method for in vitro activity of the CRISPR / Cas system.
[0103] Figure 67 is a flowchart showing an example of a detailed process for manufacturing microchamber cells (S1100) among high-throughput evaluation methods for in vitro activity of the CRISPR / Cas system.
[0104] Figure 68 is a flowchart showing an example of a detailed process for manufacturing microchamber cells (S1100) among high-throughput evaluation methods for in vitro activity of the CRISPR / Cas system.
[0105] Figure 69 is a flowchart illustrating an example of a detailed process for inducing a compartmentalized reaction (S1200) in a high-throughput evaluation method for the in vitro activity of a CRISPR / Cas system. (Top) A flowchart illustrating an example of a detailed process for inducing a compartmentalized reaction (S1200). (Bottom) A flowchart illustrating an example of a detailed process for treating one or more enzymes (S1220).
[0106] Figure 70 is a flowchart showing an example of a detailed process for evaluating the activity of a CRISPR / Cas system (S1300) among high-throughput evaluation methods for in vitro activity of a CRISPR / Cas system.
[0107] Figure 71 is a flowchart showing an example of a detailed process for obtaining sequence information from nucleic acids collected during a detailed process for evaluating the activity of a CRISPR / Cas system (S1300) among high-throughput evaluation methods for in vitro activity of a CRISPR / Cas system (S1320).
[0108] Figure 72 is a flowchart showing an example of a detailed process for evaluating target-guide activity (S1330) among detailed processes for evaluating the activity of a CRISPR / Cas system (S1300) among high-throughput evaluation methods for in vitro activity of a CRISPR / Cas system.
[0109] Figure 73 is a flowchart illustrating an example of a method for learning a cleavage activity prediction model of a CRISPR / Cas system.
[0110] Figure 74 is a flowchart showing an example of a detailed process for deriving training data from raw data (S2200) among the methods for learning a cleavage activity prediction model of a CRISPR / Cas system.
[0111] Figure 75 is a flowchart illustrating an example of a method for predicting target nucleic acid cleavage activity using an activity prediction model.
[0112] Figure 76 is a flowchart illustrating an example of a process for predicting cleavage activity by inputting a selected target-guide alignment pattern into an activity prediction model, among methods for predicting target nucleic acid cleavage activity using an activity prediction model. In particular, the process is illustrated when only one target-guide pattern is selected in the previous process (S2500).
[0113] Figure 77 is a flowchart illustrating an example of a process for predicting cleavage activity by inputting a selected target-guide alignment pattern into an activity prediction model, among methods for predicting target nucleic acid cleavage activity using an activity prediction model. In particular, the process is illustrated when two or more target-guide patterns are selected in the previous process (S2500).
[0114] Figure 78 is a flowchart showing an example of a method for predicting target nucleic acid cleavage selectivity.
[0115] Figure 79 is a flowchart illustrating an example of a guide nucleic acid selection method for selectively cleaving only one of two target nucleic acids.
[0116]
[0117] Hereinafter, the best mode for carrying out the invention is exemplified. This includes some, but not all, implementations of the invention disclosed herein. The embodiments described in this paragraph are merely exemplary, and the implementations described in this paragraph should not be construed as the "best mode for carrying out the invention." Those skilled in the art will likely envision numerous variations and more desirable implementations of the examples described in this paragraph, and such variations should also be considered to be included within the best mode for carrying out the invention.
[0118] This specification describes a method for high-throughput evaluation of the activity of the CRISPR / Cas system for various combinations of guide RNA and target DNA.
[0119] A method is disclosed comprising:
[0120] (a) A process for preparing microchamber cells, comprising:
[0121] (a-1) Process of preparing a cell population;
[0122] (a-2) Process of treating the target-guide library to the above cell population,
[0123] Here, the target-guide library includes a plurality of target-guide vectors,
[0124] Each target-guide vector comprises a target DNA, a DNA encoding a guide RNA, and a promoter operably linked to the DNA encoding the guide RNA,
[0125] The above target-guide library contains at least two types of target-guide vectors having different target DNA sequences, different guide RNA sequences, or both of the above sequences,
[0126] The above target-guided library is processed according to a predetermined multiplicity of infection (MOI); and
[0127] (a-3) A process of treating a fixed reagent to a cell population treated with the target-guide library;
[0128] Here, by process (a-2), most of the cells included in the cell population prepared in (a-1) are i) cells not transfected with the target-guide vector, or ii) cells transfected with only one target-guide vector.
[0129] By the above process (a-3), only the one target-guide vector is fixed to the transfected cell and functions as an independent reaction chamber.
[0130] Here, fixed cells transfected with only one target-guide vector are called microchamber cells;
[0131] (b) a process for inducing a compartmentalized response by microchamber cells in a cell population, comprising:
[0132] (b-1) A process of treating a cell population including the above microchamber cells with a transcription enzyme,
[0133] Here, the transcription enzyme is delivered to each of the microchamber cells,
[0134] The DNA encoding the guide RNA contained in each microchamber cell is transcribed into guide RNA by the above transcription enzyme; and
[0135] (b-2) A process of treating a Cas protein in a cell population treated with the above transcription enzyme,
[0136] Here, the Cas protein is delivered to each of the microchamber cells;
[0137] Here, through the above process (b), within each microchamber cell, the Cas protein and the guide RNA transcribed in each microchamber cell combine to form a CRISPR / Cas complex, and the CRISPR / Cas complex reacts with the target DNA within the microchamber cell.
[0138] Reactions within each microchamber cell occur independently; and
[0139] (c) a process for analyzing the reaction activity between the guide RNA and the target DNA included in each target-guide vector, including:
[0140] (c-1) A process of acquiring information about a target-guide vector contained in microchamber cells in a cell population in which a compartmentalized response is induced by microchamber cells;
[0141] Here, information about the target-guide vector includes the sequence of the guide RNA contained therein, the sequence of the target DNA, and whether the target DNA is cleaved; and
[0142] (c-2) A process for determining the reaction activity between the guide RNA and target DNA included in each target-guide vector based on the information regarding (c-1) above.
[0143] In one embodiment, in the above method,
[0144] The above process (a) further includes the following process (a-2-1) after the above process (a-2):
[0145] (a-2-1) A process for expanding a cell population treated with the above target-guided library;
[0146] The above process (a-3) may be a process of treating a fixed reagent to a cell population expanded through the above process (a-2-1).
[0147] This specification discloses a population of fixed cells comprising microchamber cells:
[0148] Here, each microchamber cell contains a Cas protein, a guide nucleic acid, and a target nucleic acid,
[0149] Each Cas protein, guide nucleic acid, and target nucleic acid is an exogenous construct,
[0150] Each microchamber cell contains only one type of guide nucleic acid and only one type of target nucleic acid,
[0151] The population of fixed cells comprises at least two different types of microchamber cells, wherein the different microchamber cells are cells that differ in the sequence of the guide nucleic acid contained therein, the sequence of the target nucleic acid contained therein, or both sequences,
[0152] Here, the endogenous enzyme activity of each microchamber cell is inhibited, while the exogenous components remain active, allowing the microchamber cells to mimic the reaction chambers of an in vitro environment.
[0153] Here, the population of fixed cells ensures compartmentalized interactions between Cas proteins, guide nucleic acids, and target nucleic acids within each microchamber cell.
[0154] As an example,
[0155] Each of the above microchamber cells contains guide RNA as a guide nucleic acid and target DNA as a target nucleic acid,
[0156] Each of the above microchamber cells comprises a vector comprising a nucleic acid encoding the target DNA and the guide RNA, or a fragment thereof,
[0157] The above vector has the structure of [Structural Formula 1] or [Structural Formula 2]:
[0158] [Structural formula 1]
[0159] [Target] - [Barcode] - [Promoter] - [Guide] - [RS1]; or
[0160] [Structural formula 2]
[0161] [RS2] - [Constant] - [Target] - [Barcode] - [Promoter] - [Guide] - [RS1]
[0162] Here, the Target is the target DNA,
[0163] The above Barcode is a barcode that 1) encrypts target-guide sequence information, 2) encrypts unique identification information of the target-guide vector, or 3) encrypts both 1) and 2).
[0164] The above Guide is a DNA encoding the above guide RNA,
[0165] The above promoter is operably linked to DNA encoding the guide RNA,
[0166] The above RS1 is the first restriction enzyme site,
[0167] The above Constant is an invariant sequence,
[0168] The above RS2 initiates a population of fixed cells containing microchamber cells, which are the second restriction enzyme site.
[0169] This specification discloses a method for learning an activity prediction model of a CRISPR / Cas system, comprising:
[0170] (a) A process for preparing learning data, including:
[0171] (a-1) Process of obtaining raw data,
[0172] Here, the raw data includes a plurality of raw data points,
[0173] Each of the above raw data points includes 1) a target nucleic acid sequence, 2) a guide nucleic acid sequence, and 3) cleavage activity.
[0174] The above cleavage activity is a value measuring the degree to which a CRISPR / Cas complex including the guide nucleic acid sequence cleaves a target nucleic acid sequence;
[0175] (a-2) A process of generating sorted data by augmenting the above raw data,
[0176] Here, the alignment data includes a plurality of alignment data points,
[0177] Applying one or more sorting algorithms to each of the above raw data points to generate one or more sorted data points per raw data point,
[0178] Each of the above alignment data points comprises 1) an aligned target nucleic acid sequence, 2) an aligned guide nucleic acid sequence, and 3) the cleavage activity,
[0179] The aligned target nucleic acid sequence and aligned guide nucleic acid sequence are the results of aligning the target nucleic acid sequence and guide nucleic acid sequence included in the raw data points according to an alignment algorithm, and include gap information; and
[0180] (a-3) A process of generating learning data by processing the above sorted data,
[0181] Here, the learning data includes a plurality of learning data points,
[0182] For each training data point, for each sorted data point,
[0183] Using the above aligned target nucleic acid sequence and the above aligned guide nucleic acid sequence as input values,
[0184] Label the above cutting activity as an output value,
[0185] Converting the aligned target nucleic acid sequence and the aligned guide nucleic acid sequence into a form suitable for learning a prediction model; and
[0186] (b) A process for training the above prediction model, performed using the above training data;
[0187] Here, the prediction model is configured to input an aligned target nucleic acid sequence and an aligned guide nucleic acid sequence and output a predicted cleavage activity.
[0188] This specification discloses a method for learning a CRISPR / Cas system activity prediction model, comprising:
[0189] The above method includes:
[0190] (a) A process for preparing learning data, including:
[0191] (a-1) Process of obtaining raw data,
[0192] Here, the raw data includes a plurality of raw data points,
[0193] Each of the above raw data points includes 1) a target nucleic acid sequence, 2) a guide nucleic acid sequence, and 3) cleavage activity.
[0194] The above cleavage activity is a value measuring the degree to which a CRISPR / Cas complex including the guide nucleic acid sequence cleaves a target nucleic acid sequence;
[0195] (a-2) A process of generating sorted data by augmenting the above raw data,
[0196] Here, the alignment data includes a plurality of alignment data points,
[0197] Applying one or more sorting algorithms to each of the above raw data points to generate one or more sorted data points per raw data point,
[0198] Each of the above alignment data points comprises 1) an aligned target nucleic acid sequence, 2) an aligned guide nucleic acid sequence, 3) mismatch information, 4) optionally, additional input variables, and 5) the truncation activity of the raw data point,
[0199] The aligned target nucleic acid sequence and aligned guide nucleic acid sequence are the result of aligning the target nucleic acid sequence and guide nucleic acid sequence included in the raw data points according to an alignment algorithm, and include gap information.
[0200] The above mismatch information includes information about each base position of the aligned target nucleic acid sequence and the aligned guide nucleic acid sequence, expressed as one of the following: bases of the two nucleic acids match; bases of the two nucleic acids do not match; a gap occurs in the target nucleic acid; or a gap occurs in the guide nucleic acid.
[0201] The above additional input variables include one or more of the following information determined from the aligned target nucleic acid sequence and the aligned guide nucleic acid sequence: Melting Point (T) of the target nucleic acid-guide nucleic acid binding m ); the minimum free energy (MFE) of target nucleic acid-guide nucleic acid binding; and the change in free energy of target nucleic acid-guide nucleic acid binding; and
[0202] (a-3) A process of generating learning data by processing the above sorted data,
[0203] Here, the learning data includes a plurality of learning data points,
[0204] For each training data point, for each sorted data point,
[0205] Using the aligned target nucleic acid sequence, the aligned guide nucleic acid sequence, the mismatch information, and optionally the additional input variable as input values,
[0206] Label the above cutting activity as an output value,
[0207] Converting the aligned target nucleic acid sequence, the aligned guide nucleic acid sequence, the mismatch information, and the additional input variables into a form suitable for learning a prediction model; and
[0208] (b) A process for training the above prediction model, performed using the above training data;
[0209] Here, the prediction model is configured to receive an aligned target nucleic acid sequence, an aligned guide nucleic acid sequence, the mismatch information, and optionally the additional variable as input and output a predicted cleavage activity.
[0210] This specification discloses a method for predicting the cleavage activity of a CRISPR / Cas system, comprising:
[0211] (a) a process for obtaining a target nucleic acid sequence and a guide nucleic acid sequence;
[0212] (b) a process of deriving one or more input variables from the target nucleic acid sequence and the guide nucleic acid sequence:
[0213] Here, each of the above input variables includes:
[0214] (i) an aligned target nucleic acid sequence; and
[0215] (ii) aligned guide nucleic acid sequence;
[0216] Here, different input variables include different aligned target nucleic acid sequences and aligned guide nucleic acid sequence information;
[0217] (c) a process of predicting the cleavage activity of the CRISPR / Cas complex including the guide nucleic acid for the target nucleic acid by utilizing the activity prediction model;
[0218] Here, the activity prediction model is a model trained to predict the cleavage activity of the CRISPR / Cas system by inputting an aligned target nucleic acid sequence and an aligned guide nucleic acid sequence.
[0219] By inputting each of the above one or more input variables into the above active prediction model, one predicted cut-off activity value is calculated per input variable,
[0220] If there is one predicted cleavage activity value, the predicted cleavage activity value is used as the prediction result,
[0221] If there are two or more predicted cutting activity values, the predicted cutting activity values are weighted averaged with a predetermined weight and used as the prediction result.
[0222] This specification discloses a method for predicting the cleavage activity of a CRISPR / Cas system, comprising:
[0223] (a) a process for obtaining a target nucleic acid sequence and a guide nucleic acid sequence;
[0224] (b) a process of deriving one or more input variables from the target nucleic acid sequence and the guide nucleic acid sequence:
[0225] Here, each of the above input variables includes:
[0226] (i) aligned target nucleic acid sequence;
[0227] (ii) aligned guide nucleic acid sequence;
[0228] (iii) inconsistent information;
[0229] The mismatch information includes information about each base position of the aligned target nucleic acid sequence and the aligned guide nucleic acid sequence, expressed as one of the following: bases of the two nucleic acids match; bases of the two nucleic acids do not match; a gap occurs in the target nucleic acid; or a gap occurs in the guide nucleic acid; and
[0230] (iv) additional input variables,
[0231] The above additional input variables include one or more of the following information, determined from the aligned target nucleic acid sequence and the aligned guide nucleic acid sequence: Melting Point (T) of the target nucleic acid-guide nucleic acid binding m ); minimum free energy (MFE) of target nucleic acid-guide nucleic acid binding; and free energy change of target nucleic acid-guide nucleic acid binding;
[0232] Here, different input variables include different aligned target nucleic acid sequences and aligned guide nucleic acid sequence information;
[0233] (c) a process of predicting the cleavage activity of the CRISPR / Cas complex including the guide nucleic acid for the target nucleic acid by utilizing the activity prediction model;
[0234] Here, the activity prediction model is a model trained to predict the cleavage activity of the CRISPR / Cas system by receiving an aligned target nucleic acid sequence, an aligned guide nucleic acid sequence, mismatch information, and additional input variables as input.
[0235] By inputting each of the above one or more input variables into the above active prediction model, one predicted cut-off activity value is calculated per input variable,
[0236] If there is one predicted cleavage activity value, the predicted cleavage activity value is used as the prediction result,
[0237] If there are two or more predicted cutting activity values, the predicted cutting activity values are weighted averaged with a predetermined weight and used as the prediction result.
[0238] This specification discloses a method for selecting a guide nucleic acid for a CRISPR / Cas system having high selective cleavage activity, comprising:
[0239] (a) a process for determining a first target nucleic acid and a second target nucleic acid;
[0240] (b) a process for deriving multiple guide nucleic acid candidates;
[0241] Here, the plurality of guide nucleic acid candidates include a guide nucleic acid having a sequence that perfectly matches the first target nucleic acid,
[0242] Based on a sequence that perfectly corresponds to the first target nucleic acid, it comprises at least one sequence to which the following modifications have been applied:
[0243] (1) Changing any one base to another base;
[0244] (2) Change any two bases to other bases;
[0245] (4) Remove any one base from the sequence;
[0246] (5) Adding one random base at any position in the sequence; or
[0247] (6) Any combination of the above (1) to (5) variations;
[0248] (c) a process of obtaining a predicted value of the first target nucleic acid cleavage activity for each of the guide nucleic acid candidates according to the above method;
[0249] (d) a process of obtaining a predicted value of the second target nucleic acid cleavage activity for each of the guide nucleic acid candidates according to the above method;
[0250] (e) for each of the above guide nucleic acid candidates, a process of deriving a selective cleavage activity for the first target nucleic acid compared to the second target nucleic acid using the first target nucleic acid cleavage activity prediction value and the second target nucleic acid cleavage activity prediction value; and
[0251] (f) A process for selecting a guide nucleic acid having high selective cleavage activity according to a predetermined standard using the selective cleavage activity value for the first target nucleic acid compared to the second target nucleic acid.
[0252]
[0253] Hereinafter, the present invention will be described in more detail through specific implementations and examples with reference to the attached drawings. It should be noted that the attached drawings include some, but not all, implementations of the invention. The invention disclosed by this specification may be implemented in various ways and is not limited to the specific implementations described herein. These implementations should be considered as provided to satisfy the legal requirements applicable to this specification. Those skilled in the art will be able to think of many modifications and other implementations of the invention disclosed herein. Therefore, the invention disclosed herein is not limited to the specific implementations described herein, and it should be understood that modifications and other implementations thereof are also included within the scope of the claims.
[0254] Definition of Terms
[0255] approximately
[0256] The term "about" as used herein means an amount, level, value, number, frequency, percentage, dimension, size, amount, weight, or length that varies by about 30, 25, 20, 15, 10, 9, 8, 7, 6, 5, 4, 3, 2, 1%, or 0% with respect to a reference amount, level, value, number, frequency, percentage, dimension, size, amount, weight, or length.
[0257] nucleic acids
[0258] The term "nucleic acid" used in this specification broadly refers to DNA, RNA, and their various forms found in living organisms, as well as chemically and physically modified nucleic acids. Unless otherwise specified, the term "nucleic acid" should not be interpreted narrowly to DNA or RNA, but rather should be interpreted appropriately according to the context. The term encompasses all other meanings recognizable to those skilled in the art.
[0259] Nucleic acid sequence notation
[0260] The symbols A, T, C, G, and U used herein are to be interpreted as meanings understood by a person skilled in the art. They may be appropriately interpreted as bases, nucleosides, or nucleotides in DNA or RNA, depending on the context and technology. For example, when referring to a base, they may be interpreted as adenine (A), thymine (T), cytosine (C), guanine (G), or uracil (U) themselves, respectively; when referring to a nucleoside, they may be interpreted as adenosine (A), thymidine (T), cytidine (C), guanosine (G), or uridine (U), respectively; and when referring to a nucleotide in a sequence, they should be interpreted to mean a nucleotide containing each of the above nucleosides.
[0261] Exogenous, or derived from exogenous
[0262] In this specification, the term "exogenous" is primarily used to describe the origin of a component within a cell. When referring to an "exogenous component," it means that the component is not a gene / protein / other component inherent to the cell, but a component introduced from outside the cell. Furthermore, the term "exogenous" in this specification also includes the meaning of "exogenously derived." An "exogenously derived component" refers to a case where the component is not directly introduced from outside the cell, but is derived (e.g., through replication, transcription, translation, or expression) from an exogenous component. The above terms encompass all other meanings that would be recognized by a person skilled in the art.
[0263] Amino acid sequence notation
[0264] Unless otherwise stated, when describing amino acid sequences in this specification, amino acid single-letter notation or three-letter notation is used, and is written in the N-terminal to C-terminal direction. For example, when written as RNVP, it means a peptide in which arginine, asparagine, valine, and proline are sequentially connected from the N-terminal to the C-terminal. Another example, when written as Thr-Leu-Lys, it means a peptide in which threonine, leucine, and lysine are sequentially connected from the N-terminal to the C-terminal. In the case of amino acids that cannot be expressed in the single-letter notation, other letters are used and additional explanations are provided.
[0265] The notation for each amino acid is as follows: Alanine (Ala, A); Arginine (Arg, R); Asparagine (Asn, N); Aspartic acid (Asp, D); Cysteine (Cys, C); Glutamic acid (Glu, E); Glutamine (Gln, Q); Glycine (Gly, G); Histidine (His, H); Isoleucine (Ile, I); Leucine (Leu, L); Lysine (Lys, K); Methionine (Met, M); Phenylalanine (Phe, F); Proline (Pro, P); Serine (Ser, S); Threonine (Thr, T); Tryptophan (Trp, W); Tyrosine (Ty, Y); and Valine (Val, V).
[0266] Chapter 1. High-throughput assessment of in vitro activity of the CRISPR / Cas system
[0267] High-throughput assessment method for in vitro activity of CRISPR / Cas systems
[0268] generalization
[0269] This specification discloses a method for high-throughput assessment of the in vitro activity of a CRISPR / Cas system. The method aims to assess the cleavage activity of various target-guide combinations in a high-throughput manner in an in vitro environment by varying the guide nucleic acid and target nucleic acid of the CRISPR / Cas system. The method can be conceptualized as the following steps: (a) creating numerous independent reaction compartments using fixed cells (manufacturing microchamber cells), (b) independently inducing a predetermined reaction in each reaction compartment (inducing a compartmentalized reaction), and (c) analyzing the reaction results (assessing the activity of the CRISPR / Cas system). Through step (a), microchamber cells are created, each capable of inducing a specific reaction in a cell, while simulating an in vitro environment. Through step (b), the reaction between the actual CRISPR / Cas complex and the target nucleic acid occurs in each microchamber cell, independently and in parallel. Through the above process (c), the reactions occurring in each microchamber cell are analyzed, and based on this, the cleavage activity exhibited when the CRISPR / Cas system is applied to various guide-target combinations is evaluated. This specification also discloses 1) microchamber cells functioning as independent reaction chambers used in implementing the above method, and 2) a target-guide library for inducing reactions under various conditions. The invention will now be described in more detail with reference to the flowcharts of FIGS. 66 to 72 and the identification numbers described therein.
[0270] Assessing the in vitro activity of the CRISPR / Cas system
[0271] In the high-throughput evaluation method of the in vitro activity of the CRISPR / Cas system, the in vitro activity of the CRISPR / Cas system refers to 1) in an in vitro environment, 2) the activity of the CRISPR / Cas system comprising a specific guide nucleic acid and a specific Cas protein, 3) cleaving a specific target nucleic acid. The cleavage activity encompasses not only the cleavage activity of a guide nucleic acid sequence that is perfectly matched to the target nucleic acid sequence, but also the cleavage activity of a guide nucleic acid sequence that contains one or more mismatches. Unless otherwise stated, when referring to a target nucleic acid and a guide nucleic acid in this specification, it refers to a target nucleic acid having any sequence and a guide nucleic acid having any sequence. Furthermore, when referring to "the reaction of a target-guide vector" or "the activity of a target-guide vector", it refers to the reaction or cleavage activity of a CRISPR / Cas complex comprising a guide nucleic acid included in the vector and a target nucleic acid included in the vector. When the target nucleic acid and guide nucleic acid are specified in context, the above terms may be abbreviated as "target-guided reaction" or "target-guided activity." A more detailed description of the target-guided vector is provided in the [Target-Guided Library] section.
[0272] microchamber cells
[0273] The high-throughput method for assessing the in vitro activity of a CRISPR / Cas system disclosed herein utilizes microchamber cells. Each microchamber cell is an independent reaction chamber designed to allow a CRISPR / Cas system comprising a specific guide nucleic acid to interact with a specific target nucleic acid. The microchamber cells are characterized by: 1) being conditioned to mimic an in vitro environment; 2) containing one type of target nucleic acid or a nucleic acid encoding the same per cell; 3) containing one type of guide nucleic acid or a nucleic acid encoding the same per cell; and 3) being free of cell-to-cell cross-reactivity. Cell-to-cell cross-reactivity is free of, or is extremely unlikely to occur, when a guide nucleic acid (i.e., a CRISPR / Cas complex comprising the same) contained in one microchamber cell reacts with a target nucleic acid contained in another microchamber cell. For example, the microchamber cells may be fixed cells transfected with a target-guide vector. Fixation of the cells causes cell death, inhibiting their vital activity and thus creating an internal environment similar to an in vitro environment. Fixed cells maintain their cell membrane structure, but some of the membrane is damaged, allowing some substances to enter and exit. Experimental results indicate that no cross-reactivity of targeting-guided vectors occurs between fixed cells. Therefore, after transfection with the targeting-guided vector, fixed cells can function as microchamber cells. Various implementations of microchamber cells that satisfy the above conditions are described in more detail in the [Microchamber Cells] section below.
[0274] Target-Guide Library
[0275] The high-throughput method for assessing the in vitro activity of the CRISPR / Cas system disclosed herein utilizes a target-guide library. The target-guide library comprises various target-guide vectors (or target-guide pairs). Each target-guide vector refers to a unit of a vector comprising a target nucleic acid or a nucleic acid encoding the target nucleic acid, and a guide nucleic acid or a nucleic acid encoding the target nucleic acid. By appropriately treating a cell population with the target-guide library, microchamber cells transfected with only one target-guide vector can be produced in a high-throughput manner. Furthermore, appropriately treating the microchamber cells can induce a reaction of the target-guide vector. The target-guide library is a key component for generating various combinations of target-guide reactions in a high-throughput manner. Various implementations of the target-guide library are described in more detail in the [Target-Guide Library] section below.
[0276] High-Throughput Evaluation of Activity Method #1 - Manufacturing Microchamber Cells (S1100)
[0277] The high-throughput method for evaluating the in vitro activity of the CRISPR / Cas system disclosed in this specification comprises the process of preparing microchamber cells. This process may include specific steps required depending on how the microchamber cells are implemented. For example, when preparing fixed microchamber cells after transfection with a targeting-guide vector, the process may include preparing a cell population; treating the cell population with a targeting-guide library; and treating the cell population with a fixation reagent. Various implementations of the above steps are described in more detail in the [Main Step of the High-Throughput Evaluation Method #1 - Preparing Microchamber Cells] section below.
[0278] High-Throughput Evaluation of Activity Method #2 - Inducing Compartmentalized Responses Using External Factors (S1200)
[0279] The high-throughput evaluation method for the in vitro activity of the CRISPR / Cas system disclosed in this specification includes a process of inducing a compartmentalized reaction using an external factor. This process is a process of inducing a target-guided reaction in each of the manufactured microchamber cells. This process may include specific steps as needed depending on how the manufactured microchamber cells are implemented. For example, if the manufactured microchamber cells are fixed cells transfected with a target-guided vector, the process of inducing a compartmentalized reaction using an external factor may include a process of processing a transcription enzyme and a process of processing a Cas protein. Various implementations of the above process are described in more detail in the [Main Process of the High-Throughput Evaluation Method #2 - Inducing a Compartmentalized Reaction] section below.
[0280] High-Throughput Evaluation of Activity Method #3 - Evaluating the Activity of the CRISPR / Cas System (S1300)
[0281] The high-throughput method for assessing the in vitro activity of a CRISPR / Cas system disclosed in this specification comprises a process for assessing the activity of the CRISPR / Cas system. This process comprises obtaining nucleic acid sequence information contained in microchamber cells in which compartmentalized reactions are induced, and processing the information to obtain the activity of various target-guide combinations. The process may include detailed steps necessary for high-throughput analysis of which reactions occurred within each microchamber cell. Various implementations of the process are described in more detail in the following section [Main Step of the High-Throughput Evaluation Method #3 - Assessing the Activity of the CRISPR / Cas System].
[0282] Method Feature #1 - Each cell is used as a reaction chamber.
[0283] This high-throughput method for assessing the in vitro activity of the CRISPR / Cas system is characterized by using the cells themselves as independent reaction chambers. Therefore, existing well-established techniques can be utilized to condition each cell to allow the intended reaction to occur and to immobilize the cells to mimic an in vitro environment. To the best of the knowledge of the inventors of this application, no previous research has attempted to use cells as reaction chambers in an in vitro environment. The inventors have intensively studied and verified 1) whether it is possible to immobilize cells to mimic an in vitro environment and 2) whether cross-reactions between each immobilized cell can be blocked, allowing the immobilized cells to function as independent reaction chambers, thereby completing the present invention.
[0284] Method Feature #2 - High-throughput activity assessment possible
[0285] Most of the steps involved in the above method treat the cell population as a single unit, rather than individual cells. In other words, the method is a high-throughput method. This method allows for the simultaneous testing of numerous combinations of Cas proteins, target nucleic acids, and guide nucleic acids, with the results readily available.
[0286] Feature #3: Unlimited Scalability
[0287] This method can be scaled up virtually without limitation, simply by selecting appropriate cells. In this specification, the method was performed using a target-guided library containing over 100,000 different species, and theoretically, it would not be difficult to perform the method using a larger number of target-guided libraries. Therefore, it is highly suitable for high-throughput exploration of target-guided reactions on a very large scale.
[0288] microchamber cells
[0289] generalization
[0290] This specification discloses microchamber cells and a method for high-throughput assessment of the in vitro activity of a CRISPR / Cas system using the same. The term "microchamber cells" refers to a plurality of microchamber cell populations, or microchamber cells included in a specific cell population. Each microchamber cell is 1) conditioned to mimic an in vitro environment; 2) comprises one type of target nucleic acid or a nucleic acid encoding the same per cell; and 3) comprises one type of guide nucleic acid or a nucleic acid encoding the same per cell. The microchamber cells do not cross-react or have a very low probability of cross-reacting between the individual microchamber cells contained therein. Each microchamber cell can induce target-guided reactions contained therein to occur in an in vitro environment. Furthermore, since cross-reactions do not occur or are negligible within the microchamber cells, various combinations of target-guided reactions can be performed simultaneously in a high-throughput manner. The term "microchamber cell" as used in this specification encompasses all cells in a state that changes as a result of high-throughput assessment of the in vitro activity of the CRISPR / Cas system. Accordingly, the microchamber cell may further comprise additional components within it.
[0291] Types of cells
[0292] The type of the above microchamber cells is not limited, as long as they satisfy the conditions described above. For example, the above microchamber cells may be prokaryotic. In another example, the above microchamber cells may be eukaryotic.
[0293] Core Component #1 - Target Nucleic Acid
[0294] The above microchamber cells contain a target nucleic acid or a nucleic acid encoding the target nucleic acid. The target nucleic acid is a nucleic acid of a certain length that is the target of a CRISPR / Cas complex reaction. The target nucleic acid includes 1) a portion that can be recognized by the Cas protein, and 2) a portion that can interact with a guide nucleic acid. The target nucleic acid can be cleaved by the CRISPR / Cas complex as a high-throughput evaluation method for in vitro activity is performed. In other words, the above microchamber cells can contain a fragment of the target nucleic acid.
[0295] Core Component #2 - Guide Nucleic Acids
[0296] The above microchamber cells contain a guide nucleic acid or a nucleic acid encoding the same. The guide nucleic acid comprises 1) a scaffold capable of interacting with a Cas protein, and 2) a guide domain capable of interacting with a target nucleic acid. The guide nucleic acid binds to the corresponding Cas protein to form a CRISPR / Cas complex, and guides the CRISPR / Cas complex to the target.
[0297] Containing one type of target nucleic acid and guide nucleic acid
[0298] A single microchamber cell is a cell that triggers a single target-guided reaction. Therefore, the microchamber cell comprises one type of target nucleic acid or a nucleic acid encoding the same; and one type of guide nucleic acid or a nucleic acid encoding the same. Here, each type of nucleic acid is distinguished based on its nucleic acid sequence.
[0299] May include additional components
[0300] The above microchamber cells may further include additional components depending on the stage of the high-throughput assessment method for the in vitro activity of the CRISPR / Cas system. That is, the components included may vary depending on the time point at which the microchamber cells are produced. The above microchamber cells may be cells 1) manufactured as microchamber cells, 2) introduced with external factors to induce target-guided reactions, or 3) completed with the intended target-guided reactions. For example, the above microchamber cells may be cells 1) introduced with a target-guided vector and 2) immobilized to mimic an in vitro environment. In another example, the above microchamber cells may be cells 1) introduced with a target-guided vector, 2) immobilized to mimic an in vitro environment, and 3) additionally introduced with external factors such as Cas proteins. Specific additional components that may be included are described in more detail below.
[0301] Additional Component #1 - Transcription Enzyme
[0302] The above microchamber cells may comprise a transcription enzyme. The above transcription enzyme may be introduced into the above microchamber cells during a process for inducing a compartmentalized reaction in a high-throughput evaluation method for the in vitro activity of the CRISPR / Cas system. The above transcription enzyme transcribes an encoding nucleic acid contained in the target-guide vector. The nucleic acid transcribed by the above transcription enzyme may vary depending on the specific implementation of the above high-throughput evaluation method for the in vitro activity of the above CRISPR / Cas system. For example, if the above target-guide vector comprises target DNA and DNA encoding a guide RNA, the above transcription enzyme may transcribe the DNA encoding the guide RNA to express the guide RNA.
[0303] Additional Component #2 - Cas Protein
[0304] The above microchamber cells may contain a Cas protein. The above Cas protein may be introduced into the above microchamber cells during a process for inducing a compartmentalized reaction in a high-throughput evaluation method for the in vitro activity of the CRISPR / Cas system. The above Cas protein binds to a guide nucleic acid within the microchamber cells to form a CRISPR / Cas complex, and the above CRISPR / Cas complex reacts with a target nucleic acid within the microchamber cells. For example, the above Cas protein may bind to a guide RNA within the microchamber cells to form a CRISPR / Cas complex, and the above CRISPR / Cas complex may react with a target DNA within the microchamber cells.
[0305] Additional Component #3 - Restriction Enzymes
[0306] The above microchamber cells may contain a restriction enzyme. The above Cas protein may be introduced into the above microchamber cells during a process for inducing a compartmentalized reaction during a high-throughput evaluation of the in vitro activity of the CRISPR / Cas system. The above restriction enzyme cleaves the restriction site, if present in the above targeting-guide vector. For example, the above restriction enzyme may cleave the restriction site located at the end of the DNA encoding the guide RNA contained in the targeting-guide vector.
[0307] Core components and additional components are exogenous components
[0308] The above core components and additional components are collectively referred to as exogenous components. Here, "exogenous components" refer to components that are contained within a cell but are not inherent genes or components of the cell, but rather originate from outside the cell. The term encompasses not only components directly introduced from outside, but also components derived from exogenous components. For example, if a nucleic acid encoding a guide nucleic acid is delivered to a cell, the nucleic acid encoding the guide nucleic acid is an exogenous component. Furthermore, if the nucleic acid encoding the guide nucleic acid is transcribed and the guide nucleic acid is expressed, the guide nucleic acid is a component derived from the exogenous component and is also referred to as an exogenous component according to this specification. For another example, if a target nucleic acid is delivered to a cell, the target nucleic acid is an exogenous component. Furthermore, if the target nucleic acid is cloned within the cell during the process of culturing and proliferating the cell, the cloned target nucleic acid is a component derived from the exogenous component and is also referred to as an exogenous component. This specification uses the term "endogenous components" to distinguish them from the exogenous components mentioned above. Endogenous components are components inherent in the cell, rather than exogenous components.
[0309] Contains components in various forms
[0310] As described above, the components contained in the microchamber cells may interact with each other and exist in various forms as the high-throughput evaluation method of the in vitro activity of the CRISPR / Cas system is performed. For example, the microchamber cells may comprise a target nucleic acid or a nucleic acid encoding the same, and a target-guide vector comprising the guide nucleic acid or a nucleic acid encoding the same. Here, the target-guide vector is described in more detail in the [Target-Guide Vector and Target-Guide Library] section below. In another example, the microchamber cells may independently comprise the target nucleic acid and the guide nucleic acid. In another example, the microchamber cells may comprise a CRISPR / Cas complex in which the guide nucleic acid is bound to a Cas protein. In another example, the microchamber cells may comprise a nucleic acid encoding the guide nucleic acid, a transcription enzyme, and a guide nucleic acid expressed by the two components. In another example, the microchamber cells may comprise a CRISPR / Cas complex in which the guide nucleic acid and the Cas protein are bound, and a fragment of the target nucleic acid.
[0311] Microchamber Cell Condition #1 - Mimicking the in vitro environment
[0312] These microchamber cells are characterized by being conditioned to mimic an in vitro environment. This mimicking of an in vitro environment means that the reactions occurring between substances within these microchamber cells are identical to, or close to, those occurring in vitro, rather than in vivo. Therefore, unlike typical cells, these microchamber cells are largely unaffected by the cell's endogenous enzymatic activity when the CRISPR / Cas system cleaves nucleic acids. For example, these microchamber cells could be fixed cells. Fixing cells halts their vital functions and largely inhibits protein activity, allowing the intracellular environment to resemble the in vitro environment. Alternatively, these microchamber cells could be cells in which all genes involved in nucleic acid repair, CRISPR / Cas activity and cleavage, and the cleavage of cleaved nucleic acids, have been knocked out, resulting in a loss of function.
[0313] Microchamber Cell State #2 - Exogenous Components Are Active
[0314] The exogenous components of the above microchamber cells are active. This means that these exogenous components can interact or react with each other. For example, if the microchamber cells contain a guide nucleic acid and a Cas protein, these are active and can bind to form a CRISPR / Cas complex. For another example, if the above microchamber cells contain a nucleic acid encoding the guide nucleic acid and a transcription enzyme, the transcription enzyme can transcribe the nucleic acid encoding the guide nucleic acid, thereby expressing the guide nucleic acid.
[0315] Microchamber Cell Status #3 - Cell Membrane Structure and Permeability
[0316] While these microchamber cells maintain their cell membrane structure, they do not completely block external substances, allowing certain exogenous factors to enter the cell interior. Therefore, as the above method is performed, these microchamber cells may incorporate additional components within their cells. For example, these microchamber cells may contain Cas9 proteins, exogenous transcription enzymes, restriction enzymes, and the like.
[0317] Microchamber cells feature blocked cell-to-cell cross-reactivity.
[0318] Microchamber cells, which are a collection of microchamber cells, do not undergo cross-reaction between individual cells, or at least not to a significant degree. When performing high-throughput, diverse target-guided reactions using a microchamber cell population, it is guaranteed that only the intended target-guided reaction occurs within each microchamber cell. For example, if the microchamber cells comprise a first cell and a second cell, the CRISPR / Cas complex of the first cell does not react with the target nucleic acid of the second cell. Even if it does react, the reaction does not occur to a degree significant enough to affect the analysis results.
[0319] Target-guide vectors and target-guide libraries
[0320] generalization
[0321] This specification discloses a target-guide vector and a target-guide library, which are collections of target-guide vectors, used in a high-throughput method for evaluating the in vitro activity of a CRISPR / Cas system. The target-guide vector is a vector designed to induce a single target-guide reaction. The target-guide vector comprises a target nucleic acid or a nucleic acid encoding the target nucleic acid, and a guide nucleic acid or a nucleic acid encoding the target nucleic acid. Furthermore, the target-guide vector may include additional components, such as a promoter, a barcode, and a restriction enzyme cleavage site, to enable expression of the target nucleic acid or guide nucleic acid in an appropriate form. The target-guide library comprises a plurality of target-guide vectors. The target-guide library encompasses various target-guide combinations. Using the target-guide library, microchamber cells capable of inducing various types of target-guide reactions can be produced in a high-throughput manner.
[0322] Target-Guide Vector #1 - Target Nucleic Acid
[0323] The target-guide vector comprises a target nucleic acid or a nucleic acid encoding the target nucleic acid. The target nucleic acid is a nucleic acid of a certain length that is the target of the CRISPR / Cas complex's reaction. The target-guide vector comprises the target nucleic acid itself or an encoding nucleic acid that can be expressed as the target nucleic acid by a transcription enzyme. The target nucleic acid may be double-stranded DNA, single-stranded RNA, or any other nucleic acid that can be cleaved by the CRISPR / Cas complex. The target nucleic acid comprises 1) a portion that interacts with the Cas protein and 2) a portion that can interact with the guide nucleic acid. For example, if the target nucleic acid is double-stranded DNA comprising a target strand and a non-target strand, the non-target strand comprises a protospacer adjacent motif (PAM) recognized by the Cas protein; and a protospacer, and the target strand is a portion that complementarily binds to the protospacer and includes a target portion that binds to the guide RNA. For another example, if the above target nucleic acid is a single-stranded RNA, it includes a protospacer flanking sequence (PFS) recognized by the Cas protein; and a target portion that binds to the guide RNA.
[0324] Target-Guide Vector #2 - Guide Nucleic Acid
[0325] The target-guide vector comprises a guide nucleic acid or a nucleic acid encoding a guide nucleic acid. The guide nucleic acid binds to a corresponding Cas protein to form a CRISPR / Cas complex, and guides the CRISPR / Cas complex to a target. The target-guide vector comprises the guide nucleic acid itself or an encoding nucleic acid capable of being expressed as a guide nucleic acid by a transcription enzyme. The guide nucleic acid comprises 1) a scaffold that interacts with the Cas protein, and 2) a guide domain that interacts with the target nucleic acid. The guide nucleic acid may be RNA or DNA. The guide nucleic acid binds to a Cas protein capable of interacting with the scaffold sequence. The CRISPR / Cas complex, in which the guide nucleic acid and the Cas protein bind, can cleave a nucleic acid targeted by the guide domain of the guide nucleic acid. Here, the term "nucleic acid targeted by the guide domain" includes not only a sequence that is perfectly complementary to the sequence of the guide domain, but also a nucleic acid that is mismatched but that the guide domain binds to and allows the CRISPR / Cas complex to cut. Here, the "mismatched sequence" encompasses all cases where some bases are different, where a bulge occurs in the target nucleic acid, where a bulge occurs in the guide nucleic acid, or any combination thereof.
[0326] Relationship between target nucleic acid sequence and guide nucleic acid sequence
[0327] The target portion of the target nucleic acid included in the target-guide vector and the guide domain of the guide nucleic acid are sequences that can interact with each other. However, the high-throughput evaluation method for the in vitro activity of the CRISPR / Cas system provided in this specification does not distinguish between on-target cleavage activity and off-target cleavage activity, but rather aims to evaluate the "activity of cleaving the target nucleic acid." Therefore, the target portion and the guide domain do not always need to have perfectly complementary sequences. That is, the target portion and the guide domain may have a mismatch of one or more bases. For example, one or more bases in the target portion and the guide domain may not be complementary bases, one or more bulges may occur in the target portion, one or more bulges may occur in the guide domain, or any combination of the above patterns may appear. The sequence of the target nucleic acid and the sequence of the guide nucleic acid included in the above target-guide vector can be freely designed without any special restrictions according to the purpose.
[0328] Other configurations of target-guide vectors #1 - Promoter
[0329] The target-guide vector may include one or more promoters. If the target-guide vector includes a nucleic acid encoding a target nucleic acid, a nucleic acid encoding a guide nucleic acid, or both, it may need to include an additional promoter to express the nucleic acid. Each of the one or more promoters may be operably linked to a nucleic acid encoding a target nucleic acid or a nucleic acid encoding a guide nucleic acid. A transcription enzyme may recognize the promoter and transcribe and express the encoded nucleic acid operably linked thereto.
[0330] Other Target-Guide Vector Configurations #2 - Barcodes and UMIs
[0331] The target-guide vector may include one or more barcodes and one or more Unique Molecular Identifiers (UMIs). The barcodes and UMIs allow for easy identification of how a given target-guide combination reacts during the evaluation of the activity of the CRISPR / Cas system in a high-throughput evaluation method for the in vitro activity of the CRISPR / Cas system. The barcodes contain information about the target nucleic acid sequence and guide nucleic acid sequence contained in the target-guide vector within the library. When the barcodes are included, the sequences of the target nucleic acid and guide nucleic acid contained in the vector containing the barcode can be identified. The UMIs contain information about a specific target-guide vector itself within the library. The UMIs represent unique identification information of the vector. This allows for identification of the vector from which the clones are derived when a single vector is amplified into multiple clones during the polymerase chain reaction (PCR) process. In this specification, both the above barcode and the above UMI portion are collectively referred to as "barcode", which should be interpreted appropriately depending on the context.
[0332] Other configurations of target-guide vectors #3 - Restriction enzyme cleavage sites
[0333] The target-guide vector may include one or more restriction enzyme cleavage sites. These restriction enzyme cleavage sites, independent of the cleavage of the target nucleic acid by the CRISPR / Cas complex, refer to sites capable of being cleaved by a restriction enzyme. These restriction enzyme cleavage sites may be added for various purposes. For example, the restriction enzyme cleavage sites may be added to linearize the circular target-guide vector and facilitate the appropriate expression of the target nucleic acid or guide nucleic acid when the site is cleaved with a restriction enzyme. In another example, the restriction enzyme cleavage sites may be added to cleave the site with a restriction enzyme and ligate an adapter to the cleaved site. Specific implementation examples thereof are disclosed in the [Possible Embodiments of the Invention] section.
[0334] Other Target-Guide Vector Components #4 - Antibiotic Resistance Genes
[0335] The target-guide vector may include one or more antibiotic resistance genes. These antibiotic resistance genes can be utilized as selection factors in high-throughput evaluation methods for the in vitro activity of the CRISPR / Cas system, when a cell selection process is optionally included. Specifically, these antibiotic resistance genes confer resistance to specific antibiotics on cells containing the target-guide vector. Therefore, when a cell population treated with the target-guide library is cultured in a medium containing the specific antibiotic, only cells transfected with the target-guide vector survive. Consequently, cells transfected with the target-guide vector can be selected using the antibiotic resistance genes as selection factors. As long as the above purpose can be achieved, the composition of the antibiotic resistance genes is not limited. Any known composition can be appropriately utilized to incorporate the antibiotic resistance genes into the target-guide vector.
[0336] Target-Guide Library
[0337] This specification discloses a target-guided library. The target-guided library comprises a plurality of target-guided vectors. Furthermore, the target-guided library comprises two or more types of target-guided vectors. As described above, "the same type of target-guided vector" refers to a vector in which the target nucleic acid sequence and the guide nucleic acid sequence contained in the target-guided vector are identical. The target-guided library can be composed of any combination of target-guides, depending on the purpose. Furthermore, once the target-guide combinations to be included in the library are determined, a person skilled in the art can prepare the target-guided library using known techniques.
[0338] Key Steps in High-Throughput Evaluation Methods #1 - Manufacturing Microchamber Cells (S1100)
[0339] generalization
[0340] A method for high-throughput assessment of the in vitro activity of a CRISPR / Cas system disclosed herein comprises the process of producing microchamber cells. The process aims to produce microchamber cells, each containing a variety of targeting-guide vectors, at high throughput. Each of the microchamber cells mimics an in vitro environment, contains a targeting-guide vector, and is capable of inducing compartmentalized reactions (parallel and independent reactions). The process comprises preparing an appropriate cell population; processing a targeting-guide library to deliver the targeting-guide vector into the cells; and fixing the cells using a fixative. As a result of the process, the cell population comprises microchamber cells. The cell population may also include cells that have not been delivered with a targeting-guide vector. The cell population may also include cells delivered with two or more targeting-guide vectors. The cell population ensures that compartmentalized reactions occur within each microchamber cell.
[0341] Preparing cell populations (S1110)
[0342] The process of manufacturing microchamber cells involves preparing a cell population in detail. This process involves preparing a suitable cell population that can become microchamber cells. The cells prepared in this process are living cells. As described above, there are no specific limitations on the cells that can become microchamber cells, and thus, there are no specific limitations on the cells prepared in this process. For example, a prokaryotic cell population or a eukaryotic cell population can be prepared. This process can be appropriately performed by a person skilled in the art using known methods.
[0343] Processing the target-guide library (S1120)
[0344] The process of manufacturing microchamber cells involves processing a target-guide library, which is a detailed process. In this process, the cell population prepared above is treated with a pre-designed target-guide library. The target-guide library can be appropriately prepared using a known method, and specific implementation examples are described in detail in the [Experimental Examples] section. Processing the target-guide library means using an appropriate method to deliver each target-guide vector contained therein into each cell. At this time, the target-guide library is processed so that only one target-guide vector is delivered to most of the cells delivered with the target-guide vector. In other words, after the process of processing the target-guide library, most cells in the above cell population will be either 1) cells that have not been delivered with a target-guide vector, or 2) cells that have been delivered with only one target-guide vector. While the above cell population may contain some cells delivered with more than one target-guide vector, this will not significantly affect the analysis results.
[0345] Processing fixed reagents (S1140)
[0346] The process of manufacturing microchamber cells involves a detailed process involving the treatment of a fixation reagent. In this process, a cell population treated with a target-guide library is treated with the fixation reagent. The cells in the above cell population are fixed with the fixation reagent. As described above, fixation of the cells suspends their vital activities, making the intracellular environment similar to an in vitro environment. In other words, the endogenous enzymatic activity of the cells is inhibited; or the protein activity of the cells may be inhibited. In contrast, the target-guide vector is not inactivated and can be utilized in the subsequent process of inducing a compartmentalized response. Cells in the above cell population that have been transfected with a single target-guide vector undergo the fixation process to become microchamber cells.
[0347] May include a process of expanding cells (S1130)
[0348] The process for manufacturing the above microchamber cells may optionally include a cell expansion process. The cell expansion process refers to a process of increasing the number of cells by proliferating them after processing the target-guide library and before treating with a fixative. This process allows for securing a larger number of cells introduced with the target-guide vector, and furthermore, a larger number of microchamber cells. This process allows for cloning of the target-guide vector within each microchamber cell, thereby increasing the number of cells. The cell expansion process may be appropriately performed using known techniques. This process may be performed at an appropriate time after processing the target-guide library and before treating with a fixative. Furthermore, this process may be performed simultaneously with the cell selection process, or may be performed simultaneously or in any order.
[0349] May include a process for selecting cells (S1130)
[0350] The process of manufacturing the above microchamber cells may optionally include a cell selection process. The cell selection process refers to a process of selecting only cells introduced with a targeting-guide vector from a cell population. Through the above process, only cells introduced with the targeting-guide vector can be obtained. The cell selection process can be appropriately performed using known techniques. For example, each targeting-guide vector can be incorporating an antibiotic resistance gene, and cells treated with the targeting-guide library can be cultured in a medium containing the antibiotic to select the cells. The above process can be performed at an appropriate time after targeting-guide library treatment and before fixation reagent treatment. Furthermore, the above process can be performed together with the cell expansion process, and can be performed simultaneously or in any order.
[0351] Key Step #2 of High-Throughput Evaluation Methods - Inducing a Compartmentalized Reaction (S1200)
[0352] generalization
[0353] The high-throughput assessment method for the in vitro activity of the CRISPR / Cas system disclosed herein comprises a process for inducing a compartmentalized reaction. The purpose of this process is to form a CRISPR / Cas complex comprising a Cas protein and a guide nucleic acid within each microchamber cell prepared in the above process, and to allow this complex to react with a target nucleic acid. As described above, each microchamber cell mimics an in vitro environment, and the microchamber cells (or a cell population including them) ensure that a compartmentalized reaction occurs. Therefore, through this process, various combinations of target-guided reactions occur in parallel within each microchamber cell. Specifically, the process comprises (optionally) processing one or more enzymes; and processing a Cas protein. Through this process, the processed Cas protein and the guide nucleic acid contained in or derived from a target-guide vector contained in each microchamber cell combine to form a CRISPR / Cas complex, which reacts with the target nucleic acid contained in or derived from the target-guide vector. In the above process, the target-guided reaction within each microchamber cell is characterized by compartmentalization. This compartmentalization means that the CRISPR / Cas complex within one microchamber cell does not react with the target nucleic acid within another microchamber cell, or only to a negligible extent.
[0354] Treating with one or more enzymes (S1220)
[0355] The process of inducing the above compartmentalized reaction may include, in detail, treating one or more enzymes. In this case, the one or more enzymes include one or more restriction enzymes, one or more transcription enzymes, or a combination thereof. The one or more enzymes are delivered to each microchamber cell and act on the target-guide vector to cleave it (restriction enzyme) or transcribe a desired portion (transcription enzyme). The process prepares the target-guide reaction to occur or prepares a portion to be used in a subsequent activity assay. For example, the process may include treating one or more restriction enzymes (S1221). The one or more restriction enzymes are delivered to each microchamber cell and cleave the restriction enzyme cleavage site of the target-guide vector. In another example, the process may include treating one or more transcription enzymes (S1222). The one or more transcription enzymes are delivered to each microchamber cell, recognize the corresponding promoter within the target-guide vector, and transcribe the encoding nucleic acid operably linked thereto. At this time, the above-mentioned encoded nucleic acid may be a nucleic acid encoding a guide nucleic acid, a nucleic acid encoding a target nucleic acid, or both.
[0356] Processing Cas proteins (S1210)
[0357] The process of inducing the above compartmentalized response involves processing Cas proteins in detail. In this process, one or more Cas proteins are administered to a cell population containing the microchamber cells. These Cas proteins are delivered to each microchamber cell. These Cas proteins bind to the guide nucleic acid contained in the microchamber cell, forming a CRISPR / Cas complex. This CRISPR / Cas complex then interacts with the target nucleic acid contained in the microchamber cell.
[0358] The order in which the elements are processed is irrelevant.
[0359] The above detailed steps involve processing specific elements, i.e., restriction enzymes, transcriptases, or Cas proteins. As long as the goal of this process is to "induce compartmentalized target-guided reactions within each microchamber cell," the order in which these detailed steps are performed is not critical. For example, in the process of inducing the above compartmentalized reactions, the restriction enzymes, transcriptases, and Cas proteins may be processed all at once, some at once, or in any order.
[0360] Key Step #3 of High-Throughput Evaluation Methods - Assessing the Activity of the CRISPR / Cas System (S1300)
[0361] generalization
[0362] A method for high-throughput assessment of the in vitro activity of a CRISPR / Cas system disclosed in this specification comprises a process for assessing the activity of the CRISPR / Cas system. The process aims to obtain sequence information from microchamber cells in which a target-guided reaction has occurred and to measure the target-guided activity of the CRISPR / Cas system by processing the sequence information. The process comprises collecting nucleic acids from a population of cells after a compartmentalized reaction has been induced; obtaining sequence information from the collected nucleic acids; and assessing the cleavage activity of a guide nucleic acid sequence of the CRISPR / Cas system against a target nucleic acid sequence (target-guided activity) from the nucleic acid sequence information.
[0363] Process for collecting nucleic acids from a cell population (S1310)
[0364] This process involves collecting nucleic acids from a cell population. With the advancement of genetic engineering, methods for collecting nucleic acids (genomic nucleic acids, exogenous nucleic acids, etc.) from individual cells in a cell population have become well-established. This nucleic acid collection can be performed using any known method. As a result of this process, nucleic acids within the cells in the cell population are collected.
[0365] Process for obtaining nucleic acid sequence information from a cell population (S1320)
[0366] In this process, nucleic acid sequence information is obtained from the collected nucleic acids. Here, the nucleic acid sequence information of one unit (one data point) includes: 1) the sequence of the target nucleic acid; 2) the sequence of the guide nucleic acid; 3) whether the target nucleic acid is cleaved; 4) (optionally) whether the target nucleic acid is not cleaved; and 5) (optionally) unique identification information of the target-guide vector. Here, the nucleic acid sequence information of one unit means information about the target-guide reaction that occurred in one microchamber cell. Depending on how the target-guide vector is designed, additional information indicated as optional above can be obtained. With the development of genetic engineering, methods for obtaining nucleic acid sequence information from a population of cells have been well established. For example, the above process can be performed by a person skilled in the art using a known method. As another example, the above process can be performed using an adapter. Specifically, the collected nucleic acids are treated with an adapter that is ligated to a site cleaved by a restriction enzyme or a CRISPR / Cas complex (S1321); Using a primer that binds to the adapter, the nucleic acid to which the adapter has been ligated is amplified by polymerase chain reaction (PCR) (S1322); the amplified nucleic acid can be sequenced to obtain nucleic acid sequence information (S1323). More specific implementation examples are disclosed in the [Possible Embodiments of the Invention] and [Experimental Examples] sections.
[0367] Target-guided activity assay (S1330)
[0368] In this process, the nucleic acid cleavage activity of the CRISPR / Cas system is evaluated based on the obtained information. Specifically, 1) nucleic acid sequence information is grouped based on target nucleic acid sequence, guide nucleic acid sequence information, and other necessary information (S1331), and 2) for each group, target-guide activity is calculated using the number of cleaved target nucleic acids, the number of unreduced target nucleic acids, or both (S1332). Following the above process, the cleavage activity for each target-guide combination can be determined.
[0369] Expandable to assess the cleavage activity of nucleic acid-derived nucleases
[0370] The high-throughput evaluation method for the in vitro activity of the CRISPR / Cas system disclosed in this specification can be performed using a nuclease that cleaves nucleic acids in a similar manner to the CRISPR / Cas system. Specifically, the cleavage activity of a nucleic acid-guided nuclease that 1) functions in conjunction with a guide nucleic acid, 2) cleaves a target nucleic acid targeted by the guide nucleic acid, and 3) can be delivered into a cell can be evaluated using the above method. The nucleic acid-guided nuclease may be referred to as a programmable nuclease. The nucleic acid-guided nuclease may be, for example, a zinc finger nuclease (ZFN); a meganuclease; a transcription activator-like effector nuclease (TALEN); an IscB and omega RNA complex; or a TnpB and omega RNA complex. The above nucleic acid-derived nucleases encompass all known and yet to be discovered nucleic acid-derived nucleases. Furthermore, they include naturally occurring nucleases and their variants, as well as artificial nucleic acid-derived nucleases developed using technologies such as artificial intelligence. For example, Jeffrey et al. (Design of highly functional genome editors by modeling the universe of CRISPR-Cas sequences, Jeffrey . Ruffolo, Stephen Nayfach, Joseph Gallagher, Aadyot Bhatnagar, Joel Beazer, Riffat Hussain, Jordan Russ, Jennifer Yip, Emily Hill, Martin Pacesa, Alexander . Meeske, Peter Cameron, Ali Madani, bioRxiv 2024).04.22.590591; doi: https: / doi.org / 10.1101 / 2024.04.22.590591), or a related technology. In the method described above, the Cas protein can be changed to a nucleic acid-induced nuclease, and the CRISPR / Cas complex can be changed to a nucleic acid-nuclease complex. The same applies to the contents described in Chapters 2 and 3.
[0371] Chapter 2. CRISPR / Cas System Cleavage Activity Prediction Model
[0372] CRISPR / Cas system cleavage activity prediction model and its application
[0373] generalization
[0374] This specification discloses a method for training a cleavage activity prediction model and a pseudo-cleavage activity prediction model for a CRISPR / Cas system. Furthermore, the specification discloses a method for selecting a guide nucleic acid that selectively cleaves a specific nucleic acid by utilizing the pseudo-cleavage activity prediction model.
[0375] The above cleavage activity prediction model is configured to receive a target nucleic acid sequence and a guide nucleic acid sequence as input and output a quantitative prediction value of how well a CRISPR / Cas complex containing a specific Cas protein and the above guide nucleic acid cleaves the target nucleic acid. Specifically, the above cleavage activity prediction model is configured to align the above target nucleic acid sequence and the above guide nucleic acid sequence, and to predict the target-guided cleavage activity by using the aligned target nucleic acid sequence and the aligned guide nucleic acid sequence as key input variables and receiving additional information as input as necessary.
[0376] The above cleavage activity prediction model is trained using data on the activity of a CRISPR / Cas complex comprising a specific Cas protein and various guide nucleic acids to cleave various target nucleic acids. The above data can be obtained, for example, by the method disclosed in Chapter 1. The above cleavage activity prediction model is trained by extracting 1) aligned target nucleic acid sequences, 2) aligned guide nucleic acid sequences, 3) (optionally) mismatch information, and 4) (optionally) additional information based on each target-guide activity data; appropriately preprocessing the above information; and labeling the preprocessed information as input variables and the cleavage activity as an output variable. The above learning method is characterized by appropriately utilizing data augmentation when multiple target-guide alignment patterns can be derived from a specific target-guide combination. Hereinafter, the terms "training a model" and "learning a model" have the same meaning and may be used interchangeably.
[0377] The cleavage activity prediction model learned in this way can be used in a method for selecting a guide nucleic acid that selectively cleaves a target nucleic acid well when a target nucleic acid to be cleaved and a noise nucleic acid that should not be cleaved are determined. The above guide nucleic acid selection method derives various guide sequence candidates and applies the above cleavage activity prediction model to each guide nucleic acid to select the guide nucleic acid that cleaves the target nucleic acid best compared to the noise nucleic acid. At this time, the guide sequence candidate group includes a sequence that completely corresponds to the target nucleic acid (perfect matched) and sequences that are mismatched with the target nucleic acid in various patterns. The above nucleic acid selection method is characterized by deriving guide nucleic acid candidates not only when the target nucleic acid and the guide nucleic acid base are different, but also when a bulge occurs in the target nucleic acid and when a bulge occurs in the guide nucleic acid, and by exploring the cleavage activity of each. Hereinafter, the invention will be described in more detail with reference to the flowcharts of FIGS. 73 to 79 and the identification numbers described therein.
[0378] CRISPR / Cas system cleavage activity prediction model
[0379] This specification discloses a cleavage activity prediction model of a CRISPR / Cas system. The cleavage activity prediction model is configured to receive a target nucleic acid sequence and a guide nucleic acid sequence as input and output a quantitative prediction value regarding how well a CRISPR / Cas complex comprising a specific Cas protein and the guide nucleic acid cleaves the target nucleic acid. More specifically, the cleavage activity prediction model utilizes as input variables 1) aligned target-guide sequence information derived by aligning the input target-guide sequences; and optionally, 2) additional information derived by referencing the target-guide sequence information and the aligned target-guide sequence information, and outputs a quantitative prediction value regarding how well a CRISPR / Cas complex comprising a specific Cas protein and the guide nucleic acid cleaves the target nucleic acid.
[0380] The above activity prediction model can be applied to the following methods: 1) a method for predicting how well a CRISPR / Cas complex will cleave a target nucleic acid given a target nucleic acid sequence and a guide nucleic acid sequence; 2) a method for predicting whether a guide nucleic acid selectively cleaves a specific target nucleic acid compared to other target nucleic acids given two target nucleic acid sequences and a guide nucleic acid sequence; and 3) a method for selecting a guide nucleic acid sequence that selectively cleaves a target nucleic acid given a target nucleic acid sequence and a noise nucleic acid sequence.
[0381] Training method for cutting activity prediction model
[0382] This specification discloses a method for training a cleavage activity prediction model of a CRISPR / Cas system. The method for training the cleavage activity prediction model is performed under the premise that sufficient experimental data measuring how well a CRISPR / Cas complex comprising a specific Cas protein and various guide nucleic acids cleaves various target nucleic acids is secured. The experimental data can be obtained, for example, using the high-throughput method disclosed in Chapter 1. From the target-guide sequence information included in the raw data, 1) aligned target nucleic acid sequences, 2) aligned guide nucleic acid sequences, 3) (optionally) mismatch information, and 4) (optionally) additional information are derived, the information is appropriately preprocessed, each target nucleic acid cleavage activity is labeled as an output value, and a cleavage activity prediction model is derived by appropriately training. Here, in some cases, the target nucleic acid and the guide nucleic acid may be aligned in multiple patterns, in which case each alignment pattern is used to augment the training data. However, multiple aligned sequence information is not always used. Which alignment pattern(s) to use for learning is determined based on the importance of the alignment pattern, such as the alignment score.
[0383] A method for predicting cutting activity using a cutting activity prediction model.
[0384] This specification discloses a method for predicting the cleavage activity of a CRISPR / Cas system. The cleavage activity prediction method predicts target-guide activity using the aforementioned cleavage activity prediction model. Specifically, a target nucleic acid sequence and a guide nucleic acid sequence are aligned to derive aligned target-guide nucleic acid information, and (optionally) additional information is derived and input into the cleavage activity prediction model to obtain a cleavage activity prediction value. In this case, when the target nucleic acid and guide nucleic acid are aligned in multiple patterns, each alignment pattern and the corresponding additional information can be used as different input variables to obtain multiple predicted values, which can then be weighted and averaged to derive a single predicted value. In this case, whether to augment input variables with multiple alignment patterns, which alignment pattern to input into the model, and the weighting of each output value in the weighted average are determined according to predetermined criteria. For example, when alignment scores are derived for target-guide alignment patterns and there are two or more alignment patterns with the highest alignment scores, all of these alignment patterns can be used as input variables, and the arithmetic average of each output value can be used as a predicted value.
[0385] Application of the Cut Activity Prediction Model #1 - Target Selectivity Prediction Method
[0386] This specification discloses a method for predicting how well a target nucleic acid is selectively cleaved when two target nucleic acids and one guide nucleic acid are given. Specifically, when the two target nucleic acids are referred to as a first target nucleic acid and a second target nucleic acid, respectively, the method comprises: obtaining a cleavage activity prediction value (a first value) for the first target nucleic acid and the guide nucleic acid using the cleavage activity prediction method; obtaining a cleavage activity prediction value (a second value) for the second target nucleic acid and the guide nucleic acid using the cleavage activity prediction method; and predicting the activity of cleaving the second target nucleic acid relative to the first target nucleic acid (or vice versa) based on the first value and the second value. The method for predicting target selectivity can further be applied to a guide nucleic acid selection method with high target selectivity.
[0387] Application of the Cleavage Activity Prediction Model #2 - A Guide Nucleic Acid Selection Method with High Target Selectivity
[0388] This specification discloses a method for selecting a guide nucleic acid that selectively cleaves only a target nucleic acid among a target nucleic acid to be cleaved and a noise nucleic acid to be left uncleaved. Specifically, guide nucleic acid candidates are derived, including a sequence that completely corresponds to the target nucleic acid (a perfect match) and various patterns of mismatches (mismatches); for each guide nucleic acid candidate, a predicted value for target nucleic acid cleavage activity is derived using a cleavage activity prediction model; for each guide nucleic acid candidate, a predicted value for noise nucleic acid cleavage activity is derived using the cleavage activity prediction model; and based on the two predicted values, a guide nucleic acid that best cleaves the target nucleic acid compared to the noise nucleic acid is selected. This guide nucleic acid selection method is characterized by utilizing a much wider range of mismatch sequences as candidates than conventionally known methods. Specifically, guide nucleic acid candidates are derived by encompassing not only cases where the target nucleic acid and the guide nucleic acid base mismatch, but also cases where a bulge occurs in the target nucleic acid and cases where a bulge occurs in the guide nucleic acid, and the cleavage selectivity of each target nucleic acid is examined.
[0389] A feature of the cutting activity prediction model learning method - Augmenting data using target-guide alignment information.
[0390] The method for learning a gastric cleavage activity prediction model is characterized by leveraging data augmentation when multiple alignment patterns are available when aligning target and guide nucleic acids. This is because the gastric cleavage activity prediction model utilizes not only cases where the target and guide nucleic acids are perfectly matched, but also cases where they are mismatched. Even when the target and guide nucleic acids are mismatched or completely identical, they can be aligned in various patterns depending on their sequences. To encompass all these cases, the gastric cleavage activity prediction model augments data using multiple alignment patterns. Data augmentation means deriving multiple aligned target and aligned guide nucleic acid sequences from the same target and guide nucleic acid sequences, and utilizing each of these as training data. However, data augmentation is not performed automatically when multiple alignment patterns are found. Rather, the importance of each alignment pattern is considered when determining the data points to be augmented. This contributes to improving the performance of the gastric cleavage activity prediction model.
[0391] Characteristics of a highly target-selective guide nucleic acid selection method - Exploring a wide range of mismatched sequences
[0392] This method for selecting guide nucleic acids with high target selectivity is characterized by including a wide range of mismatched sequences in the search. This method does not distinguish between on-target cleavage activity and off-target cleavage activity, but rather aims to determine whether a target nucleic acid is more effectively cleaved than a noise nucleic acid. Therefore, unlike conventional methods, this method includes guide nucleic acid candidates not only when one or more bases are mismatched with the target nucleic acid, but also when a bulge is present in the target nucleic acid and when a bulge is present in the guide nucleic acid. This method, by selecting from a wider pool of candidates, increases the probability of identifying the most selective guide nucleic acid.
[0393] CRISPR / Cas system cleavage activity prediction model
[0394] generalization
[0395] This specification discloses a cleavage activity prediction model for the CRISPR / Cas system. This cleavage activity prediction model focuses on predicting target-guided cleavage activity from aligned target-guide sequence information. To enhance the accuracy and reliability of this cleavage activity prediction model, various additional information can be used, and appropriate data preprocessing steps can be added. This cleavage activity prediction model can be appropriately constructed using known techniques.
[0396] Below, an implementation example of the above-described truncation activity prediction model is outlined, and its structure is explained based on this. However, the above-described truncation activity prediction model is not limited to the implementation example below, and certain processes, configurations, or information used may be omitted as needed, as long as the core features described above are maintained. Various implementation examples are described in the [Possible Embodiments of the Invention] section.
[0397] In one embodiment, a cleavage activity prediction model of a CRISPR / Cas system is configured to receive a target nucleic acid sequence and a guide nucleic acid sequence (collectively referred to as target-guide sequence information) as input and output a cleavage activity prediction value. The cleavage activity prediction model includes a feature extractor, a concatenation layer, and an output module. The cleavage activity prediction model is configured to receive the target-guide sequence information as input and perform the following: derive an aligned target nucleic acid sequence, an aligned guide nucleic acid sequence, and mismatch information (collectively referred to as aligned target-guide sequence information) from the target-guide sequence information; derive additional information by referring to the target-guide sequence information and the aligned target-guide sequence information; appropriately preprocess the aligned target-guide sequence information and input it to the feature extractor to extract sequence feature information; appropriately preprocess the additional information and combine it with the sequence feature information in a concatenation layer to derive input information; and input the input information to the output module to output a cleavage activity prediction value.
[0398] Assuming that a specific Cas protein is used
[0399] The above cleavage activity prediction models assume the use of a specific Cas protein. That is, within a single cleavage activity prediction model, the type of Cas protein is determined. Accordingly, it is also assumed that the sequence of the protospacer adjacent motif (PAM), which is information related to the Cas protein, and the scaffold sequence of the guide nucleic acid are also determined. Accordingly, various cleavage activity prediction models can be distinguished based on the type of Cas protein used.
[0400] Input Variable #1 - Target-Guide Sequence Information
[0401] The above cleavage activity prediction model is configured to receive target nucleic acid sequence information and guide nucleic acid sequence information as input. The target nucleic acid sequence includes a sequence of a portion that interacts with the guide nucleic acid. The guide nucleic acid sequence includes a sequence of a portion that interacts with the target nucleic acid, i.e., a guide domain. For example, if the target nucleic acid is double-stranded DNA and the guide nucleic acid is RNA, the target nucleic acid sequence includes a protospacer sequence, and the guide nucleic acid sequence includes a spacer sequence.
[0402] Input variable #2 - Aligned target-guide sequence information
[0403] The above cleavage activity prediction model is configured to derive aligned target-guide sequence information from the above target-guide sequence information. This is the result of aligning the target nucleic acid sequence and the guide nucleic acid sequence using an alignment algorithm. The aligned target nucleic acid sequence and the aligned guide nucleic acid sequence each include gaps. Furthermore, the aligned target-guide sequence information may further include mismatch information. The mismatch information compares the aligned target nucleic acid sequence and the aligned guide nucleic acid sequence, and provides information on whether, for each position: 1) the bases of the two nucleic acids match (correspond); 2) the bases of the two nucleic acids do not match; 3) a gap is formed in the target nucleic acid; or 4) a gap is formed in the guide nucleic acid. The length of the aligned nucleic acid sequence may vary depending on the target nucleic acid sequence and the guide nucleic acid sequence and the alignment method. However, for use in the cleavage activity prediction model, it is preferable for all input data to have a fixed size. Accordingly, for use in the above cleavage activity prediction model, a blank sequence or blank information can be added to the aligned target-guide sequence information to ensure a constant length. For example, all of the above aligned target-guide sequence information has a fixed length of 64 nt, and when the target-guide is aligned, if the sequence length including gaps is less than 64 nt, the remaining portion can be filled with empty information (e.g., 0 or a category value corresponding to empty information) to fill the 64 nt length (see the top of Figure 28).
[0404] Input Variable #3 - Additional Information
[0405] The above cleavage activity prediction model is configured to derive additional information by referring to the above target-guide sequence information and the above aligned target-guide sequence information. The above additional information may be one or more selected from the following: a Protospacer Adjacent Motif (PAM) sequence; and whether a mutation has occurred in the above PAM sequence; the type of the 20th base of the target nucleic acid sequence; the type of the guide nucleic acid; the melting temperature (T) of the guide nucleic acid sequence. m ), T of the target nucleic acid sequence m ; Minimal Free Energy (MFE) of the guide nucleic acid sequence; MFE of the guide nucleic acid sequence including the scaffold; total number of mismatches, target nucleic acid bulges, and guide nucleic acid bulges; T of matched guide nucleic acid-target nucleic acid m ; and the free energy change (ΔG) during guide nucleic acid-target nucleic acid hybridization H ).
[0406] Feature Extractor
[0407] The above-described cutting activity prediction model includes a feature extractor. The feature extractor is configured to receive aligned target-guide sequence information (or appropriately preprocessed information) and derive sequence feature information. The structure and configuration of the feature extractor are not limited, as long as it can extract appropriate sequence feature information. For example, the feature extractor may be a transformer encoder.
[0408] Concatenation Layer
[0409] The above-mentioned cut-off activity prediction model includes a connection layer. The connection layer is configured to receive the above-mentioned sequence feature information and the above-mentioned additional information (or appropriately preprocessed information) as input and derive input variables. The structure and composition of the connection layer are not otherwise restricted, as long as it can derive the input variables input to the output module.
[0410] Output Module
[0411] The above-mentioned cutting activity prediction model includes an output module. The above-mentioned output module is configured to receive input variables derived from the above-mentioned connection layer and output target-guided cutting activity. The above-mentioned output module is not limited in its structure or configuration as long as it can output the above-mentioned target-guided cutting activity. For example, the above-mentioned output module may be configured as a fully connected layer.
[0412] Output Variable - Target-Guided Cutting Active
[0413] The above cleavage activity prediction model is configured to output target-guided cleavage activity. Target-guided cleavage activity is a quantitative prediction of the degree to which a CRISPR / Cas complex, including a predetermined Cas protein and an input guide nucleic acid, cleaves the input target nucleic acid. The physical meaning of this output value can be determined based on the data used to train the above cleavage activity prediction model.
[0414] A method for learning a prediction model for the cleavage activity of the CRISPR / Cas system.
[0415] generalization
[0416] This specification discloses a learning method for a CRISPR / Cas system cleavage activity prediction model. The learning method utilizes data measuring various target-guide reactions. The learning method is characterized by augmenting the data by utilizing the alignment results between target and guide nucleic acids.
[0417] Below, an implementation example of a learning method for the above-mentioned truncation activity prediction model is outlined, and a learning method for the truncation activity prediction model is described based on this. However, the learning method for the above-mentioned truncation activity prediction model is not limited to the implementation example below, and certain processes, configurations, or information used may be omitted as needed, as long as the core features described above are maintained. Various implementation examples for this are described in the [Possible Embodiments of the Invention] section.
[0418] In one embodiment, the learning method includes: obtaining raw data; deriving training data from the raw data; and learning a cleavage activity prediction model of a CRISPR / Cas system. During the process of deriving the training data, the target nucleic acid and the guide nucleic acid are aligned to derive aligned target-guide sequence information. Furthermore, additional information is derived by referencing the target-guide sequence information and the aligned target-guide sequence information. At this time, the data is augmented based on the target-guide alignment results. When the target nucleic acid and the guide nucleic acid can be aligned in multiple patterns and the multiple patterns satisfy a certain criterion, all aligned target-guide sequence information satisfying the criterion are used. Accordingly, the additional information derived by referencing the aligned target-guide sequence information may also vary. The data is augmented by including all of the aligned target-guide sequence information and the additional information resulting from the aligned target-guide sequence information in the training data. The input and output variables of the cleavage activity prediction model, as well as the representative structure of the cleavage activity prediction model, are described in the above paragraph.
[0419] Below, the learning method is described with a focus on its features.
[0420] Process of obtaining raw data (S2100)
[0421] The learning method for the above CRISPR / Cas system cleavage activity prediction model includes a process of obtaining raw data. The raw data of the above learning method includes target-guide reaction data. The target-guide reaction data refers to data measured by varying the guide nucleic acid and target nucleic acid to determine the activity of a CRISPR / Cas complex comprising a Cas protein and a guide nucleic acid in cleaving the target nucleic acid. For convenience of description, the minimum unit of the above raw data is defined as a raw data point. In other words, a bundle of information required for one target-guide reaction is referred to as one raw data point. The above raw data point includes 1) a target nucleic acid sequence, 2) a guide nucleic acid sequence, and 3) target-guide activity.
[0422] Training Data Derivation Process #1 - Deriving Aligned Target-Guide Sequence Information (S2200)
[0423] The learning method of the cleavage activity prediction model of the above CRISPR / Cas system includes a process of deriving training data. The minimum unit of the above training data is defined as a training data point. The above training data point is derived from the corresponding raw data point. As described above, the raw data point includes information on a target nucleic acid sequence and a guide nucleic acid sequence. An alignment algorithm is applied to align the target nucleic acid sequence and the guide nucleic acid sequence of each of the above raw data points, thereby deriving an aligned target nucleic acid sequence and an aligned guide nucleic acid sequence (S2210). In addition, additional information is derived by referring to the above target-guide sequence information and the aligned target-guide sequence information, and training data including the aligned target-guide sequence information and the additional information is derived (S2220). Here, the additional information is as described in the [CRISPR / Cas system cleavage activity prediction model] paragraph.
[0424] Training Data Derivation Process #2 - Augmenting Data Based on Target-Guide Alignment Results
[0425] During the above training data derivation process, data is augmented based on the alignment results between the target nucleic acid and guide nucleic acid of each raw data point. When aligning the target nucleic acid sequence and guide nucleic acid sequence of a raw data point, if there are two or more possible target-guide alignments, and two or more alignment patterns satisfy a predetermined criterion, each result is used as training data. The predetermined criterion may be, for example, "selecting the pattern with the highest alignment score." In this case, if there are two or more alignment patterns with the highest alignment scores, each result is used for model training. Depending on the aligned target-guide sequence information, additional information referencing it may vary, and each additional information is also included in the training data point. In other words, a single raw data point can be augmented with two or more training data points. In other words, when matching a raw data point with a training data point derived from it, a single raw data point can correspond to two or more training data points. The training data point includes the following information: aligned target nucleic acid sequence; aligned guide nucleic acid; mismatch information; and additional information.
[0426] Training a model to predict the cleavage activity of the CRISPR / Cas system (S2300)
[0427] The method for training the cleavage activity prediction model of the CRISPR / Cas system described above includes a process of training the cleavage activity prediction model of the CRISPR / Cas system using the training data described above. At this time, the cleavage activity prediction model is trained by labeling each training data point included in the training data with the target-guided activity of the corresponding raw data point as an output value. The specific training method may vary depending on the structure of the cleavage activity prediction model described above, and known techniques may be appropriately utilized.
[0428] Method for predicting target nucleic acid cleavage activity using an activity prediction model
[0429] generalization
[0430] This specification discloses a method for predicting target nucleic acid cleavage activity using a cleavage activity prediction model of a CRISPR / Cas system. The above cleavage activity prediction method is a method for predicting target-guide activity when a target nucleic acid and a guide nucleic acid are given, assuming the use of a specific Cas protein.
[0431] The above target nucleic acid cleavage activity prediction method includes: obtaining target nucleic acid sequence and guide nucleic acid sequence information (S2400); aligning the target nucleic acid sequence and guide nucleic acid sequence to select a target-guide alignment pattern that satisfies a predetermined criterion (S2500); inputting the selected target-guide alignment pattern into an activity prediction model to predict cleavage activity (S2600).
[0432] Here, the details of the subsequent process (S2600) may differ slightly depending on whether there is one or multiple selected target-guide alignment patterns.
[0433] (If there is only one selected target-guide alignment pattern) Aligned target-guide nucleic acid information and necessary additional input variables are derived (S2610); this is input into a cleavage activity prediction model to obtain a result value (S2620); and the above result value is used as a prediction result (S2630).
[0434] (In case there are two or more selected target-guide alignment patterns) Target-guide nucleic acid information aligned for each pattern and necessary additional input variables are derived (S2640); each input variable is independently input into a cleavage activity prediction model to obtain multiple result values (S2650); and the multiple predicted values are weighted and averaged to be used as a prediction result (S2660).
[0435] The above method is characterized by treating each piece of information as an independent input variable and inputting it into the model when it is necessary to utilize multiple target-guide alignment pattern information according to the target nucleic acid sequence and guide nucleic acid sequence, and then calculating a weighted average of the result values to derive a single predicted value.
[0436] Multiple input variables can be utilized according to the target-guide alignment pattern.
[0437] A key feature of the above target nucleic acid cleavage activity prediction method is its ability to utilize multiple aligned target-guide sequence information. When aligning target and guide nucleic acid sequences, the method uses each aligned target-guide sequence information as an independent input variable if 1) multiple alignment patterns are possible and 2) the multiple alignment patterns satisfy predetermined criteria. This improves prediction accuracy by considering all possible interactions when the target and guide nucleic acid sequences are mismatched. For example, the predetermined criteria may be the "highest alignment score."
[0438] Weighted average of predicted results
[0439] When utilizing multiple aligned target-guide sequences in the above target nucleic acid cleavage activity prediction method, a weighted average of the results derived from each input variable is used as the final prediction value. The weights used when averaging the results can be determined based on the importance of each target-guide alignment pattern. For example, if all target-guide alignment patterns are equally important, the weights of the results can all be equal. In other words, the arithmetic average of the results can be used as the final prediction value.
[0440] Application of Activity Prediction Models #1 - Predicting Target Nucleic Acid Cleavage Selectivity
[0441] This specification discloses a method for predicting the degree to which one of two target nucleic acids is selectively cleaved, given two target nucleic acids and one guide nucleic acid. Hereinafter, the target nucleic acid to be cleaved is referred to as a "target nucleic acid," and the target nucleic acid not to be cleaved is referred to as a "noise nucleic acid." The method includes: obtaining a target nucleic acid sequence, a noise nucleic acid sequence, and a guide nucleic acid sequence (S2700); obtaining a predicted value of the cleavage activity of the guide nucleic acid toward the target nucleic acid (target cleavage activity) (S2800); obtaining the cleavage activity of the guide nucleic acid toward the noise nucleic acid (noise cleavage activity) (S2900); and determining the target nucleic acid cleavage selectivity based on the target cleavage activity and the noise cleavage activity (S3000). At this time, each predicted value of the cleavage activity is obtained by the method disclosed in [Method for Predicting Target Nucleic Acid Cleavage Activity Using an Activity Prediction Model]. Herein, the target nucleic acid cleavage selectivity refers to an index for determining how well the target nucleic acid is cleaved compared to the degree to which the noise nucleic acid is cleaved. The above cleavage selectivity is also referred to as the difference between the target nucleic acid cleavage activity and the noise nucleic acid cleavage activity.
[0442] Utilization of the Active Prediction Model #2 - A Guide Nucleic Acid Selection Method that Selectively Cleaves Only One of Two Target Nucleic Acids
[0443] generalization
[0444] This specification discloses a method for selecting a guide nucleic acid that selectively cleaves one target nucleic acid over the other, given two target nucleic acids. Hereinafter, the target nucleic acid to be cleaved is referred to as a "target nucleic acid," and the target nucleic acid not to be cleaved is referred to as a "noise nucleic acid." The method includes: obtaining a target nucleic acid sequence and a noise nucleic acid sequence (S3100); deriving guide nucleic acid candidates (S3200); determining target nucleic acid cleavage selectivity for each guide nucleic acid candidate (S3300); and selecting an optimized guide nucleic acid according to a predetermined criterion (S3400). In this case, the guide nucleic acid candidates include a sequence that is perfectly matched to the target nucleic acid, and a sequence that further has one or more mismatches and / or bulges, and the method is characterized by including various mismatch sequences in the guide nucleic acid candidates for exploration. Here, the process of predicting target nucleic acid selectivity for each guide nucleic acid candidate follows the method disclosed in the paragraph [Utilization of the activity prediction model #1 - Method for predicting target nucleic acid cleavage selectivity].
[0445] Deriving guide nucleic acid candidates (S3200)
[0446] Given the two target nucleic acids mentioned above, a method for selecting a guide nucleic acid that selectively cleaves one target nucleic acid over the other involves deriving candidate guide nucleic acids. This method is characterized by searching for suitable guide nucleic acids by broadly encompassing sequences that are either completely identical to the target sequence or mismatched. For example, the sequences of the candidate guide nucleic acids include:
[0447] 1) A sequence that perfectly matches (or corresponds) to the above target nucleic acid;
[0448] 2) A sequence in which one base in the sequence of 1) is substituted with another base;
[0449] 3) A sequence in which two bases in the sequence of 1) are substituted with different bases;
[0450] 4) A sequence in which one random base is inserted at a random position in the sequence of 1);
[0451] 5) A sequence in which one random base is deleted at any position in the sequence of 1); and
[0452] 6) Any combination of 2) to 5).
[0453] Determining target nucleic acid cleavage selectivity for each guide nucleic acid candidate (S3300)
[0454] Given the two target nucleic acids described above, a method for selecting a guide nucleic acid that selectively cleaves one target nucleic acid over the other involves predicting the target nucleic acid cleavage selectivity for each guide nucleic acid candidate. This process is performed using the target nucleic acid sequence, the noise nucleic acid sequence, and each guide nucleic acid candidate sequence, using the method disclosed in [Utilization of an Activity Prediction Model #1 - Method for Predicting Target Nucleic Acid Cleavage Selectivity]. As a result of the above process, a predicted target nucleic acid cleavage selectivity is obtained for each guide nucleic acid candidate.
[0455] Selecting the optimized guide nucleic acid (S3400)
[0456] Given the two target nucleic acids described above, a method for selecting a guide nucleic acid that selectively cleaves one target nucleic acid over the other involves selecting the guide nucleic acids based on predetermined criteria. Through the above process, one or more guide nucleic acids are selected, and their sequences are output.
[0457] Advantages of a guide nucleic acid selection method that selectively cleaves only one of two target nucleic acids
[0458] Through the above method, guide nucleic acids that selectively cleave only one nucleic acid (the target nucleic acid) among two nucleic acids (the target nucleic acid and the noise nucleic acid) with very similar sequences can be efficiently derived. Furthermore, the inventors of this application experimentally demonstrated that the guide nucleic acid selected by the above model indeed highly selectively cleaves only the target nucleic acid (see Experimental Examples). Because the CRISPR / Cas system has off-target cleavage activity, it is very difficult to find guide nucleic acids that selectively cleave only the target nucleic acid when the sequences of the two nucleic acids to be distinguished are very similar (e.g., differ by only one base). Previously, 1) there was no machine learning model that measured "cleavage activity" and 2) there was no established method for high-throughput confirmation of cleavage activity. Therefore, to find such guide nucleic acids, the only way was to manually check one or two combinations of target-guide reactions at a time in a low-throughput manner. This is a very time-consuming, laborious, and expensive task.
[0459] The method disclosed in this specification enables the efficient and low-cost identification of guide nucleic acids possessing selective cleavage activity. This leads to the application inventions described in Chapter 3 below.
[0460] Chapter 3. Application of the Invention - Method for Concentrating Rare Nucleic Acid Samples
[0461] Overview of methods for enriching rare nucleic acid samples
[0462] This specification discloses a method for enriching rare nucleic acids contained in a sample. The rare nucleic acid enrichment method is a method for selectively enriching only rare nucleic acids when 1) a small amount of a nucleic acid to be enriched (referred to as a rare nucleic acid) and a relatively large amount of another nucleic acid (referred to as a background nucleic acid) are mixed in the original sample, and 2) the rare nucleic acid and the background nucleic acid have similar sequences. According to the method, a target to be selectively cleaved between the rare nucleic acid and the background nucleic acid is determined; based on the determined content, the "guide nucleic acid selection method for selectively cleaving only one of two target nucleic acids" of this specification is performed to determine one or more optimized guide nucleic acids; and using the optimized guide nucleic acid(s), the rare nucleic acid is enriched by an appropriate method depending on the nucleic acid to be cleaved. At this time, the method for enriching rare nucleic acids using the optimized guide nucleic acid varies depending on whether the target of cleavage is the background nucleic acid or the rare nucleic acid.
[0463] When the target of cutting is a background nucleic acid, the rare nucleic acid is concentrated by referring to the method disclosed in Korean Publication No. 2015-0138074 A.
[0464] If the target of cutting is a rare nucleic acid, the rare nucleic acid is concentrated by referring to the method disclosed in U.S. Publication No. 2019-0300935 A1.
[0465] More specific details are explained in the paragraphs below.
[0466] Rare Nucleic Acid Sample Enrichment Method #1 - Enrichment by Cutting Out Background Nucleic Acids
[0467] In one embodiment, the method for enriching a rare nucleic acid sample can be performed by cleaving and removing background nucleic acids from a raw sample. Specifically, the process is performed by processing a CRISPR / Cas complex containing guide nucleic acid(s) optimized for the raw sample (selectively cleaving background nucleic acids) and selectively amplifying and enriching uncleaved rare nucleic acids. The basic idea and implementation of the enrichment method are disclosed in a prior document (KR 2015-0138074 A). Furthermore, the inventors of this application confirmed that even when there are multiple types of rare nucleic acids to be enriched and corresponding background nucleic acids in a sample, enrichment is also possible in a multiplexed manner by simultaneously processing optimized guide nucleic acids for each rare-background combination.
[0468] Rare Nucleic Acid Sample Enrichment Method #2 - Enriching Rare Nucleic Acids by Cleaving
[0469] In one embodiment, the method for enriching a rare nucleic acid sample can be performed by cleaving rare nucleic acids from a raw sample and removing uncleaved background nucleic acids. Specifically, the process is performed by: blocking the terminal portions of nucleic acids contained in the raw sample; processing a CRISPR / Cas complex containing the optimized guide nucleic acid(s) (wherein the rare nucleic acids are selectively cleaved); processing an adapter capable of linking to the nucleic acid cleavage site; and selectively amplifying and enriching fragments of the rare nucleic acids linked to the adapter. The basic idea and implementation of the enrichment method are disclosed in prior art (US 2019-0300935 A1). Furthermore, the inventors of this application confirmed that even when a sample contains multiple types of rare nucleic acids to be enriched and corresponding background nucleic acids, enrichment is possible in a multiplexed manner by simultaneously processing optimized guide nucleic acids for each rare-background combination.
[0470] Application of Rare Nucleic Acid Sample Enrichment Methods - Circulating Tumor Nucleic Acid Enrichment
[0471] This rare nucleic acid sample enrichment method can be used to enrich circulating tumor nucleic acids (CTNAs), a cancer marker present in patient samples. CTNAs include ctDNA and ctRNA.
[0472] Circulating tumor nucleic acids (CTUs) containing cancer-related variants (CRVs) are often very similar in sequence to circulating free nucleic acids (CFNAs) with corresponding wild-type sequences. Similarly, CFNAs include cfDNA and cfRNA. For example, CTU sequences may be sequences in which a single nucleotide variant (SNV) has occurred in the wild-type sequence. Furthermore, when samples, such as blood, are collected from patients, CTUNAs are often present in very small amounts. Therefore, direct detection of CTUNAs from samples is difficult and requires a concentration process.
[0473] To enrich circulating tumor nucleic acids (CTUs) in patient samples, the above rare nucleic acid sample enrichment method can be used. In this case, the rare nucleic acid corresponds to CTUs containing cancer-related variants, and the background nucleic acid corresponds to CTUs containing the wild-type sequence of the cancer-related variant portion.
[0474] This specification discloses optimized guide nucleic acid sequences that can be used to enrich known circulating tumor nucleic acids. Depending on the implementation of the method for enriching rare nucleic acid samples, the optimized guide nucleic acids are divided into: 1) optimized guide nucleic acids that selectively cleave cancer-associated variant nucleic acid sequences; or 2) optimized guide nucleic acids that selectively cleave wild-type sequences corresponding to cancer-associated variants.
[0475] More specific implementation examples are disclosed in the [Possible Embodiments of the Invention] section.
[0476]
[0477] [Possible embodiments of the invention]
[0478] Cas protein
[0479] Example 1, Cas protein
[0480] A Cas protein having nucleic acid cleavage activity.
[0481] Example 2, Cas type limitation
[0482] In Example 1, the Cas protein is selected from the following:
[0483] Cas12a (Cpf1); Cas12b1 (C2c1); Cas12c (C2c3); Cas12e (CasX); Cas12d (CasY); Cas12g; Cas12h; Cas12i; Cas1; Cas1B; Cas2; Cas3; Cas4; Cas5; Cas6; Cas7; Cas8; Cas9 (also known as Csn1 and Csx12); Cas10; Csy1; Csy2; Csy3; Cse1; Cse2; Csc1; Csc2; Csa5; Csn2; Csm2; Csm3; Csm4; Csm5; Csm6; Cmr1; Cmr3; Cmr4; Cmr5; Cmr6; Csb1; Csb2; Csb3; Csx17; Csx14; Csx10; Csx16; CsaX; Csx3; Csx1; Csx15; Csf1; Csf2; Csf3; Csf4; Cas13a (C2c2); Cas13b; Cas13c; Cas13d; Cas14; xCas9; circular permutation Cas9; Argonaute (Ago) domain; variants of the above Cas proteins; And artificial Cas proteins developed with artificial intelligence or related technologies (see, e.g., Jeffrey et. al. (Design of highly functional genome editors by modeling the universe of CRISPR-Cas sequences, Jeffrey . Ruffolo, Stephen Nayfach, Joseph Gallagher, Aadyot Bhatnagar, Joel Beazer, Riffat Hussain, Jordan Russ, Jennifer Yip, Emily Hill, Martin Pacesa, Alexander . Meeske, Peter Cameron, Ali Madani, bioRxiv 2024.04.22.590591; doi: https: / doi.org / 10.1101 / 2024.04.22.590591).
[0484] Example 3, Cas9 type limitation
[0485] In Example 2, the Cas protein is a Cas9 protein, and the Cas9 protein is derived from a microorganism selected from the following, or a variant thereof:
[0486] Streptococcus pyogenes; Streptococcus thermophilus; Streptococcus sp.; Staphylococcus aureus; Campylobacter jejuni; Nocardiopsis dassonvillei; Streptomyces pristinaespiralis; Streptomyces viridochromogenes; Streptosporangium roseum; AlicyclobacHlus acidocaldarius; Bacillus pseudomycoides; Bacillus selenitireducens; Exiguobacterium sibiricum; Lactobacillus delbrueckii; Lactobacillus salivarius; Microscilla marina; Burkholderiales bacterium; Polaromonas naphthalenivorans; Polaromonas sp.; Crocosphaera watsonii; Cyanothece sp.; Microcystis aeruginosa; Synechococcus sp.); Acetohalobium arabaticum; Ammonifex degensii; Caldicelulosiruptor bescii; Candidatus Desulforudis; Clostridium botulinum; Clostridium difficile; Finegoldia magna; Natranaerobius thermophilus; Pelotomaculum thermopropionicum; Acidithiobacillus caldus; Acidithiobacillus ferrooxidans; Allochromatium vinosum; Marinobacter sp.; Nitrosococcus halophilus; Nitrosococcus watsoni; Pseudoalteromonas haloplanktis; Ktedonobacter racemifer; Methanohalobium evestigatum; Anabaena variabilis; Nodularia spumigena; Nostoc sp.; Arthrospira maxima; Arthrospira platensis; Arthrospira sp.; Lyngbya sp.; Microcoleus chthonoplastes; Oscillatoria sp.); Petrotoga mobilis; Thermosipho africanus; and Acaryochloris marina.
[0487] Example 4, SpCas9 type limitation
[0488] In Example 3, the Cas9 protein is a Streptococcus pyogenes-derived Cas9 protein (SpCas9 protein) or a variant thereof.
[0489] Example 5, SpCas9 variant limitation
[0490] In Example 4, the Cas9 protein is a mutant of an SpCas9 protein selected from the following:
[0491] SpCas9-HF1; SpCas9-NRRH-HF1; SpCas9-NRCH-HF1; SuperFi-Cas9; and evoCas9.
[0492] Example 6, Sequence limitation
[0493] In Example 4, the Cas9 protein comprises an amino acid sequence selected from the following:
[0494] Sequence number 1 to sequence number 6.
[0495] Example 7, coding sequence limitation
[0496] In Example 4, the Cas9 protein is encoded by a nucleic acid sequence selected from the following:
[0497] Sequence number 7 to sequence number 12.
[0498] Guide nucleic acid
[0499] Example 8, Guide Nucleic Acid
[0500] A guide nucleic acid that binds to any one of the Cas proteins of Examples 1 to 7 to form a CRISPR / Cas complex.
[0501] Example 9, including scaffold and guide domains
[0502] In any one of Example 8, the guide nucleic acid comprises:
[0503] a scaffold, wherein said scaffold is capable of interacting with a Cas protein to form a complex; and
[0504] A guide domain, wherein the guide domain is capable of guiding a CRISPR / Cas complex to a target nucleic acid.
[0505] Example 10, Scaffold Sequence Limitation
[0506] In any one of Examples 8 to 9, the scaffold of the guide nucleic acid comprises a sequence selected from the following:
[0507] A sequence selected from among GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGctagtccgttatcaacttgaaaaagtggcaccgagtcggtgc (SEQ ID NO: 13), GTTTCAGAGCTATGCTGGAAACAGCATAGCAAGTTGAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTT (SEQ ID NO: 14), GTTTAAGAGCTATGCTGGAAACAGCATAGCAAGTTTAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTTT (SEQ ID NO: 15), and GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAAAGTGGCACCGAGTCGGTGCTTTTTTT (SEQ ID NO: 16); or
[0508] A sequence that is 80% or more, 81% or more, 82% or more, 83% or more, 84% or more, 85% or more, 86% or more, 87% or more, 88% or more, 89% or more, 90% or more, 91% or more, 92% or more, 93% or more, 94% or more, 95% or more, 96% or more, 97% or more, 98% or more, or 99% or more identical (homologous) to a sequence selected from SEQ ID NOs: 13 to 16.
[0509] Example 11, Relationship with target nucleic acid
[0510] In any one of Examples 8 to 10, the guide domain of the guide nucleic acid is in a relationship selected from the following with the target nucleic acid:
[0511] (1) The guide domain binds completely complementarily to all or part of the target nucleic acid;
[0512] (2) The guide domain complementarily binds to all or part of the target nucleic acid, but one or more base pairs are not a complementary combination, or a mismatch occurs;
[0513] (3) The guide domain complementarily binds to all or part of the target nucleic acid, but one or more guide domain bases do not participate in binding, or a bulge occurs in the guide domain;
[0514] (4) The guide domain complementarily binds to all or part of the target nucleic acid, but one or more target nucleic acid bases do not participate in binding, or a bulge occurs in the target nucleic acid;
[0515] (5) any combination of (1) to (4) above; or
[0516] (6) The above guide domain does not complementarily bind to all or part of the target nucleic acid.
[0517] Example 12, length of the guide domain
[0518] In any one of Examples 8 to 11, the length of the guide domain is:
[0519] 12nt, 13nt, 14nt, 15nt, 16nt, 17nt, 18nt, 19nt, 20nt, 21nt, 22nt, 23nt, 24nt, 25nt, 26nt, 27nt, 28nt, 29nt, or 30nt; or
[0520] A length within the numerical range consisting of the two numbers mentioned above (e.g., 18 nt to 22 nt).
[0521] Example 13, may include G or GG
[0522] In any one of Examples 8 to 12, the guide nucleic acid additionally includes a G or GG base at the 5' end or 3' end of the guide domain;
[0523] For example, if the guide nucleic acid is a guide RNA that binds to the Cas9 protein to form a CRISPR / Cas complex,
[0524] The above guide RNA may additionally include a G or GG base at the 5' end of the guide domain.
[0525] Example 14, Guide DNA
[0526] In any one of Examples 8 to 13, the guide nucleic acid is guide DNA or guide RNA.
[0527] Target nucleic acid
[0528] Example 15, target nucleic acid
[0529] A target nucleic acid capable of being cleaved by a CRISPR / Cas complex formed by combining the Cas protein of any one of Examples 1 to 7 and the guide nucleic acid of any one of Examples 8 to 14.
[0530] Example 16, Partial differentiation of target nucleic acid
[0531] In Example 15, the target nucleic acid comprises a portion that can be recognized by a Cas protein; and a portion that can interact with the guide domain of the guide nucleic acid.
[0532] Example 17, target nucleic acid type limitation
[0533] In Example 15, the target nucleic acid is double-stranded DNA, single-stranded DNA, double-stranded RNA, or single-stranded RNA.
[0534] Example 18, double-stranded DNA limitation
[0535] In Example 17, the target nucleic acid is double-stranded DNA,
[0536] The target nucleic acid comprises a target strand and a non-target strand,
[0537] The target strand comprises a target portion capable of interacting with the guide domain of the guide nucleic acid,
[0538] The above non-target strand comprises a portion that can be recognized by the Cas protein.
[0539] Example 19, target nucleic acid of dsDNA cleavage Cas
[0540] In Example 18, the portion of the non-target strand of the target nucleic acid that can be recognized by the nucleic acid-guided nuclease is a Protospacer Adjacent Motif (PAM) or a Protospacer Flanking Sequence (PFS).
[0541] The non-target strand comprises a protospacer adjacent to the PAM or PFS,
[0542] The protospacer of the non-target strand is a portion that complementarily binds to the target portion of the target strand.
[0543] Example 20, Relationship between target nucleic acid sequence and guide nucleic acid sequence
[0544] In any one of Examples 18 to 19, the sequence of the protospacer of the target nucleic acid and the sequence of the guide domain of the guide nucleic acid are in a relationship selected from the following:
[0545] (1) When the sequence of the protospacer and the sequence of the guide domain are aligned, they are perfectly matched or are perfectly homologous;
[0546] (2) When the sequence of the protospacer and the sequence of the guide domain are aligned, one or more bases are mismatched;
[0547] (3) When the sequence of the protospacer and the sequence of the guide domain are aligned, one or more bulges occur in the protospacer;
[0548] (4) When the sequence of the protospacer and the sequence of the guide domain are aligned, one or more bulges occur in the guide domain;
[0549] (5) any combination of (1) to (4) above; or
[0550] (6) The sequence of the above protospacer and the sequence of the above guide domain are completely inconsistent.
[0551] Example 21, single-stranded RNA limitation
[0552] In Example 17, the target nucleic acid is single-stranded RNA,
[0553] The target nucleic acid comprises a portion that can be recognized by the Cas protein and a target portion that can interact with the guide domain of the guide nucleic acid.
[0554] Example 22, PAM, PFS only
[0555] In Example 21, the portion of the target nucleic acid that can be recognized by the Cas protein is a protospacer adjacent motif (PAM) or a protospacer flanking sequence (PFS).
[0556] Example 23, Relationship between target nucleic acid sequence and guide nucleic acid sequence
[0557] In any one of Examples 21 to 22, the sequence of the target portion of the target nucleic acid and the sequence of the guide domain of the guide nucleic acid are in a relationship selected from the following:
[0558] (1) When the sequence of the target portion and the sequence of the guide domain are aligned, they are perfectly complementary;
[0559] (2) When the sequence of the target portion and the sequence of the guide domain are aligned, one or more base pairs are not complementary;
[0560] (3) When the sequence of the target portion and the sequence of the guide domain are aligned, one or more bulges occur in the target portion;
[0561] (4) When the sequence of the target portion and the sequence of the guide domain are aligned, one or more bulges occur in the guide domain; or
[0562] (5) Any combination of (1) to (4) above.
[0563] (6) The sequence of the above protospacer and the sequence of the above guide domain are not perfectly complementary.
[0564] Example 24, PAM sequence limitation
[0565] In any one of Examples 15 to 23, when the target nucleic acid comprises a PAM, the PAM is NGG, NAG, NGA, NGT, NGC, NAA, AAAG, AACG, AAGG, AATG, CAAG, CACG, CAGG, CATG, GAAG, GACG, GAGG, GATG, TAAG, TACG, TAGG, TATG, ACAG, ACCG, ACGG, CCGG, GCAG, GCCG, GCGG, TCGG, AGAG, AGCG, AGGG, AGTG, CGAG, CGCG, CGGG, CGTG, GGAG, GGCG, GGGG, GGTG, TGAG, TGCG, TGGG, TGTG, ATAG, ATCG, ATGG, ATTG, CTAG, CTCG, CTGG, CTTG, GTAG, GTCG, GTGG, GTTG, TTGG, GGAT, Selected from the group consisting of GGCT, GGGT, GGTT, TGCT, AGAA, AGCA, AGGA, AGTA, CGAA, CGCA, CGGA, CGTA, GGAA, GGCA, GGGA, GGTA, TGAA, TGCA, TGGA, TGTA, AGAC, AGCC, AGGC, AGTC, CGAC, CGCC, CGGC, CGTC, GGAC, GGCC, GGGC, GGTC, TGAC, TGCC, TGGC, TGTC and ATGA, wherein N is a base selected from A, T, G or C.
[0566] Target-guide vector
[0567] Example 25, target-guide vector
[0568] Target-guide vectors including:
[0569] A guide nucleic acid of any one of Examples 8 to 14, or a nucleic acid encoding the guide nucleic acid; and
[0570] The target nucleic acid of any one of Examples 15 to 24, or a nucleic acid encoding the target nucleic acid.
[0571] Example 26, Promoter Addition
[0572] In Example 25, the target-guide vector further comprises one or more promoters,
[0573] The above target-guide vector comprises a selected combination of:
[0574] A nucleic acid encoding a guide nucleic acid and a promoter operably linked thereto;
[0575] A nucleic acid encoding a target nucleic acid, and a promoter operably linked thereto; or
[0576] Both combinations above.
[0577] Example 27, including barcodes and UMIs
[0578] In any one of Examples 25 to 26, the target-guide vector further comprises a selected configuration from among:
[0579] A barcode, wherein the barcode encodes target nucleic acid sequence and guide nucleic acid sequence information;
[0580] A Unique Molecular Identifier (UMI), wherein the UMI encodes unique identification information of the target-guide vector; or
[0581] All of the above configurations.
[0582] Example 28, including restriction enzyme cleavage sites
[0583] In any one of Examples 25 to 27, the target-guide vector further comprises one or more restriction enzyme cleavage sites.
[0584] Example 29, restriction enzyme limitation
[0585] In Example 28, the restriction enzyme cleavage site can be cleaved by one or more restriction enzymes selected from the following:
[0586] EcoRI; EcoRV; SmaI; HpaI; PvuII; DraI; AluI; HaeIII; BsaAI; FspI; HincII; MscI; NaeI; NruI; PmlI; SfoI; BamHI; HindIII; NotI; and XhoI.
[0587] Example 30, including antibiotic resistance genes
[0588] In any one of Examples 25 to 29, the target-guide vector further comprises one or more antibiotic resistance genes.
[0589] Example 31, Example of an antibiotic resistance gene
[0590] In Example 30, the antibiotic resistance gene is at least one selected from the following:
[0591] AmpR (Ampicillin Resistance); KanR (Kanamycin Resistance); TetR (Tetracycline Resistance); CmR (Chloramphenicol Resistance); StrR (Streptomycin Resistance); SpecR (Spectinomycin Resistance); Bla (Beta-lactamase); NeoR (Neomycin Resistance); ZeoR (Zeocin Resistance); and HygR (Hygromycin Resistance).
[0592] Example 32, Cut-seq1 structure limitation
[0593] In any one of Examples 25 to 31, the target-guide vector comprises the following structure:
[0594] [Target] - [Identifier] - [Promoter] - [Guide] - [RS1]
[0595] Here, the Target is a target DNA;
[0596] The above identifier is 1) a barcode encoding target-guide sequence information, 2) a Unique Molecular Identifier (UMI) encoding unique identification information of the target-guide vector, or 3) an identifier including both 1) and 2);
[0597] The above Guide is a DNA encoding the above guide RNA;
[0598] The above promoter is operably linked to DNA encoding the guide RNA; and
[0599] The above RS1 is the first restriction enzyme cleavage site.
[0600] Example 33, Cut-seq2 structure limitation
[0601] In any one of Examples 25 to 32, the target-guide vector comprises the following structure:
[0602] [RS2] - [Constant] - [Target] - [Identifier] - [Promoter] - [Guide] - [RS1]
[0603] Here, the Target is a target DNA;
[0604] The above identifier is 1) a barcode encoding target-guide sequence information, 2) a Unique Molecular Identifier (UMI) encoding unique identification information of the target-guide vector, or 3) an identifier including both 1) and 2);
[0605] The above Guide is a DNA encoding the above guide RNA;
[0606] The above promoter is operably linked to DNA encoding the guide RNA;
[0607] The above RS1 is the first restriction enzyme cleavage site;
[0608] The above Constant is an invariant sequence; and
[0609] The above RS2 is the second restriction enzyme site.
[0610] Example 34, Promoter Only
[0611] In any one of Examples 25 to 33,
[0612] If the target-guide vector comprises a promoter, the promoter is selected from the following:
[0613] CMV; Efla; CAG; PGK; TRE; U6; T7; lacUV5; gapA; T5; recA; Ptac; Patac; pAl; lac; Sp6; araBad; trp; and their fusion promoters.
[0614] Example 35, restriction enzyme limitation
[0615] In any one of Examples 25 to 34,
[0616] When the target-guide vector comprises a restriction enzyme cleavage site, a first restriction enzyme cleavage site, or a second restriction enzyme cleavage site, the restriction enzyme cleavage sites can be cleaved by a restriction enzyme selected from the following:
[0617] AccI; ACC65; AhaII; AsuII; Asp718; AvaI; AvrII; BamH1; BelI; BglIII; BsaNI; BspMII; BssHII; BspMI; BstEII; Bsu90I; DdeI; Eco0109; EcoR1; EcoRII; HindIII; HinPI; HinfI; HpaII; MaeI; MaeII; MaeIII; MboI; MluI; NarI; NcoI; NdeI; NdeII; NheI; Not1; PpuMI; RsrII; Sal1; SauI; Sau3AI; Sau961; ScrFI; SpeI; StyI; TagI; TthMI; XbaI; XhoII; XmaI; AarI; Ascl; BbrCI; CspI; DraI; FseI; NotI; NruI; PacI; PmeI; PvuI; SapI; SdaI; SplI; SwaI; AluI; BalI; BfrBI; BsaAI; BsaBI; BsrBI; BtrI; Cac8I; CdiI; CviJI; CviRI; Eco47III; Eco78I; EcoICRI; EcoRV; FnuDII; FspAI; HaeI; HaeIII; Hpy8I; LpnI; MlyI; MslI; MstI; NaeI; NalIV; NruI; NspBII; OliI; PmaCI; PmeI; PshAI; PsiI; PvuII; RsaI; ScaI; SmaI; SnaBI; SrfI; SspI; SspD5I; StuI; XcaI; XmnI; and ZraI.
[0618] Target-Guide Library
[0619] Example 36, Target-Guided Library
[0620] A target-guide library containing multiple target-guide vectors,
[0621] Here, the target-guide vector is any one selected from Examples 25 to 35,
[0622] The above target-guide library contains two types of target-guide vectors:
[0623] The two different types of target-guide vectors differ in the sequence of the target nucleic acid contained therein, the sequence of the guide nucleic acid, or both.
[0624] Example 37, target-guided library, n expression
[0625] A target-guide library containing n types of target-guide vectors,
[0626] Here, n is an integer greater than or equal to 2,
[0627] When any integer i, j satisfies 1 ≤ i < j ≤ n,
[0628] The i-th type of target-guide vector and the j-th type of target-guide vector are each independently selected target-guide vectors from among Examples 25 to 35,
[0629] The sequence of the target nucleic acid included in the ith type of target-guide vector is called the ith target sequence,
[0630] The sequence of the guide nucleic acid included in the ith type of target-guide vector is called the ith guide sequence,
[0631] The sequence of the target nucleic acid included in the jth type of target-guide vector is called the jth target sequence,
[0632] If the sequence of the guide nucleic acid included in the jth type of target-guide vector is called the jth guide sequence,
[0633] Satisfies the following:
[0634] The ith target sequence is different from the jth target sequence;
[0635] The i-th guide sequence is different from the j-th guide sequence; or
[0636] Both of the above.
[0637] Example 38, structural limitation
[0638] In any one of Examples 36 to 37, all target-guide vectors included in the target-guide library have the same structure.
[0639] Example 39, Cut-seq1 library
[0640] In Example 38, all target-guide vectors included in the target-guide library are target-guide vectors according to Example 32.
[0641] Example 40, Cut-seq2 library
[0642] In Example 38, all target-guide vectors included in the target-guide library are target-guide vectors according to Example 33.
[0643] microchamber cells
[0644] Example 41, Microchamber Cells
[0645] Microchamber cells containing:
[0646] A guide nucleic acid of any one of Examples 8 to 14, or a nucleic acid encoding the guide nucleic acid; and
[0647] The target nucleic acid of any one of Examples 15 to 24, or a nucleic acid encoding the target nucleic acid,
[0648] Here, the guide nucleic acid and the target nucleic acid are exogenous components,
[0649] The microchamber cell comprises only one type of guide nucleic acid or a nucleic acid encoding the guide nucleic acid,
[0650] The microchamber cell comprises only one type of target nucleic acid or a nucleic acid encoding the target nucleic acid,
[0651] Here, the internal environment of the microchamber cell is configured to simulate an in vitro environment.
[0652] Example 42, Additional Components - Nucleic Acid-Guided Nucleases
[0653] In Example 41, the microchamber cell further comprises a nucleic acid-guided nuclease of any one of Examples 1 to 7,
[0654] The above nucleic acid-guided nuclease is also an exogenous construct.
[0655] Example 43, Additional Components - Various Enzymes
[0656] In any one of Examples 41 to 42, the microchamber cells further comprise one or more transcription enzymes (RNA polymerases), one or more restriction enzymes, or any combination of the two.
[0657] Example 44, shaping of components
[0658] In any one of Examples 41 to 43, the microchamber cell comprises a component selected from the following:
[0659] The guide nucleic acid of any one of Examples 8 to 14, and the target nucleic acid of any one of Examples 15 to 24, wherein the guide nucleic acid and the target nucleic acid are each independent nucleic acids;
[0660] A CRISPR / Cas complex comprising the Cas protein of any one of Examples 1 to 7 and the guide nucleic acid of any one of Examples 8 to 14, and the target nucleic acid or a fragment of the target nucleic acid of any one of Examples 15 to 24;
[0661] The target-guide vector of any one of Examples 25 to 35, or an intercept of the target-guide vector; or
[0662] The target-guide vector of any one of Examples 25 to 35 or a fragment of the target-guide vector, the transcription enzyme, and the guide nucleic acid or target nucleic acid transcribed from the transcription enzyme.
[0663] Example 45: The significance of in vitro environmental simulation
[0664] In any one of Examples 41 to 44, "the internal environment of the microchamber cell is configured to simulate an in vitro environment" is a state selected from the following:
[0665] The vital activity of the above microchamber cells ceases, and the above exogenous components remain active;
[0666] The protein activity of the above microchamber cells is inhibited, and the exogenous components remain active;
[0667] The endogenous enzyme activity of the microchamber cells is inhibited, and the exogenous components remain active;
[0668] The nucleic acid repair mechanism of the above microchamber cells is inhibited, and the exogenous components remain active;
[0669] During the activity of the above microchamber cells, the activity affecting the activity of the CRISPR / Cas system and the nucleic acid cleavage activity is suppressed, and the exogenous components remain active; or
[0670] Any combination of the above conditions;
[0671] Here, the words stopped or suppressed may be replaced by the following verbs:
[0672] Inhibited; blocked; halted; silenced, muted; restricted; or reduced.
[0673] Example 46, dead cells
[0674] In any one of Examples 41 to 45, the microchamber cells are dead cells.
[0675] Example 47, fixed cells
[0676] In any one of Examples 41 to 46, the microchamber cells are fixed cells.
[0677] Example 48, Knockout of Genes Related to Nucleic Acid Repair
[0678] In any one of Examples 41 to 47, the microchamber cells are cells in which some or all of the genes affecting:
[0679] The process by which CRISPR / Cas complexes are formed within cells;
[0680] The process by which the CRISPR / Cas complex cleaves nucleic acids within a cell;
[0681] The process by which nucleic acids cut by the CRISPR / Cas complex are repaired within a cell;
[0682] The process of repairing nucleic acid fragments in cells; or
[0683] Any combination of the above processes.
[0684] Example 49, exogenous composition
[0685] In any one of Examples 41 to 48, the exogenous components of the microchamber cells each independently originate from any one of the following:
[0686] delivered, introduced, transfected, or administered from outside the cell; or
[0687] Derived, cloned, transcribed, translated, or expressed from a substance delivered from outside the cell.
[0688] Example 50, Cell Type Restriction
[0689] In any one of Examples 41 to 49,
[0690] The above microchamber cells are either prokaryotic or eukaryotic cells.
[0691] Example 51, Prokaryotic cell type limitation
[0692] In Example 50, the microchamber cells are E. coli.
[0693] microchamber cells
[0694] Example 52, cell population
[0695] A cell population containing:
[0696] Two or more types of microchamber cells,
[0697] Here, each microchamber cell is a microchamber cell of any one of Examples 41 to 51,
[0698] Two different types of microchamber cells have different guide nucleic acids or sequences encoding said guide nucleic acids, different target nucleic acids or sequences encoding said target nucleic acids, or both; and
[0699] Other cells,
[0700] Here, the cell population ensures compartmentalized interactions within each of the microchamber cells.
[0701] Example 53, Other Cell-Limited
[0702] In Example 52, the other cells are one or more selected from the following:
[0703] Cells of the same species as the above microchamber cells;
[0704] A cell of the same species as the above microchamber cell, comprising two or more types of guide nucleic acids, two or more types of target nucleic acids, or two or more types of guide nucleic acids and target nucleic acids;
[0705] Cells not included in the above categories.
[0706] Example 54, Other cell states
[0707] In any one of Examples 52 to 53, the other cells are dead or fixed cells.
[0708] Example 55, Microchamber Cell Population
[0709] A microchamber cell population comprising two or more types of microchamber cells,
[0710] Here, each microchamber cell is a microchamber cell of any one of Examples 41 to 51,
[0711] Two different types of microchamber cells have different guide nucleic acids or sequences encoding said guide nucleic acids, different target nucleic acids or sequences encoding said target nucleic acids, or both different.
[0712] Here, the microchamber cell population ensures compartmentalized interactions within each microchamber cell.
[0713] Example 56, Microchamber Cell Composition and Significance of Interactions
[0714] In any one of Examples 52 to 55,
[0715] Each of the above microchamber cells comprises a Cas protein of any one of Examples 1 to 7, a guide nucleic acid of any one of Examples 8 to 14, and a target nucleic acid of any one of Examples 15 to 24,
[0716] Ensuring compartmentalized interactions within each of the above microchamber cells means that the following conditions are satisfied:
[0717] (1) In each microchamber cell, the Cas protein and the guide nucleic acid bind to form a CRISPR / Cas complex, and the CRISPR / Cas complex interacts with the target nucleic acid; and
[0718] (2) The CRISPR / Cas complex in one microchamber cell does not interact with the target nucleic acid in another microchamber cell, or the number is negligible.
[0719] Example 57, n types of microchamber cells
[0720] A cell population containing microchamber cells,
[0721] Here, the variables used in the description below are defined as follows:
[0722] n is an integer greater than or equal to 2;
[0723] x is any integer greater than or equal to 1 and less than or equal to n;
[0724] i, j are arbitrary integers satisfying 1 ≤ i < j ≤ n;
[0725] Here, the cell population comprises at least n types of microchamber cells,
[0726] Each microchamber cell is classified as “xth type”,
[0727] Here, the xth type of microchamber cell satisfies:
[0728] 1) A guide nucleic acid comprising a guide nucleic acid selected from Examples 25 to 35 or a nucleic acid encoding the same;
[0729] 2) A target nucleic acid comprising a target nucleic acid selected from Examples 25 to 35 or a nucleic acid encoding the same;
[0730] 3) The x guide nucleic acid can bind to any one of the Cas proteins of Examples 1 to 7 to form a x CRISPR / Cas complex;
[0731] Here, the i-th type of microchamber cell and the j-th type of microchamber cell satisfy:
[0732] a) The sequence of the i-th guide nucleic acid and the sequence of the j-th guide nucleic acid are different;
[0733] b) the sequence of the i-th target nucleic acid is different from the sequence of the j-th target nucleic acid; or
[0734] c) Satisfies both a) and b);
[0735] Here, each of the above microchamber cells is a microchamber cell of any one of Examples 41 to 51,
[0736] The above cell population ensures that the i CRISPR / Cas complex does not interact with the j target nucleic acid.
[0737] Example 58, including Cas protein
[0738] In Example 57, the x-th type of microchamber cell further comprises a Cas protein selected from Examples 1 to 7.
[0739] A high-throughput method for evaluating the in vitro cleavage activity of the CRISPR / Cas system.
[0740] Example 59, High-throughput evaluation method for in vitro cleavage activity
[0741] A method for high-throughput evaluation of the in vitro cleavage activity of a CRISPR / Cas system, comprising:
[0742] (a) Process of preparing microchamber cells;
[0743] As a result of performing the above process, a cell population of any one of Examples 52 to 58 is produced,
[0744] Here, the microchamber cells included in the above cell population each contain one type of target-guide vector of any one of Examples 25 to 35;
[0745] (b) a process for inducing a compartmentalized response in the cell population prepared in (a);
[0746] As a result of performing the above process,
[0747] In each microchamber cell included in the cell population, the Cas protein of any one of Examples 1 to 7 and the guide nucleic acid of any one of Examples 8 to 14 bind to form a CRISPR / Cas complex, and the CRISPR / Cas complex reacts with the target nucleic acid of any one of Examples 15 to 24,
[0748] The above reaction occurs independently within each microchamber cell, and no cross-reaction occurs between the microchamber cells.
[0749] (c) After the (b) process, a process of analyzing the reaction activity between the guide nucleic acid and the target nucleic acid of the target-guide vector contained in each microchamber cell,
[0750] As a result of performing the above process,
[0751] The CRISPR / Cas complex containing the guide nucleic acid of each of the target-guide vectors is evaluated to determine how well it cleaves the target nucleic acid of the target-guide vector.
[0752] Process of preparing microchamber cells
[0753] Example 60, Process for preparing microchamber cells
[0754] In Example 59, the process (a) includes:
[0755] (a-1) Process of preparing a cell population;
[0756] (a-2) Process of treating the target-guide library to the above cell population,
[0757] wherein the target-guided library is a target-guided library of any one of Examples 36 to 40; and
[0758] (a-3) Process of treating a fixation reagent to a cell group that has completed the above (a-2) process;
[0759] At this time, (a) as a result of performing the process, the above cell population includes microchamber cells and fixed cells.
[0760] Example 61, Process for preparing microchamber cells, including expansion process
[0761] In Example 60, the process (a) includes:
[0762] (a-1) Process of preparing a cell population;
[0763] (a-2) Process of treating the target-guide library to the above cell population,
[0764] Here, the target-guided library is a target-guided library of any one of Examples 36 to 40;
[0765] (a-3) Process of culturing the above cell population,
[0766] As a result of performing the above process, 1) the cell population is expanded, 2) cells containing the target-guide vector are selected from the cell population, or 3) both 1) and 2) occur; and
[0767] (a-4) Process of treating a fixation reagent to a cell population that has completed the above (a-3) process.
[0768] Example 62, Cultivation Process Specification
[0769] In Example 61,
[0770] The above process (a-3) is a process of culturing the cell population in a medium containing a specific antibiotic,
[0771] The target-guide vector included in the above target-guide library includes an antibiotic resistance gene,
[0772] As a result of performing the above process (a-3), the cell population is expanded, and cells containing the target-guide vector are selected from the cell population.
[0773] Example 63, using Cut-seq1, 2 libraries
[0774] In any one of Examples 60 to 62,
[0775] The above target-guided library is a target-guided library selected from any one of Example 39 or Example 40.
[0776] Example 64, Target-Guide Library Processing Process
[0777] In any one of Examples 60 to 63,
[0778] As a result of treating the cell population with the target-guide library, almost all of the cells included in the cell population are either 1) introduced with one target-guide vector or 2) are cells without the target-guide vector introduced.
[0779] Example 65, Conditions to consider when processing target-guided libraries
[0780] In Example 64, when the cell population is a eukaryotic cell, the amount of the target-guided library to be treated is determined by considering the multiplicity of infection (MOI).
[0781] As a result, almost all of the cells in the cell population are either 1) introduced with one targeting-guide vector or 2) not introduced with a targeting-guide vector.
[0782] Example 66, Conditions to consider when processing target-guided libraries
[0783] In Example 64, when the cell population is a prokaryotic cell, the amount of the target-guide library to be processed is determined by considering the coverage of each vector,
[0784] As a result, almost all of the cells in the cell population are either 1) introduced with one targeting-guide vector or 2) not introduced with a targeting-guide vector.
[0785] Process of inducing compartmentalized responses
[0786] Example 67, Process for inducing compartmentalized reactions
[0787] In any one of Examples 59 to 63, the step (b) comprises:
[0788] (b-2) A process of treating a cell population in which the process (a) above has been completed with any one of the Cas proteins of Examples 1 to 7;
[0789] Here, the Cas protein is included in the target-guide vector contained in the microchamber cell, or binds to the guide nucleic acid encoded by the target-guide vector to form a CRISPR / Cas complex.
[0790] Example 68, Process for inducing compartmentalized reactions
[0791] In any one of Examples 59 to 63, the step (b) comprises:
[0792] (b-1) A process of treating one or more enzymes to a cell population that has completed the process (a) above;
[0793] wherein said one or more enzymes comprise one or more restriction enzymes, one or more transcription enzymes, or a combination thereof;
[0794] As a result of performing the above process, usable guide nucleic acids are present within each microchamber cell;
[0795] (b-2) A process of treating a cell population in which the above (a) process has been completed with any one of the Cas proteins of Examples 1 to 7,
[0796] Here, the Cas protein binds to the guide nucleic acid contained in each of the microchamber cells to form a CRISPR / Cas complex;
[0797] Here, the above process (b-1) and the above process (b-2) are performed simultaneously or in any order.
[0798] Example 69, Cut-seq1 related process specification
[0799] In Example 63,
[0800] The target-guide library processed in the above process (a) is the target-guide library of Example 39,
[0801] The above process (b) includes:
[0802] (b-1) A process of treating a cell population that has completed the above (a) process with a restriction enzyme and a transcription enzyme,
[0803] Here, the restriction enzyme is delivered to each microchamber cell to cut the first restriction enzyme cutting site of the target-guide vector contained in the microchamber cell,
[0804] The above transcription enzyme is delivered to each microchamber cell and acts on DNA encoding the guide RNA contained in the target-guide vector contained in the microchamber cell to transcribe the guide RNA; and
[0805] (b-2) A process of treating a cell in which the above (a) process has been completed with any one of the Cas proteins of Examples 1 to 7,
[0806] Here, the Cas protein binds to the guide RNA to form a CRISPR / Cas complex,
[0807] The CRISPR / Cas complex reacts with the target DNA contained in the target-guide vector;
[0808] Here, the above process (b-1) and the above process (b-2) are performed simultaneously or in any order.
[0809] Example 70, Cut-seq2 related process specification
[0810] In Example 63,
[0811] The target-guide library processed in the above process (a) is the target-guide library of Example 40,
[0812] The above process (b) includes:
[0813] (b-1) a process of treating a cell population that has completed the above (a) process with one or more restriction enzymes and a transcription enzyme;
[0814] Here, the one or more restriction enzymes are delivered to each microchamber cell to cut the first restriction enzyme cutting site and the second restriction enzyme cutting site of the target-guide vector contained in the microchamber cell,
[0815] The above transcription enzyme is delivered to each microchamber cell and acts on DNA encoding the guide RNA contained in the target-guide vector contained in the microchamber cell to transcribe the guide RNA; and
[0816] (b-2) A process of treating a cell in which the above (a) process has been completed with any one of the Cas proteins of Examples 1 to 7,
[0817] Here, the Cas protein binds to the guide RNA to form a CRISPR / Cas complex,
[0818] The CRISPR / Cas complex reacts with the target DNA contained in the target-guide vector;
[0819] Here, the order in which the above process (b-1) and the above process (b-2) are performed is irrelevant.
[0820] Example 71, Transcription enzyme related
[0821] In any one of Examples 67 to 70, when a transcription enzyme is used in the process,
[0822] The above transcription enzyme is RNA polymerase, and the RNA polymerase can recognize a promoter selected from among CMV, Efla, CAG, PGK, TRE, U6, T7, lacUV5, gapA, T5, recA, Ptac, Patac, pAl, lac, Sp6, araBad, trp, and fusion promoters thereof and function on a nucleic acid operably linked thereto.
[0823] The process of evaluating activity
[0824] Example 72, Process for Evaluating Activity
[0825] In any one of Examples 59 to 70,
[0826] The above (c) process includes:
[0827] (c-1) (b) A process of collecting nucleic acids from a population of cells at the end of the process;
[0828] (c-2) A process of obtaining sequence information by sequencing the nucleic acids collected in the above process (c-1);
[0829] Here, the sequence information is obtained in units of target-guide vectors;
[0830] (c-3) A process for determining target-guide activity based on the sequence information obtained in the above process (c-2);
[0831] Here, the above target-guided activity quantitatively indicates how well a CRISPR / Cas complex containing a specific guide nucleic acid cleaves a specific target.
[0832] Example 73, Process for evaluating activity, using an adapter
[0833] In any one of Examples 59 to 70,
[0834] The above (c) process includes:
[0835] (c-1) (b) A process of collecting nucleic acids from a population of cells at the end of the process;
[0836] (c-2) A process of processing an adapter into the nucleic acid collected in the above (c-1) process,
[0837] Here, the adapter is linked to a position cut by a restriction enzyme in the above process (b) or cut by a CRISPR / Cas complex;
[0838] (c-3) A process of obtaining sequence information by sequencing the nucleic acid after the above (c-2) process is completed.
[0839] Here, the sequencing is performed including a process of amplifying the nucleic acid to which the adapter is linked,
[0840] The above sequence information is obtained in units of target-guide vectors; and
[0841] (c-4) A process for determining target-guide activity based on the sequence information obtained in the above process (c-3);
[0842] Here, the above target-guided activity quantitatively indicates how well a CRISPR / Cas complex containing a specific guide nucleic acid cleaves a specific target.
[0843] Example 74, Process for Evaluating Activity, Cut-seq1
[0844] In Example 69,
[0845] The above (c) process includes:
[0846] (c-1) (b) A process of collecting nucleic acids from a population of cells at the end of the process;
[0847] (c-2) A process of processing an adapter into the nucleic acid collected in the above (c-1) process,
[0848] Here, the adapter is linked to a position cut by a restriction enzyme in the above process (b) or cut by a CRISPR / Cas complex;
[0849] (c-3) A process of obtaining sequence information by sequencing the nucleic acid after the above (c-2) process is completed.
[0850] Here, the sequencing is performed including a process of amplifying the nucleic acid to which the adapter is linked,
[0851] The above sequence information includes a plurality of data points,
[0852] One data point is information about one target-guide vector in which the target DNA has been cleaved,
[0853] Each data point includes target nucleic acid sequence information and guide nucleic acid sequence information of the corresponding target-guide vector;
[0854] (c-4) A process for determining target-guide activity based on the sequence information obtained in the above process (c-3);
[0855] Here, the above target-guided activity quantitatively indicates how well a CRISPR / Cas complex containing a specific guide nucleic acid cleaves a specific target.
[0856] The above target-guide activity is determined based on the number of data points belonging to each group (i.e., target-guide combination) when the sequence information obtained in (c-3) is classified based on the target nucleic acid sequence and the guide nucleic acid sequence.
[0857] Example 75, Barcode, UMI-based, Cut-seq1
[0858] In Example 69, the identifier of the target-guide vector includes both a barcode and a UMI,
[0859] The above (c) process includes:
[0860] (c-1) (b) A process of collecting nucleic acids from a population of cells at the end of the process;
[0861] (c-2) A process of processing an adapter into the nucleic acid collected in the above (c-1) process,
[0862] Here, the adapter is linked to a position cut by a restriction enzyme in the above process (b) or cut by a CRISPR / Cas complex;
[0863] (c-3) A process of obtaining sequence information by sequencing the nucleic acid after the above (c-2) process is completed.
[0864] Here, the sequencing is performed including a process of amplifying the nucleic acid to which the adapter is linked,
[0865] The above sequence information includes a plurality of data points,
[0866] One data point is information about one target-guide vector in which the target DNA has been cleaved,
[0867] Each data point contains barcode sequence and UMI sequence information;
[0868] (c-4) A process for determining target-guide activity based on the sequence information obtained in the above process (c-3);
[0869] Here, the above target-guided activity quantitatively indicates how well a CRISPR / Cas complex containing a specific guide nucleic acid cleaves a specific target.
[0870] The target-guided activity is determined based on how many UMIs each group contains when grouping the data points based on the barcode.
[0871] Example 76, Process for Evaluating Activity, Cut-seq2
[0872] In Example 70,
[0873] The above (c) process includes:
[0874] (c-1) (b) A process of collecting nucleic acids from a population of cells at the end of the process;
[0875] (c-2) A process of processing an adapter into the nucleic acid collected in the above (c-1) process,
[0876] Here, the adapter is linked to a position cut by a restriction enzyme in the above process (b) or cut by a CRISPR / Cas complex;
[0877] (c-3) A process of obtaining sequence information by sequencing the nucleic acid after the above (c-2) process is completed.
[0878] Here, the sequencing is performed including a process of amplifying the nucleic acid to which the adapter is linked,
[0879] The above sequence information includes a plurality of data points,
[0880] One data point is information about one target-guide vector,
[0881] Each data point includes target nucleic acid sequence information of the corresponding target-guide vector, guide nucleic acid sequence information, and information on whether the target nucleic acid is cleaved;
[0882] (c-4) A process for determining target-guide activity based on the sequence information obtained in the above process (c-3);
[0883] Here, the above target-guided activity quantitatively indicates how well a CRISPR / Cas complex containing a specific guide nucleic acid cleaves a specific target.
[0884] The above target-guide activity is determined based on the number of data points that include information that i) the target nucleic acid was cleaved, and ii) the number of data points that include information that the target nucleic acid was not cleaved, belonging to each group (i.e., target-guide combination), when the sequence information obtained in (c-3) is classified based on the target nucleic acid sequence and the guide nucleic acid sequence.
[0885] Example 77, Barcode and UMI criteria, Cut-seq2
[0886] In Example 70, the identifier of the target-guide vector includes both a barcode and a UMI,
[0887] The above (c) process includes:
[0888] (c-1) (b) A process of collecting nucleic acids from a population of cells at the end of the process;
[0889] (c-2) A process of processing an adapter into the nucleic acid collected in the above (c-1) process,
[0890] Here, the adapter is linked to a position cut by a restriction enzyme in the above process (b) or cut by a CRISPR / Cas complex;
[0891] (c-3) A process of obtaining sequence information by sequencing the nucleic acid after the above (c-2) process is completed.
[0892] Here, the sequencing is performed including a process of amplifying the nucleic acid to which the adapter is linked,
[0893] The above sequence information includes a plurality of data points,
[0894] One data point is information about one target-guide vector,
[0895] Each data point includes the barcode sequence, the UMI sequence, and information on whether the target nucleic acid is cleaved;
[0896] (c-4) A process for determining target-guide activity based on the sequence information obtained in the above process (c-3);
[0897] Here, the above target-guided activity quantitatively indicates how well a CRISPR / Cas complex containing a specific guide nucleic acid cleaves a specific target.
[0898] The above target-guide activity is determined based on the number of UMIs of data points including information that the target nucleic acid was cleaved, and ii) the number of UMIs of data points including information that the target nucleic acid was not cleaved, for each group when the sequence information is grouped based on the barcode sequence.
[0899] Structure of the cleavage activity prediction model of the CRISPR / Cas system
[0900] Example 78, Cleavage Activity Prediction Model Structure #1 - Sequence Information is Input Without Feature Extraction
[0901] A cleavage activity prediction model configured to input information about a guide nucleic acid sequence and a target nucleic acid sequence (hereinafter, target-guide sequence information) and output a cleavage activity prediction value, comprising:
[0902] Output Module;
[0903] Here, the cleavage activity prediction model is configured to receive the target-guide sequence information as input and perform the following:
[0904] Deriving an aligned target nucleic acid sequence, an aligned guide nucleic acid sequence, and (optionally) mismatch information (collectively referred to as aligned target-guide sequence information) from the above target-guide sequence information;
[0905] Here, the aligned guide nucleic acid sequence and the aligned target nucleic acid sequence are derived by aligning the target nucleic acid sequence and the guide nucleic acid sequence,
[0906] The above mismatch information compares the aligned target nucleic acid sequence with the aligned guide nucleic acid sequence, and indicates, for each position: 1) whether the bases of the two nucleic acids match (correspond); 2) whether the bases of the two nucleic acids do not match; 3) whether a gap is formed in the target nucleic acid; or 4) whether a gap is formed in the guide nucleic acid; and
[0907] The above aligned target-guide sequence information is appropriately preprocessed and input into the output module to output a cleavage activity prediction value;
[0908] The above cleavage activity prediction value is a quantitative prediction value of the degree to which the CRISPR / Cas complex including the guide nucleic acid and a specific Cas protein cleaves the target nucleic acid.
[0909] Example 79, Cut Activity Prediction Model Structure #2 - Sequence information is input without feature extraction, and additional information is input.
[0910] A cleavage activity prediction model configured to input information about a guide nucleic acid sequence and a target nucleic acid sequence (hereinafter, target-guide sequence information) and output a cleavage activity prediction value, comprising:
[0911] Concatenation Layer; and
[0912] Output Module;
[0913] Here, the cleavage activity prediction model is configured to receive the target-guide sequence information as input and perform the following:
[0914] Deriving an aligned target nucleic acid sequence, an aligned guide nucleic acid sequence, and (optionally) mismatch information (collectively referred to as aligned target-guide sequence information) from the above target-guide sequence information;
[0915] Here, the aligned guide nucleic acid sequence and the aligned target nucleic acid sequence are derived by aligning the target nucleic acid sequence and the guide nucleic acid sequence,
[0916] The above mismatch information compares the aligned target nucleic acid sequence with the aligned guide nucleic acid sequence, and indicates, for each position: 1) whether the bases of the two nucleic acids match (correspond); 2) whether the bases of the two nucleic acids do not match; 3) whether a gap is formed in the target nucleic acid; or 4) whether a gap is formed in the guide nucleic acid;
[0917] Additional information is derived by referring to the target-guide sequence information aligned with the above target-guide sequence information.
[0918] Here, the additional information is one or more selected from the following: a Protospacer Adjacent Motif (PAM) sequence; and whether a mutation has occurred in the PAM sequence; the type of the 20th base of the target nucleic acid sequence; the type of the guide nucleic acid; the melting temperature (T) of the guide nucleic acid sequence m ), T of the target nucleic acid sequence m ; Minimal Free Energy (MFE) of the guide nucleic acid sequence; MFE of the guide nucleic acid sequence including the scaffold; total number of mismatches, target nucleic acid bulges, and guide nucleic acid bulges; T of matched guide nucleic acid-target nucleic acid m; and the free energy change (ΔG) during guide nucleic acid-target nucleic acid hybridization H );
[0919] The above aligned target-guide sequence information and the above additional information are appropriately preprocessed, and the input information is derived by combining the above aligned target-guide sequence information in the connection layer; and
[0920] Input the above input information into the above output module to output the cut activity prediction value;
[0921] The above cleavage activity prediction value is a quantitative prediction value of the degree to which the CRISPR / Cas complex including the guide nucleic acid and a specific Cas protein cleaves the target nucleic acid.
[0922] Example 80, Cutting Activity Prediction Model Structure #3 - Sequence Information is Feature Extracted and Input
[0923] A cleavage activity prediction model configured to input information about a guide nucleic acid sequence and a target nucleic acid sequence (hereinafter, target-guide sequence information) and output a cleavage activity prediction value, comprising:
[0924] Feature Extractor; and
[0925] Output Module;
[0926] Here, the cleavage activity prediction model is configured to receive the target-guide sequence information as input and perform the following:
[0927] Deriving an aligned target nucleic acid sequence, an aligned guide nucleic acid sequence, and (optionally) mismatch information (collectively referred to as aligned target-guide sequence information) from the above target-guide sequence information;
[0928] Here, the aligned guide nucleic acid sequence and the aligned target nucleic acid sequence are derived by aligning the target nucleic acid sequence and the guide nucleic acid sequence,
[0929] The above mismatch information compares the aligned target nucleic acid sequence with the aligned guide nucleic acid sequence, and indicates, for each position: 1) whether the bases of the two nucleic acids match (correspond); 2) whether the bases of the two nucleic acids do not match; 3) whether a gap is formed in the target nucleic acid; or 4) whether a gap is formed in the guide nucleic acid;
[0930] The above aligned target-guide sequence information is appropriately preprocessed and input into the feature extractor to extract sequence feature information; and
[0931] Input the above sequence feature information into the output module to output a cleavage activity prediction value;
[0932] The above cleavage activity prediction value is a quantitative prediction value of the degree to which the CRISPR / Cas complex including the guide nucleic acid and a specific Cas protein cleaves the target nucleic acid.
[0933] Example 81, Cutting Activity Prediction Model Structure #4 - Sequence information is input as feature extraction, and additional information is input.
[0934] A cleavage activity prediction model configured to input information about a guide nucleic acid sequence and a target nucleic acid sequence (hereinafter, target-guide sequence information) and output a cleavage activity prediction value, comprising:
[0935] Feature Extractor;
[0936] Concatenation Layer; and
[0937] Output Module;
[0938] Here, the cleavage activity prediction model is configured to receive the target-guide sequence information as input and perform the following:
[0939] Deriving an aligned target nucleic acid sequence, an aligned guide nucleic acid sequence, and (optionally) mismatch information (collectively referred to as aligned target-guide sequence information) from the above target-guide sequence information;
[0940] Here, the aligned guide nucleic acid sequence and the aligned target nucleic acid sequence are derived by aligning the target nucleic acid sequence and the guide nucleic acid sequence,
[0941] The above mismatch information compares the aligned target nucleic acid sequence with the aligned guide nucleic acid sequence, and indicates, for each position: 1) whether the bases of the two nucleic acids match (correspond); 2) whether the bases of the two nucleic acids do not match; 3) whether a gap is formed in the target nucleic acid; or 4) whether a gap is formed in the guide nucleic acid;
[0942] Additional information is derived by referring to the target-guide sequence information aligned with the above target-guide sequence information.
[0943] Here, the additional information is one or more selected from the following: a Protospacer Adjacent Motif (PAM) sequence; and whether a mutation has occurred in the PAM sequence; the type of the 20th base of the target nucleic acid sequence; the type of the guide nucleic acid; the melting temperature (T) of the guide nucleic acid sequence m ), T of the target nucleic acid sequence m ; Minimal Free Energy (MFE) of the guide nucleic acid sequence; MFE of the guide nucleic acid sequence including the scaffold; total number of mismatches, target nucleic acid bulges, and guide nucleic acid bulges; T of matched guide nucleic acid-target nucleic acid m ; and the free energy change (ΔG) during guide nucleic acid-target nucleic acid hybridization H );
[0944] The above aligned target-guide sequence information is appropriately preprocessed and input into the feature extractor to extract sequence feature information;
[0945] The above additional information is appropriately preprocessed and combined with the sequence feature information in the connection layer to derive input information;
[0946] Input the above input information into the output module to output a cut activity prediction value;
[0947] The above cleavage activity prediction value is a quantitative prediction value of the degree to which the CRISPR / Cas complex including the guide nucleic acid and a specific Cas protein cleaves the target nucleic acid.
[0948] Example 82, Specific limitation of model structure
[0949] In any one of Examples 78 to 81, the model has a structure selected from the following:
[0950] Fully Connected Layer;
[0951] Convolutional Neural Network (CNN);
[0952] Transformer Encoder; and
[0953] GRU (Gated Recurrent Unit);
[0954] Convolutional Neural Network (CNN);
[0955] Deep Neural Network (DNN);
[0956] Simple Recurrent Neural Network (Simple RNN);
[0957] Transformer Encoder;
[0958] GRU (Gated Recurrent Unit); and
[0959] LSTM (Long Short-Term Memory).
[0960] Example 83, Feature Extractor Limited
[0961] In any one of embodiments 78 to 82, the feature extractor is a structure selected from the following:
[0962] Convolutional Neural Network (CNN);
[0963] Transformer Encoder; and
[0964] GRU (Gated Recurrent Unit);
[0965] Example 84, Output Module Limited
[0966] In any one of embodiments 78 to 83, the output module is a structure selected from the following:
[0967] Fully Connected Layer;
[0968] Convolutional Neural Network (CNN);
[0969] Deep Neural Network (DNN);
[0970] Simple Recurrent Neural Network (Simple RNN);
[0971] Transformer Encoder;
[0972] GRU (Gated Recurrent Unit); and
[0973] LSTM (Long Short-Term Memory).
[0974] Example 85, including additional information
[0975] In any one of Examples 78 to 84, the cleavage activity prediction model is configured to output a prediction value by additionally using one or more pieces of information that influence the CRISPR / Cas complex to cleave the target nucleic acid.
[0976] Example 86, information used in actual models among target nucleic acid and guide nucleic acid
[0977] In any one of Examples 78 to 85,
[0978] Information about the sequence of the above guide nucleic acid is information about the sequence of the guide domain of the above guide nucleic acid (see Examples 8 to 14),
[0979] Information on the sequence of the target nucleic acid refers to information on the sequence of the target portion of the target nucleic acid (see Examples 15 to 24).
[0980] How to learn a cut activity prediction model
[0981] Example 87, method for learning a cutting activity prediction model
[0982] A method for training a model (hereinafter, a cleavage activity prediction model) that quantitatively predicts the extent to which a CRISPR / Cas complex comprising the guide nucleic acid and a specific Cas protein cleaves the target nucleic acid by inputting target-guide sequence information including:
[0983] (a) The process of preparing training data;
[0984] Here, the training data includes a plurality of training data points,
[0985] Each training data point contains aligned target-guide sequence information, optionally additional information, and is labeled with a target-guide activation value.
[0986] Here, the target-guide activity value is a quantitative measurement of the degree to which the CRISPR / Cas complex containing the guide nucleic acid and a specific Cas protein cleaves the target nucleic acid; and
[0987] (b) A process of training a cut activity prediction model using the above training data.
[0988] Example 88, Limited to the model to be trained
[0989] In Example 87, the cutting activity prediction model has a structure selected from any one of Examples 78 to 86.
[0990] Example 89, Segmenting the Training Data Preparation Process
[0991] In any one of Examples 87 to 88, the process (a) comprises:
[0992] (a-1) Raw data acquisition process,
[0993] Here, the raw data includes a plurality of raw data points,
[0994] The above raw data points include 1) a target nucleic acid sequence, 2) a guide nucleic acid sequence, and 3) an activity of a CRISPR / Cas complex comprising the guide nucleic acid and a specific Cas protein to cleave the target nucleic acid (hereinafter, target-guide activity);
[0995] (a-2) Alignment Data derivation process,
[0996] Here, the raw data can be augmented according to predetermined criteria; and
[0997] (a-3) Training data generation process.
[0998] Example 90, raw data acquisition process
[0999] In Example 89, the raw data is obtained through any one method selected from Examples 59 to 77.
[1000] Example 91, Process of deriving sorted data
[1001] In any one of Examples 89 to 90,
[1002] The above sorting data includes a plurality of sorting data points,
[1003] In the above process (a-2), the target nucleic acid sequence and the guide nucleic acid sequence of each raw data point are aligned to derive an aligned data point corresponding to the raw data point.
[1004] Here, the alignment data points include 1) an aligned target nucleic acid sequence, 2) an aligned guide nucleic acid sequence, and 3) a target-guide activity of the raw data points.
[1005] At this time, the target nucleic acid sequence and the guide nucleic acid sequence can be aligned into two or more patterns,
[1006] If two or more alignment patterns satisfy the above predetermined criteria,
[1007] Deriving multiple sorted data points using all sorting patterns that satisfy the above predetermined criteria.
[1008] Example 92, Training Data Generation Process, No Additional Information
[1009] In any one of Examples 89 to 91,
[1010] The above training data includes a plurality of training data points,
[1011] In the above process (a-3), training data points are derived from each sorted data point,
[1012] 1) the aligned target nucleic acid sequence, 2) the aligned guide nucleic acid sequence, and 3) (optionally) the mismatch information of each of the above alignment data points are input variables,
[1013] Label the target-guided activity of each of the above alignment data points as the output value.
[1014] Example 93, Training Data Generation Process, Additional Information Derivation
[1015] In any one of Examples 89 to 91,
[1016] The above training data includes a plurality of training data points,
[1017] In the above process (a-3), training data points are derived from each sorted data point,
[1018] 1) the aligned target nucleic acid sequence, 2) the aligned guide nucleic acid sequence, and 3) (optionally) the mismatch information of each of the above alignment data points are input variables,
[1019] Add one or more additional information selected from the following as input variables: Protospacer Adjacent Motif (PAM) sequence; and whether a mutation has occurred in the PAM sequence; the type of the 20th base of the target nucleic acid sequence; the type of guide nucleic acid; the melting temperature (T) of the guide nucleic acid sequence m ), T of the target nucleic acid sequence m ; Minimal Free Energy (MFE) of the guide nucleic acid sequence; MFE of the guide nucleic acid sequence including the scaffold; total number of mismatches, target nucleic acid bulges, and guide nucleic acid bulges; T of matched guide nucleic acid-target nucleic acid m ; and the free energy change (ΔG) during guide nucleic acid-target nucleic acid hybridization H ),
[1020] Label the target-guided activity of each of the above alignment data points as the output value.
[1021] Example 94, Additional Information Available
[1022] In any one of Examples 87 to 93, each training data point further includes as an input variable one or more pieces of information that can influence the CRISPR / Cas complex to cleave the target nucleic acid.
[1023] Example 95, Learning Process Related
[1024] In any one of Examples 87 to 94, the step (b) comprises:
[1025] (b-1) A process of appropriately preprocessing 1) aligned target nucleic acid sequence, 2) aligned guide nucleic acid sequence, and 3) (optionally) mismatch information of training data and inputting them into a feature extractor to derive sequence features;
[1026] (b-2) If the training data contains additional information, the process of appropriately preprocessing it and connecting it with the sequence features in the concatenation layer; and
[1027] (b-3) A process of training an output module using data from (b-1) or (b-2) as input variables and target-guide activity as output variables.
[1028] Example 96, Feature Extractor Limited
[1029] In Example 95, the feature extractor is selected from the following:
[1030] Fully Connected Layer;
[1031] Convolutional Neural Network (CNN);
[1032] Transformer Encoder; and
[1033] GRU (Gated Recurrent Unit);
[1034] Example 97, Output Module Limited
[1035] In any one of embodiments 95 to 96, the output module is selected from the following:
[1036] Fully Connected Layer;
[1037] Convolutional Neural Network (CNN);
[1038] Deep Neural Network (DNN);
[1039] Simple Recurrent Neural Network (Simple RNN);
[1040] Transformer Encoder;
[1041] GRU (Gated Recurrent Unit); and
[1042] LSTM (Long Short-Term Memory).
[1043] Example 98, Deriving training data from n types of raw data
[1044] In any one of Examples 87 to 97,
[1045] The above process (a) includes:
[1046] (a-1) A process of obtaining raw data containing n raw data points,
[1047] Here, when i is a random integer between 1 and n,
[1048] The ith raw data point is (t i , g i , c i ) and
[1049] Here, t i is the sequence of the target nucleic acid of the i-th raw data point,
[1050] Here, g i is the sequence of the guide nucleic acid of the i-th raw data point,
[1051] Here, c i is the i-th raw data point, g i t of the CRISPR / Cas complex containing i is cutting active; and
[1052] (a-2) A process for generating training data including m training data points from the above raw data,
[1053] Here, the target nucleic acid sequence and guide nucleic acid sequence of each of the above raw data are aligned, and a target-guide alignment pattern is selected according to a predetermined criterion.
[1054] t of the ith raw data point i Wow g i When aligning, the selected target-guide alignment pattern is a i If you say so,
[1055] Above And,
[1056] The jth training data point derived from the ith raw data point contains the following information:
[1057] (at ij , ag ij , c i ),
[1058] Here, j is from 1 to a i is an arbitrary integer between ,
[1059] at ij is t i Wow g iis the jth aligned target nucleic acid sequence derived by sorting,
[1060] ag ij is t i Wow g i The jth aligned guide nucleic acid sequence derived by sorting.
[1061] Example 99, Meaning of Sequence Information
[1062] In any one of Examples 87 to 98,
[1063] Information about the sequence of the above guide nucleic acid is information about the sequence of the guide domain of the above guide nucleic acid (see Examples 8 to 14),
[1064] Information on the sequence of the target nucleic acid refers to information on the sequence of the target portion of the target nucleic acid (see Examples 15 to 24).
[1065] Method for predicting target cleavage selectivity
[1066] Example 109, Method for Predicting Target Cleavage Selectivity
[1067] A method for predicting the extent to which one of two target nucleic acids is selectively cleaved, comprising:
[1068] (a) a process of obtaining a first target nucleic acid sequence (target nucleic acid), a second target nucleic acid sequence (noise nucleic acid), and a guide nucleic acid sequence; and
[1069] (b) A process of calculating the difference between the cleavage activity for the first target nucleic acid sequence and the cleavage activity for the second target nucleic acid sequence using a cleavage activity prediction model.
[1070] Example 110, Prediction Model Limited
[1071] In Example 109, the cutting activity prediction model has a structure of any one of Examples 78 to 86, or is a model trained by any one of Examples 87 to 99.
[1072] Example 111, Process for Obtaining Difference in Cutting Activity
[1073] In any of the No entry found cases, the above (b) process includes:
[1074] (b-1) A process for quantitatively predicting the extent to which a CRISPR / Cas complex comprising a guide nucleic acid and a specific Cas protein cleaves a first target nucleic acid by any one of the methods of Examples 100 to 108;
[1075] Here, the derived predicted value is called the first value;
[1076] (b-2) A process for quantitatively predicting the extent to which a CRISPR / Cas complex comprising a guide nucleic acid and a specific Cas protein cleaves a second target nucleic acid by any one of the methods of Examples 100 to 108;
[1077] Here, the derived predicted value is called the second value; and
[1078] (b-3) A process for obtaining the difference between the cleavage activity for the first target nucleic acid sequence and the cleavage activity for the second target nucleic acid sequence based on the first value and the second value.
[1079] Example 112: Process for Finding Selectivity
[1080] In any one of Examples 109 to 111, the difference between the cleavage activity for the first target nucleic acid sequence and the cleavage activity for the second target nucleic acid sequence is referred to as cleavage selectivity for the first target nucleic acid sequence or cleavage selectivity for the second target nucleic acid sequence.
[1081] Example 113, Meaning of Sequence Information
[1082] In any one of Examples 109 to 112,
[1083] Information about the sequence of the above guide nucleic acid is information about the sequence of the guide domain of the above guide nucleic acid (see Examples 8 to 14),
[1084] Information on the sequence of the target nucleic acid refers to information on the sequence of the target portion of the target nucleic acid (see Examples 15 to 24).
[1085] A method for determining a guide sequence with high target selectivity
[1086] Example 114, Method for Determining a Guide Sequence with High Target Selectivity
[1087] A method for determining an optimized guide nucleic acid that selectively cleaves a target nucleic acid over a noise nucleic acid, comprising:
[1088] (a) a process of obtaining a target nucleic acid sequence to be cut and a noise nucleic acid sequence to be not cut;
[1089] (b) a process for generating guide nucleic acid candidates;
[1090] (c) for each guide nucleic acid candidate, a process of predicting target nucleic acid cleavage selectivity using a cleavage activity prediction model; and
[1091] (d) A process for determining optimized guide nucleic acid(s) that satisfy predetermined criteria.
[1092] Example 115, Prediction Model Limited
[1093] In Example 114, the cutting activity prediction model has a structure of any one of Examples 78 to 86, or is a model trained by any one of Examples 87 to 99.
[1094] Example 116, Process for generating guide nucleic acid candidates
[1095] In any one of Examples 114 to 115, the guide nucleic acid candidates generated in the process (b) include one or more selected from the following:
[1096] (1) A guide nucleic acid having a sequence perfectly matched to the target nucleic acid;
[1097] (2) A guide nucleic acid in which any one or more bases are changed to a different base, compared to (1) above;
[1098] (3) A guide nucleic acid in which any one or more bases are removed at any position, compared to (1) above;
[1099] (4) compared to (1) above, a guide nucleic acid having one or more bases added at any position; and
[1100] (5) A guide nucleic acid in which the modifications disclosed in (2) to (4) are arbitrarily combined, compared to (1) above.
[1101] Example 117, Process for Generating Guide Nucleic Acid Candidates, Description Focusing on the Modification Action
[1102] In any one of Examples 114 to 116, the step (b) comprises:
[1103] The above guide nucleic acid candidates include guide nucleic acids having a sequence that perfectly matches the target nucleic acid,
[1104] Contains one or more sequences having the following modifications applied based on a sequence that perfectly corresponds to the above target nucleic acid:
[1105] (1) Changing any one or more bases to other bases;
[1106] (2) Removing any one or more bases at any one or more positions;
[1107] (3) adding any one or more bases at any one or more positions;
[1108] (4) Any combination of the above (1) to (3) variations.
[1109] Example 118, Process for Predicting Target Nucleic Acid Cleavage Selectivity
[1110] In any one of Examples 114 to 117, the step (c) comprises:
[1111] Using any one of the methods of Examples 109 to 113, the difference between the cleavage activity for the target nucleic acid and the cleavage activity for the noise nucleic acid (cleavage selectivity for the target nucleic acid) of each guide nucleic acid candidate sequence is determined.
[1112] Example 119, Process for Determining Optimized Guide Nucleic Acids
[1113] In any one of Examples 114 to 118, the step (d) comprises:
[1114] Verify that the cleavage selectivity of each guide nucleic acid candidate sequence for the target nucleic acid meets a predetermined standard,
[1115] A guide nucleic acid candidate sequence that meets predetermined criteria is determined as an optimized guide nucleic acid.
[1116] Example 120, Partial Specification of the Guide Nucleic Acid
[1117] In any one of Examples 114 to 119,
[1118] The above guide nucleic acid comprises a guide domain and a scaffold,
[1119] Determining an optimized guide nucleic acid by the above method means that the sequence of the guide domain of the guide nucleic acid has been determined.
[1120] Rare Nucleic Acid Enrichment Method #1 - Background Nucleic Acid Cleavage
[1121] Example 121, Method for Enriching Rare Nucleic Acids by Cleaving Background Nucleic Acids
[1122] A method for enriching rare nucleic acids, comprising:
[1123] (a) the process of obtaining a sample from the subject;
[1124] Here, the sample is suspected of containing a rare nucleic acid to be concentrated and a background nucleic acid having a sequence similar to the rare nucleic acid;
[1125] (b) a process of inducing a CRISPR / Cas complex comprising an optimized guide nucleic acid and a Cas protein to come into contact with the nucleic acid in the sample;
[1126] Here, the CRISPR / Cas complex selectively cleaves background nucleic acids in the sample; and
[1127] (c) a process for processing the rare nucleic acid so that it can be selectively amplified and concentrated;
[1128] (d) (c) A process for identifying whether a rare nucleic acid exists in a sample that has undergone the process.
[1129] Example 122, CRISPR / Cas processing method
[1130] In Example 121, the process (b) is performed by a method selected from the following:
[1131] Processing the guide nucleic acid and Cas protein optimized for the above sample; or
[1132] Processing a CRISPR / Cas complex containing a guide nucleic acid optimized for the sample and a Cas protein.
[1133] Example 123, Purpose of Enrichment of Cancer-Associated Mutations
[1134] In any one of Examples 121 to 122, the method is a method for enriching a nucleic acid comprising a cancer-associated mutant sequence,
[1135] The above rare nucleic acid comprises a cancer-associated mutant sequence of a specific gene,
[1136] The above background nucleic acid contains a wild type sequence corresponding to the above rare nucleic acid.
[1137] Example 124, Optimized Guide Nucleic Acid Selection Method
[1138] In any one of Examples 121 to 123, the optimized guide nucleic acid is determined by any one method selected from Examples 114 to 120,
[1139] When performing the above method, the noise nucleic acid is the rare nucleic acid, and the target nucleic acid is the background nucleic acid.
[1140] Example 125, using SpCas9-HF1
[1141] In Example 123, the Cas protein is SpCas9-HF1 protein,
[1142] The guide domain of the above specific gene and the above optimized guide nucleic acid is at least one selected from the following combinations:
[1143] When the cancer-related gene is the ABCA12 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 21 to SEQ ID NO: 25; When the cancer-related gene is the AFF3 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 38 to SEQ ID NO: 40; When the cancer-related gene is the AHNAK2 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 41 to SEQ ID NO: 42; When the cancer-related gene is the AKAP9 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 43 to SEQ ID NO: 45; When the cancer-related gene is the AMPH gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 57 to SEQ ID NO: 62; When the cancer-related gene is the APC gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 73 to SEQ ID NO: 127; When the cancer-related gene is the ARID1B gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 175 to SEQ ID NO: 180; When the cancer-related gene is the ASXL1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 190 to SEQ ID NO: 194; When the cancer-related gene is the ATP10A gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 222 to SEQ ID NO: 225; When the cancer-related gene is the BCL11A gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 245; When the cancer-related gene is the BCL7A gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 246 to SEQ ID NO: 271; When the cancer-related gene is the BRAFP1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 290 to SEQ ID NO: 292;When the cancer-related gene is the BRCA1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 293 to SEQ ID NO: 302; When the cancer-related gene is the CACNA1C gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 329 to SEQ ID NO: 334; When the cancer-related gene is the CACNA1D gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 335 to SEQ ID NO: 338; When the cancer-related gene is the CCND1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 355 to SEQ ID NO: 357; When the cancer-related gene is the CDH18 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 375 to SEQ ID NO: 383; When the cancer-related gene is a CDK12 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 391 to SEQ ID NO: 397; When the cancer-related gene is a FLT4 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 858 to SEQ ID NO: 859; When the cancer-related gene is a GATA3 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 870 to SEQ ID NO: 876; When the cancer-related gene is a MKRN2 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 1260 to SEQ ID NO: 1262; When the cancer-related gene is a MYD88 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 1291 to SEQ ID NO: 1292; When the cancer-related gene is the NCOR1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 1330 to SEQ ID NO: 1336;When the cancer-related gene is the NFE2L2 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 1377 to SEQ ID NO: 1412; When the cancer-related gene is the NOTCH1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 1423 to SEQ ID NO: 1429; When the cancer-related gene is the PKHD1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 1648 to SEQ ID NO: 1654; When the cancer-related gene is the POLE gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 1667 to SEQ ID NO: 1676; When the cancer-related gene is the POLQ gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 1677; When the cancer-related gene is the PTPN11 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 1762 to SEQ ID NO: 1769; When the cancer-related gene is the PTPN13 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 1770 to SEQ ID NO: 1775; When the cancer-related gene is the RAD51B gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 1832; When the cancer-related gene is the RAF1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 1833; When the cancer-related gene is the SPAG17 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2041 to SEQ ID NO: 2042; When the cancer-related gene is a SPEN gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2043 to SEQ ID NO: 2064; When the cancer-related gene is a SRCAP gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2073 to SEQ ID NO: 2078;When the cancer-related gene is the TET1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 2127; When the cancer-related gene is the THBS2 gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 2142; When the cancer-related gene is the TRRAP gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 2515 to SEQ ID NO: 2531; When the cancer-related gene is the UBR5 gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 2552 to SEQ ID NO: 2559; When the cancer-related gene is the VWF gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 2576 to SEQ ID NO: 2582; When the cancer-related gene is the XIRP2 gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 2595 to SEQ ID NO: 2597; and when the cancer-related gene is the ZRSR2 gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 2632.
[1144] Here, the cancer-associated mutations targeted by each guide domain are described in the description section of each sequence in the sequence listing file.
[1145] Example 126, using SpCas9-NRRH-HF1
[1146] In Example 123, the Cas protein is SpCas9-NRRH-HF1 protein,
[1147] The guide domain of the above specific gene and the above optimized guide nucleic acid is at least one selected from the following combinations:
[1148] When the cancer-related gene is the ABCA12 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2633 to SEQ ID NO: 2637; When the cancer-related gene is the AFF3 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2650 to SEQ ID NO: 2652; When the cancer-related gene is the AHNAK2 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2653 to SEQ ID NO: 2654; When the cancer-related gene is the AKAP9 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2655 to SEQ ID NO: 2657; When the cancer-related gene is the AMPH gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2669 to SEQ ID NO: 2674; When the cancer-related gene is the APC gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2685 to SEQ ID NO: 2739; When the cancer-related gene is the ARID1B gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2787 to SEQ ID NO: 2792; When the cancer-related gene is the ASXL1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2802 to SEQ ID NO: 2806; When the cancer-related gene is the ATP10A gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2834 to SEQ ID NO: 2837; When the cancer-related gene is the BCL11A gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2857; When the cancer-related gene is the BCL7A gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 2858 to SEQ ID NO: 2883;When the cancer-related gene is the BRAFP1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2902 to SEQ ID NO: 2904; When the cancer-related gene is the BRCA1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2905 to SEQ ID NO: 2914; When the cancer-related gene is the CACNA1C gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2941 to SEQ ID NO: 2946; When the cancer-related gene is the CACNA1D gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2947 to SEQ ID NO: 2950; When the cancer-related gene is the CCND1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2967 to SEQ ID NO: 2969; When the cancer-related gene is the CDH18 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 2987 to SEQ ID NO: 2995; When the cancer-related gene is the CDK12 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 3003 to SEQ ID NO: 3009; When the cancer-related gene is the FLT4 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 3470 to SEQ ID NO: 3471; When the cancer-related gene is the GATA3 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 3482 to SEQ ID NO: 3488; When the cancer-related gene is the MKRN2 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 3872 to SEQ ID NO: 3874; When the cancer-related gene is the MYD88 gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 3903 to SEQ ID NO: 3904;When the cancer-related gene is the NCOR1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 3942 to SEQ ID NO: 3948; When the cancer-related gene is the NFE2L2 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 3989 to SEQ ID NO: 4024; When the cancer-related gene is the NOTCH1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 4035 to SEQ ID NO: 4041; When the cancer-related gene is the PKHD1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 4260 to SEQ ID NO: 4266; When the cancer-related gene is the POLE gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 4279 to SEQ ID NO: 4288; When the cancer-related gene is a POLQ gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 4289; When the cancer-related gene is a PTPN11 gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 4374 to SEQ ID NO: 4381; When the cancer-related gene is a PTPN13 gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 4382 to SEQ ID NO: 4387; When the cancer-related gene is a RAD51B gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 4444; When the cancer-related gene is a RAF1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 4445; When the cancer-related gene is a SPAG17 gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 4653 to SEQ ID NO: 4654; When the cancer-related gene is a SPEN gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 4655 to SEQ ID NO: 4676;When the cancer-related gene is the SRCAP gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 4685 to SEQ ID NO: 4690; When the cancer-related gene is the TET1 gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 4739; When the cancer-related gene is the THBS2 gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 4754; When the cancer-related gene is the TRRAP gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 5127 to SEQ ID NO: 5143; When the cancer-related gene is the UBR5 gene, the guide domain of the optimized sgRNA comprises a sequence selected from SEQ ID NO: 5164 to SEQ ID NO: 5171; When the cancer-related gene is a VWF gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 5188 to SEQ ID NO: 5194; When the cancer-related gene is a XIRP2 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 5207 to SEQ ID NO: 5209; and when the cancer-related gene is a ZRSR2 gene, the guide domain of the optimized sgRNA comprises a sequence selected from among SEQ ID NO: 5244;
[1149] Here, the cancer-associated mutations targeted by each guide domain are described in the description section of each sequence in the sequence listing file.
[1150] Rare Nucleic Acid Enrichment Method #2 - Rare Nucleic Acid Cleavage
[1151] Example 127, Method for Concentrating Rare Nucleic Acids by Cleaving Rare Nucleic Acids
[1152] A method for enriching rare nucleic acids, comprising:
[1153] (a) the process of obtaining a sample from the subject;
[1154] Here, the sample is suspected of containing a rare nucleic acid to be concentrated and a background nucleic acid having a sequence similar to the rare nucleic acid;
[1155] (b) a process of blocking the terminal portions of nucleic acids contained in the sample;
[1156] (c) a process of inducing a CRISPR / Cas complex comprising an optimized guide nucleic acid and a Cas protein to come into contact with the nucleic acid in the sample;
[1157] Here, the CRISPR / Cas complex selectively cleaves rare nucleic acids in the sample;
[1158] (d) (c) After the process, a process of processing an adapter that can be connected to the nucleic acid cutting part in the sample; and
[1159] (e) A process for selectively amplifying and concentrating a fragment of a rare nucleic acid to which the above adapter is connected.
[1160] Example 128, CRISPR / Cas processing method
[1161] In Example 127, the process (c) is performed by a method selected from the following:
[1162] Processing the guide nucleic acid and Cas protein optimized for the above sample; or
[1163] Processing a CRISPR / Cas complex containing a guide nucleic acid optimized for the sample and a Cas protein.
[1164] Example 129, Purpose of Enrichment of Cancer-Associated Mutations
[1165] In any one of Examples 127 to 128, the method is a method for enriching a nucleic acid comprising a cancer-associated mutant sequence,
[1166] The above rare nucleic acid comprises a cancer-associated mutant sequence of a specific gene,
[1167] The above background nucleic acid contains a wild type sequence corresponding to the above rare nucleic acid.
[1168] Example 130, Optimized Guide Nucleic Acid Selection Method
[1169] In any one of Examples 127 to 129, the optimized guide nucleic acid is determined by any one method selected from Examples 114 to 120,
[1170] When performing the above method, the noise nucleic acid is used as the background nucleic acid, and the target nucleic acid is used as the rare nucleic acid.
[1171] Target-guide vector-centric embodiment
[1172] Example 131, target-guide vector
[1173] Nucleic acids containing:
[1174] (i) a nucleic acid encoding sgRNA;
[1175] (ii) a promoter that induces transcription of the sgRNA;
[1176] (iii) a target sequence, wherein said sgRNA can hybridize with said target sequence;
[1177] (iv) barcode oligonucleotides.
[1178] Example 132, Promoter Operating Environment
[1179] In Example 131, the promoter can induce transcription of the sgRNA in a bacterial or animal cell.
[1180] Example 133, Promoter Sequence Examples
[1181] In any one of Examples 131 to 132, the promoter is selected from the group consisting of CMV, Efla, CAG, PGK, TRE, U6, T7, lacUV5, gapA, T5, recA, Ptac, Patac, pAl, lac, Sp6, araBad, trp, and fusion promoters thereof.
[1182] Example 134, T7 only
[1183] In Example 133, the promoter is a T7 promoter.
[1184] Example 135, including the first restriction enzyme cleavage site
[1185] In any one of Examples 131 to 134, the nucleic acid further comprises a first restriction enzyme cleavage site,
[1186] The above first restriction enzyme cleavage site can be recognized and cut by the first restriction enzyme.
[1187] Example 136, length limitation
[1188] In Example 135, the first restriction enzyme cleavage site is located within 100 bp downstream of the nucleic acid encoding the sgRNA.
[1189] Example 137, limited cross-section
[1190] In Example 136, the first restriction enzyme cutting site is cut by the first restriction enzyme, resulting in the formation of a sticky end or a blunt end.
[1191] Example 138, restriction enzyme limitation of the first
[1192] In any one of Examples 135 to 137, the first restriction enzyme is AccI, Acc65, AhaII, AsuII, Asp718, AvaI, AvrII, BamH1, BelI, BglIII, BsaNI, BspMII, BssHII, BspMI, BstEII, Bsu90I, DdeI, Eco0109, EcoR1, EcoRII, HindIII, HinPI, HinfI, HpaII, MaeI, MaeII, MaeIII, MboI, MluI, NarI, NcoI, NdeI, NdeII, NheI, Not1, PpuMI, RsrII, Sal1, SauI, Sau3AI, Sau961, ScrFI, SpeI, StyI, TagI, TthMI, XbaI, XhoII, XmaI, AarI, AscI, BbrCI, CspI, DraI, FseI, NotI, NruI, PacI, PmeI, PvuI, SapI, SdaI, SplI and SwaI; optionally, the first restriction enzyme is EcoR1.
[1193] Example 139, Target-Guide Relationship
[1194] In any one of Examples 131 to 138, the sgRNA comprises a sequence that is at least 70% complementary to the target sequence.
[1195] Example 140, including mismatches
[1196] In any one of Examples 131 to 139, the sgRNA comprises 1 to 5 mismatches compared to the target sequence.
[1197] Example 141, Mismatch Pattern Limitation
[1198] In Example 140, the mismatch means that a base is substituted, a base is removed, or a base is added compared to a sequence complementary to the target sequence.
[1199] Example 142, Mismatch Location Limitation
[1200] In Example 140, the mismatch corresponds to the 4th to 8th positions of the sgRNA.
[1201] Example 143, including PAM
[1202] In any one of Examples 131 to 142, the target sequence comprises a PAM.
[1203] Example 144, PAM sequence limitation
[1204] The method of Example 143, wherein the PAM is NGG, NAG, NGA, NGT, NGC, NAA, AAAG, AACG, AAGG, AATG, CAAG, CACG, CAGG, CATG, GAAG, GACG, GAGG, GATG, TAAG, TACG, TAGG, TATG, ACAG, ACCG, ACGG, CCGG, GCAG, GCCG, GCGG, TCGG, AGAG, AGCG, AGGG, AGTG, CGAG, CGCG, CGGG, CGTG, GGAG, GGCG, GGGG, GGTG, TGAG, TGCG, TGGG, TGTG, ATAG, ATCG, ATGG, ATTG, CTAG, CTCG, CTGG, CTTG, GTAG, GTCG, GTGG, GTTG, TTGG, GGAT, GGCT, GGGT, GGTT, TGCT, Selected from the group consisting of AGAA, AGCA, AGGA, AGTA, CGAA, CGCA, CGGA, CGTA, GGAA, GGCA, GGGA, GGTA, TGAA, TGCA, TGGA, TGTA, AGAC, AGCC, AGGC, AGTC, CGAC, CGCC, CGGC, CGTC, GGAC, GGCC, GGGC, GGTC, TGAC, TGCC, TGGC, TGTC and ATGA, wherein N is a base selected from A, T, G or C, and optionally, the PAM is NGG.
[1205] Example 145, target sequence length limitation
[1206] In any one of Examples 131 to 144, the target sequence does not exceed 200 bp in length, preferably 30 bp to 140 bp, more preferably 120 bp.
[1207] Example 146, target sequence type limitation
[1208] In any one of Examples 131 to 145, the target sequence is at least a portion of a cancer-associated gene; Here, the target sequence is optionally ACVR2A, ADAM28, AHNAK2, AKT1, AKT1, ANKRD36C, ANKRD36C, APC, ATM, AR, ARID1A, AS1, BCL7A, BMPR2, BRAF, CRAF, C12orf4, CDKN2A, CHEK2, CHEK2, CTNNB1, DOCK3, DNMT3A, EGFR, EGFR, ERBB2, ERBB4, ESR1, FAM47C, FAT1, FAT4, FBXW7, FBXW7, FGFR3, FHOD3, GATA3, GNAS, HRAS, IDH1, IDH2, KIAA2026, KEAP1, KEAP1, KIT, KMT2C, KMT2D, KNT2AKRAS, KRAS, KRTAP1-5, KRTAP4-11, KRTAP4-11, LARP4B, MAP2K4, MAP3K1, MBOAT2, MET, MTORNFE2L2, NFE2L2, NF1, NRAS, NTRK3, PDE4DIPPGM5, PIK3CA, PIK3CA, PIK3R1, PLEKHA6, POLE, PTEN, PTPRT, RGPD8, RGPD8, RNF213, RNF43, RUNX1RXRA, SETD2 SF3B1, SPEN, SPOP, SMAD4, SMARCA4, STK11, SVIL, TGFBR2, TP53, TRIM48, TRRAP, TSPOAP1, U2AF1, UBR5, WNT16, XYLT2, ZAN, ZBTB20, With ZFHX3 and ZNF814 Selected from the group formed.
[1209] Example 147, including mutations
[1210] In any one of Examples 131 to 146, the target sequence comprises a sequence comprising a mutation in the cancer-related gene.
[1211] Example 148, PAM mutation
[1212] In Example 147, the mutation occurs at the PAM position.
[1213] Example 149, PAM mutation
[1214] In Example 147, the mutation occurs at a position other than PAM.
[1215] Example 150, including barcodes and UMIs
[1216] In any one of Examples 131 to 149, the barcode oligonucleotide comprises a barcode sequence and a Unique Molecular Identifier (UMI).
[1217] Example 151, Barcode length limitation
[1218] In Example 150, the barcode sequence has a length of 5 bp to 50 bp, and optionally, a length of 20 bp.
[1219] Example 152, UMI length limitation
[1220] In any one of embodiments 150 to 151, the UMI has a length of 3 bp to 15 bp, and optionally, a length of 8 bp.
[1221] Example 153, Encoding target of barcode sequence
[1222] In any one of Examples 150 to 152, barcode sequences associated with the same target sequence and sgRNA pair are identical, and barcode sequences associated with different target sequences and sgRNA pairs are different.
[1223] Example 154, Encoding target of UMI sequence
[1224] In any one of Examples 150 to 153, UMIs associated with the same target sequence and sgRNA pair are different.
[1225] Example 155, including an invariant sequence
[1226] In any one of Examples 131 to 154, the nucleic acid further comprises a constant sequence.
[1227] Example 156, including constant sites and second restriction enzyme cleavage sites
[1228] In Example 155, the constant sequence comprises a constant region and a second restriction enzyme cleavage site, and the second restriction enzyme cleavage site can be recognized and cleaved by the second restriction enzyme.
[1229] Example 157, range of second restriction enzyme cleavage sites
[1230] In Example 156, the second restriction enzyme cleavage site is located within 100 bp upstream of the target sequence.
[1231] Example 158, second restriction enzyme cut surface
[1232] In any one of Examples 155 to 157, the second restriction enzyme cleavage site forms a blunt end as a result of cleavage by the second restriction enzyme.
[1233] Example 159, Second restriction enzyme limitation
[1234] In any one of Examples 155 to 158, the second restriction enzyme is selected from the group consisting of AluI, BalI, BfrBI, BsaAI, BsaBI, BsrBI, BtrI, Cac8I, CdiI, CviJI, CviRI, Eco47III, Eco78I, EcoICRI, EcoRV, FnuDII, FspAI, HaeI, HaeIII, Hpy8I, LpnI, MlyI, MslI, MstI, NaeI, NalIV, NruI, NspBII, OliI, PmaCI, PmeI, PshAI, PsiI, PvuII, RsaI, ScaI, SmaI, SnaBI, SrfI, SspI, SspD5I, StuI, XcaI, XmnI and ZraI; optionally, the second restriction enzyme is EcoRV.
[1235] Example 160, Relationship between the first and second restriction enzymes
[1236] In any one of Examples 155 to 159, the second restriction enzyme is different from the first restriction enzyme.
[1237] Example 161, identical constant position sequence
[1238] In any one of Examples 155 to 160, the constant position is the same for all nucleic acids in the library containing the nucleic acid.
[1239] Example 162, Location of the Invariant Position
[1240] In any one of Examples 155 to 161, the constant position is located between the second restriction enzyme cleavage site and the target sequence.
[1241] Example 163, Invariant Position Detection Conditions
[1242] In any one of Examples 155 to 162, the constant position is detectable by a nucleic acid detection method only when the target sequence is not cleaved.
[1243] Example 164, Condition for Non-detection of Invariant Positions
[1244] In any one of Examples 155 to 162, the constant position is not detectable by a nucleic acid detection method when the target sequence is cleaved.
[1245] Expanding the Invention - Nucleic Acid-Guided Nucleases
[1246] Example 165, nucleic acid-guided nuclease
[1247] In any one of Examples 1 to 164,
[1248] The above Cas protein is a nucleic acid-guided nuclease.
[1249] The above CRISPR / Cas system is a nucleic acid-guided nuclease system.
[1250] An example in which the above CRISPR / Cas complex is applied by replacing it with a nucleic acid-guided nuclease complex.
[1251] Example 166, Nucleic Acid-Guided Nuclease Example
[1252] In Example 165, the nucleic acid-guided nuclease is selected from the following:
[1253] Zinc Finger Nuclease (ZFN); Meganuclease; Transcription Activator-Like Effector Nuclease (TALEN); IscB; and TnpB.
[1254]
[1255] [Experimental Example]
[1256] Hereinafter, the invention provided by this specification will be described in more detail through experimental examples and examples. These examples are intended solely to illustrate the subject matter disclosed by this specification, and it will be apparent to those skilled in the art that the scope of the subject matter disclosed by this specification is not limited by these examples.
[1257] Experimental Example 1. Experimental Method and Materials
[1258] Experimental Example 1.1. Oligonucleotide Library Design
[1259] 2,000 (T7_2k), 120,802 (T7_120k or U6_120k), 6,000 (U6_6k), 30,772 (U6_30k, Cut1_30k or Cut2_30k), 42,000 (Cut2_42k) sgRNA and corresponding target sequence pairs, 14,574 (7.8k_WT), 15,672 (7.8k_MT), 600 (600_WT), 600 (600_MT) targets, 2,386 (optimized_HF1_2k), 2,290 (optimized_NRRH_2k), 2,112 (perfectly-matched_2k), 200 or 800 (for rare variant sequence cleavage) The sgRNA sequence pool (sgRNA library) was synthesized by Twist Bioscience (San Francisco, CA).
[1260] Each oligonucleotide, T7_120k, Cut1_30k, Cut2_30k, and Cut2_42k, consists of the following components: part of a 23-nt scaffold, a 20-nt guide sequence (GN19), an 18-nt T7 promoter, BsmBI restriction site #1, an 8-nt random sequence, BsmB1 restriction site #2, a 20-nt barcode, a 35-nt wide targeting sequence including a PAM, and a 20-nt linker.
[1261] Oligonucleotides designed for U6_120k, U6_6k, and U6_30k contain the following elements: 20-nt linker 1, 20-nt guide sequence (GN19), BsmB1 restriction site #1, 20-nt random sequence, BsmB1 restriction site #2, 20-nt barcode, 65-nt wide target sequence including PAM, and 20-nt linker 2.
[1262] Each oligonucleotide of 7.8k_WT, 7.8k_MT, 600_WT, and 600_MT contains a 120-nt wide target sequence and a 20-nt barcode. Oligonucleotides of optimized_HF1_2k, optimized_NRRH_2k, perfectly-matched_2k, and the sgRNA library for rare variant sequence cleavage contain an 8-nt linker 1, an 18-nt T7 promoter, a 20-nt guide sequence, a 76-nt scaffold, and a 20-nt linker 2.
[1263] Experimental Example 1.2. T7_2k Design
[1264] To validate the Cut-seq1 analysis, we designed T7_2k, which consists of 652 sgRNAs and 641 target sequences. Validation was performed with the following libraries:
[1265] (1) Positive control consisting of 206 sgRNA and target sequence pairs.
[1266] Matching sgRNA and target sequence pairs were selected based on the indel frequency measured in a previous study, and sgRNA and target sequence pairs with indel frequencies exceeding 10% were selected at a ratio of 2-3 pairs per 1% range;
[1267] (2) Negative control consisting of 100 sgRNA and target sequence pairs.
[1268] sgRNAs were selected from off-target sgRNAs with low DeepSpCas9 (Kim HK, Kim Y, Lee S, Min S, Bae JY, Choi JW, Park J, Jung D, Yoon S, Kim HH. SpCas9 activity prediction by DeepSpCas9, a deep learning-based model with high generalization performance. Sci Adv. 2019 Nov 6;5(11):eaax9249. doi: 10.1126 / sciadv.aax9249. PMID: 31723604; PMCID: PMC6834390.) scores in the GeCKO v2 library. The targets were designed to be completely mismatched with the sgRNA and also included the NTT PAM sequence;
[1269] (3) Target sites consisting of 504 sgRNA (seed sgRNA) and target sequence pairs collected from other studies.
[1270] sgRNA and target sequence pairs were collected from Guide-seq, Discover-seq, Circle-seq, SITE-seq, TTISS, ONE-seq, MOFF, and CHANGE-seq.
[1271] For each seed sgRNA in library (3), 20 targets were designed as follows:
[1272] (a) Matching targets (4);
[1273] (b) Target sequences with mismatches (7 targets): 3 targets with 1-nt mismatches, 3 targets with 2-nt mismatches, 1 target with 3-nt mismatches;
[1274] (c) Target sequences with 1-nt insertions or 1-nt deletions (forming DNA or RNA bulges, respectively, in paired structures) (two targets): one insertion target and one deletion target;
[1275] (d) Target sequences detected in other studies (4 targets);
[1276] (e) Negative control (three targets): the targets are designed to be completely mismatched with the sgRNA; and
[1277] (4) 1,190 cancer-related sgRNA-target sequence pairs, comprising a total of 325 types of sgRNAs and 48 types of targets. The same sgRNA was designed to pair with either a wild-type (WT) or mutant (MT) target. For each MT target, three perfectly matched sgRNAs were designed, and for each perfectly matched sgRNA: 12 randomly selected sgRNAs with 1-nt mismatches; two randomly selected sgRNAs with 1-nt RNA bulges; and two randomly selected sgRNAs with 1-nt DNA bulges.
[1278] Experimental Example 1.3. T7_120k, U6_120k, U6_6k Design
[1279] To evaluate target-guided activity in vitro and in cultured cells (HEK293T, HeLa, Hepa 1-6, or B16-F10 cells), T7_120k was designed to measure sgRNA activity in vitro, U6_120k was designed to evaluate sgRNA activity in HEK293T cells, and U6_6k was designed to evaluate sgRNA activity in cultured cells. T7_120k and U6_120k shared a common pool of 120,802 sgRNA-target sequence pairs, comprising 40,994 types of sgRNAs and 111,862 types of targets. T7_120k and U6_120k consisted of seven main types of libraries:
[1280] (1) Positive control (201 sgRNA and target sequence pairs, consisting of 67 sgRNA sequences and 67 targets x3): This set contains pairs of sgRNA sequences and matching targets, and these sgRNAs induced high efficiency cleavage at their matching target sites in vitro.
[1281] (2) Negative control (consisting of 100 guide-target pairs, 100 sgRNAs, and 100 targets): sgRNAs were selected from non-targeting sgRNAs with low DeepSpCas9 scores in the GeCKO v2 library. The targets were designed to be completely mismatched with the sgRNAs and also included the NTT PAM sequence.
[1282] (3) Off-target sequences identified in previous studies (30,608 target-guide pairs, consisting of 166 sgRNAs and 29,778 targets): sgRNA and target sequences were taken from previous studies, including Guide-seq, Discover-seq, Circle-seq, SITE-seq, TTISS, ONE-seq, MOFF, and CHANGE-seq. For each sgRNA, the targets were designed as follows:
[1283] (a) Matched target (5x); (b) Off-target sequences: all sites detected in the previous studies described above, except for CHANGE-seq, 20 targets were randomly selected; (c) Off-target sequences not detected in other studies: these sequences were obtained using Cas-OFFinder and included five randomly selected sites with 1-, 2-, 3-, 4-, 5-, or 6-nt mismatches; and (d) Negative controls (3x): targets with 20-nt mismatches to the matched target sites and the NTT PAM sequence.
[1284] (4) Matched and mismatched targets associated with three seed sgRNAs targeting VEGFA, FANCF, or EMX1 (5,867 sgRNA and target sequence pairs, consisting of three sgRNAs and 5,809 targets): For the seed sgRNAs, the targets were designed as follows:
[1285] (a) Matched targets (5x); (b) all off-target sites detected in other studies; (c) off-target sites not detected in other studies: targets were obtained from Cas-OFFinder and contained 1.5x the number of off-target sequences detected in (b); and (d) negative controls (3 targets): targets with 20-nt mismatches to the matched target sites and an NTT PAM.
[1286] (5) Matched and Mismatched Targets Targeted by Randomly Generated sgRNAs (35,884 sgRNA-target sequence pairs, 600 sgRNAs and 23,549 targets): These pairs were designed with randomly generated sgRNAs to evaluate bulge and mismatch tolerance, and PAM interference effects. For a given seed sgRNA, 60 targets were randomly generated as follows:
[1287] (a) Matched targets (5x); (b) Non-target sequences without RNA bulges or DNA bulges (24 targets): 5 randomly designed targets with 1-, 2-, 3-, or 4-nt mismatches and 1 randomly designed target with 5-, 6-, 7-, or 8-nt mismatches; (c) Non-target sequences with RNA bulges (14 targets): 1 randomly designed target with 1- or 2-nt RNA bulges, respectively, with 0-, 1-, 2-, 3-, 4-, 5-, or 6-nt mismatches; (d) Non-target sequences with DNA bulges (14 targets): 1 randomly designed target with 1- or 2-nt DNA bulges, respectively, with 0-, 1-, 2-, 3-, 4-, 5-, or 6-nt mismatches; and (e) non-target sequences with non-NGG PAM (three targets): Targets with non-NGG PAM were designed randomly.
[1288] (6) Randomly generated sgRNA-matched targets (consisting of 40,028 target-guide pairs, 40,028 sgRNAs and 40,028 targets): Consisting of randomly generated sgRNA-matched targets used to evaluate the efficiency at the target site with perfectly matched sgRNAs.
[1289] (7) PAM compatibility library (consisting of 8,114 target-guide pairs, 30 sgRNAs and 7,634 targets): Contains 30 fixed protospacers, each evaluated with all possible PAM sequences in previous studies.
[1290] U6_6k is a sublibrary derived from U6_120k. Of the 6,000 sgRNA-target sequence pairs in U6_6k, 3,000 were derived from guide-target pairs associated with the three seed sgRNAs described in (4). Of these 3,000 guide-target pairs, 40% were detected in other studies, including all off-target sites with 1- or 2-nt mismatches and randomly selected off-target sites with 3-, 4-, 5-, or 6-nt mismatches. The remaining 60% were selected using candidate off-target sites predicted by Cas-OFFinder and not reported in other studies. The remaining 3,000 were obtained from randomly generated sgRNAs and their corresponding targets described in library (5) above. The prediction scores of DeepSpCas9 at the matching target sites of all sgRNAs in the library (5) were divided into 5% intervals, and four sgRNAs were randomly selected from each interval.
[1291] Experimental Example 1.4. U6_30k, Cut1_30k, and Cut2_30k Designs
[1292] A total of 30,772 target-guide pairs were designed to select sgRNAs that effectively cleave background nucleic acids (wild-type, hereinafter referred to as WT targets) but not rare nucleic acids (mutants, hereinafter referred to as MT targets). The sgRNA of U6_30k was expressed under the control of the U6 promoter, while the sgRNAs of Cut1_30k and Cut2_30k were expressed under the control of the T7 promoter. U6_30k, Cut1_30k, and Cut2_30k contained 94 MT targets used to detect cancer using ctDNA. The targets were designed as pairs consisting of a WT target and a corresponding MT target containing a mutation found in cancer cells or patients. For each MT target, all possible targets were designed by searching for an NGG sequence within + / - 22 bp of the corresponding mutation position. In addition, MT targets that disrupt the GG PAM were designed. The same sgRNA was used for each pair of WT and MT targets. The sgRNAs were designed as follows:
[1293] (1) Three identical sgRNAs that perfectly match the WT target sequence;
[1294] (2) For each sequence of (1), sgRNAs with all possible 1-nt mismatch mutations (a total of 57 sgRNAs per perfectly matched sgRNA);
[1295] (3) For each sequence of (1), 20 randomly selected sgRNAs with 2-nt mismatches;
[1296] (4) For each sequence of (1), sgRNAs with all possible 1-nt RNA bulges (a total of 76 sgRNAs per perfectly matched sgRNA); and
[1297] (5) sgRNAs with 19 possible 1-nt DNA bulges (19 sgRNAs in total per perfectly matched sgRNA).
[1298] Duplicate items in (2), (3), (4), and (5) were removed.
[1299] A total of 36,849 target-guide pairs were designed, including: 175 sgRNAs (3 perfectly matched sgRNAs + 19 sets of 3 sgRNAs with 1-nt mismatches + 20 sgRNAs with 2-nt mismatches + 19 sets of 4 sgRNAs with 1-nt RNA bulges + 19 sgRNAs with 1-nt DNA bulges) x 213 targets (52 WT targets + 161 MT targets) = 36,849 target-guide pairs. 6,077 duplicate sgRNAs and sgRNA or target sequences containing EcoRI, SspI, or BsmBI restriction sites were excluded. Therefore, the total number of target-guide pairs is 30,772 (36,849 - 6,077).
[1300] Experimental Example 1.5. Cut2_42k Design
[1301] Cut2_42k consists of 42,000 target-guide pairs, and contains a total of 1,941 mutant sequences selected as targets according to the following criteria (hereinafter referred to as target mutations):
[1302] (1) Common target mutations with high prevalence across eight cancer types (blood, breast, colon, liver, lung, ovarian, stomach, and pancreatic). Using the TCGA database, we selected the 10 most frequently mutated genes in each cancer type. For each gene, we selected the 15 most prevalent target mutations (total 877 target mutations = 8 cancers x 10 genes per cancer x 15 mutations per gene - 323 duplicate mutations).
[1303] (2) Target mutations detected in ctDNA of early-stage cancer patients in previous studies (a total of 1,042 target mutations).
[1304] (3) Target mutations associated with anticancer drug resistance in the OncoKB database (a total of 22 target mutations).
[1305] After selecting the target mutation, the target was designed using the following strategy:
[1306] (i) Targets with NGG PAM (1,601 target mutations): WT targets were designed for all possible cases by searching for NGG in the forward and reverse sequences within + / - 22 bp of the target mutation.
[1307] (ii) Targets with non-NGG PAM (340 target mutations): If there was no NGG PAM within + / - 22 bp of the target mutation, three WT targets containing the target mutation position were randomly selected, assuming a non-NGG PAM.
[1308] (iii) Targets with mutations that disrupt the GG PAM sequence.
[1309] (iv) MT targets were designed by introducing target mutations into the previously designed WT targets in (1), (2) and (3).
[1310] Cut2_42k contained 3,874 WT targets and 8,130 MT targets. Of the 42,000 target-guide pairs, 34,712 had targets with an NGG PAM, and 7,288 had non-NGG PAMs. sgRNAs were designed using the following approach:
[1311] (1) We designed a perfectly-matched sgRNA that binds to the WT target.
[1312] (2) All possible sgRNAs for the 1-nt mismatch, 2-nt mismatch, 1-nt RNA bulge, and 1-nt DNA bulge scenarios were generated for the perfect-match sgRNA sequence in step (1).
[1313] (3) Afterwards, 4 or 5 sgRNAs (divergent sgRNAs) were randomly selected from the sgRNA set of step (2).
[1314] (4) For each target, we designed (1) a perfect-matched sgRNA; and (3) 4 to 5 divergent sgRNAs.
[1315] Experimental Example 1.6 Design of 7.8k_WT and 7.8k_MT
[1316] To implement and evaluate a method for enriching rare nucleic acids (cancer-associated mutations of the wild-type nucleic acid) by cleaving background nucleic acids (wild-type nucleic acids), two libraries were generated, 7.8k_WT (wild-type target) and 7.8k_MT (MT target).
[1317] 1. 7.8k_WT (14,574 targets): A library composed of cancer-associated wild-type targets (WT targets). 2. 7.8k_MT (15,672 targets): A library composed of cancer-associated mutant targets (MT targets).
[1318] Cancer-associated mutations included in 7.8k_MT include:
[1319] 1. 94 cancer-associated mutations used in Cut1_30k or Cut2_30k.
[1320] 2. 1,941 cancer-associated mutations used in Cut2_42k.
[1321] 3. 1,030 mutations associated with levels 1, 2, and 3 of the OncoKB database.
[1322] After removing duplicate mutations, the total number of mutations is 94 + 1,941 + 1,030 - 453 = 2,612.
[1323] The above 2,612 target mutations were selected from the TCGA database as target mutations with high prevalence in eight major cancers (blood, breast, colon, liver, lung, ovarian, pancreatic, and gastric cancers), target mutations with high prevalence in various cancers (OncoKB database and TCGA database), target mutations reported in early cancer detection in other studies, and target mutations associated with anticancer drug resistance.
[1324] For 7.8k_MT, individual target mutations were placed in the forward or reverse complement orientation on the left, middle, and right parts of the target (2,612 target mutations x 3 (left, middle, or right) x 2 (forward or reverse complement orientation) = 15,672 targets).
[1325] 7.8k_WT is designed similarly to 7.8k_MT, but this library contains the corresponding wild-type sequence instead of the target mutations. Redundant targets in 7.8k_WT were removed (2,612 target mutations x 3 (left, middle, or right) x 2 (forward or reverse complement) - 1,098 redundant targets = 14,574 targets).
[1326] The oligonucleotide length of 7.8k_WT and 7.8k_MT is 140 nt, consisting of a 120-nt target and a 20-nt barcode.
[1327] Experimental Example 1.7. Design of the optimal_HF1_2k, optimal_NRRH_2k, and perfectly-matched_2k sgRNA libraries.
[1328] We generated three sgRNA libraries: optimal_HF1_2k (containing 2,386 sequences), optimal_NRRH_2k (containing 2,290 sequences), and perfectly-matched_2k (containing 2,112 sequences). In these libraries, sgRNAs were transcribed from the T7 promoter. The optimal_HF1_2k or optimal_NRRH_2k libraries were specifically designed for MT target enrichment using SpCas9-HF1 or SpCas9-NRRH-HF1, respectively. The design approach included the following steps:
[1329] (1) Target Selection: All potential targets with an NGG PAM were identified, integrating 2,612 clinically significant target mutations from cf_MT. If no target with an NGG PAM could be designed, three targets with non-NGG PAMs were randomly selected. Additionally, targets with mutations that disrupt the NGG PAM were also considered.
[1330] (2) sgRNA generation: In (1), sgRNAs that perfectly matched the WT target and all possible divergent sgRNAs (sgRNAs with all possible 1-nt mismatches (57 sgRNAs), sgRNAs with 2-nt mismatches (1,539 sgRNAs), sgRNAs with 1-nt RNA bulges (76 sgRNAs), and sgRNAs with 1-nt DNA bulges (19 sgRNAs)) were generated. This step generated a total of 1,692 sgRNAs for each target.
[1331] (3) For each designed target, all possible sgRNAs were generated and paired with WT and MT targets.
[1332] (4) Prediction of the Cleavage Ratio Index by DeepCut: The sgRNA-target pairs from step (3) were used as input to DeepCut-HF1 or DeepCut-NRRH-HF1 to predict the cleavage ratio index. This index was determined using the following formula:
[1333]
[1334] (5) Calculation of cleavage ratio index difference: Using the predicted cleavage ratio index values, the difference in cleavage ratio indices of the corresponding sgRNAs targeting WT and MT targets was calculated.
[1335] Cleavage ratio index difference =
[1336] (Cleavage ratio index for noise (WT) sequence) - (Cleavage ratio index for rare variant (MT) sequence)
[1337] (6) Selection of optimal sgRNA: For each target mutation, the optimal sgRNA for HF1 or NRRH1 was selected (the sgRNA with the largest difference in predicted cleavage ratio index among the sgRNAs with cleavage ratio index less than 0 for the MT target).
[1338] Here, the perfectly-matched_2k library consists of only sgRNAs that perfectly match 2,612 target mutations.
[1339] Experimental Example 1.8. 600_MT and 600_WT Designs
[1340] To evaluate the enrichment of rare nucleic acids (i.e., cancer-associated mutant sequences) by cleaving them, two libraries were generated, 600_WT (WT target) and 600_MT (MT target). The 600_MT library contained 200 target mutations randomly selected from the 7.8k_MT library. For 600_MT, individual target mutations were positioned on the left, middle, and right sides of the target, resulting in a total of 600 targets (200 target mutations x 3 positions = 600 targets). The 600_WT library was designed similarly to 600_MT, except that it contained the corresponding wild-type (WT) sequence instead of the target mutations.
[1341] Experimental Example 1.9. Design of an sgRNA Library to Cut Rare Variant Sequences
[1342] A library of seven sgRNAs was designed to cleave rare nucleic acids (600_MT). In this library, sgRNAs were transcribed from the T7 promoter. For 200 target mutations in 600_MT, all possible target sequences with NGG, NAG, or NGA PAMs were generated, and perfect-matched sgRNAs corresponding to the MT targets were generated. Subsequently, all possible mutant sgRNAs, including those with 1-nt mismatches, 2-nt mismatches, 1-nt RNA bulges, or 1-nt DNA bulges, were generated along with the perfect-matched sgRNAs. All these sgRNAs were input into DeepCut along with WT or MT target pairs to calculate the predicted cleavage rate indices for the WT or MT targets.
[1343] The sgRNA library was designed as follows:
[1344] (1) One sgRNA per mutation with the highest predicted cleavage rate index on the MT target (200 target mutations x 1 sgRNA = 200 sgRNAs).
[1345] (2) 4 sgRNAs per mutation with the highest predicted cleavage rate index in the MT target (200 target mutations x 4 sgRNAs = 800 sgRNAs).
[1346] (3) 1 sgRNA per mutation with the highest predicted cleavage ratio index difference (i.e., difference between MT target and WT target) (200 target mutations x 1 sgRNA = 200 sgRNAs).
[1347] (4) 4 sgRNAs per mutation with the highest predicted cleavage ratio index difference (200 target mutations x 4 sgRNAs = 800 sgRNAs).
[1348] (5) 4 sgRNAs per mutation with the highest predicted cleavage rate index difference among those with a cleavage rate index less than or equal to 0 from the WT target (200 target mutations x 4 sgRNAs = 800 sgRNAs).
[1349] (6) 4 sgRNAs per mutation with the highest predicted value of (cleavage rate index difference) x (cleavage rate index at MT target) (200 target mutations x 4 sgRNAs = 800 sgRNAs).
[1350] (7) 4 perfectly matched sgRNAs per mutation (200 target mutations x 4 sgRNAs = 800 sgRNAs).
[1351] Experimental Example 1.10. Plasmid Library Preparation
[1352] For each library, aliquots of bacterial cell culture were plated separately to calculate library coverage. For all final plasmid libraries, coverage was 5,000-fold greater than the initial number of oligonucleotides.
[1353] (1) Preparation of T7_2k, T7_120k, Cut1_30k, Cut2_30k, and Cut2_42k plasmid libraries
[1354] The cloning protocol was adapted from a previously described procedure. In summary, the cloning protocol consisted of two steps:
[1355] Step 1: An initial plasmid library containing target-guide pairs with barcodes was constructed. Specifically, an oligonucleotide pool of target-guide pairs with barcodes and an oligonucleotide of a scaffold sequence with two BsmB1 sites were PCR amplified using Q5 High-Fidelity DNA Polymerase (New England Biolabs (NEB), Ipswich, MA). The two fragments were assembled using NEBuilder® HiFi DNA Assembly Master Mix to generate an intermediate Gibson Assembly circular plasmid. The plasmid was then linearized with a restriction enzyme (Ssp1) and amplified by PCR. Gibson Assembly was used to insert the amplified fragments into a BsmBI-cleaved plasmid backbone containing a T7 terminator to terminate sgRNA transcription. The assembled Gibson product was concentrated by isopropanol precipitation using GlycoBlue Coprecipitant (Invitrogen), and the assembled product was purified using a TransforMax MicroPulser (Bio-Rad). TM EC100 TM Transformed into electrocompetent cells (Lucigen). SOC medium was added to the transformation mixture and incubated at 37°C for 1 h. Cells were then plated on Luria-Bertani (LB) agar plates and cultured in medium containing 50 μg / ml carbenicillin. The plasmid library was extracted from the harvested colonies using the Plasmid Maxiprep kit (Qiagen).
[1356] Step 2: Unique molecular identifiers (UMIs) were inserted. Specifically, the initial plasmid library generated in Step 1 was digested with BsmBI. The digested product was size-selected on a 2% agarose gel and gel-purified. The inserted DNA fragment containing the UMI sequence with eight random nucleotides was PCR-amplified using Q5 High-Fidelity DNA Polymerase (NEB) and gel-purified on a 4% agarose gel. The purified insert was assembled with the digested initial plasmid library vector using NEBuilder® HiFi DNA Assembly Master Mix. The product was concentrated by isopropanol precipitation and electroporated into TransforMax EC100 electrocompetent cells (Lucigen). Colonies, i.e., bacterial cell libraries, were harvested and used for Cut-seq1 or Cut-seq2. The plasmid library was extracted from some colonies using the Plasmid Maxiprep kit (Qiagen) and stored for future use.
[1357] (2) Preparation of U6_120k, U6_6k, and U6_30k plasmid libraries
[1358] The cloning protocol was adapted from a previously published procedure. Briefly, the cloning protocol consisted of two steps:
[1359] Step 1: An initial plasmid library containing sgRNA and target sequence pairs was generated. Specifically, the Lenti-gRNA-Puro plasmid (#84752, Addgene) was linearized with BsmBI enzyme (NEB) and gel-purified using the MEGAquick-spin Total Fragment DNA Purification kit (iNtRON Biotechnology). The oligonucleotide pool was PCR-amplified using Q5 High-Fidelity DNA Polymerase (NEB). The amplified fragments were gel-purified on a 2% agarose gel and assembled using the BsmBI1-cleaved Lenti-gRNA-Puro plasmid and the NEBuilder HiFi DNA Assembly kit (NEB). After incubation at 50°C for 1 h, the assembled product was concentrated by isopropanol precipitation using GlycoBlue Coprecipitant (Invitrogen) and transformed into EC100 electrocompetent cells (Lucigen) using a MicroPulser electroporator (Bio-Rad). The transformed cells were inoculated onto LB agar plates supplemented with carbenicillin (50 μg / ml) and incubated at 37°C. Total colonies were harvested, and plasmids were extracted using the Plasmid Maxiprep kit (Qiagen).
[1360] Step 2: Insertion of the sgRNA scaffold. The initial plasmid library generated in Step 1 was digested with BsmBI (NEB) and treated with calf intestinal alkaline phosphatase (NEB) at 37°C for 30 minutes. The digested product was size-selected by 2% agarose gel electrophoresis and purified using the MEGAquick-spin Total Fragment DNA Purification kit (iNtRON Biotechnology). Separately, a synthetic insert fragment containing the sgRNA scaffold and a UMI with eight random nucleotides was amplified by PCR, digested with BsmBI (NEB), and gel-purified on a 2% agarose gel. Ligation reactions were performed using 40 ng of this purified insert and 100 ng of the digested initial plasmid library vector. After overnight ligation, the reaction product was heat-inactivated at 65°C for 20 min and concentrated by isopropanol precipitation. The purified product was transformed into EC100 electrocompetent cells (Lucigen) using a MicroPulser electroporator (Bio-Rad). The transformed cells were inoculated onto LB agar plates supplemented with carbenicillin (50 μg / mL) and incubated at 37°C. Colonies were harvested, and the plasmid library was extracted using the Plasmid Maxiprep kit (Qiagen).
[1361] (3) sgRNA libraries for cleavage of 7.8k_WT, 7.8k_MT, 600_WT, 600_MT, optimized_HF1_2k, optimized_NRRH_2k, perfectly-matched_2k, and rare variant sequences
[1362] Oligonucleotide pools containing the target sequence, barcode sequence, or T7 promoter, sgRNA, and scaffold sequence were amplified by PCR using Q5 High-Fidelity DNA Polymerase (NEB) and purified using a PCR purification kit (Qiagen). The purified PCR products were assembled with the BsmB1 cleavage vector backbone using the NEBuilder HiFi DNA Assembly kit (NEB). The assembled products were concentrated by isopropanol precipitation, and the concentrated products were transformed into EC100 electrocompetent cells (Lucigen) using a MicroPulser electroporator (Bio-Rad). The transformed cells were inoculated onto LB agar medium supplemented with carbenicillin (50 μg / mL) and cultured at 37°C. Colonies were harvested, and the plasmid library was extracted using a Plasmid Maxiprep kit (Qiagen).
[1363] Experimental Example 1.11. Cut-seq1
[1364] T7_2k, T7_120k, or Cut1_30k were transformed into bacteria lacking endonuclease A (EndA-) (TransforMax EC100 Electrocompetent E. coli) at a coverage of ≥5,000-fold. The bacterial cell libraries were harvested, washed with cold phosphate-buffered saline (PBS), and fixed with cold methanol for 1 min. The bacterial cells were then washed with PBS and pelleted. The restriction enzyme EcoR1 was applied to cleave the sgRNA coding sequence, allowing RNA transcription to terminate. After this step, individual Cas9 variants and T7 RNA polymerase (NEB®) were added, and the mixture was incubated at 37°C for 12 hours, followed by heat inactivation at 65°C for 20 minutes. Proteinase K (Qiagen®) and RNase (Qiagen®) were then added to degrade the Cas9 protein and the transcribed sgRNA. Double-stranded DNA was purified using a PCR purification kit (Qiagen®). Blunt-ended adapters, A*C*A*CTCTTTCCCTACACGACGCTCTTCCGATC*T*G*G (SEQ ID NO: 17) and C*C*A*GATCGGAAGAG*C*C*A (SEQ ID NO: 18), were ligated to a target containing blunt ends cleaved by the CRISPR / Cas system. The "*" in the blunt-ended adapter indicates a phosphate thioate binding site. PCR was then performed using the primers listed below that bind to the adapter sequence and the sgRNA scaffold. The primer sequences are as follows:
[1365] T7pro_NGS_RP: GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCTCTTATTTTAACTTGCTATTTCTAGCTCTAAAAC (SEQ ID NO: 19); and T7pro_NGS_FP: ACACTCTTTCCCTACACGACG (SEQ ID NO: 20).
[1366] We then performed deep sequencing to identify the sequences cleaved by Cas9 (Fig. 2).
[1367] Experimental Example 1.12. Cleavage Index and Cleavage Selectivity Index Analysis for Cut-seq1 Results
[1368] Deep-sequencing data from Experimental Example 1.11 (Cut-seq1) were analyzed using an internal Python script. The Cut-seq1 results were obtained via NovaSeq sequencing and calculated based on the number of UMI barcodes, with a coverage of 5,000x or greater. Reads with incorrect sgRNA and target sequence barcodes were filtered out. Additionally, to exclude barcodes with low coverage in the library, an unprocessed control plasmid library was deep-sequencing to achieve a coverage of 5,000x or greater. Only barcodes with read counts exceeding 50 in the control library plasmid were analyzed, and counts per million (CPM) and truncation index were calculated using the following formulas:
[1369] Cleavage Index =
[1370] Count per million (CPM) = 10 6 x ( )
[1371] The degree of differentiation in cleavage activity between WT and MT targets by sgRNA was quantified using the following cleavage selectivity index formula:
[1372] Cleavage selectivity index = log( )
[1373] Experimental Example 1.13. Cut-seq2
[1374] The procedure used for Cut-seq2 was identical to that used for Cut-seq1, except for the addition of an EcoRV step to generate blunt ends. Restriction enzyme treatment was performed prior to in vitro transcription with Cas9 variants to allow detection of uncleaved targets (Figures 19-23 and 24-27). Briefly, Cut2_30k or Cut2_42k was transformed into endonuclease A-deficient (EndA-) bacteria (TransforMax EC100 Electrocompetent E. coli) at a coverage of ≥5,000-fold. The bacterial cell library was then fixed with cold methanol. Restriction enzymes EcoR1 and EcoRV were applied to the fixed bacterial cell library to terminate sgRNA transcription and detect uncleaved targets, respectively. After heat inactivation of the restriction enzyme at 65°C for 20 min, individual Cas9 variants and T7 RNA polymerase were added to the fixed bacterial cells and incubated at 37°C for 12 h. Proteinase K and RNase were applied, and double-stranded DNA was purified using a PCR purification kit (Qiagen®).
[1375] Blunt-ended adapters (A*C*A*CTCTTTCCCTACACGACGCTCTTCCGATC*T*G*G (SEQ ID NO: 17) and C*C*A*GATCGGAAGAG*C*C*A (SEQ ID NO: 18), * represents phosphate thioate bonds) were ligated to the purified double-stranded target, and PCR was performed using primers (T7pro_NGS_RP and T7pro_NGS_FP) to prepare for deep sequencing to analyze cleaved or uncleaved targets. If a consistent sequence was detected, the target was considered uncleaved, otherwise, the target was considered cleaved.
[1376] Experimental Example 1.14. Cleavage Index and Cleavage Selectivity Index Analysis for Cut-seq2 Results
[1377] Analysis of Cut-Seq2 deep-sequencing data was performed using a custom Python script. The results were obtained using NovaSeq sequencing with a coverage of 5,000x or more. Calculations were based on the number of UMIs associated with barcodes. Barcodes associated with incorrect guides and target sequences were removed, and barcodes with fewer than 50 UMIs were excluded. Target-guide activity was expressed as a cleavage ratio index determined using the following formula:
[1378]
[1379] The degree of cleavage activity differentiation between WT and MT targets by sgRNA was quantified using the following cleavage ratio index difference formula:
[1380] Cleavage ratio index difference
[1381] = (Cleavage ...
Claims
1. A method for evaluating the activity of the CRISPR / Cas system for various combinations of guide RNA and target DNA in high-throughput. The above method comprises: (a) a process for preparing microchamber cells, comprising: (a-1) Process of preparing a cell population; (a-2) Process of treating the target-guide library to the above cell population, Here, the target-guide library includes a plurality of target-guide vectors, Each target-guide vector comprises a target DNA, a DNA encoding a guide RNA, and a promoter operably linked to the DNA encoding the guide RNA, The above target-guide library contains at least two types of target-guide vectors having different target DNA sequences, different guide RNA sequences, or both of the above sequences, The above target-guided library is processed according to a predetermined multiplicity of infection (MOI); and (a-3) A process of treating a fixed reagent to a cell population treated with the target-guide library; Here, by process (a-2), most of the cells included in the cell population prepared in (a-1) are i) cells not transfected with the target-guide vector, or ii) cells transfected with only one target-guide vector. By the above process (a-3), only the one target-guide vector is fixed in the transfected cell and functions as an independent reaction chamber. Here, fixed cells transfected with only one target-guide vector are called microchamber cells; (b) a process for inducing a compartmentalized response by microchamber cells in a cell population, comprising: (b-1) a process of treating a cell population including the above microchamber cells with a transcription enzyme; Here, the transcription enzyme is delivered to each of the microchamber cells, The DNA encoding the guide RNA contained in each microchamber cell is transcribed into guide RNA by the above transcription enzyme; and (b-2) A process of treating a Cas protein in a cell population treated with the above-mentioned transcription enzyme; Here, the Cas protein is delivered to each of the microchamber cells; Here, through the process (b), within each microchamber cell, the Cas protein and the guide RNA transcribed from each microchamber cell bind to form a CRISPR / Cas complex, and the CRISPR / Cas complex reacts with the target DNA within the microchamber cell. The reactions within each microchamber cell occur independently; and (c) a process for analyzing the reaction activity between the guide RNA and the target DNA included in each target-guide vector, including: (c-1) A process of obtaining information about a target-guide vector contained in microchamber cells in a cell population in which a compartmentalized response is induced by microchamber cells; Here, information about the target-guide vector includes the sequence of the guide RNA contained therein, the sequence of the target DNA, and whether the target DNA is cleaved; and (c-2) A process for determining the reaction activity between the guide RNA and target DNA included in each target-guide vector based on the information of (c-1) above.
2. In paragraph 1, The above process (a) further includes the following process (a-2-1) after the above process (a-2): (a-2-1) A process for expanding a cell population treated with the above target-guided library; The above process (a-3) is a method in which a fixing reagent is treated to a cell population expanded through the above process (a-2-1).
3. A population of fixed cells containing microchamber cells, Here, each microchamber cell contains a Cas protein, a guide nucleic acid, and a target nucleic acid, Each Cas protein, guide nucleic acid, and target nucleic acid is an exogenous construct, Each microchamber cell contains only one type of guide nucleic acid and only one type of target nucleic acid, The population of said fixed cells comprises at least two different types of microchamber cells, wherein said different microchamber cells are cells in which the sequence of the guide nucleic acid, the sequence of the target nucleic acid, or both sequences contained therein are different, Here, the endogenous enzyme activity of each microchamber cell is inhibited, while the exogenous components remain active, allowing the microchamber cells to mimic the reaction chambers of an in vitro environment. Here, the population of fixed cells ensures compartmentalized interactions between the Cas protein, guide nucleic acid, and target nucleic acid within each microchamber cell.
4. In paragraph 3, Each of the above microchamber cells contains guide RNA as a guide nucleic acid and target DNA as a target nucleic acid, Each of the above microchamber cells comprises a vector, or a fragment thereof, comprising a nucleic acid encoding the target DNA and the guide RNA, The above vector has the structure of [Structural Formula 1] or [Structural Formula 2]: [Structural formula 1] [Target] - [Barcode] - [Promoter] - [Guide] - [RS 1 ]; or [Structural formula 2] [RS 2 ] - [Constant] - [Target] - [Barcode] - [Promoter] - [Guide] - [RS 1 ] Here, the Target is the target DNA, The above Barcode is a barcode that 1) encrypts target-guide sequence information, 2) encrypts unique identification information of the target-guide vector, or 3) encrypts both 1) and 2). The above Guide is a DNA encoding the above guide RNA, The above promoter is operably linked to DNA encoding the guide RNA, RS above 1 is the first restriction enzyme site, The above Constant is an invariant sequence, RS above 2 is the second restriction enzyme site.
5. Method for learning an active prediction model of a CRISPR / Cas system including: (a) A process for preparing learning data, including: (a-1) Process of obtaining raw data; Here, the raw data includes a plurality of raw data points, Each of the above raw data points comprises 1) a target nucleic acid sequence, 2) a guide nucleic acid sequence, and 3) cleavage activity. The above cleavage activity is a value measuring the degree to which a CRISPR / Cas complex including the guide nucleic acid sequence cleaves a target nucleic acid sequence; (a-2) A process of generating sorted data by augmenting the above raw data; Here, the alignment data includes a plurality of alignment data points, Applying one or more alignment algorithms to each of the above raw data points to generate one or more aligned data points per raw data point, Each of the above alignment data points comprises 1) an aligned target nucleic acid sequence, 2) an aligned guide nucleic acid sequence, and 3) the cleavage activity, The aligned target nucleic acid sequence and the aligned guide nucleic acid sequence are the results of aligning the target nucleic acid sequence and the guide nucleic acid sequence included in the raw data points according to an alignment algorithm, and include gap information; and (a-3) A process of generating learning data by processing the above sorted data, Here, the learning data includes a plurality of learning data points, For each training data point, for each alignment data point, Using the above aligned target nucleic acid sequence and the above aligned guide nucleic acid sequence as input values, Label the above cutting activity as an output value, Converting the aligned target nucleic acid sequence and the aligned guide nucleic acid sequence into a form suitable for learning a prediction model; and (b) A process of learning the above prediction model, performed using the above learning data; Here, the prediction model is configured to receive an aligned target nucleic acid sequence and an aligned guide nucleic acid sequence as input and output a predicted cleavage activity.
6. Method for learning a CRISPR / Cas system activity prediction model, including: The above method comprises: (a) A process for preparing learning data, including: (a-1) Process of obtaining raw data; Here, the raw data includes a plurality of raw data points, Each of the above raw data points comprises 1) a target nucleic acid sequence, 2) a guide nucleic acid sequence, and 3) cleavage activity. The above cleavage activity is a value measuring the degree to which a CRISPR / Cas complex including the guide nucleic acid sequence cleaves a target nucleic acid sequence; (a-2) A process of generating sorted data by augmenting the above raw data; Here, the alignment data includes a plurality of alignment data points, Applying one or more alignment algorithms to each of the above raw data points to generate one or more aligned data points per raw data point, Each of the above alignment data points comprises 1) an aligned target nucleic acid sequence, 2) an aligned guide nucleic acid sequence, 3) mismatch information, 4) optionally, additional input variables, and 5) truncation activity of the raw data point, The aligned target nucleic acid sequence and aligned guide nucleic acid sequence are the results of aligning the target nucleic acid sequence and guide nucleic acid sequence included in the raw data points according to an alignment algorithm, and include gap information. The above mismatch information includes information for each base position of the aligned target nucleic acid sequence and the aligned guide nucleic acid sequence, expressed as one of the following: bases of the two nucleic acids match; bases of the two nucleic acids do not match; a gap occurs in the target nucleic acid; or a gap occurs in the guide nucleic acid. The above additional input variables include one or more of the following information determined from the aligned target nucleic acid sequence and the aligned guide nucleic acid sequence: Melting Point (T) of the target nucleic acid-guide nucleic acid binding m ); the minimum free energy (MFE) of target nucleic acid-guide nucleic acid binding; and the free energy change of target nucleic acid-guide nucleic acid binding; and (a-3) A process of generating learning data by processing the above sorted data, Here, the learning data includes a plurality of learning data points, For each training data point, for each alignment data point, Using the aligned target nucleic acid sequence, the aligned guide nucleic acid sequence, the mismatch information, and optionally the additional input variables as input values, Label the above cutting activity as an output value, Converting the aligned target nucleic acid sequence, the aligned guide nucleic acid sequence, the mismatch information, and the additional input variables into a form suitable for learning a prediction model; and (b) A process of learning the above prediction model, performed using the above learning data; Here, the prediction model is configured to receive an aligned target nucleic acid sequence, an aligned guide nucleic acid sequence, the mismatch information, and optionally the additional variables as input, and output a predicted cleavage activity.
7. Method for predicting cleavage activity of a CRISPR / Cas system, including: (a) a process of obtaining a target nucleic acid sequence and a guide nucleic acid sequence; (b) a process of deriving one or more input variables from the target nucleic acid sequence and the guide nucleic acid sequence: Here, each of the above input variables includes: (i) an aligned target nucleic acid sequence; and (ii) aligned guide nucleic acid sequence; Here, different input variables include different aligned target nucleic acid sequences and aligned guide nucleic acid sequence information; (c) a process of predicting the cleavage activity of the CRISPR / Cas complex including the guide nucleic acid toward the target nucleic acid by utilizing the activity prediction model; Here, the activity prediction model is a model learned to predict the cleavage activity of the CRISPR / Cas system by inputting an aligned target nucleic acid sequence and an aligned guide nucleic acid sequence. By inputting one or more of the input variables into the above active prediction model, one predicted cut-off activation value is calculated for each input variable, If the predicted cleavage activity value is one, the predicted cleavage activity value is used as the prediction result, If there are two or more predicted cleavage activity values, the predicted cleavage activity values are weighted and averaged with a predetermined weight and used as the prediction result.
8. Method for predicting cleavage activity of a CRISPR / Cas system, including: (a) a process of obtaining a target nucleic acid sequence and a guide nucleic acid sequence; (b) a process of deriving one or more input variables from the target nucleic acid sequence and the guide nucleic acid sequence: Here, each of the above input variables includes: (i) an aligned target nucleic acid sequence; (ii) aligned guide nucleic acid sequence; (iii) inconsistent information; The above mismatch information includes information for each base position of the aligned target nucleic acid sequence and the aligned guide nucleic acid sequence, expressed as one of the following: bases of the two nucleic acids match; bases of the two nucleic acids do not match; a gap occurs in the target nucleic acid; or a gap occurs in the guide nucleic acid; and (iv) additional input variables, The above additional input variables include one or more of the following information determined from the aligned target nucleic acid sequence and the aligned guide nucleic acid sequence: Melting Point (T) of the target nucleic acid-guide nucleic acid binding m ); minimum free energy (MFE) of target nucleic acid-guide nucleic acid binding; and free energy change of target nucleic acid-guide nucleic acid binding; Here, different input variables include different aligned target nucleic acid sequences and aligned guide nucleic acid sequence information; (c) a process of predicting the cleavage activity of the CRISPR / Cas complex including the guide nucleic acid toward the target nucleic acid by utilizing the activity prediction model; Here, the activity prediction model is a model trained to predict the cleavage activity of the CRISPR / Cas system by receiving an aligned target nucleic acid sequence, an aligned guide nucleic acid sequence, mismatch information, and additional input variables. By inputting one or more of the input variables into the above active prediction model, one predicted cut-off activation value is calculated for each input variable, If the predicted cleavage activity value is one, the predicted cleavage activity value is used as the prediction result, If there are two or more predicted cleavage activity values, the predicted cleavage activity values are weighted and averaged with a predetermined weight and used as the prediction result.
9. Method for selecting guide nucleic acids of a CRISPR / Cas system having high selective cleavage activity, including: (a) a process of determining a first target nucleic acid and a second target nucleic acid; (b) a process for deriving multiple guide nucleic acid candidates; Here, the plurality of guide nucleic acid candidates include a guide nucleic acid having a sequence that perfectly matches the first target nucleic acid, At least one sequence comprising the following modifications based on a sequence that perfectly corresponds to the first target nucleic acid: (1) Changing any one base to another; (2) Changing any two bases to other bases; (4) Removing any one base from the sequence; (5) adding one random base at any position in the sequence; or (6) Any combination of the modifications (1) to (5) above; (c) a process of obtaining a predicted value of the first target nucleic acid cleavage activity for each of the guide nucleic acid candidates according to the method of clause 7 or 8; (d) a process of obtaining a predicted value of the second target nucleic acid cleavage activity for each of the guide nucleic acid candidates according to the method of clause 7 or 8; (e) for each of the above guide nucleic acid candidates, a process of deriving a selective cleavage activity for the first target nucleic acid compared to the second target nucleic acid by using the first target nucleic acid cleavage activity prediction value and the second target nucleic acid cleavage activity prediction value; and (f) A process for selecting a guide nucleic acid having high selective cleavage activity according to a predetermined criterion by using the selective cleavage activity value for the first target nucleic acid compared to the second target nucleic acid.
Citation Information
Patent Citations
Evaluation and improvement of nuclease cleavage specificity
EP3613852A2
Moving robot, system of moving robot and method for moving to charging station of moving robot
KR1020200018216A
Capacitor module and converter comprising the same
KR1020250026606A
Methods, models, systems, and apparatus for identifying target sequences for CAS enzymes or crispr-CAS systems for target sequences and conveying results thereof
US20150356239A1
KR20190048926A