Predictive models of gRNA HDR potential based on indel profiles
A method using Cas enzyme editing and HDR prediction models with indel profile analysis improves HDR rates by selecting gRNAs with higher HDR potential, addressing the challenge of limited HDR frequency in CRISPR-Cas9 editing.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-15
- Publication Date
- 2026-04-02
AI Technical Summary
The limited frequency of homologous recombination repair (HDR) in CRISPR-Cas9 genome editing makes it difficult to achieve high HDR rates, necessitating a method to predict and rank the HDR potential of guide RNAs (gRNAs) for improved editing outcomes.
A method involving Cas enzyme editing experiments, empirical and in silico indel profile generation, and analysis using an HDR prediction model to output HDR rate thresholds and rankings for preferred gRNAs and editing sites, incorporating machine learning processes to account for multidimensional indel profiles and cell-specific factors.
Enhances HDR rates by selecting gRNAs with higher HDR potential, reducing screening requirements and improving editing outcomes across various cell types, including iPSCs, through a comprehensive indel profile analysis and cell-line-specific predictions.
Smart Images

Figure 2026510402000001_ABST
Abstract
Description
[Technical Field]
[0001] Cross-reference of related applications This application claims priority to U.S. Provisional Patent Application No. 63 / 490,977, filed on 17 March 2023, which is incorporated in its entirety by reference.
[0002] Reference to sequence listings This application was filed in accordance with 37 CFR §1.831 and PCT Rule 13ter, along with a sequence listing XML in ST.26 XML format. The sequence listing XML file "013670-0019-WO01_sequence_listing_xml_7-MAR-2024.xml", filed with the USPTO Patent Center, was created on March 7, 2024, contains 1212 sequences, has a file size of 1.05 Mbytes, and is incorporated herein by reference in its entirety. [Background technology]
[0003] The CRISPR-Cas9 system is widely used to perform site-directed genome editing in eukaryotic cells. Sequence-specific guide RNA is required to recruit the Cas9 protein to the target site, and the Cas9 endonuclease cleaves both strands of the target DNA to create a double-strand break (DSB). This DSB is corrected by cell-specific DNA damage repair pathways. The main pathways for DSB repair are the error-prone non-homologous end joining (NHEJ) pathway, the alternative microhomology inter-end joining (MMEJ) pathway, and the homologous recombination repair (HDR) pathway. The dominant and rapid NHEJ pathway results in either correct repair that restores the Cas9 target site (thus allowing Cas9 to re-cleave) or a small insertion or deletion (indel) event within the target DNA. The MMEJ pathway, which relies on a short microhomology sequence at the cleavage site, typically results in a larger deletion event. The combination of NHEJ and MMEJ repair events creates a unique indel profile consistent with a given Cas9 guide RNA (gRNA) and cell type. In contrast, the HDR pathway relies on homologous DNA templates (typically sister chromatids in nature) to accurately repair double-stranded breaks (DSBs). The HDR pathway is frequently used in combination with CRISPR Cas9 to generate specific desirable mutations in target DNA. To do so, artificial repair templates for HDR are provided, containing target mutant DNA sequences that are either single-stranded or double-stranded DNA and have homologous regions on either side of the DSB. However, the limited frequency of repairs by the HDR pathway makes it difficult to achieve high HDR rates in this CRISPR application.
[0004] HDR results can be improved by selecting gRNAs with higher HDR potential, i.e., by selecting gRNAs with a higher frequency of MMEJ-based edits (i.e., large deletions) in the indel profile. [Prior art documents] [Non-patent literature]
[0005] [Non-Patent Document 1] Allen et al., Nature Biotechnol. 37: pp. 64-72 (2019) [Non-Patent Document 2] Smith and Waterman, Advances in Applied Mathematics 2: 4, pp. 82-489 (1981) [Non-Patent Document 3] Needleman and Wunsch, J. Mol. Biol. 48 (3): pp. 443-453 (1970) [Non-Patent Document 4] Tatiossian et al., Mol. Ther. 29(3): pp. 1057-1069 (2021) [Non-Patent Document 5] Moeller et al., Nature Commun. 13(1): 4550 pages (2022) [Non-Patent Document 6] Bodai et al., Nature Commun. 13(1): 2351 pages (2022) [Non-Patent Document 7] Kurgan et al., Mol. Ther. - Methods Clin. Dev. 21: pp. 478-491 (2021) [Overview of the project] [Problems that the invention aims to solve]
[0006] What is needed is a method to predict HDR results and rank the HDR potential of gRNAs. [Means for solving the problem]
[0007] One embodiment described herein is a method for predicting the homologous recombination repair (HDR) potential of one or more Cas guide RNAs (gRNAs), comprising: (a) with respect to one or more candidate gRNAs, (i) performing one or more Cas enzyme editing experiments using one or more candidate gRNAs to obtain edited genomic DNA; (ii) for each editing experiment, amplifying and sequencing the edited genomic DNA to produce sequenced, edited genomic DNA; for each editing experiment, generating an empirical indel profile by performing, in a processor, (iii) receiving the sequenced, edited genomic DNA; and (iv) analyzing the sequenced, edited genomic DNA and outputting an empirical indel profile; (b) inputting the empirical indel profile from step (a) into an HDR prediction model and analyzing the indel profile; and (c) outputting an HDR rate threshold, HDR score, or ranking list of candidate gRNAs indicating preferred candidate gRNAs for HDR editing experiments, and the optimal editing sites.
[0008] Another embodiment described herein is a method for predicting the homologous recombination repair (HDR) potential of one or more Cas guide RNAs (gRNAs), comprising the steps of: (a) generating in silica indel profiles for one or more candidate gRNAs by performing the following in a processor: (i) inputting candidate gRNA sequences and editing loci; and (ii) receiving in silica indel profiles; (b) inputting the in silica indel profiles from step (a) into an HDR prediction model and analyzing the indel profiles; and (c) outputting an HDR rate threshold, HDR score, or ranking list of candidate gRNAs indicating preferred candidate gRNAs for HDR editing experiments, and optimal editing sites.
[0009] Another embodiment described herein is a method for predicting the homologous recombination repair (HDR) potential of one or more Cas guide RNAs (gRNAs), comprising: (a) with respect to one or more candidate gRNAs, (i) performing one or more Cas enzyme editing experiments using one or more candidate gRNAs to obtain edited genomic DNA; (ii) with respect to each editing experiment, amplifying and sequencing the edited genomic DNA to produce sequenced, edited genomic DNA; (iii) receiving the sequenced, edited genomic DNA in a processor with respect to each editing experiment; and (iv) analyzing the sequenced, edited genomic DNA and outputting an empirical indel profile. The method comprises the steps of: (a) generating an empirical indel profile by performing the following: (b) inputting candidate gRNA sequences and editing loci into a processor; and (ii) receiving an in silica indel profile; (c) inputting the empirical indel profile from step (a) or the in silica indel profile from step (b) into an HDR prediction model and analyzing the indel profile; and (d) outputting an HDR rate threshold, HDR score, or ranking list of candidate gRNAs that indicate a candidate gRNA preferred for HDR editing experiments, and the optimal editing site.
[0010] In one embodiment, step (a)(ii) includes amplifying genomic DNA using RNase H-dependent PCR (rhPCR) and performing next-generation sequencing (NGS) to generate sequenced, edited genomic DNA. In another embodiment, the analysis of the sequenced, edited genomic DNA in step (a)(iv) includes merging the sequenced, edited genomic DNA, binning the merged, sequenced, edited genomic DNA by alignment to the genome, and providing alignment of the edited genomic DNA and characterization and quantification of empirical indel frequencies. In another embodiment, the analysis is performed using the rhAmpSeq CRISPR analysis system or CRISPAltRations. In another embodiment, the empirical indel profile includes one or more of the following: allele frequencies, template insertion frequencies, microhomology-mediated end-joining (MMEJ) deletion frequencies, entropy, insertion size frequencies, GC insertion motif frequencies, deletion size frequencies, or a combination thereof. In another embodiment, the step of generating an in silica indel profile includes predicting the effectiveness of a guide RNA, generating alignments and edit frequencies, and mutation results resulting from double-strand breaks. In another embodiment, the input is a guide sequence, and the output is a set of alignments and predictions of on-target base editing effectiveness. In another embodiment, the step of generating an in silica indel profile is performed using FORECasT. In another embodiment, the HDR prediction model in the step includes gradient-boosted regressors, ensemble methods, Lasso regression, structural equation modeling (SEM), or conventional machine learning processes that translate multidimensional indel profiles into HDR rate thresholds, HDR scores, or ranked outputs for candidate gRNAs.In another embodiment, the HDR prediction model is trained by the processor by (i) creating a training dataset using empirical indel profiles or in silica indel profiles; (ii) creating a test dataset using empirical indel profiles or in silica indel profiles; and (iii) training and testing the HDR prediction model, wherein the HDR prediction model is trained using the training dataset and the HDR prediction model is tested using the test dataset. In another embodiment, the HDR prediction model is capable of accurately ranking candidate gRNAs of overall HDR potential whose Spearman correlation is greater than 0.5. In another embodiment, the HDR rate and preferred candidate gRNAs are specific to a particular cell type or cell line. In another embodiment, the candidate gRNA sequence has a variable region of about 17 to about 24 nucleotides in length. In another embodiment, the candidate gRNA sequence has a variable region of about 20 nucleotides in length. In another embodiment, the candidate gRNA sequence includes one or more modifications to the 5' end, the 3' end, or a combination thereof. In another embodiment, the modification includes terminal blocking modification. In another embodiment, the editing site or editing locus is Cas enzyme-specific and contains about 1 to about 15 nucleotides. In another embodiment, the Cas enzyme is Cas9 or Cas12a. In another embodiment, the genomic DNA is derived from a cell or a population of interest. In another embodiment, the candidate gRNA sequence contains a sequence derived from one or more of SEQ ID NOs: 1-255 or 1021-1068. [Brief explanation of the drawing]
[0011] [Figure 1] A block diagram illustrating an exemplary system for predicting the homologous recombination repair (HDR) potential of one or more Cas guide RNAs (gRNAs) is shown according to various aspects of this disclosure. [Figure 2] A flowchart illustrating an exemplary process for predicting the homologous recombination repair (HDR) potential of one or more Cas guide RNAs (gRNAs) is provided according to various aspects of this disclosure. [Figure 3A] A graph showing the correlation between the HDR editing frequency and indel profile attributes of the control sample with only RNP in HAP1 cells. Top AF. Sites with N = 150. [Figure 3B] A graph showing the correlation between the HDR editing frequency and indel profile attributes of the control sample with only RNP in HAP1 cells. Entropy. Sites with N = 150. [Figure 3C] A graph showing the correlation between the HDR editing frequency and indel profile attributes of the control sample with only RNP in HAP1 cells. Deletion 3+. Sites with N = 150. [Figure 4] A graph showing the performance of the gradient boosting regression HDR prediction model for test data based on empirical indel profile data in HAP1 cells (n = 150). For all modeling, a 75 / 25 training-test split was performed. The data shown in the graph here are representative samples (Pearson R2 = 0.55). 100 bootstraps were performed with a specific training / test split to determine a more generalized evaluation criterion for this model. Pearson R2 = 0.45 ± 0.13, Spearman correlation = 0.67 ± 0.09. [Figure 5] A graph showing the performance of the HAP1 HDR prediction model using indel profiles and HDR data generated in Jurkat cells. No correlation was observed between the predicted HDR and the measured HDR. Sites with N = 188 after filtering. [Figure 6A] A graph showing the evaluation of Jurkat-specific repair factors and their potential impact on this NHEJ repair profile. A box plot showing that high expression (transcripts per million; TPM) of the DNTT gene encoding terminal deoxynucleotidyl transferase is observed compared to other commonly used experimental cell lines in the publicly available data stored in the Genotype-Tissue Expression database (GTEx v8). [Figure 6B]A graph showing the evaluation of Jurkat-specific repair factors and their potential impact on this NHEJ repair profile. Investigation of the Jurkat Cas9 indel profile at the same locus within the original HAP1 dataset has demonstrated that insertions of 2+bp are enriched, and this activity may be characteristic of a template-independent polymerase that adds nucleotides during repair. [Figure 7A] A graph showing the correlation between HDR editing frequency and indel profile attributes in the RNP-only control sample in K562 cells. Top AF. Sites with N = 40, filtered by more than 70% editing. [Figure 7B] A graph showing the correlation between HDR editing frequency and indel profile attributes in the RNP-only control sample in K562 cells. Entropy. Sites with N = 40, filtered by more than 70% editing. [Figure 7C] A graph showing the correlation between HDR editing frequency and indel profile attributes in the RNP-only control sample in K562 cells. Deletion 3+. Sites with N = 40, filtered by more than 70% editing. [Figure 8A] A graph showing the comparison of editing results at the target site in K562 cells and HAP1 cells. The evaluated attribute was complete HDR editing. Sites with N = 40, filtered by more than 70% editing. [Figure 8B] A graph showing the comparison of editing results at the target site in K562 cells and HAP1 cells. The evaluated attribute was entropy (RNP-only indel profile). Sites with N = 40, filtered by more than 70% editing. [Figure 8C] A graph showing the comparison of editing results at the target site in K562 cells and HAP1 cells. The evaluated attribute was top AF (RNP-only indel profile). Sites with N = 40, filtered by more than 70% editing. [Figure 8D]This graph compares the editing results of target sites in K562 cells and HAP1 cells. The attribute evaluated was a 3+bp deletion (RNP-only indel profile). N=40 sites were included, filtered for editing exceeding 70%. [Figure 9A] This graph shows the correlation between the frequency of HDR editing in control samples of RNPs only in iPSCs and indel profile attributes: top AF. N=40 sites, filtered for editing exceeding 70%. [Figure 9B] This graph shows the correlation between the frequency of HDR editing in control samples of RNPs only in iPSCs and indel profile attributes: entropy. N=40 sites, filtered for editing exceeding 70%. [Figure 9C] This graph shows the correlation between the frequency of HDR editing in control samples of RNPs only in iPSCs and indel profile attributes: deletion 3+. N=40 sites, filtered for editing exceeding 70%. [Figure 10A] This graph compares the editing results of target sites in iPSC cells and HAP1 cells. The attribute evaluated was complete HDR editing. N=40 sites were filtered with editing and sequencing read depths exceeding 70%. [Figure 10B] This graph compares the editing results of target sites in iPSC cells and HAP1 cells. The attribute evaluated was entropy (indel profile of RNP only). N=40 sites were included, filtered by editing and sequencing read depths exceeding 70%. [Figure 10C] This graph compares the editing results of target sites in iPSC cells and HAP1 cells. The evaluated attribute was top AF (indel profile of RNP only). N=40 sites were included, filtered by edit and sequencing read depths exceeding 70%. [Figure 10D]This graph compares the editing results of target sites in iPSC cells and HAP1 cells. The attribute evaluated was a 3+bp deletion (RNP-only indel profile). N=40 sites were included, filtered by editing and sequencing read depths exceeding 70%. [Figure 11A] This graph shows a comparison of target site editing results in K562 cells, iPSCs, and primary T cells. It represents the top AF (Abstract Fibre) sites. N=40 sites are filtered for editing exceeding 70%. [Figure 11B] This graph shows a comparison of target site editing results in K562 cells, iPSCs, and primary T cells. Entropy is shown. N=40 sites are filtered with over 70% editing. [Figure 11C] This graph shows a comparison of target site editing results in K562 cells, iPSCs, and primary T cells. Deletion 3+. N=40 sites, filtered for editing exceeding 70%. [Figure 12A] This graph compares the editing results of target sites in K562 cells, iPSCs, and primary T cells. The attribute evaluated was complete HDR editing. N=40 sites were filtered with editing and sequencing read depths exceeding 70%. [Figure 12B] This graph compares the editing results of target sites in K562 cells, iPSCs, and primary T cells. The attribute evaluated was entropy (indel profile of RNPs only). N=40 sites were included, filtered by editing and sequencing read depths exceeding 70%. [Figure 12C] This graph compares the editing results of target sites in K562 cells, iPSCs, and primary T cells. The attribute evaluated was high-level AF (indel profile of RNP only). N=40 sites were included, filtered by edit and sequencing read depths exceeding 70%. [Figure 12D]This graph compares the editing results of target sites in K562 cells, iPSCs, and primary T cells. The attribute evaluated was a 3+bp deletion (RNP-only indel profile). N=40 sites were included, filtered by editing and sequencing read depths exceeding 70%. [Figure 13A] This graph shows the performance of the HAP1 HDR prediction model using indel profiles and HDR data generated from K562 cells (N=36 data points after filtering). [Figure 13B] This graph shows the performance of the HAP1 HDR prediction model using indel profiles and HDR data generated from iPSC (filtered data points, N=76). [Figure 13C] This graph shows the performance of a HAP1 HDR prediction model using indel profiles and HDR data generated from primary T cells (N=45 data points after filtering). [Figure 14A] This graph shows the performance of the HAP1 HDR prediction model using 3+0eIFreq and HDR data generated from K562 cells (N=36 data points after filtering). [Figure 14B] This graph shows the performance of the HAP1 HDR prediction model using 3+0eIFreq and HDR data generated from iPSC (filtered data points of N=76). [Figure 14C] This graph shows the performance of the HAP1 HDR prediction model using 3+0eIFreq and HDR data generated from primary T cells (filtered data points N=45). [Modes for carrying out the invention]
[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those skilled in the art. For example, all nomenclature and techniques used in relation to biochemistry, molecular biology, immunology, microbiology, genetics, cell and tissue culture, and protein and nucleic acid chemistry as described herein are well known and commonly used in the art. In case of any conflict, the disclosure, including its definitions, shall prevail. Exemplary methods and materials are described below, but similar or equivalent methods and materials may be used in the practice or testing of the embodiments and aspects described herein.
[0013] As used herein, the terms “amino acid,” “nucleotide,” “polynucleotide,” “vector,” “polypeptide,” and “protein” have the general meanings that a biochemist with the ordinary art in this field would understand. Herein, standard single-letter nucleotides (A, C, G, T, U) and standard single-letter amino acids (A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, or Y) are used. Uppercase and lowercase single letters may be used within sequences to provide structural information such as complementary regions (e.g., “acgtACGT”). Unless otherwise specified, all polypeptides are shown in the N→C-terminal direction, and all nucleotide sequences are shown in the 5'→3' direction, respectively.
[0014] As used herein, terms such as “include,” “including,” “contain,” “containing,” and “having” mean “comprising.” This disclosure also intends other embodiments “including,” “essentially consisting of,” and “consisting of” the embodiments or elements presented herein, whether expressly described herein or otherwise.
[0015] Where used herein, “a,” “an,” “the,” and similar terms as used in the context of this disclosure (in particular, in the context of the claims) should be construed to include both singular and plural unless otherwise indicated herein or unless clearly inconsistent with the context. In addition, unless otherwise specified, “a,” “an,” or “the” means “one or more.”
[0016] As used herein, the term "or" may be conjunctive or disjunctive.
[0017] As used herein, the terms "and / or" refer to both connective and disjunctive conjunctions.
[0018] As used herein, the term “substantially” means to a large or substantial extent, and not to the extent of.
[0019] As used herein, the terms “about” or “approximately,” when applied to one or more values for a given purpose, refer to a value similar to a given reference value or a value within a tolerance range for a particular value determined by those skilled in the art, which in part depends on the method of measuring or determining the value, such as limitations of the measuring system. In one embodiment, the term “about” refers to any value including both integer and decimal elements that is within a maximum variation of ±10% of the value modified by the term “about.” Alternatively, “about” may mean within three or more standard deviations, according to convention in the art. Or, with respect to biological systems or processes, the term “about” may mean within one decimal place of the value, within five times in some embodiments, and within two times in some embodiments. As used herein, the symbol “~” means “about” or “approximately.”
[0020] All ranges disclosed herein include both the endpoints as discrete values and all integers and decimals specified within that range. For example, the range 0.1 to 0.2 includes 0.1, 0.2, 0.3, 0.4…2.0. Where an endpoint is modified with the term "approximately", the specified range is extended by a variation of up to ±10% of any value within the range containing that endpoint, or a variation within a range of 3 or more standard deviations.
[0021] As used herein, the terms “control” and “reference” are used synonymously. A “reference” or “control” level may be a predetermined value or range adopted as a baseline or benchmark for evaluating measurement results. “Control” also refers to a control experiment or control cells.
[0022] This specification describes the development and testing of large HDR datasets to confirm that HDR outcomes can be improved by selecting gRNAs with higher HDR potential (i.e., gRNAs with a higher frequency of MMEJ-based edits (i.e., large deletions) in their indel profiles) and to identify additional key characteristics of indel profiles that can predict HDR outcomes. Similarly described is the development of an HDR predictive model that provides a ranking of gRNA HDR potential using empirically determined gRNA indel profiles as input. The applicability of this model across multiple cell types, including iPSCs, is then demonstrated.
[0023] The process described herein can be used to provide a ranked classification of HDR potential based on user-generated empirical data, which is particularly useful for large-scale HDR screening projects. Through the appropriate selection of gRNAs with HDR-suitable indel profiles, HDR results can be improved, and the screening requirements can be significantly reduced. The present invention is adapted for use with the rhAmpSeq CRISPR analysis system and provides a streamlined workflow for initial characterization of gRNA activity and HDR potential, as well as downstream analysis of HDR experiments. In future iterations, this HDR prediction model can be run in conjunction with an indel profile prediction tool to eliminate the requirement for pre-generated indel profile data. In addition, future iterations can incorporate cell-specific information regarding DNA repair pathway expression (e.g., based on RNA-Seq data) to provide tunable cell-line-specific predictions.
[0024] The process described herein selects superior gRNAs for HDR that are more reliable than solutions proposed in the prior art. This HDR prediction model incorporates more comprehensive indel profile attributes, resulting in improved performance beyond the "MMEJ-based deletion frequency" described in the prior art. Furthermore, while the single-factor model of the prior art cannot maintain adjustments in a cell line-independent manner, the multi-factor approach described in the present invention enables cell line-specific prediction based on a larger indel profile.
[0025] One embodiment described herein is a computer implementation process for predicting the HDR potential of Cas9 guide RNA (gRNA) using empirically generated editing data as input, the process comprising delivering a Cas9 editing component containing the gRNA of interest to a cell line of interest and collecting genomic DNA after CRISPR editing. The editing results of the gRNA of interest are analyzed and quantified using an NGS-based approach such as the rhAmpSeq CRISPR analysis system. The HDR prediction tool uses this editing data as input to characterize the indel profile of the Cas9 gRNA by creating a set of characteristics, in particular, deletion frequency, insertion frequency, superallelicity, and superallelicity frequency. The HDR prediction tool inputs this set of characteristics into a regression model built on generalizable data (HAP1 HDR data + indel profile) to output a predicted HDR rate. The HDR rate is relative to individual cell lines, and therefore the actual HDR may differ. To screen and select target gRNAs from multiple options, this prediction tool takes the predicted HDR rate of each gRNA as input and provides a rank or score of HDR potential as output.
[0026] Another embodiment described herein is a computer implementation process for predicting the HDR potential of a Cas9 guide RNA (gRNA) using software predictive editing data as input, the process comprising providing sequence information of the Cas9 gRNA of interest to a software tool (e.g., FORECasT) that provides predictive editing results based on sequence context. For example, see Allen et al., Nature Biotechnol. 37: pp. 64-72 (2019), incorporated herein by reference for such teaching. This HDR prediction tool uses this in silico predictive editing data as input to characterize the indel profile of the Cas9 gRNA by creating a set of characteristics, in particular, deletion frequency, insertion frequency, superalleles, and superallele frequencies. This HDR prediction tool inputs this set of characteristics into a regression model built on generalizable data (HAP1 HDR data + indel profile) to output a predicted HDR rate. The HDR rate is relative to individual cell lines, and therefore the actual HDR may differ. To screen and select target gRNAs from multiple options, this prediction tool takes the predicted HDR rate of each gRNA as input and provides a rank or score of HDR potential as output.
[0027] Another embodiment described herein is a method for predicting HDR using complete indel profile characteristics (for deletion frequency only).
[0028] Another embodiment described herein is a method for predicting the HDR potential of gRNA using indel profiles.
[0029] Another embodiment described herein is a method for providing information to a cell line-specific HDR predictive model using cell line repair pathway expression.
[0030] One embodiment described herein is a method for predicting the homologous recombination repair (HDR) potential of one or more Cas guide RNAs (gRNAs), comprising: (a) with respect to one or more candidate gRNAs, (i) performing one or more Cas enzyme editing experiments using one or more candidate gRNAs to obtain edited genomic DNA; (ii) for each editing experiment, amplifying and sequencing the edited genomic DNA to produce sequenced, edited genomic DNA; for each editing experiment, generating an empirical indel profile by performing, in a processor, (iii) receiving the sequenced, edited genomic DNA; and (iv) analyzing the sequenced, edited genomic DNA and outputting an empirical indel profile; (b) inputting the empirical indel profile from step (a) into an HDR prediction model and analyzing the indel profile; and (c) outputting an HDR rate threshold, HDR score, or ranking list of candidate gRNAs indicating preferred candidate gRNAs for HDR editing experiments, and the optimal editing sites.
[0031] Another embodiment described herein is a method for predicting the homologous recombination repair (HDR) potential of one or more Cas guide RNAs (gRNAs), comprising the steps of: (a) generating in silica indel profiles for one or more candidate gRNAs by performing the following in a processor: (i) inputting candidate gRNA sequences and editing loci; and (ii) receiving in silica indel profiles; (b) inputting the in silica indel profiles from step (a) into an HDR prediction model and analyzing the indel profiles; and (c) outputting an HDR rate threshold, HDR score, or ranking list of candidate gRNAs indicating preferred candidate gRNAs for HDR editing experiments, and optimal editing sites.
[0032] Another embodiment described herein is a method for predicting the homologous recombination repair (HDR) potential of one or more Cas guide RNAs (gRNAs), comprising: (a) with respect to one or more candidate gRNAs, (i) performing one or more Cas enzyme editing experiments using one or more candidate gRNAs to obtain edited genomic DNA; (ii) with respect to each editing experiment, amplifying and sequencing the edited genomic DNA to produce sequenced, edited genomic DNA; (iii) receiving the sequenced, edited genomic DNA in a processor with respect to each editing experiment; and (iv) analyzing the sequenced, edited genomic DNA and outputting an empirical indel profile. The method comprises the steps of: (a) generating an empirical indel profile by performing the following: (b) inputting candidate gRNA sequences and editing loci into a processor; and (ii) receiving an in silica indel profile; (c) inputting the empirical indel profile from step (a) or the in silica indel profile from step (b) into an HDR prediction model and analyzing the indel profile; and (d) outputting an HDR rate threshold, HDR score, or ranking list of candidate gRNAs that indicate a candidate gRNA preferred for HDR editing experiments, and the optimal editing site.
[0033] In one embodiment, step (a)(ii) includes amplifying genomic DNA using RNase H-dependent PCR (rhPCR) and performing next-generation sequencing (NGS) to produce sequenced, edited genomic DNA. In another embodiment, the analysis of the sequenced, edited genomic DNA in step (a)(iv) includes merging the sequenced, edited genomic DNA, binning the merged, sequenced, edited genomic DNA by alignment to the genome, and providing alignment of the edited genomic DNA and characterization and quantification of empirical indel frequencies. In another embodiment, the analysis is performed using the rhAmpSeq CRISPR analysis system or CRISPAltRations. In another embodiment, the empirical indel profile includes one or more of the following: allele frequencies, template insertion frequencies, microhomology-mediated end-joining (MMEJ) deletion frequencies, entropy, insertion size frequencies, GC insertion motif frequencies, deletion size frequencies, or a combination thereof. In another embodiment, the step of generating an in silica indel profile includes predicting the effectiveness of a guide RNA, generating alignments and edit frequencies, and mutation results resulting from double-strand breaks. In another embodiment, the input is a guide sequence, and the output is a set of alignments and predictions of on-target base editing effectiveness. In another embodiment, the step of generating an in silica indel profile is performed using FORECasT. In another embodiment, the HDR prediction model in the step includes gradient-boosted regressors, ensemble methods, Lasso regression, structural equation modeling (SEM), or conventional machine learning processes that translate multidimensional indel profiles into HDR rate thresholds, HDR scores, or ranked outputs for candidate gRNAs.In another embodiment, the HDR prediction model is trained by the processor by (i) creating a training dataset using empirical indel profiles or in silica indel profiles; (ii) creating a test dataset using empirical indel profiles or in silica indel profiles; and (iii) training and testing the HDR prediction model, wherein the HDR prediction model is trained using the training dataset and the HDR prediction model is tested using the test dataset. In another embodiment, the HDR prediction model is capable of accurately ranking candidate gRNAs of overall HDR potential whose Spearman correlation is greater than 0.5. In another embodiment, the HDR rate and preferred candidate gRNAs are specific to a particular cell type or cell line. In another embodiment, the candidate gRNA sequence has a variable region of about 17 to about 24 nucleotides in length. In another embodiment, the candidate gRNA sequence has a variable region of about 20 nucleotides in length. In another embodiment, the candidate gRNA sequence includes one or more modifications to the 5' end, the 3' end, or a combination thereof. In another embodiment, the modification includes terminal blocking modification. In another embodiment, the editing site or editing locus is Cas enzyme-specific and contains about 1 to about 15 nucleotides. In another embodiment, the Cas enzyme is Cas9 or Cas12a. In another embodiment, the genomic DNA is derived from a cell or a population of interest. In another embodiment, the candidate gRNA sequence contains a sequence derived from one or more of SEQ ID NOs: 1-255 or 1021-1068.
[0034] Another embodiment described herein is a research tool comprising the nucleotide sequence described herein.
[0035] Another embodiment described herein is a reagent comprising the nucleotide sequence described herein.
[0036] Another embodiment described herein is a method for producing one or more of the nucleotide sequences described herein or a polypeptide encoded by the nucleotide sequences described herein, comprising the steps of: transforming or transfecting cells with nucleic acids containing the nucleotide sequences described herein; growing the cells; optionally isolating an additional amount of the nucleotide sequences described herein; inducing the expression of the polypeptide encoded by the nucleotide sequences described herein; and isolating the polypeptide encoded by the nucleotides described herein.
[0037] The polynucleotides described herein include variants having substitutions, deletions, and / or additions that may contain one or more nucleotides. These variants may be modified in the coding region, the non-coding region, or both. Modifications in the coding region may result in conserved or non-conserved amino acid substitutions, deletions, or additions. Of these, silent substitutions, additions, and deletions that do not alter the binding properties and activity are particularly preferred.
[0038] Further embodiments described herein include nucleic acid molecules comprising (a) a nucleotide sequence encoding a polypeptide having the amino acid sequences of SEQ ID NOs: 1 to 1212, or a degenerate variant, homologous variant, or codon-optimized variant thereof, or (b) a nucleotide sequence that is approximately 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99% identical, more preferably at least approximately 90-99% or 100%, to a nucleotide sequence that can hybridize complementary to any of the nucleotide sequences of (a).
[0039] A polynucleotide having a nucleotide sequence that is at least, for example, 9–99% "identical" to a reference nucleotide sequence is intended to have a nucleotide sequence identical to the reference sequence, except that the polynucleotide sequence may contain up to approximately 10 to 1 mutations, additions, or deletions for every 100 nucleotides of the reference nucleotide sequence.
[0040] In other words, to obtain a polynucleotide having a nucleotide sequence that is at least approximately 90-99% identical to a reference nucleotide sequence, up to 10% of the nucleotides in the reference sequence may be deleted, another nucleotide added, or another nucleotide replaced, or up to 10% of the total nucleotides in the reference sequence may be inserted into the reference sequence. These mutations in the reference sequence may occur at the 5' or 3' terminal position of the reference nucleotide sequence, or at any location between these terminal positions, which may be individually scattered among the nucleotides in the reference sequence, or scattered within one or more consecutive groups within the reference sequence. The same can be applied to polypeptide sequences that are at least approximately 90-99% identical to a reference polypeptide sequence.
[0041] As described above, two or more polynucleotide sequences can be compared by determining their percentage of identity. Similarly, two or more amino acid sequences can be compared by determining their percentage of identity. The percentage of identity between two sequences (nucleic acid sequences or peptide sequences) is generally described as the number of perfect matches between the two aligned sequences divided by the length of the shorter sequence and multiplied by 100. Alignment methods for polynucleotide sequences or polypeptide sequences are provided by the local homology algorithms of Smith and Waterman, Advances in Applied Mathematics 2: 4 pp. 82-489 (1981) or Needleman and Wunsch, J. Mol. Biol. 48 (3): pp. 443-453 (1970).
[0042] Another embodiment described herein is a polynucleotide vector comprising one or more nucleotide sequences described herein.
[0043] Another embodiment described herein is a cell comprising one or more nucleotide sequences or polynucleotide vectors described herein.
[0044] It will be apparent to those skilled in the art that appropriate modifications and adaptations to the compositions, formulations, methods, processes, and uses described herein can be made without departing from the scope of any embodiment or aspect thereof. The compositions and methods provided are illustrative and are not intended to limit the scope of any particular embodiment. All of the various embodiments, aspects, and options disclosed herein can be combined in any variation or iteration. The scope of the compositions, formulations, methods, and processes described herein includes all actual or potential combinations of the embodiments, aspects, options, examples, and preferences described herein. The exemplary compositions and formulations described herein may omit any component, replace any component disclosed herein, or include any component disclosed elsewhere herein. The ratio of the mass of any component of any composition or formulation disclosed herein to the mass of any other component in this formulation, or the total mass of the other components in this formulation, is disclosed herein as if it had been expressly disclosed. If the meaning of any term in any patent or publication incorporated by reference conflicts with the meaning of a term used herein, the meaning of the term or phrase herein shall prevail. Furthermore, the above discussions only disclose and illustrate exemplary embodiments. All patents and publications cited herein are incorporated herein by reference with respect to their specific teachings.
[0045] The various embodiments and aspects of the present invention described herein are summarized in the following clauses.
[0046] Clause 1. A method for predicting the homologous recombination repair (HDR) potential of one or more Cas guide RNAs (gRNAs), (a) With respect to one or more candidate gRNAs, (i) Perform one or more Cas enzyme editing experiments using one or more candidate gRNAs to obtain edited genomic DNA; (ii) For each editing experiment, the edited genomic DNA is amplified and sequenced to produce sequenced, edited genomic DNA; For each editing experiment, the processor: (iii) receiving sequenced, edited genomic DNA; and (iv) Analyze sequenced and edited genomic DNA and output empirical indel profiles. The process of generating an empirical indel profile by performing the following; (b) A step of inputting the empirical indel profile from step (a) into an HDR prediction model and analyzing the indel profile; and (c) A process of outputting an HDR rate threshold, HDR score, or ranking list of candidate gRNAs that are suitable for HDR editing experiments, and the optimal editing site. A method that includes this.
[0047] Clause 2. A method for predicting the homologous recombination repair (HDR) potential of one or more Cas guide RNAs (gRNAs), (a) In the processor, (i) Enter candidate gRNA sequences and edit loci; and (ii) Receiving an in silicoindel profile The process of generating in silicic coindel profiles for one or more candidate gRNAs by performing the following steps; (b) A step of inputting the in silico indel profile from step (a) into an HDR prediction model and analyzing the indel profile; and (c) A process of outputting an HDR rate threshold, HDR score, or ranking list of candidate gRNAs that are suitable for HDR editing experiments, and the optimal editing site. A method that includes this.
[0048] Clause 3. A method for predicting the homologous recombination repair (HDR) potential of one or more Cas guide RNAs (gRNAs), (a) With respect to one or more candidate gRNAs, (i) Perform one or more Cas enzyme editing experiments using one or more candidate gRNAs to obtain edited genomic DNA; (ii) For each editing experiment, the edited genomic DNA is amplified and sequenced to produce sequenced, edited genomic DNA; For each editing experiment, the processor: (iii) receiving sequenced, edited genomic DNA; and (iv) Analyze sequenced and edited genomic DNA and output empirical indel profiles. The process of generating an empirical indel profile by performing the following; or (b) Processor, (i) Enter candidate gRNA sequences and edit loci; and (ii) Receiving an in silicoindel profile The process of generating in silicic coindel profiles for one or more candidate gRNAs by performing the following steps; (c) A step of inputting an empirical indel profile from step (a) or an in silico indel profile from step (b) into an HDR prediction model and analyzing the indel profile; and (d) A process of outputting an HDR rate threshold, HDR score, or ranking list of candidate gRNAs that are suitable for HDR editing experiments, and the optimal editing site. A method that includes this.
[0049] Clause 4. The method according to Clause 1 or 3, wherein step (a)(ii) comprises amplifying genomic DNA using RNase H-dependent PCR (rhPCR) and performing next-generation sequencing (NGS) to produce sequenced, edited genomic DNA.
[0050] Clause 5. The method according to any one of Clauses 1, 3, or 4, wherein the analysis of sequenced and edited genomic DNA in step (a)(iv) comprises merging sequenced and edited genomic DNA, binning the merged, sequenced and edited genomic DNA by alignment to the genome, and providing alignment of the edited genomic DNA and characterization and quantification of empirical indel frequencies.
[0051] Clause 6. The method described in Clause 5, wherein the analysis is performed using the rhAmpSeq CRISPR analysis system or CRISPAltRations.
[0052] Clause 7. An empirical indel profile as described in any one of Clauses 1 to 6, comprising one or more of the following: allele frequency, template insertion frequency, microhomology-mediated end-joining (MMEJ) deletion frequency, entropy, insertion size frequency, GC insertion motif frequency, deletion size frequency, or a combination thereof.
[0053] Clause 8. The process for generating an in silicoindel profile, comprising predicting the effectiveness of the guide RNA, generating alignment and editing frequencies, and the mutation results resulting from double-strand breaks, as described in Clause 2 or 3.
[0054] Clause 9. The method according to Clause 8, wherein the input is a guide sequence and the output is a set of alignment and on-target base editing effectiveness predictions.
[0055] Clause 10. The method according to Clause 2 or 3, wherein the process of generating an in silicoindel profile is carried out using FORECasT.
[0056] Clause 11. The HDR prediction model in the process is any one of the methods described in Clauses 1 through 10, including gradient-boosted regressors, ensemble methods, lasso regression, structural equation modeling (SEM), or conventional machine learning processes that translate multidimensional indel profiles into HDR rate thresholds, HDR scores, or ranked outputs for candidate gRNAs.
[0057] Clause 12. The HDR prediction model is implemented on the processor. (i) Create a training dataset using empirical indel profiles or in silico indel profiles; (ii) Creating a test dataset using empirical indel profiles or in silico indel profiles; and (iii) Training and testing an HDR prediction model, which involves training the HDR prediction model using a training dataset and testing the HDR prediction model using a test dataset. The method described in any one of clauses 1 through 11, which is used to train by performing the exercise.
[0058] Clause 13. An HDR prediction model capable of accurately ranking candidate gRNAs of overall HDR potential having a Spearman correlation value greater than 0.5, as described in any one of Clauses 1 through 12.
[0059] Clause 14. The method according to any one of Clauses 1 to 13, wherein the HDR rate and preferred candidate gRNA are specific to a particular cell type or cell line.
[0060] Clause 15. The candidate gRNA sequence is the method described in any one of Clauses 1 to 14, having a variable region of approximately 17 to 24 nucleotides in length.
[0061] Clause 16. The candidate gRNA sequence is the method described in Clause 15, having a variable region of approximately 20 nucleotides in length.
[0062] Clause 17. A candidate gRNA sequence comprising one or more modifications to the 5' end, the 3' end, or a combination thereof, as described in any one of Clauses 1 to 16.
[0063] Clause 18. Modifications are those described in Clause 17, including terminal blocking modifications.
[0064] Clause 19. The method according to any one of Clauses 1 to 18, wherein the editing site or editing locus is Cas enzyme-specific and contains approximately 1 to approximately 15 nucleotides.
[0065] Clause 20. The method according to any one of Clauses 1 to 19, wherein the Cas enzyme is Cas9 or Cas12a.
[0066] Clause 21. Genomic DNA is derived from a cell or a population of interest, as described in any one of Clauses 1 to 20.
[0067] Clause 22. The method according to any one of Clauses 1 to 21, wherein the candidate gRNA sequence comprises a sequence derived from one or more of SEQ ID NOs. 1-255 or 1021-1068. [Examples]
[0068] (Example 1) Figure 1 shows a block diagram illustrating an exemplary system for predicting the homologous recombination repair (HDR) potential of one or more Cas9 guide RNAs (gRNAs) according to various aspects of this disclosure. In the example of Figure 1, system 100 includes a homologous recombination repair (HDR) server 104, a client device 130, and a network 140.
[0069] The HDR server 104 may be owned by an administrator, operated by an administrator, or for the administrator. The HDR server 104 comprises an electronic processor 106, a communication interface 108, and memory 110. The electronic processor 106 is communicated to the communication interface 108 and the memory 110. The electronic processor 106 is a microprocessor or another preferred processing device. The communication interface 108 may be implemented as either or both a wired network interface and / or a wireless network interface. The memory 110 is one or more of volatile memory (e.g., RAM) and non-volatile memory (e.g., ROM, FLASH, magnetic media, optical media, etc.). In some examples, the memory 110 is also a non-temporary computer-readable medium. Although shown within the HDR server 104, the memory 110 may be implemented as network storage located at least partially outside the HDR server 104 and accessed via the communication interface 108. For example, all or part of the memory 110 may be stored in the "cloud".
[0070] The HDR application 112 may be stored in a temporary or non-temporary portion of memory 110. The HDR application 112 includes machine-readable instructions executed by the electronic processor 106 to perform the functions of the HDR server 104, as described below with respect to Figure 2.
[0071] Memory 110 may include a database 114 that stores information about one or more Cas guide RNAs (gRNAs). Database 114 may be an RDF database, i.e., it may employ a Resource Description Framework. Alternatively, database 114 may be another preferred database with similar characteristics to the Resource Description Framework, and various non-SQL databases, knowledge graphs, etc. Database 114 may contain multiple data. The data may be associated with one or more Cas9 editing experiments using one or more candidate gRNAs, and may contain this information. For example, in the shown embodiment, database 114 includes indel profiles 115 and HDR data 116. Indel profiles 115 may contain multiple sets of raw data associated with an account user. In some cases, the raw data sets 115 are generated based on transactions (e.g., requests) associated with a user device 150, a client device 140, and / or a data source 130. HDR data 116 may contain client data received from a client device 140 associated with an account user. In some cases, the feedback data 116 may include fraud information associated with the user account. Memory 110 may also include training data 118 and a machine learning model 120. The training data 118 may include a set of past requests (request history) associated with the user account. Labels 120 may include a set of labeled training examples for training an ML model that generates scores associated with the user.
[0072] The data source 130 may be an on-premises, cloud, or edge computing system that provides data and may include an electronic processor that communicates with memory. This electronic processor may be a microprocessor or another preferred processing device, the memory may be one or more of volatile and non-volatile memory, and the communication interface may be a wireless or wired network interface. In some examples, the data source 130 may be accessed directly by the label server 104. In other examples, the data source 130 may be accessed indirectly via the network 160. For example, the data source 130 may be the source of transactions associated with a user account that are transmitted between the user device 150 and the data source 130. In some cases, this transaction may include one or more requests for the user account. In some embodiments, the label creation application 112 retrieves data from the data source 130 via the network 160.
[0073] The client device 140 may be a web-enabled mobile computer such as a laptop, tablet, smartphone, or other suitable computing device. Alternatively, or in addition, the client device 140 may be a desktop computer. The client device 140 includes an electronic processor that communicates with memory. This electronic processor is a microprocessor or another suitable processing device, the memory is one or more of volatile and non-volatile memory, and the communication interface may be a wireless or wired network interface.
[0074] An application, including software instructions implemented on the electronic processor of the client device 140 to perform the functions of the client device 140 described herein, is stored in a temporary or non-temporary portion of memory. This application may include a graphical user interface to facilitate interaction between the user and the client device 140.
[0075] The client device 140 can communicate with the label server 104 via the network 160. The network 160 is preferably a wireless network such as a wireless personal area network, a local area network, or other suitable network (however, it is not necessarily a wireless network). In some examples, the client device 140 can communicate directly with the label server 104. In other examples, the client device 140 can communicate indirectly with the label server 104 via the network 160.
[0076] Figure 2 is a flowchart illustrating an exemplary process 200 for predicting the homologous recombination repair (HDR) potential of one or more Cas9 guide RNAs (gRNAs) according to various aspects of this disclosure. In the example in Figure 2, the process 200 is described in a sequential flow, but parts of the process 200 may also be performed in parallel.
[0077] Process 200 generates indel profiles (in block 205). For example, client device 130 generates indel profiles 115 (e.g., empirical indel profiles) for one or more candidate gRNAs. In this example, the user performs one or more Cas9 editing experiments using one or more candidate gRNAs to obtain edited genomic DNA. In each experiment, the edited genomic DNA is amplified and sequenced to produce sequenced, edited genomic DNA. In addition, the user inputs the sequenced, edited genomic DNA into client device 130, which analyzes the sequenced, edited genomic DNA and outputs empirical indel profiles. In another example, HDR server 104 generates indel profiles 115 (in silica indel profiles) for one or more candidate gRNAs. In this example, the HDR server 104 receives candidate gRNA sequences and editing loci from the client device 130 and inputs the candidate gRNA sequences, and the HDR application generates an in silica coindel profile using locally hosted software (such as FORECasT).
[0078] Process 200 receives the indel profile (in block 210). For example, HDR server 104 receives an indel profile 115 (e.g., an insilica indel profile or an empirical indel profile) from client device 130. In another example, HDR server receives an indel profile 115 (e.g., an insilica indel profile) generated by HDR application 112 and stores this indel profile 115 in memory 110.
[0079] In the initial implementation, process 200 trains a predictive HDR model (in block 215). For example, the HDR application 112 uses the indel profile 115 to create training data 118 and trains the machine learning algorithm 120. In some cases, the training data 118 includes training and test datasets created using empirical or in-silica indel profiles. In other examples, the machine learning model 120 is initially trained using a client-generated empirical indel profile, resulting in improved accuracy of inferences determined by the machine learning model 120 in subsequent iterations of use. Subsequent runs of process 200 may not require further training, and therefore block 215 is optional; however, additional training may be beneficial in improving the accuracy of inferences determined by the machine learning model 120.
[0080] Process 200 inputs the indel profile (in block 220) into a predictive HDR model. For example, HDR application 112 inputs the indel profile 115 from block 210 into machine learning model 120. Machine learning model 120 analyzes the indel profile and generates an output. This output is a value for each candidate gRNA, indicating the HDR potential of each candidate gRNA.
[0081] Process 200 selects candidate gRNAs based on the output of the predicted HDR model (in block 225). For example, HDR application 112 selects candidate gRNAs from the received set of candidate gRNAs. In some cases, HDR application 112 determines an HDR rate threshold based on the value of each candidate gRNA. In other cases, HDR application 112 orders the set of candidate gRNAs based on the value of each candidate gRNA.
[0082] (Example 2) Key attributes of indel profiles for predicting HDR potential Large HDR datasets were generated by delivering CRISPR Cas9 HDR reagents targeting 263 sites to Jurkat and HAP1 cell lines. Cas9 ribonucleoprotein complexes (RNPs) were formed by mixing Alt-R® Sp Cas9 nuclease with either annealed Alt-R® modified crRNA:tracrRNA (two-component gRNA) or Alt-R® modified sgRNA (single guide gRNA) in a 1:1.2 ratio to the Cas9 protein gRNA (Alt-R® reagents are from IDT, Coralville, IA). The 4 μM Cas9 RNP complex was delivered together with 4 μM Alt-R® Cas9 Electroporation Enhancer and 3 μM Alt-R® HDR Donor Oligos using a Lonza 4D-Nucleofector 96-well system (Lonza, Basel, Switzerland). The Alt-R® modification includes proprietary 5' and 3' terminal blocking groups to prevent nucleotide degradation (IDT, Coralville, IA). The HDR donor is designed to introduce a 6bp "GAATTC" sequence into the DSB, corresponding to the non-target DNA strand for gRNA. CRISPR reagents were delivered to 3E5 cells (HAP1) or 5E5 cells (Jurkat) using cell line-appropriate nucleofection conditions (DS-120 program and CL-120 program, respectively). Test conditions included RNP only (2-component gRNA), RNP only (sgRNA), RNP + HDR donor (2-component gRNA), and an untreated control. After 72 hours, DNA was extracted using QuickExtract® DNA extract (Lucigen, Madison, WI). Editing results were quantified by NGS amplicon sequencing on the Illlumina MiSeq platform using the rhAmpSeq library preparation method. Data analysis was performed using the in-house version of the rhAmpSeq CRISPR Analysis System from IDT.The sequences of the gRNA protospacer, donor oligo, and sequencing primer are listed in Table 1.
[0083] [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6] [Table 1-7] [Table 1-8] [Table 1-9] [Table 1-10] [Table 1-11] [Table 1-12] [Table 1-13] [Table 1-14] [Table 1-15] [Table 1-16] Table 1-17 Table 1-18 Table 1-19 Table 1-20 Table 1-21 Table 1-22 Table 1-23 Table 1-24 Table 1-25 Table 1-26 Table 1-27 Table 1-28 Table 1-29 Table 1-30 Table 1-31 Table 1-32 Table 1-33 [Table 1-34] [Table 1-35] [Table 1-36] [Table 1-37] [Table 1-38] [Table 1-39]
[0084] To create predictive modeling characteristics, we developed software to describe the resulting NHEJ / MMEJ profiles of CRISPR Cas9 edits and connected it to the output of the rhAmpSeq CRISPR analysis system. This includes additional indel profile characteristics such as upper allele frequency, template insertion frequency, MMEJ deletion frequency, entropy, insertion size frequency, GC insertion motif frequency, and deletion size frequency. The definitions of these characteristics are described (Table 2). Indel profiles were characterized under the "RNP-only control" condition (i.e., no HDR template added). To exclude sites that may introduce confounding factors into the model (e.g., insufficient editing, insufficient data, etc.), we filtered for sites where Cas9 editing was less than 90% in the RNP-only control, sites where background editing was detected more than 10% in the unedited control, or sites where sequencing reads were less than 500 in either the RNP-only control or the HDR condition. After applying filters, 150 sites within HAP1 were used as input for further correlation analysis and modeling attempts.
[0085] [Table 2]
[0086] The Pearson correlation (R) between individual indel profile attributes and HDR results was calculated to first determine the key predictive properties of HDR (Table 3). Several indel profile properties were identified as candidates for HDR prediction (Figure 3). A negative correlation was observed between HDR rate and top AF (R). 2 =0.22). HDR rate and indel profile entropy (R 2 =0.44) and deletion 3+(R 2 A positive correlation was observed between (=0.43). These findings support the observation by Tatiossian et al. that large deletions based on MMEJ predict HDR. See Tatiossian et al., Mol. Ther. 29(3): pp. 1057-1069 (2021). However, as is clear from the results for upper AF and entropy, the complexity of the indel profile is another important characteristic. The concept of targeting upper alleles with recursive editing or double-tap methods to improve HDR was introduced by Moeller et al. and Bodai et al., but the predictive properties of the upper alleles for selecting good HDR gRNAs were not proposed. See Moeller et al., Nature Commun. 13(1): pp. 4550 (2022); Bodai et al., Nature Commun. 13(1): pp. 2351 (2022). Furthermore, the negative correlation between HDR frequency and top AF suggests that top candidates for recursive editing / double tap methods may be the worst initial candidates for HDR, or that recursive editing methods may need to be applied in a way that reduces the proliferation of these high-frequency restoration results before HDR can be fully enhanced.
[0087] [Table 3]
[0088] (Example 3) Development of an HDR predictive model While a single trait within the NHEJ / MMEJ indel profile was shown to correlate with the HDR results, the correlation is likely to be strengthened by comprehensively evaluating the trait in terms of the dependent variable within the constructed model. For the 150 HAP1 sites evaluated in Example 2, to identify and remove traits contributing to the multicollinearity problem according to the program, we first constructed a multilinear regression in GraphPad Prism (Dotmatics) using the trait with the HDR values of the site pairs as the dependent variable. Next, we split the dataset into a training database and a test database (75 / 25 split; 100 bootstraps), and then, using the trait, we constructed a Gradient Boosting Regressor using SciKit-Learn and evaluated it using the bootstrapped test dataset. From the analysis of the model on HAP1, we found that the model exhibited good Pearson correlation (R) of decisions over 100 bootstraps. 2 The results showed a strong Spearman correlation (=0.45±0.13) and ranking (Spearman correlation = 0.67±0.09) (Figure 4, sample test results).
[0089] To test whether the model is directly translatable to cell lines with known NHEJ / MMEJ repair differences, the HDR prediction model constructed with HAP1 data was further tested using Jurkat HDR and indel profile data generated for the same sites described in Example 2. It can be seen that the HDR rates predicted by the HAP1 model do not generalize well to the measured Jurkat HDR rates (Figure 5). However, Kurgan et al. have already observed that the Jurkat Cas9 indel profile is very different from that reported as a general NHEJ / MMEJ repair profile for other cell lines. For such teachings, see Kurgan et al., Mol. Ther. - Methods Clin. Dev. 21: pp. 478-491 (2021), incorporated herein by reference.
[0090] The expression profiles of DNA repair factors can contribute to a unique set of HDR predictors, thus influencing the model's ability to accurately predict HDR outcomes in specific cell types. In the case of Jurkat cells, the high expression of immune cell-specific terminal deoxynucleotidyltransferase (TdT) compared to other commonly used experimental cell lines (Figure 6A) contributes to a unique set of HDR predictors in Jurkat cells, thus contributing to the lack of generalization of the HAP1 model. The TdT protein is a template-dependent DNA polymerase that contributes to V(D)J recombination in lymphocytes by the random addition of nucleotides to the available 3' end at DSBs. Consequently, a higher insertion frequency was observed in the indel profiles of Jurkat cells (Figure 6B), and the Pearson correlation between HDR and indel profile attributes was altered compared to HAP1 cells (Table 4).
[0091] [Table 4]
[0092] (Example 4) Application of key attributes and HDR prediction models across all cell types. To investigate the performance of the HAP1-based HDR prediction model across additional cell types, a subset of 48 sites was selected from the initial 263 sites described in Example 2. At the selected sites, editing in the RNP-only control exceeded 90%, background editing in the unedited control was less than 10%, and the HDR rate with HAP1 ranged from 2% to 50%. CRISPR-Cas9 HDR reagents for these 48 sites were delivered to K562, iPSC, and primary T cell lines, and the editing results were evaluated.
[0093] Cas9 RNP (consisting of Alt-R® Sp Cas9 nuclease and Alt-R® sgRNA) was formed in a 1:1.2 ratio to the Cas9 protein gRNA. For K562 cells, the 2 μM Cas9 RNP complex was delivered together with 2 μM Alt-R Cas9 electroporation enhancer and 2 μM Alt-R HDR donor oligo using a Lonza 4D-Nucleofector 96-well system (Lonza, Basel, Switzerland) and conditions suitable for the cell line (FF-120). For iPSCs, a 4μM Cas9 RNP complex was delivered along with a 4μM Alt-R Cas9 electroporation enhancer (RNP-only control) and a 4μM Alt-R HDR donor oligo (HDR condition) using the Lonza 4D-Nucleofector 96-well system (Lonza, Basel, Switzerland) and cell line-appropriate conditions (CA-137). For primary T cells, a 4μM Cas9 RNP complex was delivered along with a 3μM Alt-R Cas9 electroporation enhancer and a 2μM Alt-R HDR donor oligo using the Lonza 4D-Nucleofector 96-well system (Lonza, Basel, Switzerland) and cell line-appropriate conditions (ER-115). The HDR donor was designed to introduce a 6bp "GAATTC" sequence into the DSB, corresponding to the non-target DNA strand for gRNA. The tested conditions included RNP only, RNP + HDR donors, and untreated controls. DNA was extracted using QuickExtract® DNA extract (Lucigen, Madison, WI) at 48 hours (K62, primary T cells) or 96 hours (iPSCs). Editing results were quantified by NGS amplicon sequencing on the Illlumina MiSeq platform using the rhAmpSeq library preparation method. Data analysis was performed using the IDT in-house version of the rhAmpSeq CRISPR Analysis System.The sequences of the gRNA protospacer, donor oligo, and sequencing primers are listed in Table 5.
[0094] In K562 cells, iPSCs, and primary T cells, with some notable exceptions, a similar correlation between HDR and the major indel profile attributes was observed (Figures 7, 9, 11). A negative correlation was observed between the HDR rate and the top AF (R 2 = 0.46 for K562, R 2 = 0.49 for iPSCs). Positive correlations were observed between the HDR rate and indel profile entropy (R 2 = 0.35 for K562, R 2 = 0.54 for iPSCs, R 2 = 0.35 for primary T cells) and deletion 3+ (R 2 = 0.35 for K562, R 2 = 0.27 for iPSCs, R 2 = 0.27 for primary T cells). Furthermore, when the results were compared to the original HAP1 dataset, HDR and indel profile attributes were reasonably correlated (R 2 = 0.47 - 0.81 for K562, R 2 = 0.42 - 0.92 for iPSCs, R 2 = 0.19 - 0.63 for primary T cells) (Figures 8, 10, 12).
[0095] Next, indel profile data from K562, iPSCs, and primary T cells were processed through 100 bootstrap iterations of a HAP1-based HDR prediction model and compared to the measured HDR rates in each cell type (sample results are shown in Figure 13). The model was unable to accurately predict absolute HDR% (Pearson correlation = -0.80±0.28 for K562, -0.63±0.17 for iPSCs, and 0.03±0.19 for primary T cells), but it was able to accurately rank gRNAs in terms of overall HDR potential (Spearman correlation = 0.66±0.06 for K562, 0.66±0.03 for iPSCs, and 0.53±0.10 for primary T cells). The inability to predict absolute HDR values was not surprising, given the variability in HDR rates observed between cell lines. However, the ability of a model to provide cell line-independent gRNA ranking is a crucial characteristic for CRISPR HDR applications.
[0096] To investigate the performance of the HDR prediction models described herein compared to prior art, we performed comparisons with predictions of large deletion frequencies alone. A secondary prediction model was constructed using HAP1 3+Del frequency as the sole predictive property. Using this model, the predicted HDR rates were compared with HDR rates measured from K562, iPSC, and primary T cell datasets (Figure 14). The deletion-based model successfully ranked the HDR potential of gRNAs in some cell lines (Spearman correlation = 0.52 for K562 cells and 0.53 for iPSCs), but did not achieve the same level of HDR ranking accuracy as the complete tool (Spearman correlation = 0.66±0.06 for K562 cells and 0.66±0.03 for iPSCs). In the case of primary T cells, deletion-based models failed to rank the HDR potential of gRNAs compared to a comprehensive, complete predictive tool (Spearman correlation = 0.53 ± 0.10) (Spearman correlation = 0.16 ± 0.07). This difference is largely due to the poor correlation between HDR and large deletions observed in primary T cells (Figure 11). This further demonstrates the advantage of comprehensive models over prior art, where a complete profile of indel properties can compensate for the poor correlation of individual properties that may be cell line-specific.
[0097] In summary, these data establish the capability of an HDR predictive model to provide a ranking of Cas9 gRNA's HDR potential based on indel profile characteristics such as large deletion frequencies, entropy, and top allele frequencies, as well as other factors. These data further demonstrate the advantages of a comprehensive indel profile-based model over published prior art that utilizes deletion frequencies alone. This model is applicable across multiple cell types, including clinically relevant cell types such as iPSCs and primary T cells. It may be possible to develop cell-type-specific HDR models based on the expression profiles of key DNA repair genes that contribute to unique indel profile characteristics.
[0098] [Table 5A] [Table 5B] [Table 5C] [Table 5D] [Table 5E] [Table 5F] [Table 5G] [Table 5H] [Explanation of Symbols]
[0099] 100 Systems 104 Homologous Recombination Repair (HDR) Server, Label Server 106 Electronic Processors 108 Communication Interfaces 110 memory 112 HDR applications, label creation applications 114 Databases 115 Indel Profiles, Raw Datasets 116 HDR data, feedback data 118 training data 120 Machine Learning Models, Labels, and Machine Learning Algorithms 130 client devices, data sources 140 Network, Client Devices 150 user devices 160 Networks 200 processes
Claims
1. A method for predicting the homologous recombination repair (HDR) potential of one or more Cas guide RNAs (gRNAs), (a) With respect to one or more candidate gRNAs, (i) Perform one or more Cas enzyme editing experiments using one or more candidate gRNAs to obtain edited genomic DNA; (ii) For each editing experiment, the edited genomic DNA is amplified and sequenced to produce sequenced, edited genomic DNA; For each editing experiment, the processor: (iii) receiving sequenced, edited genomic DNA; and (iv) Analyze sequenced and edited genomic DNA and output empirical indel profiles. The process of generating an empirical indel profile by performing the following; (b) A step of inputting the empirical indel profile from step (a) into an HDR prediction model and analyzing the indel profile; and (c) A process of outputting an HDR rate threshold, HDR score, or ranking list of candidate gRNAs that are suitable for HDR editing experiments, and the optimal editing site. A method that includes this.
2. A method for predicting the homologous recombination repair (HDR) potential of one or more Cas guide RNAs (gRNAs), (a) In the processor, (i) Enter candidate gRNA sequences and edit loci; and (ii) Receiving an in silicoindel profile The process of generating in silicic coindel profiles for one or more candidate gRNAs by performing the following steps; (b) A step of inputting the in silico indel profile from step (a) into an HDR prediction model and analyzing the indel profile; and (c) A process of outputting an HDR rate threshold, HDR score, or ranking list of candidate gRNAs that are suitable for HDR editing experiments, and the optimal editing site. A method that includes this.
3. A method for predicting the homologous recombination repair (HDR) potential of one or more Cas guide RNAs (gRNAs), (a) With respect to one or more candidate gRNAs, (i) Perform one or more Cas enzyme editing experiments using one or more candidate gRNAs to obtain edited genomic DNA; (ii) For each editing experiment, the edited genomic DNA is amplified and sequenced to produce sequenced, edited genomic DNA; For each editing experiment, the processor: (iii) receiving sequenced, edited genomic DNA; and (iv) Analyze sequenced and edited genomic DNA and output empirical indel profiles. The process of generating an empirical indel profile by performing the following; or (b) Processor, (i) Enter candidate gRNA sequences and edit loci; and (ii) Receiving an in silicoindel profile The process of generating in silicic coindel profiles for one or more candidate gRNAs by performing the following steps; (c) A step of inputting an empirical indel profile from step (a) or an in silico indel profile from step (b) into an HDR prediction model and analyzing the indel profile; and (d) A process of outputting an HDR rate threshold, HDR score, or ranking list of candidate gRNAs that are suitable for HDR editing experiments, and the optimal editing site. A method that includes this.
4. The method according to claim 1 or 3, wherein step (a)(ii) comprises amplifying genomic DNA using RNase H-dependent PCR (rhPCR) and performing next-generation sequencing (NGS) to generate sequenced, edited genomic DNA.
5. The method according to claim 1, 3, or 4, wherein analyzing the sequenced and edited genomic DNA in step (a)(iv) comprises merging the sequenced and edited genomic DNA, binning the merged, sequenced and edited genomic DNA by alignment to the genome, and providing alignment of the edited genomic DNA and characterization and quantification of empirical indel frequencies.
6. The method according to claim 5, wherein the analysis is performed using an rhAmpSeq CRISPR analysis system or CRISPAltRations.
7. The method according to any one of claims 1 to 6, wherein the empirical indel profile comprises one or more of the following: allele frequency, template insertion frequency, microhomology-mediated end-joining (MMEJ) deletion frequency, entropy, insertion size frequency, GC insertion motif frequency, deletion size frequency, or a combination thereof.
8. The method according to claim 2 or 3, wherein the step of generating an in silica coindel profile includes predicting the effectiveness of the guide RNA, generating alignment and edit frequencies, and mutation results resulting from double-strand breaks.
9. The method according to claim 8, wherein the input is a guide sequence and the output is a set of alignment and on-target base editing effectiveness predictions.
10. The method according to claim 2 or 3, wherein the step of generating an in silica coindel profile is performed using FORECasT.
11. The method according to any one of claims 1 to 10, wherein the HDR prediction model in the process includes a gradient-boosted regressor, an ensemble method, Lasso regression, structural equation modeling (SEM), or a conventional machine learning process that converts a multidimensional indel profile into an HDR rate threshold, HDR score, or ranked output for a candidate gRNA.
12. The HDR prediction model is implemented on the processor. (i) Create a training dataset using empirical indel profiles or in silico indel profiles; (ii) Creating a test dataset using empirical indel profiles or in silico indel profiles; and (iii) Training and testing an HDR prediction model, which involves training the HDR prediction model using a training dataset and testing the HDR prediction model using a test dataset. The method according to any one of claims 1 to 11, wherein training is performed by executing the following.
13. The method according to any one of claims 1 to 12, wherein the HDR prediction model is capable of accurately ranking candidate gRNAs of overall HDR potential having a Spearman correlation value greater than 0.
5.
14. The method according to any one of claims 1 to 13, wherein the HDR rate and preferred candidate gRNA are specific to a particular cell type or cell line.
15. The method according to any one of claims 1 to 14, wherein the candidate gRNA sequence has a variable region having a length of approximately 17 to approximately 24 nucleotides.
16. The method according to claim 15, wherein the candidate gRNA sequence has a variable region of approximately 20 nucleotides in length.
17. The method according to any one of claims 1 to 16, wherein the candidate gRNA sequence comprises one or more modifications to the 5' end, the 3' end, or a combination thereof.
18. The modification according to claim 17, wherein the modification includes a terminal blocking modification.
19. The method according to any one of claims 1 to 18, wherein the editing site or editing locus is Cas enzyme-specific and comprises about 1 to about 15 nucleotides.
20. The method according to any one of claims 1 to 19, wherein the Cas enzyme is Cas9 or Cas12a.
21. The method according to any one of claims 1 to 20, wherein the genomic DNA is derived from a cell or a population of interest.
22. The method according to any one of claims 1 to 21, wherein the candidate gRNA sequence includes a sequence derived from one or more of sequence numbers 1 to 255 or 1021 to 1068.