Identification of functional DNA variants in the human genome using high-throughput prime editing screens.
The high-throughput prime editing screen with lentiviral delivery optimizes PE efficiency, addressing inefficiencies in characterizing DNA variants, enabling precise genomic annotation and disease-related variant identification.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-09
- Publication Date
- 2026-03-17
AI Technical Summary
Current methods for characterizing the functional effects of DNA variants, particularly those in non-coding regions, are inefficient and lack high-throughput capabilities, hindering the understanding of disease etiology and personalized medicine.
A high-throughput prime editing (PE) screen using a lentiviral delivery system with stable expression of nCas9 and M-MLV reverse transcriptase, combined with structured RNA motifs, enables efficient editing and characterization of DNA variants at base-pair resolution.
Enables accurate genomic annotation and identification of functional DNA variants, facilitating disease risk prediction and therapeutic target identification through enhanced editing efficiency and stable expression.
Smart Images

Figure 2026509250000001 
Figure 2026509250000002 
Figure 2026509250000003
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims priority to U.S. Provisional Patent Application No. 63 / 489,418, filed on March 9, 2023. The entire disclosure of this provisional patent application is incorporated herein by reference for all purposes.
[0002] [Government Support] This invention was made with government support under grant numbers EY027789, HG009402, AG057497, and DA052713 awarded by the National Institutes of Health. The United States government has certain rights in this invention.
Background Art
[0003] Introduction Advances in genome sequencing have led to the identification of hundreds of millions of genetic variants in the human population, some of which confer risk for common diseases such as diabetes, neuropathy, and cancer. 1 A major barrier to understanding the genetic basis of these complex diseases is the lack of functional annotation for disease - related variants, especially since such variants are predominantly located within non - coding regions. Evidence is accumulating that non - coding risk variants can contribute to disease etiology by disrupting gene regulation. Even variants encoding proteins found in individuals with disease are often classified as variants of unknown significance (VUS). Therefore, to realize the potential of personalized medicine, there is a need for more accurate and high - throughput functional characterization methods that resolve the function of disease - related variants at base - pair resolution and are multiplexed across genomic loci.
[0004] With the development of genome editing technologies, it has become possible to perturb and evaluate DNA sequences on a large scale in desired regions. However, there are still fundamental barriers to using these methods for precise genome annotation. For example, CRISPRa, CRISPRi, CRISPR deletions, and CRISPR indels have been applied in gene screening strategies that characterize both genes and cis-regulatory regions 2 , but it has not been possible to accurately identify the variants that cause disease. Conventional methods for characterizing DNA variants (SNPs) by knock-in via homologous recombination are inefficient and have low throughput. Base editors also have limitations, and specific mutations (C→T, A→G, T→C, G→A) are introduced with various target efficiencies 3 . Therefore, there are still significant deficiencies in methods for effectively characterizing the roles of variants presumed to cause disease in human health and disease. To achieve a better understanding of the genetic basis of diseases, there is an urgent need for robust and high-throughput methods for performing desired edits at base pair resolution.
[0005] Prime editing (PE), a versatile and precise genetic engineering method, was developed to introduce all types of edits, including point mutations, insertions, and deletions 4 . In particular, in PE2, Streptococcus pyogenes Cas9 (SpCas9) H840A nickase and Moloney murine leukemia virus (M-MLV) reverse transcriptase are used. The spacer in the prime editing guide RNA (pegRNA) guides the Cas9 nickase and M-MLV complex to the target site, while the RT template sequence provides the desired editing information. Therefore, both targeting information and editing information can be easily programmed in the same pegRNA to perform single nucleotide substitutions, insertions, or deletions. A newer version of PE, PE3, can further enhance editing efficiency by using an additional single guide RNA (sgRNA) for nicking to facilitate replacement of the non-edited strand 5The precision genome editing capabilities of Prime Editor suggest the potential for high-throughput variant-level genome manipulation. Recently, VUS at the NPC1 locus was identified using PE screening based on lysosomal functional assays involving pegRNA transfection and targeted sequencing of this region. 6 Transient transfection of the PE mechanism followed by targeted sequencing of the edited locus allows for the identification of the editing event, but its scope is limited to that specific locus, making it impossible to scale up and evaluate a large number of loci in parallel. In addition to improved throughput, improvements in the control of the transgene copy number, stable expression of the PE mechanism, and direct comparison of loci are also desired. [Overview of the project]
[0006] This invention provides a method and composition for identifying functional DNA variants in the human genome using a high-throughput prime editing screen.
[0007] In one embodiment, the present invention provides a high-throughput prime editing screen for identifying functional DNA variants in the human genome, configured substantially as disclosed herein.
[0008] In one embodiment, the present invention provides a gene screening platform for identifying functional variants related to human health and disease, configured substantially as disclosed herein, and preferably for annotating the genome with nucleotide resolution for disease prediction and treatment practical for personalized medicine.
[0009] In one embodiment, the present invention provides a pooled prime editing screen system configured substantially as disclosed herein, which can characterize gene variants at base pair resolution and on a base pair scale, enabling accurate genome annotation for disease risk prediction, diagnosis, and identification of therapeutic targets.
[0010] In embodiments, the present invention provides a screen, platform, system, or method substantially disclosed herein, which involves infecting host cells with a lentivirus containing a stable expression cassette of nCas9 and M-MLV reverse transcriptase (Figure 1b) to obtain transformed cells with stable expression of nCas9 / M-MLV, enabling more efficient pegRNA / ngRNA packaging and lentiviral delivery with higher editing efficiency than co-infection methods (Figure 1d), thereby increasing PE efficiency and facilitating a pooled screening approach using a lentiviral library.
[0011] In embodiments, the present invention relates to a screen, platform, system, or method substantially disclosed herein, wherein the host cell is pegRNA 7~9 When pegRNAs containing a scaffold structure RNA motif, e.g., (EvopreQ1, MLV-PK1, and MLV-PK2), at their 3' end are treated with a screen, platform, system, or method, preferably MCF7-nCas9 / RT cells, and lentiviral delivery of both pegRNAs containing scaffold 1 and ngRNAs containing scaffold 2 in the same construct, are shown to exhibit higher editing efficiency at both the EMX1 and FANCF loci compared to using PEs without a structure RNA motif (Figure 5c). (Figure 1e)
[0012] The present invention encompasses all combinations of the specific embodiments listed herein, as if each combination were listed individually. [Brief explanation of the drawing]
[0013] [Figure 1a-e]This figure shows the optimization of PE efficiency in mammalian cells using lentiviral delivery. (a) Various strategies tested to optimize PE efficiency in the MCF7 cell line. Top: Delivery of the PE mechanism by co-infection with three different viruses. Bottom: Dual pegRNA / ngRNA virus infection of clone MCF7 lines stably expressing nickase Cas9 (nCas9) and Molony's mouse leukemia virus reverse transcriptase (M-MLV RT). Two scaffolds and three different structured RNA motifs tested are also shown. (b) Lentiviral constructs that generate MCF7 clones expressing nCas9 / RT. PuroR, puromycin resistance gene. M-MLV RT, Molony's mouse leukemia virus reverse transcriptase. (c) RT-qPCR analysis showing the relative expression of nCas9 / RT in different clones, normalized to dCas9 expression in established CRISPRi iPSC lines (yellow). (d) Editing efficiency and indel rate for the EMX1 and FANCF loci 2 and 4 weeks after PE introduction using two different RNA scaffolds. (e) Vectors improved for pegRNA and ngRNA expression in PE screening. RTT: reverse transcription template, PBS: primer binding site. [Figure 2a-h]This figure shows the functional characterization of MYC enhancers by saturated mutagenesis PE screening. (Top) The target enhancer is downstream of MYC. (Bottom) The enhancer region is highly enriched with ATAC-seq, H3K27ac, and ChIP-seq signals for H3K4me1. The blue region indicates the region selected for PE screening. (b) (Top) Figure showing the design of PE saturated mutagenesis screening with a 716 bp enhancer. Each nucleotide was subjected to PE substitution with three nucleotides. (Center) Each substitution event was covered by three uniquely designed pegRNA / ngRNA pairs. (Bottom) PE screening workflow. (c) Log2 (multiple of change) of each substitution at each base pair, sorted by genomic location. Mutations with a significant effect on cellular fitness are colored. Conservation scores calculated by ATAC-seq signals and PhastCons are shown. (d) JARVIS scores for base pairs with different numbers of significant substitutions. The box plots show the median, IQR, Q1 - 1.5×IQR, and Q3 + 1.5×IQR. Outliers are shown as gray dots. Means are shown as red dots. P-values were calculated using a two-sided two-sample t-test. (e) Creation of a functional PWM to identify potential TF binding sites. (f) (Top) ChIP-seq signals of six TFs in MCF7. The blue region indicates the core enhancer region. (Bottom) Sequence logo plot for the core enhancer region generated by the functional PWM from (e). (g) Matched TF binding sites. (h) (Top) Dense tracks showing nucleotide importance scores derived by the BPNet model for GATA3 and ELF1 binding sites. [Figure 3a-j]This figure shows how PE screening can reveal functional SNPs associated with breast cancer. (a) Overview of the design of the Alt and Ref libraries. The design included breast cancer-related variants (SNPs), clinical variants (ClinVar), stop codon introduction (iSTOP), and untargeted controls. For each variant, a pegRNA / ngRNA pair was designed to introduce either the Alt allele or the Ref allele. (b) Workflow of PE screening using the Alt and Ref libraries. MCF7-nCas9 / RT cells were infected with either lentiviral library. Cells were collected on day 2 and day 32 post-infection. The abundance of pegRNA / ngRNA pairs in the samples collected on day 2 and day 32 was deep-sequenced. The relative effect of each variant was determined based on the relative impact on cell growth between the Alt allele and the Ref allele. (c) Percentage of significant hits (FDR < 0.05) identified from Alt and Ref PE screens for Alt / Alt genotype, Het genotype, and Ref / Ref genotype in MCF7. (d) Functional SNPs (red) that have either a positive or negative effect on cell growth were determined by the relative effect in the Alt vs. Ref screen. Blue dots represent significant iSTOPs, and black dots represent controls. The red dashed line indicates an FDR of 0.05. (e) The absolute effect of identified functional iSTOPs and SNPs is higher than the effect of the negative control (P-value calculated by a two-sided two-sample t-test). (f) Genomic distance of SNPs tested at each risk locus for TSS of each gene. Red dots are functional SNPs within the gene body, blue dots are functional SNPs in the terminal region, and gray dots are SNPs with a non-significant effect. (g) Relative enrichment of genomic features of identified functional SNPs (P-values calculated by two-sided Fisher exact test). The number of SNPs overlapping with each genomic feature is indicated next to each bar. (h) Venn diagram showing the number of unique transcription factors (TFs) with different binding sites around functional SNPs. The number of SNPs that modify TF binding sites is also shown in parentheses. (i, j) Examples of functional SNPs that disrupt TF binding sites.(i) The Alt protective allele at rs12275749 (position shown in f) affects the SMAD3 binding site, and (j) the Alt risk allele at rs66473811 (position shown in f) coincides with the MAZ binding motif. [Figure 4a-h] This figure shows functional clinical variants identified using PE screening. (a) Functional clinical variants (red) that have either a positive or negative effect on cell growth were determined by the relative effect on cellular fitness between the Alt and Ref alleles. Blue dots represent significant iSTOPs, and black dots represent negative controls. The red dashed line shows the 5% FDR. (b) The effect sizes of identified functional iSTOPs and clinical variants are larger than the effect sizes of negative controls (P-values were calculated by a two-sided 2-sample t-test). The box plots show the median, IQR, Q1 - 1.5 × IQR, and Q3 + 1.5 × IQR. Red dots show the mean. (c) CADD scores for iSTOPs and clinical variants. (d) Number of identified functional VUSs (N, nonpolar, P, polar, Pc, positively charged, Nc, negatively charged) that cause transposition of each amino acid group. (e, f) Lollipop plots mapping functional VUS in RAD51C and BARD1 to their standard isoforms. Significant VUSs identified are shown in red. Their effects on cell growth are shown by change multipliers. (g, h) Lollipop plots mapping nonsense variants in BRCA1 and BRCA2 to their standard isoforms. Significant hits identified are shown in blue. Their effects on cell growth are shown by change multipliers. [Figure 5a-c]This figure shows the optimization of PE efficiency in the MCF7 cell lineage. (a) Prime editing efficiency and indel rate by co-infection with pegRNA-expressing lentivirus, ngRNA-expressing lentivirus, and nCas9 / RT-expressing lentivirus in MCF7 cells. (b) Immunofluorescence staining showing the localization of nCas9 / RT (red, FLAG-tagged) in the nucleus (blue, DAPI) in MCF7-nCas9 / RT cells. Scale bar, 1000 μm. (c) Editing efficiency and indel rate by PE using three different structured RNA motifs at the 3' end of pegRNA 2 weeks and 4 weeks after infection of MCF7-nCas9 / RT cells. [Figure 6a-f] This figure characterizes enhancer function and PE screening results in MCF7 cells. (a) CRISPR / Cas9 knockout of the MYC enhancer in MCF7 reduced MYC expression. P-values were calculated using a two-sided two-sample t-test. (b) Distribution of read counts of pegRNA / ngRNA pairs in the cloned plasmid library. (c) PCA analysis demonstrates high reproducibility of PE screening between biological replicates. (d) Correlation between the location of PE-induced stop codons and their effect sizes. Blue lines and P-values were calculated using a generalized additive model. Shaded areas indicate 95% confidence intervals. (e) (Top) Location of sensitive base pairs (SBPs) with three significant substitutions. (Bottom) Cumulative distribution plot of SBPs with three significant substitutions along the MYC enhancer, and the formula for calculating the slope of each contiguous bin. (f) Line graph of slopes for each contiguous bin along the MYC enhancer. The red dashed line is the cutoff for significant slopes based on slopes with a P-value equal to 0.05 derived from the Z-score. The red region represents the core enhancer region, derived from bin tilts greater than the cutoff (tilt > 0.43). [Figure 7a-c]This figure shows strategies for prioritizing genomic loci and clinical variants in PE screening. (a) MCF7 growth-related genes were selected from CRISPR / Cas9 knockout and base-editing screens in MCF7 cells. (b) Strategies used to select breast cancer-related SNPs for prime editing screens. (c) Strategies used to select clinical variants for prime editing screens. [Figure 8a-f] This figure shows the quality control and primary analysis of disease variant PE screens. (a) Heatmap of read counts by pairwise correlation and hierarchical clustering for prime editing screens. (b) Pearson correlation between the log2(change factor) of iSTOP in the Alt library screen and the log2(change factor) of gRNA in the CRISPR / Cas9 knockout screen for each target gene. (c) Volcano plot of results from the Alt library screen. (d) Volcano plot of results from the Ref library screen. (e) log2(change factor) of each iSTOP from the Alt library screen and the Ref library screen. (f) Violin plot showing the 5% FDR cutoff used for relative effect analysis comparing the Alt and Ref libraries. The numbers above the peaks indicate significant data points relative to the total data points in each category when using a 5% FDR. The 5% percentile of the P-value from the negative control was used as the empirical significance threshold to achieve a 5% false discovery rate (FDR) shown by the red dashed lines in d-f. [Figure 9a-c]This figure shows examples of functional VUSs and their potential consequences. (a) Sequence conservation of RAD51 family proteins. Alignment of RAD51 family proteins using MUSCLE. Functional VUSs identified by PE screening in RAD51C are labeled. (b) Figure showing the binding region between BARD1 and BRCA1. (c) Protein structure of the BARD1-BRCA1 complex predicted by Alphafold. Two hydrogen bonds were identified between wild-type His36 in BARD1 and Asp96 in BRCA1, but were lost due to the His36Pro mutation in BARD1. [Modes for carrying out the invention]
[0014] Unless otherwise specified or contraindicated, in these detailed descriptions and throughout this Spec., singular nouns that do not specify a quantity ("a" and "an") refer to one or more, and the term "or" means "and / or". The examples and embodiments described herein are for illustrative purposes only, and various modifications or changes taking them into account will be suggested to those skilled in the art and will be understood to be included in the spirit and scope of this application and the appended claims. All publications, patents and patent applications cited herein, including their references therein, are incorporated herein by reference in their entirety for any purpose.
[0015] Example: Identification of functional DNA variants in the human genome by high-throughput prime editing screening. Summary: Despite significant advances in the detection of DNA variants associated with human disease, interpreting their functional effects at high throughput and base-pair resolution remains challenging. Here, we have developed a novel pooled prime editing screen method that, when applied, allows for the highly reproducible characterization of thousands of coding and non-coding variants in a single experiment. To demonstrate its application, we first identified a nucleotide essential to the 716 bp MYC enhancer via prime editing-mediated saturation mutagenesis. Next, we applied the prime editing screen to functionally characterize 1304 non-coding variants associated with breast cancer and 3699 variants from ClinVar. 103 non-coding variants and 156 variants of unknown significance were found to function by affecting cellular fitness. Overall, this demonstrates a pooled prime editing screen technology that enables accurate genomic annotation for characterizing gene variants at base-pair resolution and on a base-pair scale, leading to disease risk prediction, diagnosis, and the identification of therapeutic targets.
[0016] Here, the inventors have optimized prime editing (PE) in mammalian cells to enable high-throughput pooled screening of thousands of DNA variants in the human genome via lentiviral delivery. The usefulness of the inventors' novel PE screening approach is demonstrated for three different applications, including saturated mutagenesis analysis of a 716 bp enhancer, functional characterization of 1304 breast cancer-related variants, and evaluation of the impact of 3699 clinical variants on cellular fitness. The inventors' results establish the generalizability of a pooled PE screen for accurately characterizing gene variants in the human genome.
[0017] Optimization of PE efficiency in mammalian cells delivered by lentiviruses To enable PE screening by lentiviral delivery, first, PE3 was introduced into MCF7 cells by infecting them with three different viruses: 1) a virus expressing Cas9(H840A) nickase (nCas9) and Moloney murine leukemia virus reverse transcriptase (M-MLV RT), 2) a virus expressing pegRNA, and 3) a virus expressing nick sgRNA (ngRNA). Unfortunately, with this strategy, only PE efficiencies of less than 1% were obtained with relatively high indel rates. This is because the efficiency of co-infecting the same cells with three different viruses is low (Figure 1a, Figure 5a).
[0018] It is difficult to package all PE3 components within the same virus. To enhance PE efficiency and facilitate a pooled screening approach using a lentiviral library, MCF7 cells were infected with a lentivirus containing a stable expression cassette for nCas9 and M-MLV RT (Figure 1b). After puromycin selection, numerous clones were isolated, and the clone with the highest nCas9 expression (Figure 1c, RT-qPCR, clone number 4, Figure 5b) was selected for subsequent experiments. Stable expression of nCas9 / M-MLV RT enables high-efficiency packaging and lentiviral delivery of pegRNA / ngRNA with higher editing efficiency than the co-infection method (Figure 1d). To further improve PE efficiency, three different structured RNA motifs (EvopreQ1, MLV-PK1, and MLV-PK2) were used at the 3' end of pegRNA 7~9 to evaluate editing efficiency. Cells treated with pegRNA containing a scaffold structure RNA motif consistently showed higher editing efficiency at both the EMX1 locus and the FANCF locus compared to using PE without a structured RNA motif (Figure 5c), so EvopreQ1 was added to the pegRNA design for all pooled screens. Scaffold 1 5 and Scaffold 2 10Since this did not have a significant effect on PE efficiency, it suggests the feasibility of dual delivery of pegRNA and ngRNA from the same viral particle (Figure 1d). All PE experiments in cloned MCF7 cells (MCF7-nCas9 / RT) showed relatively low indel rates (0.7%~1.95%). Therefore, we used MCF7-nCas9 / RT cells, as well as lentiviral delivery of both pegRNA containing scaffold 1 and ngRNA containing scaffold 2 from the same construct (Figure 1e).
[0019] The Prime Edit screen enables nucleotide-resolution analysis of the enhancer function. Enhancers can modulate cell-type-specific gene expression, where disease-related variants are highly enriched. Knowledge of the endogenous function of each nucleotide in enhancers should reveal key transcription factors influencing enhancer activation, leading to the development of better models of gene regulatory networks and easier prediction of the regulatory effects of disease-related non-coding variants. To test whether PE screens can quantify the effect of each base in enhancers, a CRISPRi screen is being considered. 11 We focused on an MCF7-specific MYC enhancer identified from [source]. This enhancer is located 405kb downstream of MYC and, in addition to forming a chromatin loop with the MYC promoter, exhibits an enhancer signature including open chromatin, H3K27ac, and H3K4me1 signaling (Figure 2a). Deletion of this enhancer resulted in 85% downregulation of MYC expression in MCF7 cells, confirming its enhancer activity in regulating MYC expression (Figure 6a). Since the downregulation of MYC correlates with the survival of MCF7 cells, 12 We then performed a high-throughput saturated mutagenesis screen using PE of this MYC enhancer in MCF7 cells, depending on the cell survival phenotype (Figure 2b).
[0020] To investigate the enhancer's function in detail at base-pair resolution, we designed a library of 6252 pegRNA / ngRNA pairs that generate 2148 single-nucleotide substitutions within a 716 bp MYC enhancer region. Specifically, the original base was replaced with three other nucleotides, and each event was independently evaluated three times on the same screen (Figure 2b). We also included 94 positive control pegRNA / ngRNA pairs and 400 negative control pegRNA / ngRNA pairs with stop codons (iSTOP) introduced into the MYC coding region. Of the negative controls, 246 targeted non-human genomes, and 154 targeted the AAVS1 safe harbor locus. Subsequently, MCF7-nCas9 / RT cells were infected with lentiviral libraries expressing these pegRNA / ngRNA pairs (Figure 6b). Two days after infection, the transduced cells were selected for hygromycin for one week and then grown in normal medium for a further three weeks. Cells were collected 2 and 30 days after infection, and integrated pegRNA / ngRNA pairs were amplified. The relative decrease or enrichment of each pegRNA / ngRNA between these two time points was determined by deep sequencing (Figure 2b). This screen was performed three times (Figure 6c), and negative controls containing non-human targeted pegRNA / ngRNA pairs and AAVS1 targeted pegRNA / ngRNA pairs were used to normalize the data. MAGeCK pipeline 13 We used [a specific method / tool] to calculate the change multiplier (FC) of each pegRNA / ngRNA pair between samples taken 2 days post-infection and 30 days post-infection. As expected, 78% (73 / 94) of iSTOP decreased after 30 days post-infection (log2FC<0). The rate of iSTOP decrease was negatively correlated with the distance from the transcription start site (TSS) of MYC, suggesting that gene knockout becomes more efficient when a perturbation is introduced at the 5' end. 14This is consistent with (Figure 6d). Furthermore, two iSTOPs (amino acid positions 350 and 355) that target the region between the nuclear localization signal (NLS) and the carboxy-terminal domain (CTD) were also significantly reduced (Figure 6d). The N-terminus of MYC contains a core transcriptional transactivation domain that binds to numerous partners. 15 These two iSTOPs may impair the function of wild-type MYC and its cofactors, as they generate truncated MYC that can still bind to cofactors but cannot bind to MYC DNA targets.
[0021] To investigate the effect of each nucleotide on enhancer function, nucleotides that affect cell fitness when at least one substitution occurs (FDR<0.05, |log2FC|>1) were defined as sensitive base pairs (SBPs). Of the 716 base pairs tested, 334 (46.6%) were SBPs with log2FC<-1, indicating that mutations at these positions reduce enhancer activity and cell fitness. For all three substitutions, 23.1% (77 / 334) of SBPs decreased at day 30 (FDR<0.05, log2FC<-1). Furthermore, none of the tested sequences were significantly enriched in terms of increased cell growth phenotype at day 30, indicating that perturbations of these sequences attenuated enhancer activity only (Figure 2c).
[0022] Deep learning models have been developed that prioritize non-coding regions and predict their association with human diseases. Encouragingly, SBPs with two or more significant substitutions (n=172) are classified as JARVIS. 16This predicted that SBPs with only one significant substitution (n=162) or non-SBPs (n=382) would be more harmful than those with only one significant substitution (Figure 2d). This supports the success of PE screening in verifying computationally predicted functional sequences. Furthermore, serial bin density analysis was established to detect changes in SBP density along enhancers and define regions where SBPs were concentrated (Figures 6e and 6f). Based on the slope value of the cumulative curve of SBPs with three significant substitutions, core enhancer regions containing high-density SBPs in the enhancer were identified. A larger slope value indicates a higher density of SBPs in that region. Core enhancer regions were defined by a minimum slope cutoff of 0.43 (P<0.05 derived from the Z score). Core enhancer regions (chr8:128142093~128142181, hg38) colocalized with the open chromatin summit. This region contains SBPs that exhibit the most extensive change multipliers when mutated, indicating its strong effect on enhancer activity (Figure 2c, highlighted in purple). Notably, the enhancer's core sequence was located adjacent to a highly conserved region (Figure 2c). This is because enhancers undergo rapid evolutionary changes compared to protein-coding sequences. 17 This is not surprising.
[0023] Our functional data provides a unique opportunity to compute and construct a position-weight matrix (PWM). We created a functional PWM using the change multiplier for each nucleotide substitution from the PE screen (Figure 2e). Our functional PWM and the JASPAR, HOCOMOCO, and SwissRegulon databases. 18~20 By comparing the curated transcription factor (TF) motifs with those from the dataset, 13 TFs with matching motif PWMs were identified (Figures 2g and 2h). ENCODE ChIP-seq dataset 21 Based on this, it has already become clear that five predicted TFs (GATA3, ELF1, FOXM1, MTA3, and RCOR1) will be coupled to the MYC enhancer, and the ENCODE project 22Through Avocado, YY1 is predicted to bind to this enhancer in MCF7 (Figure 2f). Furthermore, GATA3 and YY1 are connected to MCF7 23 Since it is an essential cell survival gene in [the cell], the usefulness of saturated mutagenesis that can use PE to investigate enhancer function in base pair resolution is confirmed. The essential nucleotides for the binding motifs of ELF1 and GATA3 identified by our PE screen are BPNet 24 The significance of the quantitative role of each nucleotide identified by our PE screen was further validated, as it was consistent with what was imputed by the other method. Together, the pooled PE screen demonstrated its usefulness in elucidating the nucleotide-resolution functional annotation of non-coding cis-regulating elements.
[0024] Characterization of breast cancer-related variants Next, we tested the feasibility of characterizing over 5,000 disease-associated DNA variants at various genomic loci, including non-coding variants from GWAS and variants detected from clinical samples. For variants identified in GWAS, we focused on breast cancer, the most common cancer among women in the United States. To test the feasibility of characterizing DNA variants associated with breast cancer, we primarily used samples from people of European descent. 25 Summary statistics from the largest GWAS to date were used, including this GWAS. 26 From a comprehensive fine mapping approach to CRISPR screens 23、27Candidate genes overlapping with growth phenotype genes prioritized by [the specified criteria] were selected. Candidate genes include CCND1, PSMD6, MYC, UBA52, DYNC1I2, ESR1, MRPS18C, NOL7, EWSR1, BRCA2, and GRHL2, which were negatively selected in the CRISPR knockout screen, as well as tumor suppressor genes CUX1, CASP8, and TNFSF10, which were positively selected in the CRISPR knockout screen (Figure 7a). Next, 1304 single nucleotide polymorphisms (SNPs) within 500 kbp upstream and downstream of these genes were selected (Figure 7b). These SNPs have been previously associated with breast cancer. 25 It had been suggested that these genes could potentially function. 26 Additionally, 3,699 variants were selected from the ClinVar database (Figure 7c). Of these, 2,840 were identified from patients who underwent screening for hereditary breast cancer. 28 To systematically evaluate the impact of variants on cellular fitness, two libraries were designed: a library containing a reference allele (Ref library) and a library containing an alternative allele targeting the selected variant (Alt library) (Figure 3a). Each library included 250 untargeted pegRNA / ngRNA pairs as negative controls. The Alt library included 115 pegRNA / ngRNA pairs with stop codons (iSTOP) introduced into 23 MCF7 growth-related genes as positive controls, while the Ref library used pegRNA / ngRNA pairs with reference sequences introduced into those loci. Cloned plasmids were packaged into lentiviral libraries and transduced into MCF7-nCas9 / RT cells. Cells were collected 2 and 32 days after infection, pegRNA / ngRNA pairs were amplified, and deep-sequenced (Figure 3b). Iterations of PE screening using either the Ref library or the Alt library (n=4) were reproducible at the read count level (Figure 8a).
[0025] Alt library screening revealed that 33.04% (38 / 115) of iSTOP showed a significant cell fitness effect (FDR < 0.05). This is comparable to the 31.8% positive rate of iSTOP for common essential genes reported from base editing screening in MCF7 cells. 29 Furthermore, the change factor for iSTOP showed a high correlation with the change factor for sgRNA from the MCF7 CRISPR knockout screen of the same gene. 23 (Figure 8b). Compared to day 2, both the Alt PE and Ref PE screens on day 32 showed a greater decrease in pegRNA / ngRNA pairs (FDR<0.05, Alt PE n=322 and Ref PE n=337) than in the enrichment of pegRNA / ngRNA pairs (FDR<0.05, Alt PE n 284=148 and Ref PE n=209) (binomial test, P=4.78×10 for Alt PE). -8 In Ref PE, P = 6.85 × 10 -16(Figures 8c and 8d). Theoretically, if the designed peg / ngRNA pairs match the wild-type MCF7 genotype, they should have no effect on cell growth. However, notably, certain pegRNAs matching the wild-type MCF7 genotype showed a significant effect on cell growth beyond prediction, while the proportion of significant hits for each genotype group was independent of the original MCF7 genotype (chi-square test P=0.9998 in the Ref library, P=0.999 in the Alt library, and Cochran-Mantel-Haenszel test P=0.9665 when the Ref and Alt libraries were combined). For example, in the Ref library, 11.2% (59 out of 528) of pegRNAs were significantly reduced at sites with the Ref / Ref MCF7 genotype, similar to 10.2% (55 out of 540) at heterozygous sites and 7.9% (18 out of 227) at Alt / Alt genotype sites (Figure 3c). These changes at sites where allele changes were not expected suggest that, when editing mechanisms are recruited to target sites, there are undesirable effects on constitutive nCas9 expression, similar to CRISPR inhibition (CRISPRi). 30 To test the potential CRISPRi activity of nCas9 in PE, we compared results between iSTOP in the Alt library and the corresponding pegRNA / ngRNA pairs in the Ref library. PegRNAs in the Ref library showed a smaller effect compared to iSTOP targeting the same locus on day 32, but these were still depleted on day 32, confirming the unintended effects of nCas9 occupation of the target genomic locus (Figure 8e). Taken together, prolonged PE expression was found to exhibit undesirable activity similar to CRISPRi, which is an important factor to consider when analyzing lentivirus-mediated PE screens.
[0026] To correct this undesirable PE activity, DESeq2 31The ratio of FCs for each pegRNA / ngRNA pair from the Alt PE screen and Ref PE screen was compared. Functional SNPs were determined based on their relative effects on cell growth between Alt PE and Ref PE. In total, 56 SNPs with the Ref allele and 47 SNPs with the Alt allele were identified as promoting cell growth (P<0.05, empirical significance threshold to suppress type I error was 5%, Figure 8f, Figure 3d). As expected, the identified functional SNPs had smaller effect sizes than stop codons and significantly larger effect sizes than negative control PE (Figure 3e). Furthermore, the iSTOP for cell growth-promoting genes such as MYC and GATA3 was decreased, while the iSTOP for the cell growth inhibitor PTEN was enriched, thus validating our analytical approach (Figure 3d).
[0027] Since risk variants can be either Ref or Alt alleles, functional SNPs were further annotated based on the gene annotation of breast cancer risk variants. Because most GWAS SNPs are unlikely to be causative, only a small fraction of the 1304 SNPs tested were expected to exhibit a biological effect. Using CAVIAR, the mean likelihood of causative variants was calculated, and assuming only one causative variant in each linkage disequilibrium (LD) clamp, the mean expected value for causative variants was approximately 8.9%. If more than one causative variant was found in each LD clamp, the mean probability of a variant being causative was approximately 13.0%. Compared to the reference allele, 50 surrogate alleles of risk SNPs promoted growth, while 53 surrogate alleles of risk SNPs reduced cell growth (Figure 3f). 18.45% (19 / 103) of the functionally validated risk SNPs were located within the risk gene itself. The remaining loci were located in the distal region, at an average distance of 185.8 kb from the TSS of the risk gene (Figure 3f). All tested loci contained at least one SNP with a significant effect on cell growth, except for the BRCA2 locus, where only two SNPs were tested. Finally, the identified functional SNPs significantly enriched active chromatin marks, including ATAC-seq signals, H3K27ac signals, H3K4me1 signals, and H3K4me3 signals, compared to their corresponding genomic background (1 Mbp surrounding the selected cell growth genes) (two-sided Fisher exact test, P<0.05) (Figure 3g).
[0028] To explore potential mechanisms for the regulation of changes in cellular fitness by functional SNPs, we used 40 bp regions centered on 103 identified functional SNPs in the human motif database HOCOMOCO. 19Candidate TF-binding motifs were searched for. 281 motifs containing Alt and Ref alleles, and 391 motifs (FDR < 0.05, and TF expression > 1 FPKM) were obtained. After removing redundant motifs for each SNP locus, 90 TF-binding sites were identified for 35 unique TFs associated with a cell growth suppression phenotype (log2FC(Alt / Ref) < 0), and 55 sites were identified for 29 unique TFs associated with a cell growth promotion phenotype (log2FC(Alt / Ref) > 0) (Figure 3h). In particular, the Alt allele (protective allele) rs12275479 (T>C) at the CCND1 locus disrupted the SMAD3-binding motif and was associated with reduced cell growth in the PE screen, and the TGFβ-SMAD3 axis was linked to MCF7 32 This is consistent with the reduction in the number of mamosphere initiation cells (Figures 3f and 3i). In another example, it was found that the MAZ binding site of MAZ is affected by the rs66473811(T>C)Alt allele at the PSMD6 locus. MAZ drives tumor-specific expression of the PPARγ1 gene and regulates MYC expression. 33、34 This transcription factor promotes the proliferation of breast cancer cells, which is consistent with the Alt allele being a risk allele (Figures 3f and 3j). Together, these results support the use of pooled PE screens to functionally characterize variants identified by GWAS.
[0029] Genetic variants detected in clinical samples provide valuable resources for understanding the pathogenesis of human diseases. However, many clinically found variants, even in well-characterized protein-coding genes, are annotated as variants of unknown significance (VUS) because their functional effects are unpredictable. To evaluate the ability of PE screens to functionally annotate VUSs using MCF7 growth phenotypes, pegRNA / ngRNA pairs were designed for 2532 VUSs, 745 pathogenic variants, and 422 benign variants across 17 genes (Figure 7c). 76.78% of the variants tested came from breast cancer patients. By comparing the relative effect sizes of each pair of Alt and Ref alleles, 236 functional clinical variants affecting cell growth were identified across 15 genes, including 49 pathogenic variants, 156 VUSs, and 31 benign variants (Figure 4a). The mean effect sizes for pathogenic variants, VUS, and benign variants were between the mean effect size of the negative control and the mean effect size of iSTOP (Figure 4b).
[0030] Several computational indicators have been used to assess the harmfulness of variants. 35、36 One such method is CADD, which integrates diverse genomic annotations into a single quantitative score to estimate the relative pathogenicity of human gene variants. 35iSTOP and pathogenic variants have similarly high CADD scores compared to other categories (Figure 4c). CADD scores for VUS and benign variants show a wide distribution, where the median score is much lower than that of iSTOP and pathogenic variants. Interestingly, the CADD scores for functional variants identified within the VUS or benign variant group did not have the expectedly higher CADD scores, which highlights the limitations of relying solely on computational prediction for variant annotation and emphasizes the importance of validating clinical variants by functional assays, even for variants located in well-studied protein-coding genes. For example, one benign variant (Arg378Ser) in BARD1 with a low CADD score (CADD=4.317) would not be classified as functional. However, based on our PE screening results, this variant showed a significant cell growth inhibitory effect in MCF7 cells. BARD1 (Arg378Ser) may impair the nuclear localization of the BRCA1 / BARD1 complex and potentially synergistically promote tumorigenesis in vivo with BARD1 (Pro24Ser). 37 Furthermore, most of the identified functional VUSs are missense variants, and approximately half of the significant VUSs from our screening exhibit a change in amino acid type within the same group based on polarity (Figure 4d), complicating the determination of their molecular effects. Our results provide novel insights into the potential role of clinical variants in disease pathogenesis through their modulation of cellular fitness and provide annotations for previously uncharacterized VUS variants and benign variants. Functional and structural domains are essential inducers for protein function. 60% of the identified functional VUSs are found in the UniProt database. 38These variants are located within annotated protein domains, and their pathogenicity is supported. For example, eight VUSs were identified in RAD51C (Figure 4e), a cancer susceptibility gene and a gene essential for MCF7 survival. Two variants, one in the RAD51C functional domain related to Holliday junction processing (amino acids 1-126) (Pro21Leu) and the other in the NLS region (amino acids 366-370) (Arg366Gln), were associated with reduced cell growth as determined by our PE screen (Figure 4e). Functional variants not found in any annotated domain were also identified, including a functional RAD51C VUS (Arg312Gln) associated with the phenotype of reduced MCF7 growth (Figure 4e). Arg312Trp in RAD51C causes homologous recombination repair deficiency and reduced colony formation phenotype in MCF10A cells, thus abolishing the RAD51C-RAD51D interaction. 39 Arg312Gln may have similar pathogenic effects on protein function. A comparison of the RAD51C sequence with other RAD51 family proteins revealed that the functional VUS is located in both conserved and non-conserved amino acids (Figure 9a), highlighting the difficulty of predicting variant function based solely on protein sequence conservation.
[0031] Protein-protein interactions (PPIs) are another essential functional activity in many biological processes. In this study, functional VUSs located in protein-binding regions that may influence PPIs were also identified. For example, BARD1 interacts with BRCA1 via its RING domain, and the BRCA1-BARD1 ubiquitin ligase activity is essential for DNA double-strand break repair. 40、41Since a functional VUS (His36Pro) was identified in the RING domain of BARD1 (Figure 4f), it is suggested that this clinical variant has a structural effect that affects BARD1-BRCA1 heterodimer formation (Figure 9b). Consistent with these findings, AlphaFold predicts that the His36Pro variant inhibits the formation of hydrogen bonds between His36 in BARD1 and Asp96 in BRCA1 (Figure 9c).
[0032] Nonsense mutations can generate new stop codons and truncated proteins. While most ClinVar mutations are annotated as pathogenic variants, many of their functional effects remain unclear. 28 In our PE screen, we tested 563 nonsense clinical variants across 13 breast cancer risk genes, identifying 38 variants as positive hits in 7 genes. Surprisingly, 39.47% (15 / 38) showed unexpected phenotypes compared to the cell death knockout phenotype of these genes. Specifically, while a similar number of functional nonsense variants were identified in BRCA1 (n=15) and BRCA2 (n=16) (Figures 4g, 4h), 60% (9 / 15) in BRCA1 were able to promote MCF7 cell growth, compared to only 25% (4 / 16) in BRCA2. After locating the variants within BRCA1 and BRCA2, we noticed that the truncated proteins resulting from all gain-of-function nonsense variants in BRCA1 still retained their NLS. These results were confirmed by a different nonsense mutation at Q858, located downstream of the NLS in BRCA1, resulting in a truncated BRCA1 with the NLS, which led to increased cell growth in MCF7. 29 However, in all functional variants identified in BRCA2, their NLS was located at the C-terminus. 42This was removed from the truncated protein, leading to a loss of nuclear localization of BRCA2. Overall, these results demonstrate that PE screening has the ability to functionally characterize some nonsense mutations.
[0033] Consideration The inventors have developed a "search and replace" prime editing tool. 5、9 This paper describes a novel genome screening method for investigating DNA function at base-pair resolution by employing and optimizing a pooled prime editing screen that identifies essential nucleotides in MYC enhancers via a saturated mutagenesis screen, demonstrating the functional characterization of 1304 breast cancer-related risk SNPs and providing accurate annotations for 3699 clinical variants. The inventors' research presents a novel strategy for elucidating genomic function with unprecedented precision and scale. The wide range of applications demonstrated in this study indicates that pooled PE screens can significantly enhance the functional characterization toolbox and increase the ability to elucidate the role of disease-related variants in the human genome.
[0034] Our analysis indicates that lentiviral introduction of PE results in prolonged expression of nCas9, pegRNA, and ngRNA, but can lead to undesirable sequence-specific repression, similar to CRISPRi. This bias must be corrected to create accurate base-pair resolution annotations. When evaluating the functional effects of variants, pegRNA controls should be included when introducing other alleles to the same locus. Our study normalized the sequence-specific repression bias by comparing the different effects of all base-pair substitutions at each locus in MYC enhancers on cell viability and by comparing Alt and Ref alleles for disease variants. Further improvements can be achieved through control of nCas9 expression duration. For example, doxycycline-inducible nCas9 can be selectively expressed when editing is needed and then reversibly turned off. In addition to establishing and optimizing PE screens, we defined sensitive base pairs (SBPs) and core sequences for MYC enhancer function. A functional PWM for this enhancer was generated by utilizing the effect size for all possible substitutions of each base from the PE screen. The functional PWM allowed for accurate prediction of TF binding sites within the enhancer and provided important annotations to accurately explain MYC activation in MCF7 cells.
[0035] Interpreting the effects of hereditary mutations dramatically improves our ability to predict an individual's disease risk. However, without substantial functional annotation, the use of GWAS data for risk prediction remains limited. In this study, 7.9% of the 1304 GWAS breast cancer variants tested and 6.2% of the 2532 VUSs tested were identified as significant hits with functions associated with the MCF7 growth phenotype. Our results demonstrate the feasibility of PE screening for functionally characterizing individual variants. Many applications are possible: for example, PE screening to identify variants associated with different drug therapy responses can help build models that better predict individual-specific benefits and risks from therapeutics. PE screening of variants related to readouts directly associated with physiological functions, such as endolysosomal activity in microglia or synaptic activity in neurons, using iPSC models, reveals functional variants associated with neuropsychiatric disorders. In summary, the present invention provides a functional genomic tool for practical disease prediction, prevention, and treatment needed to realize personalized medicine.
[0036] Data Availability Statement: The next-generation sequencing data reported in this study are available from the NCBI Sequence Read Archive database as accession PRJNA909251. Link to the reviewers created for BioProject PRJNA909251.
[0037] Methods Cell Culture MCF7 cells were cultured in Dulbecco's Modified Eagle Medium (DMEM) (Gibco, 10569010) supplemented with 10% fetal bovine serum (FBS) (HyClone, SH30396.03) and subcultured using trypsin-EDTA (Gibco, 25200072). All cells were cultured at 37°C with 5% CO2 and confirmed to be free of mycoplasma using the MycoAlert Mycoplasma Detection Kit (Lonza, LT07-218). Wild-type MCF7 cells were donated from the laboratory of Howard Y. Chang. The MCF7-nCas9 / RT cell line was generated by lentiviral transduction of cells in a cassette expressing the nickase-Cas9 (nCas9) Moloney mouse leukemia virus reverse transcriptase (M-MLV RT) fusion protein. The infected MCF7 cell pool was treated with puromycin (2.5 μg / ml) for two weeks. Then, single cells were sorted into 96-well plates (one cell per well) by fluorescence-activated cell sorting (FACS) to generate clonal MCF7-nCas9 / RT cell lines. nCas9 / RT expression levels were quantified in each clone via RT-qPCR to identify WTC11 doxycycline-inducible dCas9-KRAB iPSC lines. 43、44 The data was normalized to the dCas9 expression level in the given region.
[0038] Functional characterization of MYC enhancers due to CRISPR deletion Two sgRNAs were designed to knock out the MCF7 enhancer (chr8:128141747~128142627, hg38) (sg1:GAAGTTGTAAGTATAGCGAG, sg2:AGTGCCTGGCACAAGGCAGA). The sgRNAs were synthesized in vitro using the Precision gRNA Synthesis Kit (Invitrogen, A29377) according to the manufacturer's protocol, and their concentrations were quantified using Nanodrop. To deliver the genome editing mechanism, 100 pmol of Cas9-NLS protein (QB3 MacroLab, University of California, Berkeley) and 120 pmol of in vitro synthesized gRNA were electroporated into 250,000 MCF7 cells using the DN-100 Lonza 4D-Nucleofector program along with P3 primary nucleofection solution (Lonza, V4XP-3024). The cells were then plated into 6-well plates, cultured for 2 days, and subsequently plated into 96-well plates to select single clones. Successful knockout clones were identified by genomic PCR using forward primers: CACCAGGACTTGAAGGCAGC and reverse primers: CACTTCCCAACCTCAGTTTCC. MYC expression was quantified using RT-qPCR (MYC forward primer: GTCCTCGGATTCTCTGCTCT, reverse primer ATCTTCTTGTTCCTCCTCAGAGTC) and normalized to GAPDH expression levels (GAPDH forward primer: ATTCCATGGCACCGTCAAGG, reverse primer TTCTCCATGGTGGTGAAGACG).
[0039] Cloning of prime-edited plasmids To construct the lentiV2-EF1a-nCas9 / RT plasmid, the U6-sgRNA cassette was first excised from the lentiCRISPR v2 plasmid (Addgene, 52961) by double digestion with KpnI and EcoRI, followed by blunt-end ligation. Furthermore, the Cas9 cassette was replaced with an nCas9 / M-MLV-RT cassette from the pCMV-PE2 plasmid (Addgene, 132775). The lentiV2-pegRNA plasmid and lentiV2-ngRNA plasmid were constructed by replacing the Cas9 and puromycin sequences in the lentiCRISPR v2 plasmid (Addgene, 52961) with hygromycin B and EGFP sequences, respectively. The RNA motifs and sgRNA scaffolds were further integrated by Gibson assembly.
[0040] Testing Prime editing efficiency To evaluate the prime editing efficiency at the EMX1 and FANCF loci, pegRNA / ngRNA pairs were cloned into separate vectors. In lentivirus co-infection studies, MCF7 cells were first infected with EF1a-nCas9 / RT lentivirus, followed by treatment with puromycin (2.5 μg / ml, Sigma-Aldrich, P8833) for two weeks to remove uninfected cells. Subsequently, EF1a-nCas9 / RT-infected cells were seeded into 24-well plates at a rate of 12,500 cells per well and co-infected with pegRNA and ngRNA. Infected cells were treated with hygromycin B (200 μg / ml, Gibco, 10687010) 48 hours after infection, and collected one week after infection to evaluate editing efficiency. In the MCF7-nCas9 / RT clone lineage study, cells were seeded in 24-well plates at a rate of 12,500 cells per well, followed by lentiviral infection (pegRNA-mCherry and ngRNA-EGFP). Two days after infection, double-positive cells for mCherry and EGFP were isolated by FACS and cultured. Cultured cells were then collected at two and four weeks after infection to evaluate editing efficiency. Subsequently, 620 genomic DNA molecules were extracted from each sample using the Wizard genomic DNA purification kit (Promega, A1120). Target genomic regions were amplified from the purified genomic DNA, and the amplicons were sequenced on the Illumina NovaSeq 6000 platform. In short, sequencing libraries were prepared using DNA primers that amplified the target genomic loci, and the first round of PCR (PCR1) was performed. Next, DNA primers containing index adapters were used in the second round of PCR (PCR2) to add these adapters to the PCR1 amplicons. Finally, dual-index primers were used in the third round of PCR (PCR3) to add Illumina indices to each PCR2 amplicon. Alignment of the amplicons with the reference sequence was performed using CRISPResso2. 45The analysis was performed using [method / tool name]. For the quantification of all prime editing efficiency, the frequency of wild-type and edited amplicons was quantified using a 21 bp window centered on either a 1 bp wild-type sequence or an edited sequence. The remaining amplicons were classified as indels. SNP prioritization: Breast cancer susceptibility genes identified by GWAS. 26 Fourteen MCF7 growth-related genes that overlap with the selected genes were chosen. For each gene, the Breast Cancer Alliance Consortium 25 SNPs were selected using GWAS results. For association with breast cancer within ±500kb of each transcription start site, GWAS P < 1 × 10⁻¹⁰ -5 We identified genome-wide significant SNPs with a minor allele frequency <0.02 and an odds ratio <0.9 or >1.2 (approximately representing the upper and lower quartiles of the odds ratio distribution for SNPs that meet the location, P-value, and MAF threshold). We also examined GWAS results from Latino populations. 46 Using the GWAS P < 1 × 10 at the ESR1 locus -5 SNPs having the following characteristics were also selected separately. The 641 chain disequilibrium (LD) clamps among the selected SNPs were placed in the LD Link R package. 47 Using R 2 The LD threshold was determined to be >0.1. Next, CAVIAR 48Using this method, we prioritized the variants most likely to be causative in each haplotype block based on their posterior probability of being causative (>0.1), the highest posterior probability (≤0.1), or the most extreme odds ratio. CAVIAR was run twice for each locus. The first time, we assumed there was only one causative variant per LD clamp, and then again, we identified more than one causative variant in each LD clamp. Clinical variant prioritization: Clinical variants were retrieved from the ClinVar database (accessed December 25, 2021), and all single nucleotide variants (SNVs) were retained for prime editing screen design (Figure 7c). First, we selected only SNVs that overlapped with breast cancer risk genes and MCF7 growth-related genes. Next, we retained only SNVs in the categories of benign, pathogenic, and significance unknown. Furthermore, for SNVs related to BARD1, BRCA1, BRCA2, RAD51C, RAD51D, and PTEN, only SNVs with more than three registered individuals were retained, as these genes have thousands of identified variants. Ultimately, 5310 SNVs were obtained using our selection criteria, and pegRNA / ngRNA pairs were successfully designed for 3699 of these SNVs.
[0041] Design and construction of a prime editing library To perform nucleotide-resolution analysis of MYC enhancer function, we first target a 716bp enhancer region using a pegRNA / ngRNA pair, and then use PrimeDesign's PooledDesign-Saturation mutagenesis tool. 49The design was performed using pegIT. PegRNA / ngRNA pairs were optimized based on proximity between ngRNA and pegRNA (over 50 bp) and primer binding site (PBS) length (approximately 14 nt), and sequences containing BsmBI cleavage sites (GAGACG, CGTCTC) or TTTTT were redesigned. Next, the specificity and efficiency of the pegRNA and ngRNA spacer sequences were evaluated using GuideScan2. Spacer sequences with low specificity were redesigned to improve specificity. Finally, three different pegRNA / ngRNA pairs were designed so that 93.0% (666 / 716) of substitutions targeted the same base pair. Each repeat in the pegRNA / ngRNA pair shared the same pegRNA and sgRNA spacer sequences, with only the substituted allele having a different pegRNA extension sequence. pegIT was used to design a positive control guide. 50 Using pegIT, we generated pegRNA / ngRNA pairs that introduced a stop codon within the MYC coding region by modifying a single base pair. 50 The best pegRNA / ngRNA pair was selected for each position proposed by the AAVS1 locus, as in previous studies. 51 Based on this, select the negative control region for the targeted pegRNA / ngRNA pair and guide it as shown above in PrimeDesign. 49The design was created using the following method. For non-targeting pegRNA / ngRNA pairs, the pegRNA spacer sequence, ngRNA spacer sequence, and pegRNA elongation sequence were selected from the ENCODE non-targeting sgRNA reference data set (https: / / www.encodeproject.org / files / ENCFF058BPG / ). Guanine nucleotides were added to the 5' end of all pegRNA / ngRNAs with a first nucleotide other than G to improve transcription efficiency from the U6 promoter. These component sequences were concatenated using the following template: 5'-CTTGGAGAAAAGCCTTGTTT[ngRNA-spacer]GTTTAGAGACG[5nt-random sequence]CGTCTCACACC[pegRNA-spacer]GTTTTAGAGCTAGAAATAGCAAGTTAAAATAAGGCTAGTCCGTTATCAACTTGAAAA AGTGGCACCGAGTCGGTGC[pegRNA elongation region]CCTAACACCGCGGTTC-3'.
[0042] Library oligos for the MYC enhancer screen were synthesized by Twist Bioscience and amplified using NEBNext High-Fidelity 2× PCR Master Mix (NEB, M0541L), forward primer: GTGTTTTGAGACTATAAATATCCCTTGGAGAAAAGCCTTGTTT, and reverse primer: CTAGTTGGTTTAACGCGTAACTAGATAGAACCGCGGTGTTAGG. Emulsion PCR (ePCR) was used to amplify paired PegRNA / ngRNA library oligos for enhancer-saturated mutagenesis, reducing recombination of similar amplicons during PCR. Briefly, 96 ePCR reactions in 20 μl were performed using 0.01 fmol of pooled oligos with NEBNext High-Fidelity 2× PCR Master Mix (NEB, M0541S). Each 20 μl PCR mix is mixed with 40 μl of oil-surfactant mixture (mineral oil containing 4.5% Span 80 (vol / vol), 0.4% Tween 80 (vol / vol), and 0.05% Triton X-100 (vol / vol)). 52 The mixture was then combined. This mixture was vortexed at maximum speed for 5 minutes, lightly centrifuged, and amplified in a PCR instrument. The thermocycler settings were 98°C for 30 seconds, followed by 26 cycles (98°C for 10 seconds, 60°C for 20 seconds, 72°C for 30 seconds), followed by 72°C for 5 minutes, and finally held at 4°C. The ramp rate for each step was 2°C / second. After PCR, the individual reactants were combined according to previously established guidelines. 53The PCR products were purified using the QIAQuick PCR Purification Kit (Qiagen, 28104) according to the instructions. The purified PCR products were then treated with exonuclease I (NEB, M0568L) and purified using 1×AMPure XP beads (Beckman Coulter, A63881). The isolated ePCR products were then inserted into lentiV2-mU6-evopreQ1 vectors digested with BsmBI via a Gibson assembly (NEB, E2621L). The assembled products were electroporated into Endura electrocompetent Escherichia coli cells (Biosearch Technologies, 60242) to culture approximately 4000 independent bacterial colonies per library. The obtained plasmid DNA was linearized by BsmbI digestion, gel-purified, and ligated into DNA fragments containing the sgRNA scaffold and human U6 promoter using T4 ligase (NEB, M0202M). The resulting library was electroporated into Endura electrocompetent E. coli cells (Biosearch Technologies, 60242) and cultured as described above. The final plasmid library was extracted using the Qiagen EndoFree Plasmid Mega Kit (Qiagen, 12381).
[0043] Regarding the Alt library for SNPs and clinical variant screens, PrimeDesign 49PegRNA / ngRNA pairs were designed using [a specific method / tool]. 200 bp upstream and downstream sequences of each variant or iSTOP were used as input for PrimeDesign. Initial pegRNA / ngRNA pairs were generated using the following parameters: number of pegRNAs per edit: 10, downstream homology length: 10 nt, PBS length: 13 nt, maximum reverse transcription template (RTT) length: 50 nt, number of ngRNAs per pegRNA: 10, nicking distance from ngRNA to pegRNA: 50 bp and 75 bp. Next, guanine nucleotides were added to the 5' end of all pegRNA / ngRNAs with a non-G leading nucleotide to increase transcription efficiency from the U6 promoter. PegRNA / ngRNA pairs containing pegRNA spacers, ngRNA spacers, or BsmBI sites (GAGACG, CGTCTC) or TTTTT sequences in the pegRNA elongation region were removed. When multiple pairs are available at the same locus, pegRNA / ngRNA pairs were further selected to maximize specificity, efficiency, and the distance from ngRNA to pegRNA while simultaneously minimizing the distance from pegRNA to the edit. For untargeted pegRNA / ngRNA pairs, the pegRNA spacer, ngRNA spacer, and pegRNA extension sequence were selected from the ENCODE untargeted sgRNA reference dataset (https: / / www.encodeproject.org / files / ENCFF058BPG / ). The same pegRNA / ngRNA pairs as the Alt library were used to design the Ref library, but the substituted alleles in the pegRNA extension sequence were replaced with the reference allele sequence. The final oligo followed the following template structure: 5'-CTTGTGGAAAGGACGAAACACC[ngRNA-spacer]GTTTCGAGACG[6nt-random sequence]CGTCTCTTGTTT[pegRNA-spacer]gttttagagctagaaatagcaagttaaaataaggctagtccgttatcaacttgaaaaagtggcaccgagtcggtgc[pegRNA extension]TTGACGCGGTTCTATCTAGTTAC-3'.
[0044] Alt and Ref library oligos were synthesized by Twist Bioscience. The Alt and Ref plasmid libraries were cloned separately using two-step cloning. First, the oligo pool for each library was amplified using NEBNext High-Fidelity 2×PCR Master Mix (NEB, M0541L) and the following primers: Forward primer: TCGATTTCTTGGCTTTATATATCTTGTGGAAAGGACGAAACAC, Reverse primer: ATTTCTAGTTGGTTTAACGCGTAACTAGATAGAACCGCGTCAA. The PCR products were purified via gel excision and column purification (Promega, A9282) and subsequently inserted into a BsmBI-digested lentiV2-hU6-evopreQ1 vector by Gibson assembly (NEB, E2621L). The assembled products were electroporated into Endura electrocompetent E. coli cells (Biosearch Technologies, 60242). Approximately 25 million bacterial colonies were cultured from each library and subsequently purified using the QIAGEN Plasmid Maxi Kit (QIAGEN, 12163). In the second step, the plasmid libraries obtained from the first cloning step were linearized by BsmbI digestion, gel-purified, and ligated into DNA fragments containing sgRNA scaffolds and mouse U6 promoters using T4 ligase (NEB, M0202M). The ligated products were electroporated into Endura electrocompetent E. coli cells (Biosearch Technologies, 60242), and approximately 40 million bacterial colonies were cultured from each library. The final plasmid libraries were extracted using the Qiagen EndoFree Plasmid Mega Kit (Qiagen, 12381).
[0045] Lentivirus generation and titer measurement To generate a lentivirus library, use the method previously described.44 In short, a 5 μg plasmid library was co-transfected into 8 million HEK293T cells in a 10 cm dish supplemented with 36 μl of PolyJet (SignaGen Laboratories, SL100688) along with 3 μg of psPAX (Addgene, 12260) and 1 μg of pMD2.G (Addgene, 12259) packaging plasmid. The medium was changed 12 hours after transfection, and then sampled every 24 hours for a total of three samples. The sampled viral medium was filtered through a Millex-HV (0.45 μm) polyvinylidene fluoride filter (Millipore, SLHV033RS) and then concentrated by centrifugation using a 100,000 NMWL (nominal molecular weight limit) Ultra-15 centrifugation filter unit (Amicon, UFC910008).
[0046] Lentiviral titers were determined by transduction into 400,000 cells using concentrated virus and polyblen (6 μg / ml, Millipore, TR-1003-G) in increasing volumes (0 μl, 1 μl, 2 μl, 5 μl, 10 μl, 20 μl, and 40 μl). Forty-eight hours after transduction, the cells were dissociated using trypsin-EDTA (0.25%, Gibco, 25200056) and seeded as two separate repeats. One repeat was treated with hygromycin B (200 μg / ml, Gibco, 10687010) for four days, while the other was left untreated. Finally, hygromycin-resistant cells and control cells were counted, and the infection rate and viral titer were calculated.
[0047] Prime editing screen The MYC enhancer PE screen was performed in three consecutive sequences. MCF7-dCas9 / RT cells were transfected with a lentiviral library at an infection multiplicity (MOI) of 0.3, with coverage of 1000 transduced cells per pair of pegRNA / ngRNA. After 48 hours, approximately 10 million cells were collected as controls, and the remaining cells were treated with hygromycin B (200 μg / ml, Gibco, 10687010) for 7 days. After antibiotic selection, the cells were maintained in DMEM supplemented with 10% FBS for 30 days post-infection, and 10 million cells were collected from the final cell population. The Alt library screen and Ref library screen were performed in four consecutive sequences. For each repeat of the Alt and Ref screens, approximately 24 million MCF7-nCas9 / RT cells were individually infected with the lentiviral library at an MOI of 0.5, with cell coverage of 2000 infected cells per pair of pegRNA / ngRNA. 48 hours after infection, one-third of the infected cells were collected from each cell pool as control samples (day 2). The remaining cells were treated with hygromycin B (200 μg / ml, Gibco, 10687010) for 7 days and cultured until 32 days after infection (day 32).
[0048] Generation of Illumina sequencing libraries Genomic DNA was extracted from each sample via cell lysis and digestion (100 mM Tris-HCl (pH 8.5), 5 mM EDTA, 200 mM NaCl, 0.2% SDS, and proteinase K (100 μg / ml)), phenol:chloroform (Thermo Fisher Scientific, 17908) extraction, and isopropanol (Thermo Fisher Scientific, BP2618500) precipitation. In the MYC enhancer screen, ePCR was applied during library preparation to amplify paired pegRNA / ngRNA sequences from each sample and reduce recombination between similar sequences. In short, 30 20 μl ePCRs were performed using 400 ng of DNA and NEBNext High-Fidelity 2×PCR Master Mix (NEB, M0541S) with the following primers: Enh-lib-forward: TCCCTACACGACGCTCTTCCGATCTNNNNNCCTTGGAGAAAAGCCTTGTTT, Enh-lib-reverse: GGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNGAACCGCGGTGTTAGG. ePCR was performed as previously described to amplify pegRNA / ngRNA pairs from genomic DNA. Thermocycler settings were 98°C for 30 seconds, followed by 25 cycles (98°C for 10 seconds, 60°C for 20 seconds, 72°C for 1 minute), followed by 72°C for 5 minutes, and finally held at 4°C. The ramp rate for each step was 2°C / second. After PCR, the individual reactants were combined according to previously established guidelines. 53The PCR product was purified using the QIAQuick PCR Purification Kit (Qiagen, 28104) according to the instructions. The purified PCR product was then treated with exonuclease I (NEB, M0568L) and purified using 1×AMPure XP beads (Beckman Coulter, A63881). The PCR amplicon from the first round was used in the second round of PCR to add the Illumina adapter and index sequence. In the second round of PCR, six ePCR reactions were performed using NEBNext High-Fidelity 2×PCR Master Mix (NEB, M0541S), each containing 0.023 ng of purified DNA. The PCR mixture from the second round was prepared and purified in the same manner as in the first round. The thermocycler settings were 98°C for 30 seconds, followed by 12 cycles (98°C for 10 seconds, 60°C for 20 seconds, 72°C for 1 minute), followed by 72°C for 5 minutes, and finally held at 4°C. The ramp rate for each step was 2°C / second. For the Alt and Ref screens, pegRNA / ngRNA pair sequences were amplified from each sample using NEBNext High-Fidelity 2×PCR Master Mix (NEB, M0541L) and the following primers: Alt-Ref-lib-forward: TCCCTACACGACGCTCTTCCGATCTNNNNNCTTGTGGAAAGGACGAAACACC, Alt-Ref-lib-reverse: GGAGTTCAGACGTGTGCTCTTCCGATCTNNNNNCGTAACTAGATAGAACCGCGTCAA. For each sample, 24 50 μl PCR reactions containing 600 ng of genomic DNA were performed. The individual reaction products from each sample were combined and purified by column (Promega, A9282).Next, the purified product was amplified by index PCR, and the Illumina TruSeq adapter and sample index sequence were added using the following primers: index-forward: aatgatacggcgaccaccgagatctacac[8bp index]acactctttccctacacgacgctcttccgatct, index-reverse: caagcagaagacggcatacgagat[8bp index]gtgactggagttcagacgtgctcttccgatct. The final library was gel-purified and sequenced with 150bp paired ends on the Illumina NovaSeq 6000 platform.
[0049] Data processing and analysis of prime editing data The sequencing library was initially trimmed with 5bp random sequences from read 1 and read 2, and low-quality reads were filtered out using the fastp tool before formal mapping. Each pegRNA / ngRNA pair was included in the read count if it met the following criteria: (1) read 1 exactly matched a sequence containing a 20nt–21nt ngRNA spacer and a 5bp flanking sequence, and (2) read 2 exactly matched a reverse complementary sequence containing a complete pegRNA extension and a 5bp flanking sequence.
[0050] The MYC Enhancer PE screen uses the MAGeCK (0.5.9) pipeline. 13Using this method, we estimated the statistical significance and multiplier of change for each pegRNA / ngRNA pair at the sgRNA level, as well as the statistical significance and multiplier of change for each substitution at the gene level, in a cell population relative to the control. Untargeted pegRNA and AAVS1-targeted pegRNA were used as negative controls for normalization. To identify core enhancer regions for MYC enhancers based on the screening results, we first identified base pairs with three significant substitutions (FDR < 0.05) and calculated the slope for each contiguous bin (move step = 1 bp, bin size = 30 bp, x-axis: position of each base pair, y-axis: cumulative number of SBPs with three significant substitutions) (Figure 6e). The slope was then converted to a p-value corresponding to the Z-score. Core enhancer regions were identified by merging overlapping significant bins (p-value < 0.05).
[0051] In the Alt and Ref library screens, oligonucleotides with zero reads were removed from all samples, and then the following analysis was performed. The oligonucleotide count from all samples was calculated using DESeq2(1, 38, 0). 31The samples were passed to DESeq2 and normalized for various sequence depths using the median ratio method. The normalized read counts for each oligo were then modeled as a negative binomial distribution using DESeq2. Next, the change multiplier for each oligo in the Alt and Ref libraries was confirmed by comparing data from day 32 and day 2 using DESeq2 (design=~Replicate+Condition). Furthermore, the relative effect between the reference and surrogate alleles was estimated by adding an interaction term (design=~Replicate+Condition+Allele+Condition:Allele). Condition refers to the collection time (i.e., day 32 or day 2), and Allele refers to the allele category (i.e., Alt or Ref). Finally, the P-value was calculated by performing a Wald test via DESeq2. Then, a P-value cutoff corresponding to the 5th percentile of the P-value from the untargeted control oligo was selected to minimize false positive hits and achieve an empirical FDR of less than 5%.
[0052] Motif matrix comparative analysis Motif comparison is a method used to identify potential transcription factor (TF) binding sites within target MYC enhancers by directly comparing known TF motifs with the inventors' base-pair resolution functional data. 54 A new method based on this was established. First, MAGeCK(0.5.9) 13Using this method, the log2(change factor) was calculated for each substitution at each base pair. The log2(change factor) of the wild-type allele was set to 0. Next, the log2(change factor) of each substitution was converted to the corresponding change factor value. Furthermore, a position weight matrix was constructed by normalizing the change factor per base pair for each allele against the sum of the change factors per base pair for all unique alleles. In addition, the enhancer sequences were divided into numerous bins of 5 base pairs and 10 base pairs in length. Only bins with an information content (IC) greater than 3 and an "N" content of less than 10% were retained. Next, all TF motifs that were highly expressed in MCF7 cells (TPM>10, GSE175204) were collected from the JASPAR, HOCOMOCO, and SwissRegulon databases. Then, using Tomtom (P-value < 0.05), the filtered TF motif matrix and the enhancer bin matrix were compared to identify potential TF binding sites in the enhancers. Ultimately, only positive TF motif hits (positions with the highest probability greater than 0.5) that overlapped with at least 95% of the essential base pairs of the input sequence were retained.
[0053] Prediction of the contribution of base pairs to enhancer activity using BPNet Publicly released approach 24A convolutional neural network was trained using a BPNet compliant with the ENCODE project to describe ChIP-seq data for GATA3, ELF1, FOXM1, MTA3, and RCOR1. Briefly, the model input was a 1kb sequence across each ChIP-seq peak locus, with the corresponding ChIP-seq control peak used as a bias track for training. Regions from chromosome 2 were used as the tuning set, and chromosomes 5, 6, 7, 10, and 14 were used as the test set. X and Y chromosomes were excluded. The remaining regions from the other chromosomes were used to train the model with default parameters. Once the model was acquired for the ChIP-seq data of each TF, DeepLIFT was used to calculate the contribution of each input sequence base pair to enhancer activity. Finally, the TF-MoDISco contribution scores were used to cluster and determine the integrated TF motifs and map them to the input peak regions.
[0054] MCF7 genotyping analysis Sequence read archive (SRA) files for SRR7707725 and SRR7707726 (paired-end, two reads per locus) were obtained from BioProject PRJNA486532. Sequenced reads were individually aligned to the human reference genome hg38 for each run using bwa-mem v.0.7.17. The BAM files were then processed using Picard tools: SortSam, MarkDuplicates, and AddOrReplaceReadGroups. Finally, SNPs and indels were invoked via local haplotype reconstruction (HaplotypeCaller) using GATK v.4.2.5.0, followed by co-genotyping of single-sample GVCFs from HaplotypeCaller (GenotypeGVCFs). Lastly, CalcMatch v.1.1.2 was used to verify genotype consistency between the two runs.
[0055] Motif scanning and TF identification of alleles with functional breast cancer SNPs The 20bp upstream and downstream sequences of each SNP (Alt allele and Ref allele) were used as input sequences for TF motif analysis. FIMO software (version 5.5.0) 55 Using this, we searched for matching motifs centered around SNP regions in the human TF motif database HOCOMOCO (v11 FULL). 19 We identified the target SNP locus. We performed a full FIMO motif scan using default settings. Finally, we selected TFs (FPKM>1) that had binding motifs overlapping with the target SNP locus (FDR<0.05, P-value<0.0001).
[0056] Protein structure prediction using AlphaFold To investigate the effect of the BARD1 His36Pro mutation on the BARD1 / BRCA1 complex structure, we predicted the structures of wild-type BRAD1 / BRCA1 and the BARD1(His36Pro) / BRCA1 complex using AlphaFold. NMR spectroscopy was used as input for complex structure prediction. 40 The same amino acid chains used in the BARD1 / BRCA1 complex structure determined by (BARD1, residues 26-122; BRCA1, residues 1-103) were used. The amino acid chains of BARD1 and BRCA1 were generated using AlphaFold V2.2.4, which runs on Google Compute Engine with Python 3. 56、57 The data was imported into the Google Colab version. AlphaFold applied a multimer model based on the duo-sequence imputation, then searched the gene database to determine the optimal multiple sequence alignment (MSA) for the imported sequence, and began structure prediction. To avoid stereochemical violations, all structures were relaxed using the AMBER model (Assisted Model Building with Energy Refinement) with GPU acceleration. The resulting PDB file was then processed on UCSF Chimera X. 58、59The structure was imported and visualized. Different colors were assigned to the protein chains to distinguish individual chains, and the atomic structures and hydrogen bonds of selected amino acids were illustrated for interaction analysis. Finally, the snapshot function in Chimera X was used to export the real-time rendered complex structure at the optimal visualization angle.
[0057] References 1. Taliun, D. et al. Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program. Nature 590, 290-299 (2021). 2. Shalem, O., Sanjana, NE & Zhang, F. High-throughput functional genomics using CRISPR-Cas9. Nat Rev Genet 16, 299-311 (2015). 3. Anzalone, AV, Koblan, LW & Liu, DR Genome editing with CRISPR-Cas nucleases, base editors, transposases and prime editors. Nat Biotechnol 38, 824-844 (2020). 4. Chen, PJ & Liu, DR Prime editing for precise and highly versatile genome manipulation. Nat Rev Genet (2022). 5. Anzalone, AV et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149-157 (2019). 6. Erwood, S. et al. Saturation variant interpretation using CRISPR prime editing. Nat Biotechnol 40, 885-895 (2022). 7. Anzalone, A.V., Lin, A.J., Zairis, S., Rabadan, R. & Cornish, V.W. Reprogramming eukaryotic translation with ligand-responsive synthetic RNA switches. Nat Methods 13, 453-458 (2016). 8. Houck-Loomis, B. et al. An equilibrium-dependent retroviral mRNA switch regulates translational recoding. Nature 480, 561-564 (2011). 9. Nelson, J.W. et al. Engineered pegRNAs improve prime editing efficiency. Nat Biotechnol 40, 402-410 (2022). 10. Dang, Y. et al. Optimizing sgRNA structure to improve CRISPR-Cas9 knockout efficiency. Genome Biol 16, 280 (2015). 11. Chen, P.B. et al. Systematic discovery and functional dissection of enhancers needed for cancer cell fitness and proliferation. Cell Rep 41, 111630 (2022). 12. Cho, S.W. et al. Promoter of lncRNA Gene PVT1 Is a Tumor-Suppressor DNA Boundary Element. Cell 173, 1398-1412 e1322 (2018). 13. Li, W. et al. MAGeCK enables robust identification of essential genes from genome-scale CRISPR / Cas9 knockout screens. Genome Biol 15, 554 (2014). 14. Shalem, O. et al. Genome-scale CRISPR-Cas9 knockout screening in human cells. Science 343, 84-87 (2014). 15. Baluapuri, A., Wolf, E. & Eilers, M. Target gene-independent functions of MYC oncoproteins. Nat Rev Mol Cell Biol 21, 255-267 (2020). 16. Vitsios, D., Dhindsa, R.S., Middleton, L., Gussow, A.B. & Petrovski, S. Prioritizing non-coding regions based on human genomic constraint and sequence context with deep learning. Nat Commun 12, 1504 (2021). 17. Villar, D. et al. Enhancer evolution across 20 mammalian species. Cell 160, 554-566 (2015). 18. Fornes, O. et al. JASPAR 2020: update of the open-access database of transcription factor binding profiles. Nucleic Acids Res 48, D87-D92 (2020). 19. Kulakovskiy, I.V. et al. HOCOMOCO: towards a complete collection of transcription factor binding models for human and mouse via large-scale ChIP-Seq analysis. Nucleic Acids Res 46, D252-D259 (2018). 20. Pachkov, M., Balwierz, P.J., Arnold, P., Ozonov, E. & van Nimwegen, E. SwissRegulon, a database of genome-wide annotations of regulatory sites: recent updates. Nucleic Acids Res 41, D214-220 (2013). 21. Consortium, E.P. An integrated encyclopedia of DNA elements in the human genome. Nature 489, 57-74 (2012). 22. Schreiber, J., Durham, T., Bilmes, J. & Noble, W.S. Avocado: a multi-scale deep tensor factorization method learns a latent representation of the human epigenome. Genome Biol 21, 81 (2020). 23. Behan, F.M. et al. Prioritization of cancer therapeutic targets using CRISPR-Cas9 screens. Nature 568, 511-516 (2019). 24. Avsec, Z. et al. Base-resolution models of transcription-factor binding reveal soft motif syntax. Nat Genet 53, 354-366 (2021). 25. Michailidou, K. et al. Association analysis identifies 65 new breast cancer risk loci. Nature 51, 92-94 (2017). 26. Fachal, L. et al. Fine-mapping of 150 breast cancer risk regions identifies 191 likely target genes. Nat Genet 52, 56-73 (2020). 27. Hanna, R.E. et al. Massively parallel assessment of human variants with base editor screens. Cell 184, 1064-1080 e1020 (2021). 28. Landrum, M.J. et al. ClinVar: improvements to accessing data. Nucleic Acids Res 48, D835-D844 (2020). 29. Cuella-Martin, R. et al. Functional interrogation of DNA damage response variants with base editing screens. Cell 184, 1081-1097 e1019 (2021). 30. Qi, L.S. et al. Repurposing CRISPR as an RNA-guided platform for sequence-specific control of gene expression. Cell 152, 1173-1183 (2013). 31. Love, M.I., Huber, W. & Anders, S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol 15, 550 (2014). 32. Bruna, A. et al. TGFbeta induces the formation of tumour-initiating cells in claudinlow breast cancer. Nat Commun 3, 1055 (2012). 33. Bossone, S.A., Asselin, C., Patel, A.J. & Marcu, K.B. MAZ, a zinc finger protein, binds to c-MYC and C2 gene sequences regulating transcriptional initiation and termination. Proc Natl Acad Sci U S A 89, 7452-7456 (1992). 34. Wang, X. et al. MAZ drives tumor-specific expression of PPAR gamma 1 in breast cancer cells. Breast Cancer Res Treat 111, 103-111 (2008). 35. Kircher, M. et al. A general framework for estimating the relative pathogenicity of human genetic variants. Nat Genet 46, 310-315 (2014). 36. Pollard, K.S., Hubisz, M.J., Rosenbloom, K.R. & Siepel, A. Detection of nonneutral substitution rates on mammalian phylogenies. Genome Res 20, 110-121 (2010). 37. Li, W. et al. A synergetic effect of BARD1 mutations on tumorigenesis. Nat Commun 12, 1243 (2021). 38. UniProt, C. UniProt: the universal protein knowledgebase in 2021. Nucleic Acids Res 49, D480-D489 (2021). 39. Prakash, R. et al. Homologous recombination-deficient mutation cluster in tumor suppressor RAD51C identified by comprehensive analysis of cancer variants. Proc Natl Acad Sci U S A 119, e2202727119 (2022). 40. Brzovic, P.S., Rajagopal, P., Hoyt, D.W., King, M.C. & Klevit, R.E. Structure of a BRCA1-BARD1 heterodimeric RING-RING complex. Nat Struct Biol 8, 833-837 (2001). 41. Densham, R.M. et al. Human BRCA1-BARD1 ubiquitin ligase activity counteracts chromatin barriers to DNA resection. Nat Struct Mol Biol 23, 647-655 (2016). 42. Spain, B.H., Larson, C.J., Shihabuddin, L.S., Gage, F.H. & Verma, I.M. Truncated BRCA2 is cytoplasmic: implications for cancer-linked mutations. Proc Natl Acad Sci U S A 96, 13920-13925 (1999). 43. Mandegar, M.A. et al. CRISPR Interference Efficiently Induces Specific and Reversible Gene Silencing in Human iPSCs. Cell Stem Cell 18, 541-553 (2016). 44. Ren, X. et al. Parallel characterization of cis-regulatory elements for multiple genes using CRISPRpath. Sci Adv 7, eabi4360 (2021). 45. Clement, K. et al. CRISPResso2 provides accurate and rapid genome editing sequence analysis. Nat Biotechnol 37, 224-226 (2019). 46. Fejerman, L. et al. Genome-wide association study of breast cancer in Latinas identifies novel protective variants on 6q25. Nat Commun 5, 5260 (2014). 47. Machiela, M.J. & Chanock, S.J. LDlink: a web-based application for exploring population-specific haplotype structure and linking correlated alleles of possible functional variants. Bioinformatics 31, 3555-3557 (2015). 48. Hormozdiari, F., Kostem, E., Kang, E.Y., Pasaniuc, B. & Eskin, E. Identifying causal variants at loci with multiple signals of association. Genetics 198, 497-508 (2014). 49. Hsu, J.Y. et al. PrimeDesign software for rapid and simplified design of prime editing guide RNAs. Nat Commun 12, 1034 (2021). 50. Anderson, M.V., Haldrup, J., Thomsen, E.A., Wolff, J.H. & Mikkelsen, J.G. pegIT - a web-based design tool for prime editing. Nucleic Acids Res 49, W505-W509 (2021). 51. Chen, C.H. et al. Improved design and analysis of CRISPR knockout screens.Bioinformatics 34, 4095-4101 (2018). 52. Williams, R. et al. Amplification of complex gene libraries by emulsion PCR. Nat Methods 3, 545-550 (2006). 53. Verma, V., Gupta, A. & Chaudhary, V.K. Emulsion PCR made easy. Biotechniques 69, 421-426 (2020). 54. Gupta, S., Stamatoyannopoulos, J.A., Bailey, T.L. & Noble, W.S. Quantifying similarity between motifs. Genome Biol 8, R24 (2007). 55. Grant, C.E., Bailey, T.L. & Noble, W.S. FIMO: scanning for occurrences of a given motif. Bioinformatics 27, 1017-1018 (2011). 56. Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583-589 (2021). 57. Mirdita, M. et al. ColabFold: making protein folding accessible to all. Nat Methods 19, 679-682 (2022). 58. Goddard, T.D. et al. UCSF ChimeraX: Meeting modern challenges in visualization and analysis. Protein Sci 27, 14-25 (2018). 59. Pettersen, E.F. et al. UCSF ChimeraX: Structure visualization for researchers, educators, and developers. Protein Sci 30, 70-82 (2021).
Claims
1. A high-throughput screening method that includes identifying functional DNA variants in the human genome using a pooled prime editing screen.
2. The high-throughput screening method according to claim 1, wherein the variant is related to human health and disease, and the high-throughput screening method further comprises annotating a genome with nucleotide resolution for disease prediction or treatment practical for personalized medicine.
3. The high-throughput screening method according to claim 1, further comprising characterizing gene variants at base pair resolution and on a base pair scale, and advancing accurate genomic annotation for disease risk prediction, diagnosis, or identification of therapeutic targets.
4. The high-throughput screening method according to claim 1, wherein the screen comprises infection with a dual pegRNA / sgRNA virus of the clone MCF7 strain that stably expresses nickase Cas9 (nCas9) and Moloney mouse leukemia virus reverse transcriptase (M-MLV RT).
5. The high-throughput screening method according to claim 1, wherein the screen comprises the delivery of MCF7-nCas9 / RT cells and lentiviral delivery of both pegRNA containing scaffold 1 and ngRNA containing scaffold 2 in the same construct as shown in Figure 1e.
6. The high-throughput screening method according to claim 1, comprising transfecting host cells with a lentivirus containing a stable expression cassette of nCas9 and M-MLV reverse transcriptase (M-MLV RT) shown in Figure 1b to obtain transformed cells with stable expression of nCas9 / M-MLV RT, enabling more efficient pegRNA / ngRNA packaging and lentiviral delivery with higher editing efficiency than the co-infection method, thereby increasing PE efficiency and facilitating a pooled screening approach using a lentiviral library.
7. The high-throughput screening method according to claim 1, comprising transfecting host cells with pegRNA containing a scaffold structure RNA motif at the 3' end of the pegRNA, and exhibiting higher editing efficiency at both the EMX1 locus and the FANCF locus compared to using PE without a structure RNA motif.
8. The high-throughput screening method according to claim 1, comprising transfecting host cells with pegRNA containing a scaffold structure RNA motif at the 3' end of the pegRNA, and exhibiting higher editing efficiency at both the EMX1 locus and the FANCF locus compared to using PE without a structure RNA motif, wherein the RNA motif is selected from EvopreQ1, MLV-PK1, and MLV-PK2.
9. The high-throughput screening method according to claim 1, comprising (a) transfecting host cells with a lentivirus containing a stable expression cassette of nCas9 and M-MLV reverse transcriptase (M-MLV RT) as shown in Figure 1b to obtain transformed cells with stable expression of nCas9 / M-MLV RT, enabling more efficient pegRNA / ngRNA packaging and lentiviral delivery with higher editing efficiency than the co-infection method, thereby increasing PE efficiency and enabling a pooled screening approach using a lentiviral library, and (b) transfecting host cells with pegRNA containing a scaffold structure RNA motif at the 3' end of pegRNA, exhibiting higher editing efficiency at both the EMX1 locus and the FANCF locus compared to using PE without a structure RNA motif.
10. The high-throughput screening method according to claim 1, comprising (a) transfecting host cells with a lentivirus containing a stable expression cassette of nCas9 and M-MLV reverse transcriptase (M-MLV RT) as shown in Figure 1b to obtain transformed cells with stable expression of nCas9 / M-MLV RT, enabling more efficient pegRNA / ngRNA packaging and lentiviral delivery with higher editing efficiency than the co-infection method, thereby increasing PE efficiency and enabling a pooled screening approach using a lentiviral library, and (b) transfecting host cells with pegRNA containing a scaffold structure RNA motif at the 3' end of pegRNA, showing higher editing efficiency at both the EMX1 locus and the FANCF locus compared to using PE without a structure RNA motif, wherein the RNA motif is selected from EvopreQ1, MLV-PK1, and MLV-PK2.
11. The high-throughput screening method according to claim 1, wherein the screen comprises the delivery of MCF7-nCas9 / RT cells and lentiviral delivery of both pegRNA containing scaffold 1 and ngRNA containing scaffold 2 in the same construct as shown in Figure 1e.